{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,5]],"date-time":"2025-10-05T19:45:47Z","timestamp":1759693547355,"version":"3.41.0"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2020,1,15]],"date-time":"2020-01-15T00:00:00Z","timestamp":1579046400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61572434 and 91630206"],"award-info":[{"award-number":["61572434 and 91630206"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003399","name":"Shanghai Science and Technology Committee","doi-asserted-by":"crossref","award":["16DZ2293600"],"award-info":[{"award-number":["16DZ2293600"]}],"id":[{"id":"10.13039\/501100003399","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2020,3,31]]},"abstract":"<jats:p>Word embeddings, which map words into a unified vector space, capture rich semantic information. From a linguistic point of view, words have two carriers, speech and writing. Yet the most recent word embedding models focus on only the writing carrier and ignore the role of the speech carrier in semantic expressions. However, in the development of language, speech appears before writing and plays an important role in the development of writing. For phonetic language systems, the written forms are secondary symbols of spoken ones. Based on this idea, we carried out our work and proposed double-carrier word embedding (DCWE). We used DCWE to conduct a simulation of the generation order of speech and writing. We trained written embedding based on phonetic embedding. The final word embedding fuses writing and phonetic embedding. To illustrate that our model can be applied to most languages, we selected Chinese, English, and Spanish as examples and evaluated these models through word similarity and text classification experiments.<\/jats:p>","DOI":"10.1145\/3344920","type":"journal-article","created":{"date-parts":[[2020,2,24]],"date-time":"2020-02-24T18:12:51Z","timestamp":1582567971000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Enhanced Double-Carrier Word Embedding via Phonetics and Writing"],"prefix":"10.1145","volume":"19","author":[{"given":"Wenhao","family":"Zhu","sequence":"first","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7888-6954","authenticated-orcid":false,"given":"Xin","family":"Jin","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuang","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhiguo","family":"Lu","sequence":"additional","affiliation":[{"name":"Library of Shanghai University, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wu","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ke","family":"Yan","sequence":"additional","affiliation":[{"name":"School of Communication &amp; Information Engineering, Shanghai University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Baogang","family":"Wei","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University, Zhejiang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,1,15]]},"reference":[{"volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","year":"2014","author":"Bengio S.","key":"e_1_2_1_1_1"},{"key":"e_1_2_1_2_1","first-page":"6","article-title":"2003. Neural probabilistic language models","volume":"3","author":"Bengio Y.","year":"2003","journal-title":"Journal of Machine Learning Research"},{"volume-title":"Holt","year":"1933","author":"Bloomfield L.","key":"e_1_2_1_3_1"},{"key":"e_1_2_1_4_1","unstructured":"P. Bojanowski E. Grave A. Joulin etal 2016. Enriching word vectors with subword information. ArXiv Preprint Arxiv:1607.04606 2016.  P. Bojanowski E. Grave A. Joulin et al. 2016. Enriching word vectors with subword information. ArXiv Preprint Arxiv:1607.04606 2016."},{"key":"e_1_2_1_5_1","first-page":"1899","article-title":"2014. Compositional morphology for word representations and language modelling","volume":"2014","author":"Botha J. A.","journal-title":"Computer Science"},{"key":"e_1_2_1_6_1","first-page":"3151","article-title":"2017. Improving word embeddings with convolutional feature learning and subword information","volume":"3144","author":"Cao S.","year":"2017","journal-title":"AAAI"},{"volume-title":"Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence.","year":"2018","author":"Cao S.","key":"e_1_2_1_7_1"},{"volume-title":"Proceedings of the International Conference on Artificial Intelligence. AAAI Press","year":"2015","author":"Chen X.","key":"e_1_2_1_8_1"},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Y. C. Chen S. F. Huang C. H. Shen etal 2018. Phonetic-and-semantic embedding of spoken words with applications in spoken content retrieval. Arxiv Preprint Arxiv:1807.08089 2018.  Y. C. Chen S. F. Huang C. H. Shen et al. 2018. Phonetic-and-semantic embedding of spoken words with applications in spoken content retrieval. Arxiv Preprint Arxiv:1807.08089 2018.","DOI":"10.1109\/SLT.2018.8639553"},{"volume-title":"Proceedings of the 25th International Conference on Machine Learning. ACM, 160--167","author":"Collobert R.","key":"e_1_2_1_10_1"},{"key":"e_1_2_1_11_1","unstructured":"M. Etcheverry and D. Wonsever. 2016. Spanish word vectors from Wikipedia. LREC. 2016.  M. Etcheverry and D. Wonsever. 2016. Spanish word vectors from Wikipedia. LREC. 2016."},{"volume-title":"Liblinear: A library for large linear classification. The Journal of Machine Learning Research 9","year":"2008","author":"Fan Rong-En","key":"e_1_2_1_12_1"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/371920.372094"},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Guy Halawi Gideon Dror Evgeniy Gabrilovich and Yehuda Koren. 2012. Large-scale learning of word relatedness with constraints. In KDD.  Guy Halawi Gideon Dror Evgeniy Gabrilovich and Yehuda Koren. 2012. Large-scale learning of word relatedness with constraints. In KDD.","DOI":"10.1145\/2339530.2339751"},{"volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP).","author":"Hassan S.","key":"e_1_2_1_15_1"},{"key":"e_1_2_1_16_1","unstructured":"W. He W. Wang and K. Livescu. 2016. Multi-view recurrent neural acoustic word embeddings. Arxiv Preprint Arxiv:1611.04496 2016.  W. He W. Wang and K. Livescu. 2016. Multi-view recurrent neural acoustic word embeddings. Arxiv Preprint Arxiv:1611.04496 2016."},{"volume-title":"Proceedings of the Joint Conference on Lexical and Computational Semantics. Association for Computational Linguistics","year":"2013","author":"Jin P.","key":"e_1_2_1_17_1"},{"key":"e_1_2_1_18_1","unstructured":"A. Joulin E. Grave P. Bojanowski etal 2016. Bag of tricks for efficient text classification. Arxiv Preprint Arxiv:1607.01759 2016.  A. Joulin E. Grave P. Bojanowski et al. 2016. Bag of tricks for efficient text classification. Arxiv Preprint Arxiv:1607.01759 2016."},{"key":"e_1_2_1_19_1","unstructured":"A. Jansen M. Plakal R. Pandya etal 2017. Unsupervised learning of semantic audio representations. ArXiv Preprint ArXiv:1711.02209 2017.  A. Jansen M. Plakal R. Pandya et al. 2017. Unsupervised learning of semantic audio representations. ArXiv Preprint ArXiv:1711.02209 2017."},{"volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing. IEEE","author":"Kamper H.","key":"e_1_2_1_20_1"},{"key":"e_1_2_1_21_1","first-page":"2470","article-title":"2015. Multi-and cross-modal semantics beyond vision: Grounding in auditory perception. In Proceedings of the 2015 Conference on Empirical Methods","volume":"2461","author":"Kiela D.","year":"2015","journal-title":"Natural Language Processing"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/3207692.3207714"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7179089"},{"volume-title":"Annual Meeting of the Association for Computational Linguistics","year":"2014","author":"Levy O.","key":"e_1_2_1_24_1"},{"volume-title":"Overview of TASS 2017. TASS 2017: Workshop on Sentiment Analysis at SEPLN.","year":"2017","author":"Mart\u00ednez-C\u00e1mara E.","key":"e_1_2_1_25_1"},{"key":"e_1_2_1_26_1","unstructured":"T. Mikolov K. Chen G. Corrado etal 2013a. Efficient estimation of word representations in vector space. Arxiv Preprint Arxiv:1301.3781. 2013.  T. Mikolov K. Chen G. Corrado et al. 2013a. Efficient estimation of word representations in vector space. Arxiv Preprint Arxiv:1301.3781. 2013."},{"key":"e_1_2_1_27_1","first-page":"3111","article-title":"2013b. Distributed representations of words and phrases and their compositionality","volume":"2013","author":"Mikolov T.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_28_1","first-page":"1088","article-title":"2009. A scalable hierarchical distributed language model","volume":"1081","author":"Mnih A.","year":"2009","journal-title":"Advances in Neural Information Processing Systems"},{"volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP'14)","author":"Pennington J.","key":"e_1_2_1_29_1"},{"volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","year":"2014","author":"Qiu L.","key":"e_1_2_1_30_1"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/365628.365657"},{"key":"e_1_2_1_32_1","first-page":"26","article-title":"2017. Neural sentence embedding using only in-domain sentences for out-of-domain sentence detection in dialog systems","volume":"2017","author":"Ryu S.","journal-title":"Pattern Recognition Letters"},{"volume-title":"Harcourt","year":"1921","author":"Sapir Edward","key":"e_1_2_1_33_1"},{"volume-title":"Course in General Linguistics","year":"1915","author":"Saussure F. D.","key":"e_1_2_1_34_1"},{"key":"e_1_2_1_35_1","doi-asserted-by":"crossref","unstructured":"A. K. Vijayakumar R. Vedantam and D. Parikh. 2017. Sound-word2vec: Learning word representations grounded in sounds. ArXiv Preprint Arxiv:1703.01720 2017.  A. K. Vijayakumar R. Vedantam and D. Parikh. 2017. Sound-word2vec: Learning word representations grounded in sounds. ArXiv Preprint Arxiv:1703.01720 2017.","DOI":"10.18653\/v1\/D17-1096"},{"volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","year":"2018","author":"Wang S.","key":"e_1_2_1_36_1"},{"key":"e_1_2_1_37_1","unstructured":"Y. Xu and J. Liu. 2017. Implicitly incorporating morphological information into word embedding. arXiv preprint. 2017:1701.02481.  Y. Xu and J. Liu. 2017. Implicitly incorporating morphological information into word embedding. arXiv preprint. 2017:1701.02481."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1119"},{"key":"e_1_2_1_39_1","first-page":"291","article-title":"2017. Joint embeddings of Chinese words, characters, and fine-grained subcharacter components. In Proceedings of the 2017 Conference on Empirical Methods","volume":"286","author":"Yu J.","year":"2017","journal-title":"Natural Language Processing"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1011"},{"key":"e_1_2_1_41_1","first-page":"657","article-title":"2015. Character-level convolutional networks for text classification","volume":"649","author":"Zhang X.","year":"2015","journal-title":"Advances in Neural Information Processing Systems"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3344920","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3344920","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:54:27Z","timestamp":1750204467000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3344920"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,1,15]]},"references-count":41,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2020,3,31]]}},"alternative-id":["10.1145\/3344920"],"URL":"https:\/\/doi.org\/10.1145\/3344920","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2020,1,15]]},"assertion":[{"value":"2018-08-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-01-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}