{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:22:04Z","timestamp":1750220524957,"version":"3.41.0"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2021,3,31]],"date-time":"2021-03-31T00:00:00Z","timestamp":1617148800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2021,3,31]]},"abstract":"<jats:p>Logographic and alphabetic languages (e.g., Chinese vs. English) have different writing systems linguistically. Languages belonging to the same writing system usually exhibit more sharing information, which can be used to facilitate natural language processing tasks such as neural machine translation (NMT). This article takes advantage of the logographic characters in Chinese and Japanese by decomposing them into smaller units, thus more optimally utilizing the information these characters share in the training of NMT systems in both encoding and decoding processes. Experiments show that the proposed method can robustly improve the NMT performance of both \u201clogographic\u201d language pairs (JA\u2013ZH) and \u201clogographic + alphabetic\u201d (JA\u2013EN and ZH\u2013EN) language pairs in both supervised and unsupervised NMT scenarios. Moreover, as the decomposed sequences are usually very long, extra position features for the transformer encoder can help with the modeling of these long sequences. The results also indicate that, theoretically, linguistic features can be manipulated to obtain higher share token rates and further improve the performance of natural language processing systems.<\/jats:p>","DOI":"10.1145\/3431727","type":"journal-article","created":{"date-parts":[[2021,4,15]],"date-time":"2021-04-15T18:11:26Z","timestamp":1618510286000},"page":"1-15","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Using Sub-character Level Information for Neural Machine Translation of Logographic Languages"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2392-9192","authenticated-orcid":false,"given":"Longtu","family":"Zhang","sequence":"first","affiliation":[{"name":"Tokyo Metropolitan University, Hino, Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mamoru","family":"Komachi","sequence":"additional","affiliation":[{"name":"Tokyo Metropolitan University, Hino, Tokyo, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,4,15]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the ICLR.","author":"Artetxe Mikel","year":"2018","unstructured":"Mikel Artetxe , Gorka Labaka , Eneko Agirre , and Kyunghyun Cho . 2018 . Unsupervised neural machine translation . In Proceedings of the ICLR. Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2018. Unsupervised neural machine translation. In Proceedings of the ICLR."},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the ICLR.","author":"Bahdanau Dzmitry","year":"2015","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2015 . Neural machine translation by jointly learning to align and translate . In Proceedings of the ICLR. Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural machine translation by jointly learning to align and translate. In Proceedings of the ICLR."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/1717171"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the ISWC.","author":"Cao Shaosheng","year":"2017","unstructured":"Shaosheng Cao , Wei Lu , Jun Zhou , and Xiaolong Li . 2017 . Investigating stroke-level information for learning Chinese word embeddings . In Proceedings of the ISWC. Shaosheng Cao, Wei Lu, Jun Zhou, and Xiaolong Li. 2017. Investigating stroke-level information for learning Chinese word embeddings. In Proceedings of the ISWC."},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the AAAI\u201918","author":"Cao Shaosheng","year":"2018","unstructured":"Shaosheng Cao , Wei Lu , Jun Zhou , and Xiaolong Li . 2018 . cw2vec: Learning Chinese word embeddings with stroke n-gram information . In Proceedings of the AAAI\u201918 , IAAI\u201918, EAAI\u201918. 5053\u20135061. Shaosheng Cao, Wei Lu, Jun Zhou, and Xiaolong Li. 2018. cw2vec: Learning Chinese word embeddings with stroke n-gram information. In Proceedings of the AAAI\u201918, IAAI\u201918, EAAI\u201918. 5053\u20135061."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1461"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the LREC. European Language Resources Association (ELRA), 2149\u20132152","author":"Chu Chenhui","year":"2012","unstructured":"Chenhui Chu , Toshiaki Nakazawa , and Sadao Kurohashi . 2012 . Chinese characters mapping table of Japanese, traditional Chinese and simplified Chinese . In Proceedings of the LREC. European Language Resources Association (ELRA), 2149\u20132152 . Chenhui Chu, Toshiaki Nakazawa, and Sadao Kurohashi. 2012. Chinese characters mapping table of Japanese, traditional Chinese and simplified Chinese. In Proceedings of the LREC. European Language Resources Association (ELRA), 2149\u20132152."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1285"},{"key":"e_1_2_1_10_1","volume-title":"In Proceedings of the NAACL","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . In Proceedings of the NAACL 2019. 4171--4186. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. In Proceedings of the NAACL 2019. 4171--4186."},{"key":"e_1_2_1_11_1","volume-title":"large minibatch SGD: Training ImageNet in 1 hour. CoRR abs\/1706.02677","author":"Goyal Priya","year":"2017","unstructured":"Priya Goyal , Piotr Doll\u00e1r , Ross B. Girshick , Pieter Noordhuis , Lukasz Wesolowski , Aapo Kyrola , Andrew Tulloch , Yangqing Jia , and Kaiming He. 2017. Accurate , large minibatch SGD: Training ImageNet in 1 hour. CoRR abs\/1706.02677 ( 2017 ). Priya Goyal, Piotr Doll\u00e1r, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017. Accurate, large minibatch SGD: Training ImageNet in 1 hour. CoRR abs\/1706.02677 (2017)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_1_13_1","volume-title":"Proceedings of the ICLR. OpenReview.net.","author":"Kitaev Nikita","year":"2020","unstructured":"Nikita Kitaev , Lukasz Kaiser , and Anselm Levskaya . 2020 . Reformer: The efficient transformer . In Proceedings of the ICLR. OpenReview.net. Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. In Proceedings of the ICLR. OpenReview.net."},{"key":"e_1_2_1_14_1","volume-title":"Rush","author":"Klein Guillaume","year":"2018","unstructured":"Guillaume Klein , Yoon Kim , Yuntian Deng , Vincent Nguyen , Jean Senellart , and Alexander M . Rush . 2018 . OpenNMT: Neural machine translation toolkit. In Proceedings of the AMTA. 177\u2013184. Guillaume Klein, Yoon Kim, Yuntian Deng, Vincent Nguyen, Jean Senellart, and Alexander M. Rush. 2018. OpenNMT: Neural machine translation toolkit. In Proceedings of the AMTA. 177\u2013184."},{"key":"e_1_2_1_15_1","volume-title":"Apply Chinese radicals into neural machine translation: Deeper than character level. CoRR abs\/1805.01565","author":"Kuang Shaohui","year":"2018","unstructured":"Shaohui Kuang and Lifeng Han . 2018. Apply Chinese radicals into neural machine translation: Deeper than character level. CoRR abs\/1805.01565 ( 2018 ). Shaohui Kuang and Lifeng Han. 2018. Apply Chinese radicals into neural machine translation: Deeper than character level. CoRR abs\/1805.01565 (2018)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1549"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the EMNLP","author":"Luong Minh-Thang","year":"2015","unstructured":"Minh-Thang Luong , Hieu Pham , and Christopher D. Manning . 2015. Effective approaches to attention-based neural machine translation . In Proceedings of the EMNLP 2015 . 1412--1421. Minh-Thang Luong, Hieu Pham, and Christopher D. Manning. 2015. Effective approaches to attention-based neural machine translation. In Proceedings of the EMNLP 2015. 1412--1421."},{"key":"e_1_2_1_18_1","volume-title":"In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16)","author":"Nakazawa Toshiaki","year":"2016","unstructured":"Toshiaki Nakazawa , Manabu Yaguchi , Kiyotaka Uchimoto , Masao Utiyama , Eiichiro Sumita , Sadao Kurohashi , and Hitoshi Isahara . 2016 . In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16) . 2204--2208. Toshiaki Nakazawa, Manabu Yaguchi, Kiyotaka Uchimoto, Masao Utiyama, Eiichiro Sumita, Sadao Kurohashi, and Hitoshi Isahara. 2016. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16). 2204--2208."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-4007"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-4009"},{"key":"e_1_2_1_21_1","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. (2019).  Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. (2019)."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-1162"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the COLING. ACL, 298\u2013308","author":"Shen Mo","year":"2016","unstructured":"Mo Shen , Wingmui Li , HyunJeong Choe , Chenhui Chu , Daisuke Kawahara , and Sadao Kurohashi . 2016 . Consistent word segmentation, part-of-speech tagging and dependency labelling annotation for Chinese language . In Proceedings of the COLING. ACL, 298\u2013308 . Mo Shen, Wingmui Li, HyunJeong Choe, Chenhui Chu, Daisuke Kawahara, and Sadao Kurohashi. 2016. Consistent word segmentation, part-of-speech tagging and dependency labelling annotation for Chinese language. In Proceedings of the COLING. ACL, 298\u2013308."},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-2098"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2015.2513365"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969033.2969173"},{"key":"e_1_2_1_27_1","volume-title":"Chinese embedding via stroke and glyph information: A dual-channel view. CoRR abs\/1906.04287","author":"Tao Hanqing","year":"2019","unstructured":"Hanqing Tao , Shiwei Tong , Tong Xu , Qi Liu , and Enhong Chen . 2019. Chinese embedding via stroke and glyph information: A dual-channel view. CoRR abs\/1906.04287 ( 2019 ). Hanqing Tao, Shiwei Tong, Tong Xu, Qi Liu, and Enhong Chen. 2019. Chinese embedding via stroke and glyph information: A dual-channel view. CoRR abs\/1906.04287 (2019)."},{"key":"e_1_2_1_28_1","volume-title":"Proceedings of the IJCNLP. 378\u2013382","author":"Toyama Yota","year":"2017","unstructured":"Yota Toyama , Makoto Miwa , and Yutaka Sasaki . 2017 . Utilizing visual forms of Japanese characters for neural review classification . In Proceedings of the IJCNLP. 378\u2013382 . Yota Toyama, Makoto Miwa, and Yutaka Sasaki. 2017. Utilizing visual forms of Japanese characters for neural review classification. In Proceedings of the IJCNLP. 378\u2013382."},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_2_1_30_1","volume-title":"Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. CoRR abs\/1609.08144","author":"Wu Yonghui","year":"2016","unstructured":"Yonghui Wu , Mike Schuster , Zhifeng Chen , Quoc V. Le , Mohammad Norouzi , Wolfgang Macherey , Maxim Krikun , Yuan Cao , Qin Gao , Klaus Macherey , Jeff Klingner , Apurva Shah , Melvin Johnson , Xiaobing Liu , Lukasz Kaiser , Stephan Gouws , Yoshikiyo Kato , Taku Kudo , Hideto Kazawa , Keith Stevens , George Kurian , Nishant Patil , Wei Wang , Cliff Young , Jason Smith , Jason Riesa , Alex Rudnick , Oriol Vinyals , Greg Corrado , Macduff Hughes , and Jeffrey Dean . 2016. Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. CoRR abs\/1609.08144 ( 2016 ). Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016. Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. CoRR abs\/1609.08144 (2016)."},{"key":"e_1_2_1_31_1","volume-title":"Improving character-level Japanese-Chinese neural machine translation with radicals as an additional input feature. CoRR abs\/1805.02937","author":"Zhang Jinyi","year":"2018","unstructured":"Jinyi Zhang and Tadahiro Matsumoto . 2018. Improving character-level Japanese-Chinese neural machine translation with radicals as an additional input feature. CoRR abs\/1805.02937 ( 2018 ). Jinyi Zhang and Tadahiro Matsumoto. 2018. Improving character-level Japanese-Chinese neural machine translation with radicals as an additional input feature. CoRR abs\/1805.02937 (2018)."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-6303"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2860058"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3431727","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3431727","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:24:46Z","timestamp":1750195486000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3431727"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,31]]},"references-count":33,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,3,31]]}},"alternative-id":["10.1145\/3431727"],"URL":"https:\/\/doi.org\/10.1145\/3431727","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2021,3,31]]},"assertion":[{"value":"2019-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-04-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}