{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,15]],"date-time":"2025-11-15T10:19:52Z","timestamp":1763201992481,"version":"3.41.0"},"reference-count":23,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2016,7,22]],"date-time":"2016-07-22T00:00:00Z","timestamp":1469145600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["91224006 and 61173063"],"award-info":[{"award-number":["91224006 and 61173063"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Ministry of Science and Technology of China","award":["201303107"],"award-info":[{"award-number":["201303107"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2017,3,31]]},"abstract":"<jats:p>In natural language, people often misuse a word (called a \u201cconfused word\u201d) in place of other words (called \u201cconfusing words\u201d). In misspelling corrections, many approaches to finding and correcting misspelling errors are based on a simple notion called a \u201cconfusion set.\u201d The confusion set of a confused word consists of confusing words. In this article, we propose a new method of building Chinese character confusion sets.<\/jats:p>\n          <jats:p>Our method is composed of two major phases. In the first phase, we build a list of seed confusion sets for each Chinese character, which is based on measuring similarity in character pinyin or similarity in character shape. In this phase, all confusion sets are constructed manually, and the confusion sets are organized into a graph, called a \u201cseed confusion graph\u201d (SCG), in which vertices denote characters and edges are pairs of characters in the form (confused character, confusing character).<\/jats:p>\n          <jats:p>In the second phase, we extend the SCG by acquiring more pairs of (confused character, confusing character) from a large Chinese corpus. For this, we use several word patterns (or patterns) to generate new confusion pairs and then verify the pairs before adding them into a SCG. Comprehensive experiments show that our method of extending confusion sets is effective. Also, we shall use the confusion sets in Chinese misspelling corrections to show the utility of our method.<\/jats:p>","DOI":"10.1145\/2933396","type":"journal-article","created":{"date-parts":[[2016,7,22]],"date-time":"2016-07-22T12:55:37Z","timestamp":1469192137000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["A Seed-Based Method for Generating Chinese Confusion Sets"],"prefix":"10.1145","volume":"16","author":[{"given":"Liangliang","family":"Liu","sequence":"first","affiliation":[{"name":"School of Business Information, Shanghai University of International Business and Economics, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cungen","family":"Cao","sequence":"additional","affiliation":[{"name":"Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2016,7,22]]},"reference":[{"key":"e_1_2_1_1_1","first-page":"1328","article-title":"Pinyin-indexed method for approximate matching in chinese","volume":"49","author":"Jiang CAO, Xiaojun WU, Yunqing","year":"2009","unstructured":"Jiang CAO, Xiaojun WU, Yunqing XIA, and Fang ZHENG. 2009 . Pinyin-indexed method for approximate matching in chinese . Journal of Tsinghua University (Science and Technology) 49 , S1 (2009), 1328 -- 1332 . Jiang CAO, Xiaojun WU, Yunqing XIA, and Fang ZHENG. 2009. Pinyin-indexed method for approximate matching in chinese. Journal of Tsinghua University (Science and Technology) 49, S1 (2009), 1328--1332.","journal-title":"Journal of Tsinghua University (Science and Technology)"},{"key":"e_1_2_1_2_1","first-page":"127","article-title":"Improving the template generation or chinese character error detection with confusion sets","volume":"15","author":"Chen Yong-Zhi","year":"2010","unstructured":"Yong-Zhi Chen , Shih-Hung Wu , Ping-che Yang, and Tsun Ku . 2010 . Improving the template generation or chinese character error detection with confusion sets . Comput. Ling. Chinese Lang. Processing 15 , 2 (2010), 127 -- 144 . Yong-Zhi Chen, Shih-Hung Wu, Ping-che Yang, and Tsun Ku. 2010. Improving the template generation or chinese character error detection with confusion sets. Comput. Ling. Chinese Lang. Processing 15, 2 (2010), 127--144.","journal-title":"Comput. Ling. Chinese Lang. Processing"},{"key":"e_1_2_1_3_1","first-page":"12","article-title":"New method of character string similarity compute based on fusing multiple edit distances","volume":"27","author":"Diao Xing-Chun","year":"2010","unstructured":"Xing-Chun Diao , Ming-Chao Tan , and Jian-Jun Cao . 2010 . New method of character string similarity compute based on fusing multiple edit distances . Appl. Res. Comput. 27 (2010), 12 . Xing-Chun Diao, Ming-Chao Tan, and Jian-Jun Cao. 2010. New method of character string similarity compute based on fusing multiple edit distances. Appl. Res. Comput. 27 (2010), 12.","journal-title":"Appl. Res. Comput."},{"key":"e_1_2_1_4_1","volume-title":"1992 IEEE International Conference on","volume":"1","author":"Essen Ute","year":"1992","unstructured":"Ute Essen and Volker Steinbiss . 1992 . Cooccurrence smoothing for stochastic language modeling. In Acous tics, Speech, and Signal Processing, 1992. ICASSP-92 ., 1992 IEEE International Conference on , Vol. 1 . IEEE, 161--164. Ute Essen and Volker Steinbiss. 1992. Cooccurrence smoothing for stochastic language modeling. In Acous tics, Speech, and Signal Processing, 1992. ICASSP-92., 1992 IEEE International Conference on, Vol. 1. IEEE, 161--164."},{"key":"e_1_2_1_5_1","first-page":"266","article-title":"Sound Distinguishing Method in Speech Sound Inquiry. (July 26 2006)","volume":"1","author":"Feng Qiangze","year":"2006","unstructured":"Qiangze Feng and Cungen Cao . 2006 . Sound Distinguishing Method in Speech Sound Inquiry. (July 26 2006) . CN Patent 1 , 266 ,633. Qiangze Feng and Cungen Cao. 2006. Sound Distinguishing Method in Speech Sound Inquiry. (July 26 2006). CN Patent 1,266,633.","journal-title":"CN Patent"},{"doi-asserted-by":"publisher","key":"e_1_2_1_6_1","DOI":"10.1007\/978-3-540-70939-8_55"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the Third Workshop on Very Large Corpora","volume":"3","author":"Golding Andrew R.","year":"1995","unstructured":"Andrew R. Golding . 1995 . A bayesian hybrid method for context-sensitive spelling correction . In Proceedings of the Third Workshop on Very Large Corpora , Vol. 3 . Massachusetts Institute of Technology, Cambridge, MA, 39--53. Andrew R. Golding. 1995. A bayesian hybrid method for context-sensitive spelling correction. In Proceedings of the Third Workshop on Very Large Corpora, Vol. 3. Massachusetts Institute of Technology, Cambridge, MA, 39--53."},{"doi-asserted-by":"publisher","key":"e_1_2_1_8_1","DOI":"10.1023\/A:1007545901558"},{"doi-asserted-by":"publisher","key":"e_1_2_1_9_1","DOI":"10.3115\/981863.981873"},{"key":"e_1_2_1_10_1","volume-title":"Keith Vander Linden, and Nigel Ward","author":"Jurafsky Dan","year":"2000","unstructured":"Dan Jurafsky , James H. Martin , Andrew Kehler , Keith Vander Linden, and Nigel Ward . 2000 . Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. Vol. 2 . MIT Press . Dan Jurafsky, James H. Martin, Andrew Kehler, Keith Vander Linden, and Nigel Ward. 2000. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition. Vol. 2. MIT Press."},{"doi-asserted-by":"publisher","key":"e_1_2_1_11_1","DOI":"10.3115\/1220175.1220304"},{"key":"e_1_2_1_12_1","first-page":"318","article-title":"A stroke-segment-mesh (SSM) glyph description method of chinese characters","volume":"47","author":"Lin M.","year":"2010","unstructured":"M. Lin and R. Song . 2010 . A stroke-segment-mesh (SSM) glyph description method of chinese characters . J. Comput. Res. Dev. 47 , 2 (2010), 318 -- 327 . M. Lin and R. Song. 2010. A stroke-segment-mesh (SSM) glyph description method of chinese characters. J. Comput. Res. Dev. 47, 2 (2010), 318--327.","journal-title":"J. Comput. Res. Dev."},{"doi-asserted-by":"publisher","key":"e_1_2_1_13_1","DOI":"10.5555\/1557690.1557715"},{"doi-asserted-by":"publisher","key":"e_1_2_1_14_1","DOI":"10.5555\/1667583.1667593"},{"key":"e_1_2_1_15_1","first-page":"77","article-title":"Automatic text error detection in domain question answering","volume":"27","author":"Liu Liangliang","year":"2013","unstructured":"Liangliang Liu , Shi Wang , Dongsheng Wang , Pingze Wang , and Cungen Cao . 2013 . Automatic text error detection in domain question answering . J. Chin. Inform. Process. 27 , 3 (2013), 77 -- 83 . Liangliang Liu, Shi Wang, Dongsheng Wang, Pingze Wang, and Cungen Cao. 2013. Automatic text error detection in domain question answering. J. Chin. Inform. Process. 27, 3 (2013), 77--83.","journal-title":"J. Chin. Inform. Process."},{"doi-asserted-by":"publisher","key":"e_1_2_1_16_1","DOI":"10.1016\/0306-4573(91)90066-U"},{"key":"e_1_2_1_17_1","first-page":"31","article-title":"A hybrid method for chinese text collation","volume":"12","author":"Meng Yu","year":"1998","unstructured":"Yu Meng and Yao Tian-Shun . 1998 . A hybrid method for chinese text collation . J. Chin. Inform. Process. 12 , 2 (1998), 31 -- 36 . Yu Meng and Yao Tian-Shun. 1998. A hybrid method for chinese text collation. J. Chin. Inform. Process. 12, 2 (1998), 31--36.","journal-title":"J. Chin. Inform. Process."},{"key":"e_1_2_1_18_1","volume-title":"2001 IEEE International Conference on","volume":"3","author":"Ren Fuji","year":"2001","unstructured":"Fuji Ren , Hongchi Shi , and Qiang Zhou . 2001 . A hybrid approach to automatic chinese text checking and error correction. In Systems, Man, and Cybernetics , 2001 IEEE International Conference on , Vol. 3 . IEEE, 1693--1698. Fuji Ren, Hongchi Shi, and Qiang Zhou. 2001. A hybrid approach to automatic chinese text checking and error correction. In Systems, Man, and Cybernetics, 2001 IEEE International Conference on, Vol. 3. IEEE, 1693--1698."},{"key":"e_1_2_1_19_1","first-page":"1968","article-title":"Similarity calculation of chinese character glyph and its application in computer aided proofreading system","volume":"29","author":"Song R.","year":"1964","unstructured":"R. Song , M. Lin , and S. L. Ge . 1964 . Similarity calculation of chinese character glyph and its application in computer aided proofreading system . J. Chin. Comput. Syst. 29 , 10 (1964), 1968 . R. Song, M. Lin, and S. L. Ge. 1964. Similarity calculation of chinese character glyph and its application in computer aided proofreading system. J. Chin. Comput. Syst. 29, 10 (1964), 1968.","journal-title":"J. Chin. Comput. Syst."},{"key":"e_1_2_1_20_1","first-page":"205","article-title":"Automatic component layer analysis method for chinese characters. (Feb. 8 2012)","volume":"201","author":"Wang Shi","year":"2012","unstructured":"Shi Wang , Weimin Wang , and Jianhui Fu . 2012 a. Automatic component layer analysis method for chinese characters. (Feb. 8 2012) . CN Patent App. CN 201 ,110, 205 ,810. Shi Wang, Weimin Wang, and Jianhui Fu. 2012a. Automatic component layer analysis method for chinese characters. (Feb. 8 2012). CN Patent App. CN 201,110,205,810.","journal-title":"CN Patent App. CN"},{"key":"e_1_2_1_21_1","first-page":"205","article-title":"Chinese character pattern cognition similarity computing method. (March 28 2012)","volume":"201","author":"Wang Shi","year":"2012","unstructured":"Shi Wang , Weimin Wang , and Jianhui Fu . 2012 b. Chinese character pattern cognition similarity computing method. (March 28 2012) . CN Patent App. CN 201 ,110, 205 ,807. Shi Wang, Weimin Wang, and Jianhui Fu. 2012b. Chinese character pattern cognition similarity computing method. (March 28 2012). CN Patent App. CN 201,110,205,807.","journal-title":"CN Patent App. CN"},{"doi-asserted-by":"publisher","key":"e_1_2_1_22_1","DOI":"10.3115\/1075218.1075250"},{"volume-title":"Proceedings of the First Workshop on Natural Language Processing and Neural Networks.","author":"Zhang Lei","unstructured":"Lei Zhang , Ming Zhou , Changning Huang , and H. H. Pan . 1999. Multifeature-based approach to automatic error detection and correction of chinese text . In Proceedings of the First Workshop on Natural Language Processing and Neural Networks. Lei Zhang, Ming Zhou, Changning Huang, and H. H. Pan. 1999. Multifeature-based approach to automatic error detection and correction of chinese text. In Proceedings of the First Workshop on Natural Language Processing and Neural Networks.","key":"e_1_2_1_23_1"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2933396","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2933396","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:56:03Z","timestamp":1750222563000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2933396"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,7,22]]},"references-count":23,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2017,3,31]]}},"alternative-id":["10.1145\/2933396"],"URL":"https:\/\/doi.org\/10.1145\/2933396","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2016,7,22]]},"assertion":[{"value":"2015-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-07-22","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}