{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:44:27Z","timestamp":1750308267130,"version":"3.41.0"},"reference-count":15,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2004,6,1]],"date-time":"2004-06-01T00:00:00Z","timestamp":1086048000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Transactions on Asian Language Information Processing"],"published-print":{"date-parts":[[2004,6]]},"abstract":"<jats:p>Parsing, the task of identifying syntactic components, e.g., noun and verb phrases, in a sentence, is one of the fundamental tasks in natural language processing. Many natural language applications such as spoken-language understanding, machine translation, and information extraction, would benefit from, or even require, high accuracy parsing as a preprocessing step.<\/jats:p>\n          <jats:p>Even though most state-of-the-art statistical parsers were initially constructed for parsing in English, most of them are not language-specific, in that they do not rely on properties of the language that are specific to English. Therefore, construction of a parser in a given language becomes a matter of retraining the statistical parameters with a Treebank in the corresponding language.<\/jats:p>\n          <jats:p>The development of the Chinese treebank [Xia et al. 2000] spurred the construction of parsers for Chinese. However, Chinese as a language poses some unique problems for the development of a statistical parser, the most apparent being word segmentation. Since words in written Chinese are not delimited in the same way as in Western languages, the first problem that needs to be solved before an existing statistical method can be applied to Chinese is to identify the word boundaries. This is a step that is neglected by most pre-existing Chinese parsers, which assume that the input data has already been pre-segmented.<\/jats:p>\n          <jats:p>This article describes a character-based statistical parser, which gives the best performance to-date on the Chinese treebank data. We augment an existing maximum entropy parser with transformation-based learning, creating a parser that can operate at the character level. We present experiments that show that our parser achieves results that are close to those achievable under perfect word segmentation conditions.<\/jats:p>","DOI":"10.1145\/1034780.1034786","type":"journal-article","created":{"date-parts":[[2005,1,26]],"date-time":"2005-01-26T16:35:53Z","timestamp":1106757353000},"page":"159-168","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["A maximum-entropy chinese parser augmented by transformation-based learning"],"prefix":"10.1145","volume":"3","author":[{"given":"Pascale","family":"Fung","sequence":"first","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Grace","family":"Ngai","sequence":"additional","affiliation":[{"name":"Hong Kong Polytechnic University, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yongsheng","family":"Yang","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Benfeng","family":"Chen","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2004,6]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the Second Chinese Language Processing Workshop","author":"Bikel D. M.","year":"2000","unstructured":"Bikel , D. M. and Chiang , D . 2000. Two statistical parsing models applied to the Chinese treebank . In Proceedings of the Second Chinese Language Processing Workshop ( Hong Kong , 2000 ). 1--6. 10.3115\/1117769.1117771 Bikel, D. M. and Chiang, D. 2000. Two statistical parsing models applied to the Chinese treebank. In Proceedings of the Second Chinese Language Processing Workshop (Hong Kong, 2000). 1--6. 10.3115\/1117769.1117771"},{"key":"e_1_2_1_2_1","first-page":"543","article-title":"Transformation-based error-driven learning and natural language processing: A case study in part of speech tagging","volume":"21","author":"Brill E.","year":"1995","unstructured":"Brill , E. 1995 . Transformation-based error-driven learning and natural language processing: A case study in part of speech tagging . Computational Linguistics 21 , 4 (1995), 543 -- 565 . Brill, E. 1995. Transformation-based error-driven learning and natural language processing: A case study in part of speech tagging. Computational Linguistics 21, 4 (1995), 543--565.","journal-title":"Computational Linguistics"},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of CONLL'01","author":"Florian R.","year":"2001","unstructured":"Florian , R. and Ngai , G . 2001. Multidimensional transformation-based learning . In Proceedings of CONLL'01 ( Toulouse, France , 2001 ). 1--8. 10.3115\/1117822.1117823 Florian, R. and Ngai, G. 2001. Multidimensional transformation-based learning. In Proceedings of CONLL'01 (Toulouse, France, 2001). 1--8. 10.3115\/1117822.1117823"},{"key":"e_1_2_1_4_1","first-page":"69","article-title":"Error-driven segmentation of Chinese","volume":"8","author":"Hockenmeier J.","year":"1998","unstructured":"Hockenmeier , J. and Brew , C. 1998 . Error-driven segmentation of Chinese . Communications of COLIPS 8 , 1 (1998), 69 -- 84 . Hockenmeier, J. and Brew, C. 1998. Error-driven segmentation of Chinese. Communications of COLIPS 8, 1 (1998), 69--84.","journal-title":"Communications of COLIPS"},{"key":"e_1_2_1_5_1","volume-title":"Maximum Entropy and Bayesian Methods in Applied Statistics: Proceedings of the 4th Maximum Entropy Workshop","author":"Jaynes E. T.","year":"1984","unstructured":"Jaynes , E. T. 1984 . Monkeys, kangaroos and n . In Maximum Entropy and Bayesian Methods in Applied Statistics: Proceedings of the 4th Maximum Entropy Workshop ( Univ. of Calgary , 1984). J. H. Justice (ed.). 26--58. Jaynes, E. T. 1984. Monkeys, kangaroos and n. In Maximum Entropy and Bayesian Methods in Applied Statistics: Proceedings of the 4th Maximum Entropy Workshop (Univ. of Calgary, 1984). J. H. Justice (ed.). 26--58."},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 41st ACL","author":"Levy R.","year":"2003","unstructured":"Levy , R. and Manning , C. D . 2003. Is it harder to parse Chinese, or the Chinese treebank ? In Proceedings of the 41st ACL ( Sapporo, Japan , 2003 ). 439--446. 10.3115\/1075096.1075152 Levy, R. and Manning, C. D. 2003. Is it harder to parse Chinese, or the Chinese treebank? In Proceedings of the 41st ACL (Sapporo, Japan, 2003). 439--446. 10.3115\/1075096.1075152"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (July 2003","author":"Luo X.","year":"2003","unstructured":"Luo , X. 2003 . A maximum entropy Chinese character-based parser . In Proceedings of the Conference on Empirical Methods in Natural Language Processing (July 2003 ). 10.3115\/1119355.1119380 Luo, X. 2003. A maximum entropy Chinese character-based parser. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (July 2003). 10.3115\/1119355.1119380"},{"key":"e_1_2_1_8_1","first-page":"313","article-title":"Building a large annotated corpus of English: The Penn treebank","volume":"19","author":"Marcus M.","year":"1993","unstructured":"Marcus , M. , Marcinkiewicz , M. , and Santorini , B. 1993 . Building a large annotated corpus of English: The Penn treebank . Computational Linguistics 19 , 2 (1993), 313 -- 330 . Marcus, M., Marcinkiewicz, M., and Santorini, B. 1993. Building a large annotated corpus of English: The Penn treebank. Computational Linguistics 19, 2 (1993), 313--330.","journal-title":"Computational Linguistics"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 35th ACL (Madrid","author":"Palmer D.","year":"1997","unstructured":"Palmer , D. 1997 . A trainable rule-based algorithm for word segmentation . In Proceedings of the 35th ACL (Madrid 1997). 10.3115\/976909.979658 Palmer, D. 1997. A trainable rule-based algorithm for word segmentation. In Proceedings of the 35th ACL (Madrid 1997). 10.3115\/976909.979658"},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 3rd ACL Workshop on Very Large Corpora","author":"Ramshaw L.","year":"1995","unstructured":"Ramshaw , L. and Marcus , M . 1995. Text chunking using transformation-based learning . In Proceedings of the 3rd ACL Workshop on Very Large Corpora ( Cambridge, MA , 1995 ). Ramshaw, L. and Marcus, M. 1995. Text chunking using transformation-based learning. In Proceedings of the 3rd ACL Workshop on Very Large Corpora (Cambridge, MA, 1995)."},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the EMNLP","author":"Ratnaparkhi A.","year":"1996","unstructured":"Ratnaparkhi , A. 1996 . A maximum entropy part-of-speech tagger . In Proceedings of the EMNLP ( University of Pennsylvania , 1996). Ratnaparkhi, A. 1996. A maximum entropy part-of-speech tagger. In Proceedings of the EMNLP (University of Pennsylvania, 1996)."},{"key":"e_1_2_1_13_1","unstructured":"Sproat R. Shih C. Gale W. and Chang N. 1996. A stochastic finite-state word-segmentation algorithm for Chinese. Computational Linguistics 22 3 (1996).   Sproat R. Shih C. Gale W. and Chang N. 1996. A stochastic finite-state word-segmentation algorithm for Chinese. Computational Linguistics 22 3 (1996)."},{"volume-title":"Proceedings of the 2002 Human Language Technology Workshop.","author":"Xu J.","key":"e_1_2_1_14_1","unstructured":"Xu , J. , Miller , S. , and Weischedel , R . 2002. A statistical parser for Chinese . In Proceedings of the 2002 Human Language Technology Workshop. Xu, J., Miller, S., and Weischedel, R. 2002. A statistical parser for Chinese. In Proceedings of the 2002 Human Language Technology Workshop."},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of the 2nd International Conference on Language Resources and Evaluation (LREC-2000","author":"Xia F.","year":"2000","unstructured":"Xia , F. and Palmer , M ., et al. 2000. Developing guidelines and ensuring consistency for Chinese text annotation . In Proceedings of the 2nd International Conference on Language Resources and Evaluation (LREC-2000 , Athens , 2000 ). Xia, F. and Palmer, M., et al. 2000. Developing guidelines and ensuring consistency for Chinese text annotation. In Proceedings of the 2nd International Conference on Language Resources and Evaluation (LREC-2000, Athens, 2000)."},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the Fourth ACL Conference on Applied Natural Language Processing","author":"Wu D.","year":"1994","unstructured":"Wu , D. and Fung , P . 1994. Improving Chinese word segmentation with linguistic filters on statistical lexical acquisition . In Proceedings of the Fourth ACL Conference on Applied Natural Language Processing ( Stuttgart, Germany , 1994 ). 10.3115\/974358.974399 Wu, D. and Fung, P. 1994. Improving Chinese word segmentation with linguistic filters on statistical lexical acquisition. In Proceedings of the Fourth ACL Conference on Applied Natural Language Processing (Stuttgart, Germany,1994). 10.3115\/974358.974399"}],"container-title":["ACM Transactions on Asian Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1034780.1034786","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1034780.1034786","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:23:59Z","timestamp":1750267439000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1034780.1034786"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,6]]},"references-count":15,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2004,6]]}},"alternative-id":["10.1145\/1034780.1034786"],"URL":"https:\/\/doi.org\/10.1145\/1034780.1034786","relation":{},"ISSN":["1530-0226","1558-3430"],"issn-type":[{"type":"print","value":"1530-0226"},{"type":"electronic","value":"1558-3430"}],"subject":[],"published":{"date-parts":[[2004,6]]},"assertion":[{"value":"2004-06-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}