{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:29:01Z","timestamp":1750307341646,"version":"3.41.0"},"reference-count":25,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2011,9,1]],"date-time":"2011-09-01T00:00:00Z","timestamp":1314835200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["6.09E+15"],"award-info":[{"award-number":["6.09E+15"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002920","name":"Research Grants Council, University Grants Committee, Hong Kong","doi-asserted-by":"publisher","award":["2.01E+13"],"award-info":[{"award-number":["2.01E+13"]}],"id":[{"id":"10.13039\/501100002920","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Transactions on Asian Language Information Processing"],"published-print":{"date-parts":[[2011,9]]},"abstract":"<jats:p>Treebanks are valuable resources for syntactic parsing. For some languages such as Chinese, we can obtain multiple constituency treebanks which are developed by different organizations. However, due to discrepancies of underlying annotation standards, such treebanks in general cannot be used together through direct data combination. To enlarge training data for syntactic parsing, we focus in this article on the challenge of unifying standards of disparate treebanks by automatically converting one treebank (source treebank) to fit a different standard which is exhibited by another treebank (target treebank).<\/jats:p>\n          <jats:p>\n            We propose to convert a treebank in two sequential steps which correspond to the part-of-speech level and syntactic structure level (including tree structures and grammar labels), respectively. Approaches used in both levels can be unified as an\n            <jats:italic>informed decoding<\/jats:italic>\n            procedure, where information derived from original annotation in a source treebank is used to guide the conversion conducted by a POS tagger (or a parser in the syntactic structure level) trained on a target treebank. We take two Chinese treebanks as a case study, and experiments on these two treebanks show significant improvements in conversion accuracy over baseline systems, especially in situations where a target treebank is small in size.\n          <\/jats:p>","DOI":"10.1145\/2002980.2002982","type":"journal-article","created":{"date-parts":[[2011,9,13]],"date-time":"2011-09-13T12:43:18Z","timestamp":1315917798000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Automatic Treebank Conversion via Informed Decoding - A Case Study on Chinese Treebanks"],"prefix":"10.1145","volume":"10","author":[{"given":"Muhua","family":"Zhu","sequence":"first","affiliation":[{"name":"Northeastern University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingbo","family":"Zhu","sequence":"additional","affiliation":[{"name":"Northeastern University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tong","family":"Xiao","sequence":"additional","affiliation":[{"name":"Northeastern University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2011,9]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.3115\/1073012.1073017"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1017\/S135132490700455X"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/974305.974323"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.3115\/1034678.1034754"},{"volume-title":"Proceedings of FRAME. 19--25","author":"Ekeklint S.","key":"e_1_2_1_6_1","unstructured":"Ekeklint , S. and Nivre , J . 2007. A dependency-based conversion of propbank . In Proceedings of FRAME. 19--25 . Ekeklint, S. and Nivre, J. 2007. A dependency-based conversion of propbank. In Proceedings of FRAME. 19--25."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1162\/089120105775299177"},{"key":"e_1_2_1_8_1","volume-title":"Proceedings of the Association for Computational Linguistics (ACL\u201908)","author":"Huang L.","year":"2008","unstructured":"Huang , L. 2008 . Forest reranking: Discriminative parsing with non-local features . In Proceedings of the Association for Computational Linguistics (ACL\u201908) . 586--594. Huang, L. 2008. Forest reranking: Discriminative parsing with non-local features. In Proceedings of the Association for Computational Linguistics (ACL\u201908). 586--594."},{"volume-title":"Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Conference on Computational Natural Language Learning (EMNLP-CoNLL\u201907)","author":"Huang Z.","key":"e_1_2_1_9_1","unstructured":"Huang , Z. , Harper , M. P. , and Wang , W . 2007. Mandarin part-of-speech tagging and discriminative reranking . In Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Conference on Computational Natural Language Learning (EMNLP-CoNLL\u201907) . 1093--1102. Huang, Z., Harper, M. P., and Wang, W. 2007. Mandarin part-of-speech tagging and discriminative reranking. In Proceedings of the Joint Conference on Empirical Methods in Natural Language Processing and Conference on Computational Natural Language Learning (EMNLP-CoNLL\u201907). 1093--1102."},{"volume-title":"Proceedings of the International Conference on Parsing Technologies (IWPT\u201909)","author":"Jiang W.","key":"e_1_2_1_10_1","unstructured":"Jiang , W. and Liu , Q . 2009. Automatic adaptation of annotation standards for dependency parsing - Using projected treebank as source corpus . In Proceedings of the International Conference on Parsing Technologies (IWPT\u201909) . 25--28. Jiang, W. and Liu, Q. 2009. Automatic adaptation of annotation standards for dependency parsing - Using projected treebank as source corpus. In Proceedings of the International Conference on Parsing Technologies (IWPT\u201909). 25--28."},{"volume-title":"Proceedings of the Association for Computational Linguistics (ACL\u201909)","author":"Jiang W.","key":"e_1_2_1_11_1","unstructured":"Jiang , W. , Huang , L. , and Liu , Q . 2009. Automatic adaptation of annotation standards: Chinese word segmentation and pos tagging - A case study . In Proceedings of the Association for Computational Linguistics (ACL\u201909) . 522--539. Jiang, W., Huang, L., and Liu, Q. 2009. Automatic adaptation of annotation standards: Chinese word segmentation and pos tagging - A case study. In Proceedings of the Association for Computational Linguistics (ACL\u201909). 522--539."},{"volume-title":"Proceedings of the 16th Nordic Conference of Computational Linguistics (NODALIDA\u201907)","author":"Johansson R.","key":"e_1_2_1_12_1","unstructured":"Johansson , R. and Nugues , P . 2007. Extended constituent-to-dependency conversion for english . In Proceedings of the 16th Nordic Conference of Computational Linguistics (NODALIDA\u201907) . 105--112. Johansson, R. and Nugues, P. 2007. Extended constituent-to-dependency conversion for english. In Proceedings of the 16th Nordic Conference of Computational Linguistics (NODALIDA\u201907). 105--112."},{"volume-title":"Proceedings of the Conference on Human Language Technologies (HLT\u201902)","author":"Kingsbury P.","key":"e_1_2_1_13_1","unstructured":"Kingsbury , P. , Palmer , M. , and Marcus , M . 2002. Adding semantic annotation to the penn treebank . In Proceedings of the Conference on Human Language Technologies (HLT\u201902) . Kingsbury, P., Palmer, M., and Marcus, M. 2002. Adding semantic annotation to the penn treebank. In Proceedings of the Conference on Human Language Technologies (HLT\u201902)."},{"volume-title":"Proceedings of the 5th SIGHAN Workshop (SIGHAN\u201905)","author":"Low J. K.","key":"e_1_2_1_14_1","unstructured":"Low , J. K. , Ng , H. T. , and Guo , W . 2005. A maximum entropy approach to Chinese word segmentation . In Proceedings of the 5th SIGHAN Workshop (SIGHAN\u201905) . 161--164. Low, J. K., Ng, H. T., and Guo, W. 2005. A maximum entropy approach to Chinese word segmentation. In Proceedings of the 5th SIGHAN Workshop (SIGHAN\u201905). 161--164."},{"volume-title":"Proceedings of the Association for Computational Linguistics (ACL\u201909)","author":"Niu Z.-Y.","key":"e_1_2_1_15_1","unstructured":"Niu , Z.-Y. , Wang , H. , and Wu , H . 2009. Exploiting heterogeneous treebanks for parsing . In Proceedings of the Association for Computational Linguistics (ACL\u201909) . 46--54. Niu, Z.-Y., Wang, H., and Wu, H. 2009. Exploiting heterogeneous treebanks for parsing. In Proceedings of the Association for Computational Linguistics (ACL\u201909). 46--54."},{"volume-title":"Inductive Dependency Parsing","author":"Nivre J.","key":"e_1_2_1_16_1","unstructured":"Nivre , J. 2006. Inductive Dependency Parsing . Springer , Volume 34. Nivre, J. 2006. Inductive Dependency Parsing. Springer, Volume 34."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.3115\/1220175.1220230"},{"volume-title":"Proceedings of the 6th International Workshop on Treebanks and Linguistic Theories (TLT\u201907)","author":"Rehbein I.","key":"e_1_2_1_18_1","unstructured":"Rehbein , I. and Genabith , J . 2007. Why is it so difficult to compare treebanks? Tiger and tuba-d\/z revisited . In Proceedings of the 6th International Workshop on Treebanks and Linguistic Theories (TLT\u201907) . Rehbein, I. and Genabith, J. 2007. Why is it so difficult to compare treebanks? Tiger and tuba-d\/z revisited. In Proceedings of the 6th International Workshop on Treebanks and Linguistic Theories (TLT\u201907)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.3115\/1034678.1034712"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.3115\/981732.981766"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.3115\/1072133.1072147"},{"volume-title":"Proceedings of the 7th International Workshop on Treebanks and Linguistic Theories (TLT\u201908)","author":"Xia F.","key":"e_1_2_1_22_1","unstructured":"Xia , F. , Bhatt , R. , Rambow , O. , Palmer , M. , and Sharma , D. M . 2008. Towards a multi-representational treebank . In Proceedings of the 7th International Workshop on Treebanks and Linguistic Theories (TLT\u201908) . 159--170. Xia, F., Bhatt, R., Rambow, O., Palmer, M., and Sharma, D. M. 2008. Towards a multi-representational treebank. In Proceedings of the 7th International Workshop on Treebanks and Linguistic Theories (TLT\u201908). 159--170."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.3115\/1072228.1072373"},{"volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201909)","author":"Zhang H.","key":"e_1_2_1_24_1","unstructured":"Zhang , H. , Zhang , M. , Tan , C. L. , and Marcus , M . 2009. K-best combination of syntactic parsers . In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201909) . 1552--1560. Zhang, H., Zhang, M., Tan, C. L., and Marcus, M. 2009. K-best combination of syntactic parsers. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP\u201909). 1552--1560."},{"volume-title":"Proceedings of the Conference on Computational Linguistics (COLING\u201910)","author":"Zhu M.","key":"e_1_2_1_26_1","unstructured":"Zhu , M. and Zhu , J . 2010. Automatic treebank conversion via informed decoding . In Proceedings of the Conference on Computational Linguistics (COLING\u201910) . 1344--1352. Zhu, M. and Zhu, J. 2010. Automatic treebank conversion via informed decoding. In Proceedings of the Conference on Computational Linguistics (COLING\u201910). 1344--1352."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1645953.1646145"}],"container-title":["ACM Transactions on Asian Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2002980.2002982","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2002980.2002982","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T11:06:22Z","timestamp":1750244782000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2002980.2002982"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,9]]},"references-count":25,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2011,9]]}},"alternative-id":["10.1145\/2002980.2002982"],"URL":"https:\/\/doi.org\/10.1145\/2002980.2002982","relation":{},"ISSN":["1530-0226","1558-3430"],"issn-type":[{"type":"print","value":"1530-0226"},{"type":"electronic","value":"1558-3430"}],"subject":[],"published":{"date-parts":[[2011,9]]},"assertion":[{"value":"2010-11-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2011-09-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}