{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,25]],"date-time":"2025-12-25T09:04:03Z","timestamp":1766653443549,"version":"3.41.0"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2017,1,20]],"date-time":"2017-01-20T00:00:00Z","timestamp":1484870400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61333018"],"award-info":[{"award-number":["61333018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Strategic Priority Research Program of the CAS","award":["XDB02070007"],"award-info":[{"award-number":["XDB02070007"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2017,9,30]]},"abstract":"<jats:p>Phrase representation, an important step in many NLP tasks, involves representing phrases as continuous-valued vectors. This article presents detailed comparisons concerning the effects of word vectors, training data, and the composition and objective function used in a composition model for phrase representation. Specifically, we first discuss how the augmented word representations affect the performance of the composition model. Then, we investigate whether different types of training data influence the performance of the composition model and, if so, how they influence it. Finally, we evaluate combinations of different composition and objective functions and discuss the factors related to composition model performance. All evaluations were conducted in both English and Chinese. Our main findings are as follows: (1) The Additive model with semantic enhanced word vectors performs comparably to the state-of-the-art model; (2) The Additive model which updates augmented word vectors and the Matrix model with semantic enhanced word vectors systematically outperforms the state-of-the-art model in bigram and multi-word phrase similarity task, respectively; (3) Representing the high frequency phrases by estimating their surrounding contexts is a good training objective for bigram phrase similarity tasks; and (4) The performance gain of composition model with semantic enhanced word vectors is due to the composition function and the greater weight attached to important words. Previous works focus on the composition function; however, our findings indicate that other components in the composition model (especially word representation) make a critical difference in phrase representation.<\/jats:p>","DOI":"10.1145\/3010088","type":"journal-article","created":{"date-parts":[[2017,1,20]],"date-time":"2017-01-20T14:05:40Z","timestamp":1484921140000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Comparison Study on Critical Components in Composition Model for Phrase Representation"],"prefix":"10.1145","volume":"16","author":[{"given":"Shaonan","family":"Wang","sequence":"first","affiliation":[{"name":"National Laboratory of Pattern Recognition, Institute of Automation, University of Chinese Academy of Sciences, Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengqing","family":"Zong","sequence":"additional","affiliation":[{"name":"National Laboratory of Pattern Recognition, Institute of Automation, CAS Center for Excellence in Brain Science and Intelligence Technology, University of Chinese Academy of Sciences, Chinese Academy of Sciences"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,1,20]]},"reference":[{"volume-title":"Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing. 1183--1193","year":"2010","author":"Baroni Marco","key":"e_1_2_1_1_1"},{"volume-title":"Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics. 238--247","year":"2014","author":"Baroni Marco","key":"e_1_2_1_2_1"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944966"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/2390948.2391011"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944937"},{"volume-title":"Proceedings of the 30th AAAI Conference on Artificial Intelligence. 2690--2696","year":"2016","author":"Bollegala Danushka","key":"e_1_2_1_6_1"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1028"},{"volume-title":"Ltp: A chinese language technology platform. In Coling2010: Demonstrations.","year":"2010","author":"Che Wanxiang","key":"e_1_2_1_8_1"},{"volume-title":"Proceedings of the Main Conference on Human Language Technology Conference of the North American Chapter of the Association of Computational Linguistics. 17--24","year":"2006","author":"Chris Callison-Burch","key":"e_1_2_1_9_1"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.3115\/1690219.1690235"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390177"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2078186"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1188"},{"volume-title":"Proceedings of the ACL Workshop on Continuous Vector Space Models and Their Compositionality, 50--58","year":"2013","author":"Dinu Georgiana","key":"e_1_2_1_16_1"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.5555\/1953048.2021068"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1184"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-1113"},{"volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing. 1394--1404","year":"2011","author":"Grefenstette Edward","key":"e_1_2_1_21_1"},{"key":"e_1_2_1_22_1","unstructured":"Edward Grefenstette Georgiana Dinu Yao-Zhong Zhang Mehrnoosh Sadrzadeh and Marco Baroni. 2013. Multi-step regression learning for compositional distributional semantics. arXiv preprint arXiv:1301.6939.  Edward Grefenstette Georgiana Dinu Yao-Zhong Zhang Mehrnoosh Sadrzadeh and Marco Baroni. 2013. Multi-step regression learning for compositional distributional semantics. arXiv preprint arXiv:1301.6939."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.5555\/1870516.1870521"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1002"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1163"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N16-1162"},{"volume-title":"Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics. 873--882","author":"Huang Eric H.","key":"e_1_2_1_27_1"},{"volume-title":"Proceedings of the 29th Pacific Asia Conference on Language, Information and Computation. 328--336","year":"2015","author":"Iwai Miki","key":"e_1_2_1_28_1"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1162"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1242"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D15-1246"},{"key":"e_1_2_1_32_1","unstructured":"Omer Levy and Yoav Goldberg. 2014. Neural word embedding as implicit matrix factorization. In Advances in Neural Information Processing Systems. 2177--2185.  Omer Levy and Yoav Goldberg. 2014. Neural word embedding as implicit matrix factorization. In Advances in Neural Information Processing Systems. 2177--2185."},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-5010"},{"key":"e_1_2_1_34_1","unstructured":"Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013a. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781.  Tomas Mikolov Kai Chen Greg Corrado and Jeffrey Dean. 2013a. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781."},{"key":"e_1_2_1_35_1","unstructured":"Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013b. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems. 3111--3119.  Tomas Mikolov Ilya Sutskever Kai Chen Greg S Corrado and Jeff Dean. 2013b. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems. 3111--3119."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1079"},{"volume-title":"Proceedings of the 46th Annual Meeting of the Association for Computational Linguistics. 236--244","year":"2008","author":"Mitchell Jeff","key":"e_1_2_1_37_1"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1551-6709.2010.01106.x"},{"volume-title":"Proceedings of the 30th AAAI Conference on Artificial Intelligence.","year":"2016","author":"Mueller Jonas","key":"e_1_2_1_39_1"},{"volume-title":"Contemporary Issues in Cognitive Psychology the Loyola Symposium.","year":"1972","author":"Norman Donald A.","key":"e_1_2_1_40_1"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1094"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1045"},{"volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing. 151--161","author":"Socher Richard","key":"e_1_2_1_43_1"},{"volume-title":"Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning. 1201--1211","year":"2012","author":"Socher Richard","key":"e_1_2_1_44_1"},{"volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing. 1631","year":"2013","author":"Socher Richard","key":"e_1_2_1_45_1"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00177"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1035"},{"key":"e_1_2_1_48_1","unstructured":"Ran Tian Naoaki Okazaki and Kentaro Inui. 2015. The mechanism of additive composition. arXiv preprint arXiv:1511.08407.  Ran Tian Naoaki Okazaki and Kentaro Inui. 2015. The mechanism of additive composition. arXiv preprint arXiv:1511.08407."},{"volume-title":"Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics. 384--394","year":"2010","author":"Turian Joseph","key":"e_1_2_1_49_1"},{"volume-title":"Proceedings of the Conference of the North American Chapter of the Association of Computational Linguistics. 1142--1151","year":"2013","author":"de Cruys Tim Van","key":"e_1_2_1_50_1"},{"key":"e_1_2_1_51_1","doi-asserted-by":"crossref","unstructured":"John Wieting Mohit Bansal Kevin Gimpel Karen Livescu and Dan Roth. 2015. From paraphrase database to compositional paraphrase model and back. arXiv preprint arXiv:1506.03487.  John Wieting Mohit Bansal Kevin Gimpel Karen Livescu and Dan Roth. 2015. From paraphrase database to compositional paraphrase model and back. arXiv preprint arXiv:1506.03487.","DOI":"10.1162\/tacl_a_00246"},{"volume-title":"Proceedings of the 4th International Conference on Learning Representations.","year":"2016","author":"Wieting John","key":"e_1_2_1_52_1"},{"volume-title":"CHARAGRAM: Embedding words and sentences via character n-grams. arXiv preprint arXiv:1607.02789.","year":"2016","author":"Wieting John","key":"e_1_2_1_53_1"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P14-2089"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00135"},{"volume-title":"Proceedings of the 23rd International Conference on Computational Linguistics. 1263--1271","year":"2010","author":"Zanzotto Fabio M.","key":"e_1_2_1_56_1"},{"volume-title":"Proceedings of AAAI. 2195--2202","year":"2015","author":"Zhao Yu","key":"e_1_2_1_57_1"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3010088","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3010088","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:50:35Z","timestamp":1750218635000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3010088"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,1,20]]},"references-count":55,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2017,9,30]]}},"alternative-id":["10.1145\/3010088"],"URL":"https:\/\/doi.org\/10.1145\/3010088","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2017,1,20]]},"assertion":[{"value":"2016-05-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2017-01-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}