{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,3]],"date-time":"2025-09-03T10:44:56Z","timestamp":1756896296765,"version":"3.40.5"},"reference-count":27,"publisher":"Cambridge University Press (CUP)","issue":"5","license":[{"start":{"date-parts":[[2020,7,10]],"date-time":"2020-07-10T00:00:00Z","timestamp":1594339200000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2021,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We investigate the usage of semantic information for morphological segmentation since words that are derived from each other will remain semantically related. We use mathematical models such as maximum likelihood estimate (MLE) and maximum a posteriori estimate (MAP) by incorporating semantic information obtained from dense word vector representations. Our approach does not require any annotated data which make it fully unsupervised and require only a small amount of raw data together with pretrained word embeddings for training purposes. The results show that using dense vector representations helps in morphological segmentation especially for low-resource languages. We present results for Turkish, English, and German. Our semantic MLE model outperforms other unsupervised models for Turkish language. Our proposed models could be also used for any other low-resource language with concatenative morphology.<\/jats:p>","DOI":"10.1017\/s1351324920000406","type":"journal-article","created":{"date-parts":[[2020,7,10]],"date-time":"2020-07-10T06:38:07Z","timestamp":1594363087000},"page":"609-629","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":4,"title":["Incorporating word embeddings in unsupervised morphological segmentation"],"prefix":"10.1017","volume":"27","author":[{"given":"Ahmet","family":"\u00dcst\u00fcn","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1700-0395","authenticated-orcid":false,"given":"Burcu","family":"Can","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2020,7,10]]},"reference":[{"key":"S1351324920000406_ref6","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1140"},{"key":"S1351324920000406_ref24","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/N15-1186"},{"key":"S1351324920000406_ref14","unstructured":"Goldwater, S. , Johnson, M. and Griffiths, T.L. (2006). Interpolating between types and tokens by estimating power-law generators. In Proceedings of the Advances in Neural Information Processing Systems 18. MIT Press, pp. 459\u2013466."},{"key":"S1351324920000406_ref17","doi-asserted-by":"publisher","DOI":"10.1007\/978-94-017-6059-1_3"},{"key":"S1351324920000406_ref12","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.1984.4767596"},{"key":"S1351324920000406_ref2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-15754-7_77"},{"key":"S1351324920000406_ref3","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00318"},{"key":"S1351324920000406_ref9","unstructured":"Creutz, M. and Lagus, K. (2005b). Unsupervised Morpheme Segmentation and Morphology Induction from Text Corpora Using Morfessor 1.0. Technical Report A81."},{"key":"S1351324920000406_ref11","doi-asserted-by":"publisher","DOI":"10.3115\/981863.981907"},{"key":"S1351324920000406_ref26","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-45925-7_4"},{"key":"S1351324920000406_ref13","doi-asserted-by":"publisher","DOI":"10.1162\/089120101750300490"},{"key":"S1351324920000406_ref19","unstructured":"Lazaridou, A. , Marelli, M. , Zamparelli, R. and Baroni, M. (2013). Compositional-ly derived representations of morphologically complex words in distributional semantics. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Sofia, Bulgaria. Association for Computational Linguistics, pp. 1517\u20131526."},{"key":"S1351324920000406_ref27","doi-asserted-by":"crossref","unstructured":"\u00dcst\u00fcn, A. , Kurfal, M. and Can, B. (2018). Characters or morphemes: How to represent words? In Proceedings of The Third Workshop on Representation Learning for NLP, Melbourne, Australia. Association for Computational Linguistics, pp. 144\u2013153.","DOI":"10.18653\/v1\/W18-3019"},{"key":"S1351324920000406_ref10","doi-asserted-by":"publisher","DOI":"10.1145\/1187415.1187418"},{"key":"S1351324920000406_ref7","doi-asserted-by":"publisher","DOI":"10.3115\/1118647.1118650"},{"key":"S1351324920000406_ref22","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00130"},{"key":"S1351324920000406_ref15","unstructured":"Hankamer, J. (1986). Finite state morphology and left to right phonology. In Proceedings of the West Coast Conference on Formal Linguistics (WCCFL-5)."},{"key":"S1351324920000406_ref8","unstructured":"Creutz, M. and Lagus, K. (2005a). Inducing the morphological lexicon of a natural language from unannotated text. In Proceedings of the International and Interdisciplinary Conference on Adaptive Knowledge Representation and Reasoning (AKRR 2005), pp. 106\u2013113."},{"key":"S1351324920000406_ref21","unstructured":"Mikolov, T. , Chen, K. , Corrado, G. and Dean, J. (2013). Efficient estimation of word representations in vector space. CoRR, abs\/1301.3781."},{"key":"S1351324920000406_ref23","doi-asserted-by":"publisher","DOI":"10.3115\/1073336.1073360"},{"key":"S1351324920000406_ref5","doi-asserted-by":"publisher","DOI":"10.3115\/1117601.1117621"},{"key":"S1351324920000406_ref25","unstructured":"Team, D.D. (2016). Deeplearning4j: Open-source distributed deep learning for the JVM, Apache Software Foundation License 2.0. http:\/\/deeplearning4j.org\/ (accessed 10 February 2017)."},{"key":"S1351324920000406_ref1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"S1351324920000406_ref18","unstructured":"Kurimo, M. , Lagus, K. , Virpioja, S. and Turunen, V.T. (2011). Morpho Challenge 2010. http:\/\/research.ics.tkk.fi\/events\/morphochallenge2010\/ (accessed 10 February 2017)."},{"key":"S1351324920000406_ref20","unstructured":"Lee, Y.K. , Haghighi, A. and Barzilay, R. (2011). Modeling syntactic context improves morphological segmentation. In Proceedings of the Fifteenth Conference on Computational Natural Language Learning, CoNLL\u201911, Stroudsburg, PA, USA. Association for Computational Linguistics, pp. 1\u20139."},{"key":"S1351324920000406_ref4","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-1603"},{"key":"S1351324920000406_ref16","doi-asserted-by":"publisher","DOI":"10.2307\/411036"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324920000406","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T03:18:05Z","timestamp":1630466285000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324920000406\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,7,10]]},"references-count":27,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2021,9]]}},"alternative-id":["S1351324920000406"],"URL":"https:\/\/doi.org\/10.1017\/s1351324920000406","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"type":"print","value":"1351-3249"},{"type":"electronic","value":"1469-8110"}],"subject":[],"published":{"date-parts":[[2020,7,10]]},"assertion":[{"value":"\u00a9 The Author(s), 2020. Published by Cambridge University Press","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}}]}}