{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,4,2]],"date-time":"2024-04-02T00:18:09Z","timestamp":1712017089495},"reference-count":48,"publisher":"Cambridge University Press (CUP)","issue":"2","license":[{"start":{"date-parts":[[2023,5,10]],"date-time":"2023-05-10T00:00:00Z","timestamp":1683676800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2024,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Word order is one of the most important grammatical devices and the basis for language understanding. However, as one of the most popular NLP architectures, Transformer does not explicitly encode word order. A solution to this problem is to incorporate position information by means of position encoding\/embedding (PE). Although a variety of methods of incorporating position information have been proposed, the NLP community is still in want of detailed statistical researches on position information in real-life language. In order to understand the influence of position information on the correlation between words in more detail, we investigated the factors that affect the frequency of words and word sequences in large corpora. Our results show that absolute position, relative position, being at one of the two ends of a sentence and sentence length all significantly affect the frequency of words and word sequences. Besides, we observed that the frequency distribution of word sequences over relative position carries valuable grammatical information. Our study suggests that in order to accurately capture word\u2013word correlations, it is not enough to focus merely on absolute and relative position. Transformers should have access to more types of position-related information which may require improvements to the current architecture.<\/jats:p>","DOI":"10.1017\/s1351324923000128","type":"journal-article","created":{"date-parts":[[2023,5,10]],"date-time":"2023-05-10T17:40:04Z","timestamp":1683740404000},"page":"294-318","update-policy":"http:\/\/dx.doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":0,"title":["What should be encoded by position embedding for neural network language models?"],"prefix":"10.1017","volume":"30","author":[{"given":"Shuiyuan","family":"Yu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zihao","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haitao","family":"Liu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2023,5,10]]},"reference":[{"key":"S1351324923000128_ref23","unstructured":"Mikolov, T. , Chen, K. , Corrado, G. and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781."},{"key":"S1351324923000128_ref9","unstructured":"Goldberg, Y. (2019). Assessing bert\u2019s syntactic abilities. arXiv preprint arXiv:1901.05287."},{"key":"S1351324923000128_ref28","doi-asserted-by":"crossref","unstructured":"Pham, T.M. , Bui, T. , Mai, L. and Nguyen, A. (2020). Out of order: how important is the sequential order of words in a sentence in natural language understanding tasks? arXiv preprint arXiv:2012.15180.","DOI":"10.18653\/v1\/2021.findings-acl.98"},{"key":"S1351324923000128_ref5","unstructured":"Dufter, P. , Schmitt, M. and Sch\u00fctze, H. (2021). Position information in transformers: an overview. arXiv preprint arXiv:2102.11090."},{"key":"S1351324923000128_ref20","doi-asserted-by":"crossref","first-page":"159","DOI":"10.17791\/jcs.2008.9.2.159","article-title":"Dependency distance as a metric of language comprehension difficulty","volume":"9","author":"Liu","year":"2008","journal-title":"Journal of Cognitive Science"},{"key":"S1351324923000128_ref25","volume-title":"AACL\/IJCNLP.","author":"Park","year":"2020"},{"key":"S1351324923000128_ref4","unstructured":"Devlin, J. , Chang, M.-W. , Lee, K. and Toutanova, K. (2018). Bert: pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805."},{"key":"S1351324923000128_ref33","doi-asserted-by":"crossref","unstructured":"Schmitt, M. , Ribeiro, L.F. , Dufter, P. , Gurevych, I. and Sch\u00fctze, H. (2020). Modeling graph structure via relative position for text generation from knowledge graphs. arXiv preprint arXiv:2006.09242.","DOI":"10.18653\/v1\/11.textgraphs-1.2"},{"key":"S1351324923000128_ref42","doi-asserted-by":"crossref","unstructured":"Wang, Y.-A. and Chen, Y.-N. (2020). What do position embeddings learn? an empirical study of pre-trained language model positional encoding. arXiv preprint arXiv:2010.04903.","DOI":"10.18653\/v1\/2020.emnlp-main.555"},{"key":"S1351324923000128_ref41","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1145"},{"key":"S1351324923000128_ref21","doi-asserted-by":"publisher","DOI":"10.1016\/j.lingua.2009.10.001"},{"key":"S1351324923000128_ref31","doi-asserted-by":"publisher","DOI":"10.1109\/5.880083"},{"key":"S1351324923000128_ref44","unstructured":"Yan, H. , Deng, B. , Li, X. and Qiu, X. (2019). Tener: adapting transformer encoder for named entity recognition. arXiv preprint arXiv:1911.04474."},{"key":"S1351324923000128_ref6","doi-asserted-by":"publisher","DOI":"10.5214\/ans.0972.7531.200408"},{"key":"S1351324923000128_ref48","volume-title":"Human Behavior and the Principle of Least Effort: An Introduction to Human Ecology","author":"Zipf","year":"1949"},{"key":"S1351324923000128_ref30","unstructured":"Rosendahl, J. , Tran, V.A.K. , Wang, W. and Ney, H. (2019). Analysis of positional encodings for neural machine translation. In Proceedings of the 16th International Conference on Spoken Language Translation."},{"key":"S1351324923000128_ref40","doi-asserted-by":"publisher","DOI":"10.1017\/ATSIP.2019.12"},{"key":"S1351324923000128_ref47","volume-title":"The Psychobiology of Language","author":"Zipf","year":"1935"},{"key":"S1351324923000128_ref34","doi-asserted-by":"crossref","unstructured":"Shaw, P. , Uszkoreit, J. and Vaswani, A. (2018). Self-attention with relative position representations. arXiv preprint arXiv:1803.02155.","DOI":"10.18653\/v1\/N18-2074"},{"key":"S1351324923000128_ref32","first-page":"47","article-title":"Long range correlation in human writings","volume":"1","author":"Schenkel","year":"1993","journal-title":"Fractals-an Interdisciplinary Journal on The Complex Geometry of Nature"},{"key":"S1351324923000128_ref37","doi-asserted-by":"publisher","DOI":"10.1109\/ICCP51029.2020.9266140"},{"key":"S1351324923000128_ref10","first-page":"31","volume-title":"LREC","author":"Goldhahn","year":"2012"},{"key":"S1351324923000128_ref2","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.0510673103"},{"key":"S1351324923000128_ref13","doi-asserted-by":"publisher","DOI":"10.2307\/1421449"},{"key":"S1351324923000128_ref15","doi-asserted-by":"crossref","unstructured":"Huang, Z. , Liang, D. , Xu, P. and Xiang, B. (2020). Improve transformer models with better relative position embeddings. arXiv preprint arXiv:2009.13658.","DOI":"10.18653\/v1\/2020.findings-emnlp.298"},{"key":"S1351324923000128_ref17","doi-asserted-by":"crossref","unstructured":"Lakretz, Y. , Kruszewski, G. , Desbordes, T. , Hupkes, D. , Dehaene, S. and Baroni, M. (2019). The emergence of number and syntax units in lstm language models. arXiv preprint arXiv:1903.07435.","DOI":"10.18653\/v1\/N19-1002"},{"key":"S1351324923000128_ref27","doi-asserted-by":"crossref","unstructured":"Peters, M.E. , Neumann, M. , Iyyer, M. , Gardner, M. , Clark, C. , Lee, K. and Zettlemoyer, L. (2018). Deep contextualized word representations. arXiv preprint arXiv:1802.05365.","DOI":"10.18653\/v1\/N18-1202"},{"key":"S1351324923000128_ref1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1117723109"},{"key":"S1351324923000128_ref18","volume-title":"The Handbook of Brain Theory and Neural Networks","author":"LeCun","year":"1995"},{"key":"S1351324923000128_ref35","volume-title":"Advances in Neural Information Processing Systems","author":"Shiv","year":"2019"},{"key":"S1351324923000128_ref7","doi-asserted-by":"publisher","DOI":"10.1209\/0295-5075\/26\/4\/001"},{"key":"S1351324923000128_ref16","volume-title":"Statistical Methods for Speech Recognition","author":"Jelinek","year":"1997"},{"key":"S1351324923000128_ref38","unstructured":"Vaswani, A. , Shazeer, N. , Parmar, N. , Uszkoreit, J. , Jones, L. , Gomez, A. N. , Kaiser, L. and Polosukhin, I. (2017). Attention is all you need. arXiv preprint arXiv:1706.03762."},{"key":"S1351324923000128_ref14","doi-asserted-by":"publisher","DOI":"10.1037\/0096-3445.124.1.62"},{"key":"S1351324923000128_ref11","first-page":"1222","volume-title":"LREC","author":"Guthrie","year":"2006"},{"key":"S1351324923000128_ref19","doi-asserted-by":"crossref","unstructured":"Lin, Y. , Tan, Y.C. and Frank, R. (2019). Open sesame: getting inside bert\u2019s linguistic knowledge. arXiv preprint arXiv:1906.01698.","DOI":"10.18653\/v1\/W19-4825"},{"key":"S1351324923000128_ref46","doi-asserted-by":"crossref","unstructured":"Zhu, J. , Li, J. , Zhu, M. , Qian, L. , Zhang, M. and Zhou, G. (2019). Modeling graph structure in transformer for better amr-to-text generation. arXiv preprint arXiv:1909.00136.","DOI":"10.18653\/v1\/D19-1548"},{"key":"S1351324923000128_ref8","unstructured":"Gehring, J. , Auli, M. , Grangier, D. , Yarats, D. and Dauphin, Y.N. (2017). Convolutional sequence to sequence learning. In International Conference on Machine Learning. PMLR, pp. 1243\u20131252."},{"key":"S1351324923000128_ref3","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1097-4571(199009)41:6<391::AID-ASI1>3.0.CO;2-9"},{"key":"S1351324923000128_ref22","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1907367117"},{"key":"S1351324923000128_ref29","unstructured":"Radford, A. , Narasimhan, K. , Salimans, T. and Sutskever, I. (2018). Improving language understanding by generative pre-training."},{"key":"S1351324923000128_ref36","doi-asserted-by":"crossref","unstructured":"Takase, S. and Okazaki, N. (2019). Positional encoding to control output sequence length. arXiv preprint arXiv:1904.07418.","DOI":"10.18653\/v1\/N19-1401"},{"key":"S1351324923000128_ref39","unstructured":"Wang, B. , Shang, L. , Lioma, C. , Jiang, X. , Yang, H. , Liu, Q. and Simonsen, J.G. (2021). On position embeddings in bert. In International Conference on Learning Representations."},{"key":"S1351324923000128_ref45","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-02130-5"},{"key":"S1351324923000128_ref43","unstructured":"Wei, J. , Ren, X. , Li, X. , Huang, W. , Liao, Y. , Wang, Y. , Lin, J. , Jiang, X. , Chen, X. and Liu, Q. (2019). Nezha: neural contextualized representation for chinese language understanding. arXiv preprint arXiv:1909.00204."},{"key":"S1351324923000128_ref12","first-page":"146","article-title":"Distributional structure","volume":"10","author":"Harris","year":"1954","journal-title":"Word-Journal of The International Linguistic Association"},{"key":"S1351324923000128_ref24","doi-asserted-by":"publisher","DOI":"10.1080\/01638530802356463"},{"key":"S1351324923000128_ref26","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1162"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324923000128","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,4,1]],"date-time":"2024-04-01T13:29:03Z","timestamp":1711978143000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324923000128\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,10]]},"references-count":48,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2024,3]]}},"alternative-id":["S1351324923000128"],"URL":"https:\/\/doi.org\/10.1017\/s1351324923000128","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,10]]},"assertion":[{"value":"\u00a9 The Author(s), 2023. Published by Cambridge University Press","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (http:\/\/creativecommons.org\/licenses\/by\/4.0\/), which permits unrestricted re-use, distribution and reproduction, provided the original article is properly cited.","name":"license","label":"License","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This content has been made available to all.","name":"free","label":"Free to read"}]}}