{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:42:38Z","timestamp":1760132558874},"reference-count":53,"publisher":"MIT Press - Journals","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Transactions of the Association for Computational Linguistics"],"published-print":{"date-parts":[[2019,11]]},"abstract":"<jats:p> Contextual representation models have achieved great success in improving various downstream natural language processing tasks. However, these language-model-based encoders are difficult to train due to their large parameter size and high computational complexity. By carefully examining the training procedure, we observe that the softmax layer, which predicts a distribution of the target word, often induces significant overhead, especially when the vocabulary size is large. Therefore, we revisit the design of the output layer and consider directly predicting the pre-trained embedding of the target word for a given context. When applied to ELMo, the proposed approach achieves a 4-fold speedup and eliminates 80% trainable parameters while achieving competitive performance on downstream tasks. Further analysis shows that the approach maintains the speed advantage under various settings, even when the sentence encoder is scaled up. <\/jats:p>","DOI":"10.1162\/tacl_a_00289","type":"journal-article","created":{"date-parts":[[2019,9,30]],"date-time":"2019-09-30T19:43:11Z","timestamp":1569872591000},"page":"611-624","source":"Crossref","is-referenced-by-count":2,"title":["Efficient Contextual Representation Learning With Continuous Outputs"],"prefix":"10.1162","volume":"7","author":[{"given":"Liunian Harold","family":"Li","sequence":"first","affiliation":[{"name":"Peking University."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Patrick H.","family":"Chen","sequence":"additional","affiliation":[{"name":"University of California, Los Angeles."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cho-Jui","family":"Hsieh","sequence":"additional","affiliation":[{"name":"University of California, Los Angeles."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai-Wei","family":"Chang","sequence":"additional","affiliation":[{"name":"University of California, Los Angeles."}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","reference":[{"key":"bib1","author":"Ba Jimmy Lei","year":"2016","journal-title":"arXiv preprint arXiv:1607.06450"},{"key":"bib2","volume-title":"AISTATS","author":"Bengio Yoshua","year":"2003"},{"key":"bib3","volume-title":"EMNLP","author":"Bowman Samuel R.","year":"2015"},{"key":"bib4","author":"Bradbury James","year":"2016","journal-title":"arXiv preprint arXiv:1611.01576"},{"key":"bib5","author":"Chelba Ciprian","year":"2013","journal-title":"arXiv preprint arXiv:1312.3005"},{"key":"bib6","volume-title":"ACL","author":"Chen Mia Xu","year":"2018"},{"key":"bib7","volume-title":"ACL","author":"Chen Qian","year":"2017"},{"key":"bib8","first-page":"2493","volume":"12","author":"Collobert Ronan","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"bib9","volume-title":"ICML","author":"Dauphin Yann N.","year":"2017"},{"key":"bib10","volume-title":"NAACL-NLT","author":"Devlin Jacob","year":"2019"},{"key":"bib11","first-page":"2121","volume":"12","author":"Duchi John","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"bib12","author":"Gardner Matt","year":"2018","journal-title":"arXiv preprint arXiv:1803.07640"},{"key":"bib13","author":"Goyal Priya","year":"2017","journal-title":"arXiv preprint arXiv:1706.02677"},{"key":"bib14","author":"Grave Edouard","year":"2016","journal-title":"arXiv preprint arXiv:1609.04309"},{"key":"bib15","volume-title":"ACL","author":"He Luheng","year":"2017"},{"key":"bib16","volume-title":"ACL","author":"He Shexia","year":"2018"},{"key":"bib17","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"bib18","author":"Jernite Yacine","year":"2017","journal-title":"arXiv preprint arXiv:1705.00557"},{"key":"bib19","volume-title":"International Encyclopedia of Statistical Science","author":"Jolliffe Ian","year":"2011"},{"key":"bib20","author":"Jozefowicz Rafal","year":"2016","journal-title":"arXiv preprint arXiv:1602.02410"},{"key":"bib21","volume-title":"AAAI","author":"Kim Yoon","year":"2016"},{"key":"bib22","volume-title":"NIPS","author":"Kiros Ryan","year":"2015"},{"key":"bib23","author":"Kitaev Nikita","year":"2018","journal-title":"arXiv preprint arXiv:1812.11760"},{"key":"bib24","volume-title":"ICLR","author":"Kumar Sachin","year":"2019"},{"key":"bib25","volume-title":"EMNLP","author":"Lee Kenton","year":"2017"},{"key":"bib26","volume-title":"NAACL-HLT","author":"Lee Kenton","year":"2018"},{"key":"bib27","volume-title":"EMNLP","author":"Lei Tao","year":"2018"},{"key":"bib28","volume-title":"NIPS","author":"Levy Omer","year":"2014"},{"key":"bib29","author":"Logeswaran Lajanugen","year":"2018","journal-title":"ICLR"},{"key":"bib30","volume-title":"NIPS","author":"McCann Bryan","year":"2017"},{"key":"bib31","author":"Merity Stephen","year":"2018","journal-title":"arXiv preprint arXiv:1803.08240"},{"key":"bib32","author":"Mikolov Tomas","year":"2013","journal-title":"arXiv preprint arXiv:1301.3781"},{"key":"bib33","author":"Mikolov Tomas","year":"2017","journal-title":"arXiv preprint arXiv:1712.09405"},{"key":"bib34","volume-title":"ICML","author":"Mnih A","year":"2012"},{"key":"bib35","volume-title":"AISTATS","author":"Morin Frederic","year":"2005"},{"key":"bib36","author":"Panchenko Alexander","year":"2017","journal-title":"arXiv preprint arXiv:1710.01779"},{"key":"bib37","volume-title":"NAACL-HLT","author":"Peters Matthew E.","year":"2018"},{"key":"bib38","volume-title":"EMNLP","author":"Peters Matthew E.","year":"2018"},{"key":"bib39","volume-title":"EMNLP","author":"Pinter Yuval","year":"2017"},{"key":"bib40","volume-title":"CoNLL","author":"Pradhan Sameer","year":"2013"},{"key":"bib41","volume-title":"Joint Conference on EMNLP and CoNLL-Shared Task","author":"Pradhan Sameer","year":"2012"},{"key":"bib42","author":"Radford Alec","year":"2018","journal-title":"OpenAI Blog"},{"key":"bib43","author":"Radford Alec","year":"2019","journal-title":"OpenAI Blog"},{"key":"bib44","volume-title":"EMNLP","author":"Rajpurkar Pranav","year":"2016"},{"key":"bib45","author":"Sang Erik F","year":"2003","journal-title":"arXiv preprint cs\/0306050"},{"key":"bib46","volume-title":"ACL","author":"Sennrich Rico","year":"2016"},{"key":"bib47","volume-title":"EMNLP","author":"Socher Richard","year":"2013"},{"key":"bib48","author":"Strubell Emma","year":"2019","journal-title":"arXiv preprint arXiv:1906.02243"},{"key":"bib49","volume-title":"Workshop on Representation Learning for NLP","author":"Tang Shuai","year":"2018"},{"key":"bib50","volume-title":"NIPS","author":"Vaswani Ashish","year":"2017"},{"key":"bib51","author":"Yonghui Wu","year":"2016","journal-title":"arXiv preprint arXiv:1609.08144"},{"key":"bib52","author":"Yang Zhilin","year":"2017","journal-title":"arXiv preprint arXiv:1711.03953"},{"key":"bib53","volume-title":"Proceedings of the 47th International Conference on Parallel Processing, ICPP 2018","author":"You Yang","year":"2018"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mitpressjournals.org\/doi\/pdf\/10.1162\/tacl_a_00289","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,3,12]],"date-time":"2021-03-12T21:39:29Z","timestamp":1615585169000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/43525"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,11]]},"references-count":53,"alternative-id":["10.1162\/tacl_a_00289"],"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00289","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,11]]}}}