{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,26]],"date-time":"2025-10-26T15:08:25Z","timestamp":1761491305599,"version":"build-2065373602"},"reference-count":39,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T00:00:00Z","timestamp":1635379200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>In the last few decades, text mining has been used to extract knowledge from free texts. Applying neural networks and deep learning to natural language processing (NLP) tasks has led to many accomplishments for real-world language problems over the years. The developments of the last five years have resulted in techniques that have allowed for the practical application of transfer learning in NLP. The advances in the field have been substantial, and the milestone of outperforming human baseline performance based on the general language understanding evaluation has been achieved. This paper implements a targeted literature review to outline, describe, explain, and put into context the crucial techniques that helped achieve this milestone. The research presented here is a targeted review of neural language models that present vital steps towards a general language representation model.<\/jats:p>","DOI":"10.3390\/e23111422","type":"journal-article","created":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T23:50:28Z","timestamp":1635465028000},"page":"1422","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Language Representation Models: An Overview"],"prefix":"10.3390","volume":"23","author":[{"given":"Thorben","family":"Schomacker","sequence":"first","affiliation":[{"name":"Department of Computer Science, Hamburg University of Applied Sciences, 20099 Hamburg, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1623-5309","authenticated-orcid":false,"given":"Marina","family":"Tropmann-Frick","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Hamburg University of Applied Sciences, 20099 Hamburg, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S.R. (2018). GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv.","DOI":"10.18653\/v1\/W18-5446"},{"key":"ref_2","unstructured":"Jing, K., and Xu, J. (2019). A Survey on Neural Network Language Models. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1872","DOI":"10.1007\/s11431-020-1647-3","article-title":"Pre-Trained Models for Natural Language Processing: A Survey","volume":"63","author":"Qiu","year":"2020","journal-title":"Sci. China Technol. Sci."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Babi\u0107, K., Martin\u010di\u0107-Ip\u0161i\u0107, S., and Me\u0161trovi\u0107, A. (2020). Survey of Neural Text Representation Models. Information, 11.","DOI":"10.3390\/info11110511"},{"key":"ref_5","first-page":"74:1","article-title":"A Comprehensive Survey on Word Representation Models: From Classical to State-of-the-Art Word Representation Language Models","volume":"20","author":"Naseem","year":"2020","journal-title":"Trans. Asian Low-Resour. Lang. Inf. Process."},{"key":"ref_6","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2016). Neural Machine Translation by Jointly Learning to Align and Translate. arXiv."},{"key":"ref_7","unstructured":"Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (2017). Attention Is All You Need. Advances in Neural Information Processing Systems 30, Curran Associates, Inc."},{"key":"ref_8","unstructured":"He, P., Liu, X., Gao, J., and Chen, W. (2020). DeBERTa: Decoding-Enhanced BERT with Disentangled Attention. arXiv."},{"key":"ref_9","unstructured":"(2021, January 02). GLUE Benchmark. Available online: https:\/\/gluebenchmark.com\/."},{"key":"ref_10","unstructured":"Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2020). ALBERT: A Lite BERT for Self-Supervised Learning of Language Representations. arXiv."},{"key":"ref_11","unstructured":"Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P.J. (2020). Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer. arXiv."},{"key":"ref_12","unstructured":"Clark, K., Luong, M.T., Le, Q.V., and Manning, C.D. (2020). ELECTRA: Pre-Training Text Encoders as Discriminators Rather Than Generators. arXiv."},{"key":"ref_13","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019). BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Pennington, J., Socher, R., and Manning, C. (2014, January 25\u201329). Glove: Global Vectors for Word Representation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Stanford, CA, USA.","DOI":"10.3115\/v1\/D14-1162"},{"key":"ref_15","first-page":"3111","article-title":"Distributed Representations of Words and Phrases and Their Compositionality","volume":"Volume 2","author":"Mikolov","year":"2013","journal-title":"Proceedings of the 26th International Conference on Neural Information Processing Systems, NIPS\u201913"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Rumelhart, D.E., Hinton, G.E., and Williams, R.J. (1986). Learning Representations by Back-Propagating Errors. Neurocomputing: Foundations of Research, MIT Press 55 Hayward St.","DOI":"10.1038\/323533a0"},{"key":"ref_17","unstructured":"Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep Learning, The MIT Press. Adaptive Computation and Machine Learning."},{"key":"ref_18","unstructured":"Olah, C. (2021, January 02). Understanding LSTM Networks. Available online: https:\/\/research.google\/pubs\/pub45500\/."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merrienboer, B., Bahdanau, D., and Bengio, Y. (2014). On the Properties of Neural Machine Translation: Encoder-Decoder Approaches. arXiv.","DOI":"10.3115\/v1\/W14-4012"},{"key":"ref_20","unstructured":"Sutskever, I., Vinyals, O., and Le, Q.V. (2014). Sequence to Sequence Learning with Neural Networks. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning Phrase Representations Using RNN Encoder-Decoder for Statistical Machine Translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Li, J., Luong, M.T., and Jurafsky, D. (2015). A Hierarchical Neural Autoencoder for Paragraphs and Documents. arXiv.","DOI":"10.3115\/v1\/P15-1107"},{"key":"ref_23","unstructured":"Alammar, J. (2021, January 02). The Illustrated Transformer. Available online: https:\/\/jalammar.github.io\/illustrated-transformer\/."},{"key":"ref_24","unstructured":"Huang, C.Z.A., Vaswani, A., Uszkoreit, J., Shazeer, N., Simon, I., Hawthorne, C., Dai, A.M., Hoffman, M.D., Dinculescu, M., and Eck, D. (2018). Music Transformer. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Shaw, P., Uszkoreit, J., and Vaswani, A. (2018). Self-Attention with Relative Position Representations. arXiv.","DOI":"10.18653\/v1\/N18-2074"},{"key":"ref_26","unstructured":"Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2021, January 02). Improving Language Understanding by Generative Pre-Training. Available online: https:\/\/cdn.openai.com\/research-covers\/language-unsupervised\/language_understanding_paper.pdf."},{"key":"ref_27","unstructured":"Yang, Z., Dai, Z., Yang, Y., Carbonell, J., Salakhutdinov, R., and Le, Q.V. (2020). XLNet: Generalized Autoregressive Pretraining for Language Understanding. arXiv."},{"key":"ref_28","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv."},{"key":"ref_29","unstructured":"He, P., Liu, X., Gao, J., and Chen, W. (2021, January 02). Microsoft DeBERTa Surpasses Human Performance on SuperGLUE Benchmark. Available online: https:\/\/www.microsoft.com\/en-us\/research\/blog\/microsoft-deberta-surpasses-human-performance-on-the-superglue-benchmark."},{"key":"ref_30","unstructured":"You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., and Hsieh, C.J. (2014). Large Batch Optimization for Deep Learning: Training BERT in 76 Minutes. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Liu, X., He, P., Chen, W., and Gao, J. (2019, January 28). Multi-Task Deep Neural Networks for Natural Language Understanding. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy.","DOI":"10.18653\/v1\/P19-1441"},{"key":"ref_32","unstructured":"McCann, B., Keskar, N.S., Xiong, C., and Socher, R. (2018). The Natural Language Decathlon: Multitask Learning as Question Answering. arXiv."},{"key":"ref_33","unstructured":"Liu, Y. (2019). Fine-Tune BERT for Extractive Summarization. arXiv."},{"key":"ref_34","unstructured":"Schomacker, T., Tropmann-Frick, M., and Zukunft, O. Application of Transformer-Based Methods to Latin Text Analysis."},{"key":"ref_35","first-page":"24","article-title":"Language Models Are Unsupervised Multitask Learners","volume":"9","author":"Radford","year":"2019","journal-title":"Open AI Blog"},{"key":"ref_36","unstructured":"Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A. (2020). Language Models Are Few-Shot Learners. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Gao, T., Fisch, A., and Chen, D. (2021, January 1\u20136). Making pre-trained language models better few-shot learners. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Stroudsburg, PA, USA.","DOI":"10.18653\/v1\/2021.acl-long.295"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L. (2019). BART: Denoising Sequence-to-Sequence Pre-Training for Natural Language Generation, Translation, and Comprehension. arXiv.","DOI":"10.18653\/v1\/2020.acl-main.703"},{"key":"ref_39","unstructured":"Ziegler, D.M., Stiennon, N., Wu, J., Brown, T.B., Radford, A., Amodei, D., Christiano, P., and Irving, G. (2020). Fine-Tuning Language Models from Human Preferences. arXiv."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/11\/1422\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:22:03Z","timestamp":1760167323000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/11\/1422"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,28]]},"references-count":39,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2021,11]]}},"alternative-id":["e23111422"],"URL":"https:\/\/doi.org\/10.3390\/e23111422","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2021,10,28]]}}}