{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T03:02:31Z","timestamp":1760151751509,"version":"build-2065373602"},"reference-count":19,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2022,4,15]],"date-time":"2022-04-15T00:00:00Z","timestamp":1649980800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Large-scale pre-trained language representation and its promising performance in various downstream applications have become an area of interest in the field of natural language processing (NLP). There has been huge interest in further increasing the model\u2019s size in order to outperform the best previously obtained performances. However, at some point, increasing the model\u2019s parameters may lead to reaching its saturation point due to the limited capacity of GPU\/TPU. In addition to this, such models are mostly available in English or a shared multilingual structure. Hence, in this paper, we propose a lite BERT trained on a large corpus solely in the Romanian language, which we called \u201cA Lite Romanian BERT (ALR-BERT)\u201d. Based on comprehensive empirical results, ALR-BERT produces models that scale far better than the original Romanian BERT. Alongside presenting the performance on downstream tasks, we detail the analysis of the training process and its parameters. We also intend to distribute our code and model as an open source together with the downstream task.<\/jats:p>","DOI":"10.3390\/computers11040057","type":"journal-article","created":{"date-parts":[[2022,4,16]],"date-time":"2022-04-16T07:42:41Z","timestamp":1650094961000},"page":"57","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["A Lite Romanian BERT: ALR-BERT"],"prefix":"10.3390","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9377-5654","authenticated-orcid":false,"given":"Drago\u015f Constantin","family":"Nicolae","sequence":"first","affiliation":[{"name":"Research Institute for Artificial Intelligence, Romanian Academy, 050711 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1485-0439","authenticated-orcid":false,"given":"Rohan Kumar","family":"Yadav","sequence":"additional","affiliation":[{"name":"Department of Information and Communication, University of Agder, 4604 Grimstad, Norway"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dan","family":"Tufi\u015f","sequence":"additional","affiliation":[{"name":"Research Institute for Artificial Intelligence, Romanian Academy, 050711 Bucharest, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,4,15]]},"reference":[{"key":"ref_1","unstructured":"Sutskever, I., Vinyals, O., and Le, Q.V. (2014, January 8\u201313). Sequence to Sequence Learning with Neural Networks. Proceedings of the NIPS 2014, Montreal, QC, Canada."},{"key":"ref_2","unstructured":"Vaswani, A., Shazeer, N.M., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., and Polosukhin, I. (2017). Attention is All you Need. arXiv."},{"key":"ref_3","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 2\u20137). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the NAACL 2019, Minneapolis, MN, USA."},{"key":"ref_4","unstructured":"Lan, Z., Chen, M., Goodman, S., Gimpel, K., Sharma, P., and Soricut, R. (2020). ALBERT: A Lite BERT for Self-supervised Learning of Language Representations. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. (2018). GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, Association for Computational Linguistics.","DOI":"10.18653\/v1\/W18-5446"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Williams, A., Nangia, N., and Bowman, S.R. (2018). A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference, NAACL.","DOI":"10.18653\/v1\/N18-1101"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P. (2016, January 1\u20135). SQuAD: 100,000+ Questions for Machine Comprehension of Text. Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, Austin, TX, USA.","DOI":"10.18653\/v1\/D16-1264"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Rajpurkar, P., Jia, R., and Liang, P. (2018). Know What You Don\u2019t Know: Unanswerable Questions for SQuAD. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Association for Computational Linguistics.","DOI":"10.18653\/v1\/P18-2124"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Tjong Kim Sang, E.F., and De Meulder, F. (June, January 31). Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition. Proceedings of the Seventh Conference on Natural Language Learning at HLT\u2014-NAACL, Edmonton, AL, Canada.","DOI":"10.3115\/1119176.1119195"},{"key":"ref_10","unstructured":"Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019). RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1162\/tacl_a_00300","article-title":"SpanBERT: Improving Pre-training by Representing and Predicting Spans","volume":"8","author":"Joshi","year":"2020","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_12","unstructured":"Cui, Y., Che, W., Liu, T., Qin, B., Yang, Z., Wang, S., and Hu, G. (2019). Pre-Training with Whole Word Masking for Chinese BERT. arXiv."},{"key":"ref_13","unstructured":"Yang, Z., Dai, Z., Yang, Y., Carbonell, J.G., Salakhutdinov, R., and Le, Q.V. (2019, January 8\u201314). XLNet: Generalized Autoregressive Pretraining for Language Understanding. Proceedings of the NeurIPS 2019, Vancouver, BC, Canada."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Dumitrescu, S., Avram, A.M., and Pyysalo, S. (2020, January 5\u201310). The birth of Romanian BERT. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020, Online.","DOI":"10.18653\/v1\/2020.findings-emnlp.387"},{"key":"ref_15","unstructured":"Tiedemann, J. (2012, January 21\u201327). Parallel Data, Tools and Interfaces in OPUS. Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC\u201912), Istanbul, Turkey."},{"key":"ref_16","unstructured":"Suarez, P.J.O., Sagot, B., and Romary, L. (2019, January 22). Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures. Proceedings of the Workshop on Challenges in the Management of Large Corpora (CMLC-7) 2019, Cardiff, UK."},{"key":"ref_17","unstructured":"Hendrycks, D., and Gimpel, K. (2016). Bridging Nonlinearities and Stochastic Regularizers with Gaussian Error Linear Units. arXiv."},{"key":"ref_18","unstructured":"Mititelu, V.B., Ionn, R., Simionescu, R., Irimia, E., and Perez, C.A. (October, January 29). The romanian treebank annotated according to universal dependencies. Proceedings of the Tenth International Conference on Natural Language Processing (HRTAL\u201916), Dubrovnik, Croatia."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zeman, D., Hajic, J., Popel, M., Potthast, M., Straka, M., Ginter, F., Nivre, J., and Petrov, S. (November, January 31). CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. Proceedings of the CoNLL 2018, Brussels, Belgium.","DOI":"10.18653\/v1\/K17-3001"}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/11\/4\/57\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:54:55Z","timestamp":1760136895000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/11\/4\/57"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,15]]},"references-count":19,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,4]]}},"alternative-id":["computers11040057"],"URL":"https:\/\/doi.org\/10.3390\/computers11040057","relation":{},"ISSN":["2073-431X"],"issn-type":[{"type":"electronic","value":"2073-431X"}],"subject":[],"published":{"date-parts":[[2022,4,15]]}}}