{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,21]],"date-time":"2026-03-21T07:37:57Z","timestamp":1774078677926,"version":"3.50.1"},"reference-count":50,"publisher":"Association for Computing Machinery (ACM)","issue":"3","funder":[{"name":"Information Technology Research Center","award":["IITP-2022-RS-2022-00156354"],"award-info":[{"award-number":["IITP-2022-RS-2022-00156354"]}]},{"name":"Korean Government Ministry of Science and Information Technology","award":["RS-2019-II190231"],"award-info":[{"award-number":["RS-2019-II190231"]}]},{"name":"Institute of Information & Communications Technology Planning & Evaluation"},{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"crossref","award":["2020R1A6A1A03038540"],"award-info":[{"award-number":["2020R1A6A1A03038540"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Training language models from scratch presents a critical challenge in Natural Language Processing (NLP), primarily due to the computational demands of pre-trained Large Language Models, which are predominantly trained on English corpora using extensive resources. While offering viable solutions, existing alternatives still rely heavily on high-performance hardware. This work introduces a different approach to reducing the algorithmic complexity of Transformer-based architectures through the U-Net Encapsulated Transformer (UET), which applies dimensionality reduction to token embeddings. The UET architecture enables the development of language models with significantly reduced parameter sizes for a given set of hyperparameters. Alternatively, it allows researchers to design models of comparable size but with a substantially greater number of Transformer blocks, enhancing model depth and potential capacity. This study also outlines practical methodologies for training language models in resource-constrained environments. Experimental results illustrate the potential of the UET architecture in achieving reasonable performance under resource-constrained conditions, highlighting its promise as an accessible alternative for language model development. This work could broaden the accessibility of NLP research, empowering researchers with hardware constraints to contribute to language model development.<\/jats:p>","DOI":"10.1145\/3735653","type":"journal-article","created":{"date-parts":[[2025,5,13]],"date-time":"2025-05-13T13:09:05Z","timestamp":1747141745000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["U-Net Encapsulated Transformer for Reducing Dimensionality in Training Large Language Models"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2145-934X","authenticated-orcid":false,"given":"Marvin John","family":"Ignacio","sequence":"first","affiliation":[{"name":"Department of Computer Engineering, Sejong University, Seoul, the Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4645-1395","authenticated-orcid":false,"given":"Yong-Guk","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Sejong University, Seoul, the Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6896-2498","authenticated-orcid":false,"given":"Hulin","family":"Jin","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Anhui University, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-6249-9156","authenticated-orcid":false,"given":"Seunghee","family":"Yu","sequence":"additional","affiliation":[{"name":"Department of Business Administration, Sejong University, Seoul, the Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,3,20]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"TinyShakespeare (Shakespeare\u2019s Plays). 2024. Retrieved from https:\/\/www.kaggle.com\/datasets\/thedevastator\/the-bards-best-a-character-modeling-dataset"},{"key":"e_1_3_1_3_2","unstructured":"Google AI. 2024. Gemini\u2014Chat to Supercharge Your Ideas. Retrieved from https:\/\/gemini.google.com\/"},{"key":"e_1_3_1_4_2","doi-asserted-by":"crossref","unstructured":"Joshua Ainslie James Lee-Thorp Michiel de Jong Yury Zemlyanskiy Federico Lebr\u2019on and Sumit K. Sanghai. 2023. GQA: Training generalized multi-query transformer models from multi-head checkpoints. arXiv:2305.13245. Retrieved from https:\/\/arxiv.org\/abs\/2305.13245","DOI":"10.18653\/v1\/2023.emnlp-main.298"},{"key":"e_1_3_1_5_2","unstructured":"Iz Beltagy Matthew E. Peters and Arman Cohan. 2020. Longformer: The long-document transformer. arXiv:2004.05150. Retrieved from https:\/\/arxiv.org\/abs\/2004.05150"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6239"},{"key":"e_1_3_1_7_2","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners. arXiv:2005.14165. Retrieved from https:\/\/arxiv.org\/abs\/2005.14165","journal-title":"Language models are few-shot learners"},{"key":"e_1_3_1_8_2","article-title":"Learning phrase representations using RNN encoder\u2013decoder for statistical machine translation","author":"Cho Kyunghyun","year":"2014","unstructured":"Kyunghyun Cho, Bart van Merrienboer, \u00c7aglar G\u00fcl\u00e7ehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014. Learning phrase representations using RNN encoder\u2013decoder for statistical machine translation. In Conference on Empirical Methods in Natural Language Processing.","journal-title":"In Conference on Empirical Methods in Natural Language Processing"},{"key":"e_1_3_1_9_2","author":"Chowdhery Aakanksha","year":"2022","unstructured":"Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022. PaLM: Scaling language modeling with pathways. arXiv:2204.02311. Retrieved from https:\/\/arxiv.org\/abs\/2204.02311","journal-title":"PaLM: Scaling language modeling with pathways"},{"key":"e_1_3_1_10_2","author":"Clark Peter","year":"2018","unstructured":"Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv:1803.05457. Retrieved from https:\/\/arxiv.org\/abs\/1803.05457","journal-title":"Think you have solved question answering? Try ARC, the AI2 reasoning challenge"},{"key":"e_1_3_1_11_2","volume-title":"RedPajama: An Open Dataset for Training Large Language Models","author":"Together Computer.","year":"2023","unstructured":"Together Computer. 2023. RedPajama: An Open Dataset for Training Large Language Models. Retrieved from https:\/\/github.com\/togethercomputer\/RedPajama-Data"},{"key":"e_1_3_1_12_2","volume-title":"North American Chapter of the Association for Computational Linguistics","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In North American Chapter of the Association for Computational Linguistics."},{"key":"e_1_3_1_13_2","unstructured":"Leo Gao Jonathan Tow Baber Abbasi Stella Biderman Sid Black Anthony DiPofi Charles Foster Laurence Golding Jeffrey Hsu Alain Le Noac\u2019h et al. 2024. The Language Model Evaluation Harness. Zenodo. Retrieved from https:\/\/zenodo.org\/records\/12608602"},{"key":"e_1_3_1_14_2","unstructured":"Albert Gu and Tri Dao. 2023. Mamba: Linear-time sequence modeling with selective state spaces. arXiv:2312.00752. Retrieved from https:\/\/arxiv.org\/abs\/2312.00752"},{"key":"e_1_3_1_15_2","author":"Hendrycks Dan","year":"2016","unstructured":"Dan Hendrycks and Kevin Gimpel. 2016. Gaussian error linear units (GELUs). arXiv:1606.08415. Retrieved from https:\/\/arxiv.org\/abs\/1606.08415","journal-title":"Gaussian error linear units (GELUs)"},{"key":"e_1_3_1_16_2","author":"Hinton Geoffrey","year":"2015","unstructured":"Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv:1503.02531. Retrieved from https:\/\/arxiv.org\/abs\/1503.02531","journal-title":"Distilling the knowledge in a neural network"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_18_2","volume-title":"International Conference on Learning Representations","author":"Holtzman Ari","year":"2020","unstructured":"Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations."},{"key":"e_1_3_1_19_2","author":"Hu Edward J.","year":"2021","unstructured":"Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. LoRA: Low-rank adaptation of large language models. arXiv:2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685","journal-title":"LoRA: Low-rank adaptation of large language models"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3424883"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_3_1_22_2","author":"Jiang Albert Qiaochu","year":"2023","unstructured":"Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. arXiv:2310.06825. Retrieved from https:\/\/arxiv.org\/abs\/2310.06825","journal-title":"Mistral 7B"},{"key":"e_1_3_1_23_2","unstructured":"Jared Kaplan Sam McCandlish Tom Henighan Tom B. Brown Benjamin Chess Rewon Child Scott Gray Alec Radford Jeff Wu and Dario Amodei. 2020. Scaling laws for neural language models. arXiv:2001.08361. Retrieved from https:\/\/arxiv.org\/abs\/2001.08361"},{"key":"e_1_3_1_24_2","unstructured":"Andrej Karpathy. 2024. nanoGPT. Retrieved from https:\/\/github.com\/karpathy\/nanoGPT"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1082"},{"key":"e_1_3_1_26_2","unstructured":"Zhenzhong Lan Mingda Chen Sebastian Goodman Kevin Gimpel Piyush Sharma and Radu Soricut. 2019. ALBERT: A lite BERT for self-supervised learning of language representations. arXiv:1909.11942. Retrieved from https:\/\/arxiv.org\/abs\/1909.11942"},{"key":"e_1_3_1_27_2","unstructured":"Lightning-AI. 2024. Lit-LLaMA. Retrieved from https:\/\/github.com\/Lightning-AI\/lit-llama"},{"key":"e_1_3_1_28_2","unstructured":"Zechun Liu Changsheng Zhao Forrest N. Iandola Chen Lai Yuandong Tian Igor Fedorov Yunyang Xiong Ernie Chang Yangyang Shi Raghuraman Krishnamoorthi et al. 2024. MobileLLM: Optimizing sub-billion parameter language models for on-device use cases. arXiv:2402.14905. Retrieved from https:\/\/arxiv.org\/abs\/2402.14905"},{"key":"e_1_3_1_29_2","unstructured":"Stephen Merity Caiming Xiong James Bradbury and Richard Socher. 2016. Pointer sentinel mixture models. arXiv:1609.07843. Retrieved from https:\/\/arxiv.org\/abs\/1609.07843"},{"key":"e_1_3_1_30_2","article-title":"Can a suit of armor conduct electricity? A new dataset for open book question answering","author":"Mihaylov Todor","year":"2018","unstructured":"Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? A new dataset for open book question answering. In Conference on Empirical Methods in Natural Language Processing (EMNLP).","journal-title":"Conference on Empirical Methods in Natural Language Processing (EMNLP)"},{"key":"e_1_3_1_31_2","unstructured":"OpenAI. 2024. ChatGPT: GPT-4 Language Model. Retrieved from https:\/\/chat.openai.com\/"},{"key":"e_1_3_1_32_2","unstructured":"OpenAI. 2024. GPT-4o. Retrieved from https:\/\/platform.openai.com\/docs\/models\/gpt-4o"},{"key":"e_1_3_1_33_2","unstructured":"Long Ouyang Jeff Wu Xu Jiang Diogo Almeida Carroll L. Wainwright Pamela Mishkin Chong Zhang Sandhini Agarwal Katarina Slama Alex Ray et al. 2022. Training language models to follow instructions with human feedback. arXiv:2203.02155. Retrieved from https:\/\/arxiv.org\/abs\/2203.02155"},{"key":"e_1_3_1_34_2","unstructured":"Alec Radford Karthik Narasimhan Tim Salimans and Ilya Sutskever. 2018. Improving Language Understanding by Generative Pre-Training. Retrieved from https:\/\/openai.com\/blog\/language-unsupervised\/"},{"key":"e_1_3_1_35_2","unstructured":"Prajit Ramachandran Barret Zoph and Quoc V. Le. 2018. Searching for activation functions. arXiv:1710.05941. Retrieved from https:\/\/arxiv.org\/abs\/1710.05941"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_3_1_37_2","first-page":"318","article-title":"Learning internal representations by error propagation","author":"Rumelhart David E.","year":"1986","unstructured":"David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams. 1986. Learning internal representations by error propagation. In Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol. 1: Foundations, MIT Press, 318\u2013362.","journal-title":"In Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Vol. 1: Foundations"},{"key":"e_1_3_1_38_2","author":"Sakaguchi Keisuke","year":"2019","unstructured":"Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2019. WinoGrande: An adversarial WinoGrad schema challenge at scale. arXiv:1907.10641. Retrieved from https:\/\/arxiv.org\/abs\/1907.10641","journal-title":"WinoGrande: An adversarial WinoGrad schema challenge at scale"},{"key":"e_1_3_1_39_2","author":"Sanh Victor","year":"2019","unstructured":"Victor Sanh, L. Debut, J. Chaumond, and T. Wolf. 2019. DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter. arXiv:1910.01108. Retrieved from https:\/\/arxiv.org\/abs\/1910.01108","journal-title":"DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1454"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.semeval-1.99"},{"key":"e_1_3_1_42_2","unstructured":"Noam M. Shazeer. 2020. GLU variants improve transformer. arXiv:2002.05202. Retrieved from https:\/\/arxiv.org\/abs\/2002.05202"},{"key":"e_1_3_1_43_2","unstructured":"Daria Soboleva Faisal Al-Khateeb Robert Myers Jacob R. Steeves Joel Hestness and Nolan Dey. 2023. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. Retrieved from https:\/\/cerebras.ai\/blog\/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama"},{"key":"e_1_3_1_44_2","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar et al. 2023. LLaMA: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/arxiv.org\/abs\/2302.13971"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1162\/glep_a_00747"},{"key":"e_1_3_1_46_2","article-title":"Attention is all you need. In","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.124781"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1016\/0169-7439(87)80084-9"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.542"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/EMC2-NIPS53020.2019.00016"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.11"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3735653","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,21]],"date-time":"2026-03-21T04:48:45Z","timestamp":1774068525000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3735653"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,20]]},"references-count":50,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3735653"],"URL":"https:\/\/doi.org\/10.1145\/3735653","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,20]]},"assertion":[{"value":"2024-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-04-20","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-20","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}