{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T04:24:17Z","timestamp":1784003057462,"version":"3.55.0"},"reference-count":232,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T00:00:00Z","timestamp":1775433600000},"content-version":"vor","delay-in-days":95,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Fine-tuning large language models (LLMs) with limited data poses a practical challenge in low-resource languages, specialized domains, and constrained deployment settings. While pre-trained LLMs provide strong foundations, effective adaptation under data scarcity requires focused and efficient fine-tuning techniques. This paper presents a structured and practical survey of recent methods for fine-tuning LLMs in data-scarce scenarios. We systematically review parameter-efficient fine-tuning techniques that lower training and deployment costs, domain and cross-lingual adaptation methods for both encoder and decoder models, and model specialization strategies. We further examine preference alignment approaches that guide model behavior using limited human or synthetic feedback, emphasizing sample and compute efficiency. Throughout, we highlight empirical trade-offs, selection criteria, and best practices for choosing suitable techniques based on task constraints, including model scaling, data scaling, and the mitigation of catastrophic forgetting. The aim is to equip researchers and practitioners with actionable insights for effectively fine-tuning LLMs when data and resources are limited.<\/jats:p>","DOI":"10.1162\/tacl.a.627","type":"journal-article","created":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T16:50:03Z","timestamp":1775494203000},"page":"341-377","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":1,"title":["Fine-tuning Large Language Models with Limited Data: A Survey and Practical Guide"],"prefix":"10.1162","volume":"14","author":[{"given":"Marton","family":"Szep","sequence":"first","affiliation":[{"name":"AI in Healthcare and Medicine, Technical University of Munich (TUM) and TUM University Hospital, Munich, Germany. marton.szep@tum.de"},{"name":"Department of Orthopaedics and Sports Orthopaedics, TUM University Hospital, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Rueckert","sequence":"additional","affiliation":[{"name":"AI in Healthcare and Medicine, Technical University of Munich (TUM) and TUM University Hospital, Munich, Germany"},{"name":"Munich Center for Machine Learning (MCML), Munich, Germany"},{"name":"Department of Computing, Imperial College London, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"R\u00fcdiger","family":"von Eisenhart-Rothe","sequence":"additional","affiliation":[{"name":"Department of Orthopaedics and Sports Orthopaedics, TUM University Hospital, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Florian","family":"Hinterwimmer","sequence":"additional","affiliation":[{"name":"AI in Healthcare and Medicine, Technical University of Munich (TUM) and TUM University Hospital, Munich, Germany"},{"name":"Department of Orthopaedics and Sports Orthopaedics, TUM University Hospital, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2026,4,1]]},"reference":[{"key":"2026040612495722900_bib1","doi-asserted-by":"publisher","first-page":"7319","DOI":"10.18653\/v1\/2021.acl-long.568","article-title":"Intrinsic dimensionality explains the effectiveness of language model fine-tuning","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Aghajanyan","year":"2021"},{"issue":"9","key":"2026040612495722900_bib2","doi-asserted-by":"publisher","first-page":"224:1\u2013224:23","DOI":"10.1145\/3610289","article-title":"WAD-X: Improving zero-shot cross-lingual transfer via adapter-based word alignment","volume":"22","author":"Ahmat","year":"2023","journal-title":"ACM Transactions on Asian and Low-Resource Language Information Processing"},{"key":"2026040612495722900_bib3","first-page":"2754","article-title":"Massive vs. curated embeddings for low-resourced languages: The case of Yor\u00f9b\u00e1 and Twi","volume-title":"Proceedings of the Twelfth Language Resources and Evaluation Conference","author":"Alabi","year":"2020"},{"key":"2026040612495722900_bib4","first-page":"4336","article-title":"Adapting pre-trained language models to african languages via multilingual adaptive fine-tuning","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Alabi","year":"2022"},{"key":"2026040612495722900_bib5","doi-asserted-by":"publisher","first-page":"1778","DOI":"10.18653\/v1\/2022.acl-long.125","article-title":"Composable sparse fine-tuning for cross-lingual transfer","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Ansell","year":"2022"},{"key":"2026040612495722900_bib6","unstructured":"Argilla. 2024. Preference datasets for DPO - an Argilla collection."},{"key":"2026040612495722900_bib7","unstructured":"Argilla and MantisNLP. 2024. RLHF and Alternatives: Overview."},{"key":"2026040612495722900_bib8","article-title":"ExT5: Towards extreme multi-task scaling for transfer learning","volume-title":"International Conference on Learning Representations","author":"Aribandi","year":"2021"},{"key":"2026040612495722900_bib9","first-page":"3056","article-title":"On the multilingual capabilities of very large-scale english language models","volume-title":"Proceedings of the Thirteenth Language Resources and Evaluation Conference","author":"Armengol-Estap\u00e9","year":"2022"},{"key":"2026040612495722900_bib10","doi-asserted-by":"publisher","first-page":"4623","DOI":"10.18653\/v1\/2020.acl-main.421","article-title":"On the cross-lingual transferability of monolingual representations","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Artetxe","year":"2020"},{"key":"2026040612495722900_bib11","first-page":"4447","article-title":"A general theoretical paradigm to understand learning from human preferences","volume-title":"Proceedings of The 27th International Conference on Artificial Intelligence and Statistics","author":"Azar","year":"2024"},{"key":"2026040612495722900_bib12","doi-asserted-by":"publisher","first-page":"5002","DOI":"10.18653\/v1\/2021.emnlp-main.409","article-title":"Pre-train or annotate? Domain adaptation with a constrained budget","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Bai","year":"2021"},{"key":"2026040612495722900_bib13","article-title":"Training a helpful and harmless assistant with reinforcement learning from human feedback","author":"Bai","year":"2022"},{"key":"2026040612495722900_bib14","article-title":"Constitutional AI: Harmlessness from AI feedback","author":"Bai","year":"2022"},{"key":"2026040612495722900_bib15","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.findings-acl.39","article-title":"Comparing bad apples to good oranges: Aligning large language models via joint preference optimization","author":"Bansal","year":"2025"},{"key":"2026040612495722900_bib16","doi-asserted-by":"publisher","first-page":"522","DOI":"10.18653\/v1\/2020.emnlp-main.38","article-title":"Self-supervised meta-learning for few-shot natural language classification tasks","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Bansal","year":"2020"},{"key":"2026040612495722900_bib17","article-title":"Efficient training of language models to fill in the middle","author":"Bavarian","year":"2022"},{"key":"2026040612495722900_bib18","doi-asserted-by":"publisher","first-page":"1","DOI":"10.18653\/v1\/2022.acl-short.1","article-title":"BitFit: Simple parameter-efficient fine-tuning for transformer-based masked language-models","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Zaken","year":"2022"},{"key":"2026040612495722900_bib19","doi-asserted-by":"publisher","DOI":"10.21203\/rs.3.rs-3865391\/v1","article-title":"A comparative analysis of encoder-only and decoder-only models in intent classification and sentiment analysis: Navigating the trade-offs in model size and performance","author":"Benayas","year":"2024","journal-title":"Language Resources and Evaluation"},{"key":"2026040612495722900_bib20","doi-asserted-by":"publisher","first-page":"3563","DOI":"10.18653\/v1\/2022.emnlp-main.233","article-title":"Language contamination helps explains the cross-lingual capabilities of English pretrained models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Blevins","year":"2022"},{"key":"2026040612495722900_bib21","article-title":"Language models are few-shot learners","author":"Brown","year":"2020"},{"key":"2026040612495722900_bib22","doi-asserted-by":"publisher","first-page":"104431","DOI":"10.1016\/j.jbi.2023.104431","article-title":"Localizing in-domain adaptation of transformer-based biomedical language models","volume":"144","author":"Buonocore","year":"2023","journal-title":"Journal of Biomedical Informatics"},{"key":"2026040612495722900_bib23","article-title":"Multilingual alignment of contextual word representations","volume-title":"International Conference on Learning Representations","author":"Cao","year":"2019"},{"key":"2026040612495722900_bib24","doi-asserted-by":"publisher","first-page":"101928","DOI":"10.52202\/079017-3234","article-title":"Preference learning algorithms do not learn preference rankings","volume":"37","author":"Chen","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib25","article-title":"Maybe only 0.5% data is needed: A preliminary exploration of low training data instruction tuning","author":"Chen","year":"2023"},{"key":"2026040612495722900_bib26","doi-asserted-by":"publisher","first-page":"191","DOI":"10.1162\/tacl_a_00542","article-title":"An empirical survey of data augmentation for limited data learning in NLP","volume":"11","author":"Chen","year":"2023","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2026040612495722900_bib27","doi-asserted-by":"publisher","first-page":"2147","DOI":"10.18653\/v1\/2020.acl-main.194","article-title":"MixText: Linguistically-informed interpolation of hidden space for semi-supervised text classification","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Chen","year":"2020"},{"key":"2026040612495722900_bib28","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btad496","article-title":"Parameter-efficient fine-tuning design spaces","volume-title":"International Conference on Learning Representations","author":"Chen","year":"2022"},{"issue":"8","key":"2026040612495722900_bib29","doi-asserted-by":"crossref","first-page":"btad496","DOI":"10.1093\/bioinformatics\/btad496","article-title":"Few-shot biomedical named entity recognition via knowledge-guided instance generation and prompt contrastive learning","volume":"39","author":"Chen","year":"2023","journal-title":"Bioinformatics"},{"key":"2026040612495722900_bib30","doi-asserted-by":"publisher","first-page":"1059","DOI":"10.18653\/v1\/2022.emnlp-main.69","article-title":"Multilingual relation classification via efficient and effective prompting","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Chen","year":"2022"},{"key":"2026040612495722900_bib31","first-page":"6621","article-title":"Self- play fine-tuning converts weak language models to strong language models","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Chen","year":"2024"},{"key":"2026040612495722900_bib32","doi-asserted-by":"publisher","first-page":"3576","DOI":"10.18653\/v1\/2021.naacl-main.280","article-title":"InfoXLM: An information-theoretic framework for cross-lingual language model pre-training","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Chi","year":"2021"},{"key":"2026040612495722900_bib33","doi-asserted-by":"publisher","first-page":"3418","DOI":"10.18653\/v1\/2021.acl-long.265","article-title":"Improving pretrained cross-lingual language models via self-labeled word alignment","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Chi","year":"2021"},{"key":"2026040612495722900_bib34","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1109\/APSIPAASC58517.2023.10317500","article-title":"Learning meta soft prompt for few-shot language models","volume-title":"2023 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)","author":"Chien","year":"2023"},{"key":"2026040612495722900_bib35","article-title":"Deep reinforcement learning from human preferences","volume-title":"Advances in Neural Information Processing Systems","author":"Christiano","year":"2017"},{"key":"2026040612495722900_bib36","doi-asserted-by":"publisher","first-page":"2054","DOI":"10.18653\/v1\/2023.findings-eacl.153","article-title":"AdapterSoup: Weight averaging to improve generalization of pretrained language models","volume-title":"Findings of the Association for Computational Linguistics: EACL 2023","author":"Chronopoulou","year":"2023"},{"key":"2026040612495722900_bib37","doi-asserted-by":"publisher","first-page":"1914","DOI":"10.18653\/v1\/D18-1217","article-title":"Semi-supervised sequence modeling with cross-view training","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Clark","year":"2018"},{"key":"2026040612495722900_bib38","doi-asserted-by":"publisher","first-page":"8440","DOI":"10.18653\/v1\/2020.acl-main.747","article-title":"Unsupervised Cross-lingual representation learning at scale","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Conneau","year":"2020"},{"key":"2026040612495722900_bib39","article-title":"Cross-lingual language model pretraining","volume-title":"Advances in Neural Information Processing Systems","author":"Conneau","year":"2019"},{"key":"2026040612495722900_bib40","doi-asserted-by":"publisher","first-page":"104557","DOI":"10.1016\/j.jbi.2023.104557","article-title":"Advancing Italian biomedical information extraction with transformers-based models: Methodological insights and multicenter practical application","volume":"148","author":"Crema","year":"2023","journal-title":"Journal of Biomedical Informatics"},{"key":"2026040612495722900_bib41","doi-asserted-by":"publisher","first-page":"3610","DOI":"10.18653\/v1\/2022.naacl-main.264","article-title":"When is BERT multilingual? Isolating crucial ingredients for cross-lingual transfer","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Deshpande","year":"2022"},{"key":"2026040612495722900_bib42","doi-asserted-by":"crossref","first-page":"10088","DOI":"10.52202\/075280-0441","article-title":"QLoRA: Efficient finetuning of quantized LLMs","volume":"36","author":"Dettmers","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib43","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2026040612495722900_bib44","doi-asserted-by":"publisher","first-page":"4133","DOI":"10.18653\/v1\/2023.emnlp-main.252","article-title":"Sparse low-rank adaptation of pre-trained language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Ding","year":"2023"},{"issue":"3","key":"2026040612495722900_bib45","doi-asserted-by":"publisher","first-page":"220","DOI":"10.1038\/s42256-023-00626-4","article-title":"Parameter-efficient fine-tuning of large-scale pre-trained language models","volume":"5","author":"Ding","year":"2023","journal-title":"Nature Machine Intelligence"},{"key":"2026040612495722900_bib46","article-title":"Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping","author":"Dodge","year":"2020"},{"key":"2026040612495722900_bib47","article-title":"RAFT: Reward rAnked FineTuning for generative foundation model alignment","author":"Dong","year":"2023","journal-title":"Transactions on Machine Learning Research"},{"key":"2026040612495722900_bib48","doi-asserted-by":"publisher","first-page":"4019","DOI":"10.18653\/v1\/2020.acl-main.370","article-title":"Adversarial and domain-aware bert for cross-domain sentiment analysis","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Chunning","year":"2020"},{"key":"2026040612495722900_bib49","article-title":"KronA: Parameter efficient tuning with kronecker adapter","author":"Edalati","year":"2022"},{"key":"2026040612495722900_bib50","doi-asserted-by":"publisher","first-page":"7949","DOI":"10.18653\/v1\/2020.emnlp-main.638","article-title":"Active learning for BERT: An empirical study","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Ein-Dor","year":"2020"},{"key":"2026040612495722900_bib51","article-title":"KTO: Model alignment as prospect theoretic optimization","author":"Ethayarajh","year":"2024"},{"issue":"120","key":"2026040612495722900_bib52","first-page":"1","article-title":"Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity","volume":"23","author":"Fedus","year":"2022","journal-title":"Journal of Machine Learning Research"},{"key":"2026040612495722900_bib53","doi-asserted-by":"publisher","first-page":"968","DOI":"10.18653\/v1\/2021.findings-acl.84","article-title":"A survey of data augmentation approaches for NLP","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Feng","year":"2021"},{"key":"2026040612495722900_bib54","first-page":"1126","article-title":"Model-agnostic meta-learning for fast adaptation of deep networks","volume-title":"Proceedings of the 34th International Conference on Machine Learning","author":"Finn","year":"2017"},{"key":"2026040612495722900_bib55","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v38i20.30205","article-title":"MedAlign: A clinician-generated dataset for instruction following with electronic medical records","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Fleming","year":"2024"},{"key":"2026040612495722900_bib56","doi-asserted-by":"publisher","first-page":"233","DOI":"10.1145\/3617233.3617278","article-title":"Active learning with few-shot learning for crisis management","volume-title":"Proceedings of the 20th International Conference on Content-based Multimedia Indexing","author":"Fran\u00e7ois","year":"2023"},{"key":"2026040612495722900_bib57","article-title":"Red teaming language models to reduce harms: Methods, scaling behaviors, and lessons learned","author":"Ganguli","year":"2022","journal-title":"CoRR"},{"key":"2026040612495722900_bib58","article-title":"Towards a unified view of preference learning for large language models: A survey","author":"Gao","year":"2024"},{"key":"2026040612495722900_bib59","first-page":"10835","article-title":"Scaling laws for reward model overoptimization","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Gao","year":"2023"},{"key":"2026040612495722900_bib60","doi-asserted-by":"publisher","first-page":"6894","DOI":"10.18653\/v1\/2021.emnlp-main.552","article-title":"SimCSE: Simple contrastive learning of sentence embeddings","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Gao","year":"2021"},{"issue":"24","key":"2026040612495722900_bib61","doi-asserted-by":"publisher","first-page":"5004","DOI":"10.3390\/math11245004","article-title":"Leveraging zero and few-shot learning for enhanced model generality in hate speech detection in Spanish and English","volume":"11","author":"Antonio Garc\u00eda-D\u00edaz","year":"2023","journal-title":"Mathematics"},{"key":"2026040612495722900_bib62","doi-asserted-by":"publisher","first-page":"3020","DOI":"10.18653\/v1\/2023.findings-acl.189","article-title":"Exploring the relationship between alignment and cross-lingual transfer in multilingual transformers","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Gaschi","year":"2023"},{"key":"2026040612495722900_bib63","first-page":"3892","article-title":"Evaluation of transfer learning and domain adaptation for analyzing german-speaking job advertisements","volume-title":"Proceedings of the Thirteenth Language Resources and Evaluation Conference","author":"Gnehm","year":"2022"},{"key":"2026040612495722900_bib64","doi-asserted-by":"publisher","first-page":"477","DOI":"10.18653\/v1\/2024.emnlp-industry.36","article-title":"Arcee\u2019s MergeKit: A toolkit for merging large language models","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track","author":"Goddard","year":"2024"},{"key":"2026040612495722900_bib65","doi-asserted-by":"publisher","first-page":"2689","DOI":"10.18653\/v1\/2023.eacl-main.197","article-title":"SwitchPrompt: Learning domain-specific gated soft prompts for classification in low-resource domains","volume-title":"Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics","author":"Goswami","year":"2023"},{"key":"2026040612495722900_bib66","article-title":"The Llama 3 herd of models","author":"Grattafiori","year":"2024"},{"key":"2026040612495722900_bib67","doi-asserted-by":"publisher","first-page":"101056","DOI":"10.1016\/j.csl.2019.101056","article-title":"Low-resource text classification using domain-adversarial learning","volume":"62","author":"Grie\u00dfhaber","year":"2020","journal-title":"Computer Speech & Language"},{"key":"2026040612495722900_bib68","article-title":"Instruction tuned models are quick learners","author":"Gupta","year":"2023"},{"key":"2026040612495722900_bib69","doi-asserted-by":"publisher","first-page":"8342","DOI":"10.18653\/v1\/2020.acl-main.740","article-title":"Don\u2019t stop pretraining: Adapt language models to domains and tasks","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Gururangan","year":"2020"},{"key":"2026040612495722900_bib70","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1016\/j.aiopen.2021.08.002","article-title":"Pre-trained models: Past, present and future","volume":"2","author":"Han","year":"2021","journal-title":"AI Open"},{"key":"2026040612495722900_bib71","article-title":"Parameter-efficient fine-tuning for large models: A comprehensive survey","author":"Han","year":"2024","journal-title":"Transactions on Machine Learning Research"},{"key":"2026040612495722900_bib72","first-page":"17783","article-title":"LoRA+: Efficient low rank adaptation of large models","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Hayou","year":"2024"},{"key":"2026040612495722900_bib73","article-title":"Towards a unified view of parameter-efficient transfer learning","volume-title":"International Conference on Learning Representations","author":"He","year":"2021"},{"key":"2026040612495722900_bib74","article-title":"DeBERTa: Decoding-enhanced BERT with disentangled attention","volume-title":"International Conference on Learning Representations","author":"He","year":"2020"},{"key":"2026040612495722900_bib75","doi-asserted-by":"publisher","first-page":"2545","DOI":"10.18653\/v1\/2021.naacl-main.201","article-title":"A survey on recent approaches for natural language processing in low-resource scenarios","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Hedderich","year":"2021"},{"key":"2026040612495722900_bib76","first-page":"30016","article-title":"Training compute-optimal large language models","volume-title":"Advances in Neural Information Processing Systems","author":"Hoffmann","year":"2022"},{"key":"2026040612495722900_bib77","doi-asserted-by":"publisher","first-page":"11170","DOI":"10.18653\/v1\/2024.emnlp-main.626","article-title":"ORPO: Monolithic preference optimization without reference model","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Hong","year":"2024"},{"key":"2026040612495722900_bib78","first-page":"2790","article-title":"Parameter- efficient transfer learning for NLP","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"Houlsby","year":"2019"},{"key":"2026040612495722900_bib79","doi-asserted-by":"publisher","first-page":"328","DOI":"10.18653\/v1\/P18-1031","article-title":"Universal language model fine-tuning for text classification","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Howard","year":"2018"},{"key":"2026040612495722900_bib80","article-title":"LoRA: Low-rank adaptation of large language models","volume-title":"International Conference on Learning Representations","author":"Edward","year":"2021"},{"key":"2026040612495722900_bib81","doi-asserted-by":"publisher","first-page":"9133","DOI":"10.18653\/v1\/2023.findings-acl.581","article-title":"Language agnostic multilingual information retrieval with contrastive learning","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Xiyang","year":"2023"},{"key":"2026040612495722900_bib82","doi-asserted-by":"publisher","first-page":"5254","DOI":"10.18653\/v1\/2023.emnlp-main.319","article-title":"LLM-adapters: An adapter family for parameter-efficient fine-tuning of large language models","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Zhiqiang","year":"2023"},{"key":"2026040612495722900_bib83","doi-asserted-by":"publisher","first-page":"12365","DOI":"10.18653\/v1\/2023.findings-emnlp.826","article-title":"Not all languages are created equal in LLMs: Improving multilingual capability by cross-lingual-thought prompting","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Huang","year":"2023"},{"issue":"3","key":"2026040612495722900_bib84","doi-asserted-by":"publisher","first-page":"103279","DOI":"10.1016\/j.ipm.2023.103279","article-title":"Meta-prompt based learning for low-resource false information detection","volume":"60","author":"Huang","year":"2023","journal-title":"Information Processing & Management"},{"key":"2026040612495722900_bib85","doi-asserted-by":"publisher","first-page":"1048","DOI":"10.1145\/3539597.3570468","article-title":"Improving cross-lingual information retrieval on low-resource languages via optimal transport distillation","volume-title":"Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining","author":"Huang","year":"2023"},{"key":"2026040612495722900_bib86","doi-asserted-by":"publisher","first-page":"22771","DOI":"10.18653\/v1\/2024.emnlp-main.1267","article-title":"Instruction fine-tuning: Does prompt loss matter?","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Huerta-Enochian","year":"2024"},{"key":"2026040612495722900_bib87","doi-asserted-by":"publisher","first-page":"487","DOI":"10.18653\/v1\/2023.findings-acl.31","article-title":"TADA: Efficient task-agnostic domain adaptation for Transformers","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Hung","year":"2023"},{"key":"2026040612495722900_bib88","doi-asserted-by":"publisher","first-page":"1082","DOI":"10.18653\/v1\/2023.acl-long.61","article-title":"Glot500: Scaling multilingual corpora and language models to 500 languages","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Imani","year":"2023"},{"key":"2026040612495722900_bib89","article-title":"OPT-IML: Scaling language model instruction meta learning through the lens of generalization","author":"Iyer","year":"2023"},{"issue":"2","key":"2026040612495722900_bib90","doi-asserted-by":"publisher","first-page":"161","DOI":"10.1038\/s42256-023-00788-1","article-title":"Leveraging large language models for predictive chemistry","volume":"6","author":"Jablonka","year":"2024","journal-title":"Nature Machine Intelligence"},{"key":"2026040612495722900_bib91","article-title":"NEFTune: Noisy embeddings improve instruction finetuning","volume-title":"The Twelfth International Conference on Learning Representations","author":"Jain","year":"2024"},{"key":"2026040612495722900_bib92","first-page":"14702","article-title":"Exploring the benefits of training expert language models over instruction tuning","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Jang","year":"2023"},{"issue":"1","key":"2026040612495722900_bib93","doi-asserted-by":"publisher","first-page":"2353","DOI":"10.1038\/s41598-023-29323-3","article-title":"Information extraction from German radiological reports for general clinical text and language understanding","volume":"13","author":"Jantscher","year":"2023","journal-title":"Scientific Reports"},{"key":"2026040612495722900_bib94","article-title":"A survey on human preference learning for large language models","author":"Jiang","year":"2024"},{"key":"2026040612495722900_bib95","doi-asserted-by":"publisher","first-page":"112","DOI":"10.1109\/CogMI58952.2023.00025","article-title":"Rethinking learning rate tuning in the era of large language models","volume-title":"2023 IEEE 5th International Conference on Cognitive Machine Intelligence (CogMI)","author":"Jin","year":"2023"},{"key":"2026040612495722900_bib96","doi-asserted-by":"publisher","first-page":"5061","DOI":"10.18653\/v1\/2023.emnlp-main.307","article-title":"Parameter- efficient language model tuning with active learning in low-resource settings","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Juki\u0107","year":"2023"},{"key":"2026040612495722900_bib97","article-title":"Scaling laws for forgetting when fine-tuning large language models","author":"Kalajdzievski","year":"2024"},{"key":"2026040612495722900_bib98","article-title":"Scaling laws for neural language models","author":"Kaplan","year":"2020"},{"key":"2026040612495722900_bib99","doi-asserted-by":"publisher","first-page":"7265","DOI":"10.18653\/v1\/2021.acl-long.564","article-title":"Mind your outliers! Investigating the negative impact of outliers on active learning for visual question answering","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Karamcheti","year":"2021"},{"key":"2026040612495722900_bib100","first-page":"1022","article-title":"Compacter: Efficient low-rank hypercomplex adapter layers","volume-title":"Advances in Neural Information Processing Systems","author":"Mahabadi","year":"2021"},{"key":"2026040612495722900_bib101","doi-asserted-by":"publisher","first-page":"3638","DOI":"10.18653\/v1\/2022.acl-long.254","article-title":"Prompt-free and efficient few-shot learning with language models","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Mahabadi","year":"2022"},{"issue":"7","key":"2026040612495722900_bib102","doi-asserted-by":"publisher","first-page":"7087","DOI":"10.1609\/aaai.v36i7.20668","article-title":"Multiple-source domain adaptation via coordinated domain encoders and paired classifiers","volume":"36","author":"Karisani","year":"2022","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2026040612495722900_bib103","doi-asserted-by":"publisher","first-page":"239","DOI":"10.18653\/v1\/2023.mrl-1.18","article-title":"Contrastive learning for universal zero-shot NLI with cross-lingual sentence embeddings","volume-title":"Proceedings of the 3rd Workshop on Multi-lingual Representation Learning (MRL)","author":"Md","year":"2023"},{"key":"2026040612495722900_bib104","doi-asserted-by":"publisher","first-page":"66","DOI":"10.18653\/v1\/D18-2012","article-title":"SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Kudo","year":"2018"},{"key":"2026040612495722900_bib105","doi-asserted-by":"publisher","first-page":"5039","DOI":"10.18653\/v1\/D18-1549","article-title":"Phrase-based & neural unsupervised machine translation","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Lample","year":"2018"},{"issue":"1","key":"2026040612495722900_bib106","doi-asserted-by":"publisher","first-page":"84","DOI":"10.1017\/pan.2023.20","article-title":"Less annotating, more classifying: Addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI","volume":"32","author":"Laurer","year":"2024","journal-title":"Political Analysis"},{"key":"2026040612495722900_bib107","article-title":"Mixout: Effective regularization to finetune large-scale pretrained language models","volume-title":"International Conference on Learning Representations","author":"Lee","year":"2019"},{"key":"2026040612495722900_bib108","doi-asserted-by":"publisher","first-page":"26874","DOI":"10.3384\/VS.2001-5992.2024.11.1.1-8","article-title":"RLAIF vs. RLHF: Scaling reinforcement learning from human feedback with AI feedback","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Lee","year":"2024"},{"key":"2026040612495722900_bib109","doi-asserted-by":"publisher","first-page":"237","DOI":"10.18653\/v1\/2023.wassa-1.22","article-title":"Combining active learning and task adaptation with BERT for cost-effective annotation of social media datasets","volume-title":"Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis","author":"Lemmens","year":"2023"},{"key":"2026040612495722900_bib110","doi-asserted-by":"publisher","first-page":"3045","DOI":"10.18653\/v1\/2021.emnlp-main.243","article-title":"The power of scale for parameter- efficient prompt tuning","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Lester","year":"2021"},{"issue":"9","key":"2026040612495722900_bib111","doi-asserted-by":"publisher","first-page":"4245","DOI":"10.1109\/TKDE.2020.3038670","article-title":"Few-shot named entity recognition via meta-learning","volume":"34","author":"Li","year":"2022","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"2026040612495722900_bib112","doi-asserted-by":"publisher","first-page":"4582","DOI":"10.18653\/v1\/2021.acl-long.353","article-title":"Prefix- Tuning: Optimizing continuous prompts for generation","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Li","year":"2021"},{"key":"2026040612495722900_bib113","article-title":"AlpacaEval: An automatic evaluator of instruction-following models","author":"Li","year":"2023"},{"key":"2026040612495722900_bib114","article-title":"Scaling down to scale up: A guide to parameter-efficient fine- tuning","author":"Lialin","year":"2024"},{"key":"2026040612495722900_bib115","doi-asserted-by":"publisher","first-page":"9019","DOI":"10.18653\/v1\/2022.emnlp-main.616","article-title":"Few-shot learning with multilingual generative language models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Xi","year":"2022"},{"key":"2026040612495722900_bib116","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1007\/978-3-031-40292-0_23","article-title":"Evolutionary verbalizer search for prompt-based few shot text classification","volume-title":"Knowledge Science, Engineering and Management","author":"Ling","year":"2023"},{"key":"2026040612495722900_bib117","unstructured":"Guanghui\n              Liu\n            \n          . 2025. glgh\/awesome-llm-human-preference-datasets. Original-date: 2023-05-03T16:48:10Z."},{"key":"2026040612495722900_bib118","article-title":"Chain of hindsight aligns language models with feedback","volume-title":"International Conference on Learning Representations","author":"Liu","year":"2023"},{"key":"2026040612495722900_bib119","first-page":"1950","article-title":"Few-shot parameter- efficient fine-tuning is better and cheaper than in-context learning","volume":"35","author":"Liu","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib120","article-title":"Natural language fine-tuning","author":"Liu","year":"2024"},{"key":"2026040612495722900_bib121","doi-asserted-by":"publisher","first-page":"5035","DOI":"10.18653\/v1\/2023.findings-emnlp.335","article-title":"InteMATs: Integrating granularity-specific multilingual adapters for cross-lingual transfer","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Liu","year":"2023"},{"key":"2026040612495722900_bib122","doi-asserted-by":"publisher","first-page":"1461","DOI":"10.1145\/3583780.3614896","article-title":"GranCATs: Cross-lingual enhancement through granularity-specific contrastive adapters","volume-title":"Proceedings of the 32nd ACM International Conference on Information and Knowledge Management","author":"Liu","year":"2023"},{"issue":"9","key":"2026040612495722900_bib123","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3560815","article-title":"Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing","volume":"55","author":"Liu","year":"2023","journal-title":"ACM Computing Surveys"},{"key":"2026040612495722900_bib124","first-page":"32100","article-title":"DoRA: Weight-decomposed low-rank adaptation","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Liu","year":"2024"},{"key":"2026040612495722900_bib125","article-title":"RoBERTa: A robustly optimized bert pretraining approach","author":"Liu","year":"2019"},{"key":"2026040612495722900_bib126","first-page":"22631","article-title":"The flan collection: Designing data and methods for effective instruction tuning","volume-title":"Proceedings of the 40th International Conference on Machine Learning","author":"Longpre","year":"2023"},{"key":"2026040612495722900_bib127","doi-asserted-by":"publisher","first-page":"247","DOI":"10.18653\/v1\/2023.clinicalnlp-1.30","article-title":"Prompt discriminative language models for domain adaptation","volume-title":"Proceedings of the 5th Clinical Natural Language Processing Workshop","author":"Keming","year":"2023"},{"issue":"1","key":"2026040612495722900_bib128","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/s41524-025-01564-y","article-title":"Fine-tuning large language models for domain adaptation: Exploration of training strategies, scaling, model merging and synergistic capabilities","volume":"11","author":"Wei","year":"2025","journal-title":"NPJ Computational Materials"},{"key":"2026040612495722900_bib129","doi-asserted-by":"publisher","first-page":"942","DOI":"10.18653\/v1\/2022.emnlp-main.61","article-title":"Entity extraction in low resource domains with selective pre-training of large language models","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Mahapatra","year":"2022"},{"key":"2026040612495722900_bib130","doi-asserted-by":"publisher","first-page":"6253","DOI":"10.18653\/v1\/2022.acl-long.433","article-title":"UniPELT: A unified framework for parameter-efficient language model tuning","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Mao","year":"2022"},{"key":"2026040612495722900_bib131","doi-asserted-by":"publisher","first-page":"124198","DOI":"10.52202\/079017-3946","article-title":"SimPO: Simple preference optimization with a reference-free reward","volume":"37","author":"Meng","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib132","article-title":"The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation","author":"Meta","year":"2025"},{"key":"2026040612495722900_bib133","article-title":"Exploiting similarities among languages for machine translation","author":"Mikolov","year":"2013"},{"key":"2026040612495722900_bib134","doi-asserted-by":"publisher","first-page":"2791","DOI":"10.18653\/v1\/2022.naacl-main.201","article-title":"MetaICL: Learning to learn in context","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Min","year":"2022"},{"key":"2026040612495722900_bib135","doi-asserted-by":"publisher","first-page":"3992","DOI":"10.18653\/v1\/2022.naacl-main.293","article-title":"WECHSEL: Effective initialization of subword embeddings for cross-lingual transfer of monolingual language models","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Minixhofer","year":"2022"},{"issue":"11","key":"2026040612495722900_bib136","doi-asserted-by":"publisher","first-page":"13743","DOI":"10.1007\/s10462-023-10484-6","article-title":"Multi-task learning for few-shot biomedical relation extraction","volume":"56","author":"Moscato","year":"2023","journal-title":"Artificial Intelligence Reviews"},{"key":"2026040612495722900_bib137","first-page":"50358","article-title":"Scaling data-constrained language models","volume":"36","author":"Muennighoff","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib138","doi-asserted-by":"publisher","first-page":"15991","DOI":"10.18653\/v1\/2023.acl-long.891","article-title":"Crosslingual generalization through multitask finetuning","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Muennighoff","year":"2023"},{"key":"2026040612495722900_bib139","article-title":"Learning to route among specialized experts for zero-shot generalization","volume-title":"International Conference on Machine Learning","author":"Muqeeth","year":"2024"},{"key":"2026040612495722900_bib140","doi-asserted-by":"publisher","first-page":"7294","DOI":"10.18653\/v1\/2022.emnlp-main.492","article-title":"Eeny, meeny, miny, moe. How to choose data for morphological inflection","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Muradoglu","year":"2022"},{"key":"2026040612495722900_bib141","doi-asserted-by":"publisher","first-page":"8619","DOI":"10.18653\/v1\/2023.findings-acl.548","article-title":"Entropy-guided vocabulary augmentation of multilingual language models for low-resource tasks","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Nag","year":"2023"},{"key":"2026040612495722900_bib142","doi-asserted-by":"publisher","first-page":"4015","DOI":"10.18653\/v1\/2021.findings-acl.351","article-title":"Unsupervised domain adaptation for event detection using domain-specific adapters","volume-title":"Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021","author":"Trung","year":"2021"},{"key":"2026040612495722900_bib143","doi-asserted-by":"publisher","first-page":"1840","DOI":"10.1109\/SMC53992.2023.10394189","article-title":"A small claims court for the NLP: Judging legal text classification strategies with small datasets","volume-title":"2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC)","author":"Noguti","year":"2023"},{"key":"2026040612495722900_bib144","unstructured":"OpenAI, SandhiniAgarwal, LamaAhmad, JasonAi, SamAltman, AndyApplebaum, EdwinArbus, Rahul K.Arora, YuBai, BowenBaker, HaimingBao, BoazBarak, AllyBennett, TylerBertao, NiveditaBrett, EugeneBrevdo, GregBrockman, SebastienBubeck, CheChang, KaiChen, MarkChen, EnochCheung, AidanClark, DanCook, MaratDukhan, CaseyDvorak, KevinFives, VladFomenko, TimurGaripov, KristianGeorgiev, MiaGlaese, TarunGogineni, AdamGoucher, LukasGross, Katia GilGuzman, JohnHallman, JackieHehir, JohannesHeidecke, AlecHelyar, HaitangHu, RomainHuet, JacobHuh, SaachiJain, ZachJohnson, ChrisKoch, IrinaKofman, DominikKundel, JasonKwon, VolodymyrKyrylov, Elaine YaLe, GuillaumeLeclerc, James ParkLennon, ScottLessans, MarioLezcano-Casado, YuanzhiLi, ZhuohanLi, JiLin, JordanLiss, LilyLiu, JianchengLiu, KevinLu, ChrisLu, ZoranMartinovic, LindsayMcCallum, JoshMcGrath, ScottMcKinney, AidanMcLaughlin, SongMei, SteveMostovoy, TongMu, GideonMyles, AlexanderNeitz, AlexNichol, JakubPachocki, AlexPaino, DanaPalmie, AshleyPantuliano, GiambattistaParascandolo, JongsooPark, LeherPathak, CarolinaPaz, LudovicPeran, DmitryPimenov, MichellePokrass, ElizabethProehl, HuidaQiu, GabyRaila, FilippoRaso, HongyuRen, KimmyRichardson, DavidRobinson, BobRotsted, HadiSalman, SuvanshSanjeev, MaxSchwarzer, D.Sculley, HarshitSikchi, KendalSimon, KaranSinghal, YangSong, DaneStuckey, ZhiqingSun, PhilippeTillet, SamToizer, FoivosTsimpourlas, NikhilVyas, EricWallace, XinWang, MilesWang, OliviaWatkins, KevinWeil, AmyWendling, KevinWhinnery, CedricWhitney, HannahWong, LinYang, YuYang, MichihiroYasunaga, KristenYing, WojciechZaremba, WentingZhan, CyrilZhang, BrianZhang, EddieZhang, and ShengjiaZhao. 2025. gpt-oss-120b & gpt-oss-20b model card. ArXiv:2508.10925 [cs]."},{"key":"2026040612495722900_bib145","first-page":"38885","article-title":"Towards modular LLMs by building and reusing a library of LoRAs","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Ostapenko","year":"2024"},{"key":"2026040612495722900_bib146","first-page":"27730","article-title":"Training language models to follow instructions with human feedback","volume":"35","author":"Ouyang","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib147","article-title":"Unveiling the secret recipe: A guide for supervised fine-tuning small LLMs","author":"Pareja","year":"2024"},{"key":"2026040612495722900_bib148","doi-asserted-by":"publisher","first-page":"4998","DOI":"10.18653\/v1\/2024.findings-acl.297","article-title":"Disentangling length from quality in direct preference optimization","volume-title":"Findings of the Association for Computational Linguistics: ACL 2024","author":"Park","year":"2024"},{"key":"2026040612495722900_bib149","doi-asserted-by":"publisher","first-page":"11254","DOI":"10.18653\/v1\/2023.emnlp-main.692","article-title":"Anchoring fine-tuning of sentence transformer with semantic label information for efficient truly few-shot classification","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Pauli","year":"2023"},{"key":"2026040612495722900_bib150","doi-asserted-by":"publisher","first-page":"13603","DOI":"10.18653\/v1\/2023.findings-emnlp.908","article-title":"Less than one-shot: Named entity recognition via extremely weak supervision","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Peng","year":"2023"},{"key":"2026040612495722900_bib151","doi-asserted-by":"publisher","first-page":"487","DOI":"10.18653\/v1\/2021.eacl-main.39","article-title":"AdapterFusion: Non-destructive task composition for transfer learning","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Pfeiffer","year":"2021"},{"key":"2026040612495722900_bib152","article-title":"Modular deep learning","author":"Pfeiffer","year":"2023","journal-title":"Transactions on Machine Learning Research"},{"key":"2026040612495722900_bib153","doi-asserted-by":"publisher","first-page":"7654","DOI":"10.18653\/v1\/2020.emnlp-main.617","article-title":"MAD-X: An adapter- based framework for multi-task cross-lingual transfer","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Pfeiffer","year":"2020"},{"key":"2026040612495722900_bib154","article-title":"Sentence encoders on STILTs: Supplementary training on intermediate labeled-data tasks","author":"Phang","year":"2019"},{"key":"2026040612495722900_bib155","article-title":"Combining modular skills in multitask learning","author":"Ponti","year":"2022"},{"key":"2026040612495722900_bib156","doi-asserted-by":"publisher","first-page":"5231","DOI":"10.18653\/v1\/2020.acl-main.467","article-title":"Intermediate-task transfer learning with pretrained language models: When and why does it work?","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Pruksachatkun","year":"2020"},{"key":"2026040612495722900_bib157","doi-asserted-by":"publisher","first-page":"1910","DOI":"10.18653\/v1\/2022.acl-long.134","article-title":"Enhancing cross-lingual natural language inference by prompt-learning from cross-lingual templates","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Qi","year":"2022"},{"key":"2026040612495722900_bib158","doi-asserted-by":"publisher","first-page":"2695","DOI":"10.18653\/v1\/2023.emnlp-main.163","article-title":"Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages","volume-title":"Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing","author":"Qin","year":"2023"},{"key":"2026040612495722900_bib159","doi-asserted-by":"publisher","first-page":"16339","DOI":"10.18653\/v1\/2024.findings-acl.967","article-title":"Are decoder-only language models better than encoder-only language models in understanding word meaning?","volume-title":"Findings of the Association for Computational Linguistics: ACL 2024","author":"Qorib","year":"2024"},{"key":"2026040612495722900_bib160","doi-asserted-by":"publisher","first-page":"126207","DOI":"10.52202\/079017-4009","article-title":"Scaling laws for reward model overoptimization in direct alignment algorithms","volume":"37","author":"Rafailov","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib161","first-page":"53728","article-title":"Direct preference optimization: Your language model is secretly a reward model","volume":"36","author":"Rafailov","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib162","article-title":"Is reinforcement learning (not) for natural language processing: Benchmarks, baselines, and building blocks for natural language policy optimization","volume-title":"International Conference on Learning Representations","author":"Ramamurthy","year":"2022"},{"key":"2026040612495722900_bib163","article-title":"Effect of scale on catastrophic forgetting in neural networks","volume-title":"International Conference on Learning Representations","author":"Ramasesh","year":"2021"},{"key":"2026040612495722900_bib164","doi-asserted-by":"publisher","first-page":"7930","DOI":"10.18653\/v1\/2021.emnlp-main.626","article-title":"AdapterDrop: On the efficiency of adapters in transformers","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"R\u00fcckl\u00e9","year":"2021"},{"issue":"1","key":"2026040612495722900_bib165","doi-asserted-by":"publisher","first-page":"241","DOI":"10.1609\/aaai.v33i01.3301241","article-title":"Unsupervised neural machine translation with SMT as posterior regularization","volume":"33","author":"Ren","year":"2019","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2026040612495722900_bib166","doi-asserted-by":"publisher","first-page":"1856","DOI":"10.18653\/v1\/2023.findings-emnlp.125","article-title":"XTREME-UP: A user-centric scarce-data benchmark for under-represented languages","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Ruder","year":"2023"},{"key":"2026040612495722900_bib167","doi-asserted-by":"publisher","first-page":"64","DOI":"10.18653\/v1\/2022.mrl-1.6","article-title":"Comparative analysis of cross-lingual contextualized word embeddings","volume-title":"Proceedings of the 2nd Workshop on Multi-lingual Representation Learning (MRL)","author":"Saadi","year":"2022"},{"key":"2026040612495722900_bib168","article-title":"Multitask prompted training enables zero-shot task generalization","volume-title":"International Conference on Learning Representations","author":"Sanh","year":"2022"},{"key":"2026040612495722900_bib169","doi-asserted-by":"publisher","first-page":"255","DOI":"10.18653\/v1\/2021.eacl-main.20","article-title":"Exploiting cloze-questions for few-shot text classification and natural language inference","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Schick","year":"2021"},{"key":"2026040612495722900_bib170","doi-asserted-by":"publisher","first-page":"2339","DOI":"10.18653\/v1\/2021.naacl-main.185","article-title":"It\u2019s not just size that matters: Small language models are also few-shot learners","volume-title":"Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Schick","year":"2021"},{"key":"2026040612495722900_bib171","article-title":"Proximal policy optimization algorithms","author":"Schulman","year":"2017"},{"key":"2026040612495722900_bib172","article-title":"A critical evaluation of AI feedback for aligning large language models","volume-title":"The Thirty-eighth Annual Conference on Neural Information Processing Systems","author":"Sharma","year":"2024"},{"key":"2026040612495722900_bib173","article-title":"Outrageously large neural networks: The sparsely-gated mixture-of-experts layer","volume-title":"International Conference on Learning Representations","author":"Shazeer","year":"2017"},{"key":"2026040612495722900_bib174","article-title":"Towards data-centric RLHF: Simple metrics for preference dataset comparison","author":"Shen","year":"2024"},{"key":"2026040612495722900_bib175","doi-asserted-by":"publisher","first-page":"69176","DOI":"10.52202\/079017-2210","article-title":"Instruction tuning with loss over instructions","volume":"37","author":"Shi","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib176","doi-asserted-by":"publisher","first-page":"245","DOI":"10.1145\/325165.325242","article-title":"Animating rotation with quaternion curves","volume-title":"Proceedings of the 12th Annual Conference on Computer Graphics and Interactive Techniques","author":"Shoemake","year":"1985"},{"key":"2026040612495722900_bib177","article-title":"Prototypical networks for few-shot learning","volume-title":"Advances in Neural Information Processing Systems","author":"Snell","year":"2017"},{"key":"2026040612495722900_bib178","first-page":"596","article-title":"FixMatch: Simplifying semi- supervised learning with consistency and confidence","volume":"33","author":"Sohn","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib179","doi-asserted-by":"publisher","first-page":"110290","DOI":"10.1016\/j.knosys.2023.110290","article-title":"TaxonPrompt: Taxonomy-aware curriculum prompt learning for few-shot event classification","volume":"264","author":"Song","year":"2023","journal-title":"Knowledge-Based Systems"},{"key":"2026040612495722900_bib180","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1007\/978-3-031-28238-6_12","article-title":"Domain-aligned data augmentation for low-resource and imbalanced text classification","volume-title":"Advances in Information Retrieval","author":"Stylianou","year":"2023"},{"key":"2026040612495722900_bib181","first-page":"12991","article-title":"LST: Ladder side-tuning for parameter and memory efficient transfer learning","volume":"35","author":"Sung","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib182","first-page":"24193","article-title":"Training neural networks with fixed sparse masks","volume-title":"Advances in Neural Information Processing Systems","author":"Sung","year":"2021"},{"key":"2026040612495722900_bib183","doi-asserted-by":"publisher","first-page":"1433","DOI":"10.18653\/v1\/2020.findings-emnlp.129","article-title":"exBERT: Extending pre-trained models with domain-specific vocabulary under constrained training resources","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2020","author":"Tai","year":"2020"},{"key":"2026040612495722900_bib184","doi-asserted-by":"publisher","first-page":"8705","DOI":"10.18653\/v1\/2024.findings-emnlp.508","article-title":"Unlocking the potential of model merging for low-resource languages","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2024","author":"Tao","year":"2024"},{"key":"2026040612495722900_bib185","article-title":"Llama 2: Open foundation and fine-tuned chat models","author":"Touvron","year":"2023"},{"key":"2026040612495722900_bib186","doi-asserted-by":"publisher","first-page":"826","DOI":"10.1162\/tacl_a_00577","article-title":"Efficient methods for natural language processing: A survey","volume":"11","author":"Treviso","year":"2023","journal-title":"Transactions of the Association for Computational Linguistics"},{"issue":"4","key":"2026040612495722900_bib187","doi-asserted-by":"publisher","first-page":"433","DOI":"10.1007\/s41666-023-00140-7","article-title":"BioBERTurk: Exploring turkish biomedical language model development strategies in low-resource setting","volume":"7","author":"T\u00fcrkmen","year":"2023","journal-title":"Journal of Healthcare Informatics Research"},{"key":"2026040612495722900_bib188","doi-asserted-by":"publisher","first-page":"1278","DOI":"10.18653\/v1\/2024.findings-eacl.85","article-title":"Efficiently aligned cross-lingual transfer learning for conversational tasks using prompt-tuning","volume-title":"Findings of the Association for Computational Linguistics: EACL 2024","author":"Lifu","year":"2024"},{"key":"2026040612495722900_bib189","article-title":"Zephyr: Direct distillation of LM alignment","volume-title":"First Conference on Language Modeling","author":"Tunstall","year":"2024"},{"key":"2026040612495722900_bib190","doi-asserted-by":"publisher","first-page":"6747","DOI":"10.18653\/v1\/2023.findings-emnlp.449","article-title":"Comparing prompt-based and standard fine-tuning for urdu text classification","volume-title":"Findings of the Association for Computational Linguistics: EMNLP 2023","author":"Ullah","year":"2023"},{"key":"2026040612495722900_bib191","doi-asserted-by":"publisher","first-page":"449","DOI":"10.18653\/v1\/2023.bionlp-1.42","article-title":"RadAdapt: Radiology report summarization via lightweight domain adaptation of large language models","volume-title":"The 22nd Workshop on Biomedical Natural Language Processing and BioNLP Shared Tasks","author":"Van Veen","year":"2023"},{"key":"2026040612495722900_bib192","unstructured":"Leandro\n              von Werra\n            , YounesBelkada, LewisTunstall, EdwardBeeching, TristanThrush, NathanLambert, ShengyiHuang, KashifRasul, and QuentinGallou\u00e9dec. 2020. TRL: Transformer reinforcement learning. https:\/\/github.com\/huggingface\/trl."},{"key":"2026040612495722900_bib193","doi-asserted-by":"publisher","first-page":"7882","DOI":"10.18653\/v1\/2020.emnlp-main.635","article-title":"Exploring and predicting transferability across NLP tasks","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Tu","year":"2020"},{"key":"2026040612495722900_bib194","article-title":"Efficient large language models: A survey","author":"Wan","year":"2024","journal-title":"Transactions on Machine Learning Research"},{"key":"2026040612495722900_bib195","doi-asserted-by":"publisher","first-page":"15127","DOI":"10.18653\/v1\/2023.acl-long.843","article-title":"Towards unifying multi-lingual and cross-lingual summarization","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang","year":"2023"},{"key":"2026040612495722900_bib196","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/ISI58743.2023.10297258","article-title":"Boosting domain-specific question answering through weakly supervised self-training","volume-title":"2023 IEEE International Conference on Intelligence and Security Informatics (ISI)","author":"Wang","year":"2023"},{"key":"2026040612495722900_bib197","article-title":"InstructUIE: Multi-task instruction tuning for unified information extraction","author":"Wang","year":"2023"},{"key":"2026040612495722900_bib198","doi-asserted-by":"publisher","first-page":"5744","DOI":"10.18653\/v1\/2022.emnlp-main.388","article-title":"AdaMix: Mixture-of-adaptations for parameter-efficient model tuning","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Wang","year":"2022"},{"key":"2026040612495722900_bib199","first-page":"74764","article-title":"How far can camels go? Exploring the state of instruction tuning on open resources","volume":"36","author":"Wang","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib200","article-title":"Finetuned language models are zero- shot learners","volume-title":"International Conference on Learning Representations","author":"Wei","year":"2021"},{"key":"2026040612495722900_bib201","first-page":"24824","article-title":"Chain-of-thought prompting elicits reasoning in large language models","volume":"35","author":"Wei","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"4","key":"2026040612495722900_bib202","doi-asserted-by":"publisher","first-page":"102596","DOI":"10.1016\/j.ipm.2021.102596","article-title":"Enhanced prototypical network for few-shot relation extraction","volume":"58","author":"Wen","year":"2021","journal-title":"Information Processing & Management"},{"key":"2026040612495722900_bib203","doi-asserted-by":"publisher","first-page":"281","DOI":"10.18653\/v1\/K17-1029","article-title":"Neural domain adaptation for biomedical question answering","volume-title":"Proceedings of the 21st Conference on Computational Natural Language Learning (CoNLL 2017)","author":"Wiese","year":"2017"},{"key":"2026040612495722900_bib204","doi-asserted-by":"publisher","first-page":"38","DOI":"10.18653\/v1\/2020.emnlp-demos.6","article-title":"Transformers: State-of-the-art natural language processing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations","author":"Wolf","year":"2020"},{"key":"2026040612495722900_bib205","first-page":"23965","article-title":"Model soups: Averaging weights of multiple fine-tuned models improves accuracy without increasing inference time","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"Wortsman","year":"2022"},{"key":"2026040612495722900_bib206","doi-asserted-by":"publisher","first-page":"453","DOI":"10.1016\/j.neunet.2023.10.053","article-title":"Improving few-shot relation extraction through semantics-guided learning","volume":"169","author":"Hui","year":"2024","journal-title":"Neural Networks"},{"key":"2026040612495722900_bib207","doi-asserted-by":"publisher","first-page":"7551","DOI":"10.18653\/v1\/2023.acl-long.417","article-title":"Towards zero-shot multilingual transfer for code-switched responses","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Ting-Wei","year":"2023"},{"key":"2026040612495722900_bib208","first-page":"54104","article-title":"LESS: Selecting influential data for targeted instruction tuning","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Xia","year":"2024"},{"key":"2026040612495722900_bib209","first-page":"6256","article-title":"Unsupervised data augmentation for consistency training","volume-title":"Advances in Neural Information Processing Systems","author":"Xie","year":"2020"},{"key":"2026040612495722900_bib210","first-page":"34201","article-title":"Data selection for language models via importance resampling","volume":"36","author":"Xie","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib211","first-page":"55204","article-title":"Contrastive preference optimization: Pushing the boundaries of LLM performance in machine translation","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Haoran","year":"2024"},{"issue":"5","key":"2026040612495722900_bib212","doi-asserted-by":"publisher","first-page":"110:1\u2013110:39","DOI":"10.1145\/3706418","article-title":"Resource-efficient algorithms and systems of foundation models: A survey","volume":"57","author":"Mengwei","year":"2025","journal-title":"ACM Computing Surveys"},{"key":"2026040612495722900_bib213","doi-asserted-by":"publisher","first-page":"2089","DOI":"10.18653\/v1\/2021.acl-long.163","article-title":"Optimizing deeper transformers on small datasets","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Peng","year":"2021"},{"key":"2026040612495722900_bib214","doi-asserted-by":"publisher","first-page":"9514","DOI":"10.18653\/v1\/2021.emnlp-main.749","article-title":"Raise a child in large language model: Towards effective and generalizable fine-tuning","volume-title":"Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing","author":"Runxin","year":"2021"},{"key":"2026040612495722900_bib215","doi-asserted-by":"publisher","first-page":"5065","DOI":"10.18653\/v1\/2021.acl-long.393","article-title":"ConSERT: A contrastive framework for self- supervised sentence representation transfer","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Yan","year":"2021"},{"key":"2026040612495722900_bib216","first-page":"6032","article-title":"Unlearning bias in language models by partitioning gradients","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023","author":"Charles","year":"2023"},{"key":"2026040612495722900_bib217","doi-asserted-by":"publisher","first-page":"7935","DOI":"10.18653\/v1\/2020.emnlp-main.637","article-title":"Cold-start active learning through self-supervised language modeling","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP)","author":"Yuan","year":"2020"},{"key":"2026040612495722900_bib218","article-title":"RRHF: Rank responses to align language models with human feedback without tears","author":"Yuan","year":"2023"},{"key":"2026040612495722900_bib219","article-title":"When scaling meets LLM finetuning: The effect of data, model and finetuning method","volume-title":"The Twelfth International Conference on Learning Representations","author":"Zhang","year":"2023"},{"key":"2026040612495722900_bib220","first-page":"21442","article-title":"Fine-tuning pre-trained language models effectively by optimizing subnetworks adaptively","volume":"35","author":"Zhang","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib221","article-title":"Adaptive budget allocation for parameter-efficient fine-tuning","volume-title":"The Eleventh International Conference on Learning Representations","author":"Zhang","year":"2022"},{"key":"2026040612495722900_bib222","doi-asserted-by":"publisher","first-page":"586","DOI":"10.1109\/CVPR.2018.00068","article-title":"The unreasonable effectiveness of deep features as a perceptual metric","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zhang","year":"2018"},{"key":"2026040612495722900_bib223","article-title":"Instruction tuning for large language models: A survey","author":"Zhang","year":"2024"},{"key":"2026040612495722900_bib224","article-title":"Revisiting few-sample BERT fine-tuning","volume-title":"International Conference on Learning Representations","author":"Zhang","year":"2021"},{"key":"2026040612495722900_bib225","article-title":"Multi- task instruction tuning of LLaMa for specific scenarios: A preliminary study on writing assistance","author":"Zhang","year":"2023"},{"key":"2026040612495722900_bib226","doi-asserted-by":"publisher","first-page":"6477","DOI":"10.18653\/v1\/2023.acl-long.357","article-title":"Infusing hierarchical guidance into prompt tuning: A parameter-efficient framework for multi-level implicit discourse relation recognition","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhao","year":"2023"},{"key":"2026040612495722900_bib227","doi-asserted-by":"publisher","first-page":"1233","DOI":"10.1109\/CSCWD49262.2021.9437616","article-title":"A BERT based sentiment analysis and key entity detection approach for online financial texts","volume-title":"2021 IEEE 24th International Conference on Computer Supported Cooperative Work in Design (CSCWD)","author":"Zhao","year":"2021"},{"key":"2026040612495722900_bib228","doi-asserted-by":"publisher","first-page":"245","DOI":"10.1145\/3477495.3531933","article-title":"ADPL: Adversarial prompt-based domain adaptation for dialogue summarization with knowledge disentanglement","volume-title":"Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Zhao","year":"2022"},{"key":"2026040612495722900_bib229","article-title":"SLiC-HF: Sequence likelihood calibration with human feedback","author":"Zhao","year":"2023"},{"key":"2026040612495722900_bib230","first-page":"55006","article-title":"LIMA: Less is more for alignment","volume":"36","author":"Zhou","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040612495722900_bib231","first-page":"2223","article-title":"Unpaired image-to- image translation using cycle-consistent adversarial networks","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV)","author":"Zhu","year":"2017"},{"key":"2026040612495722900_bib232","article-title":"Fine-tuning language models from human preferences","author":"Ziegler","year":"2020"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.627\/2590819\/tacl.a.627.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/TACL.a.627\/2590819\/tacl.a.627.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,6]],"date-time":"2026-04-06T16:50:18Z","timestamp":1775494218000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/TACL.a.627\/136154\/Fine-tuning-Large-Language-Models-with-Limited"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026]]},"references-count":232,"URL":"https:\/\/doi.org\/10.1162\/tacl.a.627","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026]]},"published":{"date-parts":[[2026]]}}}