{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T22:15:38Z","timestamp":1784585738404,"version":"3.55.0"},"reference-count":160,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,10,7]],"date-time":"2024-10-07T00:00:00Z","timestamp":1728259200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"EU Horizon 2020 research and innovation programme","award":["952215"],"award-info":[{"award-number":["952215"]}]},{"name":"European Union under the Horizon Europe","award":["101079164"],"award-info":[{"award-number":["101079164"]}]},{"name":"European Union under the Horizon Europe","award":["101070093"],"award-info":[{"award-number":["101070093"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2025,1,31]]},"abstract":"<jats:p>Learning with limited labelled data, such as prompting, in-context learning, fine-tuning, meta-learning, or few-shot learning, aims to effectively train a model using only a small amount of labelled samples. However, these approaches have been observed to be excessively sensitive to the effects of uncontrolled randomness caused by non-determinism in the training process. The randomness negatively affects the stability of the models, leading to large variances in results across training runs. When such sensitivity is disregarded, it can unintentionally, but unfortunately also intentionally, create an imaginary perception of research progress. Recently, this area started to attract research attention and the number of relevant studies is continuously growing. In this survey, we provide a comprehensive overview of 415 papers addressing the effects of randomness on the stability of learning with limited labelled data. We distinguish between four main tasks addressed in the papers (investigate\/evaluate, determine, mitigate, benchmark\/compare\/report randomness effects), providing findings for each one. Furthermore, we identify and discuss seven challenges and open problems together with possible directions to facilitate further research. The ultimate goal of this survey is to emphasise the importance of this growing research area, which so far has not received an appropriate level of attention, and reveal impactful directions for future research.<\/jats:p>","DOI":"10.1145\/3691339","type":"journal-article","created":{"date-parts":[[2024,9,2]],"date-time":"2024-09-02T10:53:51Z","timestamp":1725274431000},"page":"1-40","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness"],"prefix":"10.1145","volume":"57","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0344-8620","authenticated-orcid":false,"given":"Branislav","family":"Pecher","sequence":"first","affiliation":[{"name":"Faculty of Information Technology, Brno University of Technology, Brno, Czech Republic and Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3511-5337","authenticated-orcid":false,"given":"Ivan","family":"Srba","sequence":"additional","affiliation":[{"name":"Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4105-3494","authenticated-orcid":false,"given":"Maria","family":"Bielikova","sequence":"additional","affiliation":[{"name":"Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia and Slovak.AI, Bratislava, Slovakia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,10,7]]},"reference":[{"key":"e_1_3_3_2_2","article-title":"Designing informative metrics for few-shot example selection","author":"Adiga Rishabh","year":"2024","unstructured":"Rishabh Adiga, Lakshminarayanan Subramanian, and Varun Chandrasekaran. 2024. Designing informative metrics for few-shot example selection. Retrieved from https:\/\/arXiv:2403.03861","journal-title":"Retrieved from https:\/\/arXiv:2403.03861"},{"key":"e_1_3_3_3_2","first-page":"20447","volume-title":"Advances in Neural Information Processing Systems","author":"Agarwal Mayank","year":"2021","unstructured":"Mayank Agarwal, Mikhail Yurochkin, and Yuekai Sun. 2021. On sensitivity of meta-learning to support data. In Advances in Neural Information Processing Systems, Vol. 34. Curran Associates, 20447\u201320460."},{"key":"e_1_3_3_4_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.564"},{"key":"e_1_3_3_5_2","volume-title":"R0-FoMo: Robustness of Few-shot and Zero-shot Learning in Large Foundation Models","author":"Ajith Anirudh","year":"2023","unstructured":"Anirudh Ajith, Mengzhou Xia, Ameet Deshpande, and Karthik R. Narasimhan. 2023. InstructEval: Systematic evaluation of instruction selection methods. In R0-FoMo: Robustness of Few-shot and Zero-shot Learning in Large Foundation Models. Retrieved from https:\/\/openreview.net\/forum?id=6FwaSOEeKD"},{"key":"e_1_3_3_6_2","article-title":"Reproducibility of machine learning: Terminology, recommendations and open issues","author":"Albertoni Riccardo","year":"2023","unstructured":"Riccardo Albertoni, Sara Colantonio, Piotr Skrzypczy\u0144ski, and Jerzy Stefanowski. 2023. Reproducibility of machine learning: Terminology, recommendations and open issues. Retrieved from https:\/\/arXiv:2302.12691","journal-title":"Retrieved from https:\/\/arXiv:2302.12691"},{"key":"e_1_3_3_7_2","series-title":"Proceedings of Machine Learning Research","first-page":"547","volume-title":"Proceedings of the 40th International Conference on Machine Learning","volume":"202","author":"Allingham James Urquhart","year":"2023","unstructured":"James Urquhart Allingham, Jie Ren, Michael W. Dusenberry, Xiuye Gu, Yin Cui, Dustin Tran, Jeremiah Zhe Liu, and Balaji Lakshminarayanan. 2023. A simple zero-shot prompt weighting technique to improve prompt ensembling in text-image models. In Proceedings of the 40th International Conference on Machine Learning(Proceedings of Machine Learning Research, Vol. 202). PMLR, 547\u2013568. Retrieved from https:\/\/proceedings.mlr.press\/v202\/allingham23a.html"},{"key":"e_1_3_3_8_2","article-title":"When benchmarks are targets: Revealing the sensitivity of large language model leaderboards","author":"Alzahrani Norah","year":"2024","unstructured":"Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq et\u00a0al. 2024. When benchmarks are targets: Revealing the sensitivity of large language model leaderboards. Retrieved from https:\/\/arXiv:2402.01781","journal-title":"Retrieved from https:\/\/arXiv:2402.01781"},{"key":"e_1_3_3_9_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.831"},{"key":"e_1_3_3_10_2","article-title":"Prompt design matters for computational social science tasks but in unpredictable ways","author":"Atreja Shubham","year":"2024","unstructured":"Shubham Atreja, Joshua Ashkinaze, Lingyao Li, Julia Mendelsohn, and Libby Hemphill. 2024. Prompt design matters for computational social science tasks but in unpredictable ways. Retrieved from https:\/\/arXiv:2406.11980","journal-title":"Retrieved from https:\/\/arXiv:2406.11980"},{"key":"e_1_3_3_11_2","unstructured":"Sinjini Banerjee Tim Marrinan Reilly Cannon Tony Chiang and Anand D. Sarwate. 2024. Measuring Model Variability using Robust Non-parametric Testing. Retrieved from https:\/\/arxiv.org\/abs\/2406.08307"},{"key":"e_1_3_3_12_2","article-title":"In-context learning with long-context models: An in-depth exploration","author":"Bertsch Amanda","year":"2024","unstructured":"Amanda Bertsch, Maor Ivgi, Uri Alon, Jonathan Berant, Matthew R. Gormley, and Graham Neubig. 2024. In-context learning with long-context models: An in-depth exploration. Retrieved from https:\/\/arXiv:2405.00200","journal-title":"Retrieved from https:\/\/arXiv:2405.00200"},{"key":"e_1_3_3_13_2","article-title":"Lessons from the trenches on reproducible evaluation of language models","author":"Biderman Stella","year":"2024","unstructured":"Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sidney Black, Jordan Clive et\u00a0al. 2024. Lessons from the trenches on reproducible evaluation of language models. Retrieved from https:\/\/arXiv:2405.14782","journal-title":"Retrieved from https:\/\/arXiv:2405.14782"},{"key":"e_1_3_3_14_2","volume-title":"RML@ICLR","author":"Boquet Thomas","year":"2019","unstructured":"Thomas Boquet, Laure Delisle, Denis Kochetkov, Nathan Schucher, Boris N. Oreshkin, and Julien Cornebise. 2019. Reproducibility and stability analysis in metric-based few-shot learning. In RML@ICLR, Vol. 3. Retrieved from https:\/\/openreview.net\/forum?id=B1g-SnUaUN"},{"key":"e_1_3_3_15_2","first-page":"747","volume-title":"Proceedings of Machine Learning and Systems","volume":"3","author":"Bouthillier Xavier","year":"2021","unstructured":"Xavier Bouthillier, Pierre Delaunay, Mirko Bronzi, Assya Trofimov, Brennan Nichyporuk, Justin Szeto, Naz Sepah, Edward Raff, Kanika Madan, Vikram Voleti, Samira Ebrahimi Kahou, Vincent Michalski, Dmitriy Serdyuk, Tal Arbel, Chris Pal, Ga\u00ebl Varoquaux, and Pascal Vincent. 2021. Accounting for variance in machine learning benchmarks. In Proceedings of Machine Learning and Systems, Vol. 3. 747\u2013769. Retrieved from https:\/\/proceedings.mlsys.org\/paper\/2021\/hash\/cfecdb276f634854f3ef915e2e980c31-Abstract.html"},{"key":"e_1_3_3_16_2","first-page":"725","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"Bouthillier Xavier","year":"2019","unstructured":"Xavier Bouthillier, C\u00e9sar Laurent, and Pascal Vincent. 2019. Unreproducible research is reproducible. In Proceedings of the 36th International Conference on Machine Learning. PMLR, 725\u2013734. Retrieved from https:\/\/proceedings.mlr.press\/v97\/bouthillier19a.html"},{"key":"e_1_3_3_17_2","first-page":"15787","volume-title":"Advances in Neural Information Processing Systems","author":"Bragg Jonathan","year":"2021","unstructured":"Jonathan Bragg, Arman Cohan, Kyle Lo, and Iz Beltagy. 2021. FLEX: Unifying evaluation for few-shot NLP. In Advances in Neural Information Processing Systems, Vol. 34. Curran Associates, 15787\u201315800."},{"key":"e_1_3_3_18_2","first-page":"1877","volume-title":"Advances in Neural Information Processing Systems","author":"Brown Tom","year":"2020","unstructured":"Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, Vol. 33. Curran Associates, 1877\u20131901."},{"key":"e_1_3_3_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-short.2"},{"key":"e_1_3_3_20_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.48"},{"key":"e_1_3_3_21_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.452"},{"key":"e_1_3_3_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510003.3510163"},{"key":"e_1_3_3_23_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-eacl.115"},{"key":"e_1_3_3_24_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.168"},{"key":"e_1_3_3_25_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.833"},{"key":"e_1_3_3_26_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i11.26495"},{"key":"e_1_3_3_27_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i6.20590"},{"key":"e_1_3_3_28_2","volume-title":"Advances in Neural Information Processing Systems","author":"Cutkosky Ashok","year":"2019","unstructured":"Ashok Cutkosky and Francesco Orabona. 2019. Momentum-based variance reduction in non-convex SGD. In Advances in Neural Information Processing Systems, Vol. 32. Curran Associates."},{"issue":"226","key":"e_1_3_3_29_2","first-page":"1","article-title":"Underspecification presents challenges for credibility in modern machine learning","volume":"23","author":"D\u2019Amour Alexander","year":"2022","unstructured":"Alexander D\u2019Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, and D. Sculley. 2022. Underspecification presents challenges for credibility in modern machine learning. J. Mach. Learn. Res. 23, 226 (2022), 1\u201361. Retrieved from http:\/\/jmlr.org\/papers\/v23\/20-1335.html","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_3_30_2","volume-title":"Advances in Neural Information Processing Systems","author":"Dauphin Yann N.","year":"2019","unstructured":"Yann N. Dauphin and Samuel Schoenholz. 2019. MetaInit: Initializing learning by learning to initialize. In Advances in Neural Information Processing Systems, Vol. 32. Curran Associates."},{"key":"e_1_3_3_31_2","doi-asserted-by":"publisher","unstructured":"Mostafa Dehghani Yi Tay Alexey A. Gritsenko Zhe Zhao Neil Houlsby Fernando Diaz Donald Metzler and Oriol Vinyals. 2021. The Benchmark Lottery. DOI:10.48550\/arXiv.2107.07002","DOI":"10.48550\/arXiv.2107.07002"},{"key":"e_1_3_3_32_2","doi-asserted-by":"publisher","unstructured":"Jesse Dodge Gabriel Ilharco Roy Schwartz Ali Farhadi Hannaneh Hajishirzi and Noah Smith. 2020. Fine-Tuning Pretrained Language Models: Weight Initializations Data Orders and Early Stopping. DOI:10.48550\/arXiv.2002.06305","DOI":"10.48550\/arXiv.2002.06305"},{"key":"e_1_3_3_33_2","article-title":"A survey for in-context learning","author":"Dong Qingxiu","year":"2022","unstructured":"Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Zhiyong Wu, Baobao Chang, Xu Sun, Jingjing Xu, and Zhifang Sui. 2022. A survey for in-context learning. Retrieved from https:\/\/arXiv:2301.00234","journal-title":"Retrieved from https:\/\/arXiv:2301.00234"},{"key":"e_1_3_3_34_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1266"},{"key":"e_1_3_3_35_2","doi-asserted-by":"crossref","unstructured":"Avia Efrat Or Honovich and Omer Levy. 2022. LMentry: A Language Model Benchmark of Elementary Language Tasks. Retrieved from http:\/\/arxiv.org\/abs\/2211.02069","DOI":"10.18653\/v1\/2023.findings-acl.666"},{"key":"e_1_3_3_36_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.783"},{"key":"e_1_3_3_37_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20044-1_6"},{"key":"e_1_3_3_38_2","first-page":"1","volume-title":"Proceedings of the 37th Pacific Asia Conference on Language, Information and Computation","author":"Gan Chengguang","year":"2023","unstructured":"Chengguang Gan and Tatsunori Mori. 2023. Sensitivity and robustness of large language models to prompt template in Japanese text classification tasks. In Proceedings of the 37th Pacific Asia Conference on Language, Information and Computation. ACL, Hong Kong, China, 1\u201311. https:\/\/aclanthology.org\/2023.paclic-1.1"},{"key":"e_1_3_3_39_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.679"},{"key":"e_1_3_3_40_2","article-title":"Sources of irreproducibility in machine learning: A review","author":"Gundersen Odd Erik","year":"2022","unstructured":"Odd Erik Gundersen, Kevin Coakley, and Christine Kirkpatrick. 2022. Sources of irreproducibility in machine learning: A review. Retrieved from https:\/\/arXiv:2204.07610","journal-title":"Retrieved from https:\/\/arXiv:2204.07610"},{"key":"e_1_3_3_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3589806.3600044"},{"key":"e_1_3_3_42_2","article-title":"Changing answer order can decrease mmlu accuracy","author":"Gupta Vipul","year":"2024","unstructured":"Vipul Gupta, David Pantoja, Candace Ross, Adina Williams, and Megan Ung. 2024. Changing answer order can decrease mmlu accuracy. Retrieved from https:\/\/arXiv:2406.19470","journal-title":"Retrieved from https:\/\/arXiv:2406.19470"},{"key":"e_1_3_3_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.sigdial-1.2"},{"issue":"9","key":"e_1_3_3_44_2","first-page":"5149","article-title":"Meta-learning in neural networks: A survey","volume":"44","author":"Hospedales Timothy","year":"2021","unstructured":"Timothy Hospedales, Antreas Antoniou, Paul Micaelli, and Amos Storkey. 2021. Meta-learning in neural networks: A survey. IEEE Trans. Pattern Anal. Mach. Intell. 44, 9 (2021), 5149\u20135169.","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"e_1_3_3_45_2","first-page":"3251","volume-title":"Proceedings of the 29th International Conference on Computational Linguistics","author":"Hou Yutai","year":"2022","unstructured":"Yutai Hou, Hongyuan Dong, Xinghao Wang, Bohan Li, and Wanxiang Che. 2022. MetaPrompting: Learning to learn better prompts. In Proceedings of the 29th International Conference on Computational Linguistics. International Committee on Computational Linguistics, 3251\u20133262. https:\/\/aclanthology.org\/2022.coling-1.287"},{"key":"e_1_3_3_46_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.258"},{"key":"e_1_3_3_47_2","first-page":"15398","volume-title":"Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING\u201924)","author":"Ji Baijun","year":"2024","unstructured":"Baijun Ji, Xiangyu Duan, Zhenyu Qiu, Tong Zhang, Junhui Li, Hao Yang, and Min Zhang. 2024. Submodular-based in-context example selection for LLMs-based machine translation. In Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING\u201924). ELRA and ICCL, 15398\u201315409. https:\/\/aclanthology.org\/2024.lrec-main.1337"},{"key":"e_1_3_3_48_2","unstructured":"Mingjian Jiang Yangjun Ruan Sicong Huang Saifei Liao Silviu Pitis Roger Baker Grosse and Jimmy Ba. 2023. Calibrating language models via augmented prompt ensembles. In Challenges in Deployable Generative AI. Retrieved from https:\/\/openreview.net\/forum?id=L0dc4wqbNs"},{"key":"e_1_3_3_49_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-naacl.11"},{"key":"e_1_3_3_50_2","article-title":"Challenges and applications of large language models","author":"Kaddour Jean","year":"2023","unstructured":"Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023. Challenges and applications of large language models. Retrieved from https:\/\/arXiv:2307.10169","journal-title":"Retrieved from https:\/\/arXiv:2307.10169"},{"key":"e_1_3_3_51_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.eval4nlp-1.3"},{"key":"e_1_3_3_52_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.36"},{"key":"e_1_3_3_53_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.395"},{"key":"e_1_3_3_54_2","first-page":"20571","volume-title":"Advances in Neural Information Processing Systems","author":"Levy Kfir","year":"2021","unstructured":"Kfir Levy, Ali Kavis, and Volkan Cevher. 2021. STORM+: Fully adaptive SGD with recursive momentum for nonconvex optimization. In Advances in Neural Information Processing Systems, Vol. 34. Curran Associates, 20571\u201320582."},{"key":"e_1_3_3_55_2","article-title":"MixPro: Simple yet effective data augmentation for prompt-based learning","author":"Li Bohan","year":"2023","unstructured":"Bohan Li, Longxu Dou, Yutai Hou, Yunlong Feng, Honglin Mu, and Wanxiang Che. 2023. MixPro: Simple yet effective data augmentation for prompt-based learning. Retrieved from https:\/\/arXiv:2304.09402","journal-title":"Retrieved from https:\/\/arXiv:2304.09402"},{"key":"e_1_3_3_56_2","article-title":"Large language model-aware in-context learning for code generation","author":"Li Jia","year":"2023","unstructured":"Jia Li, Ge Li, Chongyang Tao, Huangzhao Zhang, Fang Liu, and Zhi Jin. 2023. Large language model-aware in-context learning for code generation. Retrieved from https:\/\/arXiv:2310.09748","journal-title":"Retrieved from https:\/\/arXiv:2310.09748"},{"key":"e_1_3_3_57_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.411"},{"key":"e_1_3_3_58_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.495"},{"key":"e_1_3_3_59_2","article-title":"Holistic evaluation of language models","author":"Liang Percy","year":"2023","unstructured":"Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Alexander Cosgrove, Christopher D. Manning, Christopher Re, Diana Acosta-Navas, Drew Arad Hudson, Eric Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue WANG, Keshav Santhanam, Laurel Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Andrew Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. 2023. Holistic evaluation of language models. Trans. Mach. Learn. Res. (2023). Retrieved from https:\/\/openreview.net\/forum?id=iO4LZibEqW","journal-title":"Trans. Mach. Learn. Res."},{"key":"e_1_3_3_60_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.1060"},{"key":"e_1_3_3_61_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.deelio-1.10"},{"key":"e_1_3_3_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_3_63_2","unstructured":"Wei Liu Weihao Zeng Keqing He Yong Jiang and Junxian He. 2023. What makes good data for alignment? A comprehensive study of automatic data selection in instruction tuning. Retrieved from https:\/\/openreview.net\/forum?id=BTKAeLqLMw"},{"key":"e_1_3_3_64_2","article-title":"StablePT: Towards stable prompting for few-shot learning via input separation","author":"Liu Xiaoming","year":"2024","unstructured":"Xiaoming Liu, Chen Liu, Zhaohan Zhang, Chengzhengxu Li, Longtian Wang, Yu Lan, and Chao Shen. 2024. StablePT: Towards stable prompting for few-shot learning via input separation. Retrieved from https:\/\/arXiv:2404.19335","journal-title":"Retrieved from https:\/\/arXiv:2404.19335"},{"key":"e_1_3_3_65_2","article-title":"Let\u2019s learn step by step: Enhancing in-context learning ability with curriculum learning","author":"Liu Yinpeng","year":"2024","unstructured":"Yinpeng Liu, Jiawei Liu, Xiang Shi, Qikai Cheng, and Wei Lu. 2024. Let\u2019s learn step by step: Enhancing in-context learning ability with curriculum learning. Retrieved from https:\/\/arXiv:2402.10738","journal-title":"Retrieved from https:\/\/arXiv:2402.10738"},{"key":"e_1_3_3_66_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.556"},{"key":"e_1_3_3_67_2","first-page":"43136","volume-title":"Advances in Neural Information Processing Systems","volume":"36","author":"Ma Huan","year":"2023","unstructured":"Huan Ma, Changqing Zhang, Yatao Bian, Lemao Liu, Zhirui Zhang, Peilin Zhao, Shu Zhang, Huazhu Fu, Qinghua Hu, and Bingzhe Wu. 2023. Fairness-guided few-shot prompting for large language models. In Advances in Neural Information Processing Systems, Vol. 36. Curran Associates, 43136\u201343155. Retrieved from https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2023\/file\/8678da90126aa58326b2fc0254b33a8c-Paper-Conference.pdf"},{"key":"e_1_3_3_68_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.710"},{"key":"e_1_3_3_69_2","article-title":"Quantifying variance in evaluation benchmarks","author":"Madaan Lovish","year":"2024","unstructured":"Lovish Madaan, Aaditya K. Singh, Rylan Schaeffer, Andrew Poulton, Sanmi Koyejo, Pontus Stenetorp, Sharan Narang, and Dieuwke Hupkes. 2024. Quantifying variance in evaluation benchmarks. Retrieved from https:\/\/arXiv:2406.10229","journal-title":"Retrieved from https:\/\/arXiv:2406.10229"},{"key":"e_1_3_3_70_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.334"},{"key":"e_1_3_3_71_2","article-title":"Which examples to annotate for in-context learning? Towards effective and efficient selection","author":"Mavromatis Costas","year":"2023","unstructured":"Costas Mavromatis, Balasubramaniam Srinivasan, Zhengyuan Shen, Jiani Zhang, Huzefa Rangwala, Christos Faloutsos, and George Karypis. 2023. Which examples to annotate for in-context learning? Towards effective and efficient selection. Retrieved from https:\/\/arXiv:2310.20046","journal-title":"Retrieved from https:\/\/arXiv:2310.20046"},{"key":"e_1_3_3_72_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.blackboxnlp-1.21"},{"key":"e_1_3_3_73_2","unstructured":"Yu Meng Martin Michalski Jiaxin Huang Yu Zhang Tarek Abdelzaher and Jiawei Han. 2022. Tuning Language Models as Training Data Generators for Augmentation-Enhanced Few-Shot Learning. Retrieved from http:\/\/arxiv.org\/abs\/2211.03044"},{"key":"e_1_3_3_74_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.365"},{"key":"e_1_3_3_75_2","article-title":"State of what art? A call for multi-prompt llm evaluation","author":"Mizrahi Moran","year":"2023","unstructured":"Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky. 2023. State of what art? A call for multi-prompt llm evaluation. Retrieved from https:\/\/arXiv:2401.00595","journal-title":"Retrieved from https:\/\/arXiv:2401.00595"},{"key":"e_1_3_3_76_2","doi-asserted-by":"publisher","DOI":"10.7326\/0003-4819-151-4-200908180-00135"},{"key":"e_1_3_3_77_2","unstructured":"Marius Mosbach Maksym Andriushchenko and Dietrich Klakow. 2021. On the stability of fine-tuning BERT: Misconceptions explanations and strong baselines. Retrieved from https:\/\/openreview.net\/forum?id=nzpLWnVAyah"},{"key":"e_1_3_3_78_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.779"},{"key":"e_1_3_3_79_2","unstructured":"Subhabrata Mukherjee Xiaodong Liu Guoqing Zheng Saghar Hosseini Hao Cheng Greg Yang Christopher Meek Ahmed Hassan Awadallah and Jianfeng Gao. 2021. CLUES: Few-Shot Learning Evaluation in Natural Language Understanding. Retrieved from http:\/\/arxiv.org\/abs\/2111.02570"},{"key":"e_1_3_3_80_2","article-title":"In-context example selection with influences","author":"Nguyen Tai","year":"2023","unstructured":"Tai Nguyen and Eric Wong. 2023. In-context example selection with influences. Retrieved from https:\/\/arXiv:2302.11042","journal-title":"Retrieved from https:\/\/arXiv:2302.11042"},{"key":"e_1_3_3_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV51458.2022.00210"},{"key":"e_1_3_3_82_2","doi-asserted-by":"publisher","DOI":"10.1186\/s13643-021-01626-4"},{"key":"e_1_3_3_83_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.75"},{"key":"e_1_3_3_84_2","article-title":"Fighting randomness with randomness: Mitigating optimisation instability of fine-tuning using delayed ensemble and noisy interpolation","author":"Pecher Branislav","year":"2024","unstructured":"Branislav Pecher, Jan Cegin, Robert Belanec, Jakub Simko, Ivan Srba, and Maria Bielikova. 2024. Fighting randomness with randomness: Mitigating optimisation instability of fine-tuning using delayed ensemble and noisy interpolation. Retrieved from https:\/\/arXiv:2406.12471","journal-title":"Retrieved from https:\/\/arXiv:2406.12471"},{"key":"e_1_3_3_85_2","article-title":"Comparing specialised small and general large language models on text classification: 100 labelled samples to achieve break-even performance","author":"Pecher Branislav","year":"2024","unstructured":"Branislav Pecher, Ivan Srba, and Maria Bielikova. 2024. Comparing specialised small and general large language models on text classification: 100 labelled samples to achieve break-even performance. Retrieved from https:\/\/arXiv:2402.12819","journal-title":"Retrieved from https:\/\/arXiv:2402.12819"},{"key":"e_1_3_3_86_2","article-title":"On sensitivity of learning with limited labelled data to the effects of randomness: Impact of interactions and systematic choices","author":"Pecher Branislav","year":"2024","unstructured":"Branislav Pecher, Ivan Srba, and Maria Bielikova. 2024. On sensitivity of learning with limited labelled data to the effects of randomness: Impact of interactions and systematic choices. Retrieved from https:\/\/arXiv:2402.12817","journal-title":"Retrieved from https:\/\/arXiv:2402.12817"},{"key":"e_1_3_3_87_2","article-title":"Automatic combination of sample selection strategies for few-shot learning","author":"Pecher Branislav","year":"2024","unstructured":"Branislav Pecher, Ivan Srba, Maria Bielikova, and Joaquin Vanschoren. 2024. Automatic combination of sample selection strategies for few-shot learning. Retrieved from https:\/\/arXiv:2402.03038","journal-title":"Retrieved from https:\/\/arXiv:2402.03038"},{"key":"e_1_3_3_88_2","article-title":"Revisiting demonstration selection strategies in in-context learning","author":"Peng Keqin","year":"2024","unstructured":"Keqin Peng, Liang Ding, Yancheng Yuan, Xuebo Liu, Min Zhang, Yuanxin Ouyang, and Dacheng Tao. 2024. Revisiting demonstration selection strategies in in-context learning. Retrieved from https:\/\/arXiv:2401.12087","journal-title":"Retrieved from https:\/\/arXiv:2401.12087"},{"key":"e_1_3_3_89_2","article-title":"Large language models sensitivity to the order of options in multiple-choice questions","author":"Pezeshkpour Pouya","year":"2023","unstructured":"Pouya Pezeshkpour and Estevam Hruschka. 2023. Large language models sensitivity to the order of options in multiple-choice questions. Retrieved from https:\/\/arXiv:2308.11483","journal-title":"Retrieved from https:\/\/arXiv:2308.11483"},{"key":"e_1_3_3_90_2","doi-asserted-by":"publisher","DOI":"10.1145\/3324884.3416545"},{"key":"e_1_3_3_91_2","unstructured":"Jason Phang Thibault F\u00e9vry and Samuel R. Bowman. 2019. Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data Tasks. Retrieved from https:\/\/arxiv.org\/abs\/1811.01088"},{"key":"e_1_3_3_92_2","article-title":"Efficient multi-prompt evaluation of LLMs","author":"Polo Felipe Maia","year":"2024","unstructured":"Felipe Maia Polo, Ronald Xu, Lucas Weber, M\u00edrian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. 2024. Efficient multi-prompt evaluation of LLMs. Retrieved from https:\/\/arXiv:2405.17202","journal-title":"Retrieved from https:\/\/arXiv:2405.17202"},{"key":"e_1_3_3_93_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.494"},{"key":"e_1_3_3_94_2","doi-asserted-by":"crossref","unstructured":"Jian Qian Miao Sun Sifan Zhou Ziyu Zhao Ruizhi Hun and Patrick Chiang. 2024. Sub-SA: Strengthen In-context Learning via Submodular Selective Annotation. Retrieved from https:\/\/arxiv.org\/abs\/2407.05693","DOI":"10.3233\/FAIA240720"},{"key":"e_1_3_3_95_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.659"},{"key":"e_1_3_3_96_2","article-title":"In-context learning with iterative demonstration selection","author":"Qin Chengwei","year":"2023","unstructured":"Chengwei Qin, Aston Zhang, Anirudh Dagar, and Wenming Ye. 2023. In-context learning with iterative demonstration selection. Retrieved from https:\/\/arXiv:2310.09881","journal-title":"Retrieved from https:\/\/arXiv:2310.09881"},{"key":"e_1_3_3_97_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.421"},{"key":"e_1_3_3_98_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.378"},{"key":"e_1_3_3_99_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1035"},{"key":"e_1_3_3_100_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.191"},{"key":"e_1_3_3_101_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Sanh Victor","year":"2022","unstructured":"Victor Sanh, Albert Webson, Colin Raffel, Stephen Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M. Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal Nayak, Debajyoti Datta, Jonathan Chang, Mike Tian-Jian Jiang, Han Wang, Matteo Manica, Sheng Shen, Zheng Xin Yong, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Fevry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf, and Alexander M. Rush. 2022. Multitask prompted training enables zero-shot task generalization. In Proceedings of the International Conference on Learning Representations. Retrieved from https:\/\/openreview.net\/forum?id=9Vrb9D0WI4"},{"key":"e_1_3_3_102_2","doi-asserted-by":"publisher","DOI":"10.5555\/3618408.3619652"},{"key":"e_1_3_3_103_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.185"},{"key":"e_1_3_3_104_2","unstructured":"Melanie Sclar Yejin Choi Yulia Tsvetkov and Alane Suhr. 2023. Quantifying language models\u2019 sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting. Retrieved from https:\/\/openreview.net\/forum?id=RIu5lyNXjT"},{"key":"e_1_3_3_105_2","unstructured":"Thibault Sellam Steve Yadlowsky Ian Tenney Jason Wei Naomi Saphra Alexander D\u2019Amour Tal Linzen Jasmijn Bastings Iulia Turc Jacob Eisenstein Dipanjan Das and Ellie Pavlick. 2022. The MultiBERTs: BERT reproductions for robustness analysis. 30. Retrieved from https:\/\/openreview.net\/forum?id=K0E_F0gFDgA"},{"key":"e_1_3_3_106_2","unstructured":"Amrith Setlur Oscar Li and Virginia Smith. 2021. Is Support Set Diversity Necessary for Meta-Learning? Retrieved from http:\/\/arxiv.org\/abs\/2011.14048"},{"key":"e_1_3_3_107_2","first-page":"3770","volume-title":"Advances in Neural Information Processing Systems","author":"Setlur Amrith","year":"2021","unstructured":"Amrith Setlur, Oscar Li, and Virginia Smith. 2021. Two sides of meta-learning evaluation: In vs. Out of Distribution. In Advances in Neural Information Processing Systems, Vol. 34. Curran Associates, 3770\u20133783."},{"key":"e_1_3_3_108_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.811"},{"key":"e_1_3_3_109_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.277"},{"key":"e_1_3_3_110_2","doi-asserted-by":"publisher","DOI":"10.1145\/3582688"},{"key":"e_1_3_3_111_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.60"},{"key":"e_1_3_3_112_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.bsnlp-1.12"},{"key":"e_1_3_3_113_2","doi-asserted-by":"crossref","unstructured":"Hongjin Su Jungo Kasai Chen Henry Wu Weijia Shi Tianlu Wang Jiayi Xin Rui Zhang Mari Ostendorf Luke Zettlemoyer Noah A. Smith and Tao Yu. 2022. Selective annotation makes language models better few-shot learners. Retrieved from https:\/\/openreview.net\/forum?id=qY1hlv7gwg","DOI":"10.1109\/ICASSP49357.2023.10095738"},{"key":"e_1_3_3_114_2","first-page":"9913","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Summers Cecilia","year":"2021","unstructured":"Cecilia Summers and Michael J. Dinneen. 2021. Nondeterminism and instability in neural network optimization. In Proceedings of the 38th International Conference on Machine Learning. PMLR, 9913\u20139922. Retrieved from https:\/\/proceedings.mlr.press\/v139\/summers21a.html"},{"key":"e_1_3_3_115_2","unstructured":"Jiuding Sun Chantal Shaib and Byron C. Wallace. 2023. Evaluating the zero-shot robustness of instruction-tuned language models. Retrieved from https:\/\/openreview.net\/forum?id=g9diuvxN6D"},{"key":"e_1_3_3_116_2","article-title":"Mind your format: Towards consistent evaluation of in-context learning improvements","author":"Voronov Anton","year":"2024","unstructured":"Anton Voronov, Lena Wolf, and Max Ryabinin. 2024. Mind your format: Towards consistent evaluation of in-context learning improvements. Retrieved from https:\/\/arXiv:2401.06766","journal-title":"Retrieved from https:\/\/arXiv:2401.06766"},{"key":"e_1_3_3_117_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.346"},{"key":"e_1_3_3_118_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.naacl-main.206"},{"key":"e_1_3_3_119_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.877"},{"key":"e_1_3_3_120_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.conll-1.20"},{"key":"e_1_3_3_121_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.167"},{"key":"e_1_3_3_122_2","article-title":"Unveiling selection biases: Exploring order and token sensitivity in large language models","author":"Wei Sheng-Lun","year":"2024","unstructured":"Sheng-Lun Wei, Cheng-Kuang Wu, Hen-Hsen Huang, and Hsin-Hsi Chen. 2024. Unveiling selection biases: Exploring order and token sensitivity in large language models. Retrieved from https:\/\/arXiv:2406.03009","journal-title":"Retrieved from https:\/\/arXiv:2406.03009"},{"key":"e_1_3_3_123_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-short.76"},{"key":"e_1_3_3_124_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581641.3584059"},{"key":"e_1_3_3_125_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.repl4nlp-1.25"},{"key":"e_1_3_3_126_2","volume-title":"Proceedings of the International Conference on Machine Learning Workshop on In-Context Learning","author":"Wu Zhaoxuan","year":"2024","unstructured":"Zhaoxuan Wu, Xiaoqiang Lin, Zhongxiang Dai, Wenyang Hu, Yao Shu, See-Kiong Ng, Patrick Jaillet, and Bryan Kian Hsiang Low. 2024. Prompt optimization with EASE? Efficient ordering-aware automated selection of exemplars. In Proceedings of the International Conference on Machine Learning Workshop on In-Context Learning. Retrieved from https:\/\/openreview.net\/forum?id=TYxOXHYU6b"},{"key":"e_1_3_3_127_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.79"},{"key":"e_1_3_3_128_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.749"},{"key":"e_1_3_3_129_2","article-title":"Misconfidence-based demonstration selection for llm in-context learning","author":"Xu Shangqing","year":"2024","unstructured":"Shangqing Xu and Chao Zhang. 2024. Misconfidence-based demonstration selection for llm in-context learning. Retrieved from https:\/\/arXiv:2401.06301","journal-title":"Retrieved from https:\/\/arXiv:2401.06301"},{"key":"e_1_3_3_130_2","article-title":"In-context learning with retrieved demonstrations for language models: A survey","author":"Xu Xin","year":"2024","unstructured":"Xin Xu, Yue Liu, Panupong Pasupat, Mehran Kazemi et\u00a0al. 2024. In-context learning with retrieved demonstrations for language models: A survey. Retrieved from https:\/\/arXiv:2401.11624","journal-title":"Retrieved from https:\/\/arXiv:2401.11624"},{"key":"e_1_3_3_131_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.641"},{"key":"e_1_3_3_132_2","first-page":"25070","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"Yang Hansi","year":"2022","unstructured":"Hansi Yang and James Kwok. 2022. Efficient variance reduction for meta-learning. In Proceedings of the 39th International Conference on Machine Learning. PMLR, 25070\u201325095. Retrieved from https:\/\/proceedings.mlr.press\/v162\/yang22g.html"},{"key":"e_1_3_3_133_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.331"},{"key":"e_1_3_3_134_2","first-page":"39818","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Ye Jiacheng","year":"2023","unstructured":"Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong. 2023. Compositional exemplars for in-context learning. In Proceedings of the International Conference on Machine Learning. PMLR, 39818\u201339833."},{"key":"e_1_3_3_135_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.emnlp-main.572"},{"key":"e_1_3_3_136_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.141"},{"key":"e_1_3_3_137_2","article-title":"Unveiling the lexical sensitivity of LLMs: Combinatorial optimization for prompt enhancement","author":"Zhan Pengwei","year":"2024","unstructured":"Pengwei Zhan, Zhen Xu, Qian Tan, Jie Song, and Ru Xie. 2024. Unveiling the lexical sensitivity of LLMs: Combinatorial optimization for prompt enhancement. Retrieved from https:\/\/arXiv:2405.20701","journal-title":"Retrieved from https:\/\/arXiv:2405.20701"},{"key":"e_1_3_3_138_2","volume-title":"Proceedings of the 40th International Conference on Machine Learning (ICML\u201923)","author":"Zhang Biao","year":"2023","unstructured":"Biao Zhang, Barry Haddow, and Alexandra Birch. 2023. Prompting large language model for machine translation: A case study. In Proceedings of the 40th International Conference on Machine Learning (ICML\u201923). JMLR.org, Article 1722, 19 pages."},{"key":"e_1_3_3_139_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Zhang Hongyi","year":"2018","unstructured":"Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018. mixup: Beyond empirical risk minimization. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_3_140_2","unstructured":"Haojie Zhang Ge Li Jia Li Zhongjin Zhang Yuqi Zhu and Zhi Jin. 2022. Fine-tuning pre-trained language models effectively by optimizing subnetworks adaptively. 15. Retrieved from https:\/\/openreview.net\/forum?id=-r6-WNKfyhW"},{"key":"e_1_3_3_141_2","article-title":"Batch-ICL: Effective, efficient, and order-agnostic in-context learning","author":"Zhang Kaiyi","year":"2024","unstructured":"Kaiyi Zhang, Ang Lv, Yuhan Chen, Hansen Ha, Tao Xu, and Rui Yan. 2024. Batch-ICL: Effective, efficient, and order-agnostic in-context learning. Retrieved from https:\/\/arXiv:2401.06469","journal-title":"Retrieved from https:\/\/arXiv:2401.06469"},{"key":"e_1_3_3_142_2","article-title":"The impact of demonstrations on multilingual in-context learning: A multidimensional analysis","author":"Zhang Miaoran","year":"2024","unstructured":"Miaoran Zhang, Vagrant Gautam, Mingyang Wang, Jesujoba O. Alabi, Xiaoyu Shen, Dietrich Klakow, and Marius Mosbach. 2024. The impact of demonstrations on multilingual in-context learning: A multidimensional analysis. Retrieved from https:\/\/arXiv:2402.12976","journal-title":"Retrieved from https:\/\/arXiv:2402.12976"},{"key":"e_1_3_3_143_2","unstructured":"Tianyi Zhang Felix Wu Arzoo Katiyar Kilian Q. Weinberger and Yoav Artzi. 2021. Revisiting few-sample BERT fine-tuning. Retrieved from arXiv. https:\/\/openreview.net\/forum?id=cO1IH43yUF"},{"key":"e_1_3_3_144_2","doi-asserted-by":"crossref","unstructured":"Yiming Zhang Shi Feng and Chenhao Tan. 2022. Active Example Selection for In-Context Learning. Retrieved from http:\/\/arxiv.org\/abs\/2211.04486","DOI":"10.18653\/v1\/2022.emnlp-main.622"},{"key":"e_1_3_3_145_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.659"},{"key":"e_1_3_3_146_2","first-page":"4036","volume-title":"Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING\u201924)","author":"Zhao Feng","year":"2024","unstructured":"Feng Zhao, Wan Xianlin, Cheng Yan, and Chu Kiong Loo. 2024. Correcting language model bias for text classification in true zero-shot learning. In Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING\u201924). ELRA and ICCL, 4036\u20134046. Retrieved from https:\/\/aclanthology.org\/2024.lrec-main.359"},{"key":"e_1_3_3_147_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.447"},{"key":"e_1_3_3_148_2","article-title":"A survey of large language models","author":"Zhao Wayne Xin","year":"2023","unstructured":"Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong et\u00a0al. 2023. A survey of large language models. Retrieved from https:\/\/arXiv:2303.18223","journal-title":"Retrieved from https:\/\/arXiv:2303.18223"},{"key":"e_1_3_3_149_2","article-title":"NoisyICL: A little noise in model parameters calibrates in-context learning","author":"Zhao Yufeng","year":"2024","unstructured":"Yufeng Zhao, Yoshihiro Sakai, and Naoya Inoue. 2024. NoisyICL: A little noise in model parameters calibrates in-context learning. Retrieved from https:\/\/arXiv:2402.05515","journal-title":"Retrieved from https:\/\/arXiv:2402.05515"},{"key":"e_1_3_3_150_2","first-page":"12697","volume-title":"Proceedings of the 38th International Conference on Machine Learning","author":"Zhao Zihao","year":"2021","unstructured":"Zihao Zhao, Eric Wallace, Shi Feng, Dan Klein, and Sameer Singh. 2021. Calibrate before use: Improving few-shot performance of language models. In Proceedings of the 38th International Conference on Machine Learning. PMLR, 12697\u201312706. Retrieved from https:\/\/proceedings.mlr.press\/v139\/zhao21c.html"},{"key":"e_1_3_3_151_2","unstructured":"Chujie Zheng Hao Zhou Fandong Meng Jie Zhou and Minlie Huang. 2023. Large language models are not robust multiple choice selectors. Retrieved from https:\/\/openreview.net\/forum?id=shr9PXz7T0"},{"key":"e_1_3_3_152_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.acl-long.38"},{"key":"e_1_3_3_153_2","article-title":"Can chatgpt understand too? A comparative study on chatgpt and fine-tuned bert","author":"Zhong Qihuang","year":"2023","unstructured":"Qihuang Zhong, Liang Ding, Juhua Liu, Bo Du, and Dacheng Tao. 2023. Can chatgpt understand too? A comparative study on chatgpt and fine-tuned bert. Retrieved from https:\/\/arXiv:2302.10198","journal-title":"Retrieved from https:\/\/arXiv:2302.10198"},{"key":"e_1_3_3_154_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.334"},{"key":"e_1_3_3_155_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.192"},{"key":"e_1_3_3_156_2","unstructured":"Han Zhou Xingchen Wan Lev Proleev Diana Mincu Jilin Chen Katherine A. Heller and Subhrajit Roy. 2023. Batch calibration: Rethinking calibration for in-context learning and prompt engineering. Retrieved from https:\/\/openreview.net\/forum?id=L3FHMoKZcS"},{"key":"e_1_3_3_157_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.emnlp-main.659"},{"key":"e_1_3_3_158_2","article-title":"Promptbench: Towards evaluating the robustness of large language models on adversarial prompts","author":"Zhu Kaijie","year":"2023","unstructured":"Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang et\u00a0al. 2023. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. Retrieved from https:\/\/arXiv:2306.04528","journal-title":"Retrieved from https:\/\/arXiv:2306.04528"},{"key":"e_1_3_3_159_2","first-page":"16619","volume-title":"Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING\u201924)","author":"Zhu Shaolin","year":"2024","unstructured":"Shaolin Zhu, Menglong Cui, and Deyi Xiong. 2024. Towards robust in-context learning for machine translation with large language models. In Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING\u201924). ELRA and ICCL, 16619\u201316629. Retrieved from https:\/\/aclanthology.org\/2024.lrec-main.1444"},{"key":"e_1_3_3_160_2","first-page":"316","article-title":"Randomness in neural network training: Characterizing the impact of tooling","volume":"4","author":"Zhuang Donglin","year":"2022","unstructured":"Donglin Zhuang, Xingyao Zhang, Shuaiwen Song, and Sara Hooker. 2022. Randomness in neural network training: Characterizing the impact of tooling. Proc. Mach. Learn. Syst. 4 (Apr.2022), 316\u2013336. Retrieved from https:\/\/proceedings.mlsys.org\/paper_files\/paper\/2022\/hash\/427e0e886ebf87538afdf0badb805b7f-Abstract.html","journal-title":"Proc. Mach. Learn. Syst."},{"key":"e_1_3_3_161_2","article-title":"Fool your (vision and) language model with embarrassingly simple permutations","author":"Zong Yongshuo","year":"2023","unstructured":"Yongshuo Zong, Tingyang Yu, Bingchen Zhao, Ruchika Chavhan, and Timothy Hospedales. 2023. Fool your (vision and) language model with embarrassingly simple permutations. Retrieved from https:\/\/arXiv:2310.01651","journal-title":"Retrieved from https:\/\/arXiv:2310.01651"}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3691339","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3691339","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T17:49:56Z","timestamp":1750268996000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3691339"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,7]]},"references-count":160,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1,31]]}},"alternative-id":["10.1145\/3691339"],"URL":"https:\/\/doi.org\/10.1145\/3691339","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,10,7]]},"assertion":[{"value":"2023-02-24","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-08-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}