{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,10]],"date-time":"2026-04-10T10:03:05Z","timestamp":1775815385290,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":46,"publisher":"ACM","license":[{"start":{"date-parts":[[2024,5,13]],"date-time":"2024-05-13T00:00:00Z","timestamp":1715558400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100006374","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2022ZD0114804"],"award-info":[{"award-number":["2022ZD0114804"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006374","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62376154"],"award-info":[{"award-number":["62376154"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,5,13]]},"DOI":"10.1145\/3589334.3645420","type":"proceedings-article","created":{"date-parts":[[2024,5,8]],"date-time":"2024-05-08T07:08:13Z","timestamp":1715152093000},"page":"4059-4070","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Which LLM to Play? Convergence-Aware Online Model Selection with Time-Increasing Bandits"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-9800-1051","authenticated-orcid":false,"given":"Yu","family":"Xia","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University &amp; University of Michigan, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8148-8911","authenticated-orcid":false,"given":"Fang","family":"Kong","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5991-2050","authenticated-orcid":false,"given":"Tong","family":"Yu","sequence":"additional","affiliation":[{"name":"Adobe Research, San Jose, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8848-2814","authenticated-orcid":false,"given":"Liya","family":"Guo","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9758-0635","authenticated-orcid":false,"given":"Ryan A.","family":"Rossi","sequence":"additional","affiliation":[{"name":"Adobe Research, San Jose, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3580-5290","authenticated-orcid":false,"given":"Sungchul","family":"Kim","sequence":"additional","affiliation":[{"name":"Adobe Research, San Jose, CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3935-0708","authenticated-orcid":false,"given":"Shuai","family":"Li","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,5,13]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41060-017-0050-5"},{"key":"e_1_3_2_2_2_1","volume-title":"Conference on Learning Theory. PMLR, 138--158","author":"Auer Peter","year":"2019","unstructured":"Peter Auer, Pratik Gajane, and Ronald Ortner. 2019. Adaptively tracking the best bandit arm with an unknown number of distribution changes. In Conference on Learning Theory. PMLR, 138--158."},{"key":"e_1_3_2_2_3_1","volume-title":"Stochastic multi-armed-bandit problem with non-stationary rewards. Advances in neural information processing systems","author":"Besbes Omar","year":"2014","unstructured":"Omar Besbes, Yonatan Gur, and Assaf Zeevi. 2014. Stochastic multi-armed-bandit problem with non-stationary rewards. Advances in neural information processing systems, Vol. 27 (2014), 199--207."},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/3586589.3586666"},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2016.02.052"},{"key":"e_1_3_2_2_6_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. Advances in neural information processing systems Vol. 33 (2020) 1877--1901."},{"key":"e_1_3_2_2_7_1","volume-title":"The 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 418--427","author":"Cao Yang","year":"2019","unstructured":"Yang Cao, Zheng Wen, Branislav Kveton, and Yao Xie. 2019. Nearly optimal adaptive procedure with change detection for piecewise-stationary bandit. In The 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 418--427."},{"key":"e_1_3_2_2_8_1","volume-title":"Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research","volume":"1372","author":"Cella Leonardo","year":"2021","unstructured":"Leonardo Cella, Massimiliano Pontil, and Claudio Gentile. 2021. Best Model Identification: A Rested Bandit Formulation. In Proceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, 1362--1372. https:\/\/proceedings.mlr.press\/v139\/cella21a.html"},{"key":"e_1_3_2_2_9_1","unstructured":"Hyung Won Chung Le Hou Shayne Longpre Barret Zoph Yi Tay William Fedus Eric Li Xuezhi Wang Mostafa Dehghani Siddhartha Brahma et al. 2022. Scaling instruction-finetuned language models. arXiv preprint arXiv:2210.11416 (2022)."},{"key":"e_1_3_2_2_10_1","volume-title":"International Conference on Machine Learning. PMLR, 521--529","author":"Combes Richard","year":"2014","unstructured":"Richard Combes and Alexandre Proutiere. 2014. Unimodal bandits: Regret lower bounds and optimal algorithms. In International Conference on Machine Learning. PMLR, 521--529."},{"key":"e_1_3_2_2_11_1","volume-title":"Conference on learning theory. PMLR, 643--677","author":"Cutkosky Ashok","year":"2017","unstructured":"Ashok Cutkosky and Kwabena Boahen. 2017. Online learning without prior information. In Conference on learning theory. PMLR, 643--677."},{"key":"e_1_3_2_2_12_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning","volume":"80","author":"Falkner Stefan","year":"2018","unstructured":"Stefan Falkner, Aaron Klein, and Frank Hutter. 2018. BOHB: Robust and efficient hyperparameter optimization at scale. In Proceedings of the 35th International Conference on Machine Learning, Vol. 80. 1436--1445."},{"key":"e_1_3_2_2_13_1","volume-title":"The next generation. arXiv preprint arXiv:2007.04074","author":"Feurer Matthias","year":"2020","unstructured":"Matthias Feurer, Katharina Eggensperger, Stefan Falkner, Marius Lindauer, and Frank Hutter. 2020. Auto-sklearn 2.0: The next generation. arXiv preprint arXiv:2007.04074, Vol. 24 (2020)."},{"key":"e_1_3_2_2_14_1","volume-title":"Efficient and robust automated machine learning. Advances in neural information processing systems","author":"Feurer Matthias","year":"2015","unstructured":"Matthias Feurer, Aaron Klein, Katharina Eggensperger, Jost Springenberg, Manuel Blum, and Frank Hutter. 2015. Efficient and robust automated machine learning. Advances in neural information processing systems, Vol. 28 (2015)."},{"key":"e_1_3_2_2_15_1","volume-title":"Garnett (Eds.)","volume":"30","author":"Foster Dylan J","year":"2017","unstructured":"Dylan J Foster, Satyen Kale, Mehryar Mohri, and Karthik Sridharan. 2017. Parameter-Free Online Learning via Model Selection. In Advances in Neural Information Processing Systems, I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2017\/file\/a2186aa7c086b46ad4e8bf81e2a3a19b-Paper.pdf"},{"key":"e_1_3_2_2_16_1","volume-title":"Garnett (Eds.)","volume":"32","author":"Foster Dylan J","year":"2019","unstructured":"Dylan J Foster, Akshay Krishnamurthy, and Haipeng Luo. 2019. Model Selection for Contextual Bandits. In Advances in Neural Information Processing Systems, H. Wallach, H. Larochelle, A. Beygelzimer, F. dtextquotesingle Alch\u00e9-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Curran Associates, Inc. https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2019\/file\/433371e69eb202f8e7bc8ec2c8d48021-Paper.pdf"},{"key":"e_1_3_2_2_17_1","volume-title":"Proceedings of the 24th Annual Conference on Learning Theory (Proceedings of Machine Learning Research","volume":"376","author":"Garivier Aur\u00e9lien","year":"2011","unstructured":"Aur\u00e9lien Garivier and Olivier Capp\u00e9. 2011. The KL-UCB Algorithm for Bounded Stochastic Bandits and Beyond. In Proceedings of the 24th Annual Conference on Learning Theory (Proceedings of Machine Learning Research, Vol. 19), Sham M. Kakade and Ulrike von Luxburg (Eds.). PMLR, Budapest, Hungary, 359--376. https:\/\/proceedings.mlr.press\/v19\/garivier11a.html"},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/2050345.2050365"},{"key":"e_1_3_2_2_19_1","unstructured":"Hoda Heidari Michael J Kearns and Aaron Roth. 2016. Tight Policy Regret Bounds for Improving and Decaying Bandits.. In IJCAI. 1562--1570."},{"key":"e_1_3_2_2_20_1","volume-title":"8th ICML Workshop on Automated Machine Learning (AutoML). https:\/\/openreview.net\/forum?id=6tlvEH9HaX","author":"Heuillet Maxime","year":"2021","unstructured":"Maxime Heuillet, Benoit Debaque, and Audrey Durand. 2021. Sequential Automated Machine Learning: Bandits-driven Exploration using a Collaborative Filtering Representation. In 8th ICML Workshop on Automated Machine Learning (AutoML). https:\/\/openreview.net\/forum?id=6tlvEH9HaX"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-industry.24"},{"key":"e_1_3_2_2_22_1","volume-title":"Automated machine learning: methods, systems, challenges","author":"Hutter Frank","unstructured":"Frank Hutter, Lars Kotthoff, and Joaquin Vanschoren. 2019. Automated machine learning: methods, systems, challenges. Springer Nature."},{"key":"e_1_3_2_2_23_1","volume-title":"International Conference on Artificial Intelligence and Statistics. PMLR, 307--315","author":"Karimi Mohammad Reza","year":"2021","unstructured":"Mohammad Reza Karimi, Nezihe Merve G\u00fcrel, Bojan Karlavs, Johannes Rausch, Ce Zhang, and Andreas Krause. 2021. Online active model selection for pre-trained classifiers. In International Conference on Artificial Intelligence and Statistics. PMLR, 307--315."},{"key":"e_1_3_2_2_24_1","volume-title":"Auto-WEKA: Automatic model selection and hyperparameter optimization in WEKA. Automated machine learning: methods, systems, challenges","author":"Kotthoff Lars","year":"2019","unstructured":"Lars Kotthoff, Chris Thornton, Holger H Hoos, Frank Hutter, and Kevin Leyton-Brown. 2019. Auto-WEKA: Automatic model selection and hyperparameter optimization in WEKA. Automated machine learning: methods, systems, challenges (2019), 81--95."},{"key":"e_1_3_2_2_25_1","volume-title":"Bandit algorithms","author":"Lattimore Tor","unstructured":"Tor Lattimore and Csaba Szepesv\u00e1ri. 2020. Bandit algorithms. Cambridge University Press."},{"key":"e_1_3_2_2_26_1","volume-title":"Rotting bandits. Advances in neural information processing systems","author":"Levine Nir","year":"2017","unstructured":"Nir Levine, Koby Crammer, and Shie Mannor. 2017. Rotting bandits. Advances in neural information processing systems, Vol. 30 (2017)."},{"key":"e_1_3_2_2_27_1","first-page":"1","article-title":"Hyperband: A novel bandit-based approach to hyperparameter optimization","volume":"18","author":"Li Lisha","year":"2017","unstructured":"Lisha Li, Kevin Jamieson, Giulia DeSalvo, Afshin Rostamizadeh, and Ameet Talwalkar. 2017. Hyperband: A novel bandit-based approach to hyperparameter optimization. In Journal of Machine Learning Research, Vol. 18. 1--52.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5910"},{"key":"e_1_3_2_2_29_1","volume-title":"Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81.","author":"Lin Chin-Yew","year":"2004","unstructured":"Chin-Yew Lin. 2004. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out. 74--81."},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11746"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.5926"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002491"},{"key":"e_1_3_2_2_33_1","volume-title":"Stochastic Rising Bandits. In International Conference on Machine Learning. PMLR, 15421--15457","author":"Metelli Alberto Maria","year":"2022","unstructured":"Alberto Maria Metelli, Francesco Trovo, Matteo Pirola, and Marcello Restelli. 2022. Stochastic Rising Bandits. In International Conference on Machine Learning. PMLR, 15421--15457."},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D18-1206"},{"key":"e_1_3_2_2_35_1","unstructured":"OpenAI. 2023. GPT-4 Technical Report. arxiv: 2303.08774 [cs.CL]"},{"key":"e_1_3_2_2_36_1","volume-title":"Advances in Neural Information Processing Systems","volume":"27","author":"Orabona Francesco","year":"2014","unstructured":"Francesco Orabona. 2014. Simultaneous model selection and optimization through parameter-free stochastic learning. Advances in Neural Information Processing Systems, Vol. 27 (2014)."},{"key":"e_1_3_2_2_37_1","volume-title":"Koya: A Recommender System for Large Language Model Selection. In 4th Workshop on African Natural Language Processing. https:\/\/openreview.net\/forum?id=5DGm3lou3z","author":"Owodunni Abraham Toluwase","year":"2023","unstructured":"Abraham Toluwase Owodunni and Chris Chinenye Emezue. 2023. Koya: A Recommender System for Large Language Model Selection. In 4th Workshop on African Natural Language Processing. https:\/\/openreview.net\/forum?id=5DGm3lou3z"},{"key":"e_1_3_2_2_38_1","unstructured":"Baolin Peng Michel Galley Pengcheng He Hao Cheng Yujia Xie Yu Hu Qiuyuan Huang Lars Liden Zhou Yu Weizhu Chen et al. 2023. Check your facts and try again: Improving large language models with external knowledge and automated feedback. arXiv preprint arXiv:2302.12813 (2023)."},{"key":"e_1_3_2_2_39_1","unstructured":"Alec Radford Jeffrey Wu Rewon Child David Luan Dario Amodei Ilya Sutskever et al. 2019. Language models are unsupervised multitask learners. OpenAI blog Vol. 1 8 (2019) 9."},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.5555\/3455716.3455856"},{"key":"e_1_3_2_2_41_1","volume-title":"Advances in Neural Information Processing Systems","volume":"32","author":"Russac Yoan","year":"2019","unstructured":"Yoan Russac, Claire Vernade, and Olivier Capp\u00e9. 2019. Weighted linear bandits for non-stationary environments. Advances in Neural Information Processing Systems, Vol. 32 (2019)."},{"key":"e_1_3_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.240"},{"key":"e_1_3_2_2_43_1","volume-title":"The 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 2564--2572","author":"Seznec Julien","year":"2019","unstructured":"Julien Seznec, Andrea Locatelli, Alexandra Carpentier, Alessandro Lazaric, and Michal Valko. 2019. Rotting bandits are no harder than stochastic ones. In The 22nd International Conference on Artificial Intelligence and Statistics. PMLR, 2564--2572."},{"key":"e_1_3_2_2_44_1","volume-title":"International Conference on Artificial Intelligence and Statistics. PMLR, 3784--3794","author":"Seznec Julien","year":"2020","unstructured":"Julien Seznec, Pierre Menard, Alessandro Lazaric, and Michal Valko. 2020. A single algorithm for both restless and rested rotting bandits. In International Conference on Artificial Intelligence and Statistics. PMLR, 3784--3794."},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2012.2198613"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.11407"}],"event":{"name":"WWW '24: The ACM Web Conference 2024","location":"Singapore Singapore","acronym":"WWW '24","sponsor":["SIGWEB ACM Special Interest Group on Hypertext, Hypermedia, and Web"]},"container-title":["Proceedings of the ACM Web Conference 2024"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589334.3645420","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589334.3645420","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T00:28:48Z","timestamp":1755822528000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589334.3645420"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,13]]},"references-count":46,"alternative-id":["10.1145\/3589334.3645420","10.1145\/3589334"],"URL":"https:\/\/doi.org\/10.1145\/3589334.3645420","relation":{},"subject":[],"published":{"date-parts":[[2024,5,13]]},"assertion":[{"value":"2024-05-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}