{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,6]],"date-time":"2026-08-06T19:49:42Z","timestamp":1786045782452,"version":"3.56.0"},"reference-count":150,"publisher":"Emerald","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,5,23]]},"abstract":"<jats:p>Bandit algorithms, named after casino slot machines some\u00actimes known as \u201cone-armed bandits\u201d, fall into a broad cate\u00acgory of stochastic scheduling problems. In the setting with multiple arms, each arm generates a reward with a given probability. The gambler's aim is to find the arm producing the highest payoff and then continue playing in order to accumulate the maximum reward possible. However, having only a limited number of plays, the gambler is faced with a dilemma: should he play the arm currently known to produce the highest reward or should he keep on trying other arms in the hope of finding a better paying one? This problem formulation is easily applicable to many real-life scenarios, hence in recent years there has been an increased interest in developing bandit algorithms for a range of applications. In information retrieval and recommender systems, bandit algorithms, which are simple to implement and do not re\u00acquire any training data, have been particularly popular in online personalization, online ranker evaluation and search engine optimization. This survey provides a brief overview of bandit algorithms designed to tackle specific issues in infor\u00acmation retrieval and recommendation and, where applicable, it describes how they were applied in practice.<\/jats:p>","DOI":"10.1561\/1500000067","type":"journal-article","created":{"date-parts":[[2019,5,23]],"date-time":"2019-05-23T08:40:01Z","timestamp":1558600801000},"page":"299-424","source":"Crossref","is-referenced-by-count":16,"title":["Bandit Algorithms in Information Retrieval"],"prefix":"10.1108","volume":"13","author":[{"given":"Dorota","family":"G\u0142owacka","sequence":"first","affiliation":[{"name":"University of Helsinki"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2019,5,23]]},"reference":[{"key":"2026040314424724100_ref001","first-page":"17","article-title":"\u201cOnline models for content optimization\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Agarwal","year":"2009"},{"issue":"4","key":"2026040314424724100_ref002","doi-asserted-by":"crossref","first-page":"1054","DOI":"10.2307\/1427934","article-title":"\u201cSample mean based index policies by O (log n) regret for the multi-armed bandit problem\u201d","volume":"27","author":"Agrawal","year":"1995","journal-title":"Advances in Applied Probability"},{"key":"2026040314424724100_ref003","first-page":"39","article-title":"\u201cAnalysis of Thompson sampling for the multi-armed bandit problem\u201d","volume-title":"Conference on Learning Theory","author":"Agrawal","year":"2012"},{"key":"2026040314424724100_ref004","first-page":"856","article-title":"\u201cReducing dueling bandits to cardinal bandits\u201d","volume-title":"International Conference on Machine Learning","author":"Ailon","year":"2014"},{"key":"2026040314424724100_ref005","first-page":"397","article-title":"\u201cUsing confidence bounds for exploitation\u2013exploration trade-offs\u201d","volume":"3","author":"Auer","year":"2002","journal-title":"Journal of Machine Learning Research"},{"issue":"2\u20133","key":"2026040314424724100_ref006","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1023\/A:1013689704352","article-title":"\u201cFinite-time analysis of the multiarmed bandit problem\u201d","volume":"47","author":"Auer","year":"2002","journal-title":"Machine Learning"},{"issue":"1","key":"2026040314424724100_ref007","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1137\/S0097539701398375","article-title":"\u201cThe nonstochastic multiarmed bandit problem\u201d","volume":"32","author":"Auer","year":"2002","journal-title":"SIAM Journal on Computing"},{"key":"2026040314424724100_ref008","first-page":"51","article-title":"\u201cPinview: Implicit Feedback in Content-Based Image Retrieval\u201d","volume-title":"WAPA","author":"Auer","year":"2010"},{"issue":"1","key":"2026040314424724100_ref009","doi-asserted-by":"crossref","first-page":"97","DOI":"10.1016\/j.jcss.2007.04.016","article-title":"\u201cOnline linear optimization and adaptive routing\u201d","volume":"74","author":"Awerbuch","year":"2008","journal-title":"Journal of Computer and System Sciences"},{"key":"2026040314424724100_ref010","first-page":"324","article-title":"\u201cA contextual-bandit algorithm for mobile context-aware recommender system\u201d","volume-title":"International Conference on Neural Information Processing","author":"Bouneffouf","year":"2012"},{"key":"2026040314424724100_ref011","first-page":"3347","article-title":"\u201cA latent source model for online collaborative filtering\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Bresler","year":"2014"},{"key":"2026040314424724100_ref012","doi-asserted-by":"crossref","first-page":"745","DOI":"10.1145\/2911451.2914706","article-title":"\u201cAn improved multileaving algorithm for online ranker evaluation\u201d","volume-title":"Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Brost","year":"2016"},{"key":"2026040314424724100_ref013","first-page":"2161","article-title":"\u201cMulti-dueling bandits and their application to online ranker evaluation\u201d","volume-title":"Proceedings of the 25th ACM International Conference on Information and Knowledge Management","author":"Brost","year":"2016"},{"key":"2026040314424724100_ref014","article-title":"\u201cPreference-based online learning with dueling bandits: A survey\u201d","volume-title":"arXiv preprint arXiv:1807.11398","author":"Busa-Fekete","year":"2018"},{"key":"2026040314424724100_ref015","first-page":"1701","article-title":"\u201cPAC Rank Elicitation through Adaptive Sampling of Stochastic Pairwise Preferences\u201d","volume-title":"AAAI","author":"Busa-Fekete","year":"2014"},{"issue":"4","key":"2026040314424724100_ref016","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1561\/1500000055","article-title":"\u201cA survey of query auto completion in information retrieval\u201d","volume":"10","author":"Cai","year":"2016","journal-title":"Foundations and Trends in Information Retrieval"},{"key":"2026040314424724100_ref017","doi-asserted-by":"crossref","first-page":"11:1","DOI":"10.1145\/2501025.2501029","article-title":"\u201cMixing Bandits: A Recipe for Improved Cold-start Recommendations in a Social Network\u201d","volume-title":"Proceedings of the 7th Workshop on Social Network Mining and Analysis (SNAKDD \u201913)","author":"Caron","year":"2013"},{"key":"2026040314424724100_ref018","first-page":"737","article-title":"\u201cA gang of bandits\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Cesa-Bianchi","year":"2013"},{"key":"2026040314424724100_ref019","first-page":"273","article-title":"\u201cMortal multi-armed bandits\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Chakrabarti","year":"2009"},{"key":"2026040314424724100_ref020","first-page":"2249","article-title":"\u201cAn empirical evaluation of Thompson sampling\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Chapelle","year":"2011"},{"key":"2026040314424724100_ref021","first-page":"140","article-title":"\u201cInteractive Submodular Bandit\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Chen","year":"2017"},{"issue":"3","key":"2026040314424724100_ref022","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/978-3-031-02294-4","article-title":"\u201cClick models for web search\u201d","volume":"7","author":"Chuklin","year":"2015","journal-title":"Synthesis Lectures on Information Concepts, Retrieval, and Services"},{"issue":"1","key":"2026040314424724100_ref023","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1145\/2796314.2745852","article-title":"\u201cLearning to Rank: Regret Lower Bounds and Efficient Algorithms\u201d","volume":"43","author":"Combes","year":"2015","journal-title":"SIGMETRICS Performance Evaluation Review"},{"key":"2026040314424724100_ref024","doi-asserted-by":"crossref","first-page":"282","DOI":"10.1145\/290941.291009","article-title":"\u201cEfficient construction of large test collections\u201d","volume-title":"Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Cormack","year":"1998"},{"key":"2026040314424724100_ref025","first-page":"87","article-title":"\u201cAn experimental comparison of click position-bias models\u201d","volume-title":"Proceedings of the 2008 International Conference on Web Search and Data Mining","author":"Craswell","year":"2008"},{"key":"2026040314424724100_ref026","first-page":"71","article-title":"\u201cInteractive Intent Modeling from Multiple Feedback Domains\u201d","volume-title":"Proceedings of the 21st International Conference on Intelligent User Interfaces (IUI \u201916)","author":"Daee","year":"2016"},{"key":"2026040314424724100_ref027","doi-asserted-by":"crossref","first-page":"1750","DOI":"10.1109\/Allerton.2012.6483433","article-title":"\u201cLinear bandits in high dimension and recommendation systems\u201d","volume-title":"Communication, Control, and Computing (Allerton), 2012 50th Annual Allerton Conference on","author":"Deshpande","year":"2012"},{"key":"2026040314424724100_ref028","first-page":"77","article-title":"\u201cGaussian process modelling of dependencies in multi-armed bandit problems\u201d","volume-title":"Int. Symp. Op. Res","author":"Dorard","year":"2009"},{"key":"2026040314424724100_ref029","first-page":"563","article-title":"\u201cContextual dueling bandits\u201d","volume-title":"Proceedings of The 28th Conference on Learning Theory","author":"Dud\u00edk","year":"2015"},{"key":"2026040314424724100_ref030","unstructured":"Durand, A., Beaumont, J.-A., Gagn\u00e9, C., Lemay, M., and Paquet, S.\n          2017, \u201cQuery Completion Using Bandits for Engines Aggregation\u201d. arXiv preprint. arXiv:1709.04095."},{"key":"2026040314424724100_ref031","first-page":"359","article-title":"\u201cThe KL-UCB algorithm for bounded stochastic bandits and beyond\u201d","volume-title":"Proceedings of the 24th Annual Conference on Learning Theory","author":"Garivier","year":"2011"},{"key":"2026040314424724100_ref032","doi-asserted-by":"crossref","first-page":"174","DOI":"10.1007\/978-3-642-24412-4_16","article-title":"\u201cOn upper-confidence bound policies for switching bandit problems\u201d","volume-title":"International Conference on Algorithmic Learning Theory","author":"Garivier","year":"2011"},{"key":"2026040314424724100_ref033","first-page":"1253","article-title":"\u201cOn context-dependent clustering of bandits\u201d","volume-title":"International Conference on Machine Learning","author":"Gentile","year":"2017"},{"key":"2026040314424724100_ref034","first-page":"757","article-title":"\u201cOnline Clustering of Bandits\u201d","volume-title":"ICML","author":"Gentile","year":"2014"},{"key":"2026040314424724100_ref035","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1111\/j.2517-6161.1979.tb01068.x","article-title":"\u201cBandit processes and dynamic allocation indices\u201d","author":"Gittins","year":"1979","journal-title":"Journal of the Royal Statistical Society. Series B (Methodological)"},{"key":"2026040314424724100_ref036","first-page":"199","article-title":"\u201cPrior Knowledge in Learning Finite Parameter Spaces\u201d","volume-title":"International Conference on Formal Grammar","author":"G\u0142owacka","year":"2009"},{"key":"2026040314424724100_ref037","doi-asserted-by":"crossref","first-page":"124","DOI":"10.1145\/1498759.1498818","article-title":"\u201cEfficient multiple-click models in web search\u201d","volume-title":"Proceedings of the Second ACM International Conference on Web Search and Data Mining","author":"Guo","year":"2009"},{"key":"2026040314424724100_ref038","doi-asserted-by":"crossref","first-page":"2029","DOI":"10.1145\/1645953.1646293","article-title":"\u201cEvaluation of methods for relative comparison of retrieval systems based on clickthroughs\u201d","volume-title":"Proceedings of the 18th ACM Conference on Information and Knowledge Management","author":"He","year":"2009"},{"key":"2026040314424724100_ref039","first-page":"1813","article-title":"\u201cAn Efficient Bandit Algorithm for Realtime Multivariate Optimization\u201d","volume-title":"Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD \u201917)","author":"Hill","year":"2017"},{"key":"2026040314424724100_ref040","doi-asserted-by":"crossref","first-page":"249","DOI":"10.1145\/2063576.2063618","article-title":"\u201cA probabilistic method for inferring preferences from clicks\u201d","volume-title":"Proceedings of the 20th ACM International Conference on Information and Knowledge Management","author":"Hofmann","year":"2011"},{"issue":"4","key":"2026040314424724100_ref041","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1145\/2536736.2536737","article-title":"\u201cFidelity, soundness, and efficiency of interleaved comparison methods\u201d","volume":"31","author":"Hofmann","year":"2013","journal-title":"ACM Transactions on Information Systems"},{"key":"2026040314424724100_ref042","first-page":"251","article-title":"\u201cBalancing exploration and exploitation in learning to rank online\u201d","volume-title":"European Conference on Information Retrieval","author":"Hofmann","year":"2011"},{"issue":"1","key":"2026040314424724100_ref043","doi-asserted-by":"crossref","first-page":"63","DOI":"10.1007\/s10791-012-9197-9","article-title":"\u201cBalancing exploration and exploitation in listwise and pairwise online learning to rank for information retrieval\u201d","volume":"16","author":"Hofmann","year":"2013","journal-title":"Information Retrieval"},{"key":"2026040314424724100_ref044","first-page":"121","article-title":"\u201cA reinforcement learning approach to query-less image retrieval\u201d","volume-title":"International Workshop on Symbiotic Interaction","author":"Hore","year":"2015"},{"issue":"260","key":"2026040314424724100_ref045","doi-asserted-by":"crossref","first-page":"663","DOI":"10.1080\/01621459.1952.10483446","article-title":"\u201cA generalization of sampling without replacement from a finite universe\u201d","volume":"47","author":"Horvitz","year":"1952","journal-title":"Journal of the American Statistical Association"},{"key":"2026040314424724100_ref046","doi-asserted-by":"crossref","first-page":"740","DOI":"10.1145\/2695664.2695748","article-title":"\u201cEfficient Approximate Thompson Sampling for Search Query Recommendation\u201d","volume-title":"Proceedings of the 30th Annual ACM Symposium on Applied Computing (SAC \u201915)","author":"Hsieh","year":"2015"},{"key":"2026040314424724100_ref047","doi-asserted-by":"crossref","first-page":"1195","DOI":"10.1145\/2487575.2488198","article-title":"\u201cA Unified Search Federation System Based on Online User Feedback\u201d","volume-title":"Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD \u201913)","author":"Jie","year":"2013"},{"key":"2026040314424724100_ref048","article-title":"\u201cEvaluating Retrieval Performance Using Click-through Data\u201d","volume-title":"Text Mining","author":"Joachims","year":"2003"},{"key":"2026040314424724100_ref049","first-page":"1054","article-title":"\u201cNon-stochastic bandit slate problems\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Kale","year":"2010"},{"key":"2026040314424724100_ref050","unstructured":"Katariya, S., Kveton, B., Szepesv\u00e1ri, C., Vernade, C., and Wen, Z.\n          2016, \u201cStochastic rank-1 bandits\u201d. arXiv preprint arXiv:1608.03023."},{"key":"2026040314424724100_ref051","doi-asserted-by":"crossref","unstructured":"Katariya, S., Kveton, B., Szepesv\u00e1ri, C., Vernade, C., and Wen, Z.\n          2017, \u201cBernoulli Rank-1 Bandits for Click Feedback\u201d. arXiv preprint arXiv:1703.06513.","DOI":"10.24963\/ijcai.2017\/278"},{"key":"2026040314424724100_ref052","article-title":"\u201cDCM Bandits: Learning to Rank with Multiple Clicks\u201d","volume-title":"Proceedings of ICML","author":"Katariya","year":"2016"},{"key":"2026040314424724100_ref053","first-page":"592","article-title":"\u201cOn Bayesian upper confidence bounds for bandit problems\u201d","volume-title":"Artificial Intelligence and Statistics","author":"Kaufmann","year":"2012"},{"key":"2026040314424724100_ref054","first-page":"1297","article-title":"\u201cEfficient Thompson Sampling for Online Matrix-Factorization Recommendation\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Kawale","year":"2015"},{"key":"2026040314424724100_ref055","doi-asserted-by":"crossref","first-page":"681","DOI":"10.1145\/1374376.1374475","article-title":"\u201cMulti-armed bandits in metric spaces\u201d","author":"Kleinberg","year":"2008","journal-title":"Proceedings of the fortieth annual ACM symposium on Theory of Computing"},{"key":"2026040314424724100_ref056","first-page":"1911","article-title":"\u201cSpectral Thompson Sampling\u201d","author":"Koc\u00e1k","year":"2014","journal-title":"AAAI"},{"key":"2026040314424724100_ref057","first-page":"282","article-title":"\u201cBandit based Monte-Carlo planning\u201d","author":"Kocsis","year":"2006","journal-title":"European Conference on Machine Learning"},{"key":"2026040314424724100_ref058","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v27i1.8463","article-title":"\u201cA Fast Bandit Algorithm for Recommendation to Users With Heterogenous Tastes\u201d","author":"Kohli","year":"2013","journal-title":"AAAI"},{"issue":"1-3","key":"2026040314424724100_ref059","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/S0925-2312(98)00030-7","article-title":"\u201cThe self-organizing map\u201d","volume":"21","author":"Kohonen","year":"1998","journal-title":"Neurocomputing"},{"key":"2026040314424724100_ref060","first-page":"1141","article-title":"\u201cRegret lower bound and optimal algorithm in dueling bandit problem\u201d","author":"Komiyama","year":"2015","journal-title":"Proceedings of The 28th Conference on Learning Theory"},{"key":"2026040314424724100_ref061","first-page":"1235","article-title":"\u201cCopeland dueling bandit problem: regret lower bound, optimal algorithm, and computationally efficient algorithm\u201d","author":"Komiyama","year":"2016","journal-title":"Proceedings of the 33rd International Conference on Machine Learning"},{"key":"2026040314424724100_ref062","first-page":"5005","article-title":"\u201cPosition-based Multiple-play Bandit Problem with Unknown Position Bias\u201d","author":"Komiyama","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040314424724100_ref063","doi-asserted-by":"crossref","first-page":"460","DOI":"10.1007\/978-3-319-13129-0_40","article-title":"\u201cTime-decaying bandits for non-stationary systems\u201d","author":"Komiyama","year":"2014","journal-title":"International Conference on Web and Internet Economics"},{"key":"2026040314424724100_ref064","article-title":"\u201cContent-based image retrieval with hierarchical Gaussian Process bandits with self-organizing maps\u201d","author":"Konyushkova","year":"2013","journal-title":"ESANN"},{"issue":"8","key":"2026040314424724100_ref065","doi-asserted-by":"crossref","DOI":"10.1109\/MC.2009.263","article-title":"\u201cMatrix factorization techniques for recommender systems\u201d","volume":"42","author":"Koren","year":"2009","journal-title":"Computer"},{"key":"2026040314424724100_ref066","article-title":"\u201cAlgorithms for multi-armed bandit problems\u201d","author":"Kuleshov","year":"2014","journal-title":"arXiv preprint arXiv:1402.6028"},{"key":"2026040314424724100_ref067","first-page":"767","article-title":"\u201cCascading Bandits: Learning to Rank in the Cascade Model\u201d","author":"Kveton","year":"2015","journal-title":"Proceedings of the 32nd International Conference on Machine Learning"},{"key":"2026040314424724100_ref068","first-page":"1450","article-title":"\u201cCombinatorial cascading bandits\u201d","author":"Kveton","year":"2015","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040314424724100_ref069","first-page":"535","article-title":"\u201cTight regret bounds for stochastic combinatorial semi-bandits\u201d","author":"Kveton","year":"2015","journal-title":"Artificial Intelligence and Statistics"},{"key":"2026040314424724100_ref070","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1016\/j.neucom.2016.12.076","article-title":"\u201cMulti-Objective Ranked Bandits for Recommender Systems\u201d","volume":"246","author":"Lacerda","year":"2017","journal-title":"Neurocomputing"},{"key":"2026040314424724100_ref071","first-page":"1597","article-title":"\u201cMultiple-play bandits in the position-based model\u201d","author":"Lagr\u00e9e","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"1","key":"2026040314424724100_ref072","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1016\/0196-8858(85)90002-8","article-title":"\u201cAsymptotically efficient adaptive allocation rules\u201d","volume":"6","author":"Lai","year":"1985","journal-title":"Advances in Applied Mathematics"},{"key":"2026040314424724100_ref073","first-page":"528","article-title":"\u201cExploration scavenging\u201d","author":"Langford","year":"2008","journal-title":"Proceedings of the 25th International Conference on Machine Learning"},{"key":"2026040314424724100_ref074","first-page":"817","article-title":"\u201cThe epoch-greedy algorithm for multi-armed bandits with side information\u201d","author":"Langford","year":"2008","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040314424724100_ref075","first-page":"3077","article-title":"\u201cRotting bandits\u201d","author":"Levine","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040314424724100_ref076","first-page":"1089","article-title":"\u201cMultiple Queries As Bandit Arms\u201d","author":"Li","year":"2016","journal-title":"Proceedings of the 25th ACM International Conference on Information and Knowledge Management"},{"key":"2026040314424724100_ref077","doi-asserted-by":"crossref","first-page":"661","DOI":"10.1145\/1772690.1772758","article-title":"\u201cA contextual-bandit approach to personalized news article recommendation\u201d","author":"Li","year":"2010","journal-title":"Proceedings of the 19th International Conference on World Wide Web"},{"key":"2026040314424724100_ref078","doi-asserted-by":"crossref","first-page":"929","DOI":"10.1145\/2740908.2742562","article-title":"\u201cCounterfactual estimation and optimization of click metrics in search engines: A case study\u201d","author":"Li","year":"2015","journal-title":"Proceedings of the 24th International Conference on World Wide Web"},{"key":"2026040314424724100_ref079","first-page":"19","article-title":"\u201cAn Unbiased Offline Evaluation of Contextual Bandit Algorithms with Generalized Linear Models\u201d","author":"Li","year":"2011","journal-title":"JMLR Workshop and Conference Proceedings vol. 26: Online Trading of Exploration and Exploitation 2"},{"key":"2026040314424724100_ref080","doi-asserted-by":"crossref","unstructured":"Li, L., Chu, W., Langford, J., and Wang, X.\n          2011b, \u201cUnbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms\u201d. In: Proceedings of the Fourth ACM International Conference on Web Search and Data Mining (WSDM '11). Hong Kong, China: ACM. 297\u2013306. 10.1145\/1935826.1935878.","DOI":"10.1145\/1935826.1935878"},{"key":"2026040314424724100_ref081","doi-asserted-by":"crossref","DOI":"10.1145\/2911451.2911548","article-title":"\u201cCollaborative Filtering Bandits\u201d","volume-title":"The 39th International ACM SIGIR Conference on Information Retrieval (SIGIR)","author":"Li","year":"2016"},{"key":"2026040314424724100_ref082","first-page":"1245","article-title":"\u201cContextual combinatorial cascading bandits\u201d","volume-title":"International Conference on Machine Learning","author":"Li","year":"2016"},{"key":"2026040314424724100_ref083","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1145\/1835804.1835811","article-title":"\u201cExploitation and exploration in a performance based contextual advertising system\u201d","volume-title":"Proceedings of the 16th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Li","year":"2010"},{"key":"2026040314424724100_ref084","first-page":"15","article-title":"\u201cVariational Bayesian approach to movie rating prediction\u201d","volume-title":"Proceedings of KDD Cup and Workshop","author":"Lim","year":"2007"},{"key":"2026040314424724100_ref085","first-page":"225","article-title":"\u201cLearning to rank for information retrieval\u201d","volume-title":"Foundations and Trends in Information Retrieval","author":"Liu","year":"2009"},{"key":"2026040314424724100_ref086","article-title":"\u201cLetor: Benchmark dataset for research on learning to rank for information retrieval\u201d","volume-title":"Proceedings of SIGIR 2007 Workshop on Learning to Rank for Information Retrieval","author":"Liu","year":"2007"},{"key":"2026040314424724100_ref087","doi-asserted-by":"crossref","first-page":"1027","DOI":"10.1145\/2851613.2851692","article-title":"\u201cFeeling lucky?: multi-armed bandits for ordering judgements in pooling-based evaluation\u201d","volume-title":"Proceedings of the 31st Annual ACM Symposium on Applied Computing","author":"Losada","year":"2016"},{"key":"2026040314424724100_ref088","article-title":"\u201cShowing relevant ads via context multi-armed bandits\u201d","volume-title":"Proceedings of AISTATS","author":"Lu","year":"2009"},{"key":"2026040314424724100_ref089","doi-asserted-by":"crossref","unstructured":"Mahajan, D. K., Rastogi, R., Tiwari, C., and Mitra, A.\n          2012, \u201cLogUCB: An Explore-exploit Algorithm for Comments Recommendation\u201d. In: Proceedings of the 21st ACM International Conference on Information and Knowledge Management (CIKM '12). Maui, Hawaii, USA: ACM. 6\u201315. 10.1145\/2396761.2396767.","DOI":"10.1145\/2396761.2396767"},{"key":"2026040314424724100_ref090","doi-asserted-by":"crossref","unstructured":"Medlar, A., Ilves, K., Wang, P., Buntine, W., and Glowacka, D.\n          2016, \u201cPULP: A System for Exploratory Search of Scientific Literature\u201d. In: Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '16). Pisa, Italy: ACM. 1133\u20131136. 10.1145\/2911451.2911455.","DOI":"10.1145\/2911451.2911455"},{"key":"2026040314424724100_ref091","doi-asserted-by":"crossref","unstructured":"Medlar, A., Pyykko, J., and Glowacka, D.\n          2017, \u201cTowards Fine-Grained Adaptation of Exploration\/Exploitation in Information Retrieval\u201d. In: Proceedings of the 22nd International Conference on Intelligent User Interfaces (IUI '17). Limassol, Cyprus: ACM. 623\u2013627. 10.1145\/3025171.3025205.","DOI":"10.1145\/3025171.3025205"},{"key":"2026040314424724100_ref092","first-page":"1257","article-title":"\u201cProbabilistic matrix factorization\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Mnih","year":"2008"},{"key":"2026040314424724100_ref093","first-page":"315","article-title":"\u201cA UCB-Like Strategy of Collaborative Filtering\u201d","volume-title":"ACML","author":"Nakamura","year":"2014"},{"key":"2026040314424724100_ref094","doi-asserted-by":"crossref","first-page":"597","DOI":"10.1145\/2940716.2940771","article-title":"\u201cWhere to sell: Simulating auctions from learning algorithms\u201d","volume-title":"Proceedings of the 2016 ACM Conference on Economics and Computation","author":"Nazerzadeh","year":"2016"},{"key":"2026040314424724100_ref095","doi-asserted-by":"crossref","unstructured":"Nguyen, T. T., and Lauw, H. W.\n          2014, \u201cDynamic Clustering of Contextual Multi-Armed Bandits\u201d. In: Proceedings of the 23rd ACM International Conference on Information and Knowledge Management (CIKM '14). Shanghai, China: ACM. 1959\u20131962. 10.1145\/2661829.2662063.","DOI":"10.1145\/2661829.2662063"},{"key":"2026040314424724100_ref096","first-page":"216","article-title":"\u201cBandits for Taxonomies: A Model-based Approach\u201d","volume-title":"SDM","author":"Pandey","year":"2007"},{"key":"2026040314424724100_ref097","first-page":"461","article-title":"\u201cContextual combinatorial bandit and its application on diversified online recommendation\u201d","volume-title":"Proceedings of the 2014 SIAM International Conference on Data Mining","author":"Qin","year":"2014"},{"key":"2026040314424724100_ref098","doi-asserted-by":"crossref","first-page":"784","DOI":"10.1145\/1390156.1390255","article-title":"\u201cLearning diverse rankings with multi-armed bandits\u201d","volume-title":"Proceedings of the 25th International Conference on Machine Learning","author":"Radlinski","year":"2008"},{"key":"2026040314424724100_ref099","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1145\/1458082.1458092","article-title":"\u201cHow does click-through data reflect retrieval quality?\u201d","volume-title":"Proceedings of the 17th ACM conference on Information and knowledge management","author":"Radlinski","year":"2008"},{"key":"2026040314424724100_ref100","article-title":"\u201cRecommender Systems Handbook\u201d","author":"RicciLior","year":"2001"},{"key":"2026040314424724100_ref101","doi-asserted-by":"crossref","first-page":"521","DOI":"10.1145\/1242572.1242643","article-title":"\u201cPredicting clicks: estimating the click-through rate for new ads\u201d","volume-title":"Proceedings of the 16th international conference on World Wide Web","author":"Richardson","year":"2007"},{"key":"2026040314424724100_ref102","doi-asserted-by":"crossref","first-page":"169","DOI":"10.1007\/978-1-4612-5110-1_13","article-title":"\u201cSome aspects of the sequential design of experiments\u201d","volume-title":"Herbert Robbins Selected Papers","author":"Robbins","year":"1985"},{"key":"2026040314424724100_ref103","first-page":"1043","article-title":"\u201cSciNet: Interactive Intent Modeling for Information Discovery\u201d","volume-title":"Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR \u201815)","author":"Ruotsalo","year":"2015"},{"key":"2026040314424724100_ref104","first-page":"247","article-title":"\u201cTest collection based evaluation of information retrieval systems\u201d","volume-title":"Foundations and Trends in Information Retrieval","author":"Sanderson","year":"2010"},{"key":"2026040314424724100_ref105","doi-asserted-by":"crossref","first-page":"955","DOI":"10.1145\/2766462.2767838","article-title":"\u201cProbabilistic multileave for online retrieval evaluation\u201d","volume-title":"Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Schuth","year":"2015"},{"key":"2026040314424724100_ref106","first-page":"974","article-title":"\u201cPortfolio Choices with Orthogonal Bandit Learning\u201d","volume-title":"IJCAI","author":"Shen","year":"2015"},{"key":"2026040314424724100_ref107","article-title":"\u201cLearning optimally diverse rankings over large document collections\u201d","volume-title":"Proceedings of ICML","author":"Slivkins","year":"2010"},{"key":"2026040314424724100_ref108","first-page":"399","article-title":"\u201cRanked bandits in metric spaces: learning diverse rankings over large document collections\u201d","volume-title":"Journal of Machine Learning Research","author":"Slivkins","year":"2013"},{"key":"2026040314424724100_ref109","first-page":"1015","article-title":"\u201cGaussian process optimization in the bandit setting: No regret and experimental design\u201d","volume-title":"Proceedings of the 27th international conference on Machine learning","author":"Srinivas","year":"2010"},{"key":"2026040314424724100_ref110","first-page":"3386","article-title":"\u201cImproving Exploration in UCT Using Local Manifolds\u201d","volume-title":"AAAI","author":"Srinivasan","year":"2015"},{"key":"2026040314424724100_ref111","first-page":"1577","article-title":"\u201cAn online algorithm for maximizing submodular functions\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Streeter","year":"2008"},{"key":"2026040314424724100_ref112","first-page":"1794","article-title":"\u201cOnline learning of assignments\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Streeter","year":"2009"},{"key":"2026040314424724100_ref113","first-page":"2217","article-title":"\u201cLearning from logged implicit exploration data\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Strehl","year":"2010"},{"key":"2026040314424724100_ref114","article-title":"\u201cMulti-dueling Bandits with Dependent Arms\u201d","volume-title":"arXiv preprint arXiv:1705.00253","author":"Sui","year":"2017"},{"key":"2026040314424724100_ref115","first-page":"5502","article-title":"\u201cAdvancements in Dueling Bandits\u201d","volume-title":"IJCAI","author":"Sui","year":"2018"},{"key":"2026040314424724100_ref116","doi-asserted-by":"crossref","DOI":"10.1109\/TNN.1998.712192","article-title":"\u201cReinforcement learning: An introduction\u201d","author":"Sutton","year":"1998"},{"key":"2026040314424724100_ref117","first-page":"73","article-title":"\u201cEnsemble Contextual Bandits for Personalized Recommendation\u201d","volume-title":"Proceedings of the 8th ACM Conference on Recommender Systems (RecSys \u201814)","author":"Tang","year":"2014"},{"key":"2026040314424724100_ref118","first-page":"323","article-title":"\u201cPersonalized Recommendation via Parameter-Free Contextual Bandits\u201d","volume-title":"Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR \u201815)","author":"Tang","year":"2015"},{"key":"2026040314424724100_ref119","first-page":"1587","article-title":"\u201cAutomatic ad format selection via contextual bandits\u201d","volume-title":"Proceedings of the 22nd ACM International Conference on Information  Knowledge Management","author":"Tang","year":"2013"},{"key":"2026040314424724100_ref120","doi-asserted-by":"crossref","first-page":"757","DOI":"10.1145\/1277741.1277894","article-title":"\u201cCharacterizing the value of personalizing search\u201d","volume-title":"Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Teevan","year":"2007"},{"issue":"3\/4","key":"2026040314424724100_ref121","doi-asserted-by":"crossref","first-page":"285","DOI":"10.2307\/2332286","article-title":"\u201cOn the likelihood that one unknown probability exceeds another in view of the evidence of two samples\u201d","volume":"25","author":"Thompson","year":"1933","journal-title":"Biometrika"},{"key":"2026040314424724100_ref122","doi-asserted-by":"crossref","first-page":"2331","DOI":"10.1145\/3132847.3133046","article-title":"\u201cCollecting Non-Geotagged Local Tweets via Bandit Algorithms\u201d","volume-title":"Proceedings of the 2017 ACM Conference on Information and Knowledge Management","author":"Ueda","year":"2017"},{"key":"2026040314424724100_ref123","first-page":"91","article-title":"\u201cGeneric Exploration and K-armed Voting Bandits\u201d","volume-title":"ICML","author":"Urvoy","year":"2013"},{"key":"2026040314424724100_ref124","first-page":"46","article-title":"\u201cSpectral Bandits for Smooth Graph Functions\u201d","volume-title":"ICML","author":"Valko","year":"2014"},{"key":"2026040314424724100_ref125","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1145\/2645710.2645733","article-title":"\u201cExplore-exploit in top-n recommender systems via Gaussian processes\u201d","volume-title":"Proceedings of the 8th ACM Conference on Recommender Systems","author":"Vanchinathan","year":"2014"},{"issue":"2","key":"2026040314424724100_ref126","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1214\/14-STS504","article-title":"\u201cMulti-armed bandit models for the optimal design of clinical trials: benefits and challenges\u201d","volume":"30","author":"Villar","year":"2015","journal-title":"Statistical Science"},{"key":"2026040314424724100_ref127","doi-asserted-by":"publisher","first-page":"1177","DOI":"10.1145\/2736277.2741104","article-title":"\u201cGathering Additional Feedback on Search Results by Multi-Armed Bandits with Respect to Production Ranking\u201d","volume-title":"Proceedings of the 24th International Conference on World Wide Web (WWW \u201915)","author":"Vorobev","year":"2015"},{"key":"2026040314424724100_ref128","doi-asserted-by":"publisher","first-page":"1633","DOI":"10.1145\/2983323.2983847","article-title":"\u201cLearning Hidden Features for Contextual Bandits\u201d","volume-title":"Proceedings of the 25th ACM International Conference on Information and Knowledge Management (CIKM \u201916)","author":"Wang","year":"2016"},{"key":"2026040314424724100_ref129","first-page":"2695","article-title":"\u201cFactorization Bandits for Interactive Recommendation\u201d","volume-title":"AAAI","author":"Wang","year":"2017"},{"issue":"1","key":"2026040314424724100_ref130","doi-asserted-by":"publisher","first-page":"7:1","DOI":"10.1145\/2623372","article-title":"\u201cExploration in Interactive Personalized Music Recommendation: A Reinforcement Learning Approach\u201d","volume":"11","author":"Wang","year":"2014","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"issue":"11","key":"2026040314424724100_ref131","doi-asserted-by":"crossref","first-page":"2442","DOI":"10.1109\/TKDE.2017.2738639","article-title":"\u201cLearning Online Trends for Interactive Query Auto-Completion\u201d","volume":"29","author":"Wang","year":"2017","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"2026040314424724100_ref132","first-page":"2746","article-title":"\u201cEfficient Ordered Combinatorial Semi-Bandits for Whole-Page Recommendation\u201d","volume-title":"AAAI","author":"Wang","year":"2017"},{"key":"2026040314424724100_ref133","first-page":"1113","article-title":"\u201cEfficient learning in large-scale combinatorial semi-bandits\u201d","volume-title":"Proceedings of the 32nd International Conference on Machine Learning","author":"Wen","year":"2015"},{"key":"2026040314424724100_ref134","doi-asserted-by":"crossref","first-page":"136","DOI":"10.1016\/j.csda.2016.09.006","article-title":"\u201cA Bayesian adaptive design for clinical trials in rare diseases\u201d","volume":"113","author":"Williamson","year":"2017","journal-title":"Computational Statistics & Data Analysis"},{"key":"2026040314424724100_ref135","first-page":"649","article-title":"\u201cDouble thompson sampling for dueling bandits\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Wu","year":"2016"},{"key":"2026040314424724100_ref136","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1145\/2911451.2911528","article-title":"\u201cContextual Bandits in a Collaborative Environment\u201d","volume-title":"Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR \u201916)","author":"Wu","year":"2016"},{"key":"2026040314424724100_ref137","first-page":"445","article-title":"\u201cEnhancing Collaborative Filtering Music Recommendation by Balancing Exploration and Exploitation\u201d","volume-title":"ISMIR","author":"Xing","year":"2014"},{"issue":"5","key":"2026040314424724100_ref138","doi-asserted-by":"crossref","first-page":"1538","DOI":"10.1016\/j.jcss.2011.12.028","article-title":"\u201cThe k-armed dueling bandits problem\u201d","volume":"78","author":"Yue","year":"2012","journal-title":"Journal of Computer and System Sciences"},{"key":"2026040314424724100_ref139","first-page":"2483","article-title":"\u201cLinear submodular bandits and their application to diversified retrieval\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Yue","year":"2011"},{"key":"2026040314424724100_ref140","article-title":"\u201cHierarchical exploration for accelerating contextual bandits\u201d","volume-title":"Proceedings of ICML","author":"Yue","year":"2012"},{"key":"2026040314424724100_ref141","doi-asserted-by":"crossref","first-page":"1201","DOI":"10.1145\/1553374.1553527","article-title":"\u201cInteractively optimizing information retrieval systems as a dueling bandits problem\u201d","volume-title":"Proceedings of the 26th Annual International Conference on Machine Learning","author":"Yue","year":"2009"},{"key":"2026040314424724100_ref142","first-page":"241","article-title":"\u201cBeat the mean bandit\u201d","volume-title":"Proceedings of the 28th International Conference on Machine Learning (ICML-11)","author":"Yue","year":"2011"},{"key":"2026040314424724100_ref143","doi-asserted-by":"crossref","first-page":"1411","DOI":"10.1145\/2505515.2505690","article-title":"\u201cInteractive collaborative filtering\u201d","volume-title":"Proceedings of the 22nd ACM International Conference on Information & Knowledge Management","author":"Zhao","year":"2013"},{"key":"2026040314424724100_ref144","unstructured":"Zhou, L.\n          \n          2015, \u201cA survey on contextual multi-armed bandits\u201d. arXiv preprint arXiv:1508.03326."},{"key":"2026040314424724100_ref145","unstructured":"Zhou, L., and Brunskill, E.\n          2016, \u201cLatent contextual bandits and their application to personalized recommendations for new users\u201d. arXiv preprint arXiv:1604.06743."},{"key":"2026040314424724100_ref146","first-page":"307","article-title":"\u201cCopeland dueling bandits\u201d","volume-title":"Advances in Neural Information Processing Systems","author":"Zoghi","year":"2015"},{"key":"2026040314424724100_ref147","first-page":"4199","article-title":"\u201cOnline learning to rank in stochastic click models\u201d","volume-title":"International Conference on Machine Learning","author":"Zoghi","year":"2017"},{"key":"2026040314424724100_ref148","doi-asserted-by":"crossref","first-page":"73","DOI":"10.1145\/2556195.2556256","article-title":"\u201cRelative confidence sampling for efficient on-line ranker evaluation\u201d","volume-title":"Proceedings of the 7th ACM International Conference on Web Search and Data Mining","author":"Zoghi","year":"2014"},{"key":"2026040314424724100_ref149","first-page":"10","article-title":"\u201cRelative upper confidence bound for the k-armed dueling bandit problem\u201d","volume-title":"Proceedings of ICML","author":"Zoghi","year":"2014"},{"key":"2026040314424724100_ref150","article-title":"Cascading Bandits for Large-Scale Recommendation Problems","author":"Zong","year":"2016"}],"container-title":["Foundations and Trends\u00ae in Information Retrieval"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftinr\/article-pdf\/13\/4\/299\/11048218\/1500000067en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftinr\/article-pdf\/13\/4\/299\/11048218\/1500000067en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T14:32:01Z","timestamp":1777473121000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftinr\/article\/13\/4\/299\/1328680\/Bandit-Algorithms-in-Information-Retrieval"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,5,23]]},"references-count":150,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2019,5,23]]}},"URL":"https:\/\/doi.org\/10.1561\/1500000067","relation":{},"ISSN":["1554-0669","1554-0677"],"issn-type":[{"value":"1554-0669","type":"print"},{"value":"1554-0677","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,5,23]]}}}