{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,1]],"date-time":"2026-08-01T02:20:32Z","timestamp":1785550832342,"version":"3.56.0"},"reference-count":84,"publisher":"Emerald","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018,7,12]]},"abstract":"<jats:p>Thompson sampling is an algorithm for online decision problems where actions are taken sequentially in a manner that must balance between exploiting what is known to maximize immediate performance and investing to accumulate new information that may improve future performance. The algorithm addresses a broad range of problems in a computationally efficient manner and is therefore enjoying wide use. This tutorial covers the algorithm and its application, illustrating concepts through a range of examples, including Bernoulli bandit problems, shortest path problems, product recommendation, assortment, active learning with neural networks, and reinforcement learning in Markov decision processes. Most of these problems involve complex information structures, where information revealed by taking an action informs beliefs about other actions. We will also discuss when and why Thompson sampling is or is not effective and relations to alternative algorithms.<\/jats:p>","DOI":"10.1561\/2200000070","type":"journal-article","created":{"date-parts":[[2018,7,12]],"date-time":"2018-07-12T08:29:10Z","timestamp":1531384150000},"page":"1-96","source":"Crossref","is-referenced-by-count":385,"title":["A Tutorial on Thompson Sampling"],"prefix":"10.1108","volume":"11","author":[{"given":"Daniel J.","family":"Russo","sequence":"first","affiliation":[{"name":"Columbia University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Benjamin","family":"Van Roy","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Abbas","family":"Kazerouni","sequence":"additional","affiliation":[{"name":"Bentley University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ian","family":"Osband","sequence":"additional","affiliation":[{"name":"Google DeepMind"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zheng","family":"Wen","sequence":"additional","affiliation":[{"name":"Adobe Research"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2018,7,12]]},"reference":[{"key":"2026033012244353900_ref001","first-page":"2312","volume-title":"Advances in Neural Information Processing Systems 24","author":"Abbasi-Yadkori","year":"2011"},{"key":"2026033012244353900_ref002","first-page":"176","volume-title":"Proceedings of the 20th International Conference on Artificial Intel ligence and Statistics","author":"Abeille","year":"2017"},{"key":"2026033012244353900_ref003","doi-asserted-by":"crossref","first-page":"1585","DOI":"10.1145\/2505515.2514690","volume-title":"Proceedings of the 22nd ACM International Conference on Information & Knowledge Management","author":"Agarwal","year":"2013"},{"key":"2026033012244353900_ref004","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1145\/2556195.2556252","volume-title":"Proceedings of the 7th ACM international conference on Web search and data mining","author":"Agarwal","year":"2014"},{"key":"2026033012244353900_ref005","first-page":"76","volume-title":"Proceedings of the 30th Annual Conference on Learning Theory","author":"Agrawal","year":"2017"},{"key":"2026033012244353900_ref006","first-page":"39.1","volume-title":"Proceedings of the 25th Annual Conference on Learning Theory","author":"Agrawal","year":"2012"},{"key":"2026033012244353900_ref007","first-page":"99","volume-title":"Proceedings of the 16th International Conference on Artificial Intelligence and Statistics","author":"Agrawal","year":"2013"},{"key":"2026033012244353900_ref008","first-page":"127","volume-title":"Proceedings of The 30th International Conference on Machine Learning","author":"Agrawal","year":"2013"},{"issue":"2","key":"2026033012244353900_ref009","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1023\/A:1013689704352","article-title":"\u201cFinite-time analysis of the multiarmed bandit problem\u201d","volume":"47","author":"Auer","year":"2002","journal-title":"Machine Learning"},{"key":"2026033012244353900_ref010","first-page":"1646","volume-title":"Advances in Neural Information Processing Systems 26","author":"Bai","year":"2013"},{"key":"2026033012244353900_ref011","volume-title":"arXiv:1704.09011","author":"Bastani","year":"2018"},{"key":"2026033012244353900_ref012","first-page":"199","volume-title":"Advances in Neural Information Processing Systems 27","author":"Besbes","year":"2014"},{"key":"2026033012244353900_ref013","first-page":"1655","article-title":"\u201cX-armed bandits\u201d","volume":"12","author":"Bubeck","year":"2011","journal-title":"Journal of Machine Learning Research"},{"issue":"1","key":"2026033012244353900_ref014","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/2200000024","article-title":"\u201cRegret analysis of stochastic and nonstochastic multi-armed bandit problems\u201d","volume":"5","author":"Bubeck","year":"2012","journal-title":"Foundations and Trends in Machine Learning"},{"key":"2026033012244353900_ref015","first-page":"583","volume-title":"Proccedings of 29th Annual Conference on Learning Theory","author":"Bubeck","year":"2016"},{"key":"2026033012244353900_ref016","volume-title":"Discrete & Computational Geometry","author":"Bubeck","year":"2018"},{"issue":"3","key":"2026033012244353900_ref017","doi-asserted-by":"crossref","first-page":"1516","DOI":"10.1214\/13-AOS1119","article-title":"\u201cKullback-Leibler upper confidence bounds for optimal sequential allocation\u201d","volume":"41","author":"Capp\u00e9","year":"2013","journal-title":"Annals of Statistics"},{"issue":"3","key":"2026033012244353900_ref018","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1080\/00031305.1992.10475878","article-title":"\u201cExplaining the Gibbs sampler\u201d","volume":"46","author":"Casella","year":"1992","journal-title":"The American Statistician"},{"key":"2026033012244353900_ref019","first-page":"2249","volume-title":"Advances in Neural Information Processing Systems 24","author":"Chapelle","year":"2011"},{"key":"2026033012244353900_ref020","first-page":"186","volume-title":"Proceedings of the 29th International Conference on Algorithmic Learning Theory","author":"Cheng","year":"2018"},{"key":"2026033012244353900_ref021","first-page":"87","volume-title":"Proceedings of the 2008 International Conference on Web Search and Data Mining","author":"Craswell","year":"2008"},{"key":"2026033012244353900_ref022","first-page":"355","volume-title":"Proceedings of the 21st Annual Conference on Learning Theory","author":"Dani","year":"2008"},{"key":"2026033012244353900_ref023","volume-title":"arXiv:1802.01282","author":"Dimakopoulou","year":"2018"},{"key":"2026033012244353900_ref024","volume-title":"arXiv:1605.01559","author":"Durmus","year":"2016"},{"key":"2026033012244353900_ref025","volume-title":"arXiv:1410.4009","author":"Eckles","year":"2014"},{"key":"2026033012244353900_ref026","volume-title":"Working Paper","author":"Ferreira","year":"2015"},{"key":"2026033012244353900_ref027","volume-title":"preprint","author":"Francetich","year":"2017"},{"key":"2026033012244353900_ref028","volume-title":"preprint","author":"Francetich","year":"2017"},{"issue":"4","key":"2026033012244353900_ref029","doi-asserted-by":"crossref","first-page":"599","DOI":"10.1287\/ijoc.1080.0314","volume":"21","author":"Frazier","year":"2009","journal-title":"INFORMS Journal on Com\u00acputing"},{"issue":"5","key":"2026033012244353900_ref030","doi-asserted-by":"crossref","first-page":"2410","DOI":"10.1137\/070693424","article-title":"\u201cA knowledge-gradient policy for sequential information collection\u201d","volume":"47","author":"Frazier","year":"2008","journal-title":"SIAM Journal on Control and Optimization"},{"issue":"5-6","key":"2026033012244353900_ref031","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1561\/2200000049","article-title":"\u201cBayesian reinforcement learning: A survey\u201d","volume":"8","author":"Ghavamzadeh","year":"2015","journal-title":"Foundations and Trends in Ma\u00acchine Learning"},{"issue":"3","key":"2026033012244353900_ref032","doi-asserted-by":"crossref","first-page":"561","DOI":"10.1093\/biomet\/66.3.561","article-title":"\u201cA dynamic allocation index for the discounted multiarmed bandit problem\u201d","volume":"66","author":"Gittins","year":"1979","journal-title":"Biometrika"},{"key":"2026033012244353900_ref033","doi-asserted-by":"crossref","DOI":"10.1002\/9780470980033","volume-title":"Multi-armed bandit al location indices","author":"Gittins","year":"2011"},{"key":"2026033012244353900_ref034","volume-title":"arXiv:1605.05697v1","author":"Gomez-Uribe","year":"2016"},{"key":"2026033012244353900_ref035","first-page":"100","volume-title":"Proceedings of the 31st International Conference on Machine Learning","author":"Gopalan","year":"2014"},{"key":"2026033012244353900_ref036","first-page":"861","volume-title":"Proceedings of the 24th Annual Conference on Learning Theory","author":"Gopalan","year":"2015"},{"key":"2026033012244353900_ref037","first-page":"13","volume-title":"Proceedings of the 27th International Conference on Machine Learning","author":"Graepel","year":"2010"},{"key":"2026033012244353900_ref038","doi-asserted-by":"crossref","first-page":"1813","DOI":"10.1145\/3097983.3098184","volume-title":"Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Hill","year":"2017"},{"key":"2026033012244353900_ref039","first-page":"375","volume-title":"Proceedings of the 17th International Conference on Artificial Intel ligence and Statistics","author":"Honda","year":"2014"},{"key":"2026033012244353900_ref040","first-page":"1563","article-title":"\u201cNear-optimal regret bounds for reinforcement learning\u201d","volume":"11","author":"Jaksch","year":"2010","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012244353900_ref041","volume-title":"To appear in proceedings of the 22nd International Conference on Artificial Intel ligence and Statistics","author":"Kandasamy","year":"2018"},{"issue":"2","key":"2026033012244353900_ref042","doi-asserted-by":"crossref","first-page":"262","DOI":"10.1287\/moor.12.2.262","article-title":"\u201cThe multi-armed bandit problem: decomposition and computation\u201d","volume":"12","author":"Katehakis","year":"1987","journal-title":"Mathematics of Opera\u00actions Research"},{"key":"2026033012244353900_ref043","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1007\/978-3-642-34106-9_18","volume-title":"Proceedings of the 24th International Conference on Algorithmic Learning Theory","author":"Kauffmann","year":"2012"},{"key":"2026033012244353900_ref044","first-page":"592","volume-title":"Proceedings of the 15th International Conference on Artificial Intel ligence and Statistics","author":"Kaufmann","year":"2012"},{"key":"2026033012244353900_ref045","first-page":"1297","volume-title":"Advances in Neural Information Processing Systems 28","author":"Kawale","year":"2015"},{"issue":"12","key":"2026033012244353900_ref046","doi-asserted-by":"crossref","first-page":"6415","DOI":"10.1109\/TAC.2017.2653942","article-title":"\u201cThompson sampling for stochastic control: the finite parameter case\u201d","volume":"62","author":"Kim","year":"2017","journal-title":"IEEE Transactions on Automatic Control"},{"key":"2026033012244353900_ref047","first-page":"681","volume-title":"Proceedings of the 40th ACM Symposium on Theory of Computing","author":"Kleinberg","year":"2008"},{"key":"2026033012244353900_ref048","first-page":"767","volume-title":"Proceedings of the 32nd International Conference on Machine Learning","author":"Kveton","year":"2015"},{"issue":"1","key":"2026033012244353900_ref049","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1016\/0196-8858(85)90002-8","article-title":"\u201cAsymptotically efficient adaptive alloca\u00action rules\u201d","volume":"6","author":"Lai","year":"1985","journal-title":"Advances in applied mathematics"},{"key":"2026033012244353900_ref050","doi-asserted-by":"crossref","first-page":"661","DOI":"10.1145\/1772690.1772758","volume-title":"Proceedings of the 19th International Conference on World Wide Web","author":"Li","year":"2010"},{"issue":"7553","key":"2026033012244353900_ref051","doi-asserted-by":"crossref","first-page":"445","DOI":"10.1038\/nature14540","article-title":"\u201cReinforcement learning improves behaviour from evaluative feedback\u201d","volume":"521","author":"Littman","year":"2015","journal-title":"Nature"},{"key":"2026033012244353900_ref052","volume-title":"arXiv:1711.03198","author":"Liu","year":"2017"},{"key":"2026033012244353900_ref053","first-page":"3258","volume-title":"Advances in Neural Information Processing Systems 30","author":"Lu","year":"2017"},{"issue":"2","key":"2026033012244353900_ref054","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1016\/S0304-4149(02)00150-3","article-title":"\u201cErgodicity for SDEs and approximations: locally Lipschitz vector fields and degenerate noise\u201d","volume":"101","author":"Mattingly","year":"2002","journal-title":"Stochastic processes and their applications"},{"key":"2026033012244353900_ref055","first-page":"3003","volume-title":"Advances in Neural Information Processing Systems 26","author":"Osband","year":"2013"},{"key":"2026033012244353900_ref056","first-page":"4026","volume-title":"Advances in Neural Information Processing Systems 29","author":"Osband","year":"2016"},{"key":"2026033012244353900_ref057","volume-title":"arXiv:1703.07608","author":"Osband","year":"2017"},{"key":"2026033012244353900_ref058","first-page":"1466","volume-title":"Advances in Neural Information Processing Systems 27","author":"Osband","year":"2014"},{"key":"2026033012244353900_ref059","first-page":"604","volume-title":"Advances in Neural Information Processing Systems 27","author":"Osband","year":"2014"},{"key":"2026033012244353900_ref060","volume-title":"Proceedings of The Multi\u00acdisciplinary Conference on Reinforcement Learning and Decision Making","author":"Osband","year":"2017"},{"key":"2026033012244353900_ref061","first-page":"2701","volume-title":"Proceedings of the 34th International Conference on Machine Learning","author":"Osband","year":"2017"},{"key":"2026033012244353900_ref062","first-page":"2377","volume-title":"Proceedings of The 33rd International Conference on Machine Learning","author":"Osband","year":"2016"},{"key":"2026033012244353900_ref063","first-page":"1333-1342","volume-title":"Advances in Neural Information Processing Systems 30","author":"Ouyang","year":"2017"},{"issue":"1","key":"2026033012244353900_ref064","doi-asserted-by":"crossref","first-page":"255","DOI":"10.1111\/1467-9868.00123","article-title":"\u201cOptimal scaling of discrete approximations to Langevin diffusions\u201d","volume":"60","author":"Roberts","year":"1998","journal-title":"Journal of the Royal Statistical Society: Series B (Statistical Methodology)"},{"key":"2026033012244353900_ref065","first-page":"341","volume-title":"Bernoul li","author":"Roberts","year":"1996"},{"issue":"2","key":"2026033012244353900_ref066","doi-asserted-by":"crossref","first-page":"395","DOI":"10.1287\/moor.1100.0446","article-title":"\u201cLinearly parameterized bandits\u201d","volume":"35","author":"Rusmevichientong","year":"2010","journal-title":"Mathematics of Operations Research"},{"key":"2026033012244353900_ref067","first-page":"2256","volume-title":"Advances in Neural Information Processing Systems 26","author":"Russo","year":"2013"},{"key":"2026033012244353900_ref068","first-page":"1583","volume-title":"Advances in Neural Information Processing Systems 27","author":"Russo","year":"2014"},{"issue":"4","key":"2026033012244353900_ref069","doi-asserted-by":"crossref","first-page":"1221","DOI":"10.1287\/moor.2014.0650","article-title":"\u201cLearning to optimize via posterior sampling\u201d","volume":"39","author":"Russo","year":"2014","journal-title":"Mathematics of Operations Research"},{"issue":"68","key":"2026033012244353900_ref070","first-page":"1","article-title":"\u201cAn Information-Theoretic analysis of Thompson sampling\u201d","volume":"17","author":"Russo","year":"2016","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012244353900_ref071","first-page":"1417","volume-title":"Conference on Learning Theory","author":"Russo","year":"2016"},{"issue":"1","key":"2026033012244353900_ref072","doi-asserted-by":"crossref","first-page":"230","DOI":"10.1287\/opre.2017.1663","article-title":"\u201cLearning to optimize via information- directed sampling\u201d","volume":"66","author":"Russo","year":"2018","journal-title":"Operations Research"},{"key":"2026033012244353900_ref073","volume-title":"arXiv:1803.02855","author":"Russo","year":"2018"},{"issue":"4","key":"2026033012244353900_ref074","doi-asserted-by":"crossref","first-page":"500","DOI":"10.1287\/mksc.2016.1023","article-title":"\u201cCustomer acquisition via display advertising using multi-armed bandit experiments\u201d","volume":"36","author":"Schwartz","year":"2017","journal-title":"Marketing Science"},{"issue":"6","key":"2026033012244353900_ref075","doi-asserted-by":"crossref","first-page":"639","DOI":"10.1002\/asmb.874","article-title":"\u201cA modern Bayesian look at the multi-armed bandit\u201d","volume":"26","author":"Scott","year":"2010","journal-title":"Applied Stochastic Models in Business and Industry"},{"issue":"1","key":"2026033012244353900_ref076","article-title":"\u201cMulti-armed bandit experiments in the online service economy\u201d","volume":"31","author":"Scott","year":"2015","journal-title":"Applied Stochastic Models in Business and Industry"},{"issue":"5","key":"2026033012244353900_ref077","doi-asserted-by":"crossref","first-page":"3250","DOI":"10.1109\/TIT.2011.2182033","article-title":"\u201cInformation- Theoretic regret bounds for Gaussian process optimization in the bandit setting\u201d","volume":"58","author":"Srinivas","year":"2012","journal-title":"IEEE Transactions on Information Theory"},{"key":"2026033012244353900_ref078","first-page":"943","volume-title":"Proceedings of the 17th International Conference on Machine Learning","author":"Strens","year":"2000"},{"key":"2026033012244353900_ref079","volume-title":"Reinforcement learning: An introduction","author":"Sutton","year":"1998"},{"issue":"7","key":"2026033012244353900_ref080","first-page":"1","article-title":"\u201cConsistency and fluctuations for stochastic gradient Langevin dynamics\u201d","volume":"17","author":"Teh","year":"2016","journal-title":"Journal of Machine Learning Research"},{"issue":"2","key":"2026033012244353900_ref081","doi-asserted-by":"crossref","first-page":"450","DOI":"10.2307\/2371219","article-title":"\u201cOn the theory of apportionment\u201d","volume":"57","author":"Thompson","year":"1935","journal-title":"American Journal of Mathematics"},{"issue":"3\/4","key":"2026033012244353900_ref082","doi-asserted-by":"crossref","first-page":"285","DOI":"10.2307\/2332286","article-title":"\u201cOn the likelihood that one unknown probability exceeds another in view of the evidence of two samples\u201d","volume":"25","author":"Thompson","year":"1933","journal-title":"Biometrika"},{"key":"2026033012244353900_ref083","first-page":"681","volume-title":"Proceedings of the 28th International Conference on Machine Learning","author":"Welling","year":"2011"},{"key":"2026033012244353900_ref084","volume-title":"PhD thesis","author":"Wyatt","year":"1997"}],"container-title":["Foundations and Trends\u00ae in Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftmal\/article-pdf\/11\/1\/1\/11155609\/2200000070en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftmal\/article-pdf\/11\/1\/1\/11155609\/2200000070en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T18:10:49Z","timestamp":1777486249000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftmal\/article\/11\/1\/1\/1332412\/A-Tutorial-on-Thompson-Sampling"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,7,12]]},"references-count":84,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,7,12]]}},"URL":"https:\/\/doi.org\/10.1561\/2200000070","relation":{},"ISSN":["1935-8237","1935-8245"],"issn-type":[{"value":"1935-8237","type":"print"},{"value":"1935-8245","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,7,12]]}}}