{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,3]],"date-time":"2026-05-03T02:31:30Z","timestamp":1777775490264,"version":"3.51.4"},"reference-count":102,"publisher":"Emerald","issue":"6","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023,7,11]]},"abstract":"<jats:p>Reinforcement learning agents have demonstrated remarkable achievements in simulated environments. Data efficiency poses an impediment to carrying this success over to real environments. The design of data-efficient agents calls for a deeper understanding of information acquisition and representation. We discuss concepts and regret analysis that together offer principled guidance. This line of thinking sheds light on questions of what information to seek, how to seek that information, and what information to retain. To illustrate concepts, we design simple agents that build on them and present computational results that highlight data efficiency.<\/jats:p>","DOI":"10.1561\/2200000097","type":"journal-article","created":{"date-parts":[[2023,7,11]],"date-time":"2023-07-11T05:24:23Z","timestamp":1689053063000},"page":"733-865","source":"Crossref","is-referenced-by-count":12,"title":["Reinforcement Learning, Bit by Bit"],"prefix":"10.1108","volume":"16","author":[{"given":"Xiuyuan","family":"Lu","sequence":"first","affiliation":[{"name":"DeepMind","place":["USA"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Benjamin","family":"Van Roy","sequence":"additional","affiliation":[{"name":"DeepMind","place":["USA"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vikranth","family":"Dwaracherla","sequence":"additional","affiliation":[{"name":"DeepMind","place":["USA"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Morteza","family":"Ibrahimi","sequence":"additional","affiliation":[{"name":"DeepMind","place":["USA"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ian","family":"Osband","sequence":"additional","affiliation":[{"name":"DeepMind","place":["USA"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zheng","family":"Wen","sequence":"additional","affiliation":[{"name":"DeepMind","place":["USA"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","published-online":{"date-parts":[[2023,7,11]]},"reference":[{"key":"2026033012314915100_ref001","first-page":"1184","volume-title":"Advances in Neural Information Processing Systems","author":"Agrawal","year":"2017"},{"issue":"1","key":"2026033012314915100_ref002","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1214\/aoms\/1177704255","article-title":"A definition of subjective probability","volume":"34","author":"Anscombe","year":"1963","journal-title":"Annals of mathematical statistics"},{"key":"2026033012314915100_ref003","volume-title":"Deciding What to Learn: A Rate-Distortion Approach","author":"Arumugam","year":"2021"},{"issue":"2","key":"2026033012314915100_ref004","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1023\/A:1013689704352","article-title":"Finite-time analysis of the multiarmed bandit problem","volume":"47","author":"Auer","year":"2002","journal-title":"Machine learning"},{"key":"2026033012314915100_ref005","first-page":"263","volume-title":"Minimax Regret Bounds for Reinforcement Learning","author":"Azar","year":"2017"},{"issue":"5","key":"2026033012314915100_ref006","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TSMC.1983.6313077","article-title":"Neuronlike adaptive elements that can solve difficult learning control problems","volume":"SMC-13","author":"Barto","year":"1983","journal-title":"IEEE transactions on systems, man, and cybernetics"},{"key":"2026033012314915100_ref007","volume-title":"Reinforcement learning and optimal control","author":"Bertsekas","year":"2019"},{"key":"2026033012314915100_ref008","volume-title":"Neuro-dynamic programming","author":"Bertsekas","year":"1996"},{"key":"2026033012314915100_ref009","first-page":"213","article-title":"R-Max - a General Polynomial Time Algorithm for near-Optimal Reinforcement Learning","volume":"3","author":"Brafman","year":"2003","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012314915100_ref010","doi-asserted-by":"crossref","DOI":"10.1561\/9781601986276","article-title":"Regret analysis of stochastic and nonstochastic multi-armed bandit problems","volume-title":"arXiv preprint","author":"Bubeck","year":"2012"},{"key":"2026033012314915100_ref011","first-page":"266","volume-title":"Bandit Convex Optimization: T Regret in One Dimension","author":"Bubeck","year":"2015"},{"key":"2026033012314915100_ref012","first-page":"583","volume-title":"Multi-scale exploration of convex functions and bandit convex optimization","author":"Bubeck","year":"2016"},{"key":"2026033012314915100_ref013","first-page":"196","volume-title":"First-Order Bayesian Regret Analysis of Thompson Sampling","author":"Bubeck","year":"2020"},{"key":"2026033012314915100_ref014","volume-title":"Large-Scale Study of Curiosity-Driven Learning","author":"Burda","year":"2019"},{"issue":"1","key":"2026033012314915100_ref015","doi-asserted-by":"crossref","first-page":"126","DOI":"10.1287\/opre.1040.0145","article-title":"An adaptive sampling algorithm for solving Markov decision processes","volume":"53","author":"Chang","year":"2005","journal-title":"Operations Research"},{"key":"2026033012314915100_ref016","first-page":"72","volume-title":"Efficient selectivity and backup operators in Monte-Carlo tree search","author":"Coulom","year":"2006"},{"key":"2026033012314915100_ref017","volume-title":"Elements of Information Theory","author":"Cover","year":"2006"},{"key":"2026033012314915100_ref018","first-page":"213","volume-title":"Q-learning for history-based reinforcement learning","author":"Daswani","year":"2013"},{"key":"2026033012314915100_ref019","volume-title":"Feature reinforcement learning: state of the art","author":"Daswani","year":"2014"},{"key":"2026033012314915100_ref020","volume-title":"A Bit Better? Quantifying Information for Bandit Learning","author":"Devraj","year":"2021"},{"key":"2026033012314915100_ref021","first-page":"1158","volume-title":"On the performance of Thompson sampling on logistic bandits","author":"Dong","year":"2019"},{"key":"2026033012314915100_ref022","first-page":"4157","volume-title":"An Information-Theoretic Analysis for Thompson Sampling with Many Actions","author":"Dong","year":"2018"},{"issue":"9-10","key":"2026033012314915100_ref023","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1016\/S0898-1221(00)00089-4","article-title":"Some upper bounds for relative entropy and applications","volume":"39","author":"Dragomir","year":"2000","journal-title":"Computers & Mathematics with Applications"},{"key":"2026033012314915100_ref024","author":"Duff","year":"2003"},{"key":"2026033012314915100_ref025","volume-title":"Hypermodels for Exploration","author":"Dwaracherla","year":"2020"},{"key":"2026033012314915100_ref026","volume-title":"Langevin DQN","author":"Dwaracherla","year":"2021"},{"key":"2026033012314915100_ref027","volume-title":"Bayesian adaptive data analysis guarantees from subgaussianity","author":"Elder","year":"2016"},{"key":"2026033012314915100_ref028","first-page":"154","volume-title":"Bayes meets Bellman: The Gaussian process approach to temporal difference learning","author":"Engel","year":"2003"},{"key":"2026033012314915100_ref029","first-page":"201","volume-title":"Reinforcement learning with Gaussian processes","author":"Engel","year":"2005"},{"key":"2026033012314915100_ref030","article-title":"Temporal Difference Uncertainties as a Signal for Exploration","volume-title":"arXiv preprint","author":"Flennerhag","year":"2020"},{"issue":"10","key":"2026033012314915100_ref031","doi-asserted-by":"crossref","first-page":"3650","DOI":"10.1257\/aer.20181897","article-title":"Quantifying information and uncertainty","volume":"109","author":"Frankel","year":"2019","journal-title":"American Economic Review"},{"issue":"5\u20136","key":"2026033012314915100_ref032","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1561\/2200000049","article-title":"Bayesian Reinforcement Learning: A Survey","volume":"8","author":"Ghavamzadeh","year":"2015","journal-title":"Foundations and Trends\u00ae in Machine Learning"},{"key":"2026033012314915100_ref033","first-page":"241","volume-title":"A dynamic allocation index for the sequential design of experiments","author":"Gittins","year":"1974"},{"issue":"3","key":"2026033012314915100_ref034","doi-asserted-by":"crossref","first-page":"561","DOI":"10.1093\/biomet\/66.3.561","article-title":"A dynamic allocation index for the discounted multiarmed bandit problem","volume":"66","author":"Gittins","year":"1979","journal-title":"Biometrika"},{"key":"2026033012314915100_ref035","volume-title":"The Value Equivalence Principle for Model-Based Reinforcement Learning","author":"Grimm","year":"2021"},{"issue":"1","key":"2026033012314915100_ref036","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1109\/TSSC.1966.300074","article-title":"Information value theory","volume":"2","author":"Howard","year":"1966","journal-title":"IEEE Transactions on systems science and cybernetics"},{"key":"2026033012314915100_ref037","first-page":"227","volume-title":"Universal Algorithmic Intelligence: A Mathematical Top-Down Approach","author":"Hutter","year":"2007"},{"issue":"51","key":"2026033012314915100_ref038","first-page":"1563","article-title":"Near-optimal Regret Bounds for Reinforcement Learning","volume":"11","author":"Jaksch","year":"2010","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012314915100_ref039","unstructured":"Jha, A.\n           (2016). \u201cWithout Claude Shannon\u2019s information theory there would have been no internet\u201d. URL: https:\/\/www.theguardian.com\/science\/2014\/jun\/22\/shannon-information-theory."},{"key":"2026033012314915100_ref040","first-page":"1704","volume-title":"Contextual Decision Processes with low Bellman rank are PAC-Learnable","author":"Jiang","year":"2017"},{"key":"2026033012314915100_ref041","first-page":"4863","volume-title":"Advances in Neural Information Processing Systems","author":"Jin","year":"2018"},{"key":"2026033012314915100_ref042","first-page":"2137","volume-title":"Provably efficient reinforcement learning with linear function approximation","author":"Jin","year":"2020"},{"key":"2026033012314915100_ref043","first-page":"267","volume-title":"Approximately Optimal Approximate Reinforcement Learning","author":"Kakade","year":"2002"},{"key":"2026033012314915100_ref044","doi-asserted-by":"crossref","first-page":"209","DOI":"10.1023\/A:1017984413808","article-title":"Near-Optimal Reinforcement Learning in Polynomial Time","volume":"49","author":"Kearns","year":"2002","journal-title":"Machine Learning"},{"key":"2026033012314915100_ref045","volume-title":"Adam: A Method for Stochastic Optimization","author":"Kingma","year":"2015"},{"key":"2026033012314915100_ref046","author":"Kirschner","year":"2021"},{"key":"2026033012314915100_ref047","volume-title":"Information directed sampling for linear partial monitoring","author":"Kirschner","year":"2020"},{"key":"2026033012314915100_ref048","volume-title":"Asymptotically Optimal Information-Directed Sampling","author":"Kirschner","year":"2020"},{"key":"2026033012314915100_ref049","volume-title":"The Hedonistic Neuron: A Theory of Memory, Learning, and Intelligence","author":"Klopf","year":"1982"},{"key":"2026033012314915100_ref050","first-page":"282","volume-title":"Bandit based Monte-Carlo planning","author":"Kocsis","year":"2006"},{"key":"2026033012314915100_ref051","first-page":"1091","article-title":"Adaptive treatment allocation and the multi-armed bandit problem","volume-title":"The Annals of Statistics","author":"Lai","year":"1987"},{"issue":"1","key":"2026033012314915100_ref052","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1016\/0196-8858(85)90002-8","article-title":"Asymptotically efficient adaptive allocation rules","volume":"6","author":"Lai","year":"1985","journal-title":"Advances in applied mathematics"},{"key":"2026033012314915100_ref053","volume-title":"Mirror Descent and the Information Ratio","author":"Lattimore","year":"2020"},{"key":"2026033012314915100_ref054","first-page":"2111","volume-title":"An information-theoretic approach to minimax regret in partial monitoring","author":"Lattimore","year":"2019"},{"key":"2026033012314915100_ref055","first-page":"1555","volume-title":"Predictive Representations of State","author":"Littman","year":"2002"},{"key":"2026033012314915100_ref056","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v32i1.11751","volume-title":"Information directed sampling for stochastic bandits with graph feedback","author":"Liu","year":"2018"},{"key":"2026033012314915100_ref057","author":"Lu","year":"2020"},{"key":"2026033012314915100_ref058","first-page":"2461","volume-title":"Information-Theoretic Confidence Bounds for Reinforcement Learning","author":"Lu","year":"2019"},{"key":"2026033012314915100_ref059","first-page":"387","volume-title":"Instance-based utile distinctions for reinforcement learning with hidden state","author":"McCallum","year":"1995"},{"key":"2026033012314915100_ref060","volume-title":"Playing Atari With Deep Reinforcement Learning","author":"Mnih","year":"2013"},{"issue":"7540","key":"2026033012314915100_ref061","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"2026033012314915100_ref062","volume-title":"Information-Directed Exploration for Deep Reinforcement Learning","author":"Nikolov","year":"2019"},{"key":"2026033012314915100_ref063","article-title":"Variational Bayesian reinforcement learning with regret bounds","volume-title":"arXiv preprint","author":"O\u2019Donoghue","year":"2018"},{"key":"2026033012314915100_ref064","first-page":"3836","volume-title":"The uncertainty Bellman equation and exploration","author":"O\u2019Donoghue","year":"2018"},{"key":"2026033012314915100_ref065","first-page":"8617","volume-title":"Randomized prior functions for deep reinforcement learning","author":"Osband","year":"2018"},{"key":"2026033012314915100_ref066","first-page":"4026","volume-title":"Deep exploration via bootstrapped DQN","author":"Osband","year":"2016"},{"key":"2026033012314915100_ref067","volume-title":"Behaviour Suite for Reinforcement Learning","author":"Osband","year":"2020"},{"key":"2026033012314915100_ref068","first-page":"3003","volume-title":"(More) Efficient Reinforcement Learning via Posterior Sampling","author":"Osband","year":"2013"},{"key":"2026033012314915100_ref069","first-page":"2701","volume-title":"Why is Posterior Sampling Better than Optimism for Reinforcement Learning","author":"Osband","year":"2017"},{"issue":"124","key":"2026033012314915100_ref070","first-page":"1","article-title":"Deep Exploration via Randomized Value Functions","volume":"20","author":"Osband","year":"2019","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012314915100_ref071","volume-title":"Neural networks with late-phase weights","author":"Oswald","year":"2021"},{"key":"2026033012314915100_ref072","first-page":"1333","volume-title":"Learning Unknown Markov Decision Processes: A Thompson Sampling Approach","author":"Ouyang","year":"2017"},{"key":"2026033012314915100_ref073","doi-asserted-by":"crossref","DOI":"10.1002\/9781118309858","volume-title":"Optimal learning","author":"Powell","year":"2012"},{"key":"2026033012314915100_ref074","volume-title":"A Note on Reinforcement Learning, Bit by Bit","author":"Qin","year":"2023"},{"key":"2026033012314915100_ref075","first-page":"1583","article-title":"Learning to optimize via information-directed sampling","volume":"27","author":"Russo","year":"2014","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"4","key":"2026033012314915100_ref076","doi-asserted-by":"crossref","first-page":"1221","DOI":"10.1287\/moor.2014.0650","article-title":"Learning to optimize via posterior sampling","volume":"39","author":"Russo","year":"2014","journal-title":"Mathematics of Operations Research"},{"issue":"1","key":"2026033012314915100_ref077","first-page":"2442","article-title":"An information-theoretic analysis of Thompson sampling","volume":"17","author":"Russo","year":"2016","journal-title":"The Journal of Machine Learning Research"},{"issue":"1","key":"2026033012314915100_ref078","doi-asserted-by":"crossref","first-page":"230","DOI":"10.1287\/opre.2017.1663","article-title":"Learning to optimize via information-directed sampling","volume":"66","author":"Russo","year":"2018","journal-title":"Operations Research"},{"key":"2026033012314915100_ref079","volume-title":"Satisficing in Time-Sensitive Bandit Learning","author":"Russo","year":"2020"},{"issue":"1","key":"2026033012314915100_ref080","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/2200000070","article-title":"A Tutorial on Thompson Sampling","volume":"11","author":"Russo","year":"2018","journal-title":"Foundations and Trends\u00ae in Machine Learning"},{"issue":"1","key":"2026033012314915100_ref081","doi-asserted-by":"crossref","first-page":"180","DOI":"10.1287\/opre.1110.0999","article-title":"The Knowledge Gradient Algorithm for a General Class of Online Learning Problems","volume":"60","author":"Ryzhov","year":"2012","journal-title":"Operations Research"},{"key":"2026033012314915100_ref082","first-page":"1312","volume-title":"Universal Value Function Approximators","author":"Schaul","year":"2015"},{"issue":"7839","key":"2026033012314915100_ref083","doi-asserted-by":"crossref","first-page":"604609","DOI":"10.1038\/s41586-020-03051-4","article-title":"Mastering Atari, go, chess and shogi by planning with a learned model","volume":"588","author":"Schrittwieser","year":"2020","journal-title":"Nature"},{"key":"2026033012314915100_ref084","first-page":"943","volume-title":"A Bayesian Framework for Reinforcement Learning","author":"Strens","year":"2000"},{"key":"2026033012314915100_ref085","author":"Sutton","year":"1984"},{"issue":"1","key":"2026033012314915100_ref086","doi-asserted-by":"crossref","first-page":"9","DOI":"10.1023\/A:1022633531479","article-title":"Learning to predict by the methods of temporal differences","volume":"3","author":"Sutton","year":"1988","journal-title":"Machine learning"},{"key":"2026033012314915100_ref087","volume-title":"Gain adaptation beats least squares","author":"Sutton","year":"1992"},{"key":"2026033012314915100_ref088","volume-title":"Reinforcement learning: An introduction","author":"Sutton","year":"2018"},{"key":"2026033012314915100_ref089","first-page":"761","volume-title":"Horde: A scalable real-time architecture for learning knowledge from unsupervised sensorimotor interaction","author":"Sutton","year":"2011"},{"key":"2026033012314915100_ref090","first-page":"2753","volume-title":"#Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning","author":"Tang","year":"2017"},{"issue":"3","key":"2026033012314915100_ref091","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1023\/A:1022624705476","article-title":"Practical issues in temporal difference learning","volume":"8","author":"Tesauro","year":"1992","journal-title":"Machine learning"},{"issue":"2","key":"2026033012314915100_ref092","doi-asserted-by":"crossref","first-page":"215","DOI":"10.1162\/neco.1994.6.2.215","article-title":"TD-Gammon, a self-teaching backgammon program, achieves master-level play","volume":"6","author":"Tesauro","year":"1994","journal-title":"Neural computation"},{"issue":"3\/4","key":"2026033012314915100_ref093","doi-asserted-by":"crossref","first-page":"285","DOI":"10.2307\/2332286","article-title":"On the likelihood that one unknown probability exceeds another in view of the evidence of two samples","volume":"25","author":"Thompson","year":"1933","journal-title":"Biometrika"},{"issue":"2","key":"2026033012314915100_ref094","doi-asserted-by":"crossref","first-page":"450","DOI":"10.2307\/2371219","article-title":"On the theory of apportionment","volume":"57","author":"Thompson","year":"1935","journal-title":"American Journal of Mathematics"},{"key":"2026033012314915100_ref095","first-page":"5392","volume-title":"Hybrid Reward Architecture for Reinforcement Learning","author":"Van Seijen","year":"2017"},{"key":"2026033012314915100_ref096","first-page":"9310","volume-title":"Discovery of Useful Questions as Auxiliary Tasks","author":"Veeriah","year":"2019"},{"key":"2026033012314915100_ref097","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1007\/978-3-642-27645-3_11","article-title":"Bayesian reinforcement learning","volume-title":"Reinforcement learning: State of the Art","author":"Vlassis","year":"2012"},{"key":"2026033012314915100_ref098","volume-title":"Learning from delayed rewards","author":"Watkins","year":"1989"},{"key":"2026033012314915100_ref099","volume-title":"Bayesian Learning via Stochastic Gradient Langevin Dynamics","author":"Welling","year":"2011"},{"key":"2026033012314915100_ref100","volume-title":"Learning to Control","author":"Witten","year":"1976"},{"issue":"5","key":"2026033012314915100_ref101","doi-asserted-by":"crossref","first-page":"286","DOI":"10.1016\/S0019-9958(77)90354-0","article-title":"An adaptive optimal controller for discrete-time Markov environments","volume":"34","author":"Witten","year":"1977","journal-title":"Information and Control"},{"key":"2026033012314915100_ref102","article-title":"Connections between mirror descent, Thompson sampling and the information ratio","volume-title":"arXiv preprint","author":"Zimmert","year":"2019"}],"container-title":["Foundations and Trends\u00ae in Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftmal\/article-pdf\/16\/6\/733\/11155991\/2200000097en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftmal\/article-pdf\/16\/6\/733\/11155991\/2200000097en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T18:10:51Z","timestamp":1777486251000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftmal\/article\/16\/6\/733\/1332425\/Reinforcement-Learning-Bit-by-Bit"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,11]]},"references-count":102,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,7,11]]}},"URL":"https:\/\/doi.org\/10.1561\/2200000097","relation":{},"ISSN":["1935-8237","1935-8245"],"issn-type":[{"value":"1935-8237","type":"print"},{"value":"1935-8245","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,11]]}}}