{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,13]],"date-time":"2026-06-13T18:21:35Z","timestamp":1781374895790,"version":"3.54.1"},"reference-count":45,"publisher":"MIT Press","issue":"11","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,10,11]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Active inference is a theory of perception, learning, and decision making that can be applied to neuroscience, robotics, psychology, and machine learning. Recently, intensive research has been taking place to scale up this framework using Monte Carlo tree search and deep learning. The goal of this activity is to solve more complicated tasks using deep active inference. First, we review the existing literature and then progressively build a deep active inference agent as follows: we (1) implement a variational autoencoder (VAE), (2) implement a deep hidden Markov model (HMM), and (3) implement a deep critical hidden Markov model (CHMM). For the CHMM, we implemented two versions, one minimizing expected free energy, CHMM[EFE] and one maximizing rewards, CHMM[reward]. Then we experimented with three different action selection strategies: the \u03b5-greedy algorithm as well as softmax and best action selection. According to our experiments, the models able to solve the dSprites environment are the ones that maximize rewards. On further inspection, we found that the CHMM minimizing expected free energy almost always picks the same action, which makes it unable to solve the dSprites environment. In contrast, the CHMM maximizing reward keeps on selecting all the actions, enabling it to successfully solve the task. The only difference between those two CHMMs is the epistemic value, which aims to make the outputs of the transition and encoder networks as close as possible. Thus, the CHMM minimizing expected free energy repeatedly picks a single action and becomes an expert at predicting the future when selecting this action. This effectively makes the KL divergence between the output of the transition and encoder networks small. Additionally, when selecting the action down the average reward is zero, while for all the other actions, the expected reward will be negative. Therefore, if the CHMM has to stick to a single action to keep the KL divergence small, then the action down is the most rewarding. We also show in simulation that the epistemic value used in deep active inference can behave degenerately and in certain circumstances effectively lose, rather than gain, information. As the agent minimizing EFE is not able to explore its environment, the appropriate formulation of the epistemic value in deep active inference remains an open question.<\/jats:p>","DOI":"10.1162\/neco_a_01697","type":"journal-article","created":{"date-parts":[[2024,8,14]],"date-time":"2024-08-14T18:52:40Z","timestamp":1723661560000},"page":"2403-2445","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":4,"title":["Deconstructing Deep Active Inference: A Contrarian Information Gatherer"],"prefix":"10.1162","volume":"36","author":[{"given":"Th\u00e9ophile","family":"Champion","sequence":"first","affiliation":[{"name":"University of Birmingham, School of Computer Science Birmingham B15 2TT, U.K. txc314@student.bham.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marek","family":"Grze\u015b","sequence":"additional","affiliation":[{"name":"University of Kent, School of Computing Canterbury CT2 7NZ, U.K. m.grzes@kent.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lisa","family":"Bonheme","sequence":"additional","affiliation":[{"name":"University of Kent, School of Computing Canterbury CT2 7NZ, U.K. lb732@kent.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Howard","family":"Bowman","sequence":"additional","affiliation":[{"name":"University of Birmingham, School of Psychology and Computer Science, Birmingham B15 2TT, U.K."},{"name":"University College London, Wellcome Centre for Human Neuroimaging (honorary) London WC1N 3AR, U.K. h.bowman@bham.ac.uk"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,10,11]]},"reference":[{"issue":"8","key":"2024100916000914600_bib1","doi-asserted-by":"publisher","first-page":"716","DOI":"10.1073\/pnas.38.8.716","article-title":"On the theory of dynamic programming","volume":"38","author":"Bellman","year":"1952","journal-title":"Proceedings of the National Academy of Sciences USA"},{"issue":"1","key":"2024100916000914600_bib2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TCIAIG.2012.2186810","article-title":"A survey of Monte Carlo tree search methods","volume":"4","author":"Browne","year":"2012","journal-title":"IEEE Transactions on Computational Intelligence and AI in Games"},{"key":"2024100916000914600_bib3","doi-asserted-by":"crossref","DOI":"10.3389\/fncom.2020.574372","article-title":"Learning generative state space models for active inference","volume":"14","author":"\u00c7atal","year":"2020","journal-title":"Frontiers in Computational Neuroscience"},{"key":"2024100916000914600_bib4","doi-asserted-by":"publisher","first-page":"450","DOI":"10.1016\/j.neunet.2022.05.010","article-title":"Branching time active inference: Empirical study and complexity class analysis","volume":"152","author":"Champion","year":"2022","journal-title":"Neural Networks"},{"key":"2024100916000914600_bib5","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1016\/j.neunet.2022.03.036","article-title":"Branching time active inference: The theory and its generality","volume":"151","author":"Champion","year":"2022","journal-title":"Neural Networks"},{"key":"2024100916000914600_bib6","author":"Champion","year":"2023","journal-title":"Deconstructing deep active inference"},{"issue":"10","key":"2024100916000914600_bib7","doi-asserted-by":"publisher","first-page":"2762","DOI":"10.1162\/neco_a_01422","article-title":"Realizing active inference in variational message passing: The outcome-blind certainty seeker","volume":"33","author":"Champion","year":"2021","journal-title":"Neural Computation"},{"issue":"10","key":"2024100916000914600_bib8","doi-asserted-by":"publisher","first-page":"2132","DOI":"10.1162\/neco_a_01529","article-title":"Branching time active inference with Bayesian filtering","volume":"34","author":"Champion","year":"2022","journal-title":"Neural Computation"},{"key":"2024100916000914600_bib9","author":"Champion","year":"2022","journal-title":"Multi-modal and multi-factor branching time active inference"},{"key":"2024100916000914600_bib10","author":"Champion","year":"2024","journal-title":"Reframing the expected free energy: Four formulations and a unification"},{"issue":"9","key":"2024100916000914600_bib13","doi-asserted-by":"publisher","first-page":"809","DOI":"10.1016\/j.bpsc.2018.06.010","article-title":"Active inference in OpenAI gym: A paradigm for computational investigations into psychiatric illness","volume":"3","author":"Cullen","year":"2018","journal-title":"Biological Psychiatry: Cognitive Neuroscience and Neuroimaging"},{"key":"2024100916000914600_bib11","author":"Da Costa","year":"2020","journal-title":"Active inference on discrete state-spaces: A synthesis"},{"key":"2024100916000914600_bib12","author":"Da Costa","year":"2022","journal-title":"Reward maximisation through discrete active inference"},{"issue":"5","key":"2024100916000914600_bib14","doi-asserted-by":"publisher","first-page":"807","DOI":"10.1162\/neco_a_01574","article-title":"Reward maximization through discrete active inference","volume":"35","author":"Da Costa","year":"2023","journal-title":"Neural Computation"},{"key":"2024100916000914600_bib15","author":"Doersch","year":"2016","journal-title":"Tutorial on variational autoencoders"},{"key":"2024100916000914600_bib16","doi-asserted-by":"publisher","DOI":"10.3389\/fncom.2015.00136","article-title":"Dopamine, reward learning, and active inference","volume":"9","author":"FitzGerald","year":"2015","journal-title":"Frontiers in Computational Neuroscience"},{"key":"2024100916000914600_bib17","author":"Fountas","year":"2020","journal-title":"Deep active inference agents using Monte-Carlo methods"},{"key":"2024100916000914600_bib18","doi-asserted-by":"publisher","first-page":"862","DOI":"10.1016\/j.neubiorev.2016.06.022","article-title":"Active inference and learning","volume":"68","author":"Friston","year":"2016","journal-title":"Neuroscience and Biobehavioral Reviews"},{"key":"2024100916000914600_bib19","author":"Friston","year":"2020","journal-title":"Sophisticated inference"},{"issue":"10","key":"2024100916000914600_bib23","doi-asserted-by":"publisher","first-page":"1295","DOI":"10.1016\/j.visres.2008.09.007","article-title":"Bayesian surprise attracts human attention","volume":"49","author":"Itti","year":"2009","journal-title":"Vision Research"},{"key":"2024100916000914600_bib24","article-title":"Auto-encoding variational Bayes","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Kingma","year":"2014"},{"key":"2024100916000914600_bib25","article-title":"Playing FPS games with deep reinforcement learning","author":"Lample","year":"2016"},{"key":"2024100916000914600_bib26","author":"Lanillos","year":"2020","journal-title":"Robot self\/other distinction: Active inference meets neural networks learning in a mirror"},{"key":"2024100916000914600_bib27","author":"Matthey","year":"2017","journal-title":"dSprites: Disentanglement testing sprites dataset"},{"key":"2024100916000914600_bib21","article-title":"beta-VAE: Learning basic visual concepts with a constrained variational framework","volume-title":"Proceedings of the 5th International Conference on Learning Representations","author":"Matthey","year":"2017"},{"key":"2024100916000914600_bib28","doi-asserted-by":"publisher","DOI":"10.31234\/osf.io\/kf6wc","author":"Millidge","year":"2019","journal-title":"Combining active inference and hierarchical predictive coding: A tutorial introduction and case study"},{"key":"2024100916000914600_bib29","doi-asserted-by":"publisher","DOI":"10.1016\/j.jmp.2020.102348","article-title":"Deep active inference as variational policy gradients","volume":"96","author":"Millidge","year":"2020","journal-title":"Journal of Mathematical Psychology"},{"key":"2024100916000914600_bib30","author":"Mnih","year":"2016","journal-title":"Asynchronous methods for deep reinforcement learning"},{"key":"2024100916000914600_bib31","article-title":"Playing Atari with deep reinforcement learning","author":"Mnih","year":"2013"},{"key":"2024100916000914600_bib32","author":"Oliver","year":"2019","journal-title":"Active inference body perception and action for humanoid robots"},{"issue":"5\u20136","key":"2024100916000914600_bib33","doi-asserted-by":"publisher","first-page":"495","DOI":"10.1007\/s00422-019-00805-w","article-title":"Generalised free energy and active inference","volume":"113","author":"Parr","year":"2019","journal-title":"Biological Cybernetics"},{"key":"2024100916000914600_bib34","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/12441.001.0001","volume-title":"Active inference: The free energy principle in mind, brain, and behavior","author":"Parr","year":"2022"},{"key":"2024100916000914600_bib35","author":"Pezzato","year":"2020","journal-title":"Active inference and behavior trees for reactive action planning and execution in robotics"},{"key":"2024100916000914600_bib36","article-title":"Stochastic backpropagation and approximate inference in deep generative models","author":"Rezende","year":"2014","journal-title":"Proceedings of the 31st International Conference on Machine Learning"},{"key":"2024100916000914600_bib37","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1007\/978-3-030-64919-7_10","article-title":"A deep active inference model of the rubber-hand illusion","volume-title":"Active inference","author":"Rood","year":"2020"},{"key":"2024100916000914600_bib38","author":"Sancaktar","year":"2020","journal-title":"End-to-end pixel-based deep active inference for body perception and action"},{"key":"2024100916000914600_bib39","article-title":"Active inference for robotic manipulation","author":"Schneider"},{"key":"2024100916000914600_bib40","author":"Schneider","year":"2022","journal-title":"Active inference for robotic manipulation"},{"key":"2024100916000914600_bib41","author":"Schulman","year":"2017","journal-title":"Proximal policy optimization algorithms"},{"key":"2024100916000914600_bib42","doi-asserted-by":"publisher","DOI":"10.7554\/eLife.41073","article-title":"Computational mechanisms of curiosity and goal-directed exploration","author":"Schwartenbeck","year":"2018","journal-title":"eLife"},{"issue":"7587","key":"2024100916000914600_bib43","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"2024100916000914600_bib44","volume-title":"Reinforcement learning: An introduction","author":"Sutton","year":"2018"},{"issue":"6","key":"2024100916000914600_bib45","doi-asserted-by":"publisher","first-page":"547","DOI":"10.1007\/s00422-018-0785-7","article-title":"Deep active inference","volume":"112","author":"Ueltzh\u00f6ffer","year":"2018","journal-title":"Biological Cybernetics"},{"key":"2024100916000914600_bib20","author":"van Hasselt","year":"2015","journal-title":"Deep reinforcement learning with double Q-learning"},{"key":"2024100916000914600_bib22","author":"van der Himst","year":"2020","journal-title":"Deep active inference for partially observable MDPS"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/36\/11\/2403\/2474343\/neco_a_01697.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/36\/11\/2403\/2474343\/neco_a_01697.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,9]],"date-time":"2024-10-09T16:01:02Z","timestamp":1728489662000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/36\/11\/2403\/124017\/Deconstructing-Deep-Active-Inference-A-Contrarian"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,10,11]]},"references-count":45,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2024,10,11]]},"published-print":{"date-parts":[[2024,10,11]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01697","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,11]]},"published":{"date-parts":[[2024,10,11]]}}}