{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:16:01Z","timestamp":1740122161913,"version":"3.37.3"},"reference-count":50,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2021,4,19]],"date-time":"2021-04-19T00:00:00Z","timestamp":1618790400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,4,19]],"date-time":"2021-04-19T00:00:00Z","timestamp":1618790400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["1717324"],"award-info":[{"award-number":["1717324"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001711","name":"Schweizerischer Nationalfonds zur F\u00f6rderung der Wissenschaftlichen Forschung","doi-asserted-by":"publisher","award":["407540_167320"],"award-info":[{"award-number":["407540_167320"]}],"id":[{"id":"10.13039\/501100001711","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100005869","name":"Universit\u00e9 de Fribourg","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005869","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Auton Agent Multi-Agent Syst"],"published-print":{"date-parts":[[2021,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We propose a new method for learning compact state representations and policies separately but simultaneously for policy approximation in vision-based applications such as Atari games. Approaches based on deep reinforcement learning typically map pixels directly to actions to enable end-to-end training. Internally, however, the deep neural network bears the responsibility of both extracting useful information and making decisions based on it, two objectives which can be addressed independently. Separating the image processing from the action selection allows for a better understanding of either task individually, as well as potentially finding smaller policy representations which is inherently interesting. Our approach learns state representations using a compact encoder based on two novel algorithms: (i) Increasing Dictionary Vector Quantization builds a dictionary of state representations which grows in size over time, allowing our method to address new observations as they appear in an open-ended online-learning context; and (ii) Direct Residuals Sparse Coding encodes observations in function of the dictionary, aiming for highest information inclusion by disregarding reconstruction error and maximizing code sparsity. As the dictionary size increases, however, the encoder produces increasingly larger inputs for the neural network; this issue is addressed with a new variant of the Exponential Natural Evolution Strategies algorithm which adapts the dimensionality of its probability distribution along the run. We test our system on a selection of Atari games using tiny neural networks of only 6 to 18 neurons (depending on each game\u2019s controls). These are still capable of achieving results that are not much worse, and occasionally superior, to the state-of-the-art in direct policy search which uses two orders of magnitude more neurons.<\/jats:p>","DOI":"10.1007\/s10458-021-09497-8","type":"journal-article","created":{"date-parts":[[2021,4,19]],"date-time":"2021-04-19T06:02:40Z","timestamp":1618812160000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Playing Atari with few neurons"],"prefix":"10.1007","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1005-5246","authenticated-orcid":false,"given":"Giuseppe","family":"Cuccu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Julian","family":"Togelius","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Philippe","family":"Cudr\u00e9-Mauroux","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,4,19]]},"reference":[{"key":"9497_CR1","doi-asserted-by":"crossref","unstructured":"Alvernaz, S., & Togelius, J. (2017). Autoencoder-augmented neuroevolution for visual doom playing. In Computational Intelligence and Games (CIG), 2017 IEEE Conference on, IEEE, pp 1\u20138.","DOI":"10.1109\/CIG.2017.8080408"},{"key":"9497_CR2","unstructured":"Badia, A. P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, D., & Blundell, C. (2020). Agent57: Outperforming the Atari human benchmark. arXiv preprint arXiv:200313350."},{"key":"9497_CR3","doi-asserted-by":"publisher","first-page":"253","DOI":"10.1613\/jair.3912","volume":"47","author":"MG Bellemare","year":"2013","unstructured":"Bellemare, M. G., Naddaf, Y., Veness, J., & Bowling, M. (2013). The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research, 47, 253\u2013279.","journal-title":"Journal of Artificial Intelligence Research"},{"key":"9497_CR4","unstructured":"Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). Openai gym. arXiv:1606.01540."},{"key":"9497_CR5","doi-asserted-by":"crossref","unstructured":"Chrabaszcz, P., Loshchilov, I., & Hutter, F. (2018). Back to basics: Benchmarking canonical evolution strategies for playing atari. arXiv preprint arXiv:180208842.","DOI":"10.24963\/ijcai.2018\/197"},{"key":"9497_CR6","unstructured":"Coates, A., & Ng, A. Y. (2011). The importance of encoding versus training with sparse coding and vector quantization. In Proceedings of the 28th International Conference on Machine Learning (ICML-11), pp 921\u2013928."},{"key":"9497_CR7","unstructured":"Cobbe, K., Klimov, O., Hesse, C., Kim, T., & Schulman, J. (2018). Quantifying generalization in reinforcement learning. arXiv preprint arXiv:181202341."},{"key":"9497_CR8","unstructured":"Conti, E., Madhavan, V., Such, F. P., Lehman, J., Stanley, K., & Clune, J. (2018). Improving exploration in evolution strategies for deep reinforcement learning via a population of novelty-seeking agents. Advances in Neural Information Processing Systems (NIPS), 5032\u20135043."},{"key":"9497_CR9","doi-asserted-by":"crossref","unstructured":"Cuccu, G., & Gomez, F. (2012). Block diagonal natural evolution strategies. In International Conference on Parallel Problem Solving from Nature, Springer, pp 488\u2013497.","DOI":"10.1007\/978-3-642-32964-7_49"},{"key":"9497_CR10","doi-asserted-by":"crossref","unstructured":"Cuccu, G., Luciw, M., Schmidhuber, J., & Gomez, F. (2011). Intrinsically motivated neuroevolution for vision-based reinforcement learning. In Development and Learning (ICDL), 2011 IEEE International Conference on, IEEE, vol 2, pp 1\u20137.","DOI":"10.1109\/DEVLRN.2011.6037324"},{"key":"9497_CR11","unstructured":"Cuccu, G., Togelius, J., & Cudr\u00e9-Mauroux, P. (2019). Playing Atari with six neurons. In Proceedings of the 18th International Conference on Autonomous Agents and MultiAgent Systems, International Foundation for Autonomous Agents and Multiagent Systems, pp 998\u20131006."},{"key":"9497_CR12","unstructured":"Ecoffet, A., Huizinga, J., Lehman, J., Stanley, K. O., & Clune, J. (2019). Go-explore: a new approach for hard-exploration problems. arXiv preprint arXiv:190110995."},{"issue":"1","key":"9497_CR13","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1007\/s12065-007-0002-4","volume":"1","author":"D Floreano","year":"2008","unstructured":"Floreano, D., D\u00fcrr, P., & Mattiussi, C. (2008). Neuroevolution: from architectures to learning. Evolutionary Intelligence, 1(1), 47\u201362.","journal-title":"Evolutionary Intelligence"},{"key":"9497_CR14","doi-asserted-by":"crossref","unstructured":"Glasmachers, T., Schaul, T., Yi, S., Wierstra, D., & Schmidhuber, J. (2010). Exponential natural evolution strategies. In Proceedings of the 12th annual conference on Genetic and evolutionary computation, ACM, pp 393\u2013400.","DOI":"10.1145\/1830483.1830557"},{"issue":"May","key":"9497_CR15","first-page":"937","volume":"9","author":"F Gomez","year":"2008","unstructured":"Gomez, F., Schmidhuber, J., & Miikkulainen, R. (2008). Accelerated neural evolution through cooperatively coevolved synapses. Journal of Machine Learning Research, 9(May), 937\u2013965.","journal-title":"Journal of Machine Learning Research"},{"issue":"2","key":"9497_CR16","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1109\/MASSP.1984.1162229","volume":"1","author":"R Gray","year":"1984","unstructured":"Gray, R. (1984). Vector quantization. IEEE ASSP Magazine, 1(2), 4\u201329.","journal-title":"IEEE ASSP Magazine"},{"key":"9497_CR17","unstructured":"Ha, D., & Schmidhuber, J. (2018). World models. arXiv preprint arXiv:180310122."},{"issue":"2","key":"9497_CR18","doi-asserted-by":"publisher","first-page":"159","DOI":"10.1162\/106365601750190398","volume":"9","author":"N Hansen","year":"2001","unstructured":"Hansen, N., & Ostermeier, A. (2001). Completely derandomized self-adaptation in evolution strategies. Evolutionary Computation, 9(2), 159\u2013195.","journal-title":"Evolutionary Computation"},{"issue":"4","key":"9497_CR19","doi-asserted-by":"publisher","first-page":"355","DOI":"10.1109\/TCIAIG.2013.2294713","volume":"6","author":"M Hausknecht","year":"2014","unstructured":"Hausknecht, M., Lehman, J., Miikkulainen, R., & Stone, P. (2014). A neuroevolution approach to general Atari game playing. IEEE Transactions on Computational Intelligence and AI in Games, 6(4), 355\u2013366.","journal-title":"IEEE Transactions on Computational Intelligence and AI in Games"},{"key":"9497_CR20","unstructured":"Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., & Silver, D. (2017). Rainbow: Combining improvements in deep reinforcement learning. arXiv preprint arXiv:171002298."},{"key":"9497_CR21","doi-asserted-by":"crossref","unstructured":"Igel, C. (2003). Neuroevolution for reinforcement learning using evolution strategies. In Evolutionary Computation, 2003. CEC\u201903. The 2003 Congress on, IEEE, vol 4, pp 2588\u20132595.","DOI":"10.1109\/CEC.2003.1299414"},{"key":"9497_CR22","unstructured":"Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., & Kavukcuoglu, K. (2016). Reinforcement learning with unsupervised auxiliary tasks. arXiv preprint arXiv:161105397."},{"key":"9497_CR23","doi-asserted-by":"crossref","unstructured":"Juliani, A., Khalifa, A., Berges, V. P., Harper, J., Henry, H., Crespi, A., Togelius, J., & Lange, D. (2019). Obstacle tower: A generalization challenge in vision, control, and planning. arXiv preprint arXiv:190201378.","DOI":"10.24963\/ijcai.2019\/373"},{"key":"9497_CR24","doi-asserted-by":"crossref","unstructured":"Justesen, N., Torrado, R. R., Bontrager, P., Khalifa, A., Togelius, J., & Risi, S. (2018). Illuminating generalization in deep reinforcement learning through procedural level generation. In NeurIPS Workshop on Deep Reinforcement Learning.","DOI":"10.1109\/CIG.2018.8490422"},{"key":"9497_CR25","doi-asserted-by":"crossref","unstructured":"Justesen, N., Bontrager, P., Togelius, J., & Risi, S. (2019). Deep learning for video game playing. IEEE Transactions on Games.","DOI":"10.1109\/TG.2019.2896986"},{"key":"9497_CR26","doi-asserted-by":"crossref","unstructured":"Kempka, M., Wydmuch, M., Runc, G., Toczek, J., & Ja\u015bkowski, W. (2016). Vizdoom: A doom-based ai research platform for visual reinforcement learning. In 2016 IEEE Conference on Computational Intelligence and Games (CIG), IEEE, pp 1\u20138.","DOI":"10.1109\/CIG.2016.7860433"},{"key":"9497_CR27","doi-asserted-by":"crossref","unstructured":"Koutn\u00edk, J., Schmidhuber, J., & Gomez, F. (2014). Evolving deep unsupervised convolutional networks for vision-based reinforcement learning. In Proceedings of the 2014 Annual Conference on Genetic and Evolutionary Computation, pp 541\u2013548.","DOI":"10.1145\/2576768.2598358"},{"key":"9497_CR28","unstructured":"Li, C., Farkhoor, H., Liu, R., & Yosinski, J. (2018). Measuring the intrinsic dimension of objective landscapes. arXiv preprint arXiv:180408838."},{"key":"9497_CR29","doi-asserted-by":"crossref","unstructured":"Mairal, J., Bach, F., Ponce, J., et al. (2014). Sparse modeling for image and vision processing. Foundations and Trends$$\\textregistered$$in Computer Graphics and Vision, 8(2\u20133), 85\u2013283.","DOI":"10.1561\/0600000058"},{"issue":"12","key":"9497_CR30","doi-asserted-by":"publisher","first-page":"3397","DOI":"10.1109\/78.258082","volume":"41","author":"SG Mallat","year":"1993","unstructured":"Mallat, S. G., & Zhang, Z. (1993). Matching pursuits with time-frequency dictionaries. IEEE Transactions on Signal Processing, 41(12), 3397\u20133415.","journal-title":"IEEE Transactions on Signal Processing"},{"issue":"7540","key":"9497_CR31","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","volume":"518","author":"V Mnih","year":"2015","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., et al. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529.","journal-title":"Nature"},{"issue":"4","key":"9497_CR32","doi-asserted-by":"publisher","first-page":"293","DOI":"10.1109\/TCIAIG.2013.2286295","volume":"5","author":"S Ontan\u00f3n","year":"2013","unstructured":"Ontan\u00f3n, S., Synnaeve, G., Uriarte, A., Richoux, F., Churchill, D., & Preuss, M. (2013). A survey of real-time strategy game AI research and competition in starcraft. IEEE Transactions on Computational Intelligence and AI in Games, 5(4), 293\u2013311.","journal-title":"IEEE Transactions on Computational Intelligence and AI in Games"},{"key":"9497_CR33","doi-asserted-by":"crossref","unstructured":"Pathak, D., Agrawal, P., Efros, A. A., & Darrell, T. (2017). Curiosity-driven exploration by self-supervised prediction. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp 16\u201317.","DOI":"10.1109\/CVPRW.2017.70"},{"key":"9497_CR34","doi-asserted-by":"crossref","unstructured":"Pati, Y. C., Rezaiifar, R., & Krishnaprasad, P. S. (1993). Orthogonal matching pursuit: Recursive function approximation with applications to wavelet decomposition. In Signals, Systems and Computers, 1993. 1993 Conference Record of The Twenty-Seventh Asilomar Conference on, IEEE, pp 40\u201344.","DOI":"10.1109\/ACSSC.1993.342465"},{"key":"9497_CR35","unstructured":"Perez, D., Liu, J., Abdel, Samea Khalifa A., Gaina, R. D., Togelius, J., & Lucas, S. M. (2019). General video game AI: a multi-track framework for evaluating agents, games and content generation algorithms. IEEE Transactions on Games."},{"key":"9497_CR36","doi-asserted-by":"crossref","unstructured":"Perez-Liebana, D., Samothrakis, S., Togelius, J., Schaul, T., & Lucas, S.M. (2016). General video game AI: Competition, challenges and opportunities. In Thirtieth AAAI Conference on Artificial Intelligence.","DOI":"10.1609\/aaai.v30i1.9869"},{"issue":"1","key":"9497_CR37","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1109\/TCIAIG.2015.2494596","volume":"9","author":"S Risi","year":"2017","unstructured":"Risi, S., & Togelius, J. (2017). Neuroevolution in games: State of the art and open challenges. IEEE Transactions on Computational Intelligence and AI in Games, 9(1), 25\u201341.","journal-title":"IEEE Transactions on Computational Intelligence and AI in Games"},{"key":"9497_CR38","unstructured":"Salimans, T., Ho, J., Chen, X., Sidor, S., & Sutskever, I. (2017). Evolution strategies as a scalable alternative to reinforcement learning. arXiv preprint arXiv:170303864."},{"key":"9497_CR39","doi-asserted-by":"crossref","unstructured":"Schaul, T., Glasmachers, T., & Schmidhuber, J. (2011). High dimensions and heavy tails for natural evolution strategies. In Proceedings of the 13th Annual Conference on Genetic and Evolutionary Computation, ACM, pp 845\u2013852.","DOI":"10.1145\/2001576.2001692"},{"key":"9497_CR40","doi-asserted-by":"crossref","unstructured":"Sermanet, P., Lynch, C., Chebotar, Y., Hsu, J., Jang, E., Schaal, S., & Levine, S. (2018). Time-contrastive networks: Self-supervised learning from video. In 2018 IEEE International Conference on Robotics and Automation (ICRA), IEEE, pp 1134\u20131141.","DOI":"10.1109\/ICRA.2018.8462891"},{"issue":"2","key":"9497_CR41","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1162\/106365602320169811","volume":"10","author":"KO Stanley","year":"2002","unstructured":"Stanley, K. O., & Miikkulainen, R. (2002). Evolving neural networks through augmenting topologies. Evolutionary Computation, 10(2), 99\u2013127.","journal-title":"Evolutionary Computation"},{"key":"9497_CR42","unstructured":"Such, F. P., Madhavan, V., Conti, E., Lehman, J., Stanley, K. O., & Clune, J. (2017). Deep neuroevolution: Genetic algorithms are a competitive alternative for training deep neural networks for reinforcement learning. arXiv preprint arXiv:171206567"},{"issue":"3","key":"9497_CR43","first-page":"30","volume":"23","author":"J Togelius","year":"2009","unstructured":"Togelius, J., Schaul, T., Wierstra, D., Igel, C., Gomez, F., & Schmidhuber, J. (2009). Ontogenetic and phylogenetic reinforcement learning. K\u00fcnstliche Intelligenz, 23(3), 30\u201333.","journal-title":"K\u00fcnstliche Intelligenz"},{"issue":"3","key":"9497_CR44","doi-asserted-by":"publisher","first-page":"89","DOI":"10.1609\/aimag.v34i3.2492","volume":"34","author":"J Togelius","year":"2013","unstructured":"Togelius, J., Shaker, N., Karakovskiy, S., & Yannakakis, G. N. (2013). The mario AI championship 2009\u20132012. AI Magazine, 34(3), 89\u201392.","journal-title":"AI Magazine"},{"key":"9497_CR45","unstructured":"Vinyals, O., Ewalds, T., Bartunov, S., Georgiev, P., Vezhnevets, A. S., Yeo, M., Makhzani, A., K\u00fcttler, H., Agapiou, J., Schrittwieser, J., et al. (2017). Starcraft II: a new challenge for reinforcement learning. arXiv preprint arXiv:170804782."},{"key":"9497_CR46","doi-asserted-by":"crossref","unstructured":"Wierstra, D., Schaul, T., Peters, J., & Schmidhuber, J. (2008). Natural evolution strategies. In Evolutionary Computation, 2008. CEC 2008.(IEEE World Congress on Computational Intelligence). IEEE Congress on, IEEE, pp 3381\u20133387.","DOI":"10.1109\/CEC.2008.4631255"},{"issue":"1","key":"9497_CR47","first-page":"949","volume":"15","author":"D Wierstra","year":"2014","unstructured":"Wierstra, D., Schaul, T., Glasmachers, T., Sun, Y., Peters, J., & Schmidhuber, J. (2014). Natural evolution strategies. Journal of Machine Learning Research, 15(1), 949\u2013980.","journal-title":"Journal of Machine Learning Research"},{"key":"9497_CR48","doi-asserted-by":"crossref","unstructured":"Yannakakis, G. N., & Togelius, J. (2018). Artificial Intelligence and Games. Springer, http:\/\/gameaibook.org.","DOI":"10.1007\/978-3-319-63519-4"},{"issue":"9","key":"9497_CR49","doi-asserted-by":"publisher","first-page":"1423","DOI":"10.1109\/5.784219","volume":"87","author":"X Yao","year":"1999","unstructured":"Yao, X. (1999). Evolving artificial neural networks. Proceedings of the IEEE, 87(9), 1423\u20131447.","journal-title":"Proceedings of the IEEE"},{"key":"9497_CR50","doi-asserted-by":"publisher","first-page":"490","DOI":"10.1109\/ACCESS.2015.2430359","volume":"3","author":"Z Zhang","year":"2015","unstructured":"Zhang, Z., Xu, Y., Yang, J., Li, X., & Zhang, D. (2015). A survey of sparse representation: algorithms and applications. IEEE Access, 3, 490\u2013530.","journal-title":"IEEE Access"}],"container-title":["Autonomous Agents and Multi-Agent Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-021-09497-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10458-021-09497-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10458-021-09497-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,24]],"date-time":"2022-12-24T17:02:55Z","timestamp":1671901375000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10458-021-09497-8"}},"subtitle":["Improving the efficacy of reinforcement learning by decoupling feature extraction and decision making"],"short-title":[],"issued":{"date-parts":[[2021,4,19]]},"references-count":50,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2021,10]]}},"alternative-id":["9497"],"URL":"https:\/\/doi.org\/10.1007\/s10458-021-09497-8","relation":{},"ISSN":["1387-2532","1573-7454"],"issn-type":[{"type":"print","value":"1387-2532"},{"type":"electronic","value":"1573-7454"}],"subject":[],"published":{"date-parts":[[2021,4,19]]},"assertion":[{"value":"9 March 2021","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 April 2021","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"17"}}