{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T18:15:25Z","timestamp":1759342525411},"reference-count":67,"publisher":"MIT Press","issue":"4","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,3,18]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The discovery of reusable subroutines simplifies decision making and planning in complex reinforcement learning problems. Previous approaches propose to learn such temporal abstractions in an unsupervised fashion through observing state-action trajectories gathered from executing a policy. However, a current limitation is that they process each trajectory in an entirely sequential manner, which prevents them from revising earlier decisions about subroutine boundary points in light of new incoming information. In this work, we propose slot-based transformer for temporal abstraction (SloTTAr), a fully parallel approach that integrates sequence processing transformers with a slot attention module to discover subroutines in an unsupervised fashion while leveraging adaptive computation for learning about the number of such subroutines solely based on their empirical distribution. We demonstrate how SloTTAr is capable of outperforming strong baselines in terms of boundary point discovery, even for sequences containing variable amounts of subroutines, while being up to seven times faster to train on existing benchmarks.<\/jats:p>","DOI":"10.1162\/neco_a_01567","type":"journal-article","created":{"date-parts":[[2023,2,6]],"date-time":"2023-02-06T23:26:00Z","timestamp":1675725960000},"page":"593-626","update-policy":"http:\/\/dx.doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":2,"title":["Unsupervised Learning of Temporal Abstractions With Slot-Based Transformers"],"prefix":"10.1162","volume":"35","author":[{"given":"Anand","family":"Gopalakrishnan","sequence":"first","affiliation":[{"name":"The Swiss AI Lab, Lugano 6962, Switzerland"},{"name":"USI, Lugano 6900, Switzerland"},{"name":"SUPSI, Manno 6928, Switzerland anand@idsia.ch"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kazuki","family":"Irie","sequence":"additional","affiliation":[{"name":"The Swiss AI Lab, Lugano 6962, Switzerland"},{"name":"USI, Lugano 6900, Switzerland"},{"name":"SUPSI, Manno 6928, Switzerland kazuki@idsia.ch"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J\u00fcrgen","family":"Schmidhuber","sequence":"additional","affiliation":[{"name":"The Swiss AI Lab, Lugano 6962, Switzerland"},{"name":"USI, Lugano 6900, Switzerland"},{"name":"SUPSI, Manno 6928, Switzerland"},{"name":"AI Initiative, King Abdullah University of Science and Technology, Thuwal 23955, Saudi Arabia juergen@idsia.ch"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sjoerd","family":"van Steenkiste","sequence":"additional","affiliation":[{"name":"Google Research, Mountain View, CA 94043, U.S.A. svansteenkiste@google.com"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2023,3,18]]},"reference":[{"key":"2023032023153718200_","article-title":"OPAL: Offline primitive discovery for accelerating offline reinforcement learning","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Ajay","year":"2021"},{"key":"2023032023153718200_","first-page":"166","article-title":"Modular multitask reinforcement learning with policy sketches","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Andreas","year":"2017"},{"key":"2023032023153718200_","first-page":"1726","article-title":"The option-critic architecture","volume-title":"Proc. of the AAAI Conf. on Artificial Intelligence","author":"Bacon","year":"2017"},{"key":"2023032023153718200_","first-page":"438","article-title":"Hierarchical reinforcement learning based on subgoal discovery and subpolicy specialization","volume-title":"Proc. of the Conf. on Intelligent Autonomous Systems","author":"Bakker","year":"2004"},{"key":"2023032023153718200_","article-title":"Pondernet: Learning to ponder","volume-title":"ICML Workshop on Automated Reasoning","author":"Banino","year":"2021"},{"key":"2023032023153718200_","article-title":"Transdreamer: Reinforcement learning with transformer world models","author":"Chen","year":"2021","journal-title":"NeurIPS Workshop on Deep Reinforcement Learning"},{"key":"2023032023153718200_","article-title":"Decision transformer: Reinforcement learning via sequence modeling","volume-title":"Advances in neural information processing systems","author":"Chen","year":"2021"},{"article-title":"Minimalistic gridworld environment for OpenAI gym","year":"2018","author":"Chevalier-Boisvert","key":"2023032023153718200_"},{"key":"2023032023153718200_","first-page":"1724","article-title":"Learning phrase representations using RNN encoder\u2013decoder for statistical machine translation","volume-title":"Proc. of the Conf. on Empirical Methods in Natural Language Processing","author":"Cho","year":"2014"},{"key":"2023032023153718200_","first-page":"271","article-title":"Feudal reinforcement learning","volume-title":"Advances in neural information processing systems, 5","author":"Dayan","year":"1992"},{"key":"2023032023153718200_","article-title":"Universal transformers","volume-title":"Proc. of the Int. Conf. on Learning Representations","author":"Dehghani","year":"2019"},{"key":"2023032023153718200_","article-title":"Attention over learned object embeddings enables complex visual reasoning","volume-title":"Advances in neural information processing systems, 34","author":"Ding","year":"2021"},{"key":"2023032023153718200_","article-title":"An image is worth 16x16 words: Transformers for image recognition at scale","volume-title":"Proc. of the Int. Conf. on Learning Representation","author":"Dosovitskiy","year":"2021"},{"key":"2023032023153718200_","article-title":"Diversity is all you need: Learning skills without a reward function","volume-title":"Proc. of the Int. Conf. on Learning Representations","author":"Eysenbach","year":"2019"},{"journal-title":"Addressing some limitations of transformers with feedback memory","year":"2020","author":"Fan","key":"2023032023153718200_"},{"issue":"10","key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"2451","DOI":"10.1162\/089976600300015015","article-title":"Learning to forget: Continual prediction with LSTM","volume":"12","author":"Gers","year":"2000","journal-title":"Neural Computation"},{"journal-title":"Adaptive computation time for recurrent neural networks","year":"2016","author":"Graves","key":"2023032023153718200_"},{"key":"2023032023153718200_","first-page":"369","article-title":"Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"Graves","year":"2006"},{"key":"2023032023153718200_","first-page":"2424","article-title":"Multi-object representation learning with iterative variational inference","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"Greff","year":"2019"},{"key":"2023032023153718200_","first-page":"6691","article-title":"Neural expectation maximization","volume-title":"Advances in neural information processing systems, 30","author":"Greff","year":"2017"},{"journal-title":"On the binding problem in artificial neural networks","year":"2020","author":"Greff","key":"2023032023153718200_"},{"key":"2023032023153718200_","article-title":"Temporal difference variational auto-encoder","volume-title":"Proc. of the Int. Conf. on Learning Representations","author":"Gregor","year":"2019"},{"article-title":"Variational intrinsic control","year":"2017","author":"Gregor","key":"2023032023153718200_"},{"issue":"6","key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"1221","DOI":"10.3758\/BF03193267","article-title":"Making sense of abstract events: Building event schemas","volume":"34","author":"Hard","year":"2006","journal-title":"Memory and Cognition"},{"issue":"8","key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Computation"},{"issue":"1\u20132","key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"183","DOI":"10.1080\/713756773","article-title":"Event files: Evidence for automatic integration of stimulus-response episodes","volume":"5","author":"Hommel","year":"1998","journal-title":"Visual Cognition"},{"issue":"11","key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"494","DOI":"10.1016\/j.tics.2004.08.007","article-title":"Event files: Feature binding in and across perception and action","volume":"8","author":"Hommel","year":"2004","journal-title":"Trends in Cognitive Sciences"},{"key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"42","DOI":"10.1007\/s00426-005-0035-1","article-title":"Feature integration across perception and action: Event files affect response choice","volume":"71","author":"Hommel","year":"2007","journal-title":"Psychological Research"},{"key":"2023032023153718200_","article-title":"Going beyond linear transformers with recurrent fast weight programmers","volume-title":"Advances in neural information processing systems","author":"Irie","year":"2021"},{"journal-title":"Perceiver IO: A general architecture for structured inputs and outputs.","year":"2021","author":"Jaegle","key":"2023032023153718200_"},{"key":"2023032023153718200_","first-page":"4651","article-title":"Perceiver: General perception with iterative attention","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"Jaegle","year":"2021"},{"key":"2023032023153718200_","article-title":"Reinforcement learning as one big sequence modeling problem","volume-title":"Advances in neural information processing systems","author":"Janner","year":"2021"},{"issue":"2","key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1016\/0010-0285(92)90007-O","article-title":"The reviewing of object files: Object-specific integration of information","volume":"24","author":"Kahneman","year":"1992","journal-title":"Cognitive Psychology"},{"key":"2023032023153718200_","first-page":"11566","article-title":"Variational temporal abstraction","volume-title":"Advances in neural information processing systems","author":"Kim","year":"2019"},{"key":"2023032023153718200_","article-title":"Conditional object-centric learning from video","volume-title":"Proc. of the Int. Conf. on Learning Representations","author":"Kipf","year":"2022"},{"key":"2023032023153718200_","first-page":"3418","article-title":"CompILE: Compositional imitation learning and execution","volume-title":"Proc. Int. of the Conf. on Machine Learning","author":"Kipf","year":"2019"},{"volume-title":"Principles of gestalt psychology","year":"1935","author":"Koffka","key":"2023032023153718200_"},{"volume-title":"Gestalt psychology","year":"1929","author":"K\u00f6hler","key":"2023032023153718200_"},{"key":"2023032023153718200_","article-title":"Object-centric learning with slot attention","volume-title":"Advances in neural information processing systems","author":"Locatello","year":"2020"},{"key":"2023032023153718200_","article-title":"Learning task decomposition with ordered memory policy network","volume-title":"Proc. of the Int. Conf. on Learning Representations","author":"Lu","year":"2021"},{"key":"2023032023153718200_","first-page":"2295","article-title":"A Laplacian framework for option discovery in reinforcement learning","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"Machado","year":"2017"},{"key":"2023032023153718200_","first-page":"361","article-title":"Automatic discovery of subgoals in reinforcement learning using diverse density","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"McGovern","year":"2001"},{"key":"2023032023153718200_","first-page":"1928","article-title":"Asynchronous methods for deep reinforcement learning","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"Mnih","year":"2016"},{"key":"2023032023153718200_","first-page":"7487","article-title":"Stabilizing transformers for reinforcement learning","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"Parisotto","year":"2020"},{"key":"2023032023153718200_","first-page":"3681","article-title":"MCP: Learning composable hierarchical control with multiplicative compositional policies","volume-title":"Advances in neural information processing systems","author":"Peng","year":"2019"},{"key":"2023032023153718200_","doi-asserted-by":"crossref","DOI":"10.1093\/acprof:oso\/9780199898138.001.0001","volume-title":"Event cognition","author":"Radvansky","year":"2014"},{"key":"2023032023153718200_","article-title":"Compressive transformers for long-range sequence modelling","volume-title":"Proc. of the Int. Conf. on Learning Representations","author":"Rae","year":"2020"},{"article-title":"Towards compositional learning in dynamic networks","year":"1990","author":"Schmidhuber","key":"2023032023153718200_"},{"key":"2023032023153718200_","doi-asserted-by":"crossref","DOI":"10.1109\/IJCNN.1991.155375","article-title":"Learning to generate subgoals for action sequences","volume-title":"Proceedings of the Seattle International Joint Conference on Neural Networks","author":"Schmidhuber","year":"1991"},{"journal-title":"Self-delimiting neural networks","year":"2012","author":"Schmidhuber","key":"2023032023153718200_"},{"journal-title":"Reinforcement learning upside down: Don't predict rewards\u2013just map them to actions","year":"2019","author":"Schmidhuber","key":"2023032023153718200_"},{"key":"2023032023153718200_","doi-asserted-by":"crossref","first-page":"196","DOI":"10.7551\/mitpress\/3116.003.0027","article-title":"Planning simple trajectories using neural subgoal generators","volume-title":"Proc. Int. Conf. on From Animals to Animats 2: Simulation of Adaptive Behavior","author":"Schmidhuber","year":"1993"},{"key":"2023032023153718200_","first-page":"4654","article-title":"TACO: Learning task decomposition via temporal alignment for control","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"Shiarlis","year":"2018"},{"key":"2023032023153718200_","doi-asserted-by":"crossref","DOI":"10.1145\/1015330.1015353","article-title":"Using relative novelty to identify useful temporal abstractions in reinforcement learning","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"\u015eim\u015fek","year":"2004"},{"key":"2023032023153718200_","first-page":"1497","article-title":"Skill characterization based on betweenness","volume-title":"Advances in neural information processing systems","author":"\u015eim\u015fek","year":"2008"},{"key":"2023032023153718200_","first-page":"816","article-title":"Identifying useful subgoals in reinforcement learning by local graph partitioning","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"\u015eim\u015fek","year":"2005"},{"key":"2023032023153718200_","article-title":"Illiterate DALLE learns to compose","volume-title":"Proc. of the Int. Conf. on Learning Representations","author":"Singh","year":"2022"},{"issue":"1","key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1207\/s15516709cog1401_3","article-title":"Principles of object perception","volume":"14","author":"Spelke","year":"1990","journal-title":"Cognitive Science"},{"journal-title":"Training agents using upside-down reinforcement learning","year":"2019","author":"Srivastava","key":"2023032023153718200_"},{"key":"2023032023153718200_","doi-asserted-by":"crossref","first-page":"212","DOI":"10.1007\/3-540-45622-8_16","article-title":"Learning options in reinforcement learning","volume-title":"Proc. of the International Symposium on Abstraction, Reformulation, and Approximation","author":"Stolle","year":"2002"},{"issue":"1\u20132","key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"181","DOI":"10.1016\/S0004-3702(99)00052-1","article-title":"Between MDPs and semi-MDPs: A framework for temporal abstraction in reinforcement learning","volume":"112","author":"Sutton","year":"1999","journal-title":"Artificial Intelligence"},{"key":"2023032023153718200_","first-page":"358","article-title":"Hierarchical explanation-based reinforcement learning","volume-title":"Proc. of the Int. Conf. on Machine Learning","author":"Tadepalli","year":"1997"},{"key":"2023032023153718200_","first-page":"5998","article-title":"Attention is all you need","volume-title":"Advances in neural information processing systems, 30","author":"Vaswani","year":"2017"},{"key":"2023032023153718200_","article-title":"Spatial broadcast decoder: A simple architecture for learning disentangled representations in VAEs","volume-title":"Learning from Limited Labeled Data Workshop","author":"Watters","year":"2019"},{"key":"2023032023153718200_","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1037\/0096-3445.130.1.29","article-title":"Perceiving, remembering, and communicating structure in events","volume":"130","author":"Zacks","year":"2001","journal-title":"Journal of Experimental Psychology: General"},{"key":"2023032023153718200_","article-title":"Dreaming with transformers","volume-title":"Proc. of the AAAI Workshop on Reinforcement Learning in Games","author":"Zeng","year":"2022"},{"key":"2023032023153718200_","article-title":"Deformable DETR: Deformable transformers for end-to-end object detection","volume-title":"Proc. of the Int. Conf. on Learning Representations","author":"Zhu","year":"2021"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/4\/593\/2075305\/neco_a_01567.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/4\/593\/2075305\/neco_a_01567.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,13]],"date-time":"2024-10-13T18:39:31Z","timestamp":1728844771000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/35\/4\/593\/114732\/Unsupervised-Learning-of-Temporal-Abstractions"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,18]]},"references-count":67,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2023,3,18]]},"published-print":{"date-parts":[[2023,3,18]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01567","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"type":"print","value":"0899-7667"},{"type":"electronic","value":"1530-888X"}],"subject":[],"published-other":{"date-parts":[[2023,4]]},"published":{"date-parts":[[2023,3,18]]}}}