{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,12]],"date-time":"2026-07-12T02:40:22Z","timestamp":1783824022969,"version":"3.55.0"},"reference-count":64,"publisher":"MIT Press","issue":"3","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,2,17]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This article does not describe a working system. Instead, it presents a single idea about representation that allows advances made by several different groups to be combined into an imaginary system called GLOM.1 The advances include transformers, neural fields, contrastive representation learning, distillation, and capsules. GLOM answers the question: How can a neural network with a fixed architecture parse an image into a part-whole hierarchy that has a different structure for each image? The idea is simply to use islands of identical vectors to represent the nodes in the parse tree. If GLOM can be made to work, it should significantly improve the interpretability of the representations produced by transformer-like systems when applied to vision or language.<\/jats:p>","DOI":"10.1162\/neco_a_01557","type":"journal-article","created":{"date-parts":[[2022,12,22]],"date-time":"2022-12-22T00:56:31Z","timestamp":1671670591000},"page":"413-452","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":73,"title":["How to Represent Part-Whole Hierarchies in a Neural Network"],"prefix":"10.1162","volume":"35","author":[{"given":"Geoffrey","family":"Hinton","sequence":"first","affiliation":[{"name":"Google Research"},{"name":"Vector Institute, Toronto, Ontario M5G 1M1, Canada"},{"name":"Department of Computer Science, University of Toronto, Toronto, ON M5S 2E4, Canada hinton@cs.toronto.edu"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2023,2,17]]},"reference":[{"key":"2023021720442934600_","first-page":"4331","article-title":"Using fast weights to attend to the recent past","volume-title":"Advances in neural information processing systems","author":"Ba","year":"2016"},{"key":"2023021720442934600_","first-page":"15535","article-title":"Learning representations by maximizing mutual information across views","volume-title":"Advances in neural information processing systems","author":"Bachman","year":"2019"},{"key":"2023021720442934600_","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1145\/3317550.3321441","article-title":"Machine learning systems are stuck in a rut","author":"Barham","year":"2019","journal-title":"HotOS '19: Proceedings of the Workshop on Hot Topics in Operating Systems"},{"key":"2023021720442934600_","author":"Bear","year":"2020","journal-title":"Learning physical graph representations from visual scenes"},{"issue":"6356","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"161","DOI":"10.1038\/355161a0","article-title":"A self-organizing neural network that discovers surfaces in random-dot stereograms","volume":"355","author":"Becker","year":"1992","journal-title":"Nature"},{"issue":"2","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"267","DOI":"10.1162\/neco.1993.5.2.267","article-title":"Learning mixture models of spatial coherence","volume":"5","author":"Becker","year":"1993","journal-title":"Neural Computation"},{"key":"2023021720442934600_","doi-asserted-by":"crossref","first-page":"535","DOI":"10.1145\/1150402.1150464","article-title":"Model compression","author":"Bucilu\u01ce","year":"2006","journal-title":"KDD '06: Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining"},{"key":"2023021720442934600_","first-page":"1597","article-title":"A simple framework for contrastive learning of visual representations","volume-title":"Proceedings of the 37th International Conference on Machine Learning","author":"Chen","year":"2020"},{"key":"2023021720442934600_","author":"Chen","year":"2020","journal-title":"Big self-supervised models are strong semi-supervised learners"},{"key":"2023021720442934600_","author":"Chen","year":"2020","journal-title":"Exploring simple Siamese representation learning"},{"key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"111","DOI":"10.1038\/304111a0","article-title":"The function of dream sleep","volume":"304","author":"Crick","year":"1983","journal-title":"Nature"},{"key":"2023021720442934600_","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-58571-6_36","article-title":"NASA: Neural articulated shape approximation","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Deng","year":"2020"},{"key":"2023021720442934600_","article-title":"BERT: Pretraining of deep bidirectional transformers for language under standing","volume-title":"Proceedings of the NAACL-HLT","author":"Devlin","year":"2018"},{"key":"2023021720442934600_","author":"Dosovitskiy","year":"2020","journal-title":"An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale"},{"issue":"6","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"721","DOI":"10.1109\/TPAMI.1984.4767596","article-title":"Stochastic relaxation, Gibbs distributions, and the Bayesian restoration of images","volume":"6","author":"Geman","year":"1984","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2023021720442934600_","author":"Grill","year":"2020","journal-title":"Bootstrap your own latent: A new approach to self-supervised learning"},{"key":"2023021720442934600_","article-title":"Generating large images from latent vectors","author":"Ha","year":"2016","journal-title":"blog.otoro.net"},{"key":"2023021720442934600_","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR42600.2020.00975","article-title":"Momentum contrast for unsupervised visual representation learning","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He","year":"2020"},{"key":"2023021720442934600_","article-title":"Multiscale conditional random fields for image labeling","volume-title":"Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"He","year":"2004"},{"issue":"3","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"231","DOI":"10.1207\/s15516709cog0303_3","article-title":"Some demonstrations of the effects of structural descriptions in mental imagery","volume":"3","author":"Hinton","year":"1979","journal-title":"Cognitive Science"},{"key":"2023021720442934600_","first-page":"1088","article-title":"Shape representation in parallel systems","volume-title":"Proceedings of the Seventh International Joint Conference on Artificial Intelligence","author":"Hinton","year":"1981"},{"key":"2023021720442934600_","article-title":"Implementing semantic networks in parallel hardware","volume-title":"Parallel models of associative memory","author":"Hinton","year":"1981"},{"key":"2023021720442934600_","first-page":"683","article-title":"A parallel computation that assigns canonical object-based frames of reference","volume-title":"Proceedings of the 7th International Joint Conference on Artificial Intelligence","author":"Hinton","year":"1981"},{"key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1016\/0004-3702(90)90004-J","article-title":"Mapping part-whole hierarchies into connectionist networks","volume":"46","author":"Hinton","year":"1990","journal-title":"Artificial Intelligence"},{"issue":"8","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"1771","DOI":"10.1162\/089976602760128018","article-title":"Training products of experts by minimizing contrastive divergence","volume":"14","author":"Hinton","year":"2002","journal-title":"Neural Computation"},{"key":"2023021720442934600_","author":"Hinton","year":"2006","journal-title":"Grant proposal to the natural sciences and engineering research council"},{"key":"2023021720442934600_","author":"Hinton","year":"2014","journal-title":"Dark knowledge"},{"key":"2023021720442934600_","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1007\/978-3-642-21735-7_6","article-title":"Transforming auto-encoders","volume-title":"ICANN 2011: Artificial Neural Networks and Machine Learning","author":"Hinton","year":"2011"},{"key":"2023021720442934600_","article-title":"Matrix capsules with EM routing","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Hinton","year":"2018"},{"key":"2023021720442934600_","first-page":"282","article-title":"Learning and relearning in Boltzmann machines","volume-title":"Parallel distributed processing: Explorations in the microstructure of cognition, vol. 1: Foundations","author":"Hinton","year":"1986"},{"key":"2023021720442934600_","article-title":"Distilling the knowledge in a neural network","volume-title":"NIPS 2014 Deep Learning Workshop","author":"Hinton","year":"2014"},{"key":"2023021720442934600_","author":"Jabri","year":"2020","journal-title":"Space-time correspondence as a contrastive random walk"},{"key":"2023021720442934600_","first-page":"15512","article-title":"Stacked capsule autoencoders","volume-title":"Advances in neural information processing systems","author":"Kosiorek","year":"2019"},{"issue":"6","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"1183","DOI":"10.1016\/S0896-6273(02)01096-6","article-title":"Memory of sequential experience in the hippocampus during slow wave sleep","volume":"36","author":"Lee","year":"2002","journal-title":"Neuron"},{"key":"2023021720442934600_","first-page":"3744","article-title":"Set transformer: A framework for attention-based permutation-invariant neural networks","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"Lee","year":"2019"},{"key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"335","DOI":"10.1038\/s41583-020-0277-3","article-title":"Backpropagation and the brain","volume":"21","author":"Lillicrap","year":"2020","journal-title":"Nature Reviews Neuroscience"},{"key":"2023021720442934600_","author":"Locatello","year":"2020","journal-title":"Object-centric learning with slot attention"},{"key":"2023021720442934600_","first-page":"405","article-title":"NeRF: Representing scenes as neural radiance fields for view synthesis","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Mildenhall","year":"2020"},{"issue":"21","key":"2023021720442934600_","doi-asserted-by":"crossref","first-page":"9497","DOI":"10.1523\/JNEUROSCI.19-21-09497.1999","article-title":"Replay and time compression of recurring spike sequences in the hippocampus","volume":"19","author":"N\u00e1dasdy","year":"1999","journal-title":"Journal of Neuroscience"},{"key":"2023021720442934600_","first-page":"355","article-title":"A view of the EM algorithm that justifies incremental, sparse, and other variants","volume-title":"Learning in graphical models","author":"Neal","year":"1997"},{"key":"2023021720442934600_","author":"Niemeyer","year":"2020","journal-title":"GIRAFFE: Representing scenes as compositional generative neural feature fields"},{"issue":"3","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"683","DOI":"10.1162\/neco.1997.9.3.683","article-title":"A mobile robot that learns its place","volume":"9","author":"Oore","year":"1997","journal-title":"Neural Computation"},{"key":"2023021720442934600_","article-title":"Modeling im-age patches with a directed hierarchy of Markov random fields","volume-title":"Advances in neural information processing systems, 20","author":"Osindero","year":"2008"},{"issue":"2","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"232","DOI":"10.1109\/69.917563","article-title":"Learning distributed representations of concepts using linear relational embedding","volume":"13","author":"Paccanaro","year":"2001","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1038\/4580","article-title":"Predictive coding in the visual cortex: A functional interpretation of some extra-classical receptive-field effects","volume":"2","author":"Rao","year":"1999","journal-title":"Nature Neuroscience"},{"key":"2023021720442934600_","first-page":"3856","article-title":"Dynamic routing between capsules","volume-title":"Advances in neural information processing systems","author":"Sabour","year":"2017"},{"key":"2023021720442934600_","author":"Sabour","year":"2021","journal-title":"Unsupervised part representation by flow capsules"},{"issue":"8","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"3071","DOI":"10.1073\/pnas.1222618110","article-title":"Hierarchical model of natural images and the origin of scale invariance","volume":"110","author":"Saremi","year":"2013","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2023021720442934600_","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/CVPR.2007.382980","article-title":"Mapping natural image patches by explicit and implicit manifolds","volume-title":"2007 IEEE Conference on Computer Vision and Pattern Recognition","author":"Shi","year":"2007"},{"key":"2023021720442934600_","article-title":"Implicit neural representations with periodic activation functions","volume":"33","author":"Sitzmann","year":"2020","journal-title":"Advances in neural information processing systems"},{"key":"2023021720442934600_","first-page":"1121","article-title":"Scene representation networks: Continuous 3D-structure-aware neural scene representations","volume-title":"Advances in neural information processing systems","author":"Sitzmann","year":"2019"},{"key":"2023021720442934600_","author":"Srivastava","year":"2019","journal-title":"Geometric capsule autoencoders for 3D point clouds"},{"key":"2023021720442934600_","author":"Sun","year":"2020","journal-title":"Canonical capsules: Unsupervised capsules in canonical pose"},{"key":"2023021720442934600_","first-page":"11286","article-title":"ACNe: Attentive context normalization for robust permutation-equivariant learning","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sun","year":"2020"},{"key":"2023021720442934600_","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/7503.003.0173","article-title":"Modeling human motion using binary latent variables","volume-title":"Advances in neural information processing systems","author":"Taylor","year":"2007"},{"key":"2023021720442934600_","author":"Tejankar","year":"2020","journal-title":"ISD: Self-supervised learning by iterative similarity distillation"},{"key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"2109","DOI":"10.1162\/089976600300015088","article-title":"SMEM algorithm for mixture models","volume":"12","author":"Ueda","year":"2000","journal-title":"Neural Computation"},{"key":"2023021720442934600_","author":"van den Oord","year":"2018","journal-title":"Representation learning with contrastive predictive coding"},{"key":"2023021720442934600_","first-page":"5998","article-title":"Attention is all you need","volume-title":"Advances in neural information processing systems","author":"Vaswani","year":"2017"},{"key":"2023021720442934600_","article-title":"Grammar as a foreign language","volume-title":"Advances in neural information processing systems","author":"Vinyals","year":"2014"},{"key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"137","DOI":"10.1023\/B:VISI.0000013087.49260.fb","article-title":"Robust real-time face detection","volume":"57","author":"Viola","year":"2004","journal-title":"International Journal of Computer Vision"},{"issue":"5","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"1169","DOI":"10.1162\/089976602753633439","article-title":"Products of gaussians and probabilistic minor component analysis","volume":"14","author":"Williams","year":"2002","journal-title":"Neural Computation"},{"key":"2023021720442934600_","first-page":"965","article-title":"Using a neural net to instantiate a deformable model","volume-title":"Advances in neural information processing systems","author":"Williams","year":"1995"},{"issue":"4","key":"2023021720442934600_","doi-asserted-by":"publisher","first-page":"503","DOI":"10.1016\/0893-6080(94)00094-3","article-title":"Lending direction to neural networks","volume":"8","author":"Zemel","year":"1995","journal-title":"Neural Networks"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/3\/413\/2072221\/neco_a_01557.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/35\/3\/413\/2072221\/neco_a_01557.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,12,3]],"date-time":"2023-12-03T11:25:12Z","timestamp":1701602712000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/35\/3\/413\/114140\/How-to-Represent-Part-Whole-Hierarchies-in-a"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,17]]},"references-count":64,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,2,17]]},"published-print":{"date-parts":[[2023,2,17]]}},"URL":"https:\/\/doi.org\/10.1162\/neco_a_01557","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2023,3]]},"published":{"date-parts":[[2023,2,17]]}}}