{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T23:12:47Z","timestamp":1763507567419,"version":"3.45.0"},"reference-count":40,"publisher":"MIT Press","issue":"12","content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,11,18]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Humans (and many vertebrates) face the problem of fusing together multiple fixations of a scene in order to obtain a representation of the whole, where each fixation uses a high-resolution fovea and decreasing resolution in the periphery. In this letter, we explicitly represent the retinal transformation of a fixation as a linear downsampling of a high-resolution latent image of the scene, exploiting the known geometry. This linear transformation allows us to carry out exact inference for the latent variables in factor analysis (FA) and mixtures of FA models of the scene. This also allows us to formulate and solve the choice of where to look next as a Bayesian experimental design problem using the expected information gain criterion. Experiments on the Frey faces and MNIST data sets demonstrate the effectiveness of our models.<\/jats:p>","DOI":"10.1162\/neco.a.33","type":"journal-article","created":{"date-parts":[[2025,9,22]],"date-time":"2025-09-22T22:33:33Z","timestamp":1758580413000},"page":"2235-2256","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":0,"title":["Fusing Foveal Fixations Using Linear Retinal Transformations and Bayesian Experimental Design"],"prefix":"10.1162","volume":"37","author":[{"given":"Christopher K. I.","family":"Williams","sequence":"first","affiliation":[{"name":"School of Informatics, University of Edinburgh, EH8 9AB, UK c.k.i.williams@ed.ac.uk"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2025,11,18]]},"reference":[{"volume-title":"Pattern recognition and machine learning","year":"2006","author":"Bishop","key":"2025111818105388100_bib1"},{"key":"2025111818105388100_bib2","doi-asserted-by":"publisher","first-page":"124","DOI":"10.1016\/j.inffus.2021.09.005","article-title":"Real-world single image super-resolution: A brief review","volume":"79","author":"Chen","year":"2022","journal-title":"Information Fusion"},{"key":"2025111818105388100_bib3","doi-asserted-by":"publisher","first-page":"197","DOI":"10.1007\/s10514-017-9634-0","article-title":"A comparison of volumetric information gain metrics for active 3D object reconstruction","volume":"42","author":"Delmerico","year":"2018","journal-title":"Autonomous Robots"},{"key":"2025111818105388100_bib4","doi-asserted-by":"publisher","first-page":"165","DOI":"10.1016\/S0079-6123(02)40049-0","article-title":"Transsaccadic memory of position and form","volume":"140","author":"Deubel","year":"2002","journal-title":"Progress in Brain Research"},{"key":"2025111818105388100_bib5","doi-asserted-by":"publisher","first-page":"1204","DOI":"10.1126\/science.aar6170","article-title":"Neural scene representation and rendering","volume":"360","author":"Eslami","year":"2018","journal-title":"Science"},{"key":"2025111818105388100_bib6","doi-asserted-by":"crossref","DOI":"10.1093\/acprof:oso\/9780198524793.001.0001","volume-title":"Active vision: The psychology of looking and seeing","author":"Findlay","year":"2003"},{"year":"2023","author":"Gao","journal-title":"SceneHGN: Hierarchical graph networks for 3D indoor scene generation with fine-grained geometry","key":"2025111818105388100_bib7"},{"year":"1996","author":"Ghahramani","article-title":"The EM algorithm for mixtures of factor analyzers","key":"2025111818105388100_bib8"},{"key":"2025111818105388100_bib9","doi-asserted-by":"publisher","first-page":"188","DOI":"10.1016\/j.tics.2005.02.009","article-title":"Eye movements in natural behavior","volume":"9","author":"Hayhoe","year":"2005","journal-title":"Trends in Cognitive Sciences"},{"key":"2025111818105388100_bib10","first-page":"309","article-title":"In the mind\u2019s eye","volume-title":"Contemporary theory and research in visual perception","author":"Hochberg","year":"1968"},{"key":"2025111818105388100_bib11","doi-asserted-by":"crossref","DOI":"10.1109\/MFI.2008.4648062","article-title":"On entropy approximation for gaussian mixture random vectors","volume-title":"Proceedings of the 2008 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems","author":"Huber","year":"2008"},{"key":"2025111818105388100_bib12","first-page":"1957","article-title":"Practical approaches to principal component analysis in the presence of missing values","volume":"11","author":"Ilin","year":"2010","journal-title":"Journal of Machine Learning Research"},{"key":"2025111818105388100_bib13","doi-asserted-by":"publisher","first-page":"194","DOI":"10.1038\/35058500","article-title":"Computational modelling of visual attention","volume":"2","author":"Itti","year":"2001","journal-title":"Nature Reviews Neuroscience"},{"key":"2025111818105388100_bib14","first-page":"12949","article-title":"CodeNeRF: Disentangled neural radiance fields for object categories","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Jang","year":"2021"},{"volume-title":"Auto-encoding variational Bayes","year":"2014","author":"Kingma","key":"2025111818105388100_bib15"},{"key":"2025111818105388100_bib16","doi-asserted-by":"publisher","first-page":"611","DOI":"10.1007\/s11554-013-0386-6","article-title":"Efficient next-best-scan planning for autonomous 3D surface reconstruction of unknown objects","volume":"10","author":"Kriegel","year":"2015","journal-title":"Journal of Real-Time Image Processing"},{"year":"2016","author":"K\u00fcmmerer","journal-title":"DeepGaze II: Predicting fixations from deep features over time and tasks","key":"2025111818105388100_bib17"},{"key":"2025111818105388100_bib18","first-page":"1243","article-title":"Learning to combine foveal glimpses with a third-order Boltzmann machine","volume-title":"Advances in neural information processing systems","author":"Larochelle","year":"2010"},{"key":"2025111818105388100_bib19","doi-asserted-by":"publisher","first-page":"12070","DOI":"10.1109\/LRA.2022.3212668","article-title":"Uncertainty guided policy for active robotic 3D reconstruction using neural radiance fields","volume":"7","author":"Lee","year":"2022","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2025111818105388100_bib20","first-page":"834","article-title":"An information-theoretic framework for understanding saccadic eye movements","volume-title":"Advances in neural information processing systems, 12","author":"Lee","year":"2000"},{"key":"2025111818105388100_bib21","first-page":"307","article-title":"Visual stability and voluntary eye movements","volume-title":"The handbook of sensory physiology","author":"MacKay","year":"1973"},{"volume-title":"Multivariate analysis","year":"1979","author":"Mardia","key":"2025111818105388100_bib22"},{"key":"2025111818105388100_bib23","doi-asserted-by":"publisher","first-page":"221","DOI":"10.3758\/BF03202990","article-title":"Is visual information integrated across successive fixations in reading?","volume":"25","author":"McConkie","year":"1979","journal-title":"Perception and Psychophysics"},{"key":"2025111818105388100_bib24","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-58452-8_24","article-title":"NeRF: Representing scenes as neural radiance fields for view synthesis","volume-title":"Computer Vision\u2013ECCV 2020","author":"Mildenhall","year":"2020"},{"key":"2025111818105388100_bib25","article-title":"Recurrent models of visual attention","volume-title":"Advances in neural information processing systems","author":"Mnih","year":"2014"},{"key":"2025111818105388100_bib26","doi-asserted-by":"crossref","first-page":"525","DOI":"10.1016\/S0893-6080(05)80056-5","article-title":"A scaled conjugate gradient algorithm for fast supervised learning","volume":"6","author":"M\u00f8ller","year":"1993","journal-title":"Neural Networks"},{"volume-title":"NETLAB: Algorithms for pattern recognition","year":"2002","author":"Nabney","key":"2025111818105388100_bib27"},{"key":"2025111818105388100_bib28","doi-asserted-by":"publisher","first-page":"387","DOI":"10.1038\/nature03390","article-title":"Optimal eye movement strategies in visual search","volume":"434","author":"Najemnik","year":"2005","journal-title":"Nature"},{"key":"2025111818105388100_bib29","doi-asserted-by":"crossref","DOI":"10.1093\/acprof:oso\/9780199775224.001.0001","volume-title":"Why red doesn\u2019t sound like a bell: Understanding the feel of consciousness","author":"O\u2019Regan","year":"2011"},{"key":"2025111818105388100_bib30","first-page":"12013","article-title":"ATISS: Autoregressive transformers for indoor scene synthesis","volume-title":"Advances in neural information processing systems","author":"Paschalidou","year":"2021"},{"year":"2012","author":"Petersen","article-title":"Matrix cookbook","key":"2025111818105388100_bib31"},{"key":"2025111818105388100_bib32","doi-asserted-by":"publisher","first-page":"100","DOI":"10.1214\/23-STS915","article-title":"Modern Bayesian experimental design","volume":"39","author":"Rainforth","year":"2024","journal-title":"Statistical Science"},{"key":"2025111818105388100_bib33","doi-asserted-by":"publisher","first-page":"1125","DOI":"10.1109\/LRA.2023.3235686","article-title":"NeurAR: Neural uncertainty for autonomous 3D reconstruction with implicit neural representations","volume":"8","author":"Ran","year":"2023","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2025111818105388100_bib34","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1037\/0096-1523.4.4.529","article-title":"Eye movements and integrating information across fixations","volume":"4","author":"Rayner","year":"1978","journal-title":"Journal of Experimental Psychology: Human Perception and Performance"},{"key":"2025111818105388100_bib35","doi-asserted-by":"publisher","first-page":"368","DOI":"10.1111\/j.1467-9280.1997.tb00427.x","article-title":"To see or not to see: The need for attention to perceive changes in scenes","volume":"8","author":"Rensink","year":"1997","journal-title":"Psychological Science"},{"key":"2025111818105388100_bib36","doi-asserted-by":"publisher","first-page":"69","DOI":"10.1007\/BF02293851","article-title":"EM algorithms for ML factor analysis","volume":"47","author":"Rubin","year":"1982","journal-title":"Psychometrika"},{"key":"2025111818105388100_bib37","doi-asserted-by":"publisher","first-page":"611","DOI":"10.1111\/1467-9868.00196","article-title":"Probabilistic principal components analysis","volume":"61","author":"Tipping","year":"1999","journal-title":"Journal of the Royal Statistical Society B"},{"key":"2025111818105388100_bib38","doi-asserted-by":"publisher","first-page":"2845","DOI":"10.1007\/s11263-024-02316-z","article-title":"Structured generative models for scene understanding","volume":"133","author":"Williams","year":"2024","journal-title":"International Journal of Computer Vision"},{"key":"2025111818105388100_bib39","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4899-5379-7","volume-title":"Eye movements and vision","author":"Yarbus","year":"1967"},{"key":"2025111818105388100_bib40","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR42600.2020.00950","article-title":"Towards robust image classification using sequential attention models","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Zoran","year":"2020"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/37\/12\/2235\/2554624\/neco.a.33.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/neco\/article-pdf\/37\/12\/2235\/2554624\/neco.a.33.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T23:11:04Z","timestamp":1763507464000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/37\/12\/2235\/133238\/Fusing-Foveal-Fixations-Using-Linear-Retinal"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,11,18]]},"references-count":40,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2025,11,18]]},"published-print":{"date-parts":[[2025,11,18]]}},"URL":"https:\/\/doi.org\/10.1162\/neco.a.33","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"type":"print","value":"0899-7667"},{"type":"electronic","value":"1530-888X"}],"subject":[],"published-other":{"date-parts":[[2025,12]]},"published":{"date-parts":[[2025,11,18]]}}}