{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T14:03:51Z","timestamp":1783001031953,"version":"3.54.5"},"reference-count":110,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2023,6,29]],"date-time":"2023-06-29T00:00:00Z","timestamp":1687996800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000181","name":"Air Force Office of Scientific Research","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000181","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001804","name":"Canada Research Chairs","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100001804","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100000038","name":"Natural Sciences and Engineering Research Council of Canada","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Comput. Sci."],"abstract":"<jats:p>Recently, a considerable number of studies in computer vision involve deep neural architectures called vision transformers. Visual processing in these models incorporates computational models that are claimed to implement attention mechanisms. Despite an increasing body of work that attempts to understand the role of attention mechanisms in vision transformers, their effect is largely unknown. Here, we asked if the attention mechanisms in vision transformers exhibit similar effects as those known in human visual attention. To answer this question, we revisited the attention formulation in these models and found that despite the name, computationally, these models perform a special class of relaxation labeling with similarity grouping effects. Additionally, whereas modern experimental findings reveal that human visual attention involves both feed-forward and feedback mechanisms, the purely feed-forward architecture of vision transformers suggests that attention in these models cannot have the same effects as those known in humans. To quantify these observations, we evaluated grouping performance in a family of vision transformers. Our results suggest that self-attention modules group figures in the stimuli based on similarity of visual features such as color. Also, in a singleton detection experiment as an instance of salient object detection, we studied if these models exhibit similar effects as those of feed-forward visual salience mechanisms thought to be utilized in human visual attention. We found that generally, the transformer-based attention modules assign more salience either to distractors or the ground, the opposite of both human and computational salience. Together, our study suggests that the mechanisms in vision transformers perform perceptual organization based on feature similarity and not attention.<\/jats:p>","DOI":"10.3389\/fcomp.2023.1178450","type":"journal-article","created":{"date-parts":[[2023,6,29]],"date-time":"2023-06-29T19:58:15Z","timestamp":1688068695000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":40,"title":["Self-attention in vision transformers performs perceptual grouping, not attention"],"prefix":"10.3389","volume":"5","author":[{"given":"Paria","family":"Mehrani","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"John K.","family":"Tsotsos","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2023,6,29]]},"reference":[{"key":"B1","doi-asserted-by":"crossref","first-page":"4190","DOI":"10.18653\/v1\/2020.acl-main.385","article-title":"\u201cQuantifying attention flow in transformers,\u201d","author":"Abnar","year":"2020","journal-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics"},{"key":"B2","doi-asserted-by":"publisher","DOI":"10.1002\/wcs.1574","article-title":"Stop paying attention to \u201cattention\u201d","author":"Anderson","year":"2023","journal-title":"Wiley Interdiscip. Rev. Cogn. Sci"},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.7554\/eLife.74943","article-title":"Perception of an object's global shape is best described by a model of skeletal structure in human infants","author":"Ayzenberg","year":"2022","journal-title":"Elife"},{"key":"B4","doi-asserted-by":"publisher","first-page":"485","DOI":"10.3758\/BF03205306","article-title":"Overriding stimulus-driven attentional capture","volume":"55","author":"Bacon","year":"1994","journal-title":"Percept. Psychophys"},{"key":"B5","doi-asserted-by":"publisher","first-page":"46","DOI":"10.1016\/j.visres.2020.04.003","article-title":"Local features and global shape information in object classification by deep convolutional neural networks","volume":"172","author":"Baker","year":"2020","journal-title":"Vision Res"},{"key":"B6","doi-asserted-by":"publisher","first-page":"210","DOI":"10.1016\/j.tins.2011.02.003","article-title":"Mechanisms of top-down attention","volume":"34","author":"Baluch","year":"2011","journal-title":"Trends Neurosci"},{"key":"B7","article-title":"\u201cBEit: BERT pre-training of image transformers,\u201d","author":"Bao","year":"2022","journal-title":"International Conference on Learning Representations"},{"key":"B8","article-title":"\u201cAttention,\u201d","author":"Berlyne","year":"1974","journal-title":"Handbook of Perception., Chapter 8"},{"key":"B9","first-page":"10231","article-title":"\u201cUnderstanding robustness of transformers for image classification,\u201d","author":"Bhojanapalli","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)"},{"key":"B10","doi-asserted-by":"publisher","first-page":"185","DOI":"10.1109\/TPAMI.2012.89","article-title":"State-of-the-art in visual attention modeling","volume":"35","author":"Borji","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"B11","doi-asserted-by":"publisher","first-page":"95","DOI":"10.1016\/j.visres.2015.01.010","article-title":"On computational modeling of visual saliency: examining what's right, and what's left","volume":"116","author":"Bruce","year":"2015","journal-title":"Vision Res"},{"key":"B12","doi-asserted-by":"publisher","first-page":"258","DOI":"10.1016\/j.visres.2015.04.007","article-title":"Towards the quantitative evaluation of visual attention models","volume":"116","author":"Bylinskii","year":"2015","journal-title":"Vision Res"},{"key":"B13","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1003963","article-title":"Deep neural networks rival the representation of primate it cortex for core visual object recognition","author":"Cadieu","year":"2014","journal-title":"PLoS Comput. Biol"},{"key":"B14","first-page":"9650","article-title":"\u201cEmerging properties in self-supervised vision transformers,\u201d","author":"Caron","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)"},{"key":"B15","doi-asserted-by":"publisher","first-page":"1484","DOI":"10.1016\/j.visres.2011.04.012","article-title":"Visual attention: the past 25 years","volume":"51","author":"Carrasco","year":"2011","journal-title":"Vision Res"},{"key":"B16","first-page":"357","article-title":"\u201cCrossvit: cross-attention multi-scale vision transformer for image classification,\u201d","author":"Chen","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"},{"key":"B17","doi-asserted-by":"publisher","first-page":"R850","DOI":"10.1016\/j.cub.2004.09.041","article-title":"Visual attention: bottom-up versus top-down","volume":"14","author":"Connor","year":"2004","journal-title":"Curr. Biol"},{"key":"B18","article-title":"\u201cOn the relationship between self-attention and convolutional layers,\u201d","author":"Cordonnier","year":"2020","journal-title":"Eighth International Conference on Learning Representations-ICLR 2020, number CONF"},{"key":"B19","first-page":"3965","article-title":"\u201cCoAtNet: marrying convolution and attention for all data sizes,\u201d","author":"Dai","year":"2021","journal-title":"Advances in Neural Information Processing Systems, Vol. 34"},{"key":"B20","first-page":"2286","article-title":"\u201cConViT: improving vision transformers with soft convolutional inductive biases,\u201d","author":"D'Ascoli","year":"2021","journal-title":"Proceedings of the 38th International Conference on Machine Learning, Vol. 139"},{"key":"B21","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1146\/annurev.ne.18.030195.001205","article-title":"Neural mechanisms of selective visual attention","volume":"18","author":"Desimone","year":"1995","journal-title":"Ann. Rev. Neurosci"},{"key":"B22","doi-asserted-by":"publisher","first-page":"45","DOI":"10.1016\/j.concog.2018.02.005","article-title":"Attention is a sterile concept; iterative reentry is a fertile substitute","volume":"64","author":"Di Lollo","year":"2018","journal-title":"Conscious. Cogn"},{"key":"B23","first-page":"1","article-title":"\u201cA study and comparison of human and deep learning recognition performance under visual distortions,\u201d","author":"Dodge","year":"2017","journal-title":"2017 26th International Conference on Computer Communication and Networks (ICCCN)"},{"key":"B24","article-title":"\u201cAn image is worth 16x16 words: transformers for image recognition at scale,\u201d","author":"Dosovitskiy","year":"2021","journal-title":"International Conference on Learning Representations"},{"key":"B25","doi-asserted-by":"publisher","first-page":"184","DOI":"10.1016\/j.neuroimage.2016.10.001","article-title":"Seeing it all: convolutional network layers map the function of the human visual system","volume":"152","author":"Eickenberg","year":"2017","journal-title":"Neuroimage"},{"key":"B26","doi-asserted-by":"publisher","DOI":"10.32470\/CCN.2022.1147-0","article-title":"Model metamers illuminate divergences between biological and artificial neural networks","author":"Feather","year":"2022","journal-title":"bioRxiv, pages"},{"key":"B27","article-title":"\u201cHarmonizing the object recognition strategies of deep neural networks with humans,\u201d","author":"Fel","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B28","doi-asserted-by":"publisher","DOI":"10.1037\/0096-1523.18.4.1030","article-title":"Involuntary covert orienting is contingent on attentional control settings","author":"Folk","year":"1992","journal-title":"J. Exp. Psychol. Hum. Percept. Perform"},{"key":"B29","author":"Geirhos","year":"2019"},{"key":"B30","unstructured":"GhiasiA.\n            KazemiH.\n            BorgniaE.\n            ReichS.\n            ShuM.\n            GoldbulmA.\n          What do vision transformers learn? A visual exploration. arXiv [Preprint]. 2022"},{"key":"B31","doi-asserted-by":"publisher","DOI":"10.3389\/fncom.2014.00074","article-title":"Feedforward object-vision models only tolerate small image variations compared to human","author":"Ghodrati","year":"2014","journal-title":"Front. Comput. Neurosci"},{"key":"B32","doi-asserted-by":"crossref","first-page":"12165","DOI":"10.1109\/CVPR52688.2022.01186","author":"Guo","year":"2022","journal-title":"2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"B33","first-page":"87","article-title":"\u201cA survey on vision transformer,\u201d","author":"Han","year":"2022","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"B34","author":"Hanson","year":"1978","journal-title":"Computer Vision Systems: Papers from the Workshop on Computer Vision Systems"},{"key":"B35","first-page":"770","article-title":"\u201cDeep residual learning for image recognition,\u201d","author":"He","year":"2016","journal-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016"},{"key":"B36","doi-asserted-by":"publisher","DOI":"10.3389\/fncom.2014.00135","article-title":"Why vision is not both hierarchical and feedforward","author":"Herzog","year":"2014","journal-title":"Front. Comput. Neurosci"},{"key":"B37","doi-asserted-by":"publisher","first-page":"288","DOI":"10.3758\/s13414-019-01846-w","article-title":"No one knows what attention is","volume":"81","author":"Hommel","year":"2019","journal-title":"Attent. Percept. Psychophys"},{"key":"B38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/sdata.2019.12","article-title":"Characterization of deep neural network features by decodability from human brain activity","volume":"6","author":"Horikawa","year":"2019","journal-title":"Sci. Data"},{"key":"B39","first-page":"1","article-title":"\u201cFigure-ground representation in deep neural networks,\u201d","author":"Hu","year":"2019","journal-title":"2019 53rd Annual Conference on Information Sciences and Systems (CISS)"},{"key":"B40","doi-asserted-by":"publisher","DOI":"10.4249\/scholarpedia.3327","article-title":"Visual salience","author":"Itti","year":"2007","journal-title":"Scholarpedia"},{"key":"B41","author":"Itti","year":"2005","journal-title":"Neurobiology of Attention"},{"key":"B42","doi-asserted-by":"publisher","first-page":"315","DOI":"10.1146\/annurev.neuro.23.1.315","article-title":"Mechanisms of visual attention in the human cortex","volume":"23","author":"Kastner","year":"2000","journal-title":"Annu. Rev. Neurosci"},{"key":"B43","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1003915","article-title":"Deep supervised, but not unsupervised, models may explain it cortical representation","author":"Khaligh-Razavi","year":"2014","journal-title":"PLoS Comput. Biol"},{"key":"B44","doi-asserted-by":"publisher","first-page":"20180011","DOI":"10.1098\/rsfs.2018.0011","article-title":"Not-so-clevr: learning same-different relations strains feedforward neural networks","volume":"8","author":"Kim","year":"2018","journal-title":"Interface Focus"},{"key":"B45","doi-asserted-by":"publisher","first-page":"1009","DOI":"10.3758\/BF03207609","article-title":"Top-down and bottom-up attentional control: on the nature of interference from a salient distractor","volume":"61","author":"Kim","year":"1999","journal-title":"Percept. Psychophys"},{"key":"B46","article-title":"\u201cStructured attention networks,\u201d","author":"Kim","year":"2017","journal-title":"International Conference on Learning Representations"},{"key":"B47","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1146\/annurev.neuro.30.051606.094256","article-title":"Fundamental components of attention","volume":"30","author":"Knudsen","year":"2007","journal-title":"Annu. Rev. Neurosci"},{"key":"B48","article-title":"\u201cDo saliency models detect odd-one-out targets? New datasets and evaluations,\u201d","author":"Kotseruba","year":"2019","journal-title":"British Machine Vision Conference (BMVC)"},{"key":"B49","doi-asserted-by":"publisher","DOI":"10.1002\/wcs.1570","article-title":"What is attention?","author":"Krauzlis","year":"2023","journal-title":"Wiley Interdiscip. Rev. Cogn. Sci"},{"key":"B50","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1004896","article-title":"Deep neural networks as a computational model for human shape sensitivity","author":"Kubilius","year":"2016","journal-title":"PLoS Comput. Biol"},{"key":"B51","doi-asserted-by":"publisher","first-page":"621","DOI":"10.3758\/BF03196524","article-title":"Does a salient distractor capture attention early in processing?","volume":"10","author":"Lamy","year":"2003","journal-title":"Psychonom. Bull. Rev"},{"key":"B52","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2023.3261935","article-title":"How does attention work in vision transformers? A visual analytics attempt","author":"Li","year":"2023","journal-title":"IEEE Transact. Vis. Comp. Graph"},{"key":"B53","doi-asserted-by":"publisher","first-page":"13","DOI":"10.1167\/jov.22.2.13","article-title":"The role of lateral modulation in orientation-specific adaptation effect","volume":"22","author":"Lin","year":"2022","journal-title":"J. Vis"},{"key":"B54","first-page":"10012","article-title":"\u201cSwin transformer: hierarchical vision transformer using shifted windows,\u201d","author":"Liu","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)"},{"key":"B55","first-page":"11976","article-title":"\u201cA convnet for the 2020s,\u201d","author":"Liu","year":"2022","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B56","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1167\/jov.21.10.17","article-title":"A comparative biology approach to dnn modeling of vision: a focus on differences, not similarities","volume":"21","author":"Lonnqvist","year":"2021","journal-title":"J. Vis"},{"key":"B57","first-page":"7838","article-title":"\u201cOn the robustness of vision transformers to adversarial examples,\u201d","author":"Mahmood","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)"},{"key":"B58","doi-asserted-by":"publisher","first-page":"407","DOI":"10.1146\/annurev-vision-100720-031711","article-title":"Visual attention in the prefrontal cortex","volume":"8","author":"Martinez-Trujillo","year":"2022","journal-title":"Ann. Rev. Vision Sci"},{"key":"B59","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1146\/annurev-psych-122414-033400","article-title":"Neural mechanisms of selective visual attention","volume":"68","author":"Moore","year":"2017","journal-title":"Annu. Rev. Psychol"},{"key":"B60","first-page":"23296","article-title":"Intriguing properties of vision transformers","volume":"34","author":"Naseer","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst"},{"key":"B61","doi-asserted-by":"crossref","DOI":"10.1093\/oxfordhb\/9780199675111.001.0001","author":"Nobre","year":"2014","journal-title":"The Oxford Handbook of Attention"},{"key":"B62","doi-asserted-by":"crossref","first-page":"2035","DOI":"10.1609\/aaai.v36i2.20099","article-title":"\u201cLess is more: Pay less attention in vision transformers,\u201d","author":"Pan","year":"2022","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"B63","doi-asserted-by":"crossref","first-page":"629","DOI":"10.1007\/978-3-031-26284-5_38","article-title":"\u201cRdrn: recursively defined residual network for image super-resolution,\u201d","author":"Panaetov","year":"2023","journal-title":"Computer Vision-ACCV 2022"},{"key":"B64","author":"Park","year":"2022"},{"key":"B65","author":"Pashler","year":"1998","journal-title":"Attention, 1st Edn"},{"key":"B66","doi-asserted-by":"publisher","first-page":"2071","DOI":"10.1609\/aaai.v36i2.20103","article-title":"Vision transformers are robust learners","volume":"36","author":"Paul","year":"2022","journal-title":"Proc. AAAI Conf. Artif. Intell"},{"key":"B67","first-page":"259","article-title":"\u201cLow-level and high-level contributions to figure-ground organization,\u201d","author":"Peterson","year":"2015","journal-title":"The Oxford Handbook of Perceptual Organization"},{"key":"B68","doi-asserted-by":"publisher","first-page":"143","DOI":"10.1016\/j.neuron.2012.04.032","article-title":"The role of attention in figure-ground segregation in areas v1 and v4 of the visual cortex","volume":"75","author":"Poort","year":"2012","journal-title":"Neuron"},{"key":"B69","doi-asserted-by":"publisher","first-page":"1492","DOI":"10.1038\/nn1989","article-title":"Figure-ground mechanisms provide structure for selective attention","volume":"10","author":"Qiu","year":"2007","journal-title":"Nat. Neurosci"},{"key":"B70","first-page":"12116","article-title":"Do vision transformers see like convolutional neural networks?","volume":"34","author":"Raghu","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst"},{"key":"B71","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1016\/j.cobeha.2020.08.008","article-title":"Same-different conceptualization: a machine vision perspective","volume":"37","author":"Ricci","year":"2021","journal-title":"Curr. Opin. Behav. Sci"},{"key":"B72","doi-asserted-by":"publisher","first-page":"2280","DOI":"10.1109\/TPAMI.2018.2849989","article-title":"Psyphy: a psychophysics driven evaluation framework for visual recognition","volume":"41","author":"RichardWebster","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"B73","doi-asserted-by":"publisher","first-page":"106","DOI":"10.1523\/JNEUROSCI.2518-12.2013","article-title":"Different orientation tuning of near-and far-surround suppression in macaque primary visual cortex mirrors their tuning in human perception","volume":"33","author":"Shushruth","year":"2013","journal-title":"J. Neurosci"},{"key":"B74","first-page":"16519","article-title":"\u201cBottleneck transformers for visual recognition,\u201d","author":"Srinivas","year":"2021","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)"},{"key":"B75","doi-asserted-by":"publisher","first-page":"739","DOI":"10.1016\/S0896-6273(02)01029-2","article-title":"Lateral connectivity and contextual interactions in macaque primary visual cortex","volume":"36","author":"Stettler","year":"2002","journal-title":"Neuron"},{"key":"B76","doi-asserted-by":"crossref","DOI":"10.4324\/9780203968215","author":"Styles","year":"2006","journal-title":"The Psychology of Attention"},{"key":"B77","doi-asserted-by":"crossref","first-page":"350","DOI":"10.1038\/32817","article-title":"Feature selection","volume":"392","author":"Sutherland","year":"1988","journal-title":"Nature"},{"key":"B78","doi-asserted-by":"publisher","first-page":"9799","DOI":"10.1609\/aaai.v35i11.17178","article-title":"Explicitly modeled attention maps for image classification","volume":"35","author":"Tan","year":"2021","journal-title":"Proc. AAAI Conf. Artif. Intell"},{"key":"B79","first-page":"10347","article-title":"\u201cTraining data-efficient image transformers and distillation through attention,\u201d","author":"Touvron","year":"2021","journal-title":"Proceedings of the 38th International Conference on Machine Learning"},{"key":"B80","doi-asserted-by":"publisher","first-page":"423","DOI":"10.1017\/S0140525X00079577","article-title":"Analyzing vision at the complexity level","volume":"13","author":"Tsotsos","year":"1990","journal-title":"Behav. Brain Sci"},{"key":"B81","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/9780262015417.001.0001","author":"Tsotsos","year":"2011","journal-title":"A Computational Perspective on Visual Attention"},{"key":"B82","doi-asserted-by":"publisher","DOI":"10.3389\/fpsyg.2017.01216","article-title":"Complexity level analysis revisited: What can 30 years of hindsight tell us about how the brain might represent visual information?","author":"Tsotsos","year":"2017","journal-title":"Front. Psychol"},{"key":"B83","doi-asserted-by":"publisher","first-page":"212","DOI":"10.3390\/jimaging8080212","article-title":"When we study the ability to attend, what exactly are we trying to understand?","volume":"8","author":"Tsotsos","year":"2022","journal-title":"J. Imaging"},{"key":"B84","doi-asserted-by":"crossref","DOI":"10.1016\/B978-012375731-9\/50003-3","article-title":"\u201cA brief and selective history of attention,\u201d","author":"Tsotsos","year":"2005","journal-title":"Neurobiology of Attention"},{"key":"B85","doi-asserted-by":"publisher","first-page":"6201","DOI":"10.4249\/scholarpedia.6201","article-title":"Computational models of visual attention","volume":"6","author":"Tsotsos","year":"2011","journal-title":"Scholarpedia"},{"key":"B86","article-title":"\u201cAre convolutional neural networks or transformers more like human vision?,\u201d","author":"Tuli","year":"2021","journal-title":"Proceedings of the Annual Meeting of the Cognitive Science Society, Vol. 43"},{"key":"B87","doi-asserted-by":"publisher","first-page":"1075","DOI":"10.1162\/neco_a_01485","article-title":"Understanding the computational demands underlying visual reasoning","volume":"34","author":"Vaishnav","year":"2022","journal-title":"Neural Comput"},{"key":"B88","article-title":"\u201cAttention is All you need,\u201d","author":"Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B89","doi-asserted-by":"publisher","first-page":"318","DOI":"10.1167\/19.10.318","article-title":"Flipped on its head: deep learning-based saliency finds asymmetry in the opposite direction expected for singleton search of flipped and canonical targets","volume":"19","author":"Wloka","year":"2019","journal-title":"J. Vis"},{"key":"B90","first-page":"38","article-title":"\u201cTransformers: state-of-the-art natural language processing,\u201d","author":"Wolf","year":"2020","journal-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations"},{"key":"B91","first-page":"599","article-title":"\u201cVisual transformers: where do transformers really belong in vision models?,\u201d","author":"Wu","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)"},{"key":"B92","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1109\/ICCV48922.2021.00009","article-title":"\u201cCvT: introducing convolutions to vision transformers,\u201d","author":"Wu","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)"},{"key":"B93","first-page":"30392","author":"Xiao","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B94","first-page":"2048","article-title":"\u201cShow, attend and tell: neural image caption generation with visual attention,\u201d","author":"Xu","year":"2015","journal-title":"Proceedings of the 32nd International Conference on Machine Learning"},{"key":"B95","doi-asserted-by":"publisher","first-page":"4234","DOI":"10.1523\/JNEUROSCI.1993-20.2021","article-title":"Examining the coding strength of object identity and nonidentity features in human occipito-temporal cortex and convolutional neural networks","volume":"41","author":"Xu","year":"","journal-title":"J. Neurosci"},{"key":"B96","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-021-22244-7","article-title":"Limits to visual representational correspondence between convolutional neural networks and the human brain","author":"Xu","year":"","journal-title":"Nat. Commun"},{"key":"B97","doi-asserted-by":"publisher","first-page":"119635","DOI":"10.1016\/j.neuroimage.2022.119635","article-title":"Understanding transformation tolerant visual object representations in the human brain and convolutional neural networks","volume":"263","author":"Xu","year":"2022","journal-title":"Neuroimage"},{"key":"B98","first-page":"30008","article-title":"\u201cFocal attention for long-range interactions in vision transformers,\u201d","author":"Yang","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"B99","doi-asserted-by":"publisher","DOI":"10.1037\/0096-1523.25.3.661","article-title":"On the distinction between visual salience and stimulus-driven attentional capture","author":"Yantis","year":"1999","journal-title":"J. Exp. Psychol"},{"key":"B100","first-page":"558","article-title":"\u201cTokens-to-token ViT: training vision transformers from scratch on ImageNet,\u201d","author":"Yuan","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV)"},{"key":"B101","first-page":"387","article-title":"\u201cVision transformer with progressive sampling,\u201d","author":"Yue","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"},{"key":"B102","article-title":"\u201cA benchmark for compositional visual reasoning,\u201d","author":"Zerroug","year":"2022","journal-title":"Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track"},{"key":"B103","first-page":"153","article-title":"\u201cSaliency detection: a boolean map approach,\u201d","author":"Zhang","year":"2013","journal-title":"Proceedings of the IEEE International Conference on Computer Vision"},{"key":"B104","first-page":"27378","article-title":"\u201cUnderstanding the robustness in vision transformers,\u201d","author":"Zhou","year":"2022","journal-title":"Proceedings of the 39th International Conference on Machine Learning"},{"key":"B105","first-page":"2230","article-title":"\u201cConvNets vs. transformers: whose visual representations are more transferable?,\u201d","author":"Zhou","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) Workshops"},{"key":"B106","doi-asserted-by":"publisher","first-page":"439","DOI":"10.1007\/s11633-022-1348-x","article-title":"Exploring the brain-like properties of deep neural networks: a neural encoding perspective","volume":"19","author":"Zhou","year":"2022","journal-title":"Mach. Intell. Res"},{"key":"B107","first-page":"1953","article-title":"\u201cSaliency-guided transformer network combined with local embedding for no-reference image quality assessment,\u201d","author":"Zhu","year":"2021","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"},{"key":"B108","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2014196118","article-title":"Unsupervised neural network models of the ventral visual stream","author":"Zhuang","year":"2021","journal-title":"Proc. Nat. Acad. Sci. U. S. A"},{"key":"B109","first-page":"1102","article-title":"\u201cComputer vision and human perception,\u201d","author":"Zucker","year":"1981","journal-title":"Proceedings of the 7th International Joint Conference on Artificial Intelligence"},{"key":"B110","unstructured":"ZuckerS. W.\n          Vertical and horizontal processes in low level vision. 1978"}],"container-title":["Frontiers in Computer Science"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2023.1178450\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,29]],"date-time":"2023-06-29T19:58:59Z","timestamp":1688068739000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fcomp.2023.1178450\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,29]]},"references-count":110,"alternative-id":["10.3389\/fcomp.2023.1178450"],"URL":"https:\/\/doi.org\/10.3389\/fcomp.2023.1178450","relation":{},"ISSN":["2624-9898"],"issn-type":[{"value":"2624-9898","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,29]]},"article-number":"1178450"}}