{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T14:52:45Z","timestamp":1784040765323,"version":"3.55.0"},"publisher-location":"Cham","reference-count":36,"publisher":"Springer Nature Switzerland","isbn-type":[{"value":"9783032083166","type":"print"},{"value":"9783032083173","type":"electronic"}],"license":[{"start":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T00:00:00Z","timestamp":1760227200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T00:00:00Z","timestamp":1760227200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Vision Transformers (ViTs) are increasingly utilized in various computer vision tasks due to their powerful representation capabilities. However, it remains understudied how ViTs process information layer by layer. Numerous studies have shown that convolutional neural networks (CNNs) extract features of increasing complexity throughout their layers, which is crucial for tasks like domain adaptation and transfer learning. ViTs, lacking the same inductive biases as CNNs, can potentially learn global dependencies from the first layers due to their attention mechanisms. Given the increasing importance of ViTs in computer vision, there is a need to improve the layer-wise understanding of ViTs. In this work, we present a novel, layer-wise analysis of concepts encoded in state-of-the-art ViTs using neuron labeling. Our findings reveal that ViTs encode concepts with increasing complexity throughout the network. Early layers primarily encode basic features such as colors and textures, while later layers represent more specific classes, including objects and animals. As the complexity of encoded concepts increases, the number of concepts represented in each layer also rises, reflecting a more diverse and specific set of features. Additionally, different pretraining strategies influence the quantity and category of encoded concepts, with finetuning to specific downstream tasks generally reducing the number of encoded concepts and shifting the concepts to more relevant categories.\n<\/jats:p>","DOI":"10.1007\/978-3-032-08317-3_2","type":"book-chapter","created":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T03:36:39Z","timestamp":1760153799000},"page":"28-47","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["From Colors to\u00a0Classes: Emergence of\u00a0Concepts in\u00a0Vision Transformers"],"prefix":"10.1007","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-0283-2508","authenticated-orcid":false,"given":"Teresa","family":"Dorszewski","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0009-6896","authenticated-orcid":false,"given":"Lenka","family":"T\u011btkov\u00e1","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7496-8474","authenticated-orcid":false,"given":"Robert","family":"Jenssen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0442-5877","authenticated-orcid":false,"given":"Lars Kai","family":"Hansen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1395-7154","authenticated-orcid":false,"given":"Kristoffer Knutsen","family":"Wickstr\u00f8m","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2025,10,12]]},"reference":[{"key":"2_CR1","doi-asserted-by":"publisher","DOI":"10.1016\/j.dib.2020.105474","volume":"30","author":"A Acevedo","year":"2020","unstructured":"Acevedo, A., Merino, A., Alf\u00e9rez, S., Molina, \u00c1., Bold\u00fa, L., Rodellar, J.: A dataset of microscopic peripheral blood cell images for development of automatic recognition systems. Data Brief 30, 105474 (2020). https:\/\/doi.org\/10.1016\/j.dib.2020.105474","journal-title":"Data Brief"},{"key":"2_CR2","doi-asserted-by":"publisher","unstructured":"Alain, G., Bengio, Y.: Understanding intermediate layers using linear classifier probes. arXiv preprint arXiv:1610.01644 (2016). https:\/\/doi.org\/10.48550\/arXiv.1610.01644","DOI":"10.48550\/arXiv.1610.01644"},{"key":"2_CR3","doi-asserted-by":"crossref","unstructured":"Bau, D., Zhou, B., Khosla, A., Oliva, A., Torralba, A.: Network dissection: quantifying interpretability of deep visual representations. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6541\u20136549 (2017)","DOI":"10.1109\/CVPR.2017.354"},{"key":"2_CR4","unstructured":"Bykov, K., Kopf, L., Nakajima, S., Kloft, M., H\u00f6hne, M.: Labeling neural representations with inverse recognition. Adv. Neural Inf. Process. Syst. 36 (2024)"},{"key":"2_CR5","unstructured":"Dosovitskiy, A., et\u00a0al.: An image is worth 16x16 words: transformers for image recognition at scale. In: International Conference on Learning Representations (ICLR) (2021)"},{"key":"2_CR6","doi-asserted-by":"crossref","unstructured":"Fel, T., et\u00a0al.: Craft: concept recursive activation factorization for explainability. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 2711\u20132721 (2023)","DOI":"10.1109\/CVPR52729.2023.00266"},{"key":"2_CR7","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TPAMI.2022.3232328","volume":"01","author":"T Feng","year":"2023","unstructured":"Feng, T., et al.: IC9600: a benchmark dataset for automatic image complexity assessment. IEEE Trans. Pattern Anal. Mach. Intell. 01, 1\u201317 (2023). https:\/\/doi.org\/10.1109\/TPAMI.2022.3232328","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"issue":"11","key":"2_CR8","doi-asserted-by":"publisher","first-page":"665","DOI":"10.1038\/s42256-020-00257-z","volume":"2","author":"R Geirhos","year":"2020","unstructured":"Geirhos, R., et al.: Shortcut learning in deep neural networks. Nat. Mach. Intell. 2(11), 665\u2013673 (2020). https:\/\/doi.org\/10.1038\/s42256-020-00257-z","journal-title":"Nat. Mach. Intell."},{"issue":"1","key":"2_CR9","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1109\/TPAMI.2022.3152247","volume":"45","author":"K Han","year":"2022","unstructured":"Han, K., et al.: A survey on vision transformer. IEEE Trans. Pattern Anal. Mach. Intell. 45(1), 87\u2013110 (2022). https:\/\/doi.org\/10.1109\/TPAMI.2022.3152247","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"2_CR10","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., Girshick, R.: Masked autoencoders are scalable vision learners. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 16000\u201316009 (2022)","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"2_CR11","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770\u2013778 (2016)","DOI":"10.1109\/CVPR.2016.90"},{"key":"2_CR12","unstructured":"Hernandez, E., Schwettmann, S., Bau, D., Bagashvili, T., Torralba, A., Andreas, J.: Natural language descriptions of deep visual features. In: International Conference on Learning Representations (2021)"},{"key":"2_CR13","unstructured":"Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., Madry, A.: Adversarial examples are not bugs, they are features. Adv. Neural Inf. Process. Syst. 32 (2019)"},{"key":"2_CR14","unstructured":"Kim, B., et\u00a0al.: Interpretability beyond feature attribution: quantitative testing with concept activation vectors (TCAV). In: International Conference on Machine Learning, pp. 2668\u20132677. PMLR (2018)"},{"key":"2_CR15","doi-asserted-by":"publisher","unstructured":"Kirkpatrick, J., et al.: Overcoming catastrophic forgetting in neural networks. Proc. Natl. Acad. Sci. 114(13), 3521\u20133526 (2017). https:\/\/doi.org\/10.1073\/pnas.1611835114. https:\/\/www.pnas.org\/doi\/abs\/10.1073\/pnas.1611835114","DOI":"10.1073\/pnas.1611835114"},{"key":"2_CR16","first-page":"34656","volume":"37","author":"L Kopf","year":"2024","unstructured":"Kopf, L., Bommer, P.L., Hedstr\u00f6m, A., Lapuschkin, S., H\u00f6hne, M., Bykov, K.: CoSy: evaluating textual explanations of neurons. Adv. Neural. Inf. Process. Syst. 37, 34656\u201334685 (2024)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"2_CR17","doi-asserted-by":"publisher","unstructured":"Li, Y., Yosinski, J., Clune, J., Lipson, H., Hopcroft, J.: Convergent learning: do different neural networks learn the same representations? arXiv preprint arXiv:1511.07543 (2015). https:\/\/doi.org\/10.48550\/arXiv.1511.07543","DOI":"10.48550\/arXiv.1511.07543"},{"key":"2_CR18","unstructured":"Microsoft: Microsoft copilot. Generated using https:\/\/copilot.microsoft.com. Accessed 13 Feb 2025"},{"key":"2_CR19","first-page":"23296","volume":"34","author":"MM Naseer","year":"2021","unstructured":"Naseer, M.M., Ranasinghe, K., Khan, S.H., Hayat, M., Shahbaz Khan, F., Yang, M.H.: Intriguing properties of vision transformers. Adv. Neural. Inf. Process. Syst. 34, 23296\u201323308 (2021)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"key":"2_CR20","unstructured":"Oikarinen, T., Weng, T.W.: Clip-dissect: automatic description of neuron representations in deep vision networks. In: The Eleventh International Conference on Learning Representations (2023)"},{"key":"2_CR21","doi-asserted-by":"publisher","unstructured":"Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., Carter, S.: Zoom in: an introduction to circuits. Distill (2020). https:\/\/doi.org\/10.23915\/distill.00024.001. https:\/\/distill.pub\/2020\/circuits\/zoom-in","DOI":"10.23915\/distill.00024.001"},{"key":"2_CR22","doi-asserted-by":"publisher","unstructured":"Olah, C., et al.: The building blocks of interpretability. Distill 3(3), e10 (2018). https:\/\/doi.org\/10.23915\/distill.00010","DOI":"10.23915\/distill.00010"},{"key":"2_CR23","unstructured":"OpenAI: ChatGPT, model GPT-4O (2023). https:\/\/chat.openai.com. Accessed 13 Feb 2025"},{"key":"2_CR24","doi-asserted-by":"publisher","unstructured":"Oquab, M., et\u00a0al.: DINOv2: learning robust visual features without supervision. arXiv preprint arXiv:2304.07193 (2023https:\/\/doi.org\/10.48550\/arXiv.2304.07193","DOI":"10.48550\/arXiv.2304.07193"},{"key":"2_CR25","unstructured":"Radford, A., et\u00a0al.: Learning transferable visual models from natural language supervision. In: International Conference on Machine Learning, pp. 8748\u20138763. PMLR (2021)"},{"key":"2_CR26","first-page":"12116","volume":"34","author":"M Raghu","year":"2021","unstructured":"Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C., Dosovitskiy, A.: Do vision transformers see like convolutional neural networks? Adv. Neural. Inf. Process. Syst. 34, 12116\u201312128 (2021)","journal-title":"Adv. Neural. Inf. Process. Syst."},{"issue":"3","key":"2_CR27","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky, O., et al.: ImageNet large scale visual recognition challenge. Int. J. Comput. Vis. (IJCV) 115(3), 211\u2013252 (2015). https:\/\/doi.org\/10.1007\/s11263-015-0816-y","journal-title":"Int. J. Comput. Vis. (IJCV)"},{"key":"2_CR28","unstructured":"Thomas, F., B\u00e9thune, L., Lampinen, A.K., Serre, T., Hermann, K.: Understanding visual feature reliance through the lens of complexity. In: The Thirty-eighth Annual Conference on Neural Information Processing Systems (2024)"},{"key":"2_CR29","doi-asserted-by":"publisher","unstructured":"Tuli, S., Dasgupta, I., Grant, E., Griffiths, T.L.: Are convolutional neural networks or transformers more like human vision? arXiv preprint arXiv:2105.07197 (2021). https:\/\/doi.org\/10.48550\/arXiv.2105.07197","DOI":"10.48550\/arXiv.2105.07197"},{"key":"2_CR30","doi-asserted-by":"publisher","unstructured":"Vielhaben, J., Bareeva, D., Berend, J., Samek, W., Strodthoff, N.: Beyond scalars: concept-based alignment analysis in vision transformers. arXiv preprint arXiv:2412.06639 (2024). https:\/\/doi.org\/10.48550\/arXiv.2412.06639","DOI":"10.48550\/arXiv.2412.06639"},{"key":"2_CR31","unstructured":"Vielhaben, J., Bluecher, S., Strodthoff, N.: Multi-dimensional concept discovery (MCD): a unifying framework with completeness guarantees. Trans. Mach. Learn. Res. (2023)"},{"key":"2_CR32","unstructured":"Wah, C., Branson, S., Welinder, P., Perona, P., Belongie, S.: The caltech-UCSD birds-200-2011 dataset (2011). https:\/\/authors.library.caltech.edu\/27452\/?utm_campaign=The%20Batch&utm_source=hs_email &utm_medium=email &_hsenc=p2ANqtz--Sx1nvaahZe38-PWjKtUaD7qc__1GepLnIdt39_cou747ve6R6_mI2mgUTn45sU0V089Fp"},{"key":"2_CR33","doi-asserted-by":"publisher","unstructured":"Wolf, T., et\u00a0al.: Transformers: state-of-the-art natural language processing. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pp. 38\u201345. Association for Computational Linguistics, Online (2020). https:\/\/doi.org\/10.18653\/v1\/2020.emnlp-demos.6. https:\/\/www.aclweb.org\/anthology\/2020.emnlp-demos.6","DOI":"10.18653\/v1\/2020.emnlp-demos.6"},{"key":"2_CR34","doi-asserted-by":"publisher","unstructured":"Yang, J., et\u00a0al.: MedMNIST v2 - a large-scale lightweight benchmark for 2D and 3D biomedical image classification. Sci. Data 10(1), 1\u201310 (2023). https:\/\/doi.org\/10.1038\/s41597-022-01721-8. https:\/\/www.nature.com\/articles\/s41597-022-01721-8","DOI":"10.1038\/s41597-022-01721-8"},{"key":"2_CR35","unstructured":"Yosinski, J., Clune, J., Bengio, Y., Lipson, H.: How transferable are features in deep neural networks? In: Proceedings of the 28th International Conference on Neural Information Processing Systems, NIPS 2014, vol. 2, pp. 3320\u20133328 (2014)"},{"key":"2_CR36","series-title":"Lecture Notes in Computer Science","doi-asserted-by":"publisher","first-page":"818","DOI":"10.1007\/978-3-319-10590-1_53","volume-title":"Computer Vision \u2013 ECCV 2014","author":"MD Zeiler","year":"2014","unstructured":"Zeiler, M.D., Fergus, R.: Visualizing and understanding convolutional networks. In: Fleet, D., Pajdla, T., Schiele, B., Tuytelaars, T. (eds.) ECCV 2014, Part I. LNCS, vol. 8689, pp. 818\u2013833. Springer, Cham (2014). https:\/\/doi.org\/10.1007\/978-3-319-10590-1_53"}],"container-title":["Communications in Computer and Information Science","Explainable Artificial Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-032-08317-3_2","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T04:03:59Z","timestamp":1760155439000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-032-08317-3_2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,12]]},"ISBN":["9783032083166","9783032083173"],"references-count":36,"URL":"https:\/\/doi.org\/10.1007\/978-3-032-08317-3_2","relation":{},"ISSN":["1865-0929","1865-0937"],"issn-type":[{"value":"1865-0929","type":"print"},{"value":"1865-0937","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,10,12]]},"assertion":[{"value":"12 October 2025","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"The authors have no competing interests to declare that\u00a0are relevant to the content of this article.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Disclosure of Interests"}},{"value":"xAI","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"World Conference on Explainable Artificial Intelligence","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Istanbul","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"T\u00fcrkiye","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2025","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"9 July 2025","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"11 July 2025","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"3","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"xai2025","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/xaiworldconference.com\/2025\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}