{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,22]],"date-time":"2026-04-22T19:15:40Z","timestamp":1776885340196,"version":"3.51.2"},"reference-count":0,"publisher":"IOS Press","isbn-type":[{"value":"9781643685694","type":"electronic"}],"license":[{"start":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T00:00:00Z","timestamp":1734652800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,12,20]]},"abstract":"<jats:p>Histopathological image analysis remains at the forefront of computational pathology presenting numerous challenges and demanding tasks, primarily due to the complex nature of tissue structures and the extensive scale of whole slide images (WSIs). Deep learning models have been widely used in histopathology image analysis, especially convolutional neural network (CNN)-based models for classification. However, CNNs have certain limitations due to their small receptive field. Recent works employed adaptations of the classical transformer architecture to visual data [1] [2]. Models such as Vision Transformer (ViT) and Swin Transformer leverage the powerful multi-head self-attention mechanism and have demonstrated comparable or superior performance to state-of-the-art CNN-based classification models. Despite their successes, these models require huge amounts of training data to effectively learn representations as they lack the inherent inductive biases of CNNs. This work compares Vision Transformers with baseline CNN models using a breast cancer histopathological dataset. Further, we employ a novel knowledge-distillation approach to enhance the learning efficiency of vits, When trained with a limited amount of data, Unlike previous works, we aimed to minimize convolution operations when generating patch embeddings to preserve spatial information before reaching the transformer attention layers, we achieved an accuracy of 87.7% for the ViT-base trained as a student of ResNet50, which represents a 1.2% improvement in accuracy over the standalone ViT-base. [3]<\/jats:p>","DOI":"10.3233\/faia241423","type":"book-chapter","created":{"date-parts":[[2024,12,23]],"date-time":"2024-12-23T09:48:38Z","timestamp":1734947318000},"source":"Crossref","is-referenced-by-count":10,"title":["Vision Transformers and CNN-Based Knowledge-Distillation for Histopathological Image Classification"],"prefix":"10.3233","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1824-2361","authenticated-orcid":false,"given":"Seddik","family":"Boudissa","sequence":"first","affiliation":[{"name":"Graduate School of Engineering, Mie University, 1577 Kurima-machiya, Tsu, Mie 514-8507, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-3448-0535","authenticated-orcid":false,"given":"Shyam Sundar","family":"Debsarkar","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Cincinnati, OH 45221, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-6601-9213","authenticated-orcid":false,"given":"Hiroharu","family":"Kawanaka","sequence":"additional","affiliation":[{"name":"Graduate School of Engineering, Mie University, 1577 Kurima-machiya, Tsu, Mie 514-8507, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5109-6514","authenticated-orcid":false,"given":"Bruce","family":"Aronow","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Cincinnati, OH 45221, USA"},{"name":"Division of Biomedical Informatics, Cincinnati Children\u2019s Hospital Medical Center, Cincinnati, OH 45229, USA"},{"name":"Department of Pediatrics, University of Cincinnati, OH 45267, USA"},{"name":"Department of Biomedical Informatics, College of Medicine, University of Cincinnati, OH 45267, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7163-7453","authenticated-orcid":false,"given":"V.B. Surya","family":"Prasath","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Cincinnati, OH 45221, USA"},{"name":"Division of Biomedical Informatics, Cincinnati Children\u2019s Hospital Medical Center, Cincinnati, OH 45229, USA"},{"name":"Department of Pediatrics, University of Cincinnati, OH 45267, USA"},{"name":"Department of Biomedical Informatics, College of Medicine, University of Cincinnati, OH 45267, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"7437","container-title":["Frontiers in Artificial Intelligence and Applications","Fuzzy Systems and Data Mining X"],"original-title":[],"link":[{"URL":"https:\/\/ebooks.iospress.nl\/pdf\/doi\/10.3233\/FAIA241423","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,23]],"date-time":"2024-12-23T09:48:38Z","timestamp":1734947318000},"score":1,"resource":{"primary":{"URL":"https:\/\/ebooks.iospress.nl\/doi\/10.3233\/FAIA241423"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,20]]},"ISBN":["9781643685694"],"references-count":0,"URL":"https:\/\/doi.org\/10.3233\/faia241423","relation":{},"ISSN":["0922-6389","1879-8314"],"issn-type":[{"value":"0922-6389","type":"print"},{"value":"1879-8314","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,20]]}}}