{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,7]],"date-time":"2026-05-07T15:52:06Z","timestamp":1778169126595,"version":"3.51.4"},"reference-count":34,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,4,5]],"date-time":"2023-04-05T00:00:00Z","timestamp":1680652800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Research Foundation of Korea","award":["NRF-2020R1A2C1008753"],"award-info":[{"award-number":["NRF-2020R1A2C1008753"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Tomato leaf diseases can incur significant financial damage by having adverse impacts on crops and, consequently, they are a major concern for tomato growers all over the world. The diseases may come in a variety of forms, caused by environmental stress and various pathogens. An automated approach to detect leaf disease from images would assist farmers to take effective control measures quickly and affordably. Therefore, the proposed study aims to analyze the effects of transformer-based approaches that aggregate different scales of attention on variants of features for the classification of tomato leaf diseases from image data. Four state-of-the-art transformer-based models, namely, External Attention Transformer (EANet), Multi-Axis Vision Transformer (MaxViT), Compact Convolutional Transformers (CCT), and Pyramid Vision Transformer (PVT), are trained and tested on a multiclass tomato disease dataset. The result analysis showcases that MaxViT comfortably outperforms the other three transformer models with 97% overall accuracy, as opposed to the 89% accuracy achieved by EANet, 91% by CCT, and 93% by PVT. MaxViT also achieves a smoother learning curve compared to the other transformers. Afterwards, we further verified the legitimacy of the results on another relatively smaller dataset. Overall, the exhaustive empirical analysis presented in the paper proves that the MaxViT architecture is the most effective transformer model to classify tomato leaf disease, providing the availability of powerful hardware to incorporate the model.<\/jats:p>","DOI":"10.3390\/s23073751","type":"journal-article","created":{"date-parts":[[2023,4,5]],"date-time":"2023-04-05T05:42:44Z","timestamp":1680673364000},"page":"3751","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":41,"title":["Aggregating Different Scales of Attention on Feature Variants for Tomato Leaf Disease Diagnosis from Image Data: A Transformer Driven Study"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6132-1305","authenticated-orcid":false,"given":"Shahriar","family":"Hossain","sequence":"first","affiliation":[{"name":"Department of Computer Science and Engineering, BRAC University, Dhaka 1212, Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Md","family":"Tanzim Reza","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, BRAC University, Dhaka 1212, Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0306-4029","authenticated-orcid":false,"given":"Amitabha","family":"Chakrabarty","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, BRAC University, Dhaka 1212, Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6173-0857","authenticated-orcid":false,"given":"Yong Ju","family":"Jung","sequence":"additional","affiliation":[{"name":"School of Computing, Gachon University, Seongnam 13120, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,4,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Kaselimi, M., Voulodimos, A., Daskalopoulos, I., Doulamis, N., and Doulamis, A. (2022). A Vision Transformer Model for Convolution-Free Multilabel Classification of Satellite Imagery in Deforestation Monitoring. IEEE Trans. Neural Netw. Learn. Syst. Early Access, 1\u20139.","DOI":"10.1109\/TNNLS.2022.3144791"},{"key":"ref_2","first-page":"1","article-title":"Building Extraction With Vision Transformer","volume":"60","author":"Wang","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_3","first-page":"1","article-title":"Vision Transformer for Pansharpening","volume":"60","author":"Meng","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"4088","DOI":"10.1109\/TCAD.2022.3197489","article-title":"ViA: A Novel Vision-Transformer Accelerator Based on FPGA","volume":"41","author":"Wang","year":"2022","journal-title":"IEEE Trans.-Comput.-Aided Des. Integr. Circuits Syst."},{"key":"ref_5","first-page":"1","article-title":"A Survey on Vision Transformer","volume":"45","author":"Han","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An image is worth 16\u00d716 words: Transformers for image recognition at scale. Proceedings of the International Conference on Learning Representations, Vienna, Austria."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Batool, A., Hyder, S.B., Rahim, A., Waheed, N., Asghar, A., and Fawad, M. (2020, January 22\u201323). Classification and Identification of Tomato Leaf Disease Using Deep Neural Network. Proceedings of the 2020 International Conference on Engineering and Emerging Technologies (ICEET), Lahore, Pakistan.","DOI":"10.1109\/ICEET48479.2020.9048207"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1040","DOI":"10.1016\/j.procs.2018.07.070","article-title":"Tomato crop disease classification using pre-trained deep learning algorithm","volume":"133","author":"Rangarajan","year":"2018","journal-title":"Procedia Comput. Sci."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"299","DOI":"10.1080\/08839514.2017.1315516","article-title":"Deep Learning for Tomato Diseases: Classification and Symptoms Visualization","volume":"31","author":"Brahimi","year":"2017","journal-title":"Appl. Artif. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Fekri-Ershad, S. (2020). Bark Texture Classification Using Improved Local Ternary Patterns and Multilayer Neural Network, Expert Systems with Applications, Elsevier.","DOI":"10.1016\/j.eswa.2020.113509"},{"key":"ref_11","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, E.G. (2012, January 3). ImageNet classification with deep convolutional neural networks. Proceedings of the 25th International Conference on Neural Information Processing Systems-Volume 1 (NIPS\u201912), Red Hook, NY, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Qassim, H., Verma, A., and Feinzimer, D. (2018, January 8\u201310). Compressed residual-VGG16 CNN model for big data places image recognition. Proceedings of the 2018 IEEE 8th Annual Computing and Communication Workshop and Conference (CCWC), Las Vegas, NV, USA.","DOI":"10.1109\/CCWC.2018.8301729"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Carvalho, T., de Rezende, E.R.S., Alves, M.T.P., Balieiro, F.K.C., and Sovat, R.B. (2017, January 18\u201321). Exposing Computer Generated Images by Eyes Region Classification via Transfer Learning of VGG19 CNN. Proceedings of the 2017 16th IEEE International Conference on Machine Learning and Applications (ICMLA), Cancun, Mexico.","DOI":"10.1109\/ICMLA.2017.00-47"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1883","DOI":"10.4249\/scholarpedia.1883","article-title":"K-nearest neighbor","volume":"4","year":"2009","journal-title":"Scholarpedia"},{"key":"ref_15","first-page":"27","article-title":"Recent advances in image processing techniques for automated leaf pest and disease recognition\u2014A review","volume":"8","author":"Ngugi","year":"2021","journal-title":"Inf. Process. Agric."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Pandian, J., Arun, J., Kumar, V.D., Geman, O., Hnatiuc, M., Arif, M., and Kanchanadevi, K. (2022). Plant Disease Detection Using Deep Convolutional Neural Network. Appl. Sci., 12.","DOI":"10.3390\/app12146982"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"56607","DOI":"10.1109\/ACCESS.2020.2982456","article-title":"Deep Learning-Based Object Detection Improvement for Tomato Disease","volume":"8","author":"Zhang","year":"2020","journal-title":"IEEE Access"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"18568","DOI":"10.1038\/s41598-022-21498-5","article-title":"A robust deep learning approach for tomato plant leaf disease localization and classification","volume":"12","author":"Nawaz","year":"2022","journal-title":"Sci. Rep."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Albahli, S., and Nawaz, M. (2022). DCNet: DenseNet-77-based CornerNet model for the tomato plant leaf disease detection and classification. Front. Plant Sci., 13.","DOI":"10.3389\/fpls.2022.957961"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"898","DOI":"10.3389\/fpls.2020.00898","article-title":"Tomato Diseases and Pests Detection Based on Improved Yolo V3 Convolutional Neural Network","volume":"11","author":"Liu","year":"2020","journal-title":"Front. Plant Sci."},{"key":"ref_21","first-page":"33","article-title":"Pyramid methods in image processing","volume":"29","author":"Adelson","year":"1984","journal-title":"RCA Eng."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Ng, V., and Hofmann, D. (2018). Scalable Feature Extraction with Aerial and Satellite Imagery. Python Sci. Conf., 145\u2013151.","DOI":"10.25080\/Majora-4af1f417-015"},{"key":"ref_23","first-page":"100407","article-title":"Development of an Efficient CNN model for Tomato crop disease identification","volume":"28","author":"Agarwal","year":"2020","journal-title":"Sustain. Comput. Inform. Syst."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Bhujel, A., Kim, N.E., Arulmozhi, E., Basak, J.K., and Kim, H.T. (2022). A Lightweight Attention-Based Convolutional Neural Networks for Tomato Leaf Disease Classification. Agriculture, 12.","DOI":"10.3390\/agriculture12020228"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Trivedi, N.K., Gautam, V., Anand, A., Aljahdali, H.M., Villar, S.G., Anand, D., Goyal, D., and Kadry, N.S. (2021). Early Detection and Classification of Tomato Leaf Disease Using High-Performance Deep Neural Network. Sensors, 21.","DOI":"10.3390\/s21237987"},{"key":"ref_26","unstructured":"Hughes, D.P., and Salathe, M. (2020). An Open Access Repository of Images on Plant Health to Enable the Development of Mobile Disease Diagnostics. arXiv."},{"key":"ref_27","first-page":"3312","article-title":"Plants Disease Phenotyping using Quinary Patterns as Texture Descriptor","volume":"14","author":"Ahmad","year":"2020","journal-title":"Ksii Trans. Int. Inf. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"102","DOI":"10.1016\/j.compag.2018.12.042","article-title":"SLIC-SVM based leaf diseases saliency map extraction of tea plant","volume":"157","author":"Sun","year":"2019","journal-title":"Comput. Electron. Agric."},{"key":"ref_29","unstructured":"Arun, P.J., Gopal, G., Huang, M.-L., and Chang, Y.-H. (2022). Tomato Disease Multiple Sources [Data set]. Kaggle."},{"key":"ref_30","unstructured":"Huang, M.-L., and Chang, Y.-H. (2020). Dataset of Tomato Leaves. Mendeley Data."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Guo, M.H., Liu, Z.N., Mu, T.J., and Hu, S.M. (2021). Beyond self-attention: External attention using two linear layers for visual tasks. arXiv.","DOI":"10.1109\/TPAMI.2022.3211006"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Tu, Z., Talebi, H., Zhang, H., Yang, F., Milanfar, P., Bovik, A., and Li, Y. (2022). Maxvit: Multi-axis vision transformer. arXiv.","DOI":"10.1007\/978-3-031-20053-3_27"},{"key":"ref_33","unstructured":"Hassani, A., Walton, S., Shah, N., Abuduweili, A., Li, J., and Shi, H. (2021). Escaping the big data paradigm with compact transformers. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, W., Xie, E., Li, X., Fan, D.P., Song, K., Liang, D., Lu, T., Luo, P., and Shao, L. (2021, January 10\u201317). Pyramid vision transformer: A versatile backbone for dense prediction without convolutions. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00061"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/7\/3751\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:10:20Z","timestamp":1760123420000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/7\/3751"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,5]]},"references-count":34,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,4]]}},"alternative-id":["s23073751"],"URL":"https:\/\/doi.org\/10.3390\/s23073751","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,4,5]]}}}