{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T04:52:08Z","timestamp":1787028728089,"version":"3.56.0"},"reference-count":33,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T00:00:00Z","timestamp":1688947200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>The current advancement towards retinal disease detection mainly focused on distinct feature extraction using either a convolutional neural network (CNN) or a transformer-based end-to-end deep learning (DL) model. The individual end-to-end DL models are capable of only processing texture or shape-based information for performing detection tasks. However, extraction of only texture- or shape-based features does not provide the model robustness needed to classify different types of retinal diseases. Therefore, concerning these two features, this paper developed a fusion model called \u2018Conv-ViT\u2019 to detect retinal diseases from foveal cut optical coherence tomography (OCT) images. The transfer learning-based CNN models, such as Inception-V3 and ResNet-50, are utilized to process texture information by calculating the correlation of the nearby pixel. Additionally, the vision transformer model is fused to process shape-based features by determining the correlation between long-distance pixels. The hybridization of these three models results in shape-based texture feature learning during the classification of retinal diseases into its four classes, including choroidal neovascularization (CNV), diabetic macular edema (DME), DRUSEN, and NORMAL. The weighted average classification accuracy, precision, recall, and F1 score of the model are found to be approximately 94%. The results indicate that the fusion of both texture and shape features assisted the proposed Conv-ViT model to outperform the state-of-the-art retinal disease classification models.<\/jats:p>","DOI":"10.3390\/jimaging9070140","type":"journal-article","created":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T00:45:37Z","timestamp":1688949937000},"page":"140","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":97,"title":["Conv-ViT: A Convolution and Vision Transformer-Based Hybrid Feature Extraction Method for Retinal Disease Detection"],"prefix":"10.3390","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6619-4694","authenticated-orcid":false,"given":"Pramit","family":"Dutta","sequence":"first","affiliation":[{"name":"Department of Electronics and Telecommunication Engineering, Chittagong University of Engineering & Technology, Chattogram 4349, Bangladesh"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0031-9284","authenticated-orcid":false,"given":"Khaleda Akther","family":"Sathi","sequence":"additional","affiliation":[{"name":"Department of Electronics and Telecommunication Engineering, Chittagong University of Engineering & Technology, Chattogram 4349, Bangladesh"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8251-5168","authenticated-orcid":false,"given":"Md. Azad","family":"Hossain","sequence":"additional","affiliation":[{"name":"Department of Electronics and Telecommunication Engineering, Chittagong University of Engineering & Technology, Chattogram 4349, Bangladesh"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6347-7509","authenticated-orcid":false,"given":"M. Ali Akber","family":"Dewan","sequence":"additional","affiliation":[{"name":"School of Computing and Information Systems, Faculty of Science and Technology, Athabasca University, Athabasca, AB T9S 3A3, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,7,10]]},"reference":[{"key":"ref_1","unstructured":"Ram, A., and Reyes-Aldasoro, C.C. (2020). The Relationship between Fully Connected Layers and Number of Classes for the Analysis of Retinal Images. arXiv."},{"key":"ref_2","unstructured":"National Eye Institute (2023, March 17). Age-Related Macular Degeneration (AMD), Available online: https:\/\/www.nei.nih.gov\/learn-about-eye-health\/eye-conditions-and-diseases\/age-related-macular-degeneration#section-id-7323."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1107","DOI":"10.1038\/nm1010-1107","article-title":"Vascular endothelial growth factor and age-related macular degeneration: From basic science to therapy","volume":"16","author":"Ferrara","year":"2010","journal-title":"Nat. Med."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1334","DOI":"10.1001\/jamaophthalmol.2014.2854","article-title":"Prevalence of and Risk Factors for Diabetic Macular Edema in the United States","volume":"132","author":"Varma","year":"2014","journal-title":"JAMA Ophthalmol."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"564","DOI":"10.1001\/archopht.122.4.564","article-title":"Prevalence of Age-Related Macular Degeneration in the United States","volume":"122","author":"Friedman","year":"2004","journal-title":"Arch. Ophthalmol."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"3900714","DOI":"10.1109\/JPHOT.2019.2934484","article-title":"On OCT Image Classification via Deep Learning","volume":"11","author":"Wang","year":"2019","journal-title":"IEEE Photonics J."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1122","DOI":"10.1016\/j.cell.2018.02.010","article-title":"Identifying Medical Diagnoses and Treatable Diseases by Image-Based Deep Learning","volume":"172","author":"Kermany","year":"2018","journal-title":"Cell"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Khan, I.A., Sajeeb, A., and Fattah, S.A. (2020, January 20\u201321). An Automatic Ocular Disease Detection Scheme from Enhanced Fundus Images Based on Ensembling Deep CNN Networks. Proceedings of the 11th International Conference on Electrical and Computer Engineering (ICECE), Dhaka, Bangladesh.","DOI":"10.1109\/ICECE51571.2020.9393050"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2988","DOI":"10.1109\/JBHI.2020.3046771","article-title":"DeepUWF: An Automated Ultra-Wide-Field Fundus Screening System via Deep Learning","volume":"25","author":"Zhang","year":"2021","journal-title":"IEEE J. Biomed. Health Inform."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Wijesinghe, I., Gamage, C., and Chitraranjan, C. (2019, January 4\u20136). Transfer Learning with Ensemble Feature Extraction and Low-Rank Matrix Factorization for Severity Stage Classification of Diabetic Retinopathy. Proceedings of the 2019 IEEE 31st International Conference on Tools with Artificial Intelligence (ICTAI), Portland, OR, USA.","DOI":"10.1109\/ICTAI.2019.00132"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"67349","DOI":"10.1109\/ACCESS.2021.3076427","article-title":"Ensemble Learning Approach to Retinal Thickness Assessment in Optical Coherence Tomography","volume":"9","author":"Cruz","year":"2021","journal-title":"IEEE Access"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"258","DOI":"10.1016\/j.icte.2021.12.006","article-title":"Combining transformer and CNN for object detection in UAV imagery","volume":"9","author":"Hendira","year":"2023","journal-title":"ICT Express"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"13318","DOI":"10.1109\/JSEN.2022.3179535","article-title":"Movements Classification Through sEMG with Convolutional Vision Transformer and Stacking Ensemble Learning","volume":"22","author":"Shen","year":"2022","journal-title":"Sensors"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"948","DOI":"10.12688\/f1000research.73082.1","article-title":"Encoding Retina Image to Words using Ensemble of Vision Transformers for Diabetic Retinopathy Grading","volume":"10","author":"AlDahoul","year":"2021","journal-title":"F1000Research"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Gupta, A., Gautam, N., and Vishwakarma, D.K. (2022, January 29\u201331). Ensemble Learning using Vision Transformer and Convolutional Networks for Person Re-ID. Proceedings of the 2022 6th International Conference on Computing Methodologies and Communication (ICCMC), Erode, India.","DOI":"10.1109\/ICCMC53470.2022.9753761"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"106173","DOI":"10.1016\/j.engappai.2023.106173","article-title":"TransCNN: Hybrid CNN and transformer mechanism for surveillance anomaly detection. Engineering Applications of Artificial Intelligence","volume":"123 Pt A","author":"Ullah","year":"2023","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"103289","DOI":"10.1016\/j.ipm.2023.103289","article-title":"Vision transformer attention with multi-reservoir echo state network for anomaly recognition","volume":"60","author":"Ullah","year":"2023","journal-title":"Inf. Process. Manag."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"342","DOI":"10.1016\/j.neucom.2022.10.081","article-title":"Transformers and CNNs fusion network for salient object detection","volume":"520","author":"Yao","year":"2023","journal-title":"Neurocomputing"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yang, L., Yang, Y., Yang, J., Zhao, N., Wu, L., Wang, L., and Wang, T. (2022). FusionNet: A Convolution\u2013Transformer Fusion Network for Hyperspectral Image Classification. Remote Sens., 14.","DOI":"10.3390\/rs14164066"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"341","DOI":"10.3390\/signals3020022","article-title":"An Empirical Study on Ensemble of Segmentation Approaches","volume":"3","author":"Nanni","year":"2022","journal-title":"Signals"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Liu, H., and Hu, Q. (2021). TransFuse: Fusing Transformers and CNNs for Medical Image Segmentation. arXiv.","DOI":"10.1007\/978-3-030-87193-2_2"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"876065","DOI":"10.3389\/fnins.2022.876065","article-title":"O-Net (2022): A Novel Framework With Deep Fusion of CNN and Transformer for Simultaneous Segmentation and Classification","volume":"16","author":"Wang","year":"2022","journal-title":"Front. Neurosci."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the Inception Architecture for Computer Vision. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"6111","DOI":"10.1007\/s00521-019-04097-w","article-title":"A Transfer Convolutional Neural Network for Fault Diagnosis Based on ResNet-50","volume":"32","author":"Wen","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_25","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_27","unstructured":"Trockman, A., and Kolter, J.Z. (2022). Patches Are All You Need?. arXiv."},{"key":"ref_28","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Advances in Neural Information Processing Systems 30 (NIPS 2017), The MIT Press."},{"key":"ref_29","unstructured":"Kermany, D., Zhang, K., and Goldbaum, M. (2018). Large Dataset of Labeled Optical Coherence Tomography (OCT) and Chest XRay Images. Mendeley Data, 3."},{"key":"ref_30","unstructured":"Drummond, C., and Holte, R.C. (2003). Workshop on Learning from Imbalanced Datasets II, National Research Council."},{"key":"ref_31","first-page":"43","article-title":"OCTID: Optical Coherence Tomography Image Database","volume":"5","author":"Gholami","year":"2020","journal-title":"Data"},{"key":"ref_32","first-page":"2459","article-title":"Fully automated detection of retinal disorders by image-based deep learning","volume":"258","author":"Li","year":"2020","journal-title":"Graefes Arch. Clin. Exp. Ophthalmol."},{"key":"ref_33","first-page":"4736","article-title":"Iterative Fusion Convolutional Neural Networks for Classification of Optical Coherence Tomography Images","volume":"20","author":"Fang","year":"2020","journal-title":"Sensors"}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/9\/7\/140\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:09:37Z","timestamp":1760126977000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/9\/7\/140"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,10]]},"references-count":33,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["jimaging9070140"],"URL":"https:\/\/doi.org\/10.3390\/jimaging9070140","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,10]]}}}