{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T09:40:26Z","timestamp":1780738826938,"version":"3.54.1"},"reference-count":43,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2025,7,29]],"date-time":"2025-07-29T00:00:00Z","timestamp":1753747200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Big Data"],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>In the medical AI field, there is a significant gap between advances in AI technology and the challenge of applying locally trained models to diverse patient populations. This is mainly due to the limited availability of labeled medical image data, driven by privacy concerns. To address this, we have developed a self-supervised machine learning framework for detecting eye diseases from optical coherence tomography (OCT) images, aiming to achieve generalized learning while minimizing the need for large labeled datasets.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>Our framework, OCT-SelfNet, effectively addresses the challenge of data scarcity by integrating diverse datasets from multiple sources, ensuring a comprehensive representation of eye diseases. By employing a robust two-phase training strategy self-supervised pre-training with unlabeled data followed by a supervised training stage, we utilized the power of a masked autoencoder built on the SwinV2 backbone.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>Extensive experiments were conducted across three datasets with varying encoder backbones, assessing scenarios including the absence of self-supervised pre-training, the absence of data fusion, low data availability, and unseen data to evaluate the efficacy of our methodology. OCT-SelfNet outperformed the baseline model (ResNet-50, ViT) in most cases. Additionally, when tested for cross-dataset generalization, OCT-SelfNet surpassed the performance of the baseline model, further demonstrating its strong generalization ability. An ablation study revealed significant improvements attributable to self-supervised pre-training and data fusion methodologies.<\/jats:p><\/jats:sec><jats:sec><jats:title>Discussion<\/jats:title><jats:p>Our findings suggest that the OCT-SelfNet framework is highly promising for real-world clinical deployment in detecting eye diseases from OCT images. This demonstrates the effectiveness of our two-phase training approach and the use of a masked autoencoder based on the SwinV2 backbone. Our work bridges the gap between basic research and clinical application, which significantly enhances the framework's domain adaptation and generalization capabilities in detecting eye diseases.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fdata.2025.1609124","type":"journal-article","created":{"date-parts":[[2025,7,29]],"date-time":"2025-07-29T05:21:18Z","timestamp":1753766478000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["OCT-SelfNet: a self-supervised framework with multi-source datasets for generalized retinal disease detection"],"prefix":"10.3389","volume":"8","author":[{"given":"Fatema-E","family":"Jannat","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sina","family":"Gholami","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Minhaj Nur","family":"Alam","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hamed","family":"Tabkhi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2025,7,29]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","first-page":"322","DOI":"10.1097\/IAE.0000000000002373","article-title":"Quantitative OCT angiography features for objective classification and staging of diabetic retinopathy","volume":"40","author":"Alam","year":"2020","journal-title":"Retina"},{"key":"B2","doi-asserted-by":"publisher","first-page":"3998193","DOI":"10.1155\/2022\/3998193","article-title":"Olive disease classification based on vision transformer and CNN models","volume":"2022","author":"Alshammari","year":"2022","journal-title":"Comput. Intell. Neurosci"},{"key":"B3","doi-asserted-by":"crossref","first-page":"489","DOI":"10.1109\/ICSIPA.2017.8120661","article-title":"\u201cClassification OF SD-OCT images using a deep learning approach,\u201d","volume-title":"2017 IEEE International Conference on Signal and Image Processing Applications (ICSIPA)","author":"Awais","year":"2017"},{"key":"B4","doi-asserted-by":"publisher","first-page":"178","DOI":"10.3390\/diagnostics13020178","article-title":"Vision-transformer-based transfer learning for mammogram classification","volume":"13","author":"Ayana","year":"2023","journal-title":"Diagnostics"},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2106.08254","article-title":"Beit: Bert pre-training of image transformers","author":"Bao","year":"2021","journal-title":"arXiv"},{"key":"B6","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv:1810.04805","article-title":"Bert: pre-training of deep bidirectional transformers for language understanding","author":"Devlin","year":"2018","journal-title":"arXiv"},{"key":"B7","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.11929","article-title":"An image is worth 16x16 words: transformers for image recognition at scale","author":"Dosovitskiy","year":"2020","journal-title":"arXiv"},{"key":"B8","doi-asserted-by":"publisher","first-page":"2851","DOI":"10.1007\/s11517-022-02627-8","article-title":"Self-supervised patient-specific features learning for OCT image classification","volume":"60","author":"Fang","year":"2022","journal-title":"Med. Biol. Eng. Comput"},{"key":"B9","doi-asserted-by":"publisher","first-page":"369","DOI":"10.3928\/15428877-20110812-01","article-title":"Analysis of the relationship between drusen size and drusen area in eyes with age-related macular degeneration","volume":"42","author":"Friberg","year":"2011","journal-title":"Ophthalmic Surg. Lasers Imaging Retina"},{"key":"B10","doi-asserted-by":"publisher","first-page":"1259017","DOI":"10.3389\/fmed.2023.1259017","article-title":"Federated learning for diagnosis of age-related macular degeneration","volume":"10","author":"Gholami","year":"2023","journal-title":"Front. Med"},{"key":"B11","first-page":"16000","article-title":"\u201cMasked autoencoders are scalable vision learners,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer vision and Pattern Recognition","author":"He","year":"2022"},{"key":"B12","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1512.03385","article-title":"Deep residual learning for image recognition","author":"He","year":"2015","journal-title":"arXiv"},{"key":"B13","doi-asserted-by":"publisher","first-page":"1874","DOI":"10.1364\/BOE.487518","article-title":"Longitudinal deep network for consistent oct layer segmentation","volume":"14","author":"He","year":"2023","journal-title":"Biomed. Optics Express"},{"key":"B14","doi-asserted-by":"publisher","first-page":"4037","DOI":"10.1109\/TPAMI.2020.2992393","article-title":"Self-supervised visual feature learning with deep neural networks: a survey","volume":"43","author":"Jing","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell"},{"key":"B15","doi-asserted-by":"publisher","first-page":"1122","DOI":"10.1016\/j.cell.2018.02.010","article-title":"Identifying medical diagnoses and treatable diseases by image-based deep learning","volume":"172","author":"Kermany","year":"2018","journal-title":"Cell"},{"key":"B16","doi-asserted-by":"publisher","first-page":"100197","DOI":"10.1016\/j.xops.2022.100197","article-title":"Detection of nonexudative macular neovascularization on structural OCT images using vision transformers","volume":"2","author":"Kihara","year":"2022","journal-title":"Ophthalmol. Sci"},{"key":"B17","doi-asserted-by":"publisher","first-page":"14628","DOI":"10.1038\/s41598-023-41362-4","article-title":"Oct-based deep-learning models for the identification of retinal key signs","volume":"13","author":"Leandro","year":"2023","journal-title":"Sci. Rep"},{"key":"B18","doi-asserted-by":"publisher","first-page":"322","DOI":"10.1016\/j.oret.2016.12.009","article-title":"Deep learning is effective for the classification of OCT images of normal versus age-related macular degeneration","volume":"1","author":"Lee","year":"2017","journal-title":"Ophthalmol. Retina"},{"key":"B19","doi-asserted-by":"publisher","first-page":"19545","DOI":"10.1038\/s41598-023-46626-7","article-title":"Automated deep learning-based AMD detection and staging in real-world OCT datasets (pinnacle study report 5)","volume":"13","author":"Leingang","year":"2023","journal-title":"Sci. Rep"},{"key":"B20","article-title":"Medflip: medical vision-and-language self-supervised fast pre-training with masked autoencoder","author":"Li","year":"2024","journal-title":"arXiv preprint"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2012.07261","article-title":"Octa-500: a retinal dataset for optical coherence tomography angiography study","author":"Li","year":"2020","journal-title":"arXiv"},{"key":"B22","first-page":"12009","article-title":"\u201cSwin transformer v2: scaling up capacity and resolution,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu","year":"2022"},{"key":"B23","first-page":"10012","article-title":"\u201cSwin transformer: Hierarchical vision transformer using shifted windows,\u201d","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Liu","year":"2021"},{"key":"B24","article-title":"Decoupled weight decay regularization","author":"Loshchilov","year":"2017","journal-title":"arXiv preprint"},{"key":"B25","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1167\/tvst.7.6.41","article-title":"Deep learning-based automated classification of multi-categorical abnormalities from optical coherence tomography images","volume":"7","author":"Lu","year":"2018","journal-title":"Transl. Vis. Sci. Technol"},{"key":"B26","doi-asserted-by":"publisher","first-page":"3195","DOI":"10.1364\/BOE.450193","article-title":"Retinal layer segmentation in optical coherence tomography (OCT) using a 3D deep-convolutional regression network for patients with age-related macular degeneration","volume":"13","author":"Mukherjee","year":"","journal-title":"Biomed. Optics Express"},{"key":"B27","doi-asserted-by":"publisher","DOI":"10.1117\/12.2612991","article-title":"\u201cRetinal layer segmentation for age-related macular degeneration patients with 3D-UNet,\u201d","author":"Mukherjee","year":"","journal-title":"Medical Imaging 2022: Computer-Aided Diagnosis, Volume 12033"},{"key":"B28","doi-asserted-by":"publisher","first-page":"107141","DOI":"10.1016\/j.cmpb.2022.107141","article-title":"Ievit: an enhanced vision transformer architecture for chest X-ray image classification","volume":"226","author":"Okolo","year":"2022","journal-title":"Comput. Methods Programs Biomed"},{"key":"B29","doi-asserted-by":"publisher","first-page":"103327","DOI":"10.1016\/j.compbiomed.2019.103327","article-title":"Self-supervised iterative refinement learning for macular OCT volumetric data classification","volume":"111","author":"Qiu","year":"2019","journal-title":"Comput. Biol. Med"},{"key":"B30","doi-asserted-by":"publisher","first-page":"3199","DOI":"10.1167\/iovs.18-24106","article-title":"Prediction of individual disease conversion in early amd using artificial intelligence","volume":"59","author":"Schmidt-Erfurth","year":"2018","journal-title":"Invest. Ophthalmol. Vis. Sci"},{"key":"B31","doi-asserted-by":"publisher","first-page":"190","DOI":"10.1097\/ICU.0b013e32835fefee","article-title":"Long-term follow-up of vascular endothelial growth factor inhibitor therapy for neovascular age-related macular degeneration","volume":"24","author":"Scott","year":"2013","journal-title":"Curr. Opin. Ophthalmol"},{"key":"B32","doi-asserted-by":"publisher","first-page":"105368","DOI":"10.1016\/j.compbiomed.2022.105368","article-title":"Multi-scale convolutional neural network for automated AMD classification using retinal OCT images. Comput","volume":"144","author":"Sotoudeh-Paima","year":"2022","journal-title":"Biol. Med"},{"key":"B33","doi-asserted-by":"publisher","first-page":"3568","DOI":"10.1364\/BOE.5.003568","article-title":"Fully automated detection of diabetic macular edema and dry age-related macular degeneration from optical coherence tomography images","volume":"5","author":"Srinivasan","year":"2014","journal-title":"Biomed. Optics Express"},{"key":"B34","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s12886-020-01382-4","article-title":"Classification of optical coherence tomography images using a capsule network","volume":"20","author":"Tsuji","year":"2020","journal-title":"BMC Ophthalmol"},{"key":"B35","unstructured":"Attention is all you need\n          \n          \n            \n              Vaswani\n              A.\n            \n            \n              Shazeer\n              N.\n            \n            \n              Parmar\n              N.\n            \n            \n              Uszkoreit\n              J.\n            \n            \n              Jones\n              L.\n            \n            \n              Gomez\n              A. N.\n            \n            \n              Kaiser\n              L.\n            \n            \n              Polosukhin\n              I.\n            \n          \n          arXiv preprint\n          \n          2017"},{"key":"B36","doi-asserted-by":"publisher","first-page":"102469","DOI":"10.1016\/j.compmedimag.2024.102469","article-title":"Cervical OCT image classification using contrastive masked autoencoders with swin transformer","volume":"118","author":"Wang","year":"2024","journal-title":"Computerized Med. Imaging Graph"},{"key":"B37","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1038\/sigtrans.2016.23","article-title":"Genetic and environmental factors strongly influence risk, severity and progression of age-related macular degeneration","volume":"1","author":"Wang","year":"2016","journal-title":"Signal Transduct. Target. Ther"},{"key":"B38","doi-asserted-by":"publisher","first-page":"929755","DOI":"10.3389\/fphar.2022.929755","article-title":"Semi-supervised vision transformer with adaptive token sampling for breast cancer classification","volume":"13","author":"Wang","year":"2022","journal-title":"Front. Pharmacol"},{"key":"B39","doi-asserted-by":"publisher","first-page":"20260","DOI":"10.1038\/s41598-023-46433-0","article-title":"Self-supervised pre-training with contrastive and masked autoencoder methods for dealing with small datasets in deep learning for medical imaging","volume":"13","author":"Wolf","year":"2023","journal-title":"Sci. Rep."},{"key":"B40","volume-title":"Blindness and Visual Impairment","year":"2023"},{"key":"B41","doi-asserted-by":"publisher","first-page":"107660","DOI":"10.1016\/j.cmpb.2023.107660","article-title":"Resnet and its application to medical image processing: research progress and challenges","volume":"240","author":"Xu","year":"2023","journal-title":"Comput. Methods Programs Biomed."},{"key":"B42","doi-asserted-by":"publisher","first-page":"176","DOI":"10.1136\/bjo.2008.137356","article-title":"Spectral domain optical coherence tomography for quantitative evaluation of drusen and associated structural changes in non-neovascular age-related macular degeneration","volume":"93","author":"Yi","year":"2009","journal-title":"Br. J. Ophthalmol."},{"key":"B43","doi-asserted-by":"publisher","DOI":"10.1109\/ISBI53787.2023.10230477","article-title":"\u201cSelf pre-training with masked autoencoders for medical image classification and segmentation,\u201d","author":"Zhou","year":"2023","journal-title":"2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI)"}],"container-title":["Frontiers in Big Data"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2025.1609124\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,29]],"date-time":"2025-07-29T05:21:20Z","timestamp":1753766480000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fdata.2025.1609124\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,29]]},"references-count":43,"alternative-id":["10.3389\/fdata.2025.1609124"],"URL":"https:\/\/doi.org\/10.3389\/fdata.2025.1609124","relation":{},"ISSN":["2624-909X"],"issn-type":[{"value":"2624-909X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,29]]},"article-number":"1609124"}}