{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T15:06:36Z","timestamp":1784646396421,"version":"3.55.0"},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"20","license":[{"start":{"date-parts":[[2023,12,20]],"date-time":"2023-12-20T00:00:00Z","timestamp":1703030400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,12,20]],"date-time":"2023-12-20T00:00:00Z","timestamp":1703030400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100014718","name":"Innovative Research Group Project of the National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U1806202, 61533011"],"award-info":[{"award-number":["U1806202, 61533011"]}],"id":[{"id":"10.13039\/100014718","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"publisher","award":["ZR2019BF035, ZR2020ZD25, ZR2021QF042, 2022CXGC10501"],"award-info":[{"award-number":["ZR2019BF035, ZR2020ZD25, ZR2021QF042, 2022CXGC10501"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Lung cancer constitutes the most severe cause of cancer-related mortality. Recent evidence supports that early detection by means of computed tomography (CT) scans significantly reduces mortality rates. Given the remarkable progress of Vision Transformers (ViTs) in the field of computer vision, we have delved into comparing the performance of ViTs versus Convolutional Neural Networks (CNNs) for the automatic identification of lung cancer based on a dataset of 212 medical images. Importantly, neither ViTs nor CNNs require lung nodule annotations to predict the occurrence of cancer. To address the dataset limitations, we have trained both ViTs and CNNs with three advanced techniques: transfer learning, self-supervised learning, and sharpness-aware minimizer. Remarkably, we have found that CNNs achieve highly accurate prediction of a patient\u2019s cancer status, with an outstanding recall (93.4%) and area under the Receiver Operating Characteristic curve (AUC) of 98.1%, when trained with self-supervised learning. Our study demonstrates that both CNNs and ViTs exhibit substantial potential with the three strategies. However, CNNs are more effective than ViTs with the insufficient quantities of dataset.<\/jats:p>","DOI":"10.1007\/s11042-023-17644-4","type":"journal-article","created":{"date-parts":[[2023,12,20]],"date-time":"2023-12-20T04:49:16Z","timestamp":1703047756000},"page":"59253-59269","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":37,"title":["Comparing CNN-based and transformer-based models for identifying lung cancer: which is more effective?"],"prefix":"10.1007","volume":"83","author":[{"given":"Lulu","family":"Gai","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mengmeng","family":"Xing","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xu","family":"Qiao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,12,20]]},"reference":[{"issue":"5","key":"17644_CR1","doi-asserted-by":"publisher","first-page":"613","DOI":"10.1016\/j.jtho.2016.03.012","volume":"11","author":"AS Tsao","year":"2016","unstructured":"Tsao AS, Scagliotti GV, Bunn PA Jr, Carbone DP, Warren GW, Bai C, De Koning HJ, Yousaf-Khan AU, McWilliams A, Tsao MS (2016) Scientific advances in lung cancer 2015. J Thor Oncol 11(5):613\u2013638","journal-title":"J Thor Oncol"},{"issue":"3","key":"17644_CR2","doi-asserted-by":"publisher","first-page":"385","DOI":"10.1007\/s11547-010-0507-2","volume":"115","author":"F Fraioli","year":"2010","unstructured":"Fraioli F, Serra G, Passariello R (2010) CAD (computed-aided detection) and CADX (computer aided diagnosis) systems in identifying and characterising lung nodules on chest CT: overview of research, developments and new prospects. La Radiol Med 115(3):385\u2013402","journal-title":"La Radiol Med"},{"issue":"20","key":"17644_CR3","doi-asserted-by":"publisher","first-page":"28651","DOI":"10.1007\/s11042-022-12644-2","volume":"81","author":"V Kukreja","year":"2022","unstructured":"Kukreja V, Sakshi (2022) Machine learning models for mathematical symbol recognition: a stem to stern literature analysis. Multimedia Tools Appl 81(20):28651\u201328687","journal-title":"Multimedia Tools Appl"},{"issue":"7","key":"17644_CR4","first-page":"182","volume":"3","author":"G Vijaya","year":"2014","unstructured":"Vijaya G, Suhasini A, Priya R (2014) Automatic detection of lung cancer in CT images. IJRET: Int J Res Eng Technol 3(7):182\u2013186","journal-title":"IJRET: Int J Res Eng Technol"},{"issue":"7","key":"17644_CR5","doi-asserted-by":"publisher","first-page":"7047","DOI":"10.1007\/s10462-022-10330-1","volume":"56","author":"Sakshi","year":"2023","unstructured":"Sakshi, Kukreja V (2023) A dive in white and grey shades of ml and non-ml literature: a multivocal analysis of mathematical expressions. Artif Intell Rev 56(7):7047\u20137135","journal-title":"Artif Intell Rev"},{"issue":"5","key":"17644_CR6","doi-asserted-by":"publisher","first-page":"1299","DOI":"10.1109\/TMI.2016.2535302","volume":"35","author":"N Tajbakhsh","year":"2016","unstructured":"Tajbakhsh N, Shin JY, Gurudu SR, Hurst RT, Kendall CB, Gotway MB, Liang J (2016) Convolutional neural networks for medical image analysis: full training or fine tuning? IEEE Trans Med Imag 35(5):1299\u20131312","journal-title":"IEEE Trans Med Imag"},{"key":"17644_CR7","doi-asserted-by":"publisher","first-page":"105539","DOI":"10.1016\/j.compbiomed.2022.105539","volume":"146","author":"NF Aurna","year":"2022","unstructured":"Aurna NF, Yousuf MA, Taher KA, Azad A, Moni MA (2022) A classification of MRI brain tumor based on two stage feature level ensemble of deep CNN models. Comput Biol Med 146:105539","journal-title":"Comput Biol Med"},{"key":"17644_CR8","doi-asserted-by":"publisher","first-page":"104536","DOI":"10.1016\/j.compbiomed.2021.104536","volume":"134","author":"B Rostami","year":"2021","unstructured":"Rostami B, Anisuzzaman D, Wang C, Gopalakrishnan S, Niezgoda J, Yu Z (2021) Multiclass wound image classification using an ensemble deep CNN-based classifier. Comput Biol Med 134:104536","journal-title":"Comput Biol Med"},{"key":"17644_CR9","doi-asserted-by":"publisher","first-page":"103345","DOI":"10.1016\/j.compbiomed.2019.103345","volume":"111","author":"S Deepak","year":"2019","unstructured":"Deepak S, Ameer P (2019) Brain tumor classification using deep CNN features via transfer learning. Comput Biol Med 111:103345","journal-title":"Comput Biol Med"},{"key":"17644_CR10","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S et al (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv:2010.11929"},{"key":"17644_CR11","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R, Li L-J, Li K, Fei-Fei L (2009) Imagenet: a large-scale hierarchical image database. Paper presented at the 2009 IEEE conference on computer vision and pattern recognition, pp 248\u2013255","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"17644_CR12","unstructured":"Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. Adv Neural Inform Process Syst 25"},{"key":"17644_CR13","doi-asserted-by":"crossref","unstructured":"Hershey S, Chaudhuri S, Ellis DP, Gemmeke JF, Jansen A, Moore RC, Plakal M, Platt D, Saurous RA, Seybold B (2017) CNN architectures for large-scale audio classification. Paper presented at the 2017 IEEE international conference on acoustics, speech and signal processing (icassp), pp 131\u2013135","DOI":"10.1109\/ICASSP.2017.7952132"},{"key":"17644_CR14","doi-asserted-by":"publisher","first-page":"58","DOI":"10.1016\/j.artmed.2018.04.008","volume":"88","author":"D Bardou","year":"2018","unstructured":"Bardou D, Zhang K, Ahmad SM (2018) Lung sounds classification using convolutional neural networks. Artif Intell Med 88:58\u201369","journal-title":"Artif Intell Med"},{"key":"17644_CR15","unstructured":"Kukreja V, Lodhi S et al (2023) Impact of varying strokes on recognition rate: a case study on handwritten mathematical expressions. Int J Comput Digit Sys"},{"key":"17644_CR16","doi-asserted-by":"publisher","first-page":"104292","DOI":"10.1016\/j.engappai.2021.104292","volume":"103","author":"V Kukreja","year":"2021","unstructured":"Kukreja V (2021) A retrospective study on handwritten mathematical symbols and expressions: classification and recognition. Eng Appl Artif Intell 103:104292","journal-title":"Eng Appl Artif Intell"},{"key":"17644_CR17","doi-asserted-by":"crossref","unstructured":"Ronneberger O, Fischer P, Brox T (2015) U-net: convolutional networks for biomedical image segmentation. Paper presented at the international conference on medical image computing and computer-assisted intervention, pp 234\u2013241","DOI":"10.1007\/978-3-319-24574-4_28"},{"issue":"1","key":"17644_CR18","doi-asserted-by":"publisher","first-page":"457","DOI":"10.1007\/s11831-022-09805-9","volume":"30","author":"Kukreja V Sakshi","year":"2023","unstructured":"Sakshi Kukreja V (2023) Image segmentation techniques: statistical, comprehensive, semi-automated analysis and an application perspective analysis of mathematical expressions. Archiv Computat Methods Eng 30(1):457\u2013495","journal-title":"Archiv Computat Methods Eng"},{"key":"17644_CR19","doi-asserted-by":"crossref","unstructured":"\u00c7i\u00e7ek \u00d6, Abdulkadir A, Lienkamp SS, Brox T, Ronneberger O (2016) 3d u-net: learning dense volumetric segmentation from sparse annotation. Paper presented at the international conference on medical image computing and computer-assisted intervention, pp 424\u2013432","DOI":"10.1007\/978-3-319-46723-8_49"},{"key":"17644_CR20","doi-asserted-by":"crossref","unstructured":"Girshick R (2015) Fast R-CNN. In: Proceedings of the IEEE international conference on computer vision, pp 1440\u20131448","DOI":"10.1109\/ICCV.2015.169"},{"key":"17644_CR21","unstructured":"Ren S, He K, Girshick R, Sun J (2015) Faster R-CNN: towards real-time object detection with region proposal networks. Adv Neural Inform Process Syst 28"},{"key":"17644_CR22","doi-asserted-by":"crossref","unstructured":"Redmon J, Divvala S, Girshick R, Farhadi A (2016) You only look once: unified, real-time object detection. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 779\u2013788","DOI":"10.1109\/CVPR.2016.91"},{"key":"17644_CR23","doi-asserted-by":"publisher","first-page":"103912","DOI":"10.1016\/j.compbiomed.2020.103912","volume":"123","author":"R Rosati","year":"2020","unstructured":"Rosati R, Romeo L, Silvestri S, Marcheggiani F, Tiano L, Frontoni E (2020) Faster R-CNN approach for detection and quantification of DNA damage in COMET assay images. Comput Biol Med 123:103912","journal-title":"Comput Biol Med"},{"key":"17644_CR24","unstructured":"Simonyan K, Zisserman A (2014) Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556"},{"key":"17644_CR25","doi-asserted-by":"crossref","unstructured":"Szegedy C, Liu W, Jia Y, Sermanet P, Reed S, Anguelov D, Erhan D, Vanhoucke V, Rabinovich A (2015) Going deeper with convolutions. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 1\u20139","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"17644_CR26","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"17644_CR27","unstructured":"Tan M, Le Q (2019) Efficientnet: rethinking model scaling for convolutional neural networks. In: International conference on machine learning, pp 6105\u20136114"},{"key":"17644_CR28","unstructured":"Ba JL, Kiros JR, Hinton GE (2016) Layer normalization. arXiv:1607.06450"},{"key":"17644_CR29","unstructured":"Touvron H, Cord M, Douze M, Massa F, Sablayrolles A, J\u00e9gou H (2021) Training data-efficient image transformers & distillation through attention. In: International conference on machine learning, pp 10347\u201310357"},{"key":"17644_CR30","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez A.N, Kaiser \u0141, Polosukhin I (2017) Attention is all you need. Adv Neural Inform Process Syst 30"},{"key":"17644_CR31","unstructured":"Raghu M, Zhang C, Kleinberg J, Bengio S (2019) Transfusion: understanding transfer learning for medical imaging. arXiv:1902.07208"},{"key":"17644_CR32","first-page":"21271","volume":"33","author":"J-B Grill","year":"2020","unstructured":"Grill J-B, Strub F, Altch\u00e9 F, Tallec C, Richemond P, Buchatskaya E, Doersch C, Avila Pires B, Guo Z, Gheshlaghi Azar M (2020) Bootstrap your own latent-a new approach to self-supervised learning. Adv Neural Inform Process Syst 33:21271\u201321284","journal-title":"Adv Neural Inform Process Syst"},{"key":"17644_CR33","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Identity mappings in deep residual networks. In: European conference on computer vision, pp 630\u2013645","DOI":"10.1007\/978-3-319-46493-0_38"},{"issue":"5","key":"17644_CR34","doi-asserted-by":"publisher","first-page":"3248","DOI":"10.1109\/TII.2021.3107785","volume":"18","author":"Y Hua","year":"2021","unstructured":"Hua Y, Yi D (2021) Synthetic to realistic imbalanced domain adaption for urban scene perception. IEEE Trans Ind Inform 18(5):3248\u20133255","journal-title":"IEEE Trans Ind Inform"},{"issue":"3","key":"17644_CR35","doi-asserted-by":"publisher","first-page":"508","DOI":"10.1109\/TIV.2020.2980671","volume":"5","author":"U Michieli","year":"2020","unstructured":"Michieli U, Biasetton M, Agresti G, Zanuttigh P (2020) Adversarial learning and self-teaching techniques for domain adaptation in semantic segmentation. IEEE Trans Intell Veh 5(3):508\u2013518","journal-title":"IEEE Trans Intell Veh"},{"key":"17644_CR36","doi-asserted-by":"crossref","unstructured":"Caron M, Touvron H, Misra I, J\u00e9gou H, Mairal J, Bojanowski P, Joulin A (2021) Emerging properties in self-supervised vision transformers. arXiv:2104.14294","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"17644_CR37","unstructured":"Hendrycks D, Dietterich T (2019) Benchmarking neural network robustness to common corruptions and perturbations. arXiv:1903.12261"},{"key":"17644_CR38","doi-asserted-by":"crossref","unstructured":"Hendrycks D, Basart S, Mu N, Kadavath S, Wang F, Dorundo E, Desai R, Zhu T, Parajuli S, Guo M (2021) The many faces of robustness: a critical analysis of out-of-distribution generalization. In: Proceedings of the IEEE\/CVF international conference on computer vision, pp 8340\u20138349","DOI":"10.1109\/ICCV48922.2021.00823"},{"key":"17644_CR39","unstructured":"Chen X, Hsieh C-J, Gong B (2021) When vision transformers outperform resnets without pretraining or strong data augmentations. arXiv:2106.01548"},{"key":"17644_CR40","doi-asserted-by":"publisher","unstructured":"P\u00e9rez-Garc\u00eda F, Sparks R, Ourselin S (2021) Torchio: a python library for efficient loading, preprocessing, augmentation and patch-based sampling of medical images in deep learning. Comput Methods Programs Biomed 106236. https:\/\/doi.org\/10.1016\/j.cmpb.2021.106236","DOI":"10.1016\/j.cmpb.2021.106236"},{"key":"17644_CR41","unstructured":"Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L et al (2019) Pytorch: an imperative style, high-performance deep learning library. Adv Neural Inform Process Syst 32"},{"key":"17644_CR42","unstructured":"Liu L, Jiang H, He P, Chen W, Liu X, Gao J, Han J (2019) On the variance of the adaptive learning rate and beyond. arXiv:1908.03265"},{"key":"17644_CR43","unstructured":"Loshchilov I, Hutter F (2016) SGDR: Stochastic gradient descent with warm restarts. arXiv:1608.03983"},{"key":"17644_CR44","unstructured":"Powers DM (2020) Evaluation: from precision, recall and f-measure to roc, informedness, markedness and correlation. arXiv:2010.16061"},{"issue":"8","key":"17644_CR45","doi-asserted-by":"publisher","first-page":"861","DOI":"10.1016\/j.patrec.2005.10.010","volume":"27","author":"T Fawcett","year":"2006","unstructured":"Fawcett T (2006) An introduction to roc analysis. Pattern Recogn Lett 27(8):861\u2013874","journal-title":"Pattern Recogn Lett"},{"key":"17644_CR46","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2015) Delving deep into rectifiers: surpassing human-level performance on imagenet classification. In: Proceedings of the IEEE international conference on computer vision, pp 1026\u20131034","DOI":"10.1109\/ICCV.2015.123"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-023-17644-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-023-17644-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-023-17644-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,6,3]],"date-time":"2024-06-03T09:16:48Z","timestamp":1717406208000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-023-17644-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,20]]},"references-count":46,"journal-issue":{"issue":"20","published-online":{"date-parts":[[2024,6]]}},"alternative-id":["17644"],"URL":"https:\/\/doi.org\/10.1007\/s11042-023-17644-4","relation":{},"ISSN":["1573-7721"],"issn-type":[{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,20]]},"assertion":[{"value":"21 April 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 September 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 October 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 December 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that we have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}