{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,5]],"date-time":"2026-01-05T12:20:28Z","timestamp":1767615628671,"version":"3.48.0"},"reference-count":24,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,1,3]],"date-time":"2026-01-03T00:00:00Z","timestamp":1767398400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001786","name":"the University of Adelaide Paul Kwok Lee Bequest","doi-asserted-by":"publisher","award":["350-75134777"],"award-info":[{"award-number":["350-75134777"]}],"id":[{"id":"10.13039\/501100001786","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>To assess the efficiency of vision\u2013language models in detecting and classifying carious and non-carious lesions from intraoral photo imaging. A dataset of 172 annotated images were classified for microcavitation, cavitated lesions, staining, calculus, and non-carious lesions. Florence-2, PaLI-Gemma, and YOLOv8 models were trained on the dataset and model performance. The dataset was divided into 80:10:10 split, and the model performance was evaluated using mean average precision (mAP), mAP50-95, class-specific precision and recall. YOLOv8 outperformed the vision\u2013language models, achieving a mean average precision (mAP) of 37% with a precision of 42.3% (with 100% for cavitation detection) and 31.3% recall. PaLI-Gemma produced a recall of 13% and 21%. Florence-2 yielded a mean average precision of 10% with a precision and recall was 51% and 35%. YOLOv8 achieved the strongest overall performance. Florence-2 and PaLI-Gemma models underperformed relative to YOLOv8 despite the potential for multimodal contextual understanding, highlighting the need for larger, more diverse datasets and hybrid architectures to achieve improved performance.<\/jats:p>","DOI":"10.3390\/jimaging12010022","type":"journal-article","created":{"date-parts":[[2026,1,5]],"date-time":"2026-01-05T10:03:48Z","timestamp":1767607428000},"page":"22","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Comparative Evaluation of Vision\u2013Language Models for Detecting and Localizing Dental Lesions from Intraoral Images"],"prefix":"10.3390","volume":"12","author":[{"given":"Maria","family":"Jahan","sequence":"first","affiliation":[{"name":"Department of Electrical and Computer Science, North South University, Dhaka 1229, Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Al Ibne","family":"Siam","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Science, North South University, Dhaka 1229, Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lamim Zakir","family":"Pronay","sequence":"additional","affiliation":[{"name":"Department of CSE, National Institute of Technology Andhra Pradesh, Tadepalligudem 534101, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5782-9894","authenticated-orcid":false,"given":"Saif","family":"Ahmed","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Science, North South University, Dhaka 1229, Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7661-3570","authenticated-orcid":false,"given":"Nabeel","family":"Mohammed","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Science, North South University, Dhaka 1229, Bangladesh"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2004-0531","authenticated-orcid":false,"given":"James","family":"Dudley","sequence":"additional","affiliation":[{"name":"Adelaide Dental School, University of Adelaide, Adelaide, SA 5000, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5905-1572","authenticated-orcid":false,"given":"Taseef Hasan","family":"Farook","sequence":"additional","affiliation":[{"name":"Adelaide Dental School, University of Adelaide, Adelaide, SA 5000, Australia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2026,1,3]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"e101","DOI":"10.1002\/ail2.101","article-title":"An Application of 3D Vision Transformers and Explainable AI in Prosthetic Dentistry","volume":"5","author":"Sifat","year":"2024","journal-title":"Appl. AI Lett."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Kuznetsova, A., Maleva, T., and Soloviev, V. (2020, January 4\u20136). Detecting apples in orchards using YOLOv3 and YOLOv5 in general and close-up images. Proceedings of the Advances in Neural Networks\u2013ISNN 2020: 17th International Symposium on Neural Networks, ISNN 2020, Cairo, Egypt. Proceedings 17.","DOI":"10.1007\/978-3-030-64221-1_20"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"e32690","DOI":"10.2196\/32690","article-title":"Vision-language model for generating textual descriptions from clinical images: Model development and validation study","volume":"8","author":"Ji","year":"2024","journal-title":"JMIR Form. Res."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Sun, L., Ahuja, C., Chen, P., D\u2019Zmura, M., Batmanghelich, K., and Bontrager, P. (March, January 28). Multi-Modal Large Language Models are Effective Vision Learners. Proceedings of the 2025 IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Tucson, AZ, USA.","DOI":"10.1109\/WACV61041.2025.00835"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"176","DOI":"10.3390\/oral3020016","article-title":"One-Stage Methods of Computer Vision Object Detection to Classify Carious Lesions from Smartphone Imaging","volume":"3","author":"Salahin","year":"2023","journal-title":"Oral"},{"key":"ref_6","unstructured":"Verma, V. (2024, November 24). Introducing Moondream2: A Tiny Vision-Language Model. Available online: https:\/\/www.analyticsvidhya.com\/blog\/2024\/03\/introducing-moondream2-a-tiny-vision-language-model\/."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Thanh, M.T.G., Van Toan, N., Ngoc, V.T.N., Tra, N.T., Giap, C.N., and Nguyen, D.M. (2022). Deep Learning Application in Dental Caries Detection Using Intraoral Photos Taken by Smartphones. Appl. Sci., 12.","DOI":"10.3390\/app12115504"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Sohan, M., Sai Ram, T., and Rami Reddy, C.V. (2024, January 18\u201320). A review on yolov8 and its advancements. Proceedings of the International Conference on Data Intelligence and Cognitive Informatics, Tirunelveli, India.","DOI":"10.1007\/978-981-99-7962-2_39"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Bai, R., Shen, F., Wang, M., Lu, J., and Zhang, Z. (2023). Improving Detection Capabilities of YOLOv8-n for Small Objects in Remote Sensing Imagery: Towards Better Precision with Simplified Model Complexity. Res. Sq.","DOI":"10.21203\/rs.3.rs-3085871\/v1"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"109778","DOI":"10.1016\/j.compbiomed.2025.109778","article-title":"Optimized Yolov8 feature fusion algorithm for dental disease detection","volume":"187","author":"Wang","year":"2025","journal-title":"Comput. Biol. Med."},{"key":"ref_11","unstructured":"Beyer, L., Steiner, A., Pinto, A.S., Kolesnikov, A., Wang, X., Salz, D., Neumann, M., Alabdulmohsin, I., Tschannen, M., and Bugliarello, E. (2024). Paligemma: A versatile 3b vlm for transfer. arXiv."},{"key":"ref_12","unstructured":"Yuan, L., Chen, D., Chen, Y.-L., Codella, N., Dai, X., Gao, J., Hu, H., Huang, X., Li, B., and Li, C. (2021). Florence: A new foundation model for computer vision. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Xiao, B., Wu, H., Xu, W., Dai, X., Hu, H., Lu, Y., Zeng, M., Liu, C., and Yuan, L. (2024, January 17\u201321). Florence-2: Advancing a unified representation for a variety of vision tasks. Proceedings of the IEEE\/CVF Conference, on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00461"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1320","DOI":"10.1038\/s41591-020-1041-y","article-title":"Minimum information about clinical artificial intelligence modeling: The MI-CLAIM checklist","volume":"26","author":"Norgeot","year":"2020","journal-title":"Nat. Med."},{"key":"ref_15","first-page":"28","article-title":"PEP 8\u2013style guide for python code","volume":"1565","author":"Warsaw","year":"2001","journal-title":"Python. Org"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Yu, Y., Yang, C.-H.H., Kolehmainen, J., Shivakumar, P.G., Gu, Y., Ren, S.R.R., Luo, Q., Gourav, A., Chen, I.-F., and Liu, Y.-C. (2023, January 16\u201320). Low-rank adaptation of large language model rescoring for parameter-efficient speech recognition. Proceedings of the 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Taipei, Taiwan.","DOI":"10.1109\/ASRU57964.2023.10389632"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Zhai, X., Mustafa, B., Kolesnikov, A., and Beyer, L. (2023, January 2\u20136). Sigmoid loss for language image pre-training. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01100"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kudo, T., and Richardson, J. (2018). SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv.","DOI":"10.18653\/v1\/D18-2012"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"103911","DOI":"10.1016\/j.imavis.2020.103911","article-title":"IoU-aware single-stage object detector for accurate localization","volume":"97","author":"Wu","year":"2020","journal-title":"Image Vis. Comput."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wenkel, S., Alhazmi, K., Liiv, T., Alrshoud, S., and Simon, M. (2021). Confidence score: The forgotten dimension of object detection performance evaluation. Sensors, 21.","DOI":"10.3390\/s21134350"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"TaheriNejad, N., and Jantsch, A. (2019, January 5\u20138). Improved machine learning using confidence. Proceedings of the 2019 IEEE Canadian Conference of Electrical and Computer Engineering (CCECE), Edmonton, AB, Canada.","DOI":"10.1109\/CCECE.2019.8861962"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"508","DOI":"10.1111\/j.1600-0722.1997.tb00238.x","article-title":"Dental calculus: Recent insights into occurrence, formation, prevention, removal and oral health effects of supragingival and subgingival deposits","volume":"105","author":"White","year":"1997","journal-title":"Eur. J. Oral Sci."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"809","DOI":"10.1007\/s13534-025-00484-6","article-title":"Vision-language foundation models for medical imaging: A review of current practices and innovations","volume":"15","author":"Ryu","year":"2025","journal-title":"Biomed. Eng. Lett."},{"key":"ref_24","unstructured":"Lin, H., Xu, C., and Qin, J. (2025). Taming Vision-Language Models for Medical Image Analysis: A Comprehensive Review. arXiv."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/1\/22\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,5]],"date-time":"2026-01-05T10:36:32Z","timestamp":1767609392000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/12\/1\/22"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,3]]},"references-count":24,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["jimaging12010022"],"URL":"https:\/\/doi.org\/10.3390\/jimaging12010022","relation":{},"ISSN":["2313-433X"],"issn-type":[{"value":"2313-433X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,3]]}}}