{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,8]],"date-time":"2026-08-08T16:48:39Z","timestamp":1786207719325,"version":"3.56.0"},"reference-count":18,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T00:00:00Z","timestamp":1778803200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computers"],"abstract":"<jats:p>Large language models (LLMs) are increasingly used in professional examinations, but their relative performance on dental board-style questions remains unclear. This study compared two reasoning-optimized models, GPT-o3 and GPT-5T, with a general-purpose multimodal model, GPT-4o, using 399 Japanese dental board-style multiple-choice questions from 2018 to 2022. All questions were presented in Japanese, and items originally accompanied by charts, photographs, or other figures were analyzed separately from items without visual materials. Accuracy and item-level agreement were assessed using pairwise McNemar tests, stratified analyses according to the original presence of visual materials, the Breslow\u2013Day test for homogeneity of odds ratios, and two-proportion z-tests. GPT-5T achieved the highest overall accuracy (294\/399, 73.7%), followed by GPT-o3 (257\/399, 64.4%) and GPT-4o (255\/399, 63.9%). Pairwise McNemar tests showed that GPT-5T outperformed both GPT-4o (Holm-adjusted p = 0.00098) and GPT-o3 (Holm-adjusted p = 0.00072), whereas GPT-o3 and GPT-4o did not differ significantly (Holm-adjusted p = 0.920). Accuracy was lower for questions originally containing visual materials than for questions without such materials across all three models (GPT-4o: 49.7% vs. 72.2%; GPT-o3: 55.1% vs. 69.8%; GPT-5T: 59.9% vs. 81.8%). The advantage of GPT-5T was more evident in questions without visual materials, and heterogeneity across question formats was observed for GPT-5T versus GPT-o3. GPT-5T showed the strongest performance in this dataset. Questions originally containing visual materials were associated with lower accuracy across all models. Because the comparison was based on distinct item groups rather than experimentally manipulated visual conditions, this result should be interpreted as a difference across question formats and may also reflect differences in item composition and difficulty between the two groups.<\/jats:p>","DOI":"10.3390\/computers15050317","type":"journal-article","created":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T14:19:29Z","timestamp":1778854769000},"page":"317","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Comparative Performance of Three GPT Models on Japanese Dental Board-Style Multiple-Choice Questions"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4811-0051","authenticated-orcid":false,"given":"Hikaru","family":"Fukuda","sequence":"first","affiliation":[{"name":"Division of Maxillofacial Surgery, Department of Science of Physical Functions, Kyushu Dental University, Fukuoka 803-8580, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6335-5947","authenticated-orcid":false,"given":"Masaki","family":"Morishita","sequence":"additional","affiliation":[{"name":"Department of Comprehensive Dental Practice Education Development, Kyushu Dental University, Fukuoka 803-8580, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1756-6305","authenticated-orcid":false,"given":"Kosuke","family":"Muraoka","sequence":"additional","affiliation":[{"name":"School of Oral Health Sciences, Kyushu Dental University, Fukuoka 803-8580, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-7850-705X","authenticated-orcid":false,"given":"Shino","family":"Maeda","sequence":"additional","affiliation":[{"name":"School of Oral Health Sciences, Kyushu Dental University, Fukuoka 803-8580, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-3318-9324","authenticated-orcid":false,"given":"Taiji","family":"Nakamura","sequence":"additional","affiliation":[{"name":"Department of Comprehensive Dental Practice Education Development, Kyushu Dental University, Fukuoka 803-8580, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6586-2539","authenticated-orcid":false,"given":"Manabu","family":"Habu","sequence":"additional","affiliation":[{"name":"Division of Maxillofacial Surgery, Department of Science of Physical Functions, Kyushu Dental University, Fukuoka 803-8580, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6297-6138","authenticated-orcid":false,"given":"Shuji","family":"Awano","sequence":"additional","affiliation":[{"name":"Department of Comprehensive Dental Practice Education Development, Kyushu Dental University, Fukuoka 803-8580, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kentaro","family":"Ono","sequence":"additional","affiliation":[{"name":"Division of Physiology, Department of Health Promotion, Kyushu Dental University, Fukuoka 803-8580, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,5,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wang, Y., Shi, Y., Yang, T., Wang, W., Sun, Z., and Zhang, Y. (2026). Structural Performance Warning Based on Computer Intelligent Monitoring and Fractional-Order Multi-Rate Kalman Fusion Method. Fractal Fract., 10.","DOI":"10.3390\/fractalfract10030186"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Tarabanis, C., Zahid, S., Mamalis, M., Zhang, K., Kalampokis, E., and Jankelson, L. (2024). Performance of publicly available large language models on internal medicine board-style questions. PLoS Digit. Health, 3.","DOI":"10.1371\/journal.pdig.0000604"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"e72062","DOI":"10.2196\/72062","article-title":"Large language models in medical diagnostics: Scoping review with bibliometric analysis","volume":"27","author":"Su","year":"2025","journal-title":"J. Med. Internet Res."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"e48291","DOI":"10.2196\/48291","article-title":"Large language models in medical education: Opportunities, challenges, and future directions","volume":"9","author":"AlSaad","year":"2023","journal-title":"JMIR Med. Educ."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"100847","DOI":"10.1016\/j.identj.2025.100847","article-title":"Evaluating the accuracy and performance of ChatGPT-4o in solving Japanese national dental technician examination","volume":"75","author":"Fukuda","year":"2025","journal-title":"Int. Dent. J."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"e64486","DOI":"10.2196\/64486","article-title":"Accuracy of large language models when answering clinical research questions: Systematic review and network meta-analysis","volume":"27","author":"Wang","year":"2025","journal-title":"J. Med. Internet Res."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"61","DOI":"10.1038\/s41586-024-07930-y","article-title":"Larger and more instructable language models become less reliable","volume":"634","author":"Zhou","year":"2024","journal-title":"Nature"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"e74142","DOI":"10.2196\/74142","article-title":"Evaluating the reasoning capabilities of large language models for medical coding and hospital readmission risk stratification: Zero-shot prompting approach","volume":"27","author":"Naliyatthaliyazchayil","year":"2025","journal-title":"J. Med. Internet Res."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1964","DOI":"10.1093\/jamia\/ocae131","article-title":"Reasoning with large language models for medical question answering","volume":"31","author":"Lucas","year":"2024","journal-title":"J. Am. Med. Inform. Assoc."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"e240153","DOI":"10.1148\/radiol.240153","article-title":"Performance of GPT-4 with vision on text- and image-based ACR diagnostic radiology in-training examination questions","volume":"312","author":"Hayden","year":"2024","journal-title":"Radiology"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Altermatt, F.R., Neyem, A., Sumonte, N.I., Villagr\u00e1n, I., Mendoza, M., Lacassie, H.J., and Delfino, A.E. (2025). Evaluating GPT-4o in high-stakes medical assessments: Performance and error analysis on a Chilean anesthesiology exam. BMC Med. Educ., 25.","DOI":"10.1186\/s12909-025-08084-9"},{"key":"ref_12","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., and Altenschmidt, J. (2023). GPT-4 Technical Report. arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1185","DOI":"10.1111\/j.1541-0420.2010.01408.x","article-title":"Multiple McNemar tests","volume":"66","author":"Westfall","year":"2010","journal-title":"Biometrics"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Chau, R.C.W., Thu, K.M., Yu, O.Y., Hsung, R.T.-C., Wang, D.C.P., Man, M.W.H., Wang, J.J., and Lam, W.Y.H. (2025). Evaluation of Chatbot Responses to Text-Based Multiple-Choice Questions in Prosthodontic and Restorative Dentistry. Dent. J., 13.","DOI":"10.3390\/dj13070279"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1111\/j.1365-2923.2009.03425.x","article-title":"A primer on classical test theory and item response theory for assessments in medical education","volume":"44","author":"Champlain","year":"2010","journal-title":"Med. Educ."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"739","DOI":"10.1046\/j.1365-2923.2003.01587.x","article-title":"Item response theory: Applications of modern test theory in medical education","volume":"37","author":"Downing","year":"2003","journal-title":"Med. Educ."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Kelly, C.J., Karthikesalingam, A., and Suleyman, M. (2019). Key challenges for delivering clinical impact with artificial intelligence. BMC Med., 17.","DOI":"10.1186\/s12916-019-1426-2"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"104938","DOI":"10.1016\/j.jdent.2024.104938","article-title":"Accuracy and consistency of chatbots versus clinicians for answering pediatric dentistry questions: A pilot study","volume":"144","author":"Rokhshad","year":"2024","journal-title":"J. Dent."}],"container-title":["Computers"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/5\/317\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,15]],"date-time":"2026-05-15T14:28:09Z","timestamp":1778855289000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-431X\/15\/5\/317"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,15]]},"references-count":18,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2026,5]]}},"alternative-id":["computers15050317"],"URL":"https:\/\/doi.org\/10.3390\/computers15050317","relation":{},"ISSN":["2073-431X"],"issn-type":[{"value":"2073-431X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,15]]}}}