{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T02:57:30Z","timestamp":1781233050019,"version":"3.54.1"},"reference-count":46,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,6,11]],"date-time":"2026-06-11T00:00:00Z","timestamp":1781136000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MAKE"],"abstract":"<jats:p>Clinical reasoning over 3D CT scans is inherently compositional, requiring the integration of anatomical measurement, pathology assessment, spatial comparison, and clinical interpretation. We introduce MedToolica, a finetuning-free, role-based agentic framework for quantitative 3D abdominal CT reasoning that decomposes complex queries into structured sub-tasks coordinated through specialized expert tools. Empirical evaluation across quantitative reasoning benchmarks demonstrates that MedToolica is particularly effective in organ-centric measurement tasks when supported by reliable expert tools, achieving strong quantitative agreement (e.g., CCC=0.99 for organ HU estimation versus 0.46 for finetuned baselines) and notable gains on multi-step visual reasoning tasks. In contrast, lesion-oriented tasks remain constrained by upstream tool limitations, indicating that reasoning sophistication alone cannot compensate for unreliable perception. Furthermore, we observe that the capability of the core language model substantially influences orchestration quality: smaller LLM orchestrators exhibit reduced overall accuracy due to higher execution failure rates (25% vs. 79%) and increased susceptibility to hallucination (43% vs. 2%). Collectively, these findings identify expert tool reliability and orchestration capability as critical determinants of performance in compositional medical AI and highlight both the promise and current limitations of finetuning-free agentic reasoning for quantitative 3D CT analysis.<\/jats:p>","DOI":"10.3390\/make8060162","type":"journal-article","created":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T01:52:08Z","timestamp":1781229128000},"page":"162","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["MedToolica: Finetuning-Free Agentic Compositional Tool Learning for 3D CT Reasoning"],"prefix":"10.3390","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2967-3033","authenticated-orcid":false,"given":"Abdullah","family":"Hosseini","sequence":"first","affiliation":[{"name":"AI Innovation Lab, Weill Cornell Medicine-Qatar, Doha P.O. Box 24144, Qatar"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4145-5509","authenticated-orcid":false,"given":"Ahmed","family":"Serag","sequence":"additional","affiliation":[{"name":"AI Innovation Lab, Weill Cornell Medicine-Qatar, Doha P.O. Box 24144, Qatar"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,6,11]]},"reference":[{"key":"ref_1","unstructured":"Xu, M., Amiranashvili, T., Navarro, F., Fritsak, M., Hamamci, I.E., Shit, S., Wittmann, B., Er, S., Christ, S.M., and de la Rosa, E. (2025). CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"He, Y., Guo, P., Tang, Y., Myronenko, A., Nath, V., Xu, Z., Yang, D., Zhao, C., Simon, B., and Belue, M. (2025, January 11\u201315). VISTA3D: A unified segmentation foundation model for 3D medical imaging. Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.01943"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Chen, Q., Chen, X., Song, H., Xiong, Z., Yuille, A., Wei, C., and Zhou, Z. (2024, January 16\u201322). Towards generalizable tumor synthesis. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01060"},{"key":"ref_4","unstructured":"Wu, L., Zhuang, J., Ni, X., and Chen, H. (2024). Freetumor: Advance tumor segmentation via large-scale tumor synthesis. arXiv."},{"key":"ref_5","unstructured":"Di Piazza, T., Lazarus, C., Nempont, O., and Boussel, L. (2025). Structured Spectral Graph Representation Learning for Multi-label Abnormality Analysis from 3D CT Scans. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Hamamci, I.E., Er, S., and Menze, B. (2024). Ct2rep: Automated radiology report generation for 3d medical imaging. International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer.","DOI":"10.1007\/978-3-031-72390-2_45"},{"key":"ref_7","first-page":"ubaf011","article-title":"M3: Multimodal artificial intelligence for medical report generation and visual question answering from 3D abdominal CT scans","volume":"2","author":"Hosseini","year":"2025","journal-title":"BJR| Artif. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Di Piazza, T., Lazarus, C., Nempont, O., and Boussel, L. (2025, January 14\u201317). Ct-agrg: Automated abnormality-guided report generation from 3d chest ct volumes. Proceedings of the 2025 IEEE 22nd International Symposium on Biomedical Imaging (ISBI), Houston, TX, USA.","DOI":"10.1109\/ISBI60581.2025.10981073"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Hamamci, I.E., Er, S., Wang, C., Almas, F., Simsek, A.G., Esirgun, S.N., Dogan, I., Durugol, O.F., Hou, B., and Shit, S. (2026). Generalist foundation models from a multimodal dataset for 3D computed tomography. Nat. Biomed. Eng., 1\u201319.","DOI":"10.1038\/s41551-025-01599-y"},{"key":"ref_10","unstructured":"Chen, H., Zhao, W., Li, Y., Zhong, T., Wang, Y., Shang, Y., Guo, L., Han, J., Liu, T., and Liu, J. (2024). 3d-ct-gpt: Generating 3d radiology reports through integration of large vision-language models. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"e70921","DOI":"10.2196\/70921","article-title":"Hype vs reality in the integration of artificial intelligence in clinical workflows","volume":"9","author":"Solaiman","year":"2025","journal-title":"JMIR Form. Res."},{"key":"ref_12","unstructured":"Moor, M., Huang, Q., Wu, S., Yasunaga, M., Dalmia, Y., Leskovec, J., Zakka, C., Reis, E.P., and Rajpurkar, P. (2023). Med-flamingo: A multimodal medical few-shot learner. Machine Learning for Health (ML4H), PMLR."},{"key":"ref_13","first-page":"28541","article-title":"Llava-med: Training a large language-and-vision assistant for biomedicine in one day","volume":"36","author":"Li","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_14","unstructured":"Xu, W., Chan, H.P., Li, L., Aljunied, M., Yuan, R., Wang, J., Xiao, C., Chen, G., Liu, C., and Li, Z. (2025). Lingshu: A generalist foundation model for unified multimodal medical understanding and reasoning. arXiv."},{"key":"ref_15","unstructured":"Sellergren, A., Kazemzadeh, S., Jaroensri, T., Kiraly, A., Traverse, M., Kohlberger, T., Xu, S., Jamil, F., Hughes, C., and Lau, C. (2025). Medgemma technical report. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Pan, J., Liu, C., Wu, J., Liu, F., Zhu, J., Li, H.B., Chen, C., Ouyang, C., and Rueckert, D. (2025). Medvlm-r1: Incentivizing medical reasoning capability of vision-language models (vlms) via reinforcement learning. International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer.","DOI":"10.1007\/978-3-032-04981-0_32"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"2727","DOI":"10.1109\/TMI.2026.3661001","article-title":"Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models","volume":"45","author":"Lai","year":"2026","journal-title":"IEEE Trans. Med. Imaging"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"7866","DOI":"10.1038\/s41467-025-62385-7","article-title":"Towards generalist foundation model for radiology by leveraging web-scale 2d&3d medical data","volume":"16","author":"Wu","year":"2025","journal-title":"Nat. Commun."},{"key":"ref_19","unstructured":"Lai, H., Jiang, Z., Yao, Q., Wang, R., He, Z., Tao, X., Wei, W., Lv, W., and Zhou, S.K. (2024). E3D-GPT: Enhanced 3D visual foundation for medical vision-language model. arXiv."},{"key":"ref_20","unstructured":"Bai, F., Du, Y., Huang, T., Meng, M.Q.H., and Zhao, B. (2024). M3d: Advancing 3d medical image analysis with multi-modal large language models. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2524","DOI":"10.1109\/JBHI.2025.3604595","article-title":"Med3dvlm: An efficient vision-language model for 3d medical image analysis","volume":"30","author":"Xin","year":"2025","journal-title":"IEEE J. Biomed. Health Inform."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Hosseini, A., Ibrahim, A., and Serag, A. (2025). From Slices to Volumes: Multi-scale Fusion of 2D and 3D Features for CT Scan Report Generation. International Conference on Medical Image Computing and Computer-Assisted Intervention, Springer.","DOI":"10.1007\/978-3-032-04978-0_26"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"2626117","DOI":"10.1080\/08839514.2026.2626117","article-title":"SPINE: Segmentation-guided Processing and Integration of Multimodal Spinal MRI for Natural-Language Enhanced Report Generation","volume":"40","author":"Helmy","year":"2026","journal-title":"Appl. Artif. Intell."},{"key":"ref_24","unstructured":"Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y., and Wu, Y. (2024). Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv."},{"key":"ref_25","unstructured":"Lai, H., Jiang, Z., Zhang, K., Yao, Q., Wang, R., He, Z., Tao, X., Wei, W., and Zhou, S.K. (2026). Med3D-R1: Incentivizing Clinical Reasoning in 3D Medical Vision-Language Models for Abnormality Diagnosis. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Fathi, N., Kumar, A., and Arbel, T. (2025). Aura: A multi-modal medical agent for understanding, reasoning and annotation. International Workshop on Agentic AI for Medicine, Springer.","DOI":"10.1007\/978-3-032-06004-4_11"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Li, B., Yan, T., Pan, Y., Luo, J., Ji, R., Ding, J., Xu, Z., Liu, S., Dong, H., and Lin, Z. (2024). Mmedagent: Learning to use medical tools with multi-modal agent. Findings of the Association for Computational Linguistics: EMNLP 2024, Association for Computational Linguistics.","DOI":"10.18653\/v1\/2024.findings-emnlp.510"},{"key":"ref_28","unstructured":"Fallahpour, A., Ma, J., Munim, A., Lyu, H., and Wang, B. (2025). Medrax: Medical reasoning agent for chest x-ray. arXiv."},{"key":"ref_29","unstructured":"Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K.R., and Cao, Y. (2022, January 25\u201329). React: Synergizing reasoning and acting in language models. Proceedings of the The Eleventh International Conference on Learning Representations, Virtual."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Nath, V., Li, W., Yang, D., Myronenko, A., Zheng, M., Lu, Y., Liu, Z., Yin, H., Law, Y.M., and Tang, Y. (2025, January 11\u201315). Vila-m3: Enhancing vision-language models with medical expert knowledge. Proceedings of the Computer Vision and Pattern Recognition Conference, Nashville, TN, USA.","DOI":"10.1109\/CVPR52734.2025.01378"},{"key":"ref_31","unstructured":"Xia, P., Wang, J., Peng, Y., Zeng, K., Dong, Z., Wu, X., Tang, X., Zhu, H., Li, Y., and Zhang, L. (2025). Mmedagent-rl: Optimizing multi-agent collaboration for multimodal medical reasoning. arXiv."},{"key":"ref_32","unstructured":"Fan, Y., Hao, J., Chen, H., Bao, J., Shao, Y., Liang, Y., Hung, K.F., and Tang, H. (2026). OralGPT-Plus: Learning to Use Visual Tools via Reinforcement Learning for Panoramic X-ray Analysis. arXiv."},{"key":"ref_33","unstructured":"Hoopes, A., Dey, N., Butoi, V.I., Guttag, J.V., and Dalca, A.V. (2024). VoxelPrompt: A Vision Agent for End-to-End Medical Image Analysis. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, Z., Wu, J., Cai, L., Low, C.H., Yang, X., Li, Q., and Jin, Y. (2025). Medagent-pro: Towards evidence-based multi-modal medical diagnosis via reasoning agentic workflow. arXiv.","DOI":"10.20944\/preprints202503.1751.v2"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Raza, M., Salem, S., Kwon, H., Hussain, J., Gu, Y.H., and Al-Antari, M.A. (IEEE J. Biomed. Health Inform., 2025). Multimodal Knowledge-Infused VLM for Respiratory Disease Prediction and Clinical Report Generation, IEEE J. Biomed. Health Inform., early access.","DOI":"10.1109\/JBHI.2025.3631264"},{"key":"ref_36","unstructured":"Lin, Y., Ding, Y., Wu, Y., and Peng, Y. (2026). MARCH: Multi-Agent Radiology Clinical Hierarchy for CT Report Generation. arXiv."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Mao, Y., Xu, W., Qin, Y., and Gao, Y. (2025). CT-Agent: A multimodal-LLM agent for 3D CT radiology question answering. arXiv.","DOI":"10.1007\/s11432-025-4818-7"},{"key":"ref_38","unstructured":"Roschewitz, M., Styppa, K., Tao, Y., Sohn, J., Delbrouck, J.B., Gundersen, B., Deperrois, N., Bluethgen, C., Vogt, J., and Menze, B. (2026). RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Feng, J., Zheng, Q., Wu, C., Zhao, Z., Zhang, Y., Wang, Y., and Xie, W. (2025). M 3 builder: A multi-agent system for automated machine learning in medical imaging. International Workshop on Agentic AI for Medicine, Springer.","DOI":"10.1007\/978-3-032-06004-4_12"},{"key":"ref_40","unstructured":"Sellergren, A., Gao, C., Mahvar, F., Kohlberger, T., Jamil, F., Traverse, M., Tono, A., Sadjad, B., Yang, L., and Lau, C. (2026). Medgemma 1.5 technical report. arXiv."},{"key":"ref_41","unstructured":"Erdur, A.C., Scholz, D., Pan, J., Wiestler, B., Rueckert, D., and Peeken, J.C. (2026). Agentic Large Language Models for Training-Free Neuro-Radiological Image Analysis. arXiv."},{"key":"ref_42","unstructured":"Lu, P., Chen, B., Liu, S., Thapa, R., Boen, J., and Zou, J. (2025). Octotools: An agentic framework with extensible tools for complex reasoning. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"e230024","DOI":"10.1148\/ryai.230024","article-title":"TotalSegmentator: Robust segmentation of 104 anatomic structures in CT images","volume":"5","author":"Wasserthal","year":"2023","journal-title":"Radiol. Artif. Intell."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Tian, J., Liu, L., Shi, Z., and Xu, F. (2019). Automatic couinaud segmentation from CT volumes on liver using GLC-UNet. International Workshop on Machine Learning in Medical Imaging, Springer.","DOI":"10.1007\/978-3-030-32692-0_32"},{"key":"ref_45","unstructured":"Chen, Y., Xiao, W., Bassi, P.R., Zhou, X., Er, S., Hamamci, I.E., Zhou, Z., and Yuille, A. (2025). Are vision language models ready for clinical diagnosis? a 3d medical benchmark for tumor-centric visual question answering. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"255","DOI":"10.2307\/2532051","article-title":"A concordance correlation coefficient to evaluate reproducibility","volume":"45","author":"Lawrence","year":"1989","journal-title":"Biometrics"}],"container-title":["Machine Learning and Knowledge Extraction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/6\/162\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T02:33:26Z","timestamp":1781231606000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/6\/162"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,11]]},"references-count":46,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["make8060162"],"URL":"https:\/\/doi.org\/10.3390\/make8060162","relation":{},"ISSN":["2504-4990"],"issn-type":[{"value":"2504-4990","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,11]]}}}