{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T15:02:36Z","timestamp":1784559756133,"version":"3.55.0"},"reference-count":48,"publisher":"Oxford University Press (OUP)","issue":"5","license":[{"start":{"date-parts":[[2026,4,17]],"date-time":"2026-04-17T00:00:00Z","timestamp":1776384000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100002462","name":"Chungnam National University","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100002462","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100003725","name":"National Research Foundation of Korea","doi-asserted-by":"publisher","award":["RS-2026-25480008"],"award-info":[{"award-number":["RS-2026-25480008"]}],"id":[{"id":"10.13039\/501100003725","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,5,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Object comparison, a core cognitive ability for human understanding and decision-making, is essential in complex three-dimensional (3D) environments, such as enabling intelligent robotic systems to autonomously select appropriate tools and plan sequences of object manipulations, supporting computer-aided design (CAD) in assessing modifications against previous models, and ensuring quality control by verifying whether manufactured products match their original designs. However, there is a lack of research on directly comparing and explaining complex 3D objects in various formats, such as mesh, voxel, and point cloud. To address this gap, we propose Comp-PointLLM, a novel multimodal large language model framework for explainable comparison of 3D point cloud objects. Comp-PointLLM enhances 3D object understanding through a hybrid architecture that integrates 3D geometric features and two-dimensional visual features. Furthermore, we propose an automated data generation pipeline to construct comparison-captioning and question-answering datasets based on LLMs Experimental results demonstrate that Comp-PointLLM significantly outperforms baseline models across diverse object categories and comparison criteria, and exhibits strong zero-shot generalization to unseen categories. Ablation studies confirm that the hybrid architecture, two-stage training, and integrated data strategy all contribute to performance gains. Comp-PointLLM lays a solid foundation for comparing real-world 3D objects in product design, engineering, and intelligent robotic systems, paving the way for more advanced AI applications.<\/jats:p>","DOI":"10.1093\/jcde\/qwag039","type":"journal-article","created":{"date-parts":[[2026,4,16]],"date-time":"2026-04-16T11:36:37Z","timestamp":1776339397000},"page":"122-147","source":"Crossref","is-referenced-by-count":0,"title":["Comp-PointLLM: An explainable multimodal large language model for 3D object comparison"],"prefix":"10.1093","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-9381-0842","authenticated-orcid":false,"given":"Mingi","family":"Kim","sequence":"first","affiliation":[{"name":"Chungnam National University Department of Computer Science and Engineering, , 99 Daehak-ro, Yuseong-gu, Daejeon 34134 ,","place":["Republic of Korea"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9013-2338","authenticated-orcid":false,"given":"Hyungki","family":"Kim","sequence":"additional","affiliation":[{"name":"Chungnam National University Department of Computer Science and Engineering, , 99 Daehak-ro, Yuseong-gu, Daejeon 34134 ,","place":["Republic of Korea"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2026,4,17]]},"reference":[{"key":"2026072010153098600_bib1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2303.08774","article-title":"GPT-4 Technical Report","author":"Achiam","year":"2023"},{"key":"2026072010153098600_bib2","doi-asserted-by":"publisher","first-page":"23716","DOI":"10.48550\/arXiv.2204.14198","article-title":"Flamingo: A visual language model for few-shot learning","volume":"35","author":"Alayrac","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026072010153098600_bib3","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1512.03012","article-title":"ShapeNet: An information-rich 3D model repository","author":"Chang","year":"2015"},{"key":"2026072010153098600_bib4","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1504.00325","article-title":"Microsoft COCO captions: Data collection and evaluation server","author":"Chen","year":"2015"},{"key":"2026072010153098600_bib5","doi-asserted-by":"publisher","first-page":"49250","DOI":"10.48550\/arXiv.2305.06500","article-title":"InstructBLIP: Towards general-purpose vision-language models with instruction tuning","volume":"36","author":"Dai","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026072010153098600_bib6","doi-asserted-by":"publisher","first-page":"13142","DOI":"10.1109\/CVPR52729.2023.01263","article-title":"Objaverse: A universe of annotated 3D objects","author":"Deitke","year":"2023","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"2026072010153098600_bib7","doi-asserted-by":"publisher","first-page":"24199","DOI":"10.1109\/CVPR52733.2024.02284","article-title":"Describing differences in image sets with natural language","author":"Dunlap","year":"2024","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"2026072010153098600_bib8","doi-asserted-by":"crossref","unstructured":"Fan R., He F., Liu Y., Song Y., Fan L., Yan X. (2025). A parametric and feature-based CAD dataset to support human-computer interaction for advanced 3D shape learning. Integrated Computer-Aided Engineering, 32, 75\u201396. 10.3233\/ICA-240744.","DOI":"10.3233\/ICA-240744"},{"key":"2026072010153098600_bib9","doi-asserted-by":"publisher","first-page":"162","DOI":"10.1093\/jcde\/qwaf091","article-title":"Automatic reconstruction of 3D topology optimization to editable CAD model with rotation minimizing frames","volume":"12","author":"Feng","year":"2025","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026072010153098600_bib10","doi-asserted-by":"publisher","first-page":"312","DOI":"10.1093\/jcde\/qwaf001","article-title":"A large-scale multiview point cloud registration method based on distance statistical distribution of weak features","volume":"12","author":"Feng","year":"2025","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026072010153098600_bib11","doi-asserted-by":"publisher","first-page":"15180","DOI":"10.1109\/CVPR52729.2023.01457","article-title":"ImageBind: One embedding space to bind them all","author":"Girdhar","year":"2023","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"2026072010153098600_bib12","doi-asserted-by":"publisher","first-page":"4338","DOI":"10.1109\/TPAMI.2020.3005434","article-title":"Deep learning for 3D point clouds: A survey","volume":"43","author":"Guo","year":"2020","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026072010153098600_bib13","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2309.00615","article-title":"Point-bind & point-LLM: aligning point cloud with multi-modality for 3D understanding, generation, and instruction following","author":"Guo","year":"2023"},{"key":"2026072010153098600_bib14","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2309.03905","article-title":"ImageBind-LLM: Multi-modality Instruction Tuning","author":"Han","year":"2023"},{"key":"2026072010153098600_bib15","doi-asserted-by":"publisher","first-page":"20482","DOI":"10.48550\/arXiv.2307.12981","article-title":"3D-LLM: Injecting the 3D world into large language models","volume":"36","author":"Hong","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026072010153098600_bib16","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2106.09685","article-title":"LoRA: Low-rank adaptation of large language models","author":"Hu","year":"2022","journal-title":"International Conference on Learning Representations"},{"key":"2026072010153098600_bib17","doi-asserted-by":"publisher","first-page":"253","DOI":"10.1016\/j.jcde.2015.06.008","article-title":"An automatic 3D CAD model errors detection method of aircraft structural part for NC machining","volume":"2","author":"Huang","year":"2015","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026072010153098600_bib18","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2312.08168","article-title":"Chat-3d v2: Bridging 3d scene and large language models with object identifiers","author":"Huang","year":"2023"},{"key":"2026072010153098600_bib19","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1808.10584","article-title":"Learning to Describe Differences Between Pairs of Similar Images","author":"Jhamtani","year":"2018"},{"key":"2026072010153098600_bib20","doi-asserted-by":"publisher","first-page":"2075","DOI":"10.1109\/ICCV48922.2021.00210","article-title":"Viewpoint-Agnostic Change Captioning with Cycle Consistency","author":"Kim","year":"2021","journal-title":"IEEE\/CVF International Conference on Computer Vision"},{"key":"2026072010153098600_bib21","doi-asserted-by":"publisher","first-page":"2332","DOI":"10.1093\/jcde\/qwad102","article-title":"Improved semantic segmentation network using normal vector guidance for LiDAR point clouds","volume":"10","author":"Kim","year":"2023","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026072010153098600_bib22","doi-asserted-by":"publisher","first-page":"239","DOI":"10.1093\/jcde\/qwaf134","article-title":"Robust tessellation of CAD models without self-intersections","volume":"13","author":"Li","year":"2026","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026072010153098600_bib23","doi-asserted-by":"publisher","first-page":"19730","DOI":"10.48550\/arXiv.2301.12597","article-title":"BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models","author":"Li","year":"2023","journal-title":"International Conference on Machine Learning"},{"key":"2026072010153098600_bib24","doi-asserted-by":"publisher","first-page":"12888","DOI":"10.48550\/arXiv.2201.12086","article-title":"BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation","author":"Li","year":"2022","journal-title":"International Conference on Machine Learning"},{"key":"2026072010153098600_bib25","doi-asserted-by":"publisher","first-page":"740","DOI":"10.1007\/978-3-319-10602-1_48","article-title":"Microsoft COCO: Common objects in context","author":"Lin","year":"2014","journal-title":"European Conference on Computer Vision"},{"key":"2026072010153098600_bib26","doi-asserted-by":"publisher","first-page":"2973","DOI":"10.1109\/CVPRW67362.2025.00280","article-title":"Comparison Visual Instruction Tuning","author":"Lin","year":"2025","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops"},{"key":"2026072010153098600_bib27","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2402.03327","article-title":"Uni3D-LLM: Unifying point cloud perception, generation and editing with large language models","author":"Liu","year":"2024"},{"key":"2026072010153098600_bib28","doi-asserted-by":"publisher","first-page":"34892","DOI":"10.48550\/arXiv.2304.08485","article-title":"Visual Instruction Tuning","volume":"36","author":"Liu","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026072010153098600_bib29","doi-asserted-by":"publisher","first-page":"75307","DOI":"10.48550\/arXiv.2306.07279","article-title":"Scalable 3D captioning with pretrained models","volume":"36","author":"Luo","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026072010153098600_bib30","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1016\/j.jcde.2017.11.003","article-title":"Multi-criteria retrieval of CAD assembly models","volume":"5","author":"Lupinetti","year":"2018","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026072010153098600_bib31","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2405.10255","article-title":"When LLMs step into the 3D World: A survey and meta-analysis of 3D tasks via Multi-modal large language models","author":"Ma","year":"2024"},{"key":"2026072010153098600_bib32","doi-asserted-by":"publisher","first-page":"4624","DOI":"10.1109\/ICCV.2019.00472","article-title":"Robust change captioning","author":"Park","year":"2019","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"},{"key":"2026072010153098600_bib33","doi-asserted-by":"publisher","first-page":"214","DOI":"10.1007\/978-3-031-72775-7_13","article-title":"Shapellm: Universal 3d object understanding for embodied interaction","author":"Qi","year":"2024","journal-title":"European Conference on Computer Vision"},{"key":"2026072010153098600_bib34","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2103.00020","article-title":"Learning transferable visual models from natural language supervision","author":"Radford","year":"2021"},{"key":"2026072010153098600_bib35","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2409.12191","article-title":"Qwen2-vl: Enhancing vision-language model\u2019s perception of the world at any resolution","author":"Wang","year":"2024"},{"key":"2026072010153098600_bib36","doi-asserted-by":"publisher","first-page":"175","DOI":"10.1093\/jcde\/qwaf051","article-title":"Aerodynamic performance prediction of 3D aircraft based on point cloud deep learning","volume":"12","author":"Wang","year":"2025","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026072010153098600_bib37","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2308.08769","article-title":"Chat-3d: Data-efficiently tuning large language model for universal dialogue of 3d scenes","author":"Wang","year":"2023"},{"key":"2026072010153098600_bib38","doi-asserted-by":"publisher","first-page":"15849","DOI":"10.1109\/CVPR52688.2022.01539","article-title":"JoinABLe: Learning bottom-up assembly of parametric CAD joints","author":"Willis","year":"2022","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"2026072010153098600_bib39","doi-asserted-by":"publisher","first-page":"6772","DOI":"10.1109\/ICCV48922.2021.00670","article-title":"Deepcad: A deep generative network for computer-aided design models","author":"Wu","year":"2021","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision"},{"key":"2026072010153098600_bib40","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2411.04954","article-title":"Cad-mllm: Unifying multimodality-conditioned cad generation with mllm","author":"Xu","year":"2024"},{"key":"2026072010153098600_bib41","doi-asserted-by":"publisher","first-page":"131","DOI":"10.1007\/978-3-031-72698-9_8","article-title":"PointLLM: Empowering large language models to understand point clouds","author":"Xu","year":"2024","journal-title":"European Conference on Computer Vision"},{"key":"2026072010153098600_bib42","doi-asserted-by":"publisher","first-page":"300","DOI":"10.1093\/jcde\/qwae115","article-title":"NCFDet: Enhanced point cloud features using the neural collapse phenomenon in multimodal fusion for 3D object detection","volume":"12","author":"Xu","year":"2025","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026072010153098600_bib43","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2503.16416","article-title":"Survey on evaluation of LLM-based Agents","author":"Yehudai","year":"2025"},{"key":"2026072010153098600_bib44","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2205.01917","article-title":"CoCa: Contrastive Captioners are Image-Text Foundation Models","author":"Yu","year":"2022"},{"key":"2026072010153098600_bib45","doi-asserted-by":"publisher","first-page":"19313","DOI":"10.1109\/CVPR52688.2022.01871","article-title":"Point-BERT: Pre-training 3D point cloud transformers with masked point modeling","author":"Yu","year":"2022","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"2026072010153098600_bib46","doi-asserted-by":"publisher","first-page":"274","DOI":"10.1016\/j.jcde.2016.04.002","article-title":"A framework for similarity recognition of CAD models","volume":"3","author":"Zehtaban","year":"2016","journal-title":"Journal of Computational Design and Engineering"},{"key":"2026072010153098600_bib47","doi-asserted-by":"publisher","first-page":"5625","DOI":"10.1109\/TPAMI.2024.3369699","article-title":"Vision-language models for vision tasks: A survey","volume":"46","author":"Zhang","year":"2024","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026072010153098600_bib48","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2304.10592","article-title":"MiniGPT-4: Enhancing vision-language understanding with advanced large language models","author":"Zhu","year":"2023"}],"container-title":["Journal of Computational Design and Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/jcde\/advance-article-pdf\/doi\/10.1093\/jcde\/qwag039\/68106651\/qwag039.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/13\/5\/122\/68106651\/qwag039.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/jcde\/article-pdf\/13\/5\/122\/68106651\/qwag039.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T14:15:46Z","timestamp":1784556946000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/jcde\/article\/13\/5\/122\/8658277"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,17]]},"references-count":48,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,5,7]]}},"URL":"https:\/\/doi.org\/10.1093\/jcde\/qwag039","relation":{},"ISSN":["2288-5048"],"issn-type":[{"value":"2288-5048","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026,5]]},"published":{"date-parts":[[2026,4,17]]}}}