{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T02:24:33Z","timestamp":1783131873417,"version":"3.54.6"},"reference-count":0,"publisher":"SPIE-Intl Soc Optical Eng","issue":"04","funder":[{"name":"Science and Technology Project of Shandong Provincial Department of Transportation","award":["2021B120"],"award-info":[{"award-number":["2021B120"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"publisher","award":["2021M702030"],"award-info":[{"award-number":["2021M702030"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"publisher"},{"id":"https:\/\/ror.org\/0426zh255","id-type":"ROR","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Electron. Imag."],"published-print":{"date-parts":[[2026,7,4]]},"abstract":"<jats:p>Cameras and LiDAR are currently the two most commonly used sensors in autonomous driving. The effectiveness of feature fusion between camera and LiDAR modalities directly impacts the stability and safety of autonomous driving systems. Existing fusion methods typically concatenate bird\u2019s eye view (BEV) features from both modalities under a unified resolution, followed by convolutional or self-attention operations to integrate them. However, these approaches suffer from limited receptive fields and insufficient feature alignment. In this paper, we propose a multimodal and multiscale feature fusion framework tailored for 3D object detection and map segmentation tasks in autonomous driving. First, we design an uncertainty-aware prediction branch combined with a global alignment module to enable adaptive weighting and accurate alignment of LiDAR and camera BEV features. Then, we introduce an inter-head feature interaction strategy integrated into an efficient attention mechanism to facilitate information complementation and conflict correction among attention heads, enhancing global context awareness. Finally, a coarse-to-fine multimodal fusion strategy is presented, leveraging efficient attention to reduce the computational burden of the fusion process. Extensive experiments on the nuScenes and Waymo datasets demonstrate the superiority of our approach, achieving up to a 0.5% improvement in detection accuracy and a 2.4% increase in intersection over union for map segmentation compared with state-of-the-art methods.<\/jats:p>","DOI":"10.1117\/1.jei.35.4.043001","type":"journal-article","created":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T02:10:28Z","timestamp":1783131028000},"source":"Crossref","is-referenced-by-count":0,"title":["C2FMFusion: a multimodal BEV feature fusion method for autonomous driving"],"prefix":"10.1117","volume":"35","author":[{"given":"Yali","family":"Liu","sequence":"first","affiliation":[{"name":"Shandong Jiaotong University, School of Information Science and Electrical Engineering (Artificial Intelligence College), Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-8878-1201","authenticated-orcid":false,"given":"Cui","family":"Ni","sequence":"additional","affiliation":[{"name":"Shandong Jiaotong University, School of Information Science and Electrical Engineering (Artificial Intelligence College), Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dongqing","family":"Yang","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Advanced Technology Research Institute, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wen","family":"Rong","sequence":"additional","affiliation":[{"name":"Shandong Hi-speed Group Co.LTD Innovation Research Institute, Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guangyuan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shandong Jiaotong University, School of Information Science and Electrical Engineering (Artificial Intelligence College), Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4771-7984","authenticated-orcid":false,"given":"Peng","family":"Wang","sequence":"additional","affiliation":[{"name":"Shandong Jiaotong University, School of Information Science and Electrical Engineering (Artificial Intelligence College), Jinan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"189","container-title":["Journal of Electronic Imaging"],"original-title":[],"deposited":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T02:10:28Z","timestamp":1783131028000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.spiedigitallibrary.org\/journals\/journal-of-electronic-imaging\/volume-35\/issue-04\/043001\/C2FMFusion--a-multimodal-BEV-feature-fusion-method-for-autonomous\/10.1117\/1.JEI.35.4.043001.full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,4]]},"references-count":0,"journal-issue":{"issue":"04","published-online":{"date-parts":[[2026,7,1]]}},"URL":"https:\/\/doi.org\/10.1117\/1.jei.35.4.043001","relation":{},"ISSN":["1017-9909"],"issn-type":[{"value":"1017-9909","type":"print"}],"subject":[],"published":{"date-parts":[[2026,7,4]]}}}