{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:16:54Z","timestamp":1760145414924,"version":"build-2065373602"},"reference-count":38,"publisher":"MDPI AG","issue":"14","license":[{"start":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T00:00:00Z","timestamp":1721088000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Education Research Project for Young and Middle-aged Teachers of Fujian Provincial Department of Education","award":["JAT220310","MJY22025"],"award-info":[{"award-number":["JAT220310","MJY22025"]}]},{"name":"Minjiang University Scientific Research Promotion Fund","award":["JAT220310","MJY22025"],"award-info":[{"award-number":["JAT220310","MJY22025"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In autonomous driving, the fusion of multiple sensors is considered essential to improve the accuracy and safety of 3D object detection. Currently, a fusion scheme combining low-cost cameras with highly robust radars can counteract the performance degradation caused by harsh environments. In this paper, we propose the IRBEVF-Q model, which mainly consists of BEV (Bird\u2019s Eye View) fusion coding module and an object decoder module.The BEV fusion coding module solves the problem of unified representation of different modal information by fusing the image and radar features through 3D spatial reference points as a medium. The query in the object decoder, as a core component, plays an important role in detection. In this paper, Heat Map-Guided Query Initialization (HGQI) and Dynamic Position Encoding (DPE) are proposed in query construction to increase the a priori information of the query. The Auxiliary Noise Query (ANQ) then helps to stabilize the matching. The experimental results demonstrate that the proposed fusion model IRBEVF-Q achieves an NDS of 0.575 and a mAP of 0.476 on the nuScenes test set. Compared to recent state-of-the-art methods, our model shows significant advantages, thus indicating that our approach contributes to improving detection accuracy.<\/jats:p>","DOI":"10.3390\/s24144602","type":"journal-article","created":{"date-parts":[[2024,7,16]],"date-time":"2024-07-16T09:46:57Z","timestamp":1721123217000},"page":"4602","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["IRBEVF-Q: Optimization of Image\u2013Radar Fusion Algorithm Based on Bird\u2019s Eye View Features"],"prefix":"10.3390","volume":"24","author":[{"given":"Ganlin","family":"Cai","sequence":"first","affiliation":[{"name":"School of Computer and Big Data, Minjiang University, Fuzhou 350108, China"},{"name":"College of Physics and Information Engineering, Fuzhou University, Fuzhou 350108, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Feng","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Physics and Information Engineering, Fuzhou University, Fuzhou 350108, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6772-3954","authenticated-orcid":false,"given":"Ente","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Computer and Big Data, Minjiang University, Fuzhou 350108, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,7,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Wang, T., Zhu, X., Pang, J., and Lin, D. (2021, January 10\u201317). Fcos3d: Fully convolutional one-stage monocular 3d object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00107"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1341","DOI":"10.1109\/TITS.2020.2972974","article-title":"Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges","volume":"22","author":"Feng","year":"2020","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Vora, S., Lang, A.H., Helou, B., and Beijbom, O. (2020, January 13\u201319). Pointpainting: Sequential fusion for 3d object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00466"},{"key":"ref_4","first-page":"16494","article-title":"Multimodal virtual point 3d detection","volume":"34","author":"Yin","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Wang, C., Ma, C., Zhu, M., and Yang, X. (2021, January 20\u201325). Pointaugmenting: Cross-modal augmentation for 3d object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01162"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Xu, S., Zhou, D., Fang, J., Yin, J., Bin, Z., and Zhang, L. (2021, January 19\u201322). Fusionpainting: Multimodal fusion with adaptive attention for 3d object detection. Proceedings of the 2021 IEEE International Intelligent Transportation Systems Conference (ITSC), Indianapolis, IN, USA.","DOI":"10.1109\/ITSC48978.2021.9564951"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Bai, X., Hu, Z., Zhu, X., Huang, Q., Chen, Y., Fu, H., and Tai, C.L. (2022, January 18\u201324). Transfusion: Robust lidar-camera fusion for 3d object detection with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00116"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"11182","DOI":"10.1109\/LRA.2022.3193465","article-title":"EZFusion: A Close Look at the Integration of LiDAR, Millimeter-Wave Radar, and Camera for Accurate 3D Object Detection and Tracking","volume":"7","author":"Li","year":"2022","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ma, X., Zhang, Y., Xu, D., Zhou, D., Yi, S., Li, H., and Ouyang, W. (2021, January 20\u201325). Delving into localization errors for monocular 3d object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00469"},{"key":"ref_10","unstructured":"Lim, T.Y., Ansari, A., Major, B., Fontijne, D., Hamilton, M., Gowaikar, R., and Subramanian, S. (2019, January 8\u201314). Radar and camera early fusion for vehicle detection in advanced driver assistance systems. Proceedings of the Machine Learning for Autonomous Driving Workshop at the 33rd Conference on Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Kim, Y., Choi, J.W., and Kum, D. (2020, January 25\u201329). Grif net: Gated region of interest fusion network for robust 3d object detection from radar point cloud and monocular image. Proceedings of the 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA.","DOI":"10.1109\/IROS45743.2020.9341177"},{"key":"ref_12","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). Pointnet: Deep learning on point sets for 3d classification and segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA."},{"key":"ref_13","first-page":"5099","article-title":"Pointnet++: Deep hierarchical feature learning on point sets in a metric space","volume":"30","author":"Qi","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Mao, J., Xue, Y., Niu, M., Bai, H., Feng, J., Liang, X., Xu, H., and Xu, C. (2021, January 11\u201317). Voxel transformer for 3d object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00315"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Lang, A.H., Vora, S., Caesar, H., Zhou, L., Yang, J., and Beijbom, O. (2019, January 15\u201320). Pointpillars: Fast encoders for object detection from point clouds. Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01298"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Shi, S., Guo, C., Jiang, L., Wang, Z., Shi, J., Wang, X., and Li, H. (2020, January 13\u201319). Pv-rcnn: Point-voxel feature set abstraction for 3d object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01054"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Nabati, R., and Qi, H. (2021, January 5\u20139). Centerfusion: Center-based radar and camera fusion for 3d object detection. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Virtual.","DOI":"10.1109\/WACV48630.2021.00157"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kim, Y., Kim, S., Choi, J.W., and Kum, D. (2023, January 20\u201327). Craft: Camera-radar 3d object detection with spatio-contextual fusion transformer. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v37i1.25198"},{"key":"ref_19","first-page":"1992","article-title":"Deepinteraction: 3d object detection via modality interaction","volume":"35","author":"Yang","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Cai, Y., Zhang, W., Wu, Y., and Jin, C. (2024, January 20\u201324). FusionFormer: A Concise Unified Feature Fusion Transformer for 3D Pose Estimation. Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada.","DOI":"10.1609\/aaai.v38i2.27849"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Chen, Y., Yu, Z., Chen, Y., Lan, S., An kumar, A., Jia, J., and Alvarez, J.M. (2023, January 2\u20136). Focalformer3d: Focusing on hard instance for 3d object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00771"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Liu, H., Teng, Y., Lu, T., Wang, H., and Wang, L. (2023, January 2\u20136). Sparsebev: High-performance sparse 3d object detection from multi-camera videos. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01703"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, H., Tang, H., Shi, S., Li, A., Li, Z., Schiele, B., and Wang, L. (2023, January 2\u20136). UniTR: A Unified and Efficient Multi-Modal Transformer for Bird\u2019s-Eye-View Representation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00625"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Yan, J., Liu, Y., Sun, J., Jia, F., Li, S., Wang, T., and Zhang, X. (2023, January 2\u20136). Cross modal transformer: Towards fast and robust 3d object detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01675"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Dong, X., Zhuang, B., Mao, Y., and Liu, L. (2021, January 20\u201325). Radar camera fusion via representation learning in autonomous driving. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPRW53098.2021.00183"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Li, Z., Wang, W., Li, H., Xie, E., Sima, C., Lu, T., Qiao, Y., and Dai, J. (2022, January 23\u201327). Bevformer: Learning bird\u2019s-eye-view representation from multi-camera images via spatiotemporal transformers. Proceedings of the European conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20077-9_1"},{"key":"ref_27","unstructured":"Ma, Y., Wang, T., Bai, X., Yang, H., Hou, Y., Wang, Y., Qiao, Y., Yang, R., Manocha, D., and Zhu, X. (2022). Vision-centric bev perception: A survey. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_30","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_31","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. (2020). Deformable detr: Deformable transformers for end-to-end object detection. arXiv."},{"key":"ref_32","unstructured":"Tang, Y., Dorn, S., and Savani, C. (October, January 28). Center3d: Center-based monocular 3d object detection with joint depth understanding. Proceedings of the DAGM German Conference on Pattern Recognition, T\u00fcbingen, Germany."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. (2020, January 13\u201319). nuscenes: A multimodal dataset for autonomous driving. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1523","DOI":"10.1109\/TIV.2023.3240287","article-title":"Bridging the view disparity between radar and camera features for multi-modal fusion 3d object detection","volume":"8","author":"Zhou","year":"2023","journal-title":"IEEE Trans. Intell. Veh."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Yu, F., Wang, D., Shelhamer, E., and Darrell, T. (2018, January 18\u201323). Deep layer aggregation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00255"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Wu, Z., Chen, G., Gan, Y., Wang, L., and Pu, J. (June, January 29). Mvfusion: Multi-view 3d object detection with semantic-aligned radar and camera fusion. Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA), London, UK.","DOI":"10.1109\/ICRA48891.2023.10161329"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lee, Y., Hwang, J.W., Lee, S., Bae, Y., and Park, J. (2019, January 16\u201317). An energy and GPU-computation efficient backbone network for real-time object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Long Beach, CA, USA.","DOI":"10.1109\/CVPRW.2019.00103"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/14\/4602\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:17:30Z","timestamp":1760109450000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/14\/4602"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,16]]},"references-count":38,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2024,7]]}},"alternative-id":["s24144602"],"URL":"https:\/\/doi.org\/10.3390\/s24144602","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2024,7,16]]}}}