{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T14:30:25Z","timestamp":1781706625692,"version":"3.54.5"},"reference-count":40,"publisher":"MDPI AG","issue":"15","license":[{"start":{"date-parts":[[2024,7,24]],"date-time":"2024-07-24T00:00:00Z","timestamp":1721779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"the National Natural Science Foundation of China","award":["62271311"],"award-info":[{"award-number":["62271311"]}]},{"name":"the National Natural Science Foundation of China","award":["62071333"],"award-info":[{"award-number":["62071333"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Robust object detection in complex environments, poor visual conditions, and open scenarios presents significant technical challenges in autonomous driving. These challenges necessitate the development of advanced fusion methods for millimeter-wave (mmWave) radar point cloud data and visual images. To address these issues, this paper proposes a radar\u2013camera robust fusion network (RCRFNet), which leverages self-supervised learning and open-set recognition to effectively utilise the complementary information from both sensors. Specifically, the network uses matched radar\u2013camera data through a frustum association approach to generate self-supervised signals, enhancing network training. The integration of global and local depth consistencies between radar point clouds and visual images, along with image features, helps construct object class confidence levels for detecting unknown targets. Additionally, these techniques are combined with a multi-layer feature extraction backbone and a multimodal feature detection head to achieve robust object detection. Experiments on the nuScenes public dataset demonstrate that RCRFNet outperforms state-of-the-art (SOTA) methods, particularly in conditions of low visual visibility and when detecting unknown class objects.<\/jats:p>","DOI":"10.3390\/s24154803","type":"journal-article","created":{"date-parts":[[2024,7,24]],"date-time":"2024-07-24T14:55:47Z","timestamp":1721832947000},"page":"4803","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["RCRFNet: Enhancing Object Detection with Self-Supervised Radar\u2013Camera Fusion and Open-Set Recognition"],"prefix":"10.3390","volume":"24","author":[{"given":"Minwei","family":"Chen","sequence":"first","affiliation":[{"name":"Shanghai Key Laboratory of Intelligent Sensing and Recognition, Shanghai Jiao Tong University, Shanghai 200240, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yajun","family":"Liu","sequence":"additional","affiliation":[{"name":"Shanghai Key Laboratory of Intelligent Sensing and Recognition, Shanghai Jiao Tong University, Shanghai 200240, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1238-8538","authenticated-orcid":false,"given":"Zenghui","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shanghai Key Laboratory of Intelligent Sensing and Recognition, Shanghai Jiao Tong University, Shanghai 200240, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5037-0972","authenticated-orcid":false,"given":"Weiwei","family":"Guo","sequence":"additional","affiliation":[{"name":"Center of Digital Innovation, Tongji University, Shanghai 200092, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,7,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"6401","DOI":"10.1109\/TSMC.2023.3283021","article-title":"Milestones in Autonomous Driving and Intelligent Vehicles\u2014Part II: Perception and Planning","volume":"53","author":"Chen","year":"2023","journal-title":"IEEE Trans. Syst. Man. Cybern. Syst."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Wei, Z., Zhang, F., Chang, S., Liu, Y., Wu, H., and Feng, Z. (2022). MmWave Radar and Vision Fusion for Object Detection in Autonomous Driving: A Review. Sensors, 22.","DOI":"10.3390\/s22072542"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Huo, B., Li, C., Zhang, J., Xue, Y., and Lin, Z. (2023). SAFF-SSD: Self-Attention Combined Feature Fusion-Based SSD for Small Object Detection in Remote Sensing. Remote Sens., 15.","DOI":"10.3390\/rs15123027"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Carion, N., and Massa, F. (2020, January 23\u201328). End-to-End Object Detection with Transformers. Proceedings of the Computer Vision\u2014ECCV 2020, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2151","DOI":"10.1109\/TPAMI.2023.3333838","article-title":"Delving Into the Devils of Bird\u2019s-Eye-View Perception: A Review, Evaluation and Recipe","volume":"46","author":"Li","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","first-page":"1","article-title":"Transferable SAR Image Classification Crossing Different Satellites under Open Set Condition","volume":"19","author":"Zhao","year":"2022","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TGRS.2022.3230378","article-title":"An Automatic Ship Detection Method Adapting to Different Satellites SAR Images with Feature Alignment and Compensation Loss","volume":"60","author":"Zhao","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"16296","DOI":"10.1109\/ACCESS.2021.3052506","article-title":"Feature Fusion Based on Bayesian Decision Theory for Radar Deception Jamming Recognition","volume":"9","author":"Zhou","year":"2021","journal-title":"IEEE Access"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Bendale, A., and Boult, T.E. (July, January 27). Towards Open Set Deep Networks. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.173"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Bao, W., Yu, Q., and Kong, Y. (2021, January 11\u201317). Evidential Deep Learning for Open Set Action Recognition. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01310"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Svenningsson, P., Fioranelli, F., and Yarovoy, A. (2021, January 8\u201314). Radar-PointGNN: Graph Based Object Recognition for Unstructured Radar Point-cloud Data. Proceedings of the 2021 IEEE Radar Conference (RadarConf21), Atlanta, GA, USA.","DOI":"10.1109\/RadarConf2147009.2021.9455172"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Meyer, M., Kuschk, G., and Tomforde, S. (2021, January 11\u201317). Graph Convolutional Networks for 3D Object Detection on Radar Data. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00340"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Zhang, L., Wu, J., and Guo, W. (2024). Optical and Synthetic Aperture Radar Image Fusion for Ship Detection and Recognition: Current state, challenges, and future prospects. IEEE Geosci. Remote Sens. Mag., 2\u201338.","DOI":"10.1109\/MGRS.2024.3404506"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Bi, X., Tan, B., Xu, Z., and Huang, L. (1977). A New Method of Target Detection Based on Autonomous Radar and Camera Data Fusion, SAE International. 2017-01-1977.","DOI":"10.4271\/2017-01-1977"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Chang, S., Zhang, Y., Zhang, F., Zhao, X., Huang, S., Feng, Z., and Wei, Z. (2020). Spatial Attention Fusion for Obstacle Detection Using MmWave Radar and Vision Sensor. Sensors, 20.","DOI":"10.3390\/s20040956"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Drews, F., Feng, D., Faion, F., Rosenbaum, L., Ulrich, M., and Gl\u00e4ser, C. (2022, January 23\u201327). DeepFusion: A Robust and Modular 3D Object Detector for Lidars, Cameras and Radars. Proceedings of the 2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Kyoto, Japan.","DOI":"10.1109\/IROS47612.2022.9981778"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Nabati, R., and Qi, H. (2021, January 5\u20139). CenterFusion: Center-based Radar and Camera Fusion for 3D Object Detection. Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA.","DOI":"10.1109\/WACV48630.2021.00157"},{"key":"ref_19","unstructured":"Kim, J., Kim, Y., and Kum, D. (December, January 30). Low-level sensor fusion network for 3d vehicle detection using radar range-azimuth heatmap and monocular image. Proceedings of the Asian Conference on Computer Vision, Kyoto, Japan."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Chavez-Garcia, R.O., Burlet, J., Vu, T.D., and Aycard, O. (2012, January 3\u20137). Frontal object perception using radar and mono-vision. Proceedings of the 2012 IEEE Intelligent Vehicles Symposium, Alcala de Henares, Spain.","DOI":"10.1109\/IVS.2012.6232307"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"338","DOI":"10.1016\/j.robot.2016.05.001","article-title":"Radar and stereo vision fusion for multitarget tracking on the special Euclidean group","volume":"83","year":"2016","journal-title":"Robot. Auton. Syst."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Wang, X., Xu, L., Sun, H., Xin, J., and Zheng, N. (2014, January 8\u201311). Bionic vision inspired on-road obstacle detection and tracking using radar and visual information. Proceedings of the 17th International IEEE Conference on Intelligent Transportation Systems (ITSC), Qingdao, China.","DOI":"10.1109\/ITSC.2014.6957663"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Guo, X., Du, J., Gao, J.Y., and Wang, W. (2018, January 18\u201320). Pedestrian Detection Based on Fusion of Millimeter Wave Radar and Vision. Proceedings of the 2018 International Conference on Artificial Intelligence and Pattern Recognition, Beijing, China.","DOI":"10.1145\/3268866.3268868"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Kadow, U., Schneider, G., and Vukotich, A. (2007, January 13\u201315). Radar-Vision Based Vehicle Recognition with Evolutionary Optimized and Boosted Features. Proceedings of the 2007 IEEE Intelligent Vehicles Symposium, Istanbul, Turkey.","DOI":"10.1109\/IVS.2007.4290206"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"131","DOI":"10.1109\/MITS.2023.3298534","article-title":"Collaborative Perception in Autonomous Driving: Methods, Datasets, and Challenges","volume":"15","author":"Han","year":"2023","journal-title":"IEEE Intell. Transp. Syst. Mag."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Nobis, F., Geisslinger, M., Weber, M., Betz, J., and Lienkamp, M. (2019, January 15\u201317). A Deep Learning-based Radar and Camera Sensor Fusion Architecture for Object Detection. Proceedings of the 2019 Sensor Data Fusion: Trends, Solutions, Applications (SDF), Bonn, Germany.","DOI":"10.1109\/SDF.2019.8916629"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Maiti, A., Elberink, S.O., and Vosselman, G. (2023, January 17\u201324). TransFusion: Multi-modal Fusion Network for Semantic Segmentation. Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW59228.2023.00695"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Aziz, K., De Greef, E., Rykunov, M., Bourdoux, A., and Sahli, H. (2020, January 21\u201325). Radar-camera Fusion for Road Target Classification. Proceedings of the 2020 IEEE Radar Conference (RadarConf20), Florence, Italy.","DOI":"10.1109\/RadarConf2043947.2020.9266510"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Kim, J., Emer\u0161i\u010d, \u017d., and Han, D.S. (2019, January 11\u201313). Vehicle Path Prediction based on Radar and Vision Sensor Fusion for Safe Lane Changing. Proceedings of the 2019 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), Okinawa, Japan.","DOI":"10.1109\/ICAIIC.2019.8669081"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Bertozzi, M., Bombini, L., Cerri, P., Medici, P., Antonello, P.C., and Miglietta, M. (2008, January 4\u20136). Obstacle detection and classification fusing radar and vision. Proceedings of the 2008 IEEE Intelligent Vehicles Symposium, Eindhoven, The Netherlands.","DOI":"10.1109\/IVS.2008.4621304"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Obrvan, M., Cesic, J., and Petrovi\u0107, I. (2015, January 19\u201321). Appearance Based Vehicle Detection by Radar-Stereo Vision Integration. Proceedings of the ROBOT, Lisbon, Portugal.","DOI":"10.1007\/978-3-319-27146-0_34"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Yu, F., Wang, D., Shelhamer, E., and Darrell, T. (2018, January 18\u201323). Deep Layer Aggregation. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00255"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Yoshihashi, R., Shao, W., Kawakami, R., You, S., Iida, M., and Naemura, T. (2019, January 15\u201320). Classification-Reconstruction Learning for Open-Set Recognition. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00414"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"137065","DOI":"10.1109\/ACCESS.2019.2942382","article-title":"Extending Reliability of mmWave Radar Tracking and Detection via Fusion with Camera","volume":"7","author":"Zhang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Zhao, Z., and Dong, M. (2023, January 21\u201323). Channel-Spatial Dynamic Convolution: An Exquisite Omni-dimensional Dynamic Convolution. Proceedings of the 2023 8th International Conference on Intelligent Computing and Signal Processing (ICSP), Xi\u2019an, China.","DOI":"10.1109\/ICSP58490.2023.10248781"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Chen, X., Kundu, K., Zhang, Z., Ma, H., Fidler, S., and Urtasun, R. (2016, January 27\u201330). Monocular 3D Object Detection for Autonomous Driving. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.236"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Duan, K., Bai, S., Xie, L., Qi, H., Huang, Q., and Tian, Q. (November, January 27). CenterNet: Keypoint Triplets for Object Detection. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00667"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Wang, T., Zhu, X., Pang, J., and Lin, D. (2021, January 11\u201317). FCOS3D: Fully Convolutional One-Stage Monocular 3D Object Detection. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW), Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00107"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. (2020, January 13\u201319). nuScenes: A Multimodal Dataset for Autonomous Driving. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01164"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/15\/4803\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:22:37Z","timestamp":1760109757000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/15\/4803"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,24]]},"references-count":40,"journal-issue":{"issue":"15","published-online":{"date-parts":[[2024,8]]}},"alternative-id":["s24154803"],"URL":"https:\/\/doi.org\/10.3390\/s24154803","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,24]]}}}