{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,16]],"date-time":"2026-02-16T20:47:38Z","timestamp":1771274858409,"version":"3.50.1"},"reference-count":45,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,12,21]],"date-time":"2024-12-21T00:00:00Z","timestamp":1734739200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62272035"],"award-info":[{"award-number":["62272035"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,1,31]]},"abstract":"<jats:p>Recently, object detection methods based on multi-modal fusion have gained widespread adoption in autonomous driving, proving to be valuable for detecting objects in dynamic environments. Among them, millimeter wave (mmWave) radar is commonly utilized as an effective complement to cameras, as it is almost unaffected by harsh weather conditions. However, current approaches that fuse mmWave radar and camera often overlook the correlation between the two modalities, failing to fully exploit their complementary features. To address this, we propose a temporal-enhanced radar and camera fusion network to explore the correlation between these two modalities and learn a comprehensive representation for object detection. In our model, a temporal fusion model is introduced to fuse mmWave radar features from different moments, thus mitigating the problem of mmWave radar point-object mismatch due to object movement. Moreover, a new correlation-based fusion strategy using the dedicated mask cross-attention is proposed to fuse mmWave radar and vision features more effectively. Finally, we design a gate feature pyramid network that selects shallow texture information based on deep semantic information to obtain more representative features. The experimental results on the nuScenes benchmark demonstrate the effectiveness of our proposed method.<\/jats:p>","DOI":"10.1145\/3700442","type":"journal-article","created":{"date-parts":[[2024,10,14]],"date-time":"2024-10-14T13:57:49Z","timestamp":1728914269000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Temporal-Enhanced Radar and Camera Fusion for Object Detection"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0006-0453-0879","authenticated-orcid":false,"given":"Linhua","family":"Kong","sequence":"first","affiliation":[{"name":"Institute of Information Science, Visual Intellgence +X International Cooperation Joint Laboratory of MOE, Beijing Jiaotong University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8765-7640","authenticated-orcid":false,"given":"Yiming","family":"Wang","sequence":"additional","affiliation":[{"name":"Nanjing University of Posts and Telecommunications, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9718-4277","authenticated-orcid":false,"given":"Dongxia","family":"Chang","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Visual Intellgence +X International Cooperation Joint Laboratory of MOE, Beijing Jiaotong University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8581-9554","authenticated-orcid":false,"given":"Yao","family":"Zhao","sequence":"additional","affiliation":[{"name":"Institute of Information Science, Visual Intellgence +X International Cooperation Joint Laboratory of MOE, Beijing Jiaotong University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,12,21]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00116"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.robot.2016.05.001"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8794312"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.3390\/s20040956"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00041"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01212"},{"key":"e_1_3_1_10_2","first-page":"8394","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chen Yilun","year":"2023","unstructured":"Yilun Chen, Zhiding Yu, Yukang Chen, Shiyi Lan, Anima Anandkumar, Jiaya Jia, and Jose M. Alvarez. 2023. FocalFormer3D: Focusing on hard instance for 3D object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 8394\u20138405."},{"key":"e_1_3_1_11_2","first-page":"15263","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Cheng Yuwei","year":"2021","unstructured":"Yuwei Cheng, Hu Xu, and Yimin Liu. 2021. Robust small object detection on the water surface through fusion of camera and millimeter wave radar. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 15263\u201315272."},{"key":"e_1_3_1_12_2","first-page":"16344","volume-title":"Proceedings of the 36th International Conference on Advances in Neural Information Processing System","volume":"35","author":"Dao Tri","year":"2022","unstructured":"Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher R\u00e9. 2022. FlashAttention: Fast and memory-efficient exact attention with IO-awareness. In Proceedings of the 36th International Conference on Advances in Neural Information Processing System, Vol. 35, 16344\u201316359."},{"key":"e_1_3_1_13_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_14_2","first-page":"1","volume-title":"Proceedings of the 2022 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC \u201922)","author":"Gu Yaqing","year":"2022","unstructured":"Yaqing Gu, Shiyuan Meng, and Kun Shi. 2022. Radar-enhanced image fusion-based object detection for autonomous driving. In Proceedings of the 2022 IEEE International Conference on Signal Processing, Communications and Computing (ICSPCC \u201922). IEEE, 1\u20136."},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00548"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00599"},{"key":"e_1_3_1_17_2","first-page":"75782","article-title":"Query-based temporal fusion with explicit motion for 3D object detection","volume":"36","author":"Hou Jinghua","year":"2024","unstructured":"Jinghua Hou, Zhe Liu, Zhikang Zou, Dingkang Liang, Xiaoqing Ye, and Xiang Bai. 2024. Query-based temporal fusion with explicit motion for 3D object detection. In Advances in Neural Information Processing Systems, Vol. 36, 75782\u201375797.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00069"},{"key":"e_1_3_1_19_2","doi-asserted-by":"crossref","first-page":"351","DOI":"10.1007\/978-3-030-34879-3_27","volume-title":"Proceedings of the Pacific-Rim Symposium on Image and Video Technology","author":"John Vijay","year":"2019","unstructured":"Vijay John and Seiichi Mita. 2019. RVNet: Deep sensor fusion of monocular camera and radar for image-based obstacle detection in challenging environments. In Proceedings of the Pacific-Rim Symposium on Image and Video Technology. Springer, 351\u2013364."},{"key":"e_1_3_1_20_2","doi-asserted-by":"crossref","first-page":"749","DOI":"10.1109\/IVS.2007.4290206","volume-title":"Proceedings of the 2007 IEEE Intelligent Vehicles Symposium","author":"Kadow Ulrich","year":"2007","unstructured":"Ulrich Kadow, Georg Schneider, and Alejandro Vukotich. 2007. Radar-vision based vehicle recognition with evolutionary optimized and boosted features. In Proceedings of the 2007 IEEE Intelligent Vehicles Symposium. IEEE, 749\u2013754."},{"key":"e_1_3_1_21_2","first-page":"234","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Kim Seung-Wook","year":"2018","unstructured":"Seung-Wook Kim, Hyong-Keun Kook, Jee-Young Sun, Mun-Cheon Kang, and Sung-Jea Ko. 2018. Parallel feature pyramid network for object detection. In Proceedings of the European Conference on Computer Vision (ECCV \u201918), 234\u2013250."},{"key":"e_1_3_1_22_2","first-page":"17615","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Kim Youngseok","year":"2023","unstructured":"Youngseok Kim, Juyeb Shin, Sanmin Kim, In-Jae Lee, Jun Won Choi, and Dongsuk Kum. 2023. CRN: Camera radar net for accurate, robust, efficient 3D perception. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 17615\u201317626."},{"key":"e_1_3_1_23_2","first-page":"366","volume-title":"Proceedings of the 2020 15th IEEE International Conference on Signal Processing (ICSP \u201920)","volume":"1","author":"Li Liang-qun","year":"2020","unstructured":"Liang-qun Li and Yuan-liang Xie. 2020. A feature pyramid fusion detection algorithm based on radar and camera sensor. In Proceedings of the 2020 15th IEEE International Conference on Signal Processing (ICSP \u201920), Vol. 1, IEEE, 366\u2013370."},{"key":"e_1_3_1_24_2","first-page":"17182","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li Yingwei","year":"2022","unstructured":"Yingwei Li, Adams Wei Yu, Tianjian Meng, Ben Caine, Jiquan Ngiam, Daiyi Peng, Junyang Shen, Yifeng Lu, Denny Zhou, Quoc V. Le, et al. 2022. DeepFusion: Lidar-camera deep fusion for multi-modal 3D object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 17182\u201317191."},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.106"},{"key":"e_1_3_1_26_2","first-page":"657","volume-title":"Proceedings of the Progress in Automobile Lighting, Held Laboratory of Lighting Technology (PAL \u201901)","volume":"9","author":"Milch Stefan","year":"2001","unstructured":"Stefan Milch and Marc Behrens. 2001. Pedestrian detection with radar and computer vision. In Proceedings of the Progress in Automobile Lighting, Held Laboratory of Lighting Technology (PAL \u201901), Vol. 9, 657\u2013664."},{"key":"e_1_3_1_27_2","first-page":"3093","volume-title":"Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP \u201919)","author":"Nabati Ramin","year":"2019","unstructured":"Ramin Nabati and Hairong Qi. 2019. RRPN: Radar region proposal network for object detection in autonomous vehicles. In Proceedings of the 2019 IEEE International Conference on Image Processing (ICIP \u201919). IEEE, 3093\u20133097."},{"key":"e_1_3_1_28_2","unstructured":"Ramin Nabati and Hairong Qi. 2020. Radar-camera sensor fusion for joint object detection and distance estimation in autonomous vehicles. arXiv:2009.08428. Retrieved from https:\/\/arxiv.org\/abs\/2009.08428"},{"key":"e_1_3_1_29_2","first-page":"1","volume-title":"Proceedings of the 2019 Sensor Data Fusion: Trends, Solutions, Applications (SDF \u201919)","author":"Nobis Felix","year":"2019","unstructured":"Felix Nobis, Maximilian Geisslinger, Markus Weber, Johannes Betz, and Markus Lienkamp. 2019. A deep learning-based radar and camera sensor fusion architecture for object detection. In Proceedings of the 2019 Sensor Data Fusion: Trends, Solutions, Applications (SDF \u201919). IEEE, 1\u20137."},{"key":"e_1_3_1_30_2","first-page":"91","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster R-CNN: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems, Vol. 28, 91\u201399.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_31_2","first-page":"802","article-title":"Convolutional LSTM network: A machine learning approach for precipitation nowcasting","volume":"28","author":"Shi Xingjian","year":"2015","unstructured":"Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wang-chun Woo. 2015. Convolutional LSTM network: A machine learning approach for precipitation nowcasting. In Advances in Neural Information Processing Systems, Vol. 28, 802\u2013810.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_32_2","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"e_1_3_1_33_2","first-page":"3358","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Song Ziying","year":"2023","unstructured":"Ziying Song, Haiyue Wei, Lin Bai, Lei Yang, and Caiyan Jia. 2023. GraphAlign: Enhancing accurate feature alignment by graph matching for multi-modal 3D object detection. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 3358\u20133369."},{"key":"e_1_3_1_34_2","first-page":"3087","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"St\u00e4cker Lukas","year":"2022","unstructured":"Lukas St\u00e4cker, Philipp Heidenreich, Jason Rambach, and Didier Stricker. 2022. Fusion point pruning for optimized 2D object detection with radar-camera fusion. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 3087\u20133094."},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"e_1_3_1_36_2","first-page":"6000","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30, 6000\u20136010.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466780"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TITS.2009.2032769"},{"key":"e_1_3_1_39_2","first-page":"3282","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Wu Zizhang","year":"2024","unstructured":"Zizhang Wu, Yunzhe Wu, Xiaoquan Wang, Yuanzhu Gan, and Jian Pu. 2024. A robust diffusion modeling framework for radar camera 3D object detection. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, 3282\u20133292."},{"key":"e_1_3_1_40_2","first-page":"8461","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Xia Yan","year":"2023","unstructured":"Yan Xia, Mariia Gladkova, Rui Wang, Qianyun Li, Uwe Stilla, Joao F. Henriques, and Daniel Cremers. 2023. CASSPR: Cross attention single scan place recognition. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, 8461\u20138472."},{"key":"e_1_3_1_41_2","first-page":"1986","volume-title":"Proceedings of the 2020 IEEE International Conference on Image Processing (ICIP \u201920)","author":"Yadav Ritu","year":"2020","unstructured":"Ritu Yadav, Axel Vierling, and Karsten Berns. 2020. Radar+ RGB fusion for robust object detection in autonomous vehicle. In Proceedings of the 2020 IEEE International Conference on Image Processing (ICIP \u201920). IEEE, 1986\u20131990."},{"key":"e_1_3_1_42_2","first-page":"11794","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang Zongxin","year":"2020","unstructured":"Zongxin Yang, Linchao Zhu, Yu Wu, and Yi Yang. 2020. Gated channel transformation for visual recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, 11794\u201311803."},{"issue":"1","key":"e_1_3_1_43_2","first-page":"1524","article-title":"SparseFusion3D: Sparse sensor fusion for 3D object detection by radar and camera in environmental perception","volume":"9","author":"Yu Zedong","year":"2023","unstructured":"Zedong Yu, Weibing Wan, Maiyu Ren, Xiuyuan Zheng, and Zhijun Fang. 2023. SparseFusion3D: Sparse sensor fusion for 3D object detection by radar and camera in environmental perception. IEEE Transactions on Intelligent Vehicles 9, 1 (2023), 1524\u20131536.","journal-title":"IEEE Transactions on Intelligent Vehicles"},{"key":"e_1_3_1_44_2","first-page":"286","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV \u201918)","author":"Zhang Yulun","year":"2018","unstructured":"Yulun Zhang, Kunpeng Li, Kai Li, Lichen Wang, Bineng Zhong, and Yun Fu. 2018. Image super-resolution using very deep residual channel attention networks. In Proceedings of the European Conference on Computer Vision (ECCV \u201918), 286\u2013301."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00276"},{"key":"e_1_3_1_46_2","first-page":"251","article-title":"Camera radar fusion for increased reliability in ADAS applications","volume":"2018","author":"Zhong Z.","year":"2018","unstructured":"Z. Zhong, S. Liu, M. Mathew, and A. Dubey. 2018. Camera radar fusion for increased reliability in ADAS applications. Electronic Imaging 2018 (2018), 251\u2013254.","journal-title":"Electronic Imaging"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3700442","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3700442","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:17:38Z","timestamp":1750295858000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3700442"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,21]]},"references-count":45,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1,31]]}},"alternative-id":["10.1145\/3700442"],"URL":"https:\/\/doi.org\/10.1145\/3700442","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,21]]},"assertion":[{"value":"2024-02-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-06","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-21","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}