{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T07:03:06Z","timestamp":1781593386264,"version":"3.54.5"},"reference-count":59,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2020,2,11]],"date-time":"2020-02-11T00:00:00Z","timestamp":1581379200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61801052, 61227801, 61941102, 61631003"],"award-info":[{"award-number":["61801052, 61227801, 61941102, 61631003"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>For autonomous driving, it is important to detect obstacles in all scales accurately for safety consideration. In this paper, we propose a new spatial attention fusion (SAF) method for obstacle detection using mmWave radar and vision sensor, where the sparsity of radar points are considered in the proposed SAF. The proposed fusion method can be embedded in the feature-extraction stage, which leverages the features of mmWave radar and vision sensor effectively. Based on the SAF, an attention weight matrix is generated to fuse the vision features, which is different from the concatenation fusion and element-wise add fusion. Moreover, the proposed SAF can be trained by an end-to-end manner incorporated with the recent deep learning object detection framework. In addition, we build a generation model, which converts radar points to radar images for neural network training. Numerical results suggest that the newly developed fusion method achieves superior performance in public benchmarking. In addition, the source code will be released in the GitHub.<\/jats:p>","DOI":"10.3390\/s20040956","type":"journal-article","created":{"date-parts":[[2020,2,11]],"date-time":"2020-02-11T11:45:30Z","timestamp":1581421530000},"page":"956","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":142,"title":["Spatial Attention Fusion for Obstacle Detection Using MmWave Radar and Vision Sensor"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5663-2793","authenticated-orcid":false,"given":"Shuo","family":"Chang","sequence":"first","affiliation":[{"name":"School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yifan","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fan","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Beijing Information Science and Technology University, Beijing 100101, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaotong","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sai","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhiyong","family":"Feng","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhiqing","family":"Wei","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Beijing University of Posts and Telecommunications, Beijing 100876, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2020,2,11]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., and Erhan, D. (2016, January 11\u201314). SSD: Single shot multibox detector. Proceedings of the European Conference on Computer Vision Workshops (ECCV 2016), Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR2014), Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision (ICCV2015), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"9:1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"6:1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Lin, T.L., Piotr, D., Ross, G., He, K.M., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR2017), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K.M., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision (ICCV2017), Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"134014","DOI":"10.1109\/ACCESS.2019.2941892","article-title":"Aggregated Residual Dilation-Based Feature Pyramid Network for Object Detection","volume":"7","author":"Zhao","year":"2019","journal-title":"IEEE Access"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Tian, Z., Shen, C.H., Chen, H., and He, T. (2019). FCOS: Fully Convolutional One-Stage Object Detection. arXiv.","DOI":"10.1109\/ICCV.2019.00972"},{"key":"ref_11","unstructured":"Simon, C., Will, M., and Paul, N. (2019, January 20\u201324). Distant Vehicle Detection Using Radar and Vision. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA2019), Montreal, QC, Canada."},{"key":"ref_12","unstructured":"Langer, D., and Jochem, T. (1996, January 19\u201320). Fusing radar and vision for detecting, classifying and avoiding roadway obstacles. Proceedings of the IEEE Conference on Intelligent Vehicles, Tokyo, Japan."},{"key":"ref_13","unstructured":"Cou\u00e9, C., Fraichard, T., Bessiere, P., and Mazer, E. (2002, January 17\u201321). Multi-sensor data fusion using Bayesian programming: An automotive application. Proceedings of the IEEE 2002 Intelligent Vehicles Symposium, Versailles, France."},{"key":"ref_14","unstructured":"Kawasaki, N., and Kiencke, U. (2004, January 14\u201317). Standard platform for sensor fusion on advanced driver assistance system using bayesian network. Proceedings of the IEEE 2004 Intelligent Vehicles Symposium, Parma, Italy."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"338","DOI":"10.1016\/j.robot.2016.05.001","article-title":"Radar and stereo vision fusion for multitarget tracking on the special Euclidean group","volume":"83","year":"2016","journal-title":"Robot. Auton. Syst."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Obrvan, M., \u0106esi\u0107, J., and Petrovi\u0107, I. (2015, January 19\u201321). Appearance based vehicle detection by radar-stereo vision integration. Proceedings of the Robot 2015: Second Iberian Robotics Conference, Lisbon, Portugal.","DOI":"10.1007\/978-3-319-27146-0_34"},{"key":"ref_17","first-page":"4:606","article-title":"Collision sensing by stereo vision and radar sensor fusion","volume":"10","author":"Wu","year":"2009","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Chavez-Garcia, R.O., Burlet, J., Vu, T.D., and Aycard, O. (2012, January 3\u20137). Frontal object perception using radar and mono-vision. Proceedings of the IEEE 2012 Intelligent Vehicles Symposium, Alcala de Henares, Spain.","DOI":"10.1109\/IVS.2012.6232307"},{"key":"ref_19","first-page":"17:258-1","article-title":"Camera radar fusion for increased reliability in ADAS applications","volume":"2018","author":"Zhong","year":"2018","journal-title":"Electron. Imaging"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"641","DOI":"10.1016\/j.ins.2014.03.080","article-title":"Data fusion of radar and image measurements for multi-object tracking via Kalman filtering","volume":"278","author":"Kim","year":"2014","journal-title":"Inf. Sci."},{"key":"ref_21","unstructured":"Steux, B., Laurgeau, C., Salesse, L., and Wautier, D. (2002, January 17\u201321). Fade: A vehicle detection and tracking system featuring monocular color vision and radar data fusion. Proceedings of the IEEE 2002 Intelligent Vehicles Symposium, Versailles, France."},{"key":"ref_22","unstructured":"Streubel, R., and Yang, B. (2016, January 5\u20138). Fusion of stereo camera and MIMO-FMCW radar for pedestrian tracking in indoor environments. Proceedings of the IEEE 19th International Conference on Information Fusion, Heidelberg, Germany."},{"key":"ref_23","unstructured":"Long, N.B., Wang, K.W., Cheng, R.Q., Yang, K.L., and Bai, J.S. (2018, January 10\u201313). Fusion of millimeter wave radar and RGB-depth sensors for assisted navigation of the visually impaired. Proceedings of the Millimetre Wave and Terahertz Sensors and Technology XI, Berlin, Germany."},{"key":"ref_24","first-page":"4:044102-1","article-title":"Unifying obstacle detection, recognition, and fusion based on millimeter wave radar and RGB-depth sensors for the visually impaired","volume":"90","author":"Long","year":"2019","journal-title":"Revi. Sci. Instrum."},{"key":"ref_25","unstructured":"Milch, S., and Behrens, M. (2001, January 25\u201326). Pedestrian detection with radar and computer vision. Proceedings of the 2001 PAL Symposium\u2014Progress in Automobile Lighting, Darmstadt, Germany."},{"key":"ref_26","unstructured":"Bombini, L., Cerri, P., Medici, P., and Alessandretti, G. (2006, January 14\u201315). Radar-vision fusion for vehicle detection. Proceedings of the 3rd International Workshop on Intelligent Transportation, Hamburg, Germany."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1:95","DOI":"10.1109\/TITS.2006.888597","article-title":"Vehicle and guard rail detection using radar and vision data fusion","volume":"8","author":"Alessandretti","year":"2007","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Kadow, U., Schneider, G., and Vukotich, A. (2007, January 13\u201315). Radar-vision based vehicle recognition with evolutionary optimized and boosted features. Proceedings of the IEEE 2007 Intelligent Vehicles Symposium, Istanbul, Turkey.","DOI":"10.1109\/IVS.2007.4290206"},{"key":"ref_29","unstructured":"Haselhoff, A., Kummert, A., and Schneider, G. (2007, January 3\u20137). Radar-vision fusion for vehicle detection by means of improved haar-like feature and adaboost approach. Proceedings of the IEEE 2007 15th European Signal Processing Conference, Poznan, Poland."},{"key":"ref_30","unstructured":"Ji, Z.P., and Prokhorov, D. (July, January 30). Radar-vision fusion for object classification. Proceedings of the IEEE 2008 11th International Conference on Information Fusion, Cologne, Germany."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Serfling, M., Loehlein, O., Schweiger, R., and Dietmayer, K. (2009, January 3\u20135). Camera and imaging radar feature level sensor fusion for night vision pedestrian recognition. Proceedings of the IEEE 2009 Intelligent Vehicles Symposium, Xi\u2019an, China.","DOI":"10.1109\/IVS.2009.5164345"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"3:182","DOI":"10.1109\/TITS.2002.802932","article-title":"An obstacle detection method by fusion of radar and motion stereo","volume":"3","author":"Kato","year":"2002","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"9:8992","DOI":"10.3390\/s110908992","article-title":"Integrating millimeter wave radar with a monocular vision sensor for on-road obstacle detection applications","volume":"11","author":"Wang","year":"2011","journal-title":"Sensors"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Guo, X.P., Du, J.S., Gao, J., and Wang, W. (2018, January 18\u201320). Pedestrian Detection Based on Fusion of Millimeter Wave Radar and Vision. Proceedings of the 2018 International Conference on Artificial Intelligence and Pattern Recognition, Beijing, China.","DOI":"10.1145\/3268866.3268868"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"John, V., and Mita, S. (2019, January 18\u201322). RVNet: Deep Sensor Fusion of Monocular Camera and Radar for Image-Based Obstacle Detection in Challenging Environments. Proceedings of the 2019 Pacific-Rim Symposium on Image and Video Technology, Sydney, Australia.","DOI":"10.1007\/978-3-030-34879-3_27"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Nobis, F., Geisslinger, M., Weber, M., Betz, J., and Lienkamp, M. (2019, January 15\u201317). A Deep Learning-based Radar and Camera Sensor Fusion Architecture for Object Detection. Proceedings of the 2019 Sensor Data Fusion: Trends, Solutions, Applications (SDF2019), Bonn, Germany.","DOI":"10.1109\/SDF.2019.8916629"},{"key":"ref_37","unstructured":"Holger, C., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. (2019). nuScenes: A multimodal dataset for autonomous driving. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision (ECCV2014), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_39","unstructured":"(2020, February 10). FCOS Model. Available online: https:\/\/github.com\/tianzhi0549\/FCOS."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"2:154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective search for object recognition","volume":"104","author":"Uijlings","year":"2013","journal-title":"Int. J. Comput. Vis."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zitnick, C.L., and Doll\u00e1r, P. (2014, January 6\u201312). Edge boxes: Locating object proposals from edges. Proceedings of the European Conference on Computer Vision Workshops (ECCV 2014), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_26"},{"key":"ref_42","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G.E. (2012, January 3\u20136). Imagenet classification with deep convolutional neural networks. Proceedings of the Advances in Neural Information Processing Systems (NIPS 2012), Lake Tahoe, NV, USA."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"2:303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The pascal visual object classes (voc) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis."},{"key":"ref_44","first-page":"3:1093","article-title":"Joint integrated probabilistic data association: JIPDA","volume":"40","author":"Musicki","year":"2004","journal-title":"IEEE Trans. Aerosp. Electro. Syst."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Sugimoto, S., Tateda, H., Takahashi, H., and Okutomi, M. (2004, January 26). Obstacle detection using millimeter-wave radar and its visualization on image sequence. Proceedings of the 17th International Conference on Pattern Recognition (ICPR2004), Cambridge, UK.","DOI":"10.1109\/ICPR.2004.1334537"},{"key":"ref_46","first-page":"8:790","article-title":"Mean shift, mode seeking, and clustering","volume":"17","author":"Cheng","year":"1995","journal-title":"IEEE Trans. Pattern Anal. Mach. Intel."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"He, K.M., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Cision (ICCV2017), Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_48","unstructured":"Weng, J.Y., and Zhang, N. (2006, January 16\u201321). Optimal in-place learning and the lobe component analysis. Proceedings of the IEEE International Joint Conference on Neural Network, Vancouver, BC, Canada."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"He, K.M., Zhang, X.Y., Ren, S.Q., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR2016), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_50","unstructured":"Fizyr (2020, February 10). Keras Retinanet. Available online: https:\/\/github:com\/fizyr\/keras-retinanet."},{"key":"ref_51","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_52","first-page":"1:19","article-title":"Three-Level Early Fusion for Road User Detection","volume":"1","author":"Lindl","year":"2006","journal-title":"PReVENT Fus. Forum e-J."},{"key":"ref_53","first-page":"2:525","article-title":"Multiple sensor fusion and classification for moving object detection and tracking","volume":"17","author":"Aycard","year":"2015","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"7:2075","DOI":"10.1109\/TITS.2016.2533542","article-title":"On-road vehicle detection and tracking using MMW radar and monovision fusion","volume":"17","author":"Wang","year":"2016","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Liu, X., Sun, Z.P., and He, H.G. (2011, January 10\u201312). On-road vehicle detection fusing radar and vision. Proceedings of the IEEE 2011 International Conference on Vehicular Electronics and Safety, Beijing, China.","DOI":"10.1109\/ICVES.2011.5983805"},{"key":"ref_56","doi-asserted-by":"crossref","unstructured":"Yu, J.H., Jiang, Y.N., Wang, Z.Y., Cao, Z.M., and Huang, T. (2016, January 15\u201319). Unitbox: An advanced object detection network. Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam, The Netherlands.","DOI":"10.1145\/2964284.2967274"},{"key":"ref_57","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., and Antiga, L. (2019, January 10\u201312). PyTorch: An imperative style, high-performance deep learning library. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_58","doi-asserted-by":"crossref","first-page":"3:211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"He, K.M., Zhang, X.Y., Ren, S.Q., and Sun, J. (2015, January 7\u201313). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. Proceedings of the IEEE International Conference on Computer Cision (ICCV2015), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.123"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/4\/956\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T08:56:41Z","timestamp":1760173001000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/4\/956"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,2,11]]},"references-count":59,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2020,2]]}},"alternative-id":["s20040956"],"URL":"https:\/\/doi.org\/10.3390\/s20040956","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,2,11]]}}}