{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T12:24:50Z","timestamp":1781612690427,"version":"3.54.5"},"reference-count":51,"publisher":"MDPI AG","issue":"20","license":[{"start":{"date-parts":[[2023,10,12]],"date-time":"2023-10-12T00:00:00Z","timestamp":1697068800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Key Project of National Nature Science Foundation of China","award":["61932012"],"award-info":[{"award-number":["61932012"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Multi-modal sensors are the key to ensuring the robust and accurate operation of autonomous driving systems, where LiDAR and cameras are important on-board sensors. However, current fusion methods face challenges due to inconsistent multi-sensor data representations and the misalignment of dynamic scenes. Specifically, current fusion methods either explicitly correlate multi-sensor data features by calibrating parameters, ignoring the feature blurring problems caused by misalignment, or find correlated features between multi-sensor data through global attention, causing rapidly escalating computational costs. On this basis, we propose a transformer-based end-to-end multi-sensor fusion framework named the adaptive fusion transformer (AFTR). The proposed AFTR consists of the adaptive spatial cross-attention (ASCA) mechanism and the spatial temporal self-attention (STSA) mechanism. Specifically, ASCA adaptively associates and interacts with multi-sensor data features in 3D space through learnable local attention, alleviating the problem of the misalignment of geometric information and reducing computational costs, and STSA interacts with cross-temporal information using learnable offsets in deformable attention, mitigating displacements due to dynamic scenes. We show through numerous experiments that the AFTR obtains SOTA performance in the nuScenes 3D object detection task (74.9% NDS and 73.2% mAP) and demonstrates strong robustness to misalignment (only a 0.2% NDS drop with slight noise). At the same time, we demonstrate the effectiveness of the AFTR components through ablation studies. In summary, the proposed AFTR is an accurate, efficient, and robust multi-sensor data fusion framework.<\/jats:p>","DOI":"10.3390\/s23208400","type":"journal-article","created":{"date-parts":[[2023,10,12]],"date-time":"2023-10-12T03:14:32Z","timestamp":1697080472000},"page":"8400","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["AFTR: A Robustness Multi-Sensor Fusion Model for 3D Object Detection Based on Adaptive Fusion Transformer"],"prefix":"10.3390","volume":"23","author":[{"given":"Yan","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Artificial Intelligence, China University of Mining and Technology-Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8413-123X","authenticated-orcid":false,"given":"Kang","family":"Liu","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, China University of Mining and Technology-Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hong","family":"Bao","sequence":"additional","affiliation":[{"name":"College of Robotics, Beijing Union University, Beijing 100027, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xu","family":"Qian","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, China University of Mining and Technology-Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zihan","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, China University of Mining and Technology-Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shiqing","family":"Ye","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, China University of Mining and Technology-Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weicen","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, China University of Mining and Technology-Beijing, Beijing 100083, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,10,12]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"eaav9843","DOI":"10.1126\/scirobotics.aav9843","article-title":"Self-Driving Cars: A City Perspective","volume":"4","author":"Duarte","year":"2019","journal-title":"Sci. Robot."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"3135","DOI":"10.1109\/TITS.2019.2926042","article-title":"Is It Safe to Drive? An Overview of Factors, Metrics, and Datasets for Driveability Assessment in Autonomous Driving","volume":"21","author":"Guo","year":"2020","journal-title":"IEEE Trans. Intell. Transport. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"E1","DOI":"10.1038\/s41586-020-1987-4","article-title":"Life and Death Decisions of Autonomous Vehicles","volume":"579","author":"Bigman","year":"2020","journal-title":"Nature"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Vora, S., Lang, A.H., Helou, B., and Beijbom, O. (2020, January 13\u201319). PointPainting: Sequential Fusion for 3D Object Detection. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00466"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Liu, Z., Tang, H., Amini, A., Yang, X., Mao, H., Rus, D.L., and Han, S. (June, January 29). BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird\u2019s-Eye View Representation. Proceedings of the 2023 IEEE International Conference on Robotics and Automation (ICRA), London, UK.","DOI":"10.1109\/ICRA48891.2023.10160968"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Gao, X., Wang, Z., Feng, Y., Ma, L., Chen, Z., and Xu, B. (2023). Benchmarking Robustness of AI-Enabled Multi-Sensor Fusion Systems: Challenges and Opportunities. arXiv.","DOI":"10.1145\/3611643.3616278"},{"key":"ref_7","first-page":"180","article-title":"DETR3D: 3D Object Detection from Multi-View Images via 3D-to-2D Queries","volume":"Volume 164","author":"Wang","year":"2022","journal-title":"Proceedings of the 5th Conference on Robot Learning"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.-M. (2020, January 23\u201328). End-to-End Object Detection with Transformers. Proceedings of the Computer Vision\u2014ECCV, Glasgow, UK.","DOI":"10.1007\/978-3-030-58598-3"},{"key":"ref_9","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention Is All You Need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Avidan, S., Brostow, G., Ciss\u00e9, M., Farinella, G.M., and Hassner, T. (2022, January 23\u201327). BEVFormer: Learning Bird\u2019s-Eye-View Representation from Multi-Camera Images via Spatiotemporal Transformers. Proceedings of the Computer Vision\u2014ECCV, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19818-2"},{"key":"ref_11","first-page":"10421","article-title":"BEVFusion: A Simple and Robust LiDAR-Camera Fusion Framework","volume":"35","author":"Liang","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Vedaldi, A., Bischof, H., Brox, T., and Frahm, J.-M. (2020, January 23\u201328). Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D. Proceedings of the Computer Vision\u2014ECCV 2020, Glasgow, UK.","DOI":"10.1007\/978-3-030-58598-3"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Yan, J., Liu, Y., Sun, J., Jia, F., Li, S., Wang, T., and Zhang, X. (2023). Cross Modal Transformer: Towards Fast and Robust 3D Object Detection. arXiv.","DOI":"10.1109\/ICCV51070.2023.01675"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Li, Y., Yu, A.W., Meng, T., Caine, B., Ngiam, J., Peng, D., Shen, J., Lu, Y., Zhou, D., and Le, Q.V. (2022, January 21). DeepFusion: Lidar-Camera Deep Fusion for Multi-Modal 3D Object Detection. Proceedings of the CVPR, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01667"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chen, X., Zhang, T., Wang, Y., Wang, Y., and Zhao, H. (2023, January 17\u201324). FUTR3D: A Unified Sensor Fusion Framework for 3D Detection. Proceedings of the 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Vancouver, BC, Canada.","DOI":"10.1109\/CVPRW59228.2023.00022"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wang, T., Zhu, X., Pang, J., and Lin, D. (2021). FCOS3D: Fully Convolutional One-Stage Monocular 3D Object Detection. arXiv.","DOI":"10.1109\/ICCVW54120.2021.00107"},{"key":"ref_17","unstructured":"Tian, Z., Shen, C., Chen, H., and He, T. (November, January 27). FCOS: Fully Convolutional One-Stage Object Detection. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_18","unstructured":"Luo, Z., Zhou, C., Zhang, G., and Lu, S. (2022). DETR4D: Direct Multi-View 3D Object Detection with Sparse Attention. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Avidan, S., Brostow, G., Ciss\u00e9, M., Farinella, G.M., and Hassner, T. (2022, January 23\u201327). PETR: Position Embedding Transformation for Multi-View 3D Object Detection. Proceedings of the Computer Vision\u2014ECCV 2022, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19836-6"},{"key":"ref_20","first-page":"1042","article-title":"PolarFormer: Multi-Camera 3D Object Detection with Polar Transformer","volume":"37","author":"Jiang","year":"2023","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_21","unstructured":"Huang, J., Huang, G., Zhu, Z., Ye, Y., and Du, D. (2022). BEVDet: High-Performance Multi-Camera 3D Object Detection in Bird-Eye-View. arXiv."},{"key":"ref_22","unstructured":"Huang, J., and Huang, G. (2022). BEVDet4D: Exploit Temporal Cues in Multi-Camera 3D Object Detection. arXiv."},{"key":"ref_23","first-page":"1477","article-title":"BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object Detection","volume":"37","author":"Li","year":"2023","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_24","first-page":"1486","article-title":"BEVStereo: Enhancing Depth Estimation in Multi-View 3D Object Detection with Temporal Stereo","volume":"37","author":"Li","year":"2023","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Yang, C., Chen, Y., Tian, H., Tao, C., Zhu, X., Zhang, Z., Huang, G., Li, H., Qiao, Y., and Lu, L. (2023, January 18\u201322). BEVFormer v2: Adapting Modern Image Backbones to Bird\u2019s-Eye-View Recognition via Perspective Supervision. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01710"},{"key":"ref_26","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., and Dai, J. (2020, January 22\u201324). Deformable DETR: Deformable Transformers for End-to-End Object Detection. Proceedings of the International Conference on Learning Representations, Beijing, China."},{"key":"ref_27","unstructured":"Qi, C.R., Su, H., Mo, K., and Guibas, L.J. (2017, January 21\u201326). PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA."},{"key":"ref_28","unstructured":"Qi, C.R., Yi, L., Su, H., and Guibas, L.J. (2017, January 4). PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space. Proceedings of the Advances in Neural Information Processing Systems, Red Hook, NY, USA."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhou, Y., and Tuzel, O. (2018, January 18\u201323). VoxelNet: End-to-End Learning for Point Cloud Based 3D object detection. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00472"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yan, Y., Mao, Y., and Li, B. (2018). SECOND: Sparsely Embedded Convolutional Detection. Sensors, 18.","DOI":"10.3390\/s18103337"},{"key":"ref_31","unstructured":"Graham, B., Engelcke, M., and van der Maaten, L. 3D Semantic Segmentation With Submanifold Sparse Convolutional Networks. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Lang, A.H., Vora, S., Caesar, H., Zhou, L., Yang, J., and Beijbom, O. (2019, January 15\u201320). PointPillars: Fast Encoders for object detection From Point Clouds. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01298"},{"key":"ref_33","unstructured":"Shi, S., Guo, C., Yang, J., and Li, H. (2020). PV-RCNN: The Top-Performing LiDAR-Only Solutions for 3D Detection \/ 3D Tracking \/ Domain Adaptation of Waymo Open Dataset Challenges. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Qi, C.R., Liu, W., Wu, C., Su, H., and Guibas, L.J. (2018, January 18\u201323). Frustum PointNets for 3D object detection From RGB-D Data. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00102"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Xu, S., Zhou, D., Fang, J., Yin, J., Bin, Z., and Zhang, L. (2021, January 19\u201322). FusionPainting: Multi-sensor Fusion with Adaptive Attention for 3D Object Detection. Proceedings of the 2021 IEEE International Intelligent Transportation Systems Conference (ITSC), Indianapolis, IN, USA.","DOI":"10.1109\/ITSC48978.2021.9564951"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Sindagi, V.A., Zhou, Y., and Tuzel, O. (2019, January 20\u201324). MVX-Net: Multi-sensor VoxelNet for 3D object detection. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8794195"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Bai, X., Hu, Z., Zhu, X., Huang, Q., Chen, Y., Fu, H., and Tai, C.-L. (2022, January 18\u201324). TransFusion: Robust LiDAR-Camera Fusion for 3D Object Detection with Transformers. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00116"},{"key":"ref_38","first-page":"1992","article-title":"DeepInteraction: 3D Object Detection via Modality Interaction","volume":"Volume 35","author":"Koyejo","year":"2022","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_39","first-page":"18442","article-title":"Unifying Voxel-Based Representation with Transformer for 3D Object Detection","volume":"Volume 35","author":"Koyejo","year":"2022","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Park, D., Ambrus, R., Guizilini, V., Li, J., and Gaidon, A. (2021, January 10\u201317). Is Pseudo-Lidar Needed for Monocular 3D Object Detection?. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00313"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., and Wei, Y. (2017, January 22\u201329). Deformable Convolutional Networks. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.89"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"108796","DOI":"10.1016\/j.patcog.2022.108796","article-title":"3D Object Detection for Autonomous Driving: A Survey","volume":"130","author":"Qian","year":"2022","journal-title":"Pattern Recognit."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollar, P. (2017, January 22\u201329). Focal Loss for Dense Object Detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Caesar, H., Bankiti, V., Lang, A.H., Vora, S., Liong, V.E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O. (2020, January 13\u201319). NuScenes: A Multi-sensor Dataset for Autonomous Driving. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01164"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The Pascal Visual Object Classes (VOC) Challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"Int. J. Comput. Vis"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are We Ready for Autonomous Driving? The KITTI Vision Benchmark Suite. In Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_49","unstructured":"Loshchilov, I., and Hutter, F. (2019). Decoupled Weight Decay Regularization. arXiv."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Peng, C., Xiao, T., Li, Z., Jiang, Y., Zhang, X., Jia, K., Yu, G., and Sun, J. (2018, January 18\u201323). MegDet: A Large Mini-Batch Object Detector. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00647"},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Yin, T., Zhou, X., and Krahenbuhl, P. Center-Based 3D Object Detection and Tracking. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR46437.2021.01161"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/20\/8400\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:05:20Z","timestamp":1760130320000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/20\/8400"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,12]]},"references-count":51,"journal-issue":{"issue":"20","published-online":{"date-parts":[[2023,10]]}},"alternative-id":["s23208400"],"URL":"https:\/\/doi.org\/10.3390\/s23208400","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,10,12]]}}}