{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T11:42:01Z","timestamp":1782301321842,"version":"3.54.5"},"reference-count":44,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,5,1]],"date-time":"2024-05-01T00:00:00Z","timestamp":1714521600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["6227010741"],"award-info":[{"award-number":["6227010741"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100005046","name":"Natural Science Foundation of Heilongjiang Province","doi-asserted-by":"publisher","award":["LH2022E114"],"award-info":[{"award-number":["LH2022E114"]}],"id":[{"id":"10.13039\/501100005046","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:p>The object detection method serves as the core technology within the unmanned driving perception module, extensively employed for detecting vehicles, pedestrians, traffic signs, and various objects. However, existing object detection methods still encounter three challenges in intricate unmanned driving scenarios: unsatisfactory performance in multi-scale object detection, inadequate accuracy in detecting small objects, and occurrences of false positives and missed detections in densely occluded environments. Therefore, this study proposes an improved object detection method for unmanned driving, leveraging Transformer architecture to address these challenges. First, a multi-scale Transformer feature extraction method integrated with channel attention is used to enhance the network's capability in extracting features across different scales. Second, a training method incorporating Query Denoising with Gaussian decay was employed to enhance the network's proficiency in learning representations of small objects. Third, a hybrid matching method combining Optimal Transport and Hungarian algorithms was used to facilitate the matching process between predicted and actual values, thereby enriching the network with more informative positive sample features. Experimental evaluations conducted on datasets including KITTI demonstrate that the proposed method achieves 3% higher mean Average Precision (mAP) than that of the existing methodologies.<\/jats:p>","DOI":"10.3389\/fnbot.2024.1342126","type":"journal-article","created":{"date-parts":[[2024,5,1]],"date-time":"2024-05-01T05:07:23Z","timestamp":1714540043000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Improved object detection method for unmanned driving based on Transformers"],"prefix":"10.3389","volume":"18","author":[{"given":"Huaqi","family":"Zhao","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiang","family":"Peng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Su","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun-Bao","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jeng-Shyang","family":"Pan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaoguang","family":"Su","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaomin","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2024,5,1]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv:2012.09958","article-title":"Toward transformer-based object detection","author":"Beal","year":"2020","journal-title":"arXiv"},{"key":"B2","first-page":"213","article-title":"\u201cEnd-to-end object detection with transformers,\u201d","volume-title":"European conference on computer vision","author":"Carion","year":"2020"},{"key":"B3","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1007\/BF00994018","article-title":"Support-vector networks","volume":"20","author":"Cortes","year":"1995","journal-title":"Mach. Learn"},{"key":"B4","first-page":"886","article-title":"\u201cHistograms of oriented gradients for human detection,\u201d","volume-title":"2005 IEEE computer society conference on computer vision and pattern recognition (CVPR'05), Vol. 1","author":"Dalal","year":"2005"},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.11929","article-title":"An image is worth 16x16 words: transformers for image recognition at scale","author":"Dosovitskiy","year":"2020","journal-title":"arXiv"},{"key":"B6","first-page":"303","article-title":"\u201cOta: optimal transport assignment for object detection,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ge","year":""},{"key":"B7","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2107.08430","article-title":"Yolox: exceeding yolo series in 2021","author":"Ge","year":"","journal-title":"arXiv"},{"key":"B8","doi-asserted-by":"publisher","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: the kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Rob. Res"},{"key":"B9","first-page":"580","article-title":"\u201cRich feature hierarchies for accurate object detection and semantic segmentation,\u201d","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Girshick","year":"2014"},{"key":"B10","first-page":"7132","article-title":"\u201cSqueeze-and-excitation networks,\u201d","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Hu","year":"2018"},{"key":"B11","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01780","article-title":"\u201cLite DETR: an interleaved multi-scale encoder for efficient DETR,\u201d","author":"Li","year":"2023","journal-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition"},{"key":"B12","first-page":"13619","article-title":"\u201cDN-DETR: accelerate detr training by introducing query denoising,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li","year":"2022"},{"key":"B13","doi-asserted-by":"publisher","first-page":"603","DOI":"10.1109\/TIV.2022.3165353","article-title":"Cross-domain object detection for autonomous driving: a stepwise domain adaptative YOLO approach","volume":"7","author":"Li","year":"2022","journal-title":"IEEE Trans. Intell. Veh."},{"key":"B14","first-page":"510","article-title":"\u201cSelective kernel networks,\u201d","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Li","year":"2019"},{"key":"B15","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.02068","article-title":"\u201cDPM-OT: a new diffusion probabilistic model based on optimal transport,\u201d","author":"Li","year":"2023","journal-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision"},{"key":"B16","first-page":"2980","article-title":"\u201cFocal loss for dense object detection,\u201d","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Lin","year":"2017"},{"key":"B17","doi-asserted-by":"crossref","first-page":"740","DOI":"10.1007\/978-3-319-10602-1_48","article-title":"\u201cMicrosoft coco: common objects in context,\u201d","volume-title":"Computer Vision-ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13","author":"Lin","year":"2014"},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2306.13643","article-title":"Lightglue: local feature matching at light speed","author":"Lindenberger","year":"2023","journal-title":"arXiv"},{"key":"B19","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2201.12329","article-title":"DAB-DETR: dynamic anchor boxes are better queries for detr","author":"Liu","year":"2022","journal-title":"arXiv"},{"key":"B20","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1007\/978-3-319-46448-0_2","article-title":"\u201cSSD: single shot multibox detector,\u201d","volume-title":"Computer Vision-ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part I 14","author":"Liu","year":"2016"},{"key":"B21","first-page":"4463","article-title":"\u201cSemantic correspondence as an optimal transport problem,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu","year":"2020"},{"key":"B22","first-page":"12009","article-title":"\u201cSwin transformer v2: scaling up capacity and resolution,\u201d","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Liu","year":"2022"},{"key":"B23","first-page":"10012","article-title":"\u201cSwin transformer: hierarchical vision transformer using shifted windows,\u201d","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Liu","year":"2021"},{"key":"B24","first-page":"1","article-title":"\u201cEfficient multi-scale attention module with cross-spatial learning,\u201d","volume-title":"ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Ouyang","year":"2023"},{"key":"B25","first-page":"783","article-title":"\u201cFcanet: frequency channel attention networks,\u201d","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Qin","year":"2021"},{"key":"B26","first-page":"779","article-title":"\u201cYou only look once: unified, real-time object detection,\u201d","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Redmon","year":"2016"},{"key":"B27","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: towards real-time object detection with region proposal networks","volume":"28","author":"Ren","year":"2015","journal-title":"Adv. Neural Inf. Process. Syst"},{"key":"B28","first-page":"658","article-title":"\u201cGeneralized intersection over union: a metric and a loss for bounding box regression,\u201d","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Rezatofighi","year":"2019"},{"key":"B29","first-page":"4938","article-title":"\u201cSuperglue: learning feature matching with graph neural networks,\u201d","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Sarlin","year":"2020"},{"key":"B30","first-page":"2446","article-title":"\u201cScalability in perception for autonomous driving: Waymo open dataset,\u201d","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Sun","year":"2020"},{"key":"B31","first-page":"14454","article-title":"\u201cSparse R-CNN: end-to-end object detection with learnable proposals,\u201d","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Sun","year":"2021"},{"key":"B32","first-page":"7464","article-title":"\u201cYOLOV7: trainable bag-of-freebies sets new state-of-the-art for real-time object detectors,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang","year":"2023"},{"key":"B33","first-page":"11534","article-title":"\u201cECA-NET: efficient channel attention for deep convolutional neural networks,\u201d","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Wang","year":"2020"},{"key":"B34","doi-asserted-by":"publisher","first-page":"1626","DOI":"10.13229\/j.cnki.jdxbgxb20210652","article-title":"Transfer learning of medical image segmentation based on optimal transport feature selection","volume":"52","author":"Wang","year":"2022","journal-title":"Jilin Daxue Xuebao"},{"key":"B35","first-page":"568","article-title":"\u201cPyramid vision transformer: a versatile backbone for dense prediction without convolutions,\u201d","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Wang","year":"2021"},{"key":"B36","doi-asserted-by":"publisher","first-page":"415","DOI":"10.1007\/s41095-022-0274-8","article-title":"PVT V2: improved baselines with pyramid vision transformer","volume":"8","author":"Wang","year":"2022","journal-title":"Comput. Vis. Media"},{"key":"B37","first-page":"2567","article-title":"\u201cAnchor DETR: query design for transformer-based detector,\u201d","volume-title":"Proceedings of the AAAI conference on artificial intelligence, Vol. 36","author":"Wang","year":"2022"},{"key":"B38","first-page":"11794","article-title":"\u201cGated channel transformation for visual recognition,\u201d","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Yang","year":"2020"},{"key":"B39","doi-asserted-by":"crossref","first-page":"164","DOI":"10.1109\/GECOST55694.2022.10010490","article-title":"\u201cSafety helmet detection using deep learning: implementation and comparative study using YOLOV5, YOLOV6, and YOLOV7,\u201d","volume-title":"2022 International Conference on Green Energy, Computing and Sustainable Technology (GECOST)","author":"Yung","year":"2022"},{"key":"B40","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1007\/978-3-030-58555-6_16","article-title":"\u201cDynamic R-CNN: towards high quality object detection via dynamic training,\u201d","volume-title":"Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XV 16","author":"Zhang","year":"2020"},{"key":"B41","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2203.03605","article-title":"Dino: DETR with improved denoising anchor boxes for end-to-end object detection","author":"Zhang","year":"2022","journal-title":"arXiv"},{"key":"B42","doi-asserted-by":"publisher","first-page":"380","DOI":"10.1109\/TMM.2019.2929005","article-title":"Widerperson: a diverse dataset for dense pedestrian detection in the wild","volume":"22","author":"Zhang","year":"2019","journal-title":"IEEE Trans. Multimed"},{"key":"B43","first-page":"10323","article-title":"\u201cBiformer: vision transformer with bi-level routing attention,\u201d","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhu","year":"2023"},{"key":"B44","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2010.04159","article-title":"Deformable detr: deformable transformers for end-to-end object detection","author":"Zhu","year":"2020","journal-title":"arXiv"}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2024.1342126\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,5,1]],"date-time":"2024-05-01T05:07:40Z","timestamp":1714540060000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2024.1342126\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,1]]},"references-count":44,"alternative-id":["10.3389\/fnbot.2024.1342126"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2024.1342126","relation":{},"ISSN":["1662-5218"],"issn-type":[{"value":"1662-5218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,1]]},"article-number":"1342126"}}