{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T00:36:11Z","timestamp":1768264571506,"version":"3.49.0"},"reference-count":54,"publisher":"MDPI AG","issue":"13","license":[{"start":{"date-parts":[[2022,6,26]],"date-time":"2022-06-26T00:00:00Z","timestamp":1656201600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["61802019"],"award-info":[{"award-number":["61802019"]}]},{"name":"National Natural Science Foundation of China","award":["61932012"],"award-info":[{"award-number":["61932012"]}]},{"name":"National Natural Science Foundation of China","award":["61871039"],"award-info":[{"award-number":["61871039"]}]},{"name":"National Natural Science Foundation of China","award":["61906017"],"award-info":[{"award-number":["61906017"]}]},{"name":"National Natural Science Foundation of China","award":["62006020"],"award-info":[{"award-number":["62006020"]}]},{"name":"National Natural Science Foundation of China","award":["KM201911417003"],"award-info":[{"award-number":["KM201911417003"]}]},{"name":"National Natural Science Foundation of China","award":["KM201911417009"],"award-info":[{"award-number":["KM201911417009"]}]},{"name":"National Natural Science Foundation of China","award":["KM201911417001"],"award-info":[{"award-number":["KM201911417001"]}]},{"name":"Beijing Municipal Education Commission Science and Technology Program","award":["61802019"],"award-info":[{"award-number":["61802019"]}]},{"name":"Beijing Municipal Education Commission Science and Technology Program","award":["61932012"],"award-info":[{"award-number":["61932012"]}]},{"name":"Beijing Municipal Education Commission Science and Technology Program","award":["61871039"],"award-info":[{"award-number":["61871039"]}]},{"name":"Beijing Municipal Education Commission Science and Technology Program","award":["61906017"],"award-info":[{"award-number":["61906017"]}]},{"name":"Beijing Municipal Education Commission Science and Technology Program","award":["62006020"],"award-info":[{"award-number":["62006020"]}]},{"name":"Beijing Municipal Education Commission Science and Technology Program","award":["KM201911417003"],"award-info":[{"award-number":["KM201911417003"]}]},{"name":"Beijing Municipal Education Commission Science and Technology Program","award":["KM201911417009"],"award-info":[{"award-number":["KM201911417009"]}]},{"name":"Beijing Municipal Education Commission Science and Technology Program","award":["KM201911417001"],"award-info":[{"award-number":["KM201911417001"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Object detection plays a vital role in autonomous driving systems, and the accurate detection of surrounding objects can ensure the safe driving of vehicles. This paper proposes a category-assisted transformer object detector called DetectFormer for autonomous driving. The proposed object detector can achieve better accuracy compared with the baseline. Specifically, ClassDecoder is assisted by proposal categories and global information from the Global Extract Encoder (GEE) to improve the category sensitivity and detection performance. This fits the distribution of object categories in specific scene backgrounds and the connection between objects and the image context. Data augmentation is used to improve robustness and attention mechanism added in backbone network to extract channel-wise spatial features and direction information. The results obtained by benchmark experiment reveal that the proposed method can achieve higher real-time detection performance in traffic scenes compared with RetinaNet and FCOS. The proposed method achieved a detection performance of 97.6% and 91.4% in AP50 and AP75 on the BCTSDB dataset, respectively.<\/jats:p>","DOI":"10.3390\/s22134833","type":"journal-article","created":{"date-parts":[[2022,6,26]],"date-time":"2022-06-26T22:50:23Z","timestamp":1656283823000},"page":"4833","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":31,"title":["DetectFormer: Category-Assisted Transformer for Traffic Scene Object Detection"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6062-6166","authenticated-orcid":false,"given":"Tianjiao","family":"Liang","sequence":"first","affiliation":[{"name":"Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing 100101, China"},{"name":"College of Robotics, Beijing Union University, Beijing 100101, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hong","family":"Bao","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing 100101, China"},{"name":"College of Robotics, Beijing Union University, Beijing 100101, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2293-1004","authenticated-orcid":false,"given":"Weiguo","family":"Pan","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing 100101, China"},{"name":"College of Robotics, Beijing Union University, Beijing 100101, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyue","family":"Fan","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing 100101, China"},{"name":"College of Robotics, Beijing Union University, Beijing 100101, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Han","family":"Li","sequence":"additional","affiliation":[{"name":"Beijing Key Laboratory of Information Service Engineering, Beijing Union University, Beijing 100101, China"},{"name":"College of Robotics, Beijing Union University, Beijing 100101, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,6,26]]},"reference":[{"key":"ref_1","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_2","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv."},{"key":"ref_3","first-page":"213","article-title":"End-to-End Object Detection with Transformers","volume":"Volume 12346","author":"Vedaldi","year":"2020","journal-title":"Computer Vision\u2014ECCV 2020"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"886","DOI":"10.1109\/CVPR.2005.177","article-title":"Histograms of Oriented Gradients for Human Detection","volume":"Volume 1","author":"Dalal","year":"2005","journal-title":"Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905)"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1627","DOI":"10.1109\/TPAMI.2009.167","article-title":"Object Detection with Discriminatively Trained Part-Based Models","volume":"32","author":"Felzenszwalb","year":"2010","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/5254.708428","article-title":"Support vector machines","volume":"13","author":"Hearst","year":"1998","journal-title":"IEEE Intell. Syst. Their Appl."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chen, Z., Shi, Q., and Huang, X. (July, January 28). Automatic detection of traffic lights using support vector machine. Proceedings of the 2015 IEEE Intelligent Vehicles Symposium (IV), Seoul, Korea.","DOI":"10.1109\/IVS.2015.7225659"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Cai, Z., and Vasconcelos, N. (2018, January 18\u201323). Cascade R-CNN: Delving Into High Quality Object Detection. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00644"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You Only Look Once: Unified, Real-Time Object Detection. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, Faster, Stronger. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_14","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An Incremental Improvement. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"318","DOI":"10.1109\/TPAMI.2018.2858826","article-title":"Focal Loss for Dense Object Detection","volume":"42","author":"Lin","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_16","first-page":"1922","article-title":"FCOS: A Simple and Strong Anchor-free Object Detector","volume":"44","author":"Tian","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","unstructured":"Zhou, X., Wang, D., and Kr\u00e4henb\u00fchl, P. (2019). Objects as Points. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"642","DOI":"10.1007\/s11263-019-01204-1","article-title":"CornerNet: Detecting Objects as Paired Keypoints","volume":"128","author":"Law","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"43253","DOI":"10.1109\/ACCESS.2021.3059052","article-title":"Automatic Recognition of Traffic Signs Based on Visual Inspection","volume":"9","author":"He","year":"2021","journal-title":"IEEE Access"},{"key":"ref_20","unstructured":"Sabour, S., Frosst, N., and Hinton, G.E. (2017, January 4\u20139). Dynamic Routing Between Capsules. Proceedings of the NIPS, Long Beach, CA, USA."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1016\/j.patrec.2021.02.003","article-title":"A method of cross-layer fusion multi-object detection and recognition based on improved faster R-CNN model in complex traffic environment","volume":"145","author":"Li","year":"2021","journal-title":"Pattern Recognit. Lett."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Lian, J., Yin, Y., Li, L., Wang, Z., and Zhou, Y. (2021). Small Object Detection in Traffic Scenes Based on Attention Feature Fusion. Sensors, 21.","DOI":"10.3390\/s21093031"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"40701","DOI":"10.1109\/ACCESS.2022.3166923","article-title":"ALODAD: An Anchor-Free Lightweight Object Detector for Autonomous Driving","volume":"10","author":"Liang","year":"2022","journal-title":"IEEE Access"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1007\/978-3-642-24797-2_4","article-title":"Long Short-Term Memory","volume":"Volume 385","author":"Graves","year":"2012","journal-title":"Supervised Sequence Labelling with Recurrent Neural Networks"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Cho, K., van Merrienboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014). Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_26","unstructured":"Zaremba, W., Sutskever, I., and Vinyals, O. (2014). Recurrent Neural Network Regularization. arXiv."},{"key":"ref_27","unstructured":"Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. arXiv."},{"key":"ref_28","unstructured":"Wu, Y., Schuster, M., Chen, Z., Le, Q.V., Norouzi, M., Macherey, W., Krikun, M., Cao, Y., Gao, Q., and Macherey, K. (2016). Google\u2019s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wang, H., Wu, Z., Liu, Z., Cai, H., Zhu, L., Gan, C., and Han, S. (2020, January 5\u201310). HAT: Hardware-Aware Transformers for Efficient Natural Language Processing. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online.","DOI":"10.18653\/v1\/2020.acl-main.686"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"681","DOI":"10.1007\/s11023-020-09548-1","article-title":"GPT-3: Its Nature, Scope, Limits, and Consequences","volume":"30","author":"Floridi","year":"2020","journal-title":"Minds Mach."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Dong, L., Xu, S., and Xu, B. (2018, January 5\u201320). Speech-Transformer: A No-Recurrence Sequence-to-Sequence Model for Speech Recognition. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8462506"},{"key":"ref_32","unstructured":"Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2016). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv."},{"key":"ref_33","unstructured":"Yan, H., Ma, X., and Pu, Z. (2021). Learning Dynamic and Hierarchical Traffic Spatiotemporal Features With Transformer. IEEE Trans. Intell. Transport. Syst., 1\u201314."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"736","DOI":"10.1111\/tgis.12644","article-title":"Traffic transformer: Capturing the continuity and periodicity of time series for traffic forecasting","volume":"24","author":"Cai","year":"2020","journal-title":"Trans. GIS"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-030-01234-2_1","article-title":"CBAM: Convolutional Block Attention Module","volume":"Volume 11211","author":"Ferrari","year":"2018","journal-title":"Computer Vision\u2014ECCV 2018"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Hou, Q., Zhou, D., and Feng, J. (2021, January 20\u201325). Coordinate Attention for Efficient Mobile Network Design. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01350"},{"key":"ref_37","unstructured":"DeVries, T., and Taylor, G.W. (2017). Improved Regularization of Convolutional Neural Networks with Cutout. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Yun, S., Han, D., Chun, S., Oh, S.J., Yoo, Y., and Choe, J. (November, January 27). CutMix: Regularization Strategy to Train Strong Classifiers With Localizable Features. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00612"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"3825532","DOI":"10.1155\/2022\/3825532","article-title":"Traffic Sign Detection via Improved Sparse R-CNN for Autonomous Vehicles","volume":"2022","author":"Liang","year":"2022","journal-title":"J. Adv. Transp."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The KITTI dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"740","DOI":"10.1007\/978-3-319-10602-1_48","article-title":"Microsoft COCO: Common Objects in Context","volume":"Volume 8693","author":"Fleet","year":"2014","journal-title":"Computer Vision\u2014ECCV 2014"},{"key":"ref_42","unstructured":"Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., and Xu, J. (2019). MMDetection: Open MMLab Detection Toolbox and Benchmark. arXiv."},{"key":"ref_43","unstructured":"Loshchilov, I., and Hutter, F. (2019). Decoupled Weight Decay Regularization. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/3065386","article-title":"ImageNet classification with deep convolutional neural networks","volume":"60","author":"Krizhevsky","year":"2017","journal-title":"Commun. ACM"},{"key":"ref_45","unstructured":"Glorot, X., and Bengio, Y. (2010, January 13\u201315). Understanding the difficulty of training deep feedforward neural networks. Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, Sardinia, Italy. JMLR Workshop and Conference Proceedings."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature Pyramid Networks for Object Detection. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Koonce, B. (2021). MobileNetV3. Convolutional Neural Networks with Swift for Tensorflow, Apress.","DOI":"10.1007\/978-1-4842-6168-2"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"2071","DOI":"10.1109\/TPAMI.2015.2389830","article-title":"Regionlets for Generic Object Detection","volume":"37","author":"Wang","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Chen, X., Kundu, K., Zhang, Z., Ma, H., Fidler, S., and Urtasun, R. (2016, January 27\u201330). Monocular 3D Object Detection for Autonomous Driving. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.236"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1007\/978-3-319-46493-0_22","article-title":"A Unified Multi-scale Deep Convolutional Neural Network for Fast Object Detection","volume":"Volume 9908","author":"Leibe","year":"2016","journal-title":"Computer Vision\u2014ECCV 2016"},{"key":"ref_52","doi-asserted-by":"crossref","first-page":"21","DOI":"10.1007\/978-3-319-46448-0_2","article-title":"SSD: Single Shot MultiBox Detector","volume":"Volume 9905","author":"Leibe","year":"2016","journal-title":"Computer Vision\u2014ECCV 2016"},{"key":"ref_53","doi-asserted-by":"crossref","first-page":"102827","DOI":"10.1016\/j.cviu.2019.102827","article-title":"ASSD: Attentive single shot multibox detector","volume":"189","author":"Yi","year":"2019","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"404","DOI":"10.1007\/978-3-030-01252-6_24","article-title":"Receptive Field Block Net for Accurate and Fast Object Detection","volume":"Volume 11215","author":"Ferrari","year":"2018","journal-title":"Computer Vision\u2014ECCV 2018"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/13\/4833\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:38:44Z","timestamp":1760139524000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/13\/4833"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,26]]},"references-count":54,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2022,7]]}},"alternative-id":["s22134833"],"URL":"https:\/\/doi.org\/10.3390\/s22134833","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,26]]}}}