{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T17:28:11Z","timestamp":1777656491799,"version":"3.51.4"},"reference-count":39,"publisher":"MDPI AG","issue":"21","license":[{"start":{"date-parts":[[2022,10,22]],"date-time":"2022-10-22T00:00:00Z","timestamp":1666396800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62072378"],"award-info":[{"award-number":["62072378"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent years, intelligent driving technology based on vehicle\u2013road cooperation has gradually become a research hotspot in the field of intelligent transportation. There are many studies regarding vehicle perception, but fewer studies regarding roadside perception. As sensors are installed at different heights, the roadside object scale varies violently, which burdens the optimization of networks. Moreover, there is a large amount of overlapping and occlusion in complex road environments, which leads to a great challenge of object distinction. To solve the two problems raised above, we propose RD-YOLO. Based on YOLOv5s, we reconstructed the feature fusion layer to increase effective feature extraction and improve the detection capability of small targets. Then, we replaced the original pyramid network with a generalized feature pyramid network (GFPN) to improve the adaptability of the network to different scale features. We also integrated a coordinate attention (CA) mechanism to find attention regions in scenarios with dense objects. Finally, we replaced the original Loss with Focal-EIOU Loss to improve the speed of the bounding box regression and the positioning accuracy of the anchor box. Compared to the YOLOv5s, the RD-YOLO improves the mean average precision (mAP) by 5.5% on the Rope3D dataset and 2.9% on the UA-DETRAC dataset. Meanwhile, by modifying the feature fusion layer, the weight of RD-YOLO is decreased by 55.9% while the detection speed is almost unchanged. Nevertheless, the proposed algorithm is capable of real-time detection at faster than 71.9 frames\/s (FPS) and achieves higher accuracy than the previous approaches with a similar FPS.<\/jats:p>","DOI":"10.3390\/s22218097","type":"journal-article","created":{"date-parts":[[2022,10,24]],"date-time":"2022-10-24T10:09:23Z","timestamp":1666606163000},"page":"8097","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":37,"title":["RD-YOLO: An Effective and Efficient Object Detector for Roadside Perception System"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7060-7391","authenticated-orcid":false,"given":"Lei","family":"Huang","sequence":"first","affiliation":[{"name":"School of Electronic Information, Xijing University, Xi\u2019an 710123, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenzhun","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Electronic Information, Xijing University, Xi\u2019an 710123, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,10,22]]},"reference":[{"key":"ref_1","first-page":"28177","article-title":"A review on road traffic accident and related factors","volume":"10","author":"Muthusamy","year":"2015","journal-title":"Int. J. Appl. Eng. Res."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1109\/MVT.2017.2670018","article-title":"Cooperative intelligent transport systems in Europe: Current deployment status and outlook","volume":"12","author":"Sjoberg","year":"2017","journal-title":"IEEE Veh. Technol. Mag."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1016\/j.vlsi.2017.07.007","article-title":"Algorithm and hardware implementation for visual perception system in autonomous vehicle: A survey","volume":"59","author":"Shi","year":"2017","journal-title":"Integration"},{"key":"ref_4","first-page":"595","article-title":"Road condition detection using smartphone sensors: A survey","volume":"7","author":"Chugh","year":"2014","journal-title":"Int. J. Electron. Electr. Eng."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Tsukada, M., Oi, T., Kitazawa, M., and Esaki, H. (2020). Networked roadside perception units for autonomous driving. Sensors, 20.","DOI":"10.3390\/s20185320"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chtourou, A., Merdrignac, P., and Shagdar, O. (May, January 25). Collective perception service for connected vehicles and roadside infrastructure. Proceedings of the 2021 IEEE 93rd Vehicular Technology Conference (VTC2021-Spring), Online.","DOI":"10.1109\/VTC2021-Spring51267.2021.9448753"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Ardianto, S., Chen, C.-J., and Hang, H.-M. (2017, January 22\u201324). Real-time traffic sign recognition using color segmentation and SVM. Proceedings of the 2017 International Conference on Systems, Signals and Image Processing (IWSSIP), Pozna\u0144, Poland.","DOI":"10.1109\/IWSSIP.2017.7965570"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"1080","DOI":"10.1049\/cje.2021.08.005","article-title":"Traffic Sign Recognition Using an Attentive Context Region-Based Detection Framework","volume":"30","author":"Zhigang","year":"2021","journal-title":"Chin. J. Electron."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Bai, Z., Wu, G., Qi, X., Liu, Y., Oguchi, K., and Barth, M.J. (2022, January 5\u20139). Infrastructure-based object detection and tracking for cooperative driving automation: A survey. Proceedings of the 2022 IEEE Intelligent Vehicles Symposium (IV), Aachen, Germany.","DOI":"10.1109\/IV51971.2022.9827461"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_12","first-page":"91","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren","year":"2015","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016, January 11\u201314). Ssd: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_15","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_16","unstructured":"Bochkovskiy, A., Wang, C.-Y., and Liao, H.-Y.M. (2004). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_18","unstructured":"Murugan, V., Vijaykumar, V., and Nidhila, A. (2019, January 18\u201319). A deep learning RCNN approach for vehicle recognition in traffic surveillance system. Proceedings of the 2019 International Conference on Communication and Signal Processing (ICCSP), Kuala Lumpur, Malaysia."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"3825532","DOI":"10.1155\/2022\/3825532","article-title":"Traffic sign detection via improved sparse R-CNN for autonomous vehicles","volume":"2022","author":"Liang","year":"2022","journal-title":"J. Adv. Transp."},{"key":"ref_20","unstructured":"Benjumea, A., Teeti, I., Cuzzolin, F., and Bradley, A. (2021). YOLO-Z: Improving small object detection in YOLOv5 for autonomous vehicles. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Du, F.-J., and Jiao, S.-J. (2022). Improvement of Lightweight Convolutional Neural Network Model Based on YOLO Algorithm and Its Research in Pavement Defect Detection. Sensors, 22.","DOI":"10.3390\/s22093537"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Wang, X., Hua, X., Xiao, F., Li, Y., Hu, X., and Sun, P. (2018). Multi-object detection in traffic scenes based on improved SSD. Electronics, 7.","DOI":"10.3390\/electronics7110302"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zhu, J., Li, X., Jin, P., Xu, Q., Sun, Z., and Song, X. (2020). Mme-yolo: Multi-sensor multi-level enhanced yolo for robust vehicle detection in traffic surveillance. Sensors, 21.","DOI":"10.3390\/s21010027"},{"key":"ref_24","first-page":"1","article-title":"YOLOv4-5D: An effective and efficient object detector for autonomous driving","volume":"70","author":"Cai","year":"2021","journal-title":"IEEE Trans. Instrum. Meas."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Li, L., Fang, M., Yin, Y., Lian, J., and Wang, Z. (2021, January 15\u201319). A Traffic Scene Object Detection Method Combining Deep Learning and Stereo Vision Algorithm. Proceedings of the 2021 IEEE International Conference on Real-Time Computing and Robotics (RCAR), Xining, China.","DOI":"10.1109\/RCAR52367.2021.9517460"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, C.-Y., Liao, H.-Y.M., Wu, Y.-H., Chen, P.-Y., Hsieh, J.-W., and Yeh, I.-H. (2020, January 14\u201319). CSPNet: A new backbone that can enhance learning capability of CNN. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00203"},{"key":"ref_27","unstructured":"Lazebnik, S., Schmid, C., and Ponce, J. (2006, January 17\u201322). Beyond bags of features: Spatial pyramid matching for recognizing natural scene categories. Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201906), New York, NY, USA."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Grauman, K., and Darrell, T. (2005, January 17\u201321). The pyramid match kernel: Discriminative classification with sets of image features. Proceedings of the Tenth IEEE International Conference on Computer Vision (ICCV\u201905), Beijing, China.","DOI":"10.1109\/ICCV.2005.239"},{"key":"ref_29","unstructured":"Neubeck, A., and Van Gool, L. (, January 20\u201324). Efficient non-maximum suppression. Proceedings of the 18th International Conference on Pattern Recognition (ICPR\u201906), Hong Kong, China."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path aggregation network for instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 13\u201319). Efficientdet: Scalable and efficient object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_34","unstructured":"Jiang, Y., Tan, Z., Wang, J., Sun, X., Lin, M., and Li, H. (2022). GiraffeDet: A Heavy-Neck Paradigm for Object Detection. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hou, Q., Zhou, D., and Feng, J. (2021, January 20\u201325). Coordinate attention for efficient mobile network design. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01350"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1016\/j.neucom.2022.07.042","article-title":"Focal and efficient IOU loss for accurate bounding box regression","volume":"506","author":"Zhang","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zheng, Z., Wang, P., Liu, W., Li, J., Ye, R., and Ren, D. (2020, January 7\u201312). Distance-IoU loss: Faster and better learning for bounding box regression. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Ye, X., Shu, M., Li, H., Shi, Y., Li, Y., Wang, G., Tan, X., and Ding, E. (2022, January 18\u201324). Rope3D: The Roadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection Task. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.02065"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Zhu, X., Lyu, S., Wang, X., and Zhao, Q. (2021, January 10\u201317). TPH-YOLOv5: Improved YOLOv5 based on transformer prediction head for object detection on drone-captured scenarios. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00312"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/21\/8097\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:00:50Z","timestamp":1760144450000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/21\/8097"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,22]]},"references-count":39,"journal-issue":{"issue":"21","published-online":{"date-parts":[[2022,11]]}},"alternative-id":["s22218097"],"URL":"https:\/\/doi.org\/10.3390\/s22218097","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,22]]}}}