{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T04:35:18Z","timestamp":1781757318634,"version":"3.54.5"},"reference-count":34,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2024,2,28]],"date-time":"2024-02-28T00:00:00Z","timestamp":1709078400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["62071071"],"award-info":[{"award-number":["62071071"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>A vision-based autonomous driving perception system necessitates the accomplishment of a suite of tasks, including vehicle detection, drivable area segmentation, and lane line segmentation. In light of the limited computational resources available, multi-task learning has emerged as the preeminent methodology for crafting such systems. In this article, we introduce a highly efficient end-to-end multi-task learning model that showcases promising performance on all fronts. Our approach entails the development of a reliable feature extraction network by introducing a feature extraction module called C2SPD. Moreover, to account for the disparities among various tasks, we propose a dual-neck architecture. Finally, we present an optimized design for the decoders of each task. Our model evinces strong performance on the demanding BDD100K dataset, attaining remarkable accuracy (Acc) in vehicle detection and superior precision in drivable area segmentation (mIoU). In addition, this is the first work that can process these three visual perception tasks simultaneously in real time on an embedded device Atlas 200I A2 and maintain excellent accuracy.<\/jats:p>","DOI":"10.3390\/s24051547","type":"journal-article","created":{"date-parts":[[2024,2,28]],"date-time":"2024-02-28T06:14:22Z","timestamp":1709100862000},"page":"1547","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["A Multi-Task Network Based on Dual-Neck Structure for Autonomous Driving Perception"],"prefix":"10.3390","volume":"24","author":[{"given":"Guopeng","family":"Tan","sequence":"first","affiliation":[{"name":"School of Information & Electrical Engineering, Hebei University of Engineering, Handan 056038, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chao","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information & Electrical Engineering, Hebei University of Engineering, Handan 056038, China"},{"name":"Hebei Key Laboratory of Security & Protection Information Sensing and Processing, Handan 056038, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhihua","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information & Electrical Engineering, Hebei University of Engineering, Handan 056038, China"},{"name":"Hebei Key Laboratory of Security & Protection Information Sensing and Processing, Handan 056038, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuanbiao","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Information & Electrical Engineering, Hebei University of Engineering, Handan 056038, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ruikai","family":"Li","sequence":"additional","affiliation":[{"name":"School of Information & Electrical Engineering, Hebei University of Engineering, Handan 056038, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2024,2,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1137","DOI":"10.1109\/TPAMI.2016.2577031","article-title":"Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks","volume":"39","author":"Ren","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_2","unstructured":"Redmon, J., and Farhadi, A. (2018). Yolov3: An increme-ntal improvement. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Wang, C.-Y., Bochkovskiy, A., and Liao, H.-Y.M. (2021, January 20\u201325). Scaled-YOLOv4: Scaling Cross Stage Partial Network. Proceedings of the IEEE International Conference on Computer Vision, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01283"},{"key":"ref_4","unstructured":"Jocher, G. (2024, February 25). YOLOv5 Release v6.2. Available online: https:\/\/github.com\/ultralytics\/yolov5\/releases\/tag\/v6.2."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Wang, C.-Y., Bochkovskiy, A., and Liao, H.-Y.M. (2022). Yolov7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv.","DOI":"10.1109\/CVPR52729.2023.00721"},{"key":"ref_6","unstructured":"Jocher, G. (2024, February 25). Ultralytics YOLOv8. Available online: https:\/\/github.com\/ultralytics\/ultralytics."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Wu, D., Liao, M., Zhang, W., and Wang, X. (2021). Yolop: You only look once for panoptic driving perception. arXiv.","DOI":"10.1007\/s11633-022-1339-y"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Huang, H., Lin, L., Tong, R., and Hu, H. (2020, January 4\u20138). UNet 3+: A Full-Scale Connected UNet for Medical Image Segmentation. Proceedings of the ICA-SSP 2020\u20142020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053405"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"SegNet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid Scene Parsing Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Neven, D., Brabandere, B.D., Georgoulis, S., Proesmans, M., and Gool, L.V. (2018, January 26\u201330). Towards End-to-End Lane Detection: An Instance Segmentation Approach. Proceedings of the IEEE Intelligent Vehicles Symposium, Changshu, China.","DOI":"10.1109\/IVS.2018.8500547"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Pan, X., Shi, J., Luo, P., Wang, X., and Tang, X. (2018, January 2\u20137). Spatial as deep: Spatial cnn for traffic scene understanding. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12301"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1109\/TITS.2019.2890870","article-title":"Line-CNN: End-to-End Traffic Line Detection With Line Proposal Unit","volume":"21","author":"Li","year":"2020","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_14","unstructured":"Yu, F., Xian, W., Chen, Y., Liu, F., Liao, M., Madhavan, V., and Darrell, T. (2018). Bdd100k: A diverse driving video database with scalable annotation tooling. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Liu, W., Anguelov, D., Erhan, D., Szegedy, C., Reed, S., Fu, C.-Y., and Berg, A.C. (2016, January 11\u201314). SSD: Single shot multibox detector. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_2"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 13\u201319). EfficientDet Scalable and Efficient Object Detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1041","DOI":"10.1109\/TITS.2019.2962094","article-title":"Using Channel-Wise Attention for Deep CNN Based Real-Time Semantic Segmentation with Class-Aware Edge Information","volume":"22","author":"Han","year":"2021","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_18","unstructured":"Hou, Y., Ma, Z., Liu, C., and Loy, C.C. (November, January 27). Learning Lightweight Lane Detection CNNs by Self Attention Distillation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"7649","DOI":"10.1109\/TIP.2021.3107210","article-title":"Location Sensitive Network for Human Instance Segmentation","volume":"30","author":"Zhang","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Teichmann, M., Weber, M., Z\u00f6llner, M., Cipolla, R., and Urtasun, R. (2018, January 26\u201330). MultiNet: Real-time Joint Semantic Reasoning for Autonomous Driving. Proceedings of the IEEE Intelligent Vehicles Symposium, Changshu, China.","DOI":"10.1109\/IVS.2018.8500504"},{"key":"ref_22","unstructured":"Vu, D., Ngo, B., and Phan, H. (2022). Hybridnets: End-to-end perception network. arXiv."},{"key":"ref_23","unstructured":"Tan, M., and Le, Q.V. (2019). EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Zhu, W., Li, H., Cheng, X., and Jiang, Y. (2023). A Multi-Task Road Feature Extraction Network with Grouped Convolution and Attention Mechanisms. Sensors, 23.","DOI":"10.3390\/s23198182"},{"key":"ref_25","unstructured":"Han, C., Zhao, Q., Zhang, S., Chen, Y., Zhang, Z., and Yuan, J. (2022). YOLOPv2: Better, Faster, Stronger for Panoptic Driving Perception. arXiv."},{"key":"ref_26","unstructured":"Raja, S., and Luo, T. (2022). No More Strided Convolutions or Pooling: A New CNN Building Block for Low-Resolution Images and Small Objects. arXiv."},{"key":"ref_27","unstructured":"Park, H., Yoo, Y., Seo, G., Han, D., Yun, S., and Kwak, N. (2018). C3: Concentrated-Comprehensive Convolution and its application to semantic segmentation. arXiv."},{"key":"ref_28","unstructured":"Wang, C.-Y., Liao, H.-Y.M., and Yeh, I.-H. (2022). Designing Network Design Strategies Through Gradient Path Analysis. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., and Jia, J. (2018, January 18\u201323). Path Aggregation Network for Instance Segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00913"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"318","DOI":"10.1109\/TPAMI.2018.2858826","article-title":"Focal Loss for Dense Object Detection","volume":"42","author":"Lin","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Salehi, S.S.M., Erdogmus, D., and Gholipour, A. (2017). Tversky Loss Function for Image Segmentation Using 3D Fully Convolutional Deep Networks. arXiv.","DOI":"10.1007\/978-3-319-67389-9_44"},{"key":"ref_32","unstructured":"Loshchilov, I., and Hutter, F. (2016). SGDR: Stochastic gradient descent with warm restarts. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"4670","DOI":"10.1109\/TITS.2019.2943777","article-title":"DLT-Net: Joint Detection of Drivable Areas, Lane Lines, and Traffic Objects","volume":"21","author":"Qian","year":"2020","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_34","unstructured":"Paszke, A., Chaurasia, A., and Kim, S. (2016). ENet: A Deep Neural Network Architecture for Real-Time Semantic Segmentation. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/5\/1547\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T14:06:15Z","timestamp":1760105175000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/24\/5\/1547"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,28]]},"references-count":34,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2024,3]]}},"alternative-id":["s24051547"],"URL":"https:\/\/doi.org\/10.3390\/s24051547","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,28]]}}}