{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,19]],"date-time":"2026-08-19T19:39:21Z","timestamp":1787168361087,"version":"build-2736575974"},"reference-count":33,"publisher":"MDPI AG","issue":"19","license":[{"start":{"date-parts":[[2023,9,30]],"date-time":"2023-09-30T00:00:00Z","timestamp":1696032000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["42130112"],"award-info":[{"award-number":["42130112"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>To cope with the challenges of autonomous driving in complex road environments, the need for collaborative multi-tasking has been proposed. This research direction explores new solutions at the application level and has become a hot topic of great interest. In the field of natural language processing and recommendation algorithms, the use of multi-task learning networks has been proven to reduce time, computing power, and storage usage in various task coupling cases. Due to the characteristics of the multi-task learning network, it has also been applied to visual road feature extraction in recent years. This article proposes a multi-task road feature extraction network that combines group convolution with transformer and squeeze excitation attention mechanisms. The network can simultaneously perform drivable area segmentation, lane line segmentation, and traffic object detection tasks. The experimental results of the BDD-100K dataset show that the proposed method performs well for different tasks and has a higher accuracy than similar algorithms. The proposed method provides new ideas and methods for the autonomous road perception of vehicles and the generation of highly accurate maps in visual-based autonomous driving processes.<\/jats:p>","DOI":"10.3390\/s23198182","type":"journal-article","created":{"date-parts":[[2023,10,2]],"date-time":"2023-10-02T04:39:30Z","timestamp":1696221570000},"page":"8182","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["A Multi-Task Road Feature Extraction Network with Grouped Convolution and Attention Mechanisms"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-2902-4703","authenticated-orcid":false,"given":"Wenjie","family":"Zhu","sequence":"first","affiliation":[{"name":"School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongwei","family":"Li","sequence":"additional","affiliation":[{"name":"School of Geo-Science & Technology, Zhengzhou University, Zhengzhou 450052, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xianglong","family":"Cheng","sequence":"additional","affiliation":[{"name":"School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yirui","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,9,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/2300000059","article-title":"Semantics for robotic mapping, perception and interaction: A survey","volume":"8","author":"Garg","year":"2020","journal-title":"Found. Trends Robot."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_4","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (July, January 26). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Ali, F. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_6","unstructured":"Redmon, J., and Ali, F. (2018). Yolov3: An incremental improvement. arXiv."},{"key":"ref_7","unstructured":"Bochkovskiy, A., Wang, C.-Y., Hong, Y., and Liao, M. (2020). Yolov4: Optimal speed and accuracy of object detection. arXiv."},{"key":"ref_8","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Medical Image Computing and Computer-Assisted Intervention\u2013MICCAI 2015. Proceedings of the 18th International Conference, Munich, Germany. Proceedings, Part III 18."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Dollar, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern RecognIition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_10","unstructured":"Zheng, T., Fag, H., Zhang, Y., Tang, W., Yang, Z., Liu, H., and Cai, D. (2023, January 7\u201314). Resa: Recurrent feature-shift aggregator for lane detection. Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Pan, X., Shi, J., and Luo, P. (2018, January 2\u20133). Spatial as deep: Spatial cnn for traffic scene understanding. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12301"},{"key":"ref_12","unstructured":"Davy, N., De Brabandere, B., Georgoulis, S., Proesmans, M., and Van Gool, L. (2018, January 26\u201330). Towards end-to-end lane detection: An instance segmentation approach. Proceedings of the IEEE Intelligent Vehicles Symposium (IV), Changshu, China."},{"key":"ref_13","unstructured":"Caruana, R. (1993, January 27\u201329). Multitask learning: A knowledge-based source of inductive bias1. Proceedings of the Tenth International Conference on Machine Learning, San Francisco, CA, USA."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Ma, J., Zhao, Z., Yi, X., Chen, J., Hong, L., and Chi, E.H. (2018, January 19\u201323). Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, London, UK.","DOI":"10.1145\/3219819.3220007"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Qin, Z., Chen, Y., Zhao, Z., Chen, Z., Metzler, D., and Qin, J. (2020, January 6\u201310). Multitask mixture of sequential experts for user activity streams. Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Virtual Event, CA USA.","DOI":"10.1145\/3394486.3403359"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhao, Z., Hong, L., Wei, L., Chen, S., Nath, A., Andrews, S., Jumthekar, M., Sathiamoorthy, M., Yi, X., and Chi, E. (2019, January 16\u201320). Recommending what video to watch next: A multitask ranking system. Proceedings of the 13th ACM Conference on Recommender Systems, Copenhagen, Denmark.","DOI":"10.1145\/3298689.3346997"},{"key":"ref_17","unstructured":"Ren, S., He, K., Girshick, R., and Sun, J. (2015, January 7\u201312). Faster R-CNN: Towards real-time object detection with region proposal networks. Proceedings of the Advances in Neural Information Processing Systems 28, Montreal, QC, Canada."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_19","unstructured":"Duan, K., Xie, L., Qi, H., Bai, S., Huang, Q., and Tian, Q. (2021). Location-sensitive visual recognition with cross-iou loss. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Teichmann, M., Weber, M., Zoellner, M., Cipolla, R., and Urtasun, R. (2018, January 26\u201330). Multinet: Real-time joint semantic reasoning for autonomous driving. Proceedings of the IEEE Intelligent Vehicles Symposium (IV), Changshu, China.","DOI":"10.1109\/IVS.2018.8500504"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"550","DOI":"10.1007\/s11633-022-1339-y","article-title":"Yolop: You only look once for panoptic driving perception","volume":"19","author":"Wu","year":"2022","journal-title":"Mach. Intell. Res."},{"key":"ref_22","unstructured":"Vu, D., Bao, N., and Hung, P. (2022). Hybridnets: End-to-end perception network. arXiv."},{"key":"ref_23","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kiaser, \u0141., and Polosukhin, I. (2017, January 4\u20137). Attention is all you need. Proceedings of the Annual Conference on Neural Information Processing System: Advances in Neural Information Processing Systems 30, Long Beach, CA, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Hu, J., Li, S., and Gang, S. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","article-title":"Spatial pyramid pooling in deep convolutional networks for visual recognition","volume":"37","author":"He","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, C.-Y., Lioa, H.-Y.M., Wu, Y.-H., Chen, P.-Y., Hsieh, J.-W., and Yeh, I.-H. (2020, January 14\u201319). CSPNet: A new backbone that can enhance learning capability of CNN. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00203"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Dollar, P., Tu, Z., and He, K. (2017, January 21\u201326). Aggregated residual transformations for deep neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., and Darrell, T. (2020, January 13\u201319). Bdd100k: A diverse driving dataset for heterogeneous multitask learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00271"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid scene parsing network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.660"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"4670","DOI":"10.1109\/TITS.2019.2943777","article-title":"DLT-Net: Joint detection of drivable areas, lane lines, and traffic objects","volume":"21","author":"Qian","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1145\/3140659.3080254","article-title":"SCNN: An accelerator for compressed-sparse convolutional neural networks","volume":"45","author":"Parashar","year":"2017","journal-title":"ACM SIGARCH Comput. Archit. News"},{"key":"ref_32","unstructured":"Paszke, A., Chaurasia, A., Kim, S., and Culurciello, E. (2016). Enet: A deep neural network architecture for real-time semantic segmentation. arXiv."},{"key":"ref_33","unstructured":"Hou, Y., Ma, Z., Liu, C., and Loy, C.C. (November, January 27). Learning lightweight lane detection cnns by self attention distillation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/19\/8182\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:02:41Z","timestamp":1760130161000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/19\/8182"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,30]]},"references-count":33,"journal-issue":{"issue":"19","published-online":{"date-parts":[[2023,10]]}},"alternative-id":["s23198182"],"URL":"https:\/\/doi.org\/10.3390\/s23198182","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,30]]}}}