{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T16:46:46Z","timestamp":1784738806825,"version":"3.55.0"},"reference-count":32,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2022,11,7]],"date-time":"2022-11-07T00:00:00Z","timestamp":1667779200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,11,7]],"date-time":"2022-11-07T00:00:00Z","timestamp":1667779200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach. Intell. Res."],"published-print":{"date-parts":[[2022,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>A panoptic driving perception system is an essential part of autonomous driving. A high-precision and real-time perception system can assist the vehicle in making reasonable decisions while driving. We present a panoptic driving perception network (you only look once for panoptic (YOLOP)) to perform traffic object detection, drivable area segmentation, and lane detection simultaneously. It is composed of one encoder for feature extraction and three decoders to handle the specific tasks. Our model performs extremely well on the challenging BDD100K dataset, achieving state-of-the-art on all three tasks in terms of accuracy and speed. Besides, we verify the effectiveness of our multi-task learning model for joint training via ablative studies. To our best knowledge, this is the first work that can process these three visual perception tasks simultaneously in real-time on an embedded device Jetson TX2(23 FPS), and maintain excellent accuracy. To facilitate further research, the source codes and pre-trained models are released at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/hustvl\/YOLOP\">https:\/\/github.com\/hustvl\/YOLOP<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s11633-022-1339-y","type":"journal-article","created":{"date-parts":[[2022,11,7]],"date-time":"2022-11-07T05:02:47Z","timestamp":1667797367000},"page":"550-562","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":400,"title":["YOLOP: You Only Look Once for Panoptic Driving Perception"],"prefix":"10.1007","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2306-5769","authenticated-orcid":false,"given":"Dong","family":"Wu","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Man-Wen","family":"Liao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wei-Tian","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6732-7823","authenticated-orcid":false,"given":"Xing-Gang","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiang","family":"Bai","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wen-Qing","family":"Cheng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wen-Yu","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,11,7]]},"reference":[{"key":"1339_CR1","first-page":"91","volume":"1","author":"S Q Ren","year":"2015","unstructured":"S. Q. Ren, K. M. He, R. Girshick, J. Sun. Faster R-CNN: Towards real-time object detection with region proposal networks. In Proceedings of the 28th International Conference on Neural Information Processing Systems, Montreal, Canada, vol. 1, pp. 91\u201399, 2015.","journal-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems"},{"key":"1339_CR2","doi-asserted-by":"publisher","unstructured":"A. Bochkovskiy, C. Y. Wang, H. Y. M. Liao. YOLOv4: Optimal speed and accuracy of object detection. DOI: https:\/\/doi.org\/10.48550\/arXiv.2004.10934., 2020","DOI":"10.48550\/arXiv.2004.10934"},{"key":"1339_CR3","doi-asserted-by":"publisher","unstructured":"A. Paszke, A. Chaurasia, S. Kim, E. Culurciello. ENet: A deep neural network architecture for real-time semantic segmentation. DOI: https:\/\/doi.org\/10.48550\/arXiv.1606.02147., 2016","DOI":"10.48550\/arXiv.1606.02147"},{"key":"1339_CR4","doi-asserted-by":"publisher","first-page":"6230","DOI":"10.1109\/CVPR.2017.660","volume-title":"Pyramid scene parsing network","author":"H S Zhao","year":"2017","unstructured":"H. S. Zhao, J. P. Shi, X. J. Qi, X. G. Wang, J. Y. Jia. Pyramid scene parsing network. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, IEEE, Honolulu, USA, pp. 6230\u20136239, 2017. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.660."},{"key":"1339_CR5","doi-asserted-by":"publisher","unstructured":"X. G. Pan, J. P. Shi, P. Luo, X. G. Wang, X. O. Tang. Spatial as deep: Spatial CNN for traffic scene understanding. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence, New Orleans, USA, pp. 7276\u20137283, 2018. DOI: https:\/\/doi.org\/10.1609\/aaai.v32i1.12301.","DOI":"10.1609\/aaai.v32i1.12301"},{"key":"1339_CR6","doi-asserted-by":"publisher","first-page":"1013","DOI":"10.1109\/ICCV.2019.00110","volume-title":"Learning lightweight lane detection CNNs by self attention distillation","author":"Y N Hou","year":"2019","unstructured":"Y. N. Hou, Z. Ma, C. X. Liu, C. C. Loy. Learning lightweight lane detection CNNs by self attention distillation. In Proceedings of IEEE\/CVF International Conference on Computer Vision, IEEE, Seoul, Korea, pp.1013\u20131021, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00110."},{"key":"1339_CR7","doi-asserted-by":"publisher","first-page":"13024","DOI":"10.1109\/CVPR46437.2021.01283","volume-title":"Scaled-YOLOv4: Scaling cross stage partial network","author":"C Y Wang","year":"2021","unstructured":"C. Y. Wang, A. Bochkovskiy, H. Y. M. Liao. Scaled-YOLOv4: Scaling cross stage partial network. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Nashville, USA, pp. 13024\u201313033, 2021. DOI: https:\/\/doi.org\/10.1109\/CVPR46437.2021.01283."},{"key":"1339_CR8","doi-asserted-by":"publisher","first-page":"2980","DOI":"10.1109\/ICCV.2017.322","volume-title":"Mask R-CNN","author":"K M He","year":"2017","unstructured":"K. M. He, G. Gkioxari, P. Doll\u00e1r, R. Girshick. Mask R-CNN. In Proceedings of IEEE International Conference on Computer Vision, IEEE, Venice, Italy, pp. 2980\u20132988, 2017. DOI: https:\/\/doi.org\/10.1109\/ICCV.2017.322."},{"key":"1339_CR9","doi-asserted-by":"publisher","unstructured":"F. Yu, W. Q. Xian, Y. Y. Chen, F. C. Liu, M. K. Liao, V. Madhavan, T. Darrell. BDD100K: A diverse driving video database with scalable annotation tooling. DOI: https:\/\/doi.org\/10.48550\/arXiv.1805.04687., 2018","DOI":"10.48550\/arXiv.1805.04687"},{"key":"1339_CR10","doi-asserted-by":"publisher","first-page":"580","DOI":"10.1109\/CVPR.2014.81","volume-title":"Rich feature hierarchies for accurate object detection and semantic segmentation","author":"R Girshick","year":"2014","unstructured":"R. Girshick, J. Donahue, T. Darrell, J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, IEEE, Columbus, USA, pp. 580\u2013587, 2014. DOI: https:\/\/doi.org\/10.1109\/CVPR.2014.81."},{"key":"1339_CR11","doi-asserted-by":"publisher","first-page":"1440","DOI":"10.1109\/ICCV.2015.169","volume-title":"Fast R-CNN","author":"R Girshick","year":"2015","unstructured":"R. Girshick. Fast R-CNN. In Proceedings of the IEEE International Conference on Computer Vision, IEEE, Santiago, Chile, pp. 1440\u20131448, 2015. DOI: https:\/\/doi.org\/10.1109\/ICCV.2015.169."},{"key":"1339_CR12","unstructured":"J. F. Dai, Y. Li, K. M. He, J. Sun. R-FCN: Object detection via region-based fully convolutional networks. In Proceedings of the 30th International Conference on Neural Information Processing Systems, Barcelona, Spain, pp. 379\u2013387, 2016."},{"key":"1339_CR13","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1007\/978-3-319-46448-0_2","volume-title":"SSD: Single shot MultiBox detector","author":"W Liu","year":"2016","unstructured":"W. Liu, D. Anguelov, D. Erhan, C. Szegedy, S. Reed, C. Y. Fu, A. C. Berg. SSD: Single shot MultiBox detector. In Proceedings of the 14th European Conference on Computer Vision, Springer, Amsterdam, The Netherlands, pp. 21\u201337, 2016. DOI: https:\/\/doi.org\/10.1007\/978-3-319-46448-0_2."},{"key":"1339_CR14","doi-asserted-by":"publisher","first-page":"779","DOI":"10.1109\/CVPR.2016.91","volume-title":"You only look once: Unified, real-time object detection","author":"J Redmon","year":"2016","unstructured":"J. Redmon, S. Divvala, R. Girshick, A. Farhadi. You only look once: Unified, real-time object detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, IEEE, Las Vegas, USA, pp. 779\u2013788, 2016. DOI: https:\/\/doi.org\/10.1109\/CVPR.2016.91."},{"key":"1339_CR15","doi-asserted-by":"publisher","first-page":"6517","DOI":"10.1109\/CVPR.2017.690","volume-title":"YOLO9000: Better, faster, stronger","author":"J Redmon","year":"2017","unstructured":"J. Redmon, A. Farhadi. YOLO9000: Better, faster, stronger. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, IEEE, Honolulu, USA, pp. 6517\u20136525, 2017. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.690."},{"key":"1339_CR16","unstructured":"J. Redmon, A. Farhadi. YOLOv3: An incremental improvement. [Online], Availabe: https:\/\/arxiv.org\/abs\/1804.02767, 2018."},{"key":"1339_CR17","doi-asserted-by":"publisher","first-page":"3431","DOI":"10.1109\/CVPR.2015.7298965","volume-title":"Fully convolutional networks for semantic segmentation","author":"J Long","year":"2015","unstructured":"J. Long, E. Shelhamer, T. Darrell. Fully convolutional networks for semantic segmentation. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, IEEE, Boston, USA, pp. 3431\u20133440, 2015. DOI: https:\/\/doi.org\/10.1109\/CVPR.2015.7298965."},{"issue":"2","key":"1339_CR18","doi-asserted-by":"publisher","first-page":"1041","DOI":"10.1109\/TITS.2019.2962094","volume":"22","author":"H Y Han","year":"2021","unstructured":"H. Y. Han, Y. C. Chen, P. Y. Hsiao, L. C. Fu. Using channel-wise attention for deep CNN based real-time semantic segmentation with class-aware edge information. IEEE Transactions on Intelligent Transportation Systems, vol. 22, no. 2, pp. 1041\u20131051, 2021. DOI: https:\/\/doi.org\/10.1109\/TITS.2019.2962094.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"1339_CR19","doi-asserted-by":"publisher","first-page":"286","DOI":"10.1109\/IVS.2018.8500547","volume-title":"Towards end-to-end lane detection: An instance segmentation approach","author":"D Neven","year":"2018","unstructured":"D. Neven, B. De Brabandere, S. Georgoulis, M. Proesmans, L. Van Gool. Towards end-to-end lane detection: An instance segmentation approach. In Proceedings of IEEE Intelligent Vehicles Symposium, IEEE, Changshu, China, pp. 286\u2013291, 2018. DOI: https:\/\/doi.org\/10.1109\/IVS.2018.8500547."},{"key":"1339_CR20","unstructured":"K. W. Duan, L. X. Xie, H. G. Qi, S. Bai, Q. M. Huang, Q. Tian. Location-sensitive visual recognition with cross-IOU loss. [Online], Available: https:\/\/arxiv.org\/abs\/2104.04899v1, 2021."},{"key":"1339_CR21","doi-asserted-by":"publisher","first-page":"1013","DOI":"10.1109\/IVS.2018.8500504","volume-title":"MultiNet: Real-time joint semantic reasoning for autonomous driving","author":"M Teichmann","year":"2018","unstructured":"M. Teichmann, M. Weber, M. Z\u00f6llner, R. Cipolla, R. Urtasun. MultiNet: Real-time joint semantic reasoning for autonomous driving. In Proceedings of IEEE Intelligent Vehicles Symposium, IEEE, Changshu, China, pp. 1013\u20131020, 2018. DOI: https:\/\/doi.org\/10.1109\/IVS.2018.8500504."},{"issue":"11","key":"1339_CR22","doi-asserted-by":"publisher","first-page":"4670","DOI":"10.1109\/TITS.2019.2943777","volume":"21","author":"Y Q Qian","year":"2020","unstructured":"Y. Q. Qian, J. M. Dolan, M. Yang. DLT-Net: Joint detection of drivable areas, lane lines, and traffic objects. IEEE Transactions on Intelligent Transportation Systems, vol. 21, no. 11, pp. 4670\u20134679, 2020. DOI: https:\/\/doi.org\/10.1109\/TITS.2019.2943777.","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"1339_CR23","doi-asserted-by":"publisher","first-page":"502","DOI":"10.1007\/978-3-030-01246-5_30","volume-title":"Geometric constrained joint lane segmentation and lane boundary detection","author":"J Zhang","year":"2018","unstructured":"J. Zhang, Y. Xu, B. B. Ni, Z. Y. Duan. Geometric constrained joint lane segmentation and lane boundary detection. In Proceedings of the 15th European Conference on Computer Vision, Springer, Munich, Germany, pp. 502\u2013518, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01246-5_30."},{"key":"1339_CR24","unstructured":"Z. L. Kang, K. Grauman, F. Sha. Learning with whom to share in multi-task feature learning. In Proceedings of the 28th International Conference on International Conference on Machine Learning, Bellevue, USA, pp. 521\u2013528, 2011."},{"key":"1339_CR25","doi-asserted-by":"publisher","first-page":"1571","DOI":"10.1109\/CVPRW50498.2020.00203","volume-title":"CSPNet: A new backbone that can enhance learning capability of CNN","author":"C Y Wang","year":"2020","unstructured":"C. Y. Wang, H. Y. M. Liao, Y. H. Wu, P. Y. Chen, J. W. Hsieh, I. H. Yeh. CSPNet: A new backbone that can enhance learning capability of CNN. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, IEEE, Seattle, USA, pp. 1571\u20131580, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPRW50498.2020.00203."},{"issue":"9","key":"1339_CR26","doi-asserted-by":"publisher","first-page":"1904","DOI":"10.1109\/TPAMI.2015.2389824","volume":"37","author":"K M He","year":"2015","unstructured":"K. M. He, X. Y. Zhang, S. Q. Ren, J. Sun. Spatial pyramid pooling in deep convolutional networks for visual recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 37, no. 9, pp. 1904\u20131916, 2015. DOI: https:\/\/doi.org\/10.1109\/TPAMI.2015.2389824.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1339_CR27","doi-asserted-by":"publisher","first-page":"936","DOI":"10.1109\/CVPR.2017.106","volume-title":"Feature pyramid networks for object detection","author":"T Y Lin","year":"2017","unstructured":"T. Y. Lin, P. Doll\u00e1r, R. Girshick, K. M. He, B. Hariharan, S. Belongie. Feature pyramid networks for object detection. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, IEEE, Honolulu, USA, pp. 936\u2013944, 2017. DOI: https:\/\/doi.org\/10.1109\/CVPR.2017.106."},{"key":"1339_CR28","doi-asserted-by":"publisher","first-page":"8759","DOI":"10.1109\/CVPR.2018.00913","volume-title":"Path aggregation network for instance segmentation","author":"S Liu","year":"2018","unstructured":"S. Liu, L. Qi, H. F. Qin, J. P. Shi, J. Y. Jia. Path aggregation network for instance segmentation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, USA, pp. 8759\u20138768, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00913."},{"key":"1339_CR29","doi-asserted-by":"publisher","first-page":"2999","DOI":"10.1109\/ICCV.2017.324","volume-title":"Focal loss for dense object detection","author":"T Y Lin","year":"2017","unstructured":"T. Y. Lin, P. Goyal, R. Girshick, K. M. He, P. Doll\u00e1r. Focal loss for dense object detection. In Proceedings of IEEE International Conference on Computer Vision, IEEE, Venice, Italy, pp. 2999\u20133007, 2017. DOI: https:\/\/doi.org\/10.1109\/ICCV.2017.324."},{"key":"1339_CR30","doi-asserted-by":"publisher","unstructured":"Z. H. Zheng, P. Wang, W. Liu, J. Z. Li, R. G. Ye, D. W. Ren. Distance-IoU loss: Faster and better learning for bounding box regression. In Proceedings of the 34th AAAI Conference on Artificial Intelligence, New York, USA, pp. 12993\u201313000, 2020. DOI: https:\/\/doi.org\/10.1609\/aaai.v34i07.6999.","DOI":"10.1609\/aaai.v34i07.6999"},{"key":"1339_CR31","unstructured":"I. Loshchilov, F. Hutter. SGDR: Stochastic gradient descent with warm restarts. In Proceedings of the 5th International Conference on Learning Representations, Toulon, France, 2017."},{"key":"1339_CR32","doi-asserted-by":"publisher","first-page":"770","DOI":"10.1109\/CVPR.2016.90","volume-title":"Deep residual learning for image recognition","author":"K M He","year":"2016","unstructured":"K. M. He, X. Y. Zhang, S. Q. Ren, J. Sun. Deep residual learning for image recognition. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, IEEE, Las Vegas, USA, pp. 770\u2013778, 2016. DOI: https:\/\/doi.org\/10.1109\/CVPR.2016.90."}],"updated-by":[{"DOI":"10.1007\/s11633-023-1452-6","type":"correction","label":"Correction","source":"publisher","updated":{"date-parts":[[2023,5,27]],"date-time":"2023-05-27T00:00:00Z","timestamp":1685145600000}}],"container-title":["Machine Intelligence Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-022-1339-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11633-022-1339-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-022-1339-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,6,1]],"date-time":"2023-06-01T10:22:30Z","timestamp":1685614950000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11633-022-1339-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,7]]},"references-count":32,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022,12]]}},"alternative-id":["1339"],"URL":"https:\/\/doi.org\/10.1007\/s11633-022-1339-y","relation":{"correction":[{"id-type":"doi","id":"10.1007\/s11633-023-1452-6","asserted-by":"object"}]},"ISSN":["2731-538X","2731-5398"],"issn-type":[{"value":"2731-538X","type":"print"},{"value":"2731-5398","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,11,7]]},"assertion":[{"value":"25 March 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 May 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 November 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 May 2023","order":4,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Correction","order":5,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"A Correction to this paper has been published:","order":6,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"https:\/\/doi.org\/10.1007\/s11633-023-1452-6","URL":"https:\/\/doi.org\/10.1007\/s11633-023-1452-6","order":7,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}}]}}