{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,17]],"date-time":"2025-11-17T03:00:00Z","timestamp":1763348400698,"version":"build-2065373602"},"reference-count":34,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2022,1,20]],"date-time":"2022-01-20T00:00:00Z","timestamp":1642636800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Romanian National Authority for Scientific Research and Control for Autonomous Systems","award":["PN-III-P4-ID-PCCF-2016-0180"],"award-info":[{"award-number":["PN-III-P4-ID-PCCF-2016-0180"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Panoptic segmentation provides a rich 2D environment representation by unifying semantic and instance segmentation. Most current state-of-the-art panoptic segmentation methods are built upon two-stage detectors and are not suitable for real-time applications, such as automated driving, due to their high computational complexity. In this work, we introduce a novel, fast and accurate single-stage panoptic segmentation network that employs a shared feature extraction backbone and three network heads for object detection, semantic segmentation, instance-level attention masks. Guided by object detections, our new panoptic segmentation head learns instance specific soft attention masks based on spatial embeddings. The semantic masks for stuff classes and soft instance masks for things classes are pixel-wise coherent and can be easily integrated in a panoptic output. The training and inference pipelines are simplified and no post-processing of the panoptic output is necessary. Benefiting from fast inference speed, the network can be deployed in automated vehicles or robotic applications. We perform extensive experiments on COCO and Cityscapes datasets and obtain competitive results in both accuracy and time. On the Cityscapes dataset we achieve 59.7 panoptic quality with an inference speed of more than 10 FPS on high resolution 1024 \u00d7 2048 images.<\/jats:p>","DOI":"10.3390\/s22030783","type":"journal-article","created":{"date-parts":[[2022,1,20]],"date-time":"2022-01-20T22:51:06Z","timestamp":1642719066000},"page":"783","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["Fast Panoptic Segmentation with Soft Attention Embeddings"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4036-6336","authenticated-orcid":false,"given":"Andra","family":"Petrovai","sequence":"first","affiliation":[{"name":"Computer Science Department, Technical University of Cluj-Napoca, Memorandumului 28, 400114 Cluj-Napoca, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sergiu","family":"Nedevschi","sequence":"additional","affiliation":[{"name":"Computer Science Department, Technical University of Cluj-Napoca, Memorandumului 28, 400114 Cluj-Napoca, Romania"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,1,20]]},"reference":[{"doi-asserted-by":"crossref","unstructured":"Kirillov, A., He, K., Girshick, R., Rother, C., and Doll\u00e1r, P. (2019, January 15\u201320). Panoptic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","key":"ref_1","DOI":"10.1109\/CVPR.2019.00963"},{"doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2016, January 27\u201330). The cityscapes dataset for semantic urban scene understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","key":"ref_2","DOI":"10.1109\/CVPR.2016.350"},{"doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the European Conference on Computer Vision, Zurich, Switzerland.","key":"ref_3","DOI":"10.1007\/978-3-319-10602-1_48"},{"doi-asserted-by":"crossref","unstructured":"Neuhold, G., Ollmann, T., Rota Bulo, S., and Kontschieder, P. (2017, January 21\u201326). The mapillary vistas dataset for semantic understanding of street scenes. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","key":"ref_4","DOI":"10.1109\/ICCV.2017.534"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1341","DOI":"10.1109\/TITS.2020.2972974","article-title":"Deep multi-modal object detection and semantic segmentation for autonomous driving: Datasets, methods, and challenges","volume":"22","author":"Feng","year":"2020","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"doi-asserted-by":"crossref","unstructured":"Porzi, L., Bulo, S.R., Colovic, A., and Kontschieder, P. (2019, January 15\u201320). Seamless scene segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","key":"ref_6","DOI":"10.1109\/CVPR.2019.00847"},{"doi-asserted-by":"crossref","unstructured":"Petrovai, A., and Nedevschi, S. (2019, January 27\u201330). Multi-task Network for Panoptic Segmentation in Automated Driving. Proceedings of the 2019 IEEE Intelligent Transportation Systems Conference (ITSC), Auckland, NZ, USA.","key":"ref_7","DOI":"10.1109\/ITSC.2019.8917422"},{"doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask r-cnn. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","key":"ref_8","DOI":"10.1109\/ICCV.2017.322"},{"doi-asserted-by":"crossref","unstructured":"Xiong, Y., Liao, R., Zhao, H., Hu, R., Bai, M., Yumer, E., and Urtasun, R. (2019, January 15\u201320). Upsnet: A unified panoptic segmentation network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","key":"ref_9","DOI":"10.1109\/CVPR.2019.00902"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1742","DOI":"10.1109\/LRA.2020.2969919","article-title":"Fast panoptic segmentation network","volume":"5","author":"Meletis","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"doi-asserted-by":"crossref","unstructured":"Hou, R., Li, J., Bhargava, A., Raventos, A., Guizilini, V., Fang, C., Lynch, J., and Gaidon, A. (2020, January 13\u201319). Real-Time Panoptic Segmentation from Dense Detections. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","key":"ref_11","DOI":"10.1109\/CVPR42600.2020.00855"},{"doi-asserted-by":"crossref","unstructured":"Cheng, B., Collins, M.D., Zhu, Y., Liu, T., Huang, T.S., Adam, H., and Chen, L.C. (2020, January 13\u201319). Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","key":"ref_12","DOI":"10.1109\/CVPR42600.2020.01249"},{"doi-asserted-by":"crossref","unstructured":"Tian, Z., Shen, C., Chen, H., and He, T. (2019, January 27\u201328). Fcos: Fully convolutional one-stage object detection. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea.","key":"ref_13","DOI":"10.1109\/ICCV.2019.00972"},{"doi-asserted-by":"crossref","unstructured":"Kirillov, A., Girshick, R., He, K., and Doll\u00e1r, P. (2019, January 15\u201320). Panoptic feature pyramid networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","key":"ref_14","DOI":"10.1109\/CVPR.2019.00656"},{"doi-asserted-by":"crossref","unstructured":"Sofiiuk, K., Barinova, O., and Konushin, A. (2019, January 27\u201328). Adaptis: Adaptive instance selection network. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea.","key":"ref_15","DOI":"10.1109\/ICCV.2019.00745"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1551","DOI":"10.1007\/s11263-021-01445-z","article-title":"Efficientps: Efficient panoptic segmentation","volume":"129","author":"Mohan","year":"2021","journal-title":"Int. J. Comput. Vis."},{"doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","key":"ref_17","DOI":"10.1109\/ICCV.2017.324"},{"doi-asserted-by":"crossref","unstructured":"Weber, M., Luiten, J., and Leibe, B. (2020, January 25\u201329). Single-shot panoptic segmentation. Proceedings of the 2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA.","key":"ref_18","DOI":"10.1109\/IROS45743.2020.9341546"},{"doi-asserted-by":"crossref","unstructured":"Gao, N., Shan, Y., Wang, Y., Zhao, X., Yu, Y., Yang, M., and Huang, K. (2019, January 27\u201328). Ssap: Single-shot instance segmentation with affinity pyramid. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea.","key":"ref_19","DOI":"10.1109\/ICCV.2019.00073"},{"doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","key":"ref_20","DOI":"10.1109\/CVPR.2017.106"},{"doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 17\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","key":"ref_21","DOI":"10.1109\/CVPR.2016.90"},{"doi-asserted-by":"crossref","unstructured":"Lee, Y., and Park, J. (2020, January 13\u201319). CenterMask: Real-time anchor-free instance segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","key":"ref_22","DOI":"10.1109\/CVPR42600.2020.01392"},{"unstructured":"(2022, January 14). NVIDIA TensorRT. Available online: https:\/\/developer.nvidia.com\/tensorrt.","key":"ref_23"},{"doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., and Jia, J. (2017, January 21\u201326). Pyramid scene parsing network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","key":"ref_24","DOI":"10.1109\/CVPR.2017.660"},{"unstructured":"Kendall, A., Gal, Y., and Cipolla, R. (2018, January 18\u201323). Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","key":"ref_25"},{"doi-asserted-by":"crossref","unstructured":"Neven, D., Brabandere, B.D., Proesmans, M., and Gool, L.V. (2019, January 15\u201320). Instance segmentation by jointly optimizing spatial embeddings and clustering bandwidth. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","key":"ref_26","DOI":"10.1109\/CVPR.2019.00904"},{"doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","key":"ref_27","DOI":"10.1109\/CVPR.2009.5206848"},{"doi-asserted-by":"crossref","unstructured":"Petrovai, A., and Nedevschi, S. (November, January 19). Real-Time Panoptic Segmentation with Prototype Masks for Automated Driving. Proceedings of the 2020 IEEE Intelligent Vehicles Symposium (IV), Las Vegas, NV, USA.","key":"ref_28","DOI":"10.1109\/IV47402.2020.9304836"},{"unstructured":"Yang, T.J., Collins, M.D., Zhu, Y., Hwang, J.J., Liu, T., Zhang, X., Sze, V., Papandreou, G., and Chen, L.C. (2019). Deeperlab: Single-shot image parser. arXiv.","key":"ref_29"},{"doi-asserted-by":"crossref","unstructured":"Li, Q., Qi, X., and Torr, P.H. (2020, January 13\u201319). Unifying training and inference for panoptic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","key":"ref_30","DOI":"10.1109\/CVPR42600.2020.01333"},{"doi-asserted-by":"crossref","unstructured":"Wang, H., Luo, R., Maire, M., and Shakhnarovich, G. (2020, January 13\u201319). Pixel consensus voting for panoptic segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","key":"ref_31","DOI":"10.1109\/CVPR42600.2020.00948"},{"doi-asserted-by":"crossref","unstructured":"Varga, R., Costea, A., Florea, H., Giosan, I., and Nedevschi, S. (2017, January 16\u201319). Super-sensor for 360-degree environment perception: Point cloud segmentation using image features. Proceedings of the 2017 IEEE 20th International Conference on Intelligent Transportation Systems (ITSC), Yokohama, Japan.","key":"ref_32","DOI":"10.1109\/ITSC.2017.8317846"},{"unstructured":"(2022, January 14). Urban Parking and Driving H2020 European Project (UP-Drive). Available online: https:\/\/up-drive.ethz.ch\/.","key":"ref_33"},{"unstructured":"(2022, January 14). The Automotive Data and Time Triggered Framework. Available online: https:\/\/www.elektrobit.com\/products\/automated-driving\/eb-assist\/adtf.","key":"ref_34"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/3\/783\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:04:42Z","timestamp":1760133882000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/3\/783"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,20]]},"references-count":34,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["s22030783"],"URL":"https:\/\/doi.org\/10.3390\/s22030783","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2022,1,20]]}}}