{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T03:24:14Z","timestamp":1783481054101,"version":"3.55.0"},"reference-count":87,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2021,2,26]],"date-time":"2021-02-26T00:00:00Z","timestamp":1614297600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,2,26]],"date-time":"2021-02-26T00:00:00Z","timestamp":1614297600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100010663","name":"H2020 European Research Council","doi-asserted-by":"publisher","award":["871449"],"award-info":[{"award-number":["871449"]}],"id":[{"id":"10.13039\/100010663","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002347","name":"Bundesministerium f\u00fcr Bildung und Forschung","doi-asserted-by":"publisher","award":["ISA4.0"],"award-info":[{"award-number":["ISA4.0"]}],"id":[{"id":"10.13039\/501100002347","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100006785","name":"Google","doi-asserted-by":"publisher","award":["Google Cloud research grant"],"award-info":[{"award-number":["Google Cloud research grant"]}],"id":[{"id":"10.13039\/100006785","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2021,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Understanding the scene in which an autonomous robot operates is critical for its competent functioning. Such scene comprehension necessitates recognizing instances of traffic participants along with general scene semantics which can be effectively addressed by the panoptic segmentation task. In this paper, we introduce the Efficient Panoptic Segmentation (EfficientPS) architecture that consists of a shared backbone which efficiently encodes and fuses semantically rich multi-scale features. We incorporate a new semantic head that aggregates fine and contextual features coherently and a new variant of Mask R-CNN as the instance head. We also propose a novel panoptic fusion module that congruously integrates the output logits from both the heads of our EfficientPS architecture to yield the final panoptic segmentation output. Additionally, we introduce the KITTI panoptic segmentation dataset that contains panoptic annotations for the popularly challenging KITTI benchmark. Extensive evaluations on Cityscapes, KITTI, Mapillary Vistas and Indian Driving Dataset demonstrate that our proposed architecture consistently sets the new state-of-the-art on all these four benchmarks while being the most efficient and fast panoptic segmentation architecture to date.\n<\/jats:p>","DOI":"10.1007\/s11263-021-01445-z","type":"journal-article","created":{"date-parts":[[2021,2,26]],"date-time":"2021-02-26T06:02:38Z","timestamp":1614319358000},"page":"1551-1579","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":208,"title":["EfficientPS: Efficient Panoptic Segmentation"],"prefix":"10.1007","volume":"129","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5067-4279","authenticated-orcid":false,"given":"Rohit","family":"Mohan","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4710-3114","authenticated-orcid":false,"given":"Abhinav","family":"Valada","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,2,26]]},"reference":[{"key":"1445_CR1","doi-asserted-by":"crossref","unstructured":"Arbel\u00e1ez, P., Pont-Tuset, J., Barron, J. T., Marques, F., & Malik, J. (2014). Multiscale combinatorial grouping. In Proceedings of the conference on computer vision and pattern recognition (pp. 328\u2013335).","DOI":"10.1109\/CVPR.2014.49"},{"issue":"12","key":"1445_CR2","doi-asserted-by":"publisher","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","volume":"39","author":"V Badrinarayanan","year":"2017","unstructured":"Badrinarayanan, V., Kendall, A., & Cipolla, R. (2017). Segnet: A deep convolutional encoder-decoder architecture for image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 39(12), 2481\u20132495.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1445_CR3","doi-asserted-by":"crossref","unstructured":"Bai, M., & Urtasun, R. (2017). Deep watershed transform for instance segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 5221\u20135229).","DOI":"10.1109\/CVPR.2017.305"},{"key":"1445_CR4","volume-title":"Theories of infant development","author":"JG Bremner","year":"2008","unstructured":"Bremner, J. G., & Slater, A. (2008). Theories of infant development. London: Wiley."},{"key":"1445_CR5","doi-asserted-by":"crossref","unstructured":"Brostow, G.J., Shotton, J., Fauqueur, J., & Cipolla, R. (2008). Segmentation and recognition using structure from motion point clouds. In European conference on computer vision, Springer (pp. 44\u201357).","DOI":"10.1007\/978-3-540-88682-2_5"},{"key":"1445_CR6","doi-asserted-by":"crossref","unstructured":"Bulo, S. R., Neuhold, G., & Kontschieder, P. (2017). Loss max-pooling for semantic image segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 7082\u20137091).","DOI":"10.1109\/CVPR.2017.749"},{"issue":"4","key":"1445_CR7","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","volume":"40","author":"LC Chen","year":"2017","unstructured":"Chen, L. C., Papandreou, G., Kokkinos, I., Murphy, K., & Yuille, A. L. (2017a). Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 40(4), 834\u2013848.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1445_CR8","unstructured":"Chen, L. C., Papandreou, G., Schroff, F., & Adam, H. (2017b). Rethinking atrous convolution for semantic image segmentation. arXiv preprint arXiv:1706.05587."},{"key":"1445_CR9","unstructured":"Chen, L. C., Collins, M., Zhu, Y., Papandreou, G., Zoph, B., Schroff, F., Adam, H., & Shlens, J. (2018a). Searching for efficient multi-scale architectures for dense image prediction. In Advances in neural information processing systems (pp. 8713\u20138724)."},{"key":"1445_CR10","doi-asserted-by":"crossref","unstructured":"Chen, L. C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018b) Encoder-decoder with atrous separable convolution for semantic image segmentation. arXiv preprint arXiv:1802.02611.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"1445_CR11","doi-asserted-by":"crossref","unstructured":"Cheng, B., Collins, M. D., Zhu, Y., Liu, T., Huang, T. S., Adam, H., & Chen, L. C. (2020). Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 12475\u201312485).","DOI":"10.1109\/CVPR42600.2020.01249"},{"key":"1445_CR12","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017). Xception: Deep learning with depthwise separable convolutions. In Proceedings of the conference on computer vision and pattern recognition (pp. 1251\u20131258).","DOI":"10.1109\/CVPR.2017.195"},{"key":"1445_CR13","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., & Schiele, B. (2016). The cityscapes dataset for semantic urban scene understanding. In Proceedings of the conference on computer vision and pattern recognition (pp. 3213\u20133223).","DOI":"10.1109\/CVPR.2016.350"},{"key":"1445_CR14","doi-asserted-by":"crossref","unstructured":"Dai, J., He, K., Li, Y., Ren, S., & Sun, J. (2016). Instance-sensitive fully convolutional networks. In European conference on computer vision (pp. 534\u2013549).","DOI":"10.1007\/978-3-319-46466-4_32"},{"key":"1445_CR15","doi-asserted-by":"crossref","unstructured":"Dai, J., Qi, H., Xiong, Y., Li, Y., Zhang, G., Hu, H., & Wei, Y. (2017). Deformable convolutional networks. In Proceedings of the international conference on computer vision (pp. 764\u2013773).","DOI":"10.1109\/ICCV.2017.89"},{"key":"1445_CR16","unstructured":"de\u00a0Geus, D., Meletis, P., & Dubbelman, G. (2018). Panoptic segmentation with a joint semantic and instance segmentation network. arXiv preprint arXiv:1809.02110."},{"key":"1445_CR17","doi-asserted-by":"crossref","unstructured":"Gao, N., Shan, Y., Wang, Y., Zhao, X., Yu, Y., Yang, M., & Huang, K. (2019). Ssap: Single-shot instance segmentation with affinity pyramid. In Proceedings of the international conference on computer vision (pp. 642\u2013651).","DOI":"10.1109\/ICCV.2019.00073"},{"key":"1445_CR18","first-page":"79","volume":"5","author":"A Geiger","year":"2013","unstructured":"Geiger, A., Lenz, P., Stiller, C., & Urtasun, R. (2013). Vision meets robotics: The kitti dataset. International Journal of Robotics Research., 5, 79.","journal-title":"International Journal of Robotics Research."},{"key":"1445_CR19","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015). Fast r-cnn. In Proceedings of the international conference on computer vision (pp. 1440\u20131448).","DOI":"10.1109\/ICCV.2015.169"},{"key":"1445_CR20","unstructured":"Glorot, X., & Bengio, Y. (2010). Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the thirteenth international conference on artificial intelligence and statistics (pp. 249\u2013256)."},{"key":"1445_CR21","doi-asserted-by":"crossref","unstructured":"Hariharan, B., Arbel\u00e1ez, P., Girshick, R., & Malik, J. (2014). Simultaneous detection and segmentation. In European conference on computer vision (pp. 297\u2013312).","DOI":"10.1007\/978-3-319-10584-0_20"},{"key":"1445_CR22","doi-asserted-by":"crossref","unstructured":"Hariharan, B., Arbel\u00e1ez, P., Girshick, R., & Malik, J. (2015). Hypercolumns for object segmentation and fine-grained localization. In Proceedings of the conference on computer vision and pattern recognition (pp. 447\u2013456).","DOI":"10.1109\/CVPR.2015.7298642"},{"key":"1445_CR23","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the conference on computer vision and pattern recognition (pp. 770\u2013778).","DOI":"10.1109\/CVPR.2016.90"},{"key":"1445_CR24","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., & Girshick, R. (2017). Mask r-cnn. In Proceedings of the international conference on computer vision (pp. 2961\u20132969).","DOI":"10.1109\/ICCV.2017.322"},{"key":"1445_CR25","doi-asserted-by":"crossref","unstructured":"He, X., & Gould, S. (2014a). An exemplar-based crf for multi-instance object segmentation. In Proceedings of the conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2014.45"},{"key":"1445_CR26","doi-asserted-by":"crossref","unstructured":"He, X., & Gould, S. (2014b). An exemplar-based crf for multi-instance object segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 296\u2013303).","DOI":"10.1109\/CVPR.2014.45"},{"key":"1445_CR27","doi-asserted-by":"crossref","unstructured":"Howard, A., Sandler, M., Chu, G., Chen, L. C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., Vasudevan, V., et\u00a0al. (2019). Searching for mobilenetv3. In Proceedings of the international conference on computer vision (pp. 1314\u20131324).","DOI":"10.1109\/ICCV.2019.00140"},{"key":"1445_CR28","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the conference on computer vision and pattern recognition (pp. 7132\u20137141).","DOI":"10.1109\/CVPR.2018.00745"},{"key":"1445_CR29","unstructured":"Ioffe, S., & Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of the 32nd international conference on international conference on machine learning, JMLR.org, ICML\u201915 (Vol. 37, pp. 448\u2013456)."},{"key":"1445_CR30","unstructured":"Kaiser, L., Gomez, A. N., & Chollet, F. (2017). Depthwise separable convolutions for neural machine translation. arXiv preprint arXiv:1706.03059."},{"key":"1445_CR31","unstructured":"Kang, B. R., & Kim, H. Y. (2018). Bshapenet: Object detection and instance segmentation with bounding shape masks. arXiv preprint arXiv:1810.10327."},{"key":"1445_CR32","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Girshick, R., He, K., & Doll\u00e1r, P. (2019a) Panoptic feature pyramid networks. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 6399\u20136408).","DOI":"10.1109\/CVPR.2019.00656"},{"key":"1445_CR33","doi-asserted-by":"crossref","unstructured":"Kirillov, A., He, K., Girshick, R., Rother, C., & Doll\u00e1r, P. (2019b). Panoptic segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 9404\u20139413).","DOI":"10.1109\/CVPR.2019.00963"},{"key":"1445_CR34","doi-asserted-by":"crossref","unstructured":"Kontschieder, P., Bulo, S. R., Bischof, H., & Pelillo, M. (2011). Structured class-labels in random forests for semantic image labelling. In Proceedings of the international conference on computer vision (pp. 2190\u20132197).","DOI":"10.1109\/ICCV.2011.6126496"},{"key":"1445_CR35","unstructured":"Kr\u00e4henb\u00fchl, P., & Koltun, V. (2011). Efficient inference in fully connected crfs with gaussian edge potentials. In Advances in neural information processing systems (pp. 109\u2013117)."},{"key":"1445_CR36","unstructured":"Li, J., Raventos, A., Bhargava, A., Tagawa, T., & Gaidon, A. (2018a). Learning to fuse things and stuff. arXiv preprint arXiv:1812.01192."},{"key":"1445_CR37","doi-asserted-by":"crossref","unstructured":"Li, Q., Arnab, A., & Torr, P. H. (2018b). Weakly-and semi-supervised panoptic segmentation. In Proceedings of the European conference on computer vision (ECCV) (pp. 102\u2013118).","DOI":"10.1007\/978-3-030-01267-0_7"},{"key":"1445_CR38","unstructured":"Li, X., Zhang, L., You, A., Yang, M., Yang, K., & Tong, Y. (2019a). Global aggregation then local distribution in fully convolutional networks. arXiv preprint arXiv:1909.07229."},{"key":"1445_CR39","doi-asserted-by":"crossref","unstructured":"Li, Y., Qi, H., Dai, J., Ji, X., & Wei, Y. (2017). Fully convolutional instance-aware semantic segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 2359\u20132367).","DOI":"10.1109\/CVPR.2017.472"},{"key":"1445_CR40","doi-asserted-by":"crossref","unstructured":"Li, Y., Chen, X., Zhu, Z., Xie, L., Huang, G., Du, D., & Wang, X. (2019b). Attention-guided unified network for panoptic segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 7026\u20137035).","DOI":"10.1109\/CVPR.2019.00719"},{"key":"1445_CR41","doi-asserted-by":"crossref","unstructured":"Lin, T. Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., & Zitnick, C. L. (2014). Microsoft coco: Common objects in context. In European conference on computer vision (pp. 740\u2013755), Springer.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"1445_CR42","doi-asserted-by":"crossref","unstructured":"Lin, T. Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. In Proceedings of the conference on computer vision and pattern recognition (pp. 2117\u20132125).","DOI":"10.1109\/CVPR.2017.106"},{"key":"1445_CR43","doi-asserted-by":"crossref","unstructured":"Liu, H., Peng, C., Yu, C., Wang, J., Liu, X., Yu, G., & Jiang, W. (2019). An end-to-end network for panoptic segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 6172\u20136181).","DOI":"10.1109\/CVPR.2019.00633"},{"key":"1445_CR44","doi-asserted-by":"crossref","unstructured":"Liu, S., Jia, J., Fidler, S., & Urtasun, R. (2017). Sgn: Sequential grouping networks for instance segmentation. In Proceedings of the international conference on computer vision (pp. 3496\u20133504).","DOI":"10.1109\/ICCV.2017.378"},{"key":"1445_CR45","doi-asserted-by":"crossref","unstructured":"Liu, S., Qi, L., Qin, H., Shi, J., & Jia, J. (2018). Path aggregation network for instance segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 8759\u20138768).","DOI":"10.1109\/CVPR.2018.00913"},{"key":"1445_CR46","unstructured":"Liu, W., Rabinovich, A., & Berg, A. C. (2015). Parsenet: Looking wider to see better. arXiv preprint arXiv:1506.04579."},{"key":"1445_CR47","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., & Darrell, T. (2015). Fully convolutional networks for semantic segmentation. In Proceedings of the conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"1445_CR48","doi-asserted-by":"crossref","unstructured":"Neuhold, G., Ollmann, T., Rota, B. S., & Kontschieder, P. (2017). The mapillary vistas dataset for semantic understanding of street scenes. In Proceedings of the international conference on computer vision (pp. 4990\u20134999).","DOI":"10.1109\/ICCV.2017.534"},{"key":"1445_CR49","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et\u00a0al. (2019). Pytorch: An imperative style, high-performance deep learning library. In Advances in neural information processing systems (pp. 8024\u20138035)."},{"key":"1445_CR50","unstructured":"Pinheiro, P. O., Collobert, R., & Doll\u00e1r, P. (2015). Learning to segment object candidates. In Advances in neural information processing systems (pp. 1990\u20131998)."},{"key":"1445_CR51","doi-asserted-by":"crossref","unstructured":"Plath, N., Toussaint, M., & Nakajima, S. (2009). Multi-class image segmentation using conditional random fields and global classification. In Proceedings of the international conference on machine learning (pp. 817\u2013824).","DOI":"10.1145\/1553374.1553479"},{"key":"1445_CR52","doi-asserted-by":"crossref","unstructured":"Porzi, L., Bulo, S. R., Colovic, A., & Kontschieder, P. (2019). Seamless scene segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 8277\u20138286).","DOI":"10.1109\/CVPR.2019.00847"},{"key":"1445_CR53","unstructured":"Radwan, N., Valada, A., & Burgard, W. (2018). Multimodal interaction-aware motion prediction for autonomous street crossing. arXiv preprint arXiv:1808.06887."},{"key":"1445_CR54","doi-asserted-by":"crossref","unstructured":"Ren, M., & Zemel, R. S. (2017). End-to-end instance segmentation with recurrent attention. In Proceedings of the conference on computer vision and pattern recognition (pp. 6656\u20136664).","DOI":"10.1109\/CVPR.2017.39"},{"key":"1445_CR55","doi-asserted-by":"crossref","unstructured":"Romera-Paredes, B., & Torr, P. H. S. (2016). Recurrent instance segmentation. In European conference on computer vision (pp. 312\u2013329), Springer.","DOI":"10.1007\/978-3-319-46466-4_19"},{"key":"1445_CR56","doi-asserted-by":"crossref","unstructured":"Ros, G., Ramos, S., Granados, M., Bakhtiary, A., Vazquez, D., & Lopez, A. M. (2015). Vision-based offline-online perception paradigm for autonomous driving. In IEEE winter conference on applications of computer vision (pp. 231\u2013238).","DOI":"10.1109\/WACV.2015.38"},{"key":"1445_CR57","unstructured":"Rota, B. S., Porzi, L., & Kontschieder, P. (2018). In-place activated batchnorm for memory-optimized training of dnns. In Proceedings of the conference on computer vision and pattern recognition (pp. 5639\u20135647)."},{"issue":"3","key":"1445_CR58","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., et al. (2015). Imagenet large scale visual recognition challenge. International Journal of Computer Vision, 115(3), 211\u2013252.","journal-title":"International Journal of Computer Vision"},{"key":"1445_CR59","doi-asserted-by":"crossref","unstructured":"Shotton, J., Johnson, M., & Cipolla, R. (2008). Semantic texton forests for image categorization and segmentation. In Proceedings of the conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2008.4587503"},{"key":"1445_CR60","doi-asserted-by":"crossref","unstructured":"Silberman, N., Sontag, D., & Fergus, R. (2014). Instance segmentation of indoor scenes using a coverage loss. In European conference on computer vision (pp. 616\u2013631).","DOI":"10.1007\/978-3-319-10590-1_40"},{"key":"1445_CR61","unstructured":"Simonyan, K., & Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556."},{"key":"1445_CR62","doi-asserted-by":"crossref","unstructured":"Sofiiuk, K., Barinova, O., & Konushin, A. (2019). Adaptis: Adaptive instance selection network. In Proceedings of the international conference on computer vision (pp. 7355\u20137363).","DOI":"10.1109\/ICCV.2019.00745"},{"key":"1445_CR63","doi-asserted-by":"crossref","unstructured":"Sturgess, P., Alahari, K., Ladicky, L., & Torr, P. H. (2009). Combining appearance and structure from motion features for road scene understanding. In British machine vision conference.","DOI":"10.5244\/C.23.62"},{"issue":"7","key":"1445_CR64","doi-asserted-by":"publisher","first-page":"1370","DOI":"10.1109\/TPAMI.2013.193","volume":"36","author":"M Sun","year":"2013","unstructured":"Sun, M., Bs, K., Kohli, P., & Savarese, S. (2013). Relating things and stuff via objectproperty interactions. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7), 1370\u20131383.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"1445_CR65","unstructured":"Tan, M., & Le, Q.V. (2019). Efficientnet: Rethinking model scaling for convolutional neural networks. arXiv preprint arXiv:1905.11946."},{"key":"1445_CR66","doi-asserted-by":"crossref","unstructured":"Tian, Z., He, T., Shen, C., & Yan, Y. (2019). Decoders matter for semantic segmentation: Data-dependent decoding enables flexible feature aggregation. In Proceedings of the conference on computer vision and pattern recognition (pp. 3126\u20133135).","DOI":"10.1109\/CVPR.2019.00324"},{"key":"1445_CR67","doi-asserted-by":"crossref","unstructured":"Tighe, J., & Lazebnik, S. (2013). Finding things: Image parsing with regions and per-exemplar detectors. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 3001\u20133008).","DOI":"10.1109\/CVPR.2013.386"},{"key":"1445_CR68","doi-asserted-by":"crossref","unstructured":"Tighe, J., Niethammer, M., & Lazebnik, S. (2014). Scene parsing with object instances and occlusion ordering. In Proceedings of the conference on computer vision and pattern recognition (pp. 3748\u20133755).","DOI":"10.1109\/CVPR.2014.479"},{"issue":"2","key":"1445_CR69","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1007\/s11263-005-6642-x","volume":"63","author":"Z Tu","year":"2005","unstructured":"Tu, Z., Chen, X., Yuille, A. L., & Zhu, S. C. (2005). Image parsing: Unifying segmentation, detection, and recognition. International Journal of Computer Vision, 63(2), 113\u2013140.","journal-title":"International Journal of Computer Vision"},{"key":"1445_CR70","doi-asserted-by":"crossref","unstructured":"Uhrig, J., Cordts, M., Franke, U., & Brox, T. (2016). Pixel-level encoding and depth layering for instance-level semantic labeling. In German conference on pattern recognition (pp. 14\u201325).","DOI":"10.1007\/978-3-319-45886-1_2"},{"key":"1445_CR71","unstructured":"Valada, A., Dhall, A., & Burgard, W. (2016a). Convoluted mixture of deep experts for robust semantic segmentation. In IEEE\/RSJ international conference on intelligent robots and systems (IROS) workshop, state estimation and terrain perception for all terrain mobile robots."},{"key":"1445_CR72","unstructured":"Valada, A., Oliveira, G., Brox, T., & Burgard, W. (2016b). Towards robust semantic segmentation using deep fusion. In Robotics: Science and systems (RSS 2016) workshop, are the sceptics right? Limits and potentials of deep learning in robotics."},{"key":"1445_CR73","doi-asserted-by":"crossref","unstructured":"Valada, A., Vertens, J., Dhall, A., & Burgard, W. (2017). Adapnet: Adaptive semantic segmentation in adverse environmental conditions. In Proceedings of the IEEE international conference on robotics and automation (pp. 4644\u20134651).","DOI":"10.1109\/ICRA.2017.7989540"},{"key":"1445_CR74","unstructured":"Valada, A., Radwan, N., & Burgard, W. (2018). Incorporating semantic and geometric priors in deep pose regression. In Workshop on learning and inference in robotics: Integrating structure, priors and models at robotics: Science and systems (RSS)."},{"key":"1445_CR75","doi-asserted-by":"publisher","unstructured":"Valada, A., Mohan, R., & Burgard, W. (2019). Self-supervised model adaptation for multimodal semantic segmentation. International Journal of Computer Vision,. https:\/\/doi.org\/10.1007\/s11263-019-01188-y, special Issue: Deep Learning for Robotic VisionD","DOI":"10.1007\/s11263-019-01188-y"},{"key":"1445_CR76","doi-asserted-by":"crossref","unstructured":"Varma, G., Subramanian, A., Namboodiri, A., Chandraker, M., & Jawahar, C. (2019). Idd: A dataset for exploring problems of autonomous navigation in unconstrained environments. In IEEE winter conference on applications of computer vision (WACV) (pp. 1743\u20131751).","DOI":"10.1109\/WACV.2019.00190"},{"key":"1445_CR77","doi-asserted-by":"crossref","unstructured":"Wu, Y., & He, K. (2018). Group normalization. In Proceedings of the European conference on computer vision (ECCV) (pp. 3\u201319).","DOI":"10.1007\/978-3-030-01261-8_1"},{"key":"1445_CR78","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Doll\u00e1r, P., Tu, Z., & He, K. (2017). Aggregated residual transformations for deep neural networks. In Proceedings of the conference on computer vision and pattern recognition (pp. 1492\u20131500).","DOI":"10.1109\/CVPR.2017.634"},{"key":"1445_CR79","doi-asserted-by":"crossref","unstructured":"Xiong, Y., Liao, R., Zhao, H., Hu, R., Bai, M., Yumer, E., & Urtasun, R. (2019). Upsnet: A unified panoptic segmentation network. In Proceedings of the conference on computer vision and pattern recognition (pp. 8818\u20138826).","DOI":"10.1109\/CVPR.2019.00902"},{"issue":"3","key":"1445_CR80","doi-asserted-by":"publisher","first-page":"331","DOI":"10.1007\/s00138-014-0649-7","volume":"27","author":"P Xu","year":"2016","unstructured":"Xu, P., Davoine, F., Bordes, J. B., Zhao, H., & Den\u0153ux, T. (2016). Multimodal information fusion for urban scene understanding. Machine Vision and Applications, 27(3), 331\u2013349.","journal-title":"Machine Vision and Applications"},{"key":"1445_CR81","unstructured":"Yang, T. J., Collins, M. D., Zhu, Y., Hwang, J. J., Liu, T., Zhang, X., Sze, V., Papandreou, G., & Chen, L. C. (2019). Deeperlab: Single-shot image parser. arXiv preprint arXiv:1902.05093."},{"key":"1445_CR82","unstructured":"Yao, J., Fidler, S., & Urtasun, R. (2012). Describing the scene as a whole: Joint object detection, scene classification and semantic segmentation. In Proceedings of the conference on computer vision and pattern recognition (pp. 702\u2013709)."},{"key":"1445_CR83","unstructured":"Yu, F., & Koltun, V. (2015). Multi-scale context aggregation by dilated convolutions. arXiv preprint arXiv:1511.07122."},{"key":"1445_CR84","doi-asserted-by":"crossref","unstructured":"Zhang. C., Wang, L., & Yang, R. (2010). Semantic segmentation of urban scenes using dense depth maps. In European conference on computer vision (pp. 708\u2013721), Springer.","DOI":"10.1007\/978-3-642-15561-1_51"},{"key":"1445_CR85","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Fidler, S., & Urtasun, R. (2016). Instance-level segmentation for autonomous driving with deep densely connected mrfs. In Proceedings of the conference on computer vision and pattern recognition (pp. 669\u2013677).","DOI":"10.1109\/CVPR.2016.79"},{"key":"1445_CR86","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., & Jia, J. (2017). Pyramid scene parsing network. In Proceedings of the conference on computer vision and pattern recognition (pp. 2881\u20132890).","DOI":"10.1109\/CVPR.2017.660"},{"key":"1445_CR87","unstructured":"Z\u00fcrn, J., Burgard, W., & Valada, A. (2019). Self-supervised visual terrain classification from unsupervised acoustic feature learning. arXiv preprint arXiv:1912.03227."}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-021-01445-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-021-01445-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-021-01445-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,5,5]],"date-time":"2021-05-05T06:15:47Z","timestamp":1620195347000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-021-01445-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,2,26]]},"references-count":87,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2021,5]]}},"alternative-id":["1445"],"URL":"https:\/\/doi.org\/10.1007\/s11263-021-01445-z","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,2,26]]},"assertion":[{"value":"5 April 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 January 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 February 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}