{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T03:16:33Z","timestamp":1787022993682,"version":"3.56.0"},"reference-count":457,"publisher":"Emerald","issue":"2-3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,10,5]]},"abstract":"<jats:p>Image segmentation is the task of associating pixels in an image with their respective object class labels. It has a wide range of applications in many industries including healthcare, transportation, robotics, fashion, home improvement, and tourism. Many deep learning-based approaches have been developed for image-level object recognition and pixel-level scene understanding \u2014 with the latter requiring a much denser annotation of scenes with a large set of objects. Extensions of image segmentation tasks include 3D and video segmentation, where units of voxels, point clouds, and video frames are classified into different objects. We use \u201cObject Segmentation\u201d to refer to the union of these segmentation tasks. In this monograph, we investigate both traditional and modern object segmentation approaches, comparing their strengths, weaknesses, and utilities. We examine in detail the wide range of deep learning-based segmentation techniques developed in recent years, provide a review of the widely used datasets and evaluation metrics, and discuss potential future research directions.<\/jats:p>","DOI":"10.1561\/0600000097","type":"journal-article","created":{"date-parts":[[2022,10,5]],"date-time":"2022-10-05T04:06:46Z","timestamp":1664942806000},"page":"111-283","source":"Crossref","is-referenced-by-count":42,"title":["A Comprehensive Review of Modern Object Segmentation Approaches"],"prefix":"10.1108","volume":"13","author":[{"given":"Yuanbo","family":"Wang","sequence":"first","affiliation":[{"name":"Invitae Corporation ,","place":["USA"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Unaiza","family":"Ahsan","sequence":"additional","affiliation":[{"name":"The Home Depot ,","place":["USA"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hanyan","family":"Li","sequence":"additional","affiliation":[{"name":"Indeed Inc. ,","place":["USA"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Matthew","family":"Hagen","sequence":"additional","affiliation":[{"name":"Amazon.com, Inc. ,","place":["USA"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2022,10,5]]},"reference":[{"key":"2026032615164764700_ref001","first-page":"71","volume-title":"International Workshop on Simulation and Synthesis in Medical Imaging","author":"Abhishek","year":"2019"},{"key":"2026032615164764700_ref002","article-title":"Slic superpixels","volume-title":"Tech. rep","author":"Achanta","year":"2010"},{"issue":"6","key":"2026032615164764700_ref003","doi-asserted-by":"crossref","first-page":"641","DOI":"10.1109\/34.295913","article-title":"Seeded region growing","volume":"16","author":"Adams","year":"1994","journal-title":"IEEE Transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref004","doi-asserted-by":"publisher","DOI":"10.1007\/s11760-022-02134-1","article-title":"Deep active contours using locally controlled distance vector flow","volume-title":"Signal, Image and Video Processing","author":"Akbarimoghaddam","year":"2022"},{"key":"2026032615164764700_ref005","doi-asserted-by":"crossref","first-page":"14410","DOI":"10.1109\/ACCESS.2018.2807385","article-title":"Threat of adversarial attacks on deep learning in computer vision: A survey","volume":"6","author":"Akhtar","year":"2018","journal-title":"Ieee Access"},{"issue":"11","key":"2026032615164764700_ref006","doi-asserted-by":"crossref","first-page":"2189","DOI":"10.1109\/TPAMI.2012.28","article-title":"Measuring the objectness of image windows","volume":"34","author":"Alexe","year":"2012","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref007","first-page":"376","volume-title":"European Conference on Computer Vision","author":"Alvarez","year":"2012"},{"issue":"5","key":"2026032615164764700_ref008","doi-asserted-by":"crossref","first-page":"898","DOI":"10.1109\/TPAMI.2010.161","article-title":"Contour detection and hierarchical image segmentation","volume":"33","author":"Arbelaez","year":"2010","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref009","first-page":"328","article-title":"Multiscale combinatorial grouping","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Arbel\u00e1ez","year":"2014"},{"key":"2026032615164764700_ref010","article-title":"Joint 2d-3d-semantic data for indoor scene understanding","volume-title":"arXiv preprint arXiv:1702.01105","author":"Armeni","year":"2017"},{"key":"2026032615164764700_ref011","first-page":"1534","article-title":"3d semantic parsing of large-scale indoor spaces","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Armeni","year":"2016"},{"issue":"12","key":"2026032615164764700_ref012","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref013","first-page":"5221","article-title":"Deep watershed transform for instance segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Bai","year":"2017"},{"key":"2026032615164764700_ref014","first-page":"9297","article-title":"Semantickitti: A dataset for semantic scene understanding of lidar sequences","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Behley","year":"2019"},{"key":"2026032615164764700_ref015","first-page":"3479","article-title":"Material recognition in the wild with the materials in context database","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Bell","year":"2015"},{"key":"2026032615164764700_ref016","first-page":"1","volume-title":"2016 IEEE winter conference on applications of computer vision (WACV)","author":"Bian","year":"2016"},{"issue":"1","key":"2026032615164764700_ref017","doi-asserted-by":"crossref","first-page":"29","DOI":"10.1016\/0734-189X(85)90002-7","article-title":"Human image understanding: Recent research and a theory","volume":"32","author":"Biederman","year":"1985","journal-title":"Computer vision, graphics, and image processing"},{"key":"2026032615164764700_ref018","first-page":"328","article-title":"Outdoor Scenes Pixel-wise Semantic Segmentation using Polarimetry and Fully Convolutional Network","volume-title":"VISIGRAPP (5: VISAPP)","author":"Blanchon","year":"2019"},{"key":"2026032615164764700_ref019","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1109\/ITSC.2019.8916853","volume-title":"2019 IEEE Intelligent Transportation Systems Conference (ITSC)","author":"Blin","year":"2019"},{"key":"2026032615164764700_ref020","first-page":"216","article-title":"A new multimodal RGB and polarimetric image dataset for road scenes analysis","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Blin","year":"2020"},{"key":"2026032615164764700_ref021","first-page":"0","article-title":"Semantic segmentation of fisheye images","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV) Workshops","author":"Blott","year":"2018"},{"key":"2026032615164764700_ref022","first-page":"9157","article-title":"Yolact: Realtime instance segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Bolya","year":"2019"},{"key":"2026032615164764700_ref023","article-title":"Yolact++: Better real-time instance segmentation","volume-title":"IEEE transactions on pattern analysis and machine intelligence","author":"Bolya","year":"2020"},{"key":"2026032615164764700_ref024","doi-asserted-by":"crossref","first-page":"401","DOI":"10.1145\/1282280.1282340","article-title":"Representing shape with a spatial pyramid kernel","volume-title":"Proceedings of the 6th ACM international conference on Image and video retrieval","author":"Bosch","year":"2007"},{"key":"2026032615164764700_ref025","article-title":"Home: A household multimodal environment","volume-title":"arXiv preprint arXiv:1711.11017","author":"Brodeur","year":"2017"},{"issue":"2","key":"2026032615164764700_ref026","doi-asserted-by":"crossref","first-page":"88","DOI":"10.1016\/j.patrec.2008.04.005","article-title":"Semantic object classes in video: A high-definition ground truth database","volume":"30","author":"Brostow","year":"2009","journal-title":"Pattern Recognition Letters"},{"key":"2026032615164764700_ref027","first-page":"3547","article-title":"Scene labeling with lstm recurrent neural networks","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Byeon","year":"2015"},{"key":"2026032615164764700_ref028","first-page":"11621","article-title":"nuscenes: A multimodal dataset for autonomous driving","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Caesar","year":"2020"},{"key":"2026032615164764700_ref029","first-page":"1209","article-title":"Coco-stuff: Thing and stuff classes in context","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Caesar","year":"2018"},{"key":"2026032615164764700_ref030","article-title":"Cascade r-cnn: High quality object detection and instance segmentation","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","author":"Cai","year":"2019"},{"issue":"6","key":"2026032615164764700_ref031","doi-asserted-by":"crossref","first-page":"679","DOI":"10.1109\/TPAMI.1986.4767851","article-title":"A computational approach to edge detection","author":"Canny","year":"1986","journal-title":"IEEE Transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref032","first-page":"1","volume-title":"European Conference on Computer Vision","author":"Cao","year":"2020"},{"key":"2026032615164764700_ref033","first-page":"7088","article-title":"ShapeConv: Shape-aware Convolutional Layer for Indoor RGB-D Semantic Segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Cao","year":"2021"},{"key":"2026032615164764700_ref034","first-page":"213","volume-title":"European Conference on Computer Vision","author":"Carion","year":"2020"},{"issue":"2","key":"2026032615164764700_ref035","doi-asserted-by":"crossref","first-page":"266","DOI":"10.1109\/83.902291","article-title":"Active contours without edges","volume":"10","author":"Chan","year":"2001","journal-title":"IEEE Transactions on image processing"},{"key":"2026032615164764700_ref036","first-page":"8915","article-title":"Deep spatio-temporal random fields for efficient video segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chandra","year":"2018"},{"key":"2026032615164764700_ref037","article-title":"Shapenet: An information-rich 3d model repository","volume-title":"arXiv preprint arXiv:1512.03012","author":"Chang","year":"2015"},{"key":"2026032615164764700_ref038","article-title":"EPSNet: Efficient Panoptic Segmentation Network with Cross-layer Attention Fusion","volume-title":"Proceedings of the Asian Conference on Computer Vision (ACCV)","author":"Chang","year":"2020"},{"key":"2026032615164764700_ref039","first-page":"6","article-title":"Learning more with less: GAN-based medical image augmentation","volume-title":"Med. Imaging Technol","author":"Changhee","year":"2019"},{"key":"2026032615164764700_ref040","first-page":"3","volume-title":"International workshop on simulation and synthesis in medical imaging","author":"Chartsias","year":"2017"},{"key":"2026032615164764700_ref041","article-title":"Unifying Instance and Panoptic Segmentation with Dynamic Rank-1 Convolutions","volume-title":"arXiv preprint arXiv:2011.09796","author":"Chen","year":"2020"},{"key":"2026032615164764700_ref042","first-page":"8573","article-title":"BlendMask: Top-down meets bottom-up for instance segmentation","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Chen","year":"2020"},{"key":"2026032615164764700_ref043","first-page":"4974","article-title":"Hybrid task cascade for instance segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen","year":"2019"},{"key":"2026032615164764700_ref044","first-page":"4545","article-title":"Semantic image segmentation with task-specific edge detection using cnns and a discriminatively trained domain transform","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Chen","year":"2016"},{"key":"2026032615164764700_ref045","first-page":"31","article-title":"Searching for efficient multiscale architectures for dense image prediction","volume-title":"Advances in neural information processing systems","author":"Chen","year":"2018"},{"key":"2026032615164764700_ref046","first-page":"4013","article-title":"Masklab: Instance segmentation by refining object detection with semantic and direction features","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Chen","year":"2018"},{"issue":"4","key":"2026032615164764700_ref047","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref048","article-title":"Rethinking atrous convolution for semantic image segmentation","volume-title":"arXiv preprint arXiv:1706.05587","author":"Chen","year":"2017"},{"key":"2026032615164764700_ref049","article-title":"Scaling Wide Residual Networks for Panoptic Segmentation","volume-title":"arXiv preprint arXiv:2011.11675","author":"Chen","year":"2020"},{"key":"2026032615164764700_ref050","first-page":"3640","article-title":"Attention to scale: Scale-aware semantic image segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Chen","year":"2016"},{"key":"2026032615164764700_ref051","first-page":"801","article-title":"Encoder-decoder with atrous separable convolution for semantic image segmentation","volume-title":"Proceedings of the European conference on computer vision (ECCV)","author":"Chen","year":"2018"},{"key":"2026032615164764700_ref052","doi-asserted-by":"crossref","first-page":"2313","DOI":"10.1109\/TIP.2021.3049332","article-title":"Spatial information guided convolution for real-time RGBD semantic segmentation","volume":"30","author":"Chen","year":"2021","journal-title":"IEEE Transactions on Image Processing"},{"key":"2026032615164764700_ref053","first-page":"e5798","article-title":"Harnessing semantic segmentation masks for accurate facial attribute editing","volume-title":"Concurrency and Computation: practice and experience","author":"Chen","year":"2020"},{"issue":"6","key":"2026032615164764700_ref054","doi-asserted-by":"crossref","first-page":"2288","DOI":"10.1109\/TCSVT.2020.3020257","article-title":"Spatialflow: Bridging all tasks for panoptic segmentation","volume":"31","author":"Chen","year":"2020","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"2026032615164764700_ref055","article-title":"PanoNet: Real-time Panoptic Segmentation through Position-Sensitive Feature Embedding","volume-title":"arXiv preprint arXiv:2008.00192","author":"Chen","year":"2020"},{"key":"2026032615164764700_ref056","first-page":"1971","article-title":"Detect what you can: Detecting and representing objects using holistic models and body parts","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Chen","year":"2014"},{"key":"2026032615164764700_ref057","first-page":"561","volume-title":"European Conference on Computer Vision","author":"Chen","year":"2020"},{"key":"2026032615164764700_ref058","first-page":"2061","article-title":"Tensormask: A foundation for dense object segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Chen","year":"2019"},{"key":"2026032615164764700_ref059","first-page":"11632","article-title":"Learning active contour models for medical image segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Chen","year":"2019"},{"key":"2026032615164764700_ref060","first-page":"12475","article-title":"Panoptic-deeplab: A simple, strong, and fast baseline for bottom-up panoptic segmentation","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Cheng","year":"2020"},{"key":"2026032615164764700_ref061","first-page":"1290","article-title":"Masked-attention mask transformer for universal image segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cheng","year":"2022"},{"key":"2026032615164764700_ref062","first-page":"17864","article-title":"Per-pixel classification is not all you need for semantic segmentation","volume":"34","author":"Cheng","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026032615164764700_ref063","first-page":"7431","article-title":"Darnet: Deep active ray network for building segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cheng","year":"2019"},{"key":"2026032615164764700_ref064","first-page":"686","article-title":"Segflow: Joint learning for video object segmentation and optical flow","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Cheng","year":"2017"},{"key":"2026032615164764700_ref065","first-page":"4433","article-title":"Sparse Instance Activation for Real-Time Instance Segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Cheng","year":"2022"},{"key":"2026032615164764700_ref066","first-page":"3029","article-title":"Localitysensitive deconvolution networks with gated fusion for rgb-d indoor semantic segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Cheng","year":"2017"},{"key":"2026032615164764700_ref067","doi-asserted-by":"crossref","first-page":"155","DOI":"10.1109\/3DV.2019.00026","volume-title":"2019 International Conference on 3D Vision (3DV)","author":"Chiang","year":"2019"},{"key":"2026032615164764700_ref068","first-page":"3075","article-title":"4d spatio-temporal convnets: Minkowski convolutional neural networks","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Choy","year":"2019"},{"key":"2026032615164764700_ref069","first-page":"1","volume-title":"2007 IEEE Conference on Computer Vision and Pattern Recognition","author":"Chum","year":"2007"},{"key":"2026032615164764700_ref070","first-page":"24","article-title":"A Robust Approach toward Feature Space Analysis","volume-title":"IEEE Trans. Patt. An. Mach. Intell","author":"Comaniciu","year":"2002"},{"key":"2026032615164764700_ref071","first-page":"3213","article-title":"The cityscapes dataset for semantic urban scene understanding","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Cordts","year":"2016"},{"key":"2026032615164764700_ref072","first-page":"5828","article-title":"Scannet: Richly-annotated 3d reconstructions of indoor scenes","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Dai","year":"2017"},{"key":"2026032615164764700_ref073","first-page":"452","article-title":"3dmv: Joint 3d-multi-view prediction for 3d semantic scene segmentation","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Dai","year":"2018"},{"key":"2026032615164764700_ref074","first-page":"4578","article-title":"Scancomplete: Large-scale scene completion and semantic segmentation for 3d scans","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Dai","year":"2018"},{"issue":"5","key":"2026032615164764700_ref075","doi-asserted-by":"crossref","first-page":"1182","DOI":"10.1007\/s11263-019-01182-4","article-title":"Curriculum model adaptation with synthetic and real data for semantic foggy scene understanding","volume":"128","author":"Dai","year":"2020","journal-title":"International Journal of Computer Vision"},{"key":"2026032615164764700_ref076","doi-asserted-by":"crossref","first-page":"3819","DOI":"10.1109\/ITSC.2018.8569387","volume-title":"2018 21st International Conference on Intelligent Transportation Systems (ITSC)","author":"Dai","year":"2018"},{"key":"2026032615164764700_ref077","first-page":"3150","article-title":"Instance-aware semantic segmentation via multi-task network cascades","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Dai","year":"2016"},{"key":"2026032615164764700_ref078","first-page":"764","article-title":"Deformable convolutional networks","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Dai","year":"2017"},{"key":"2026032615164764700_ref079","first-page":"1601","article-title":"Up-detr: Unsupervised pre-training for object detection with transformers","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Dai","year":"2021"},{"key":"2026032615164764700_ref080","first-page":"886","volume-title":"2005 IEEE computer society conference on computer vision and pattern recognition (CVPR\u201905)","author":"Dalal","year":"2005"},{"key":"2026032615164764700_ref081","article-title":"Panoptic segmentation with a joint semantic and instance segmentation network","volume-title":"arXiv preprint arXiv:1809.02110","author":"De Geus","year":"2018"},{"key":"2026032615164764700_ref082","doi-asserted-by":"crossref","first-page":"248","DOI":"10.1109\/CVPR.2009.5206848","volume-title":"2009 IEEE conference on computer vision and pattern recognition","author":"Deng","year":"2009"},{"issue":"10","key":"2026032615164764700_ref083","doi-asserted-by":"crossref","first-page":"4350","DOI":"10.1109\/TITS.2019.2939832","article-title":"Restricted deformable convolution-based road scene semantic segmentation using surround view cameras","volume":"21","author":"Deng","year":"2019","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"2026032615164764700_ref084","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1109\/IVS.2017.7995725","volume-title":"2017 IEEE Intelligent Vehicles Symposium (IV)","author":"Deng","year":"2017"},{"key":"2026032615164764700_ref085","first-page":"2393","article-title":"Context contrasted feature and gated multi-scale aggregation for scene segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Ding","year":"2018"},{"key":"2026032615164764700_ref086","first-page":"21898","article-title":"Solq: Segmenting objects by learning queries","volume":"34","author":"Dong","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026032615164764700_ref087","article-title":"An image is worth 16x16 words: Transformers for image recognition at scale","volume-title":"arXiv preprint arXiv:2010.11929","author":"Dosovitskiy","year":"2020"},{"key":"2026032615164764700_ref088","first-page":"6569","article-title":"Centernet: Keypoint triplets for object detection","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Duan","year":"2019"},{"key":"2026032615164764700_ref089","first-page":"1","volume-title":"2020 Tenth International Conference on Image Processing Theory, Tools and Applications (IPTA)","author":"Dufour","year":"2020"},{"key":"2026032615164764700_ref090","first-page":"5912","article-title":"Sstvos: Sparse spatiotemporal transformers for video object segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Duke","year":"2021"},{"key":"2026032615164764700_ref091","first-page":"6144","article-title":"Segan: Segmenting and generating the invisible","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Ehsani","year":"2018"},{"issue":"1","key":"2026032615164764700_ref092","first-page":"1997","article-title":"Neural architecture search: A survey","volume":"20","author":"Elsken","year":"2019","journal-title":"The Journal of Machine Learning Research"},{"key":"2026032615164764700_ref093","doi-asserted-by":"crossref","first-page":"9463","DOI":"10.1109\/ICRA40945.2020.9197503","volume-title":"2020 IEEE International Conference on Robotics and Automation (ICRA)","author":"Engelmann","year":"2020"},{"key":"2026032615164764700_ref094","first-page":"2147","article-title":"Scalable object detection using deep neural networks","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Erhan","year":"2014"},{"issue":"2","key":"2026032615164764700_ref095","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/s11263-009-0275-4","article-title":"The pascal visual object classes (voc) challenge","volume":"88","author":"Everingham","year":"2010","journal-title":"International journal of computer vision"},{"key":"2026032615164764700_ref096","first-page":"14504","article-title":"SCF-Net: Learning Spatial Contextual Features for Large-Scale Point Cloud Segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fan","year":"2021"},{"key":"2026032615164764700_ref097","article-title":"Semantic instance segmentation via deep metric learning","volume-title":"arXiv preprint arXiv:1703.10277","author":"Fathi","year":"2017"},{"issue":"1","key":"2026032615164764700_ref098","doi-asserted-by":"crossref","first-page":"36","DOI":"10.1109\/TPAMI.2007.1144","article-title":"Groups of adjacent contour segments for object detection","volume":"30","author":"Ferrari","year":"2007","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref099","first-page":"4083","article-title":"Learning to segment moving objects in videos","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Fragkiadaki","year":"2015"},{"key":"2026032615164764700_ref100","doi-asserted-by":"crossref","first-page":"698","DOI":"10.1109\/ISBI.2016.7493362","volume-title":"2016 IEEE 13th international symposium on biomedical imaging (ISBI)","author":"Fu","year":"2016"},{"issue":"6","key":"2026032615164764700_ref101","doi-asserted-by":"crossref","first-page":"2547","DOI":"10.1109\/TNNLS.2020.3006524","article-title":"Scene segmentation with dual relation-aware attention network","volume":"32","author":"Fu","year":"2020","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"2026032615164764700_ref102","first-page":"3146","article-title":"Dual attention network for scene segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Fu","year":"2019"},{"key":"2026032615164764700_ref103","article-title":"Stacked deconvolutional network for semantic segmentation","volume-title":"IEEE Transactions on Image Processing","author":"Fu","year":"2019"},{"key":"2026032615164764700_ref104","first-page":"4453","article-title":"Semantic video cnns through representation warping","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Gadde","year":"2017"},{"key":"2026032615164764700_ref105","doi-asserted-by":"publisher","first-page":"6013","DOI":"10.1109\/TIP.2021.3090522","article-title":"Learning Category- and Instance-Aware Pixel Embedding for Fast Panoptic Segmentation","volume":"30","author":"Gao","year":"2021","journal-title":"IEEE Transactions on Image Processing"},{"key":"2026032615164764700_ref106","article-title":"Rethink Dilated Convolution for Real-time Semantic Segmentation","volume-title":"arXiv preprint arXiv:2111.09957","author":"Gao","year":"2021"},{"issue":"2","key":"2026032615164764700_ref107","doi-asserted-by":"crossref","first-page":"3216","DOI":"10.1109\/LRA.2021.3060405","article-title":"Panoster: End-to-end panoptic segmentation of lidar point clouds","volume":"6","author":"Gasperini","year":"2021","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2026032615164764700_ref108","doi-asserted-by":"crossref","first-page":"3354","DOI":"10.1109\/CVPR.2012.6248074","volume-title":"2012 IEEE conference on computer vision and pattern recognition","author":"Geiger","year":"2012"},{"key":"2026032615164764700_ref109","article-title":"Use of the stair vision library within the ISPRS 2D semantic labeling benchmark (Vaihingen)","author":"Gerke","year":"2014"},{"issue":"2","key":"2026032615164764700_ref110","doi-asserted-by":"crossref","first-page":"1742","DOI":"10.1109\/LRA.2020.2969919","article-title":"Fast panoptic segmentation network","volume":"5","author":"Geus","year":"2020","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2026032615164764700_ref111","first-page":"5485","article-title":"Part-aware panoptic segmentation","volume-title":"Proceedings of the IEEE\/ CVF Conference on Computer Vision and Pattern Recognition","author":"Geus","year":"2021"},{"key":"2026032615164764700_ref112","article-title":"A2d2: Audi autonomous driving dataset","volume-title":"arXiv preprint arXiv:2004.06320","author":"Geyer","year":"2020"},{"key":"2026032615164764700_ref113","first-page":"1440","article-title":"Fast r-cnn","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Girshick","year":"2015"},{"key":"2026032615164764700_ref114","first-page":"580","article-title":"Rich feature hierarchies for accurate object detection and semantic segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Girshick","year":"2014"},{"key":"2026032615164764700_ref115","doi-asserted-by":"crossref","first-page":"1487","DOI":"10.1145\/3097983.3098043","article-title":"Google vizier: A service for black-box optimization","volume-title":"Proceedings of the 23rd ACM SIGKDD international conference on knowledge discovery and data mining","author":"Golovin","year":"2017"},{"key":"2026032615164764700_ref116","first-page":"932","article-title":"Look into person: Self-supervised structure-sensitive learning and a new benchmark for human parsing","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Gong","year":"2017"},{"key":"2026032615164764700_ref117","first-page":"1","volume-title":"2009 IEEE 12th international conference on computer vision","author":"Gould","year":"2009"},{"key":"2026032615164764700_ref118","first-page":"12517","article-title":"Panoptic segmentation forecasting","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Graber","year":"2021"},{"key":"2026032615164764700_ref119","first-page":"9224","article-title":"3d semantic segmentation with submanifold sparse convolutional networks","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Graham","year":"2018"},{"key":"2026032615164764700_ref120","first-page":"7157","article-title":"SOTR: Segmenting Objects with Transformers","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Guo","year":"2021"},{"key":"2026032615164764700_ref121","first-page":"5356","article-title":"LVIS: A dataset for large vocabulary instance segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Gupta","year":"2019"},{"key":"2026032615164764700_ref122","first-page":"10722","article-title":"Unsupervised microvascular image segmentation using an active contours mimicking neural network","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Gur","year":"2019"},{"key":"2026032615164764700_ref123","doi-asserted-by":"publisher","DOI":"10.5194\/isprs-annals-IV-1-W1-91-2017","article-title":"Semantic3D.net: A new Large-scale Point Cloud Classification Benchmark","volume-title":"ISPRS Annals of Photogrammetry, Remote Sensing and Spatial Information Sciences","author":"Hackel","year":"2017"},{"issue":"3","key":"2026032615164764700_ref124","doi-asserted-by":"crossref","first-page":"e0213539","DOI":"10.1371\/journal.pone.0213539","article-title":"Deep convolutional neural networks for segmenting 3D in vivo multiphoton images of vasculature in Alzheimer disease mouse models","volume":"14","author":"Haft-Javaherian","year":"2019","journal-title":"PloS one"},{"key":"2026032615164764700_ref125","doi-asserted-by":"crossref","first-page":"3675","DOI":"10.1109\/ITSC.2019.8917518","volume-title":"2019 IEEE Intelligent Transportation Systems Conference (ITSC)","author":"Hahner","year":"2019"},{"key":"2026032615164764700_ref126","doi-asserted-by":"crossref","first-page":"302","DOI":"10.1016\/j.neucom.2019.11.118","article-title":"A brief survey on semantic segmentation with deep learning","volume":"406","author":"Hao","year":"2020","journal-title":"Neurocomputing"},{"key":"2026032615164764700_ref127","first-page":"297","volume-title":"European conference on computer vision","author":"Hariharan","year":"2014"},{"key":"2026032615164764700_ref128","first-page":"98","volume-title":"International Workshop on Machine Learning in Medical Imaging","author":"Hatamizadeh","year":"2019"},{"key":"2026032615164764700_ref129","article-title":"End-to-end deep convolutional active contours for image segmentation","volume-title":"arXiv preprint arXiv:1909.13359","author":"Hatamizadeh","year":"2019"},{"key":"2026032615164764700_ref130","first-page":"730","volume-title":"European Conference on Computer Vision","author":"Hatamizadeh","year":"2020"},{"key":"2026032615164764700_ref131","first-page":"5696","article-title":"Boundary-aware instance segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hayder","year":"2017"},{"key":"2026032615164764700_ref132","first-page":"3562","article-title":"Dynamic multi-scale filters for semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"He","year":"2019"},{"key":"2026032615164764700_ref133","first-page":"7519","article-title":"Adaptive pyramid context network for semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"He","year":"2019"},{"key":"2026032615164764700_ref134","first-page":"2961","article-title":"Mask r-cnn","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"He","year":"2017"},{"key":"2026032615164764700_ref135","first-page":"770","article-title":"Deep residual learning for image recognition","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"He","year":"2016"},{"issue":"5","key":"2026032615164764700_ref136","doi-asserted-by":"crossref","first-page":"647","DOI":"10.1177\/0278364911434148","article-title":"RGB-D mapping: Using Kinect-style depth cameras for dense 3D modeling of indoor environments","volume":"31","author":"Henry","year":"2012","journal-title":"The International Journal of Robotics Research"},{"issue":"8","key":"2026032615164764700_ref137","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural computation"},{"key":"2026032615164764700_ref138","first-page":"13090","article-title":"Lidar-based panoptic segmentation via dynamic shifting network","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hong","year":"2021"},{"key":"2026032615164764700_ref139","first-page":"1335","article-title":"Conditional generative adversarial network for structured domain adaptation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Hong","year":"2018"},{"key":"2026032615164764700_ref140","article-title":"Deep dual-resolution networks for real-time and accurate semantic segmentation of road scenes","volume-title":"arXiv preprint arXiv:2101.06085","author":"Hong","year":"2021"},{"issue":"3","key":"2026032615164764700_ref141","doi-asserted-by":"crossref","first-page":"781","DOI":"10.1109\/TMI.2016.2628084","article-title":"Adaptive estimation of active contour parameters using convolutional neural networks and texture analysis","volume":"36","author":"Hoogi","year":"2016","journal-title":"IEEE transactions on medical imaging"},{"issue":"4","key":"2026032615164764700_ref142","doi-asserted-by":"crossref","first-page":"814","DOI":"10.1109\/TPAMI.2015.2465908","article-title":"What makes for effective detection proposals?","volume":"38","author":"Hosang","year":"2015","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref143","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR42600.2020.00855","article-title":"Real-Time Panoptic Segmentation From Dense Detections","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Hou","year":"2020"},{"key":"2026032615164764700_ref144","first-page":"7132","article-title":"Squeeze-and-excitation networks","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Hu","year":"2018"},{"key":"2026032615164764700_ref145","first-page":"2300","article-title":"Deep level sets for salient object detection","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Hu","year":"2017"},{"key":"2026032615164764700_ref146","first-page":"11108","article-title":"Randla-net: Efficient semantic segmentation of large-scale point clouds","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Hu","year":"2020"},{"key":"2026032615164764700_ref147","first-page":"4233","article-title":"Learning to segment every thing","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hu","year":"2018"},{"key":"2026032615164764700_ref148","first-page":"108","volume-title":"European Conference on Computer Vision","author":"Hu","year":"2016"},{"key":"2026032615164764700_ref149","doi-asserted-by":"crossref","first-page":"1440","DOI":"10.1109\/ICIP.2019.8803025","volume-title":"2019 IEEE International Conference on Image Processing (ICIP)","author":"Hu","year":"2019"},{"key":"2026032615164764700_ref150","first-page":"4700","article-title":"Densely connected convolutional networks","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Huang","year":"2017"},{"key":"2026032615164764700_ref151","first-page":"590","article-title":"Domain transfer through deep activation matching","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Huang","year":"2018"},{"key":"2026032615164764700_ref152","first-page":"10133","article-title":"Cross-View Regularization for Domain Adaptive Panoptic Segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Huang","year":"2021"},{"key":"2026032615164764700_ref153","first-page":"10133","article-title":"Cross-view regularization for domain adaptive panoptic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Huang","year":"2021"},{"key":"2026032615164764700_ref154","doi-asserted-by":"crossref","first-page":"2670","DOI":"10.1109\/ICPR.2016.7900038","volume-title":"2016 23rd International Conference on Pattern Recognition (ICPR)","author":"Huang","year":"2016"},{"key":"2026032615164764700_ref155","first-page":"2626","article-title":"Recurrent slice networks for 3d segmentation of point clouds","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Huang","year":"2018"},{"key":"2026032615164764700_ref156","first-page":"864","article-title":"FaPN: Feature-Aligned Pyramid Network for Dense Image Prediction","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Huang","year":"2021"},{"issue":"10","key":"2026032615164764700_ref157","doi-asserted-by":"crossref","first-page":"2702","DOI":"10.1109\/TPAMI.2019.2926463","article-title":"The apolloscape open dataset for autonomous driving and its application","volume":"42","author":"Huang","year":"2019","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref158","first-page":"520","article-title":"Efficient uncertainty estimation for semantic segmentation in videos","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Huang","year":"2018"},{"key":"2026032615164764700_ref159","first-page":"6409","article-title":"Mask scoring r-cnn","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Huang","year":"2019"},{"key":"2026032615164764700_ref160","first-page":"603","article-title":"Ccnet: Criss-cross attention for semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Huang","year":"2019"},{"key":"2026032615164764700_ref161","doi-asserted-by":"crossref","first-page":"2374","DOI":"10.1109\/ICIP.2019.8803360","volume-title":"2019 IEEE International Conference on Image Processing (ICIP)","author":"Hung","year":"2019"},{"key":"2026032615164764700_ref162","article-title":"Adversarial learning for semi-supervised semantic segmentation","volume-title":"arXiv preprint arXiv:1802.07934","author":"Hung","year":"2018"},{"key":"2026032615164764700_ref163","first-page":"163","volume-title":"European Conference on Computer Vision","author":"Hur","year":"2016"},{"key":"2026032615164764700_ref164","first-page":"448","volume-title":"International conference on machine learning","author":"Ioffe","year":"2015"},{"key":"2026032615164764700_ref165","doi-asserted-by":"crossref","first-page":"2117","DOI":"10.1109\/CVPR.2017.228","volume-title":"2017 IEEE conference on computer vision and pattern recognition (CVPR)","author":"Jain","year":"2017"},{"key":"2026032615164764700_ref166","doi-asserted-by":"crossref","first-page":"1421","DOI":"10.1109\/IV48863.2021.9575904","volume-title":"2021 IEEE Intelligent Vehicles Symposium (IV)","author":"Jaus","year":"2021"},{"key":"2026032615164764700_ref167","article-title":"Bipartite conditional random flelds for panoptic segmentation","volume-title":"arXiv preprint arXiv:1912.05307","author":"Jayasumana","year":"2019"},{"key":"2026032615164764700_ref168","first-page":"9865","article-title":"Invariant information clustering for unsupervised image classification and segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Ji","year":"2019"},{"key":"2026032615164764700_ref169","first-page":"5580","article-title":"Video scene parsing with predictive feature learning","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Jin","year":"2017"},{"key":"2026032615164764700_ref170","first-page":"597","volume-title":"International conference on information processing in medical imaging","author":"Kamnitsas","year":"2017"},{"issue":"4","key":"2026032615164764700_ref171","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1007\/BF00133570","article-title":"Snakes: Active contour models","volume":"1","author":"Kass","year":"1988","journal-title":"International journal of computer vision"},{"key":"2026032615164764700_ref172","doi-asserted-by":"publisher","DOI":"10.5244\/C.31.57","article-title":"Bayesian SegNet: Model Uncertainty in Deep Convolutional Encoder-Decoder Architectures for Scene Understanding","author":"Kendall","year":"2017"},{"key":"2026032615164764700_ref173","first-page":"876","article-title":"Simple does it: Weakly supervised instance and semantic segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Khoreva","year":"2017"},{"key":"2026032615164764700_ref174","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1007\/978-3-030-20870-7_8","article-title":"Video Object Segmentation with Language Referring Expressions","author":"Khoreva","year":"2019"},{"key":"2026032615164764700_ref175","first-page":"9859","article-title":"Video panoptic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Kim","year":"2020"},{"key":"2026032615164764700_ref176","first-page":"6399","article-title":"Panoptic feature pyramid networks","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Kirillov","year":"2019"},{"key":"2026032615164764700_ref177","first-page":"9404","article-title":"Panoptic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Kirillov","year":"2019"},{"key":"2026032615164764700_ref178","first-page":"109","article-title":"Efficient inference in fully connected crfs with gaussian edge potentials","volume":"24","author":"Kr\u00e4henb\u00fchl","year":"2011","journal-title":"Advances in neural information processing systems"},{"key":"2026032615164764700_ref179","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Advances in neural information processing systems"},{"key":"2026032615164764700_ref180","first-page":"3168","article-title":"Feature space optimization for semantic video segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Kundu","year":"2016"},{"key":"2026032615164764700_ref181","first-page":"4558","article-title":"Large-scale point cloud semantic segmentation with superpoint graphs","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Landrieu","year":"2018"},{"key":"2026032615164764700_ref182","first-page":"10720","article-title":"Learning instance occlusion for panoptic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lazarow","year":"2020"},{"issue":"5","key":"2026032615164764700_ref183","doi-asserted-by":"crossref","first-page":"2393","DOI":"10.1109\/TIP.2018.2794205","article-title":"Reformulating level sets as deep recurrent neural network approach to semantic segmentation","volume":"27","author":"Le","year":"2018","journal-title":"IEEE Transactions on Image Processing"},{"issue":"11","key":"2026032615164764700_ref184","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradientbased learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proceedings of the IEEE"},{"key":"2026032615164764700_ref185","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1007\/3-540-46805-6_19","volume-title":"Shape, contour and grouping in computer vision","author":"LeCun","year":"1999"},{"key":"2026032615164764700_ref186","first-page":"13906","article-title":"Centermask: Real-time anchor-free instance segmentation","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Lee","year":"2020"},{"key":"2026032615164764700_ref187","first-page":"4399","article-title":"Zero-Shot Day-Night Domain Adaptation with a Physics Prior","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Lengyel","year":"2021"},{"key":"2026032615164764700_ref188","article-title":"Pyramid attention network for semantic segmentation","volume-title":"arXiv preprint arXiv:1805.10180","author":"Li","year":"2018"},{"key":"2026032615164764700_ref189","first-page":"9522","article-title":"Dfanet: Deep feature aggregation for real-time semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li","year":"2019"},{"key":"2026032615164764700_ref190","first-page":"7274","article-title":"Motion guided attention for video salient object detection","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li","year":"2019"},{"key":"2026032615164764700_ref191","first-page":"1417","article-title":"Primary video object segmentation via complementary CNNs and neighborhood reversible flow","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Li","year":"2017"},{"key":"2026032615164764700_ref192","first-page":"9397","article-title":"So-net: Self-organizing network for point cloud analysis","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Li","year":"2018"},{"key":"2026032615164764700_ref193","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-030-01267-0_7","article-title":"Weakly- and Semi-Supervised Panoptic Segmentation","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Li","year":"2018"},{"key":"2026032615164764700_ref194","first-page":"13320","article-title":"Unifying training and inference for panoptic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li","year":"2020"},{"key":"2026032615164764700_ref195","first-page":"6526","article-title":"Instance embedding transfer to unsupervised video object segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Li","year":"2018"},{"key":"2026032615164764700_ref196","first-page":"207","article-title":"Unsupervised video object segmentation with motion-based bilateral networks","volume-title":"Proceedings of the European conference on computer vision (ECCV)","author":"Li","year":"2018"},{"key":"2026032615164764700_ref197","first-page":"9167","article-title":"Expectation-maximization attention networks for semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Li","year":"2019"},{"key":"2026032615164764700_ref198","first-page":"775","volume-title":"European Conference on Computer Vision","author":"Li","year":"2020"},{"key":"2026032615164764700_ref199","article-title":"Global aggregation then local distribution in fully convolutional networks","volume-title":"arXiv preprint arXiv:1909.07229","author":"Li","year":"2019"},{"issue":"12","key":"2026032615164764700_ref200","doi-asserted-by":"crossref","first-page":"2663","DOI":"10.1109\/TMI.2018.2845918","article-title":"H-DenseUNet: hybrid densely connected UNet for liver and tumor segmentation from CT volumes","volume":"37","author":"Li","year":"2018","journal-title":"IEEE transactions on medical imaging"},{"key":"2026032615164764700_ref201","first-page":"7026","article-title":"Attention-guided unified network for panoptic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Li","year":"2019"},{"key":"2026032615164764700_ref202","first-page":"214","article-title":"Fully Convolutional Networks for Panoptic Segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Li","year":"2021"},{"key":"2026032615164764700_ref203","article-title":"Unconventional Visual Sensors for Autonomous Vehicles","volume-title":"arXiv preprint arXiv:2205. 09383","author":"Li","year":"2022"},{"key":"2026032615164764700_ref204","first-page":"5997","article-title":"Low-latency video semantic segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Li","year":"2018"},{"key":"2026032615164764700_ref205","first-page":"541","volume-title":"European conference on computer vision","author":"Li","year":"2016"},{"issue":"12","key":"2026032615164764700_ref206","doi-asserted-by":"crossref","first-page":"2402","DOI":"10.1109\/TPAMI.2015.2408360","article-title":"Deep human parsing with active template regression","volume":"37","author":"Liang","year":"2015","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref207","first-page":"125","volume-title":"European Conference on Computer Vision","author":"Liang","year":"2016"},{"key":"2026032615164764700_ref208","first-page":"254","volume-title":"International Conference on Medical image computing and computer-assisted intervention","author":"Liao","year":"2013"},{"key":"2026032615164764700_ref209","first-page":"603","article-title":"Multiscale context intertwining for semantic segmentation","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Lin","year":"2018"},{"key":"2026032615164764700_ref210","first-page":"1925","article-title":"Refinenet: Multi-path refinement networks for high-resolution semantic segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Lin","year":"2017"},{"key":"2026032615164764700_ref211","first-page":"3194","article-title":"Efficient piecewise training of deep structured models for semantic segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Lin","year":"2016"},{"key":"2026032615164764700_ref212","first-page":"4203","article-title":"Graph-guided architecture search for real-time semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lin","year":"2020"},{"key":"2026032615164764700_ref213","first-page":"2117","article-title":"Feature pyramid networks for object detection","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Lin","year":"2017"},{"key":"2026032615164764700_ref214","first-page":"2980","article-title":"Focal loss for dense object detection","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Lin","year":"2017"},{"key":"2026032615164764700_ref215","first-page":"740","volume-title":"European conference on computer vision","author":"Lin","year":"2014"},{"issue":"2","key":"2026032615164764700_ref216","doi-asserted-by":"crossref","first-page":"359","DOI":"10.1016\/j.media.2013.12.002","article-title":"Evaluation of prostate segmentation algorithms for MRI: the PROMISE12 challenge","volume":"18","author":"Litjens","year":"2014","journal-title":"Medical image analysis"},{"issue":"12","key":"2026032615164764700_ref217","doi-asserted-by":"crossref","first-page":"2368","DOI":"10.1109\/TPAMI.2011.131","article-title":"Nonparametric scene parsing via label transfer","volume":"33","author":"Liu","year":"2011","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026032615164764700_ref218","first-page":"5678","article-title":"3DCNN-DQN-RNN: A deep reinforcement learning framework for semantic parsing of large-scale 3D point clouds","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Liu","year":"2017"},{"key":"2026032615164764700_ref219","first-page":"6172","article-title":"An end-to-end network for panoptic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu","year":"2019"},{"key":"2026032615164764700_ref220","article-title":"CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers","volume-title":"arXiv preprint arXiv:2203.04838","author":"Liu","year":"2022"},{"key":"2026032615164764700_ref221","doi-asserted-by":"crossref","first-page":"2097","DOI":"10.1109\/CVPR.2011.5995323","volume-title":"CVPR 2011","author":"Liu","year":"2011"},{"key":"2026032615164764700_ref222","first-page":"8759","article-title":"Path aggregation network for instance segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Liu","year":"2018"},{"key":"2026032615164764700_ref223","first-page":"21","volume-title":"European conference on computer vision","author":"Liu","year":"2016"},{"key":"2026032615164764700_ref224","article-title":"ParseNet: Looking Wider to See Better","author":"Liu","year":"2016"},{"issue":"01","key":"2026032615164764700_ref225","doi-asserted-by":"crossref","first-page":"8778","DOI":"10.1609\/aaai.v33i01.33018778","article-title":"Point2sequence: Learning the shape representation of 3d point clouds with an attention-based sequence to sequence network","volume":"33","author":"Liu","year":"2019","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2026032615164764700_ref226","first-page":"352","volume-title":"European Conference on Computer Vision","author":"Liu","year":"2020"},{"key":"2026032615164764700_ref227","first-page":"8895","article-title":"Relation-shape convolutional neural network for point cloud analysis","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Liu","year":"2019"},{"key":"2026032615164764700_ref228","doi-asserted-by":"publisher","first-page":"9992","DOI":"10.1109\/ICCV48922.2021.00986","article-title":"Swin Transformer: Hierarchical Vision Transformer using Shifted Windows","author":"Liu","year":"2021"},{"key":"2026032615164764700_ref229","first-page":"1377","article-title":"Semantic image segmentation via deep parsing network","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Liu","year":"2015"},{"key":"2026032615164764700_ref230","first-page":"1","article-title":"Efficient dense modules of asymmetric convolution for real-time semantic segmentation","volume-title":"Proceedings of the ACM Multimedia Asia","author":"Lo","year":"2019"},{"key":"2026032615164764700_ref231","first-page":"3431","article-title":"Fully convolutional networks for semantic segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Long","year":"2015"},{"key":"2026032615164764700_ref232","article-title":"Zeroshot video object segmentation with co-attention siamese networks","volume-title":"IEEE transactions on pattern analysis and machine intelligence","author":"Lu","year":"2020"},{"key":"2026032615164764700_ref233","article-title":"Semantic segmentation using adversarial networks","volume-title":"arXiv preprint arXiv:1611.08408","author":"Luc","year":"2016"},{"key":"2026032615164764700_ref234","first-page":"613","article-title":"Convex shape prior for multi-object segmentation using a single level set function","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Luo","year":"2019"},{"key":"2026032615164764700_ref235","first-page":"4429","article-title":"Towards Robust Semantic Segmentation of Accident Scenes via Multi-Source Mixed Sampling and Meta-Learning","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Luo","year":"2022"},{"key":"2026032615164764700_ref236","doi-asserted-by":"crossref","first-page":"2766","DOI":"10.1109\/ITSC48978.2021.9564920","volume-title":"2021 IEEE International Intelligent Transportation Systems Conference (ITSC)","author":"Ma","year":"2021"},{"key":"2026032615164764700_ref237","first-page":"598","volume-title":"2017IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Ma","year":"2017"},{"key":"2026032615164764700_ref238","doi-asserted-by":"crossref","DOI":"10.5244\/C.34.65","article-title":"Making a case for 3D convolutions for object segmentation in Videos","volume-title":"arXiv preprint arXiv:2008.11516","author":"Mahadevan","year":"2020"},{"key":"2026032615164764700_ref239","first-page":"1029","article-title":"Budget-aware deep semantic video segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Mahasseni","year":"2017"},{"issue":"2","key":"2026032615164764700_ref240","doi-asserted-by":"crossref","first-page":"158","DOI":"10.1109\/34.368173","article-title":"Shape modeling with front propagation: A level set approach","volume":"17","author":"Malladi","year":"1995","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref241","first-page":"616","article-title":"Deep extreme cut: From extreme points to object segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Maninis","year":"2018"},{"key":"2026032615164764700_ref242","first-page":"8877","article-title":"Learning deep structured active contours end-to-end","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Marcos","year":"2018"},{"key":"2026032615164764700_ref243","doi-asserted-by":"crossref","DOI":"10.1201\/9781420036114","volume-title":"Chebyshev polynomials","author":"Mason","year":"2002"},{"key":"2026032615164764700_ref244","first-page":"552","article-title":"Espnet: Efficient spatial pyramid of dilated convolutions for semantic segmentation","volume-title":"Proceedings of the european conference on computer vision (ECCV)","author":"Mehta","year":"2018"},{"issue":"10","key":"2026032615164764700_ref245","doi-asserted-by":"crossref","first-page":"1993","DOI":"10.1109\/TMI.2014.2377694","article-title":"The multimodal brain tumor image segmentation benchmark (BRATS)","volume":"34","author":"Menze","year":"2014","journal-title":"IEEE transactions on medical imaging"},{"key":"2026032615164764700_ref246","first-page":"0","article-title":"Sensor fusion for joint 3d object detection and semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Meyer","year":"2019"},{"key":"2026032615164764700_ref247","doi-asserted-by":"publisher","first-page":"8505","DOI":"10.1109\/IROS45743.2020.9340837","article-title":"LiDAR Panoptic Segmentation for Autonomous Driving","volume-title":"2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Milioto","year":"2020"},{"key":"2026032615164764700_ref248","doi-asserted-by":"crossref","first-page":"565","DOI":"10.1109\/3DV.2016.79","volume-title":"2016 fourth international conference on 3D vision (3DV)","author":"Milletari","year":"2016"},{"key":"2026032615164764700_ref249","first-page":"909","article-title":"Partnet: A large-scale benchmark for fine-grained and hierarchical part-level 3d object understanding","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Mo","year":"2019"},{"key":"2026032615164764700_ref250","doi-asserted-by":"crossref","first-page":"1965","DOI":"10.1109\/ICIP.2019.8803161","volume-title":"2019 IEEE International Conference on Image Processing (ICIP)","author":"Mohajerani","year":"2019"},{"issue":"5","key":"2026032615164764700_ref251","doi-asserted-by":"crossref","first-page":"1551","DOI":"10.1007\/s11263-021-01445-z","article-title":"Efficientps: Efficient panoptic segmentation","volume":"129","author":"Mohan","year":"2021","journal-title":"International Journal of Computer Vision"},{"key":"2026032615164764700_ref252","first-page":"1","volume-title":"2008 IEEE conference on computer vision and pattern recognition","author":"Moore","year":"2008"},{"key":"2026032615164764700_ref253","first-page":"891","article-title":"The role of context for object detection and semantic segmentation in the wild","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Mottaghi","year":"2014"},{"issue":"4","key":"2026032615164764700_ref254","first-page":"349","article-title":"A focused back-propagation algorithm for temporal pattern recognition","volume":"3","author":"Mozer","year":"1989","journal-title":"Complex systems"},{"key":"2026032615164764700_ref255","doi-asserted-by":"crossref","DOI":"10.1002\/cpa.3160420503","article-title":"Optimal approximations by piecewise smooth functions and associated variational problems","volume-title":"Communications on pure and applied mathematics","author":"Mumford","year":"1989"},{"key":"2026032615164764700_ref256","doi-asserted-by":"crossref","first-page":"2996","DOI":"10.1109\/ICIP.2019.8803299","volume-title":"2019 IEEE International Conference on Image Processing (ICIP)","author":"Nag","year":"2019"},{"key":"2026032615164764700_ref257","article-title":"Light-weight refinenet for real-time semantic segmentation","volume-title":"arXiv preprint arXiv:1810.03272","author":"Nekrasov","year":"2018"},{"key":"2026032615164764700_ref258","first-page":"4990","article-title":"The mapillary vistas dataset for semantic understanding of street scenes","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Neuhold","year":"2017"},{"key":"2026032615164764700_ref259","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1109\/RAM.2013.6758588","article-title":"3D point cloud segmentation: A survey","volume-title":"2013 6th IEEE Conference on Robotics, Automation and Mechatronics (RAM)","author":"Nguyen","year":"2013"},{"key":"2026032615164764700_ref260","first-page":"6819","article-title":"Semantic video segmentation by gated recurrent flow propagation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Nilsson","year":"2018"},{"key":"2026032615164764700_ref261","first-page":"4061","article-title":"Hyperseg: Patch-wise hypernetwork for real-time semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Nirkin","year":"2021"},{"key":"2026032615164764700_ref262","first-page":"1520","article-title":"Learning deconvolution network for semantic segmentation","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Noh","year":"2015"},{"key":"2026032615164764700_ref263","first-page":"28","article-title":"Learning to segment object candidates","volume-title":"Advances in neural information processing systems","author":"O Pinheiro","year":"2015"},{"key":"2026032615164764700_ref264","first-page":"12607","article-title":"In defense of pretrained imagenet architectures for real-time semantic segmentation of road-driving images","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Orsic","year":"2019"},{"issue":"1","key":"2026032615164764700_ref265","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1016\/0021-9991(88)90002-2","article-title":"Fronts propagating with curvaturedependent speed: Algorithms based on Hamilton-Jacobi formulations","volume":"79","author":"Osher","year":"1988","journal-title":"Journal of computational physics"},{"issue":"1","key":"2026032615164764700_ref266","doi-asserted-by":"crossref","first-page":"62","DOI":"10.1109\/TSMC.1979.4310076","article-title":"A threshold selection method from gray-level histograms","volume":"9","author":"Otsu","year":"1979","journal-title":"IEEE transactions on systems, man, and cybernetics"},{"key":"2026032615164764700_ref267","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1016\/j.patrec.2021.01.010","article-title":"Deep learning for real-time semantic segmentation: Application in ultrasound imaging","volume":"144","author":"Ouahabi","year":"2021","journal-title":"Pattern Recognition Letters"},{"key":"2026032615164764700_ref268","volume-title":"Vision science: Photons to phenomenology","author":"Palmer","year":"1999"},{"key":"2026032615164764700_ref269","first-page":"1742","article-title":"Weakly-and semi-supervised learning of a deep convolutional network for semantic image segmentation","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Papandreou","year":"2015"},{"key":"2026032615164764700_ref270","first-page":"4980","article-title":"Rdfnet: Rgb-d multilevel residual feature fusion for indoor semantic segmentation","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Park","year":"2017"},{"key":"2026032615164764700_ref271","article-title":"Enet: A deep neural network architecture for real-time semantic segmentation","volume-title":"arXiv preprint arXiv:1606.02147","author":"Paszke","year":"2016"},{"key":"2026032615164764700_ref272","article-title":"PP-LiteSeg: A Superior Real-Time Semantic Segmentation Model","volume-title":"arXiv preprint arXiv:2204.02681","author":"Peng","year":"2022"},{"key":"2026032615164764700_ref273","doi-asserted-by":"crossref","DOI":"10.1109\/TITS.2022.3145588","article-title":"MASS: Multi-attentional semantic segmentation of LiDAR data for dense top-view understanding","volume-title":"IEEE Transactions on Intelligent Transportation Systems","author":"Peng","year":"2022"},{"key":"2026032615164764700_ref274","first-page":"8533","article-title":"Deep snake for real-time instance segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Peng","year":"2020"},{"key":"2026032615164764700_ref275","first-page":"724","article-title":"A benchmark dataset and evaluation methodology for video object segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Perazzi","year":"2016"},{"key":"2026032615164764700_ref276","doi-asserted-by":"crossref","first-page":"1089","DOI":"10.1109\/WACV.2019.00121","volume-title":"2019 IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Pham","year":"2019"},{"key":"2026032615164764700_ref277","first-page":"82","volume-title":"International conference on machine learning","author":"Pinheiro","year":"2014"},{"key":"2026032615164764700_ref278","first-page":"75","volume-title":"European conference on computer vision","author":"Pinheiro","year":"2016"},{"key":"2026032615164764700_ref279","article-title":"The 2017 davis challenge on video object segmentation","volume-title":"arXiv preprint arXiv:1704.00675","author":"Pont-Tuset","year":"2017"},{"key":"2026032615164764700_ref280","first-page":"7302","article-title":"Improving Panoptic Segmentation at All Scales","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Porzi","year":"2021"},{"key":"2026032615164764700_ref281","first-page":"652","article-title":"Pointnet: Deep learning on point sets for 3d classification and segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Qi","year":"2017"},{"key":"2026032615164764700_ref282","first-page":"3014","article-title":"Amodal instance segmentation with kins dataset","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Qi","year":"2019"},{"key":"2026032615164764700_ref283","first-page":"3997","article-title":"VIP-DeepLab: Learning Visual Perception With Depth-Aware Video Panoptic Segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Qiao","year":"2021"},{"key":"2026032615164764700_ref284","doi-asserted-by":"crossref","first-page":"3619","DOI":"10.1109\/ICNC.2010.5584032","volume-title":"2010 Sixth International Conference on Natural Computation","author":"Qin","year":"2010"},{"key":"2026032615164764700_ref285","first-page":"3813","article-title":"Dense-resolution network for point cloud classification and segmentation","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision","author":"Qiu","year":"2021"},{"key":"2026032615164764700_ref286","article-title":"Geometric back-projection network for point cloud classification","volume-title":"IEEE Transactions on Multimedia","author":"Qiu","year":"2021"},{"key":"2026032615164764700_ref287","first-page":"1757","article-title":"Semantic Segmentation for Real Point Cloud Scenes via Bilateral Augmentation and Adaptive Fusion","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Qiu","year":"2021"},{"key":"2026032615164764700_ref288","article-title":"Fusion-net: A deep fully residual convolutional neural network for image segmentation in connectomics","author":"Quan","year":"2016"},{"key":"2026032615164764700_ref289","article-title":"Multi-scale convolutional architecture for semantic segmentation","volume-title":"Robotics Institute, Carnegie Mellon University, Tech. Rep. CMU-RITR-15-21","author":"Raj","year":"2015"},{"issue":"3","key":"2026032615164764700_ref290","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1080\/10095020.2020.1805366","article-title":"Homogeneous tree height derivation from tree crown delineation using Seeded Region Growing (SRG) segmentation","volume":"23","author":"Ramli","year":"2020","journal-title":"Geo-spatial Information Science"},{"key":"2026032615164764700_ref291","first-page":"779","article-title":"You only look once: Unified, real-time object detection","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Redmon","year":"2016"},{"key":"2026032615164764700_ref292","first-page":"91","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren","year":"2015","journal-title":"Advances in neural information processing systems"},{"key":"2026032615164764700_ref293","first-page":"15455","article-title":"Reciprocal Transformations for Unsupervised Video Object Segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Ren","year":"2021"},{"key":"2026032615164764700_ref294","doi-asserted-by":"crossref","first-page":"7833","DOI":"10.1109\/ICPR48806.2021.9413048","volume-title":"2020 25th International Conference on Pattern Recognition (ICPR)","author":"Riaz","year":"2021"},{"key":"2026032615164764700_ref295","first-page":"2213","article-title":"Playing for benchmarks","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Richter","year":"2017"},{"key":"2026032615164764700_ref296","first-page":"102","volume-title":"European conference on computer vision","author":"Richter","year":"2016"},{"key":"2026032615164764700_ref297","first-page":"3577","article-title":"Octnet: Learning deep 3d representations at high resolutions","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Riegler","year":"2017"},{"issue":"1","key":"2026032615164764700_ref298","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1109\/TITS.2017.2750080","article-title":"Erfnet: Efficient residual factorized convnet for real-time semantic segmentation","volume":"19","author":"Romera","year":"2017","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"key":"2026032615164764700_ref299","doi-asserted-by":"crossref","first-page":"1312","DOI":"10.1109\/IVS.2019.8813888","volume-title":"2019 IEEE Intelligent Vehicles Symposium (IV)","author":"Romera","year":"2019"},{"key":"2026032615164764700_ref300","doi-asserted-by":"crossref","first-page":"577","DOI":"10.1109\/ICMA.2014.6885761","volume-title":"2014 IEEE international conference on mechatronics and automation","author":"Rong","year":"2014"},{"key":"2026032615164764700_ref301","first-page":"234","volume-title":"International Conference on Medical image computing and computer-assisted intervention","author":"Ronneberger","year":"2015"},{"key":"2026032615164764700_ref302","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1109\/WACV.2015.38","volume-title":"2015 IEEE Winter Conference on Applications of Computer Vision","author":"Ros","year":"2015"},{"key":"2026032615164764700_ref303","first-page":"3234","article-title":"The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Ros","year":"2016"},{"issue":"3","key":"2026032615164764700_ref304","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1145\/1015706.1015720","article-title":"\u201cGrabCut\u201d interactive foreground extraction using iterated graph cuts","volume":"23","author":"Rother","year":"2004","journal-title":"ACM transactions on graphics (TOG)"},{"key":"2026032615164764700_ref305","volume-title":"Human face detection in visual scenes","author":"Rowley","year":"1995"},{"key":"2026032615164764700_ref306","first-page":"186","volume-title":"European conference on computer vision","author":"Roy","year":"2016"},{"key":"2026032615164764700_ref307","article-title":"Deep active contours","volume-title":"arXiv preprint arXiv:1607.05074","author":"Rupprecht","year":"2016"},{"key":"2026032615164764700_ref308","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TPAMI.2020.3045882","article-title":"Map-Guided Curriculum Domain Adaptation and Uncertainty-Aware Evaluation for Semantic Nighttime Image Segmentation","volume-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence","author":"Sakaridis","year":"2020"},{"key":"2026032615164764700_ref309","first-page":"7374","article-title":"Guided curriculum model adaptation and uncertainty-aware evaluation for semantic nighttime image segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Sakaridis","year":"2019"},{"key":"2026032615164764700_ref310","first-page":"687","article-title":"Model adaptation with synthetic and real data for semantic dense foggy scene understanding","volume-title":"Proceedings of the european conference on computer vision (ECCV)","author":"Sakaridis","year":"2018"},{"issue":"9","key":"2026032615164764700_ref311","doi-asserted-by":"crossref","first-page":"973","DOI":"10.1007\/s11263-018-1072-8","article-title":"Semantic foggy scene understanding with synthetic data","volume":"126","author":"Sakaridis","year":"2018","journal-title":"International Journal of Computer Vision"},{"key":"2026032615164764700_ref312","first-page":"3752","article-title":"Learning from synthetic data: Addressing domain shift for semantic segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Sankaranarayanan","year":"2018"},{"key":"2026032615164764700_ref313","first-page":"1","article-title":"Learning depth from single monocular images","volume":"18","author":"Saxena","year":"2005","journal-title":"NIPS"},{"key":"2026032615164764700_ref314","article-title":"Stacked u-nets: a no-frills approach to natural image segmentation","volume-title":"arXiv preprint arXiv:1804.10343","author":"Shah","year":"2018"},{"key":"2026032615164764700_ref315","first-page":"852","volume-title":"European Conference on Computer Vision","author":"Shelhamer","year":"2016"},{"key":"2026032615164764700_ref316","first-page":"4574","article-title":"Spsequencenet: Semantic segmentation network on 4d point clouds","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Shi","year":"2020"},{"issue":"8","key":"2026032615164764700_ref317","doi-asserted-by":"crossref","first-page":"888","DOI":"10.1109\/34.868688","article-title":"Normalized cuts and image segmentation","volume":"22","author":"Shi","year":"2000","journal-title":"IEEE Transactions on pattern analysis and machine intelligence"},{"key":"2026032615164764700_ref318","first-page":"1","volume-title":"International workshop on simulation and synthesis in medical imaging","author":"Shin","year":"2018"},{"key":"2026032615164764700_ref319","first-page":"3620","article-title":"Dag-recurrent neural networks for scene labeling","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Shuai","year":"2016"},{"key":"2026032615164764700_ref320","first-page":"746","volume-title":"European conference on computer vision","author":"Silberman","year":"2012"},{"key":"2026032615164764700_ref321","first-page":"3693","article-title":"Dynamic edge-conditioned filters in convolutional neural networks on graphs","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Simonovsky","year":"2017"},{"key":"2026032615164764700_ref322","article-title":"Very deep convolutional networks for large-scale image recognition","volume-title":"arXiv preprint arXiv:1409.1556","author":"Simonyan","year":"2014"},{"key":"2026032615164764700_ref323","first-page":"715","article-title":"Pyramid dilated deeper convlstm for video salient object detection","volume-title":"Proceedings of the European conference on computer vision (ECCV)","author":"Song","year":"2018"},{"key":"2026032615164764700_ref324","first-page":"567","article-title":"Sun rgb-d: A rgb-d scene understanding benchmark suite","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Song","year":"2015"},{"key":"2026032615164764700_ref325","first-page":"1746","article-title":"Semantic scene completion from a single depth image","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Song","year":"2017"},{"key":"2026032615164764700_ref326","first-page":"5688","article-title":"Semi supervised semantic segmentation using generative adversarial network","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Souly","year":"2017"},{"issue":"4","key":"2026032615164764700_ref327","doi-asserted-by":"crossref","first-page":"501","DOI":"10.1109\/TMI.2004.825627","article-title":"Ridge-based vessel segmentation in color images of the retina","volume":"23","author":"Staal","year":"2004","journal-title":"IEEE transactions on medical imaging"},{"key":"2026032615164764700_ref328","first-page":"7262","article-title":"Segmenter: Transformer for semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Strudel","year":"2021"},{"key":"2026032615164764700_ref329","first-page":"1","volume-title":"2021 European Conference on Mobile Robots (ECMR)","author":"Sun","year":"2021"},{"key":"2026032615164764700_ref330","first-page":"111690A","volume-title":"Artificial Intelligence and Machine Learning in Defense Applications","author":"Sun","year":"2019"},{"issue":"4","key":"2026032615164764700_ref331","doi-asserted-by":"crossref","first-page":"5558","DOI":"10.1109\/LRA.2020.3007457","article-title":"Real-time fusion network for RGB-D semantic segmentation incorporating unexpected obstacle detection for road-driving images","volume":"5","author":"Sun","year":"2020","journal-title":"IEEE Robotics and Automation Letters"},{"key":"2026032615164764700_ref332","first-page":"14454","article-title":"Sparse r-cnn: End-to-end object detection with learnable proposals","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Sun","year":"2021"},{"key":"2026032615164764700_ref333","first-page":"3611","article-title":"Rethinking transformer-based set prediction for object detection","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Sun","year":"2021"},{"key":"2026032615164764700_ref334","first-page":"1","article-title":"Going deeper with convolutions","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Szegedy","year":"2015"},{"key":"2026032615164764700_ref335","volume-title":"Computer vision: algorithms and applications","author":"Szeliski","year":"2010"},{"key":"2026032615164764700_ref336","first-page":"5229","article-title":"Gated-scnn: Gated shape cnns for semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Takikawa","year":"2019"},{"key":"2026032615164764700_ref337","article-title":"Hierarchical multi-scale attention for semantic segmentation","volume-title":"arXiv preprint arXiv:2005.10821","author":"Tao","year":"2020"},{"key":"2026032615164764700_ref338","article-title":"Deep learning convolutional networks for multiphoton microscopy vasculature segmentation","volume-title":"arXiv preprint arXiv:1606.02382","author":"Teikari","year":"2016"},{"key":"2026032615164764700_ref339","first-page":"6411","article-title":"Kpconv: Flexible and deformable convolution for point clouds","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Thomas","year":"2019"},{"key":"2026032615164764700_ref340","doi-asserted-by":"crossref","first-page":"282","DOI":"10.1007\/978-3-030-58452-8_17","volume-title":"Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part 116","author":"Tian","year":"2020"},{"key":"2026032615164764700_ref341","first-page":"9627","article-title":"Fcos: Fully convolutional one-stage object detection","volume-title":"Proceedings of the IEEE\/CVF international conference on computer vision","author":"Tian","year":"2019"},{"key":"2026032615164764700_ref342","first-page":"4481","article-title":"Learning video object segmentation with visual memory","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Tokmakov","year":"2017"},{"key":"2026032615164764700_ref343","first-page":"4489","article-title":"Learning spatiotemporal features with 3d convolutional networks","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Tran","year":"2015"},{"key":"2026032615164764700_ref344","first-page":"17","article-title":"Deep end2end voxel2voxel prediction","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition workshops","author":"Tran","year":"2016"},{"key":"2026032615164764700_ref345","article-title":"Speeding up semantic segmentation for autonomous driving","author":"Treml","year":"2016"},{"key":"2026032615164764700_ref346","first-page":"3899","article-title":"Video segmentation via object flow","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Tsai","year":"2016"},{"issue":"2","key":"2026032615164764700_ref347","doi-asserted-by":"crossref","first-page":"154","DOI":"10.1007\/s11263-013-0620-5","article-title":"Selective search for object recognition","volume":"104","author":"Uijlings","year":"2013","journal-title":"International journal of computer vision"},{"key":"2026032615164764700_ref348","doi-asserted-by":"crossref","first-page":"1743","DOI":"10.1109\/WACV.2019.00190","volume-title":"2019 IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Varma","year":"2019"},{"key":"2026032615164764700_ref349","first-page":"5998","article-title":"Attention is all you need","volume-title":"Advances in neural information processing systems","author":"Vaswani","year":"2017"},{"key":"2026032615164764700_ref350","first-page":"3224","article-title":"Gaussian conditional random field network for semantic segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Vemulapalli","year":"2016"},{"key":"2026032615164764700_ref351","article-title":"A recurrent neural network based alternative to convolutional networks","volume-title":"arXiv preprint arXiv:1505.00393","author":"Visin","year":"2015"},{"key":"2026032615164764700_ref352","first-page":"41","article-title":"Reseg: A recurrent neural network-based model for semantic segmentation","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops","author":"Visin","year":"2016"},{"key":"2026032615164764700_ref353","doi-asserted-by":"publisher","first-page":"2254","DOI":"10.1109\/ICIP42928.2021.9506731","article-title":"Temporal Memory Attention for Video Semantic Segmentation","author":"Wang","year":"2021"},{"key":"2026032615164764700_ref354","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR42600.2020.00948","article-title":"Pixel Consensus Voting for Panoptic Segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Wang","year":"2020"},{"key":"2026032615164764700_ref355","first-page":"5463","article-title":"MaXDeepLab: End-to-End Panoptic Segmentation With Mask Transformers","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Wang","year":"2021"},{"key":"2026032615164764700_ref356","first-page":"108","volume-title":"European Conference on Computer Vision","author":"Wang","year":"2020"},{"issue":"3","key":"2026032615164764700_ref357","doi-asserted-by":"crossref","first-page":"035101","DOI":"10.1117\/1.OE.61.3.035101","article-title":"High-performance panoramic annular lens design for realtime semantic segmentation on aerial imagery","volume":"61","author":"Wang","year":"2022","journal-title":"Optical Engineering"},{"key":"2026032615164764700_ref358","first-page":"1788","article-title":"Semantic part segmentation using compositional model combining shape and appearance","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Wang","year":"2015"},{"key":"2026032615164764700_ref359","doi-asserted-by":"crossref","first-page":"1451","DOI":"10.1109\/WACV.2018.00163","volume-title":"2018 IEEE winter conference on applications of computer vision (WACV)","author":"Wang","year":"2018"},{"key":"2026032615164764700_ref360","first-page":"9236","article-title":"Zeroshot video object segmentation via attentive graph neural networks","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Wang","year":"2019"},{"key":"2026032615164764700_ref361","article-title":"A survey on deep learning technique for video segmentation","volume-title":"arXiv preprint arXiv:2107.01153","author":"Wang","year":"2021"},{"key":"2026032615164764700_ref362","first-page":"649","volume-title":"European Conference on Computer Vision","author":"Wang","year":"2020"},{"key":"2026032615164764700_ref363","first-page":"17721","article-title":"Solov2: Dynamic and fast instance segmentation","volume":"33","author":"Wang","year":"2020","journal-title":"Advances in Neural information processing systems"},{"key":"2026032615164764700_ref364","doi-asserted-by":"crossref","first-page":"1860","DOI":"10.1109\/ICIP.2019.8803154","volume-title":"2019 IEEE International Conference on Image Processing (ICIP)","author":"Wang","year":"2019"},{"key":"2026032615164764700_ref365","first-page":"41","volume-title":"Chinese Conference on Pattern Recognition and Computer Vision (PRCV)","author":"Wang","year":"2019"},{"issue":"5","key":"2026032615164764700_ref366","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3326362","article-title":"Dynamic graph cnn for learning on point clouds","volume":"38","author":"Wang","year":"2019","journal-title":"Acm Transactions On Graphics (tog)"},{"key":"2026032615164764700_ref367","first-page":"7500","article-title":"Object instance annotation with deep extreme level set evolution","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang","year":"2019"},{"key":"2026032615164764700_ref368","first-page":"8476","volume-title":"2020 IEEE\/RSJ International Conference on Intel ligent Robots and Systems (IROS)","author":"Weber","year":"2020"},{"key":"2026032615164764700_ref369","doi-asserted-by":"crossref","first-page":"1169","DOI":"10.1109\/TIP.2020.3042065","article-title":"Cgnet: A light-weight context guided network for semantic segmentation","volume":"30","author":"Wu","year":"2020","journal-title":"IEEE Transactions on Image Processing"},{"key":"2026032615164764700_ref370","first-page":"15769","article-title":"Dannet: A one-stage domain adaptation network for unsupervised nighttime semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wu","year":"2021"},{"key":"2026032615164764700_ref371","first-page":"9080","article-title":"Bidirectional graph reasoning network for panoptic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wu","year":"2020"},{"key":"2026032615164764700_ref372","article-title":"Auto-Panoptic: Cooperative Multi-Component Architecture Search for Panoptic Segmentation","volume-title":"arXiv preprint arXiv:2010.16119","author":"Wu","year":"2020"},{"key":"2026032615164764700_ref373","article-title":"Building generalizable agents with a realistic and rich 3d environment","volume-title":"arXiv preprint arXiv:1801.02209","author":"Wu","year":"2018"},{"key":"2026032615164764700_ref374","unstructured":"Wu, Y., A.Kirillov, F.Massa, W.-Y.Lo, and R.Girshick. (2019). \u201cDetectron2\u201d. URL: https:\/\/github.com\/facebookresearch\/detectron2."},{"key":"2026032615164764700_ref375","first-page":"6769","article-title":"Joint multi-person pose estimation and semantic part segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Xia","year":"2017"},{"issue":"4","key":"2026032615164764700_ref376","doi-asserted-by":"crossref","first-page":"4802","DOI":"10.1364\/OE.416130","article-title":"Polarization-driven semantic segmentation via efficient attention-bridged fusion","volume":"29","author":"Xiang","year":"2021","journal-title":"Optics Express"},{"key":"2026032615164764700_ref377","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2017.XIII.013","article-title":"DA-RNN: Semantic Mapping with Data Associated Recurrent Neural Networks","author":"Xiang","year":"2017"},{"key":"2026032615164764700_ref378","first-page":"12193","article-title":"Polarmask: Single shot instance segmentation with polar representation","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Xie","year":"2020"},{"key":"2026032615164764700_ref379","first-page":"12077","article-title":"SegFormer: Simple and efficient design for semantic segmentation with transformers","volume":"34","author":"Xie","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026032615164764700_ref380","first-page":"1492","article-title":"Aggregated residual transformations for deep neural networks","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Xie","year":"2017"},{"key":"2026032615164764700_ref381","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2019.00902","article-title":"UPSNet: A Unified Panoptic Segmentation Network","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Xiong","year":"2019"},{"key":"2026032615164764700_ref382","article-title":"PIDNet: A Realtime Semantic Segmentation Network Inspired from PID Controller","volume-title":"arXiv preprint arXiv:2206.02066","author":"Xu","year":"2022"},{"key":"2026032615164764700_ref383","first-page":"3060","article-title":"End-to-end semi-supervised object detection with soft teacher","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Xu","year":"2021"},{"key":"2026032615164764700_ref384","first-page":"2970","article-title":"Deep image matting","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Xu","year":"2017"},{"key":"2026032615164764700_ref385","article-title":"Youtube-vos: A large-scale video object segmentation benchmark","volume-title":"arXiv preprint arXiv:1809.03327","author":"Xu","year":"2018"},{"key":"2026032615164764700_ref386","first-page":"6556","article-title":"Dynamic video segmentation network","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Xu","year":"2018"},{"key":"2026032615164764700_ref387","first-page":"5168","article-title":"Explicit shape encoding for real-time instance segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Xu","year":"2019"},{"issue":"3","key":"2026032615164764700_ref388","doi-asserted-by":"crossref","first-page":"383","DOI":"10.1007\/s12021-018-9377-x","article-title":"Segan: Adversarial network with multi-scale l 1 loss for medical image segmentation","volume":"16","author":"Xue","year":"2018","journal-title":"Neuroinformatics"},{"key":"2026032615164764700_ref389","doi-asserted-by":"crossref","first-page":"3570","DOI":"10.1109\/CVPR.2012.6248101","volume-title":"2012 IEEE Conference on Computer vision and pattern recognition","author":"Yamaguchi","year":"2012"},{"key":"2026032615164764700_ref390","doi-asserted-by":"crossref","first-page":"1129","DOI":"10.1109\/ROBIO54168.2021.9739390","volume-title":"2021 IEEE International Conference on Robotics and Biomimetics (ROBIO)","author":"Yan","year":"2021"},{"key":"2026032615164764700_ref391","doi-asserted-by":"crossref","first-page":"7141","DOI":"10.1109\/TIP.2020.2998981","article-title":"Convexity shape prior for level set-based image segmentation method","volume":"29","author":"Yan","year":"2020","journal-title":"IEEE Transactions on Image Processing"},{"key":"2026032615164764700_ref392","first-page":"5589","article-title":"Pointasnl: Robust point clouds processing using nonlocal neural networks with adaptive sampling","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yan","year":"2020"},{"key":"2026032615164764700_ref393","doi-asserted-by":"crossref","first-page":"2159","DOI":"10.1109\/ICIP.2017.8296664","volume-title":"2017 IEEE International Conference on Image Processing (ICIP)","author":"Yang","year":"2017"},{"key":"2026032615164764700_ref394","doi-asserted-by":"crossref","first-page":"446","DOI":"10.1109\/IVS.2019.8814042","volume-title":"2019 IEEE Intelligent Vehicles Symposium (IV)","author":"Yang","year":"2019"},{"key":"2026032615164764700_ref395","doi-asserted-by":"crossref","first-page":"457","DOI":"10.1109\/IV47402.2020.9304706","volume-title":"2020 IEEE Intelligent Vehicles Symposium (IV)","author":"Yang","year":"2020"},{"key":"2026032615164764700_ref396","article-title":"Omnisupervised omnidirectional semantic segmentation","volume-title":"IEEE Transactions on Intelligent Transportation Systems","author":"Yang","year":"2020"},{"key":"2026032615164764700_ref397","doi-asserted-by":"crossref","first-page":"1866","DOI":"10.1109\/TIP.2020.3048682","article-title":"Is context-aware cnn ready for the surroundings? panoramic semantic segmentation in the wild","volume":"30","author":"Yang","year":"2021","journal-title":"IEEE Transactions on Image Processing"},{"key":"2026032615164764700_ref398","first-page":"1376","article-title":"Capturing omni-range context for omnidirectional segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Yang","year":"2021"},{"key":"2026032615164764700_ref399","first-page":"3684","article-title":"Denseaspp for semantic segmentation in street scenes","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Yang","year":"2018"},{"key":"2026032615164764700_ref400","doi-asserted-by":"publisher","first-page":"3753","DOI":"10.1038\/s41598-020-60520-6","article-title":"MRI Cross-Modality Image-to-Image Translation","volume":"10","author":"Yang","year":"2020","journal-title":"Scientific Reports"},{"key":"2026032615164764700_ref401","article-title":"Tracking Instances as Queries","volume-title":"arXiv preprint arXiv:2106.11963","author":"Yang","year":"2021"},{"key":"2026032615164764700_ref402","article-title":"Deeperlab: Singleshot image parser","volume-title":"arXiv preprint arXiv:1902.05093","author":"Yang","year":"2019"},{"issue":"07","key":"2026032615164764700_ref403","doi-asserted-by":"crossref","first-page":"12637","DOI":"10.1609\/aaai.v34i07.6955","article-title":"Sognet: Scene overlap graph network for panoptic segmentation","volume":"34","author":"Yang","year":"2020","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2026032615164764700_ref404","doi-asserted-by":"crossref","first-page":"648","DOI":"10.1109\/SMC42975.2020.9283099","volume-title":"2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC)","author":"Ye","year":"2020"},{"issue":"6","key":"2026032615164764700_ref405","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2980179.2980238","article-title":"A scalable active framework for region annotation in 3d shape collections","volume":"35","author":"Yi","year":"2016","journal-title":"ACM Transactions on Graphics (ToG)"},{"key":"2026032615164764700_ref406","first-page":"191","volume-title":"European Conference on Computer Vision","author":"Yin","year":"2020"},{"key":"2026032615164764700_ref407","doi-asserted-by":"publisher","first-page":"626","DOI":"10.1109\/ISBI.2018.8363653","article-title":"3D cGAN based cross-modality MR image synthesis for brain tumor segmentation","author":"Yu","year":"2018"},{"key":"2026032615164764700_ref408","first-page":"325","article-title":"Bisenet: Bilateral segmentation network for real-time semantic segmentation","volume-title":"Proceedings of the European conference on computer vision (ECCV)","author":"Yu","year":"2018"},{"key":"2026032615164764700_ref409","first-page":"1857","article-title":"Learning a discriminative feature network for semantic segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Yu","year":"2018"},{"key":"2026032615164764700_ref410","first-page":"2636","article-title":"Bdd100k: A diverse driving dataset for heterogeneous multitask learning","volume-title":"Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","author":"Yu","year":"2020"},{"key":"2026032615164764700_ref411","article-title":"Multi-scale context aggregation by dilated convolutions","volume-title":"arXiv preprint arXiv:1511.07122","author":"Yu","year":"2015"},{"key":"2026032615164764700_ref412","doi-asserted-by":"crossref","first-page":"1883","DOI":"10.1109\/WACV45572.2020.9093264","volume-title":"2020 IEEE Winter Conference on Applications of Computer Vision (WACV)","author":"Yuan","year":"2020"},{"key":"2026032615164764700_ref413","article-title":"Segmentation transformer: Object-contextual representations for semantic segmentation","volume":"1","author":"Yuan","year":"2021","journal-title":"European Conference on Computer Vision (ECCV)"},{"key":"2026032615164764700_ref414","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1007\/978-3-030-58539-6_11","volume-title":"Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part VI 16","author":"Yuan","year":"2020"},{"key":"2026032615164764700_ref415","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-021-01465-9","article-title":"OCNet: Object Context for Semantic Segmentation","volume-title":"International Journal of Computer Vision","author":"Yuan","year":"2021"},{"key":"2026032615164764700_ref416","doi-asserted-by":"publisher","first-page":"15.1","DOI":"10.5244\/C.30.15","article-title":"A MultiPath Network for Object Detection","author":"Zagoruyko","year":"2016"},{"key":"2026032615164764700_ref417","unstructured":"Zech, J.\n           (2018). \u201cWhat are radiological deep learning models actually learning\u201d. Medium. URL: https:\/\/medium.com\/@jrzech\/what-are-radiological-deep-learning-models-actually-learning-f97a546c5b98."},{"key":"2026032615164764700_ref418","first-page":"402","article-title":"Wilddash-creating hazard-aware benchmarks","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Zendel","year":"2018"},{"key":"2026032615164764700_ref419","doi-asserted-by":"crossref","first-page":"1386","DOI":"10.1109\/ICRA.2017.7989165","volume-title":"2017 IEEE international conference on robotics and automation (ICRA)","author":"Zeng","year":"2017"},{"key":"2026032615164764700_ref420","article-title":"Ada-Segment: Automated Multi-loss Adaptation for Panoptic Segmentation","volume-title":"arXiv preprint arXiv:2012.03603","author":"Zhang","year":"2020"},{"key":"2026032615164764700_ref421","doi-asserted-by":"crossref","first-page":"658","DOI":"10.1109\/LSP.2021.3066071","article-title":"Nonlocal aggregation for RGB-D semantic segmentation","volume":"28","author":"Zhang","year":"2021","journal-title":"IEEE Signal Processing Letters"},{"key":"2026032615164764700_ref422","first-page":"7151","article-title":"Context encoding for semantic segmentation","volume-title":"Proceedings of the IEEE conference on Computer Vision and Pattern Recognition","author":"Zhang","year":"2018"},{"key":"2026032615164764700_ref423","article-title":"Transfer beyond the Field of View: Dense Panoramic Semantic Segmentation via Unsupervised Domain Adaptation","volume-title":"IEEE Transactions on Intelligent Transportation Systems","author":"Zhang","year":"2021"},{"key":"2026032615164764700_ref424","first-page":"1760","article-title":"Trans4Trans: Efficient Transformer for Transparent Object Segmentation To Help Visually Impaired People Navigate in the Real World","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) Workshops","author":"Zhang","year":"2021"},{"key":"2026032615164764700_ref425","first-page":"16917","article-title":"Bending reality: Distortion-aware transformers for adapting to panoramic semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang","year":"2022"},{"key":"2026032615164764700_ref426","article-title":"Exploring Event-Driven Dynamic Context for Accident Scene Segmentation","volume-title":"IEEE Transactions on Intelligent Transportation Systems","author":"Zhang","year":"2021"},{"key":"2026032615164764700_ref427","doi-asserted-by":"crossref","first-page":"490","DOI":"10.1007\/978-3-030-58568-6_29","volume-title":"Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part XIV 16","author":"Zhang","year":"2020"},{"key":"2026032615164764700_ref428","doi-asserted-by":"crossref","first-page":"1850","DOI":"10.1109\/ICRA.2015.7139439","volume-title":"2015 IEEE international conference on robotics and automation (ICRA)","author":"Zhang","year":"2015"},{"key":"2026032615164764700_ref429","first-page":"10226","article-title":"Mask encoding for single shot instance segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang","year":"2020"},{"key":"2026032615164764700_ref430","article-title":"K-Net: Towards Unified Image Segmentation","volume-title":"arXiv preprint arXiv:2106.14855","author":"Zhang","year":"2021"},{"key":"2026032615164764700_ref431","first-page":"13956","article-title":"Dcnas: Densely connected neural architecture search for semantic image segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang","year":"2021"},{"key":"2026032615164764700_ref432","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR42600.2020.00962","article-title":"PolarNet: An Improved Grid Representation for Online LiDAR Point Clouds Semantic Segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhang","year":"2020"},{"key":"2026032615164764700_ref433","first-page":"269","article-title":"Exfuse: Enhancing feature fusion for semantic segmentation","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Zhang","year":"2018"},{"key":"2026032615164764700_ref434","first-page":"4106","article-title":"Pattern-affinitive propagation across depth, surface normal and semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhang","year":"2019"},{"key":"2026032615164764700_ref435","first-page":"9242","article-title":"Translating and segmenting multimodal medical volumes with cycle-and shape-consistency generative adversarial network","volume-title":"Proceedings of the IEEE conference on computer vision and pattern Recognition","author":"Zhang","year":"2018"},{"key":"2026032615164764700_ref436","first-page":"16259","article-title":"Point transformer","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision","author":"Zhao","year":"2021"},{"key":"2026032615164764700_ref437","first-page":"405","article-title":"Icnet for real-time semantic segmentation on high-resolution images","volume-title":"Proceedings of the European conference on computer vision (ECCV)","author":"Zhao","year":"2018"},{"key":"2026032615164764700_ref438","first-page":"2881","article-title":"Pyramid scene parsing network","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Zhao","year":"2017"},{"key":"2026032615164764700_ref439","first-page":"267","article-title":"Psanet: Point-wise spatial attention network for scene parsing","volume-title":"Proceedings of the European Conference on Computer Vision (ECCV)","author":"Zhao","year":"2018"},{"key":"2026032615164764700_ref440","first-page":"1009","article-title":"3D point capsule networks","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhao","year":"2019"},{"key":"2026032615164764700_ref441","first-page":"445","volume-title":"European Conference on Computer Vision","author":"Zhen","year":"2020"},{"key":"2026032615164764700_ref442","article-title":"End-to-end object detection with adaptive clustering transformer","volume-title":"arXiv preprint arXiv:2011.09315","author":"Zheng","year":"2020"},{"key":"2026032615164764700_ref443","first-page":"1529","article-title":"Conditional random fields as recurrent neural networks","volume-title":"Proceedings of the IEEE international conference on computer vision","author":"Zheng","year":"2015"},{"key":"2026032615164764700_ref444","first-page":"6881","article-title":"Rethinking Semantic Segmentation From a Sequence-to-Sequence Perspective With Transformers","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zheng","year":"2021"},{"key":"2026032615164764700_ref445","first-page":"13065","article-title":"Squeeze-and-attention networks for semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhong","year":"2020"},{"key":"2026032615164764700_ref446","first-page":"633","article-title":"Scene parsing through ade20k dataset","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Zhou","year":"2017"},{"issue":"07","key":"2026032615164764700_ref447","doi-asserted-by":"crossref","first-page":"13066","DOI":"10.1609\/aaai.v34i07.7008","article-title":"Motion-attentive transition for zero-shot video object segmentation","volume":"34","author":"Zhou","year":"2020","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"2026032615164764700_ref448","first-page":"850","article-title":"Bottom-up object detection by grouping extreme and center points","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhou","year":"2019"},{"key":"2026032615164764700_ref449","first-page":"13194","article-title":"Panoptic-PolarNet: Proposal-Free LiDAR Point Cloud Panoptic Segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhou","year":"2021"},{"key":"2026032615164764700_ref450","first-page":"9939","article-title":"Cylindrical and asymmetrical 3d convolution networks for lidar segmentation","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Zhu","year":"2021"},{"key":"2026032615164764700_ref451","article-title":"Deformable detr: Deformable transformers for end-to-end object detection","volume-title":"arXiv preprint arXiv:2010.0f159","author":"Zhu","year":"2020"},{"key":"2026032615164764700_ref452","first-page":"2349","article-title":"Deep feature flow for video recognition","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Zhu","year":"2017"},{"key":"2026032615164764700_ref453","first-page":"1464","article-title":"Semantic amodal segmentation","volume-title":"Proceedings of the IEEE conference on computer vision and pattern recognition","author":"Zhu","year":"2017"},{"key":"2026032615164764700_ref454","first-page":"0","article-title":"Shelfnet for fast semantic segmentation","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops","author":"Zhuang","year":"2019"},{"key":"2026032615164764700_ref455","first-page":"391","volume-title":"European conference on computer vision","author":"Zitnick","year":"2014"},{"key":"2026032615164764700_ref456","first-page":"3833","article-title":"Rethinking pre-training and self-training","volume":"33","author":"Zoph","year":"2020","journal-title":"Advances in neural information processing systems"},{"key":"2026032615164764700_ref457","article-title":"Neural architecture search with reinforcement learning","volume-title":"arXiv preprint arXiv:1611.01578","author":"Zoph","year":"2016"}],"container-title":["Foundations and Trends\u00ae in Computer Graphics and Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftcgv\/article-pdf\/13\/2-3\/111\/11097769\/0600000097en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftcgv\/article-pdf\/13\/2-3\/111\/11097769\/0600000097en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T14:09:09Z","timestamp":1777471749000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftcgv\/article\/13\/2-3\/111\/1330342\/A-Comprehensive-Review-of-Modern-Object"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,5]]},"references-count":457,"journal-issue":{"issue":"2-3","published-print":{"date-parts":[[2022,10,5]]}},"URL":"https:\/\/doi.org\/10.1561\/0600000097","relation":{},"ISSN":["1572-2740","1572-2759"],"issn-type":[{"value":"1572-2740","type":"print"},{"value":"1572-2759","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,10,5]]}}}