{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T07:06:20Z","timestamp":1777705580084,"version":"3.51.4"},"reference-count":15,"publisher":"SAGE Publications","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IFS"],"published-print":{"date-parts":[[2022,2,2]]},"abstract":"<jats:p>\u00a0The ubiquitous availability of cost-effective cameras has rendered large scale collection of street view data a straightforward endeavour. Yet, the effective use of these data to assist autonomous driving remains a challenge, especially lack of exploration and exploitation of stereo images with abundant perceptible depth. In this paper, we propose a novel Depth-embedded Instance Segmentation Network (DISNet) which can effectively improve the performance of instance segmentation by incorporating the depth information of stereo images. The proposed network takes binocular images as input to observe the displacement of the object and estimate the corresponding depth perception without additional supervisions. Furthermore, we introduce a new module for computing the depth cost-volume, which can be integrated with the colour cost-volume to jointly capture useful disparities of stereo images. The shared-weights structure of Siamese Network is applied to learn the intrinsic information of stereo images while reducing the computational burden. Extensive experiments have been carried out on publicly available datasets (i.e., Cityscapes and KITTI), and the obtained results clearly demonstrate the superiority in segmenting instances with different depths.<\/jats:p>","DOI":"10.3233\/jifs-202230","type":"journal-article","created":{"date-parts":[[2021,12,24]],"date-time":"2021-12-24T10:26:29Z","timestamp":1640341589000},"page":"1269-1279","source":"Crossref","is-referenced-by-count":1,"title":["Depth-embedded instance segmentation network for urban scene parsing"],"prefix":"10.1177","volume":"42","author":[{"given":"Zhifan","family":"Wang","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tong","family":"Xin","sequence":"additional","affiliation":[{"name":"Institute of Coding (IoC), School of Computing, Newcastle University, Newcastle upon Tyne, NE17RU, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shidong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Engineering, Newcastle University, Newcastle upon Tyne, NE17RU, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haofeng","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"issue":"4","key":"10.3233\/JIFS-202230_ref4","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"issue":"11","key":"10.3233\/JIFS-202230_ref11","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"The International Journal of Robotics Research"},{"key":"10.3233\/JIFS-202230_ref15","doi-asserted-by":"crossref","unstructured":"Gupta S. , Girshick R. , Arbel\u00e1ez P. and Malik J. , Learning rich features from rgb-d images for object detection and segmentation, In European conference on computer vision, 345\u2013360. Springer, (2014).","DOI":"10.1007\/978-3-319-10584-0_23"},{"issue":"3","key":"10.3233\/JIFS-202230_ref17","first-page":"226","article-title":"Image segmentation based on support vector machine","volume":"3","author":"Hai-xiang","year":"2005","journal-title":"Journal of Electronic Science and Technology"},{"issue":"2","key":"10.3233\/JIFS-202230_ref21","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1109\/TPAMI.2007.1166","article-title":"Stereo processing by semiglobal matching and mutual information","volume":"30","author":"Hirschmuller","year":"2007","journal-title":"IEEE Transactions on pattern analysis and machine intelligence"},{"key":"10.3233\/JIFS-202230_ref22","unstructured":"Kamencay P. , Breznan M. , Jarina R. , Lukac P. and Zachariasova M. , Improved depth map estimation from stereo images based on hybrid method, Radioengineering 21(1) (2012)."},{"issue":"6","key":"10.3233\/JIFS-202230_ref26","doi-asserted-by":"crossref","first-page":"3397","DOI":"10.3233\/JIFS-162254","article-title":"Study on semantic image segmentation based on convolutional neural network","volume":"33","author":"Li","year":"2017","journal-title":"Journal of Intelligent & Fuzzy Systems"},{"key":"10.3233\/JIFS-202230_ref29","unstructured":"Liaw A. , Wiener M. , et al., Classification and regression by randomforest. R news, 2(3) (2002), 18\u201322."},{"key":"10.3233\/JIFS-202230_ref30","doi-asserted-by":"crossref","unstructured":"Lin T.-Y. , Maire M. , Belongie S. , Hays J. , Perona P. , Ramanan D. , Doll\u00e1r P. and Zitnick C.L. , Microsoft coco: Common objects in context, In European conference on computer vision, pages 740\u2013755. Springer, (2014).","DOI":"10.1007\/978-3-319-10602-1_48"},{"issue":"6","key":"10.3233\/JIFS-202230_ref34","doi-asserted-by":"crossref","first-page":"4259","DOI":"10.3233\/JIFS-16653","article-title":"Decompose image into meaningful regions based on contour detector and watershed algorithm","volume":"32","author":"Luo","year":"2017","journal-title":"Journal of Intelligent & Fuzzy Systems"},{"key":"10.3233\/JIFS-202230_ref36","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1016\/j.cag.2019.09.002","article-title":"Human pose regression by combining indirect part detection and contextual information","volume":"85","author":"Luvizon","year":"2019","journal-title":"Computers & Graphics"},{"key":"10.3233\/JIFS-202230_ref40","doi-asserted-by":"crossref","unstructured":"Ramirez P.Z. , Poggi M. , Tosi F. , Mattoccia S. and Di Stefano L. , Geometry meets semantics for semi-supervised monocular depth estimation, In Asian Conference on Computer Vision, 298\u2013313. Springer, (2018).","DOI":"10.1007\/978-3-030-20893-6_19"},{"key":"10.3233\/JIFS-202230_ref46","doi-asserted-by":"crossref","unstructured":"Uhrig J. , Cordts M. , Franke U. and Brox T. , Pixel-level encoding and depth layering for instance-level semantic labeling, In German Conference on Pattern Recognition, 14\u201325. Springer, (2016).","DOI":"10.1007\/978-3-319-45886-1_2"},{"issue":"3","key":"10.3233\/JIFS-202230_ref52","doi-asserted-by":"crossref","first-page":"1853","DOI":"10.1109\/TITS.2020.3027556","article-title":"When visual disparity generation meets semantic segmentation: A mutual encouragement approach","volume":"22","author":"Zhang","year":"2020","journal-title":"IEEE Transactions on Intelligent Transportation Systems"},{"issue":"12","key":"10.3233\/JIFS-202230_ref57","doi-asserted-by":"crossref","first-page":"4643","DOI":"10.1109\/TITS.2019.2909053","article-title":"Depth embedded recurrent predictive parsing network for video scenes","volume":"20","author":"Zhou","year":"2019","journal-title":"IEEE Transactions on Intelligent Transportation Systems"}],"container-title":["Journal of Intelligent &amp; Fuzzy Systems"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/JIFS-202230","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T09:44:18Z","timestamp":1777455858000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/JIFS-202230"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,2]]},"references-count":15,"journal-issue":{"issue":"3"},"URL":"https:\/\/doi.org\/10.3233\/jifs-202230","relation":{},"ISSN":["1064-1246","1875-8967"],"issn-type":[{"value":"1064-1246","type":"print"},{"value":"1875-8967","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,2]]}}}