{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T17:50:28Z","timestamp":1779126628384,"version":"3.51.4"},"reference-count":45,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2020,1,23]],"date-time":"2020-01-23T00:00:00Z","timestamp":1579737600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61571390"],"award-info":[{"award-number":["61571390"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100017054","name":"NSFC-Zhejiang Joint Fund for the Integration of Industrialization and Informatization","doi-asserted-by":"publisher","award":["U1709214"],"award-info":[{"award-number":["U1709214"]}],"id":[{"id":"10.13039\/100017054","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>As the core task of scene understanding, semantic segmentation and depth completion play a vital role in lots of applications such as robot navigation, AR\/VR and autonomous driving. They are responsible for parsing scenes from the angle of semantics and geometry, respectively. While great progress has been made in both tasks through deep learning technologies, few works have been done on building a joint model by deeply exploring the inner relationship of the above tasks. In this paper, semantic segmentation and depth completion are jointly considered under a multi-task learning framework. By sharing a common encoder part and introducing boundary features as inner constraints in the decoder part, the two tasks can properly share the required information from each other. An extra boundary detection sub-task is responsible for providing the boundary features and constructing cross-task joint loss functions for network training. The entire network is implemented end-to-end and evaluated with both RGB and sparse depth input. Experiments conducted on synthesized and real scene datasets show that our proposed multi-task CNN model can effectively improve the performance of every single task.<\/jats:p>","DOI":"10.3390\/s20030635","type":"journal-article","created":{"date-parts":[[2020,1,23]],"date-time":"2020-01-23T10:36:02Z","timestamp":1579775762000},"page":"635","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["Simultaneous Semantic Segmentation and Depth Completion with Constraint of Boundary"],"prefix":"10.3390","volume":"20","author":[{"given":"Nan","family":"Zou","sequence":"first","affiliation":[{"name":"College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3329-7037","authenticated-orcid":false,"given":"Zhiyu","family":"Xiang","sequence":"additional","affiliation":[{"name":"Zhejiang Provincial Key Laboratory of Information Processing, Communication and Networking, Zhejiang University, Hangzhou 310000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yiman","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuya","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chengyu","family":"Qiao","sequence":"additional","affiliation":[{"name":"College of Information Science and Electronic Engineering, Zhejiang University, Hangzhou 310027, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,1,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (2016, January 27\u201330). You only look once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.91"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201328). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Ladick\u00fd, L.U., Russell, C., Kohli, P., and Torr, P.H. (October, January 29). Associative hierarchical crfs for object class image segmentation. Proceedings of the 2009 IEEE 12th International Conference on Computer Vision, Kyoto, Japan.","DOI":"10.1109\/ICCV.2009.5459248"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Noh, H., Hong, S., and Han, B. (2015, January 7\u201313). Learning deconvolution network for semantic segmentation. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.178"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"2481","DOI":"10.1109\/TPAMI.2016.2644615","article-title":"Segnet: A deep convolutional encoder-decoder architecture for image segmentation","volume":"39","author":"Badrinarayanan","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_8","unstructured":"Chen, L., Papandreou, G., Kokkinos, I., Murphy, K., and Yuille, A.L. (2014). Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1226","DOI":"10.1109\/TPAMI.2002.1033214","article-title":"Depth estimation from image structure","volume":"24","author":"Torralba","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Liu, B., Gould, S., and Koller, D. (2010, January 13\u201318). Single image depth estimation from predicted semantic labels. Proceedings of the 2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, San Francisco, CA, USA.","DOI":"10.1109\/CVPR.2010.5539823"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Hua, J., and Gong, X. (2018, January 13\u201319). A Normalized Convolutional Neural Network for Guided Sparse Depth Upsampling. Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, Stockholm, Sweden.","DOI":"10.24963\/ijcai.2018\/316"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Uhrig, J., Schneider, N., Schneider, L., Franke, U., Brox, T., and Geiger, A. (2017, January 10\u201312). Sparsity invariant cnns. Proceedings of the 2017 International Conference on 3D Vision (3DV), Qingdao, China.","DOI":"10.1109\/3DV.2017.00012"},{"key":"ref_13","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs","volume":"40","author":"Chen","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1109\/TITS.2017.2750080","article-title":"Erfnet: Efficient residual factorized convnet for real-time semantic segmentation","volume":"19","author":"Romera","year":"2017","journal-title":"IEEE Trans. Intell. Trans. Syst."},{"key":"ref_17","unstructured":"Couprie, C., Farabet, C., Najman, L., and LeCun, Y. (2013). Indoor semantic segmentation using depth information. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Gupta, S., Girshick, R., Arbel\u00e1ez, P., and Malik, J. (2014). Learning Rich Features from RGB-D Images for Object Detection and Segmentation. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-10584-0_23"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Wang, W., and Neumann, U. (2018, January 8\u201314). Depth-aware CNN for RGB-D segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_9"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"114520","DOI":"10.1109\/ACCESS.2019.2935816","article-title":"Boundary-Aware CNN for Semantic Segmentation","volume":"7","author":"Zou","year":"2019","journal-title":"IEEE Access"},{"key":"ref_21","unstructured":"Eigen, D., Puhrsch, C., and Fergus, R. (2014, January 8\u201313). Depth map prediction from a single image using a multi-scale deep network. Proceedings of the 27th International Conference on Neural Information Processing Systems\u2014Volume 2 (NIPS\u201914), Montreal, QC, Canada."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Eigen, D., and Fergus, R. (2015, January 7\u201313). Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.304"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ku, J., Harakeh, A., and Waslander, S.L. (2018, January 9\u201311). In defense of classical image processing: Fast depth completion on the cpu. Proceedings of the 2018 15th Conference on Computer and Robot Vision(CRV), Toronto, ON, Canada.","DOI":"10.1109\/CRV.2018.00013"},{"key":"ref_24","unstructured":"Hazirbas, C., Ma, L., Domokos, C., and Cremers, D. (2016). Fusenet: Incorporating Depth into Semantic Segmentation via Fusion-Based CNN Architecture. Asian Conference on Computer Vision, Springer."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Mal, F., and Karaman, S. (2018, January 21\u201325). Sparse-to-dense: Depth prediction from sparse depth samples and a single image. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, QLD, Australia.","DOI":"10.1109\/ICRA.2018.8460184"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Jaritz, M., De Charette, R., Wirbel, E., Perrotton, X., and Nashashibi, F. (2018, January 5\u20138). Sparse and dense data with cnns: Depth completion and semantic segmentation. Proceedings of the 2018 International Conference on 3D Vision (3DV), Verona, Italy.","DOI":"10.1109\/3DV.2018.00017"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zoph, B., Vasudevan, V., Shlens, J., and Le, Q.V. (2018, January 18\u201323). Learning transferable architectures for scalable image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00907"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zhang, Y., and Funkhouser, T. (2018, January 18\u201323). Deep depth completion of a single RGB-D image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00026"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Qiu, J., Cui, Z., Zhang, Y., Zhang, X., Liu, S., Zeng, B., and Pollefeys, M. (2019, January 15\u201320). Deeplidar: Deep surface normal guided depth prediction for outdoor scene from sparse lidar data and single color image. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00343"},{"key":"ref_30","unstructured":"Murphy, K.P., Torralba, A., and Freeman, W.T. (2004). Using the forest to see the trees: A graphical model relating features, objects, and scenes. Advances in Neural Information Processing Systems, MIT Press."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Teichmann, M., Weber, M., Zoellner, M., Cipolla, R., and Urtasun, R. (2018, January 26\u201330). Multinet: Real-time joint semantic reasoning for autonomous driving. Proceedings of the 2018 IEEE Intelligent Vehicles Symposium (IV), Changshu, China.","DOI":"10.1109\/IVS.2018.8500504"},{"key":"ref_32","unstructured":"Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., and LeCun, Y. (2013). Overfeat: Integrated recognition, localization and detection using convolutional networks. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Misra, I., Shrivastava, A., Gupta, A., and Hebert, M. (2016, January 27\u201330). Cross-stitch networks for multi-task learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.433"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Uhrig, J., Cordts, M., Franke, U., and Brox, T. (2016). Pixel-level Encoding and Depth Layering for Instance-Level Semantic Labeling. German Conference on Pattern Recognition, Springer.","DOI":"10.1007\/978-3-319-45886-1_2"},{"key":"ref_35","unstructured":"Kendall, A., Gal, Y., and Cipolla, R. (2018, January 18\u201323). Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Gaidon, A., Wang, Q., Cabon, Y., and Vig, E. (2016, January 27\u201330). Virtual Worlds as proxy for multi-object tracking analysis. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.470"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B. (2016, January 27\u201330). The cityscapes dataset for semantic urban scene understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.350"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"600","DOI":"10.1109\/TIP.2003.819861","article-title":"Image quality assessment: From error visibility to structural similarity","volume":"13","author":"Wang","year":"2004","journal-title":"IEEE Trans. Image Process."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Harrison, A., and Newman, P. (2010). Image and Sparse Laser Fusion for Dense Scene Reconstruction. Field and Service Robotics, Springer.","DOI":"10.1007\/978-3-642-13408-1_20"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Ferstl, D., Reinbacher, C., Ranftl, R., R\u00fcther, M., and Bischof, H. (2013, January 1\u20138). Image guided depth upsampling using anisotropic total generalized variation. Proceedings of the IEEE International Conference on Computer Vision, Sydney, NSW, Australia.","DOI":"10.1109\/ICCV.2013.127"},{"key":"ref_41","unstructured":"Paszke, A., Chaurasia, A., Kim, S., and Culurciello, E. (2016). Enet: A deep neural network architecture for real-time semantic segmentation. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Berman, M., Rannen Triki, A., and Blaschko, M.B. (2018, January 18\u201323). The Lov\u00e1sz-Softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00464"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Mehta, S., Rastegari, M., Caspi, A., Shapiro, L., and Hajishirzi, H. (2018, January 19). Espnet: Efficient spatial pyramid of dilated convolutions for semantic segmentation. Proceedings of the European Conference on Computer Vision (ECCV), Seattle, WA, USA.","DOI":"10.1007\/978-3-030-01249-6_34"},{"key":"ref_44","unstructured":"Poudel, R.P., Liwicki, S., and Cipolla, R. (2019). Fast-SCNN: Fast semantic segmentation network. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Pilzer, A., Xu, D., Puscas, M., Ricci, E., and Sebe, N. (2018, January 5\u20138). Unsupervised adversarial depth estimation using cycled generative networks. Proceedings of the 2018 International Conference on 3D Vision (3DV), Verona, Italy.","DOI":"10.1109\/3DV.2018.00073"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/3\/635\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T13:20:29Z","timestamp":1760361629000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/3\/635"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,1,23]]},"references-count":45,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2020,2]]}},"alternative-id":["s20030635"],"URL":"https:\/\/doi.org\/10.3390\/s20030635","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,1,23]]}}}