{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T03:26:05Z","timestamp":1783740365388,"version":"3.55.0"},"reference-count":95,"publisher":"Springer Science and Business Media LLC","issue":"6","license":[{"start":{"date-parts":[[2022,4,28]],"date-time":"2022-04-28T00:00:00Z","timestamp":1651104000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,4,28]],"date-time":"2022-04-28T00:00:00Z","timestamp":1651104000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000287","name":"Royal Academy of Engineering","doi-asserted-by":"crossref","award":["RF-201718-17177"],"award-info":[{"award-number":["RF-201718-17177"]}],"id":[{"id":"10.13039\/501100000287","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100000266","name":"Engineering and Physical Sciences Research Council","doi-asserted-by":"publisher","award":["EP\/P022529"],"award-info":[{"award-number":["EP\/P022529"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2022,6]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We introduce the first approach to solve the challenging problem of automatic 4D visual scene understanding for complex dynamic scenes with multiple interacting people from multi-view video. Our approach simultaneously estimates a detailed model that includes a per-pixel semantically and temporally coherent reconstruction, together with instance-level segmentation exploiting photo-consistency, semantic and motion information. We further leverage recent advances in 3D pose estimation to constrain the joint semantic instance segmentation and 4D temporally coherent reconstruction. This enables per person semantic instance segmentation of multiple interacting people in complex dynamic scenes. Extensive evaluation of the joint visual scene understanding framework against state-of-the-art methods on challenging indoor and outdoor sequences demonstrates a significant (<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\approx 40\\%$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mrow>\n                    <mml:mo>\u2248<\/mml:mo>\n                    <mml:mn>40<\/mml:mn>\n                    <mml:mo>%<\/mml:mo>\n                  <\/mml:mrow>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>) improvement in semantic segmentation, reconstruction and scene flow accuracy. In addition to the evaluation on several indoor and outdoor scenes, the proposed joint 4D scene understanding framework is applied to challenging outdoor sports scenes in the wild captured with manually operated wide-baseline broadcast cameras.<\/jats:p>","DOI":"10.1007\/s11263-022-01599-4","type":"journal-article","created":{"date-parts":[[2022,4,28]],"date-time":"2022-04-28T18:17:15Z","timestamp":1651169835000},"page":"1583-1606","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["4D Temporally Coherent Multi-Person Semantic Reconstruction and Segmentation"],"prefix":"10.1007","volume":"130","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1779-2775","authenticated-orcid":false,"given":"Armin","family":"Mustafa","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chris","family":"Russell","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Adrian","family":"Hilton","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,4,28]]},"reference":[{"key":"1599_CR1","unstructured":"4d repository, http:\/\/4drepository.inrialpes.fr\/. In: Institut national de recherche en informatique et en automatique (INRIA) Rhone Alpes."},{"key":"1599_CR2","unstructured":"Multiview video repository, http:\/\/cvssp.org\/data\/cvssp3d\/. In: Centre for Vision Speech and Signal Processing, University of Surrey, UK."},{"key":"1599_CR3","doi-asserted-by":"crossref","unstructured":"Kundu, A., Yin, X., Fathi, A., Ross, D., Brewington, B., Funkhouser, T., & Pantofaru, C. (2020). Virtual multi-view fusion for 3d semantic segmentation. In: ECCV.","DOI":"10.1007\/978-3-030-58586-0_31"},{"key":"1599_CR4","unstructured":"Gilbert, A.,\u00a0Trumble, M., Hilton, A. & Collomosse, J. (2020) Semantic estimation of 3d body shape and pose using minimal cameras. In: BMVC."},{"key":"1599_CR5","doi-asserted-by":"crossref","unstructured":"Badrinarayanan, V., Kendall, A., Cipolla, R. (2017). Segnet: A deep convolutional encoder-decoder architecture for image segmentation. TPAMI.","DOI":"10.1109\/TPAMI.2016.2644615"},{"key":"1599_CR6","doi-asserted-by":"crossref","unstructured":"Ballan, L., Brostow, G. J., Puwein, J., & Pollefeys, M. (2010). Unstructured video-based rendering: Interactive exploration of casually captured videos. Graph: ACM Trans.","DOI":"10.1145\/1833349.1778824"},{"key":"1599_CR7","doi-asserted-by":"crossref","unstructured":"Basha, T., Moses, Y., Kiryati, N. (2010). Multi-view scene flow estimation: A view centered variational approach. In: CVPR, pp. 1506\u20131513.","DOI":"10.1109\/CVPR.2010.5539791"},{"issue":"11","key":"1599_CR8","doi-asserted-by":"publisher","first-page":"1124","DOI":"10.1109\/TPAMI.2004.60","volume":"26","author":"Y Boykov","year":"2004","unstructured":"Boykov, Y., & Kolmogorov, V. (2004). An experimental comparison of min-cut\/max- flow algorithms for energy minimization in vision. TPAMI, 26(11), 1124\u20131137.","journal-title":"TPAMI"},{"key":"1599_CR9","doi-asserted-by":"crossref","unstructured":"Boykov, Y., Veksler, O., & Zabih, R. (2001). Fast approximate energy minimization via graph cuts. TPAMI,23(11), 1222\u20131239.","DOI":"10.1109\/34.969114"},{"key":"1599_CR10","doi-asserted-by":"crossref","unstructured":"Cai, Y., Huang, L., Wang, Y., Cham, T.J., Cai, J., Yuan, J., Liu, J., Yang, X., Zhu, Y., Shen, X., Liu, D., Liu, J., Thalmann, N.M. (2020). Learning progressive joint propagation for human motion prediction. In: A.\u00a0Vedaldi, H.\u00a0Bischof, T.\u00a0Brox, J.M. Frahm (eds.) Computer Vision \u2013 ECCV 2020, pp. 226\u2013242.","DOI":"10.1007\/978-3-030-58571-6_14"},{"key":"1599_CR11","unstructured":"Caliskan, A., Mustafa, A., Imre, E., Hilton, A. (2020). Multi-view consistency loss for improved single-image 3d reconstruction of clothed people. In: Asian Conference on Computer Vision (ACCV)."},{"key":"1599_CR12","doi-asserted-by":"crossref","unstructured":"Cao, Z., Simon, T., Wei, S.E., Sheikh, Y. (2017). Realtime multi-person 2d pose estimation using part affinity fields. In: CVPR.","DOI":"10.1109\/CVPR.2017.143"},{"key":"1599_CR13","doi-asserted-by":"crossref","unstructured":"Chen, H., Sun, K., Tian, Z., Shen, C., Huang, Y., Yan, Y. (2020). Blendmask: Top-down meets bottom-up for instance segmentation. In: IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR42600.2020.00860"},{"key":"1599_CR14","unstructured":"Chen, L., Papandreou, G., Kokkinos, I., Murphy, K., Yuille, A.L. (2016). Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. CoRR arXiv:1606.00915"},{"key":"1599_CR15","doi-asserted-by":"crossref","unstructured":"Chen, L., Zhu, Y., Papandreou, G., Schroff, F., Adam, H. (2018). Encoder-decoder with atrous separable convolution for semantic image segmentation.","DOI":"10.1007\/978-3-030-01234-2_49"},{"key":"1599_CR16","doi-asserted-by":"crossref","unstructured":"Chen, P.Y., Liu, A.H., Liu, Y.C., Wang, Y. (2019). Towards scene understanding: Unsupervised monocular depth estimation with semantic-aware representation. In: CVPR.","DOI":"10.1109\/CVPR.2019.00273"},{"key":"1599_CR17","doi-asserted-by":"crossref","unstructured":"Chiu, W.C., Fritz, M. (2013). Multi-class video co-segmentation with a generative multi-video model. In: CVPR.","DOI":"10.1109\/CVPR.2013.48"},{"key":"1599_CR18","doi-asserted-by":"crossref","unstructured":"Dai, A., Nie\u00dfner, M. (2018). 3dmv: Joint 3d-multi-view prediction for 3d semantic scene segmentation. In: ECCV.","DOI":"10.1007\/978-3-030-01249-6_28"},{"key":"1599_CR19","doi-asserted-by":"crossref","unstructured":"Djelouah, A., Franco, J.S., Boyer, E., Perez, P., Drettakis, G. (2016). Cotemporal Multi-View Video Segmentation. In: 3DV.","DOI":"10.1109\/3DV.2016.45"},{"key":"1599_CR20","doi-asserted-by":"crossref","unstructured":"Dosovitskiy, A., Fischery, M., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., Smagt, P., Cremers, D., Brox, T. (2015). Flownet: Learning optical flow with convolutional networks. In: ICCV.","DOI":"10.1109\/ICCV.2015.316"},{"key":"1599_CR21","doi-asserted-by":"crossref","unstructured":"Dou, M., Khamis, S., Degtyarev, Y., Davidson, P., Fanello, S.R., Kowdle, A., Escolano, S.O., Rhemann, C., Kim, D., Taylor, J., Kohli, P., Tankovich, V., Izadi, S. (2016). Fusion4d: Real-time performance capture of challenging scenes. ACM Trans. Graph. 35(4).","DOI":"10.1145\/2897824.2925969"},{"key":"1599_CR22","doi-asserted-by":"crossref","unstructured":"Eigen, D., Fergus, R. (2015). Predicting depth, surface normals and semantic labels with a common multi-scale convolutional architecture. In: ICCV.","DOI":"10.1109\/ICCV.2015.304"},{"key":"1599_CR23","doi-asserted-by":"crossref","unstructured":"Engelmann, F., St\u00fcckler, J., Leibe, B. (2016). Joint object pose estimation and shape reconstruction in urban street scenes using 3D shape priors. In: GCPR.","DOI":"10.1007\/978-3-319-45886-1_18"},{"issue":"10","key":"1599_CR24","doi-asserted-by":"publisher","first-page":"1858","DOI":"10.1109\/TPAMI.2008.113","volume":"30","author":"GD Evangelidis","year":"2008","unstructured":"Evangelidis, G. D., & Psarakis, E. Z. (2008). Parametric image alignment using enhanced correlation coefficient maximization. TPAMI, 30(10), 1858\u20131865.","journal-title":"TPAMI"},{"key":"1599_CR25","unstructured":"Everingham, M., Van\u00a0Gool, L., Williams, C.K.I., Winn, J., Zisserman, A. (2012). The PASCAL Visual Object Classes Challenge 2012 (VOC2012) Results. http:\/\/www.pascal-network.org\/challenges\/VOC\/voc2012\/workshop\/index.html"},{"issue":"8","key":"1599_CR26","doi-asserted-by":"publisher","first-page":"1915","DOI":"10.1109\/TPAMI.2012.231","volume":"35","author":"C Farabet","year":"2013","unstructured":"Farabet, C., Couprie, C., Najman, L., & LeCun, Y. (2013). Learning hierarchical features for scene labeling. TPAMI, 35(8), 1915\u20131929.","journal-title":"TPAMI"},{"key":"1599_CR27","doi-asserted-by":"crossref","unstructured":"Floros, G., Leibe, B. (2012). Joint 2d-3d temporally consistent semantic segmentation of street scenes. In: CVPR, pp. 2823\u20132830.","DOI":"10.1109\/CVPR.2012.6248007"},{"key":"1599_CR28","doi-asserted-by":"crossref","unstructured":"Godard, C., Mac Aodha, O., Brostow, G.J. (2017). Unsupervised monocular depth estimation with left-right consistency. In: CVPR.","DOI":"10.1109\/CVPR.2017.699"},{"key":"1599_CR29","doi-asserted-by":"crossref","unstructured":"Guerry, J., Boulch, A., Saux, B.L., Moras, J., Plyer, A., Filliat, D. (2017). Snapnet-r: Consistent 3d multi-view semantic labeling for robotics. In: ICCVW.","DOI":"10.1109\/ICCVW.2017.85"},{"key":"1599_CR30","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1007\/s11263-010-0413-z","volume":"93","author":"JY Guillemaut","year":"2010","unstructured":"Guillemaut, J. Y., & Hilton, A. (2010). Joint multi-layer segmentation and reconstruction for free-viewpoint video applications. IJCV, 93, 73\u2013100.","journal-title":"IJCV"},{"key":"1599_CR31","doi-asserted-by":"crossref","unstructured":"Gupta, S., Girshick, R.B., Arbelaez, P., Malik, J. (2014). Learning rich features from RGB-D images for object detection and segmentation, pp. 345\u2013360.","DOI":"10.1007\/978-3-319-10584-0_23"},{"key":"1599_CR32","doi-asserted-by":"crossref","unstructured":"Hane, C., Zach, C., Cohen, A., Pollefeys, M. (2016). Dense semantic 3d reconstruction. TPAMI p.\u00a01.","DOI":"10.1109\/TPAMI.2016.2613051"},{"key":"1599_CR33","doi-asserted-by":"crossref","unstructured":"Hariharan, B., Arbel\u00e1ez, P.A., Girshick, R.B., Malik, J. (2015). Hypercolumns for object segmentation and fine-grained localization. In: CVPR, pp. 447\u2013456.","DOI":"10.1109\/CVPR.2015.7298642"},{"key":"1599_CR34","doi-asserted-by":"publisher","unstructured":"Hasler, N., Rosenhahn, B., Thormahlen, T., Wand, M., Gall, J., Seidel, H.P. (2009). Markerless motion capture with unsynchronized moving cameras. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 224\u2013231. https:\/\/doi.org\/10.1109\/CVPR.2009.5206859.","DOI":"10.1109\/CVPR.2009.5206859"},{"key":"1599_CR35","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., Girshick, R. (2017). Mask R-CNN. In: ICCV.","DOI":"10.1109\/ICCV.2017.322"},{"key":"1599_CR36","doi-asserted-by":"crossref","unstructured":"Huang, Y., Bogo, F., Lassner, C., Kanazawa, A., Gehler, P.V., Romero, J., Akhter, I., Black, M. J. (2017). Towards accurate marker-less human shape and pose estimation over time. In: 3DV.","DOI":"10.1109\/3DV.2017.00055"},{"key":"1599_CR37","doi-asserted-by":"crossref","unstructured":"Ionescu, C., Papava, D., Olaru, V., Sminchisescu, C. (2014). Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. TPAMI, 36(7), 1325\u20131339.","DOI":"10.1109\/TPAMI.2013.248"},{"key":"1599_CR38","unstructured":"Kazhdan, M., Bolitho, M., Hoppe, H. (2006). Poisson surface reconstruction. In: Eurographics Symposium on Geometry Processing, pp. 61\u201370"},{"key":"1599_CR39","unstructured":"Kendall, A., Gal, Y., Cipolla, R. (2017). Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. CoRR arXiv:1705.07115."},{"key":"1599_CR40","unstructured":"Kendall, A., Gal, Y., Cipolla, R. (2018). Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In: CVPR."},{"key":"1599_CR41","doi-asserted-by":"crossref","unstructured":"Kim, H., Sarim, M., Takai, T., yves Guillemaut, J., Hilton, A. (2012). Outdoor dynamic 3-D scene reconstruction. T-CSVT, 22(11), 1611\u20131622.","DOI":"10.1109\/TCSVT.2012.2202185"},{"key":"1599_CR42","doi-asserted-by":"crossref","unstructured":"Klodt, M., Vedaldi, A. (2018). Supervising the new with the old: learning sfm from sfm. In: ECCV.","DOI":"10.1007\/978-3-030-01249-6_43"},{"key":"1599_CR43","doi-asserted-by":"crossref","unstructured":"Kundu, A., Li, Y., Dellaert, F., Li, F., Rehg, J.M. (2014). Joint semantic segmentation and 3d reconstruction from monocular video. In: ECCV, vol. 8694, pp. 703\u2013718.","DOI":"10.1007\/978-3-319-10599-4_45"},{"key":"1599_CR44","doi-asserted-by":"crossref","unstructured":"Kundu, A., Vineet, V., Koltun, V. (2016). Feature space optimization for semantic video segmentation. In: CVPR, pp. 3168\u20133175.","DOI":"10.1109\/CVPR.2016.345"},{"key":"1599_CR45","doi-asserted-by":"crossref","unstructured":"Lai, H., Tsai, Y., Chiu, W. (2019). Bridging stereo matching and optical flow via spatiotemporal correspondence. In: CVPR.","DOI":"10.1109\/CVPR.2019.00199"},{"key":"1599_CR46","doi-asserted-by":"crossref","unstructured":"Langguth, F., Sunkavalli, K., Hadap, S., Goesele, M. (2016). Shading-aware multi-view stereo. In: ECCV.","DOI":"10.1007\/978-3-319-46487-9_29"},{"key":"1599_CR47","doi-asserted-by":"crossref","unstructured":"Larsen, E.S., Mordohai, P., Pollefeys, M., Fuchs, H. (2007). Temporally consistent reconstruction from multiple video streams using enhanced belief propagation. In: ICCV, pp. 1\u20138.","DOI":"10.1109\/ICCV.2007.4409013"},{"key":"1599_CR48","doi-asserted-by":"crossref","unstructured":"Li, X., You, A., Zhu, Z., Zhao, H., Yang, M., Yang, K., Tong, Y. (2020). Semantic flow for fast and accurate scene parsing. In: ECCV.","DOI":"10.1007\/978-3-030-58452-8_45"},{"key":"1599_CR49","doi-asserted-by":"crossref","unstructured":"Lin, T., Maire, M., Belongie, S.J., Bourdev, L.D., Girshick, R.B., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., Zitnick, C.L. (2014). Microsoft COCO: common objects in context. CoRR arXiv:1405.0312.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"1599_CR50","doi-asserted-by":"crossref","unstructured":"Luo, B., Li, H., Song, T., Huang, C. (2015). Object segmentation from long video sequences. In: ACM Multimedia, pp. 1187\u20131190.","DOI":"10.1145\/2733373.2806313"},{"key":"1599_CR51","doi-asserted-by":"crossref","unstructured":"Menze, M., Heipke, C., Geiger, A. (2015). Discrete optimization for optical flow. In: German Conference on Pattern Recognition (GCPR), vol. 9358, (pp. 16\u201328). Springer International Publishing.","DOI":"10.1007\/978-3-319-24947-6_2"},{"key":"1599_CR52","doi-asserted-by":"crossref","unstructured":"Mostajabi, M., Yadollahpour, P., Shakhnarovich, G. (2015). Feedforward semantic segmentation with zoom-out features. In: CVPR, pp. 3376\u20133385.","DOI":"10.1109\/CVPR.2015.7298959"},{"key":"1599_CR53","doi-asserted-by":"crossref","unstructured":"Mustafa, A., Hilton, A. (2017). Semantically coherent co-segmentation and reconstruction of dynamic scenes. In: CVPR.","DOI":"10.1109\/CVPR.2017.592"},{"key":"1599_CR54","doi-asserted-by":"crossref","unstructured":"Mustafa, A., Kim, H., Guillemaut, J., Hilton, A. (2016). Temporally coherent 4d reconstruction of complex dynamic scenes. In: CVPR.","DOI":"10.1109\/CVPR.2016.504"},{"key":"1599_CR55","doi-asserted-by":"crossref","unstructured":"Mustafa, A., Kim, H., Hilton, A. (2016). 4d match trees for non-rigid surface alignment. In: ECCV.","DOI":"10.1007\/978-3-319-46448-0_13"},{"key":"1599_CR56","doi-asserted-by":"publisher","first-page":"1118","DOI":"10.1109\/TIP.2018.2872906","volume":"28","author":"A Mustafa","year":"2019","unstructured":"Mustafa, A., Kim, H., & Hilton, A. (2019). Msfd: Multi-scale segmentation-based feature detection for wide-baseline scene reconstruction. IEEE Transactions on Image Processing, 28, 1118\u20131132.","journal-title":"IEEE Transactions on Image Processing"},{"key":"1599_CR57","doi-asserted-by":"crossref","unstructured":"Mustafa, A., Russell, C., Hilton, A. (2019). U4d: Unsupervised 4d dynamic scene understanding. In: ICCV.","DOI":"10.1109\/ICCV.2019.01052"},{"key":"1599_CR58","doi-asserted-by":"crossref","unstructured":"Mustafa, A., Volino, M., Guillemaut, J., Hilton, A. (2017). 4d temporally coherent light-field video. In: 3DV.","DOI":"10.1109\/3DV.2017.00014"},{"key":"1599_CR59","doi-asserted-by":"crossref","unstructured":"Newcombe, R.A., Fox, D., Seitz, S.M. (2015). Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time. In CVPR pp. 343\u2013352.","DOI":"10.1109\/CVPR.2015.7298631"},{"key":"1599_CR60","doi-asserted-by":"crossref","unstructured":"Ranjan, A., Jampani, V., Kim, K., Sun, D., Wulff, J., Black, M.J. (2019). Adversarial collaboration: Joint unsupervised learning of depth, camera motion, optical flow and motion segmentation. In: CVPR.","DOI":"10.1109\/CVPR.2019.01252"},{"key":"1599_CR61","unstructured":"Ranjan, A., Romero, J., Black, M.J. (2018). Learning human optical flow. In: BMVC."},{"key":"1599_CR62","unstructured":"Rodriguez, A.L., Mikolajczyk, K. (2020). Desc: Domain adaptation for depth estimation via semantic consistency. In: BMVC."},{"key":"1599_CR63","doi-asserted-by":"crossref","unstructured":"Rossi, M., Gheche, M.E., Kuhn, A., Frossard, P. (2020). Joint graph-based depth refinement and normal estimation. In: IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR42600.2020.01217"},{"key":"1599_CR64","doi-asserted-by":"crossref","unstructured":"Roussos, A., Russell, C., Garg, R., Agapito, L. (2012). Dense multibody motion estimation and reconstruction from a handheld camera. In: ISMAR.","DOI":"10.1109\/ISMAR.2012.6402535"},{"key":"1599_CR65","doi-asserted-by":"crossref","unstructured":"Rusu, R.B. (2009). Semantic 3d object maps for everyday manipulation in human living environments. Ph.D. thesis, Computer Science department, Technische Universitaet Muenchen, Germany.","DOI":"10.1007\/s13218-010-0059-6"},{"key":"1599_CR66","doi-asserted-by":"crossref","unstructured":"Bi, S., Xu, Z., Sunkavalli, K., Hasan, M., Hold-Geoffroy, Y., Kriegman, D., & Ramamoorthi, R. (2020). Deep reflectance volumes: Relightable reconstructions from multi-view photometric images. In: ECCV.","DOI":"10.1007\/978-3-030-58580-8_18"},{"key":"1599_CR67","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nberger, J.L., Frahm, J.M. (2016). Structure-from-motion revisited. In: Conference on Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR.2016.445"},{"key":"1599_CR68","doi-asserted-by":"crossref","unstructured":"Sch\u00f6nberger, J.L., Zheng, E., Pollefeys, M., Frahm, J.M. (2016). Pixelwise view selection for unstructured multi-view stereo. In: European Conference on Computer Vision (ECCV).","DOI":"10.1007\/978-3-319-46487-9_31"},{"key":"1599_CR69","doi-asserted-by":"crossref","unstructured":"Sevilla-Lara, L., Sun, D., Jampani, V., Black, M.J. (2016). Optical flow with semantic segmentation and localized layers. In: CVPR, pp. 3889\u20133898.","DOI":"10.1109\/CVPR.2016.422"},{"key":"1599_CR70","unstructured":"Shelhamer, E., Long, J., Darrell, T. (2015). Fully convolutional networks for semantic segmentation. In: CVPR."},{"key":"1599_CR71","doi-asserted-by":"crossref","unstructured":"Siam, M., Gamal, M., Abdel-Razek, M., Yogamani, S., J\u00e4gersand, M. (2018). Rtseg: Real-time semantic segmentation comparative study. In: ICIP.","DOI":"10.1109\/ICIP.2018.8451495"},{"key":"1599_CR72","unstructured":"Sorkine, O., Alexa, M. (2007). As-rigid-as-possible surface modeling. In: SGP, pp. 109\u2013116."},{"key":"1599_CR73","unstructured":"Szeliski, R. (1999). A multi-view approach to motion and stereo. In: CVPR."},{"issue":"11","key":"1599_CR74","doi-asserted-by":"publisher","first-page":"2725","DOI":"10.1109\/TPAMI.2017.2766072","volume":"40","author":"T Taniai","year":"2018","unstructured":"Taniai, T., Matsushita, Y., Sato, Y., & Naemura, T. (2018). Continuous 3D label stereo matching using local expansion moves. TPAMI, 40(11), 2725\u20132739. https:\/\/doi.org\/10.1109\/TPAMI.2017.2766072.","journal-title":"TPAMI"},{"key":"1599_CR75","doi-asserted-by":"crossref","unstructured":"Tao, M.W., Bai, J., Kohli, P., Paris, S. (2012). Simpleflow: A non-iterative, sublinear optical flow algorithm. Computer Graphics Forum (Eurographics 2012), 31(2).","DOI":"10.1111\/j.1467-8659.2012.03013.x"},{"key":"1599_CR76","doi-asserted-by":"crossref","unstructured":"Tome, D., Russell, C., Agapito, L. (2017). Lifting from the deep: Convolutional 3d pose estimation from a single image. In: CVPR.","DOI":"10.1109\/CVPR.2017.603"},{"key":"1599_CR77","doi-asserted-by":"crossref","unstructured":"Tom\u00e8, D., Toso, M., Agapito, L., Russell, C. (2018). Rethinking pose in 3d: Multi-stage refinement and recovery for markerless motion capture. In: 3DV.","DOI":"10.1109\/3DV.2018.00061"},{"key":"1599_CR78","doi-asserted-by":"crossref","unstructured":"Trager, M., Hebert, M., Ponce, J. (2019). Coordinate-free carlsson-weinshall duality and relative multi-viewgeometry. In: CVPR.","DOI":"10.1109\/CVPR.2019.00031"},{"key":"1599_CR79","doi-asserted-by":"crossref","unstructured":"Tsai, Y.H., Zhong, G., Yang, M.-H., e.B., Matas, J., Sebe, N., Welling, M. (2016). Semantic co-segmentation in videos. In: ECCV, pp. 760\u2013775.","DOI":"10.1007\/978-3-319-46493-0_46"},{"key":"1599_CR80","doi-asserted-by":"crossref","unstructured":"Ulusoy, A.O., Black, M.J., Geiger, A. (2017). Semantic multi-view stereo: Jointly estimating objects and voxels. In: CVPR.","DOI":"10.1109\/CVPR.2017.482"},{"key":"1599_CR81","doi-asserted-by":"crossref","unstructured":"Vineet, V., Miksik, O., Lidegaard, M., Nie\u00dfner, M., Golodetz, S., Prisacariu, V.A., K\u00e4hler, O., Murray, D.W., Izadi, S., Perez, P., Torr, P.H.S. (2015). Incremental dense semantic stereo fusion for large-scale semantic scene reconstruction. In: ICRA.","DOI":"10.1109\/ICRA.2015.7138983"},{"key":"1599_CR82","doi-asserted-by":"crossref","unstructured":"Vlasic, D., Baran, I., Matusik, W., Popovi\u0107, J. (2008). Articulated mesh animation from multi-view silhouettes. ACM Trans. Graph., 27(3).","DOI":"10.1145\/1360612.1360696"},{"key":"1599_CR83","doi-asserted-by":"crossref","unstructured":"Vogel, C., Schindler, K., Roth, S. (2015). 3d scene flow estimation with a piecewise rigid scene model pp. 1\u201328.","DOI":"10.1007\/s11263-015-0806-0"},{"key":"1599_CR84","doi-asserted-by":"crossref","unstructured":"Wang, L., Zhang, J., Wang, O., Lin, Z., Lu, H. (2020). Sdc-depth: Semantic divide-and-conquer network for monocular depth estimation. In: IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR).","DOI":"10.1109\/CVPR42600.2020.00062"},{"issue":"1","key":"1599_CR85","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1007\/s11263-010-0404-0","volume":"95","author":"A Wedel","year":"2011","unstructured":"Wedel, A., Brox, T., Vaudrey, T., Rabe, C., Franke, U., & Cremers, D. (2011). Stereoscopic scene flow computation for 3d motion understanding. IJCV, 95(1), 29\u201351.","journal-title":"IJCV"},{"key":"1599_CR86","doi-asserted-by":"crossref","unstructured":"Wei\u00a0Zeng, S.K., Gevers, T. (2020). Pano2scene: 3d indoor semantic scene reconstruction from a single indoor panorama image. In: BMVC.","DOI":"10.1007\/978-3-030-58517-4_39"},{"key":"1599_CR87","doi-asserted-by":"crossref","unstructured":"Weinzaepfel, P., Revaud, J., Harchaoui, Z., Schmid, C. (2013). Deepflow: Large displacement optical flow with deep matching. In: ICCV, pp. 1385\u20131392.","DOI":"10.1109\/ICCV.2013.175"},{"key":"1599_CR88","doi-asserted-by":"crossref","unstructured":"Xia, F., Wang, P., Chen, X., Yuille, A.L. (2017). Joint multi-person pose estimation and semantic part segmentation. In: CVPR.","DOI":"10.1109\/CVPR.2017.644"},{"key":"1599_CR89","doi-asserted-by":"crossref","unstructured":"Xie, J., Kiefel, M., Sun, M.T., Geiger, A. (2016). Semantic instance annotation of street scenes by 3d to 2d label transfer. In: CVPR.","DOI":"10.1109\/CVPR.2016.401"},{"key":"1599_CR90","doi-asserted-by":"crossref","unstructured":"Xu, J., Ranftl, R., Koltun, V. (2017). Accurate optical flow via direct cost volume processing. In: CVPR.","DOI":"10.1109\/CVPR.2017.615"},{"key":"1599_CR91","doi-asserted-by":"crossref","unstructured":"Yao, Y., Luo, Z., Li, S., Fang, T., Quan, L. (2018). Mvsnet: Depth inference for unstructured multi-view stereo. In: ECCV.","DOI":"10.1007\/978-3-030-01237-3_47"},{"key":"1599_CR92","doi-asserted-by":"crossref","unstructured":"Zanfir, A., Sminchisescu, C. (2015). Large displacement 3d scene flow with occlusion reasoning. In: ICCV.","DOI":"10.1109\/ICCV.2015.502"},{"key":"1599_CR93","doi-asserted-by":"crossref","unstructured":"Zhao, H., Shi, J., Qi, X., Wang, X., Jia, J. (2017). Pyramid scene parsing network. In: CVPR.","DOI":"10.1109\/CVPR.2017.660"},{"key":"1599_CR94","doi-asserted-by":"crossref","unstructured":"Zheng, S., Jayasumana, S., Romera-Paredes, B., Vineet, V., Su, Z., Du, D., Huang, C., Torr, P.H.S. (2015). Conditional random fields as recurrent neural networks. In: ICCV.","DOI":"10.1109\/ICCV.2015.179"},{"key":"1599_CR95","doi-asserted-by":"crossref","unstructured":"Zhong, Y., Ji, P., Wang, J., Dai, Y., Li, H. (2019). Unsupervised deep epipolar flow for stationary or dynamic scenes. In: CVPR.","DOI":"10.1109\/CVPR.2019.01237"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-022-01599-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-022-01599-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-022-01599-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,5,27]],"date-time":"2022-05-27T21:09:43Z","timestamp":1653685783000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-022-01599-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,4,28]]},"references-count":95,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2022,6]]}},"alternative-id":["1599"],"URL":"https:\/\/doi.org\/10.1007\/s11263-022-01599-4","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,4,28]]},"assertion":[{"value":"11 December 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 February 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 April 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}