{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T05:22:23Z","timestamp":1784092943834,"version":"3.55.0"},"reference-count":53,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2023,7,12]],"date-time":"2023-07-12T00:00:00Z","timestamp":1689120000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2021YFB2802300"],"award-info":[{"award-number":["2021YFB2802300"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62125110, 62101379, 61931014"],"award-info":[{"award-number":["62125110, 62101379, 61931014"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"DiDi GAIA Research Cooperation Initiative, Natural Science Foundation of Tianjin","award":["21JCQNJC01520"],"award-info":[{"award-number":["21JCQNJC01520"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["2022M712371, 2021TQ0244"],"award-info":[{"award-number":["2022M712371, 2021TQ0244"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2023,11,30]]},"abstract":"<jats:p>This article proposes a network, referred to as Multi-View Stereo TRansformer (MVSTR) for depth estimation from multi-view images. By modeling long-range dependencies and epipolar geometry, the proposed MVSTR is capable of extracting dense features with global context and 3D consistency, which are crucial for reliable matching in multi-view stereo (MVS). Specifically, to tackle the problem of the limited receptive field of existing CNN-based MVS methods, a global-context Transformer module is designed to establish intra-view long-range dependencies so that global contextual features of each view are obtained. In addition, to further enable features of each view to be 3D consistent, a 3D-consistency Transformer module with an epipolar feature sampler is built, where epipolar geometry is modeled to effectively facilitate cross-view interaction. Experimental results show that the proposed MVSTR achieves the best overall performance on the DTU dataset and demonstrates strong generalization on the Tanks &amp; Temples benchmark dataset.<\/jats:p>","DOI":"10.1145\/3596445","type":"journal-article","created":{"date-parts":[[2023,5,5]],"date-time":"2023-05-05T12:27:38Z","timestamp":1683289658000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":30,"title":["Modeling Long-range Dependencies and Epipolar Geometry for Multi-view Stereo"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4081-2073","authenticated-orcid":false,"given":"Jie","family":"Zhu","sequence":"first","affiliation":[{"name":"School of Electrical and Information Engineering, Tianjin University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6616-453X","authenticated-orcid":false,"given":"Bo","family":"Peng","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, Tianjin University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4427-2687","authenticated-orcid":false,"given":"Wanqing","family":"Li","sequence":"additional","affiliation":[{"name":"Advanced Multimedia Research Lab, University of Wollongong, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2654-3084","authenticated-orcid":false,"given":"Haifeng","family":"Shen","sequence":"additional","affiliation":[{"name":"AIoT Platform, Didi Chuxing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7542-296X","authenticated-orcid":false,"given":"Qingming","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3171-7680","authenticated-orcid":false,"given":"Jianjun","family":"Lei","sequence":"additional","affiliation":[{"name":"School of Electrical and Information Engineering, Tianjin University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,7,12]]},"reference":[{"key":"e_1_3_1_2_2","first-page":"1","volume-title":"British Machine Vision Conference (BMVC\u201911)","volume":"11","author":"Bleyer Michael","year":"2011","unstructured":"Michael Bleyer, Christoph Rhemann, and Carsten Rother. 2011. PatchMatch stereo-stereo matching with slanted support windows. In British Machine Vision Conference (BMVC\u201911), Vol. 11. 1\u201311."},{"key":"e_1_3_1_3_2","first-page":"5639","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201918)","author":"Bulo Samuel Rota","year":"2018","unstructured":"Samuel Rota Bulo, Lorenzo Porzi, and Peter Kontschieder. 2018. In-place activated batchnorm for memory-optimized training of DNNs. In Conference on Computer Vision and Pattern Recognition (CVPR\u201918). 5639\u20135647."},{"key":"e_1_3_1_4_2","first-page":"766","volume-title":"European Conference on Computer Vision (ECCV\u201908)","author":"Campbell Neill D. F.","year":"2008","unstructured":"Neill D. F. Campbell, George Vogiatzis, Carlos Hern\u00e1ndez, and Roberto Cipolla. 2008. Using multiple hypotheses to improve depth-maps for multi-view stereo. In European Conference on Computer Vision (ECCV\u201908). 766\u2013779."},{"key":"e_1_3_1_5_2","first-page":"213","volume-title":"European Conference on Computer Vision (ECCV\u201920)","author":"Carion Nicolas","year":"2020","unstructured":"Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In European Conference on Computer Vision (ECCV\u201920). 213\u2013229."},{"key":"e_1_3_1_6_2","first-page":"12299","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201921)","author":"Chen Hanting","year":"2021","unstructured":"Hanting Chen, Yunhe Wang, Tianyu Guo, Chang Xu, Yiping Deng, Zhenhua Liu, Siwei Ma, Chunjing Xu, Chao Xu, and Wen Gao. 2021. Pre-trained image processing transformer. In Conference on Computer Vision and Pattern Recognition (CVPR\u201921). 12299\u201312310."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3000611"},{"key":"e_1_3_1_8_2","first-page":"1538","volume-title":"International Conference on Computer Vision (ICCV\u201919)","author":"Chen Rui","year":"2019","unstructured":"Rui Chen, Songfang Han, Jing Xu, and Hao Su. 2019. Point-based multi-view stereo network. In International Conference on Computer Vision (ICCV\u201919). 1538\u20131547."},{"key":"e_1_3_1_9_2","first-page":"2524","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Cheng Shuo","year":"2020","unstructured":"Shuo Cheng, Zexiang Xu, Shilin Zhu, Zhuwen Li, Li Erran Li, Ravi Ramamoorthi, and Hao Su. 2020. Deep stereo using adaptive thin volume representation with uncertainty awareness. In Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 2524\u20132534."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1179"},{"key":"e_1_3_1_11_2","first-page":"8585","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201922)","author":"Ding Yikang","year":"2022","unstructured":"Yikang Ding, Wentao Yuan, Qingtian Zhu, Haotian Zhang, Xiangyue Liu, Yuanjiang Wang, and Xiao Liu. 2022. TransMVSNet: Global context-aware multi-view stereo network with transformers. In Conference on Computer Vision and Pattern Recognition (CVPR\u201922). 8585\u20138594."},{"key":"e_1_3_1_12_2","volume-title":"International Conference on Learning Representations (ICLR\u201921)","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An image is worth 16x16 words: Transformers for image recognition at scale. In International Conference on Learning Representations (ICLR\u201921)."},{"issue":"5","key":"e_1_3_1_13_2","doi-asserted-by":"crossref","first-page":"2305","DOI":"10.1109\/TIP.2018.2885229","article-title":"Deep3DSaliency: Deep stereoscopic video saliency detection model by 3D convolutional networks","volume":"28","author":"Fang Yuming","year":"2019","unstructured":"Yuming Fang, Guanqun Ding, Jia Li, and Zhijun Fang. 2019. Deep3DSaliency: Deep stereoscopic video saliency detection model by 3D convolutional networks. IEEE Transactions on Image Processing 28, 5 (2019), 2305\u20132318.","journal-title":"IEEE Transactions on Image Processing"},{"issue":"8","key":"e_1_3_1_14_2","doi-asserted-by":"crossref","first-page":"1362","DOI":"10.1109\/TPAMI.2009.161","article-title":"Accurate, dense, and robust multiview stereopsis","volume":"32","author":"Furukawa Yasutaka","year":"2009","unstructured":"Yasutaka Furukawa and Jean Ponce. 2009. Accurate, dense, and robust multiview stereopsis. IEEE Transactions on Pattern Analysis and Machine Intelligence 32, 8 (2009), 1362\u20131376.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_15_2","first-page":"873","volume-title":"International Conference on Computer Vision (ICCV\u201915)","author":"Galliani Silvano","year":"2015","unstructured":"Silvano Galliani, Katrin Lasinger, and Konrad Schindler. 2015. Massively parallel multiview stereopsis by surface normal diffusion. In International Conference on Computer Vision (ICCV\u201915). 873\u2013881."},{"key":"e_1_3_1_16_2","first-page":"2495","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Gu Xiaodong","year":"2020","unstructured":"Xiaodong Gu, Zhiwen Fan, Siyu Zhu, Zuozhuo Dai, Feitong Tan, and Ping Tan. 2020. Cascade cost volume for high-resolution multi-view stereo and stereo matching. In Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 2495\u20132504."},{"key":"e_1_3_1_17_2","first-page":"3273","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201919)","author":"Guo Xiaoyang","year":"2019","unstructured":"Xiaoyang Guo, Kai Yang, Wukui Yang, Xiaogang Wang, and Hongsheng Li. 2019. Group-wise correlation stereo network. In Conference on Computer Vision and Pattern Recognition (CVPR\u201919). 3273\u20133282."},{"key":"e_1_3_1_18_2","first-page":"7779","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"He Yihui","year":"2020","unstructured":"Yihui He, Rui Yan, Katerina Fragkiadaki, and Shoou-I Yu. 2020. Epipolar transformers. In Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 7779\u20137788."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/3362101"},{"key":"e_1_3_1_21_2","first-page":"406","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201914)","author":"Jensen Rasmus","year":"2014","unstructured":"Rasmus Jensen, Anders Dahl, George Vogiatzis, Engin Tola, and Henrik Aan\u00e6s. 2014. Large scale multi-view stereopsis evaluation. In Conference on Computer Vision and Pattern Recognition (CVPR\u201914). 406\u2013413."},{"key":"e_1_3_1_22_2","first-page":"2307","volume-title":"International Conference on Computer Vision (ICCV\u201917)","author":"Ji Mengqi","year":"2017","unstructured":"Mengqi Ji, Juergen Gall, Haitian Zheng, Yebin Liu, and Lu Fang. 2017. SurfaceNet: An end-to-end 3D neural network for multiview stereopsis. In International Conference on Computer Vision (ICCV\u201917). 2307\u20132315."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.3390\/s22197659"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3146714"},{"key":"e_1_3_1_25_2","first-page":"5156","volume-title":"International Conference on Machine Learning (ICML\u201920)","author":"Katharopoulos Angelos","year":"2020","unstructured":"Angelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, and Fran\u00e7ois Fleuret. 2020. Transformers are RNNs: Fast autoregressive transformers with linear attention. In International Conference on Machine Learning (ICML\u201920). 5156\u20135165."},{"key":"e_1_3_1_26_2","first-page":"66","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201917)","author":"Kendall Alex","year":"2017","unstructured":"Alex Kendall, Hayk Martirosyan, Saumitro Dasgupta, Peter Henry, Ryan Kennedy, Abraham Bachrach, and Adam Bry. 2017. End-to-end learning of geometry and context for deep stereo regression. In Conference on Computer Vision and Pattern Recognition (CVPR\u201917). 66\u201375."},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073599"},{"issue":"7","key":"e_1_3_1_28_2","doi-asserted-by":"crossref","first-page":"2686","DOI":"10.1109\/TCSVT.2020.3027616","article-title":"Deep spatial-spectral subspace clustering for hyperspectral image","volume":"31","author":"Lei Jianjun","year":"2021","unstructured":"Jianjun Lei, Xinyu Li, Bo Peng, Leyuan Fang, Nam Ling, and Qingming Huang. 2021. Deep spatial-spectral subspace clustering for hyperspectral image. IEEE Transactions on Circuits and Systems for Video Technology 31, 7 (2021), 2686\u20132697.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2791810"},{"key":"e_1_3_1_30_2","first-page":"6197","volume-title":"International Conference on Computer Vision (ICCV\u201921)","author":"Li Zhaoshuo","year":"2021","unstructured":"Zhaoshuo Li, Xingtong Liu, Nathan Drenkow, Andy Ding, Francis X. Creighton, Russell H. Taylor, and Mathias Unberath. 2021. Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers. In International Conference on Computer Vision (ICCV\u201921). 6197\u20136206."},{"key":"e_1_3_1_31_2","article-title":"WT-MVSNet: Window-based transformers for multi-view stereo","author":"Liao Jinli","year":"2022","unstructured":"Jinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang, Shihao Ren, Jia Guo, Wensen Feng, and Kai Zhang. 2022. WT-MVSNet: Window-based transformers for multi-view stereo. arXiv preprint arXiv:2205.14319 (2022).","journal-title":"arXiv preprint arXiv:2205.14319"},{"issue":"1","key":"e_1_3_1_32_2","first-page":"1","article-title":"Dense 3D-convolutional neural network for person re-identification in videos","volume":"15","author":"Liu Jiawei","year":"2019","unstructured":"Jiawei Liu, Zheng-Jun Zha, Xuejin Chen, Zilei Wang, and Yongdong Zhang. 2019. Dense 3D-convolutional neural network for person re-identification in videos. ACM Transactions on Multimedia Computing, Communications, and Applications 15, 1s (2019), 1\u201319.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_33_2","first-page":"10012","volume-title":"International Conference on Computer Vision (ICCV\u201921)","author":"Liu Ze","year":"2021","unstructured":"Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In International Conference on Computer Vision (ICCV\u201921). 10012\u201310022."},{"key":"e_1_3_1_34_2","first-page":"10452","volume-title":"International Conference on Computer Vision (ICCV\u201919)","author":"Luo Keyang","year":"2019","unstructured":"Keyang Luo, Tao Guan, Lili Ju, Haipeng Huang, and Yawei Luo. 2019. P-MVSNet: Learning patch-wise matching confidence aggregation for multi-view stereo. In International Conference on Computer Vision (ICCV\u201919). 10452\u201310461."},{"key":"e_1_3_1_35_2","first-page":"5732","volume-title":"International Conference on Computer Vision (ICCV\u201921)","author":"Ma Xinjun","year":"2021","unstructured":"Xinjun Ma, Yue Gong, Qirui Wang, Jingwei Huang, Lei Chen, and Fan Yu. 2021. EPP-MVSNet: Epipolar-assembling based depth prediction for multi-view stereo. In International Conference on Computer Vision (ICCV\u201921). 5732\u20135740."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3503927"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3550485"},{"issue":"12","key":"e_1_3_1_38_2","doi-asserted-by":"crossref","first-page":"8342","DOI":"10.1109\/TCSVT.2022.3190916","article-title":"LVE-S2D: Low-light video enhancement from static to dynamic","volume":"32","author":"Peng Bo","year":"2022","unstructured":"Bo Peng, Xuanyu Zhang, Jianjun Lei, Zhe Zhang, Nam Ling, and Qingming Huang. 2022. LVE-S2D: Low-light video enhancement from static to dynamic. IEEE Transactions on Circuits and Systems for Video Technology 32, 12 (2022), 8342\u20138352.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_39_2","first-page":"234","volume-title":"International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI\u201915)","author":"Ronneberger Olaf","year":"2015","unstructured":"Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-Net: Convolutional networks for biomedical image segmentation. In International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI\u201915). 234\u2013241."},{"key":"e_1_3_1_40_2","first-page":"501","volume-title":"European Conference on Computer Vision (ECCV\u201916)","author":"Sch\u00f6nberger Johannes L.","year":"2016","unstructured":"Johannes L. Sch\u00f6nberger, Enliang Zheng, Jan-Michael Frahm, and Marc Pollefeys. 2016. Pixelwise view selection for unstructured multi-view stereo. In European Conference on Computer Vision (ECCV\u201916). 501\u2013518."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00138-011-0346-8"},{"key":"e_1_3_1_42_2","first-page":"5998","volume-title":"Advances in Neural Information Processing Systems (NeurIPS\u201917)","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems (NeurIPS\u201917). 5998\u20136008."},{"key":"e_1_3_1_43_2","first-page":"14194","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201921)","author":"Wang Fangjinhua","year":"2021","unstructured":"Fangjinhua Wang, Silvano Galliani, Christoph Vogel, Pablo Speciale, and Marc Pollefeys. 2021. PatchmatchNet: Learned multi-view patchmatch stereo. In Conference on Computer Vision and Pattern Recognition (CVPR\u201921). 14194\u201314203."},{"key":"e_1_3_1_44_2","first-page":"573","volume-title":"European Conference on Computer Vision (ECCV\u201922)","author":"Wang Xiaofeng","year":"2022","unstructured":"Xiaofeng Wang, Zheng Zhu, Guan Huang, Fangbo Qin, Yun Ye, Yijia He, Xu Chi, and Xingang Wang. 2022. MVSTER: Epipolar transformer for efficient multi-view stereo. In European Conference on Computer Vision (ECCV\u201922). 573\u2013591."},{"key":"e_1_3_1_45_2","first-page":"6187","volume-title":"International Conference on Computer Vision (ICCV\u201921)","author":"Wei Zizhuang","year":"2021","unstructured":"Zizhuang Wei, Qingtian Zhu, Chen Min, Yisong Chen, and Guoping Wang. 2021. AA-RMVSNet: Adaptive aggregation recurrent multi-view stereo network. In International Conference on Computer Vision (ICCV\u201921). 6187\u20136196."},{"key":"e_1_3_1_46_2","first-page":"5483","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201919)","author":"Xu Qingshan","year":"2019","unstructured":"Qingshan Xu and Wenbing Tao. 2019. Multi-scale geometric consistency guided multi-view stereo. In Conference on Computer Vision and Pattern Recognition (CVPR\u201919). 5483\u20135492."},{"key":"e_1_3_1_47_2","first-page":"12508","volume-title":"AAAI Conference on Artificial Intelligence (AAAI\u201920)","volume":"34","author":"Xu Qingshan","year":"2020","unstructured":"Qingshan Xu and Wenbing Tao. 2020. Learning inverse depth regression for multi-view stereo with correlation cost volume. In AAAI Conference on Artificial Intelligence (AAAI\u201920), Vol. 34. 12508\u201312515."},{"key":"e_1_3_1_48_2","first-page":"4312","volume-title":"International Conference on Computer Vision (ICCV\u201919)","author":"Xue Youze","year":"2019","unstructured":"Youze Xue, Jiansheng Chen, Weitao Wan, Yiqing Huang, Cheng Yu, Tianpeng Li, and Jiayu Bao. 2019. MVSCRF: Learning multi-view stereo with conditional random fields. In International Conference on Computer Vision (ICCV\u201919). 4312\u20134321."},{"key":"e_1_3_1_49_2","first-page":"4877","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Yang Jiayu","year":"2020","unstructured":"Jiayu Yang, Wei Mao, Jose M. Alvarez, and Miaomiao Liu. 2020. Cost volume pyramid based depth inference for multi-view stereo. In Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 4877\u20134886."},{"key":"e_1_3_1_50_2","first-page":"767","volume-title":"European Conference on Computer Vision (ECCV\u201918)","author":"Yao Yao","year":"2018","unstructured":"Yao Yao, Zixin Luo, Shiwei Li, Tian Fang, and Long Quan. 2018. MVSNet: Depth inference for unstructured multi-view stereo. In European Conference on Computer Vision (ECCV\u201918). 767\u2013783."},{"key":"e_1_3_1_51_2","first-page":"5525","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201919)","author":"Yao Yao","year":"2019","unstructured":"Yao Yao, Zixin Luo, Shiwei Li, Tianwei Shen, Tian Fang, and Long Quan. 2019. Recurrent MVSNet for high-resolution multi-view stereo depth inference. In Conference on Computer Vision and Pattern Recognition (CVPR\u201919). 5525\u20135534."},{"key":"e_1_3_1_52_2","first-page":"766","volume-title":"European Conference on Computer Vision (ECCV\u201920)","author":"Yi Hongwei","year":"2020","unstructured":"Hongwei Yi, Zizhuang Wei, Mingyu Ding, Runze Zhang, Yisong Chen, Guoping Wang, and Yu-Wing Tai. 2020. Pyramid multi-view stereo net with self-adaptive view aggregation. In European Conference on Computer Vision (ECCV\u201920). 766\u2013782."},{"issue":"1","key":"e_1_3_1_53_2","first-page":"1","article-title":"Multi-view shape generation for 3D human-like body","volume":"19","author":"Yu Hang","year":"2022","unstructured":"Hang Yu, Chilam Cheang, Yanwei Fu, and Xiangyang Xue. 2022. Multi-view shape generation for 3D human-like body. ACM Transactions on Multimedia Computing, Communications, and Applications 19, 1 (2022), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_54_2","first-page":"1949","volume-title":"Conference on Computer Vision and Pattern Recognition (CVPR\u201920)","author":"Yu Zehao","year":"2020","unstructured":"Zehao Yu and Shenghua Gao. 2020. Fast-MVSNet: Sparse-to-dense multi-view stereo with learned propagation and gauss-newton refinement. In Conference on Computer Vision and Pattern Recognition (CVPR\u201920). 1949\u20131958."}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3596445","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3596445","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:48:00Z","timestamp":1750178880000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3596445"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,12]]},"references-count":53,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2023,11,30]]}},"alternative-id":["10.1145\/3596445"],"URL":"https:\/\/doi.org\/10.1145\/3596445","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,12]]},"assertion":[{"value":"2022-10-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-20","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}