{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T00:39:27Z","timestamp":1760229567674,"version":"build-2065373602"},"reference-count":38,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2022,6,16]],"date-time":"2022-06-16T00:00:00Z","timestamp":1655337600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100012554","name":"Hubei Provincial Department of Education","doi-asserted-by":"publisher","award":["21D031"],"award-info":[{"award-number":["21D031"]}],"id":[{"id":"10.13039\/100012554","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In order to avoid the direct depth reconstruction of the original image pair and improve the accuracy of the results, we proposed a coarse-to-fine stereo matching network combining multi-level residual optimization and depth map super-resolution (ASR-Net). First, we used the u-net feature extractor to obtain the multi-scale feature pair. Second, we reconstructed global disparity in the lowest resolution. Then, we regressed the residual disparity using the higher-resolution feature pair. Finally, the lowest-resolution depth map was refined by using the disparity residual. In addition, we introduced deformable convolution and group-wise cost volume into the network to achieve adaptive cost aggregation. Further, the network uses ABPN instead of the traditional interpolation method. The network was evaluated on three datasets: scene flow, kitti2015, and kitti2012 and the experimental results showed that the speed and accuracy of our method were excellent. On the kitti2015 dataset, the three-pixel error converged to 2.86%, and the speed was about six times and two times that of GC-net and GWC-net.<\/jats:p>","DOI":"10.3390\/s22124548","type":"journal-article","created":{"date-parts":[[2022,6,19]],"date-time":"2022-06-19T21:19:26Z","timestamp":1655673566000},"page":"4548","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Adaptive Aggregate Stereo Matching Network with Depth Map Super-Resolution"],"prefix":"10.3390","volume":"22","author":[{"given":"Botao","family":"Liu","sequence":"first","affiliation":[{"name":"School of Computer Science, Yangtze University, Jingzhou 434023, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kai","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Computer Science, Yangtze University, Jingzhou 434023, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9484-6677","authenticated-orcid":false,"given":"Sheng-Lung","family":"Peng","sequence":"additional","affiliation":[{"name":"Department of Creative Technologies and Product Design, National Taipei University of Business, Taipei 10051, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1647-1769","authenticated-orcid":false,"given":"Ming","family":"Zhao","sequence":"additional","affiliation":[{"name":"School of Computer Science, Yangtze University, Jingzhou 434023, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,6,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"6917","DOI":"10.1109\/ACCESS.2017.2698164","article-title":"Mobile augmented reality survey: From where we are to where we go","volume":"5","author":"Chatzopoulos","year":"2017","journal-title":"IEEE Access"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Costanza, E., Kunz, A., and Fjeld, M. (2009). Mixed reality: A survey. Human Machine Interaction, Springer.","DOI":"10.1007\/978-3-642-00437-7_3"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"900","DOI":"10.1109\/TITS.2019.2901817","article-title":"Autonomous vehicles that interact with pedestrians: A survey of theory and practice","volume":"21","author":"Rasouli","year":"2019","journal-title":"IEEE Trans. Intell. Transp. Syst."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"237","DOI":"10.1109\/34.982903","article-title":"Vision for mobile robot navigation: A survey","volume":"24","author":"Desouza","year":"2002","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Wang, Y., Lai, Z., Huang, G., Wang, B.H., Van Der Maaten, L., Campbell, M., and Weinberger, K.Q. (2019, January 20\u201324). Anytime stereo image depth estimation on mobile devices. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8794003"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Kendall, A., Martirosyan, H., Dasgupta, S., Henry, P., Kennedy, R., Bachrach, A., and Bry, A. (2017, January 22\u201329). End-to-end learning of geometry and context for deep stereo regression. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.17"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Liu, Z.S., Wang, L.W., Li, C.T., Siu, W.C., and Chan, Y.L. (2019, January 27\u201328). Image super-resolution via attention based back projection networks. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision Workshop (ICCVW), Seoul, Korea.","DOI":"10.1109\/ICCVW.2019.00436"},{"key":"ref_8","first-page":"702","article-title":"Literature survey on stereo vision matching algorithms","volume":"41","author":"Chen","year":"2020","journal-title":"J. Graph."},{"key":"ref_9","unstructured":"Hirschmuller, H. (2005, January 20\u201325). Accurate and efficient stereo processing by semi-global matching and mutual information. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_10","first-page":"1","article-title":"Patchmatch stereo-stereo matching with slanted support windows","volume":"11","author":"Bleyer","year":"2011","journal-title":"Bmvc"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Mayer, N., Ilg, E., Hausser, P., Fischer, P., Cremers, D., Dosovitskiy, A., and Brox, T. (2016, January 27\u201330). A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.438"},{"key":"ref_12","unstructured":"Luo, K., Guan, T., Ju, L., Huang, H., and Luo, Y. (November, January 27). P-mvsnet: Learning patch-wise matching confidence aggregation for multi-view stereo. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chang, J.R., and Chen, Y.S. (2018, January 18\u201322). Pyramid stereo matching network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00567"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Guo, X., Yang, K., Yang, W., Wang, X., and Li, H. (2019, January 16\u201320). Group-wise correlation stereo network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00339"},{"key":"ref_15","unstructured":"Bahdanau, D., Cho, K., and Bengio, Y. (2014). Neural machine translation by jointly learning to align and translate. arXiv."},{"key":"ref_16","unstructured":"Zhang, Y., Li, K., Li, K., Zhong, B., and Fu, Y. (2019). Residual non-local attention networks for image restoration. arXiv."},{"key":"ref_17","first-page":"186","article-title":"Stereo Matching Network with Multi-Cost Fusion","volume":"48","author":"Zhang","year":"2022","journal-title":"Comput. Eng."},{"key":"ref_18","first-page":"877","article-title":"Pixel attention based siamese convolutionneural network for stereo matching","volume":"42","author":"Sang","year":"2020","journal-title":"Comput. Eng. Sci."},{"key":"ref_19","first-page":"216","article-title":"Real-time Binocular Depth Estimation Algorithm Based on Semantic Edge Drive","volume":"48","author":"Zhang","year":"2021","journal-title":"Comput. Sci."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Xu, H., and Zhang, J. (2020, January 13\u201319). Aanet: Adaptive aggregation network for efficient stereo matching. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00203"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Khamis, S., Fanello, S., Rhemann, C., Kowdle, A., Valentin, J., and Izadi, S. (2018, January 8\u201314). Stereonet: Guided hierarchical refinement for real-time edge-aware depth prediction. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_35"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Li, K., Li, K., Wang, L., Zhong, B., and Fu, Y. (2018, January 8\u201314). Image super-resolution using very deep residual channel attention networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_18"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"4121","DOI":"10.1109\/JSTARS.2020.3009352","article-title":"Channel-attention-based DenseNet network for remote sensing image scene classification","volume":"13","author":"Tong","year":"2020","journal-title":"IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Song, X., Dai, Y., Zhou, D., Liu, L., Li, W., Li, H., and Yang, R. (2020, January 13\u201319). Channel attention based iterative residual learning for depth map super-resolution. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00567"},{"key":"ref_25","first-page":"1479","article-title":"Exudate Detection for Retinal Fundus Image Based on U-net Incorporating Residual Module","volume":"42","author":"Fu","year":"2021","journal-title":"J. Chin. Comput. Syst."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Kong, F., Picetti, F., Lipari, V., Bestagini, P., and Tubaro, S. (2020, January 14). Deep prior-based seismic data interpolation via multi-res u-net. Proceedings of the SEG International Exposition and Annual Meeting, Virtual.","DOI":"10.1190\/segam2020-3426173.1"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-net: Convolutional networks for biomedical image segmentation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Xiao, X., Lian, S., Luo, Z., and Li, S. (2018, January 19\u201321). Weighted res-unet for height-quality retina vessel segmentation. Proceedings of the 2018 9th International Conference on Information Technology in Medicine and Education (ITME), Hangzhou, China.","DOI":"10.1109\/ITME.2018.00080"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","article-title":"DeepLab: Semantic Image Segmenta-tion with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs","volume":"40","author":"Chen","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"\u00c7i\u00e7ek, \u00d6., Abdulkadir, A., Lienkamp, S.S., Brox, T., and Ronneberger, O. (2016, January 17\u201321). 3D U-Net: Learning dense volumetric segmentation from sparse annotation. Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention, Athens, Greece.","DOI":"10.1007\/978-3-319-46723-8_49"},{"key":"ref_33","first-page":"307","article-title":"Network Model for Lung Nodule Segmentation Based on Double Attention 3D-UNet","volume":"47","author":"Wang","year":"2021","journal-title":"Comput. Eng."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, F., Galliani, S., Vogel, C., Speciale, P., and Pollefeys, M. (2021, January 20\u201325). Patchmatchnet: Learned multi-view patchmatch stereo. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01397"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Scharstein, D., Hirschm\u00fcller, H., Kitajima, Y., Krathwohl, G., Ne\u0161i\u0107, N., Wang, X., and Westling, P. (2014, January 2\u20135). Height-resolution stereo datasets with subpixel-accurate ground truth. Proceedings of the German Conference on Pattern Recognition, M\u00fcnster, Germany.","DOI":"10.1007\/978-3-319-11752-2_3"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"2323","DOI":"10.1109\/TCSVT.2018.2866399","article-title":"Deeply supervised depth map super-resolution as novel view synthesis","volume":"29","author":"Song","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Zbontar, J., and LeCun, Y. (2015, January 7\u201312). Computing the stereo matching cost with a convolutional neural network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298767"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/12\/4548\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:33:21Z","timestamp":1760139201000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/12\/4548"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,16]]},"references-count":38,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2022,6]]}},"alternative-id":["s22124548"],"URL":"https:\/\/doi.org\/10.3390\/s22124548","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2022,6,16]]}}}