{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T14:25:13Z","timestamp":1782397513705,"version":"3.54.5"},"reference-count":42,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2025,8,1]],"date-time":"2025-08-01T00:00:00Z","timestamp":1754006400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Research Fund of R&amp;D and application demonstration of trusted multimodal large-scale model technology for industrial situation awareness and decision-making","award":["Z241100001324010"],"award-info":[{"award-number":["Z241100001324010"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Accurate and efficient stereo matching is fundamental to real-time depth estimation from symmetric stereo cameras in autonomous driving systems. However, existing high-accuracy stereo matching networks typically rely on computationally expensive 3D convolutions, which limit their practicality in real-world environments. In contrast, real-time methods often sacrifice accuracy or generalization capability. To address these challenges, we propose FAMNet (Fusion Attention Multi-Scale Network), a lightweight and generalizable stereo matching framework tailored for real-time depth estimation in autonomous driving applications. FAMNet consists of two novel modules: Fusion Attention-based Cost Volume (FACV) and Multi-scale Attention Aggregation (MAA). FACV constructs a compact yet expressive cost volume by integrating multi-scale correlation, attention-guided feature fusion, and channel reweighting, thereby reducing reliance on heavy 3D convolutions. MAA further enhances disparity estimation by fusing multi-scale contextual cues through pyramid-based aggregation and dual-path attention mechanisms. Extensive experiments on the KITTI 2012 and KITTI 2015 benchmarks demonstrate that FAMNet achieves a favorable trade-off between accuracy, efficiency, and generalization. On KITTI 2015, with the incorporation of FACV and MAA, the prediction accuracy of the baseline model is improved by 37% and 38%, respectively, and a total improvement of 42% is achieved by our final model. These results highlight FAMNet\u2019s potential for practical deployment in resource-constrained autonomous driving systems requiring real-time and reliable depth perception.<\/jats:p>","DOI":"10.3390\/sym17081214","type":"journal-article","created":{"date-parts":[[2025,8,1]],"date-time":"2025-08-01T06:37:47Z","timestamp":1754030267000},"page":"1214","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["FAMNet: A Lightweight Stereo Matching Network for Real-Time Depth Estimation in Autonomous Driving"],"prefix":"10.3390","volume":"17","author":[{"given":"Jingyuan","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Computer Science, Beijing Information Science & Technology University, Beijing 102206, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0508-0727","authenticated-orcid":false,"given":"Qiang","family":"Tong","sequence":"additional","affiliation":[{"name":"College of Computer Science, Beijing Information Science & Technology University, Beijing 102206, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Na","family":"Yan","sequence":"additional","affiliation":[{"name":"College of Computer Science, Beijing Information Science & Technology University, Beijing 102206, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiulei","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Computer Science, Beijing Information Science & Technology University, Beijing 102206, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,8,1]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Chen, C., Seff, A., Kornhauser, A., and Xiao, J. (2015, January 7\u201313). Deepdriving: Learning affordance for direct perception in autonomous driving. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.312"},{"key":"ref_2","unstructured":"Biswas, J., and Veloso, M. (2011, January 9\u201313). Depth camera based localization and navigation for indoor mobile robots. Proceedings of the IEEE International Conference on Robotics and Automation, Shanghai, China."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"961","DOI":"10.1007\/s11263-018-1070-x","article-title":"Augmented reality meets computer vision: Efficient data generation for urban driving scenes","volume":"126","author":"Alhaija","year":"2011","journal-title":"Int. J. Comput. Vis."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Kendall, A., Martirosyan, H., Dasgupta, S., Henry, P., Kennedy, R., Bachrach, A., and Bry, A. (2017, January 22\u201329). End-to-end learning of geometry and context for deep stereo regression. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.17"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Chang, J.R., and Chen, Y.S. (2018, January 18\u201323). Pyramid stereo matching network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00567"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Guo, X., Yang, K., Yang, W., Wang, X., and Li, H. (2019, January 15\u201320). Group-wise correlation stereo network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00339"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhang, F., Prisacariu, V., Yang, R., and Torr, P.H.S. (2019, January 15\u201320). Ga-net: Guided aggregation net for end-to-end stereo matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00027"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Xu, G., Cheng, J., Guo, P., and Yang, X. (2022, January 18\u201324). Attention concatenation volume for accurate and efficient stereo matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01264"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Lipson, L., Teed, Z., and Deng, J. (2021, January 1\u20133). Raft-stereo: Multilevel recurrent field transforms for stereo matching. Proceedings of the International Conference on 3D Vision, Prague, Czech Republic.","DOI":"10.1109\/3DV53792.2021.00032"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Xu, G., Wang, X., Ding, X., and Yang, X. (2023, January 18\u201322). Iterative geometry encoding volume for stereo matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02099"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Xu, G., Wang, X., Zhang, Z., Cheng, J., Liao, C., and Yang, X. (2024). IGEV++: Iterative multi-range geometry encoding volumes for stereo matching. arXiv.","DOI":"10.1109\/TPAMI.2025.3569218"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"108902","DOI":"10.1016\/j.engappai.2024.108902","article-title":"Stereo matching on images based on volume fusion and disparity space attention","volume":"136","author":"Liao","year":"2024","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"110565","DOI":"10.1016\/j.engappai.2025.110565","article-title":"Fast stereo conformer: Real-time stereo matching with enhanced feature fusion for autonomous driving","volume":"149","author":"Lu","year":"2025","journal-title":"Eng. Appl. Artif. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"129002","DOI":"10.1016\/j.neucom.2024.129002","article-title":"DCVSMNet: Double cost volume stereo matching network","volume":"618","author":"Tahmasebi","year":"2025","journal-title":"Neurocomputing"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/JPROC.2023.3238524","article-title":"Object Detection in 20 Years: A Survey","volume":"111","author":"Zou","year":"2023","journal-title":"Proc. IEEE"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Tan, M., Pang, R., and Le, Q.V. (2020, January 13\u201319). EfficientDet: Scalable and efficient object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01079"},{"key":"ref_17","unstructured":"Bhar, S.F., Alhashim, I., and Wonka, P. (2021, January 19\u201325). AdaBins: Depth estimation using adaptive bins. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Virtual."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Long, J., Shelhamer, E., and Darrell, T. (2015, January 7\u201312). Fully convolutional networks for semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yu, C., Wang, J., Peng, C., Gao, C., Yu, G., and Sang, N. (2018, January 8\u201314). BiSeNet: Bilateral segmentation network for real-time semantic segmentation. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01261-8_20"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Li, Z., Liu, X., Drenkow, N., Ding, A., Creighton, F.X., Taylor, R.H., and Unberath, M. (2021, January 11\u201317). Revisiting stereo depth estimation from a sequence-to-sequence perspective with transformers. Proceedings of the IEEE International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00614"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Khamis, S., Fanello, S., Rhemann, C., Kowdle, A., Valentin, J.L., and Izadi, S. (2018, January 8\u201314). Stereonet: Guided hierarchical refinement for real-time edge-aware depth prediction. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_35"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Xu, H., and Zhang, J. (2020, January 13\u201319). Aanet: Adaptive aggregation network for efficient stereo matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00203"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, Y., Lai, Z., Huang, G., Wang, B., Maaten, L., Campbell, M., and Weinberger, K.Q. (2019, January 20\u201324). Anytime stereo image depth estimation on mobile devices. Proceedings of the IEEE International Conference on Robotics and Automation, Montreal, BC, Canada.","DOI":"10.1109\/ICRA.2019.8794003"},{"key":"ref_24","unstructured":"Duggal, S., Wang, S., Ma, W., Hu, R., and Urtasun, R. (November, January 27). Deeppruner: Learning efficient stereo matching via differentiable patchmatch. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Xu, B., Xu, Y., Yang, X., Jia, W., and Guo, Y. (2021, January 19\u201325). Bilateral grid learning for stereo matching networks. Proceedings of the IEEE International Conference on Computer Vision, Virtual.","DOI":"10.1109\/CVPR46437.2021.01231"},{"key":"ref_26","unstructured":"Bangunharcana, A., Cho, J.W., Lee, S., Kweon, I.S., Kim, K.S., and Kim, S. (October, January 27). Correlate-and-excite: Real-time stereo matching via guided cost volume excitation. Proceedings of the IEEE International Conference on Intelligent Robots and Systems, Prague, Czech Republic."},{"key":"ref_27","unstructured":"Xu, G., Zhou, H., and Yang, X. (2023). CGI-Stereo: Accurate and real-time stereo matching via context and geometry interaction. arXiv."},{"key":"ref_28","unstructured":"Jiang, X., Bian, X., and Guo, C. (2024). Ghost-Stereo: GhostNet-based cost volume enhancement and aggregation for stereo matching networks. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Pan, B., Jiao, J., Pang, J., and Cheng, J. (2024, January 13\u201317). Distill-then-prune: An efficient compression framework for real-time stereo matching network on edge devices. Proceedings of the IEEE International Conference on Robotics and Automation, Yokohama, Japan.","DOI":"10.1109\/ICRA57147.2024.10611085"},{"key":"ref_30","unstructured":"Guo, X., Zhang, C., Zhang, Y., Zheng, W., Nie, D., Poggi, M., and Chen, L. (2025). Light Stereo: Channel boost is all you need for efficient 2D cost aggregation. arXiv."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Ronneberger, O., Fischer, P., and Brox, T. (2015, January 5\u20139). U-Net: Convolutional networks for biomedical image segmentation. Proceedings of the Medical Image Computing and Computer-Assisted Intervention, Munich, Germany.","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M.L., Zhmoginov, A., and Chen, L. (2018, January 18\u201323). MobileNetV2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE International Conference on Computer Vision, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., and Urtasun, R. (2012, January 16\u201321). Are we ready for autonomous driving? The KITTI vision benchmark suite. Proceedings of the IEEE International Conference on Computer Vision, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Menze, M., and Geiger, A. (2015, January 7\u201312). Object scene flow for autonomous vehicles. Proceedings of the IEEE International Conference on Com puter Vision, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298925"},{"key":"ref_35","unstructured":"Kingma, D.P., and Ba, J. (2015, January 7\u20139). Adam: A method for stochastic optimization. Proceedings of the 3rd International Conference on Learning Representations, San Diego, CA, USA."},{"key":"ref_36","unstructured":"Guo, X., Zhang, C., Lu, J., Duan, Y., Wang, Y., Yang, T., Zhu, Z., and Chen, L. (2023). OpenStereo: A comprehensive benchmark for stereo matching and strong baseline. arXiv."},{"key":"ref_37","unstructured":"Cheng, J., Liu, L., Xu, G., Wang, X., Zhang, Z., Deng, Y., Zang, J., Chen, Y., Cai, Z., and Yang, X. (2025). MonSter: Marry monodepth to stereo unleashes power. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"4467","DOI":"10.1007\/s11760-024-03087-3","article-title":"Guided aggregation and disparity refinement for real-time stereo matching","volume":"18","author":"Yang","year":"2024","journal-title":"Signal Image Video Process."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"24253","DOI":"10.1007\/s10489-023-04646-w","article-title":"Real-time stereo matching with high accuracy via spatial attention-guided upsampling","volume":"53","author":"Wu","year":"2023","journal-title":"Appl. Intell."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Shen, Z., Dai, Y., and Rao, Z. (2021, January 19\u201325). CFNet: Cascade and fused cost volume for robust stereo matching. Proceedings of the IEEE International Conference on Computer Vision, Virtual.","DOI":"10.1109\/CVPR46437.2021.01369"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Scharstein, D., Hirschm\u00fcller, H., Kitajima, Y., Krathwohl, G., Ne\u0161ic, N., Wang, X., and Westling, P. (2014, January 2\u20135). High-resolution stereo datasets with subpixel-accurate ground truth. Proceedings of the German Conference on Pattern Recognition, Cham, Switzerland.","DOI":"10.1007\/978-3-319-11752-2_3"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Sch\u00f6ps, T., Sch\u00f6nberger, J.L., Galliani, S., Sattler, T., Schindler, K., Pollefeys, M., and Geiger, A. (2017, January 21\u201326). A multi-view stereo benchmark with high-resolution images and multi-camera videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.272"}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/8\/1214\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:20:33Z","timestamp":1760034033000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/8\/1214"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,1]]},"references-count":42,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2025,8]]}},"alternative-id":["sym17081214"],"URL":"https:\/\/doi.org\/10.3390\/sym17081214","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,1]]}}}