{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,31]],"date-time":"2025-12-31T22:28:27Z","timestamp":1767220107405,"version":"build-2065373602"},"reference-count":48,"publisher":"MDPI AG","issue":"20","license":[{"start":{"date-parts":[[2021,10,13]],"date-time":"2021-10-13T00:00:00Z","timestamp":1634083200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Despite recent stereo matching algorithms achieving significant results on public benchmarks, the problem of requiring heavy computation remains unsolved. Most works focus on designing an architecture to reduce the computational complexity, while we take aim at optimizing 3D convolution kernels on the Pyramid Stereo Matching Network (PSMNet) for solving the problem. In this paper, we design a series of comparative experiments exploring the performance of well-known convolution kernels on PSMNet. Our model saves the computational complexity from 256.66 G MAdd (Multiply-Add operations) to 69.03 G MAdd (198.47 G MAdd to 10.84 G MAdd for only considering 3D convolutional neural networks) without losing accuracy. On Scene Flow and KITTI 2015 datasets, our model achieves results comparable to the state-of-the-art with a low computational cost.<\/jats:p>","DOI":"10.3390\/s21206808","type":"journal-article","created":{"date-parts":[[2021,10,13]],"date-time":"2021-10-13T21:48:39Z","timestamp":1634161719000},"page":"6808","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Optimizing 3D Convolution Kernels on Stereo Matching for Resource Efficient Computations"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5054-0318","authenticated-orcid":false,"given":"Jianqiang","family":"Xiao","sequence":"first","affiliation":[{"name":"Division of Electrical Engineering and Computer Science, Kanazawa University, Kanazawa 920-1192, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dianbo","family":"Ma","sequence":"additional","affiliation":[{"name":"Division of Electrical Engineering and Computer Science, Kanazawa University, Kanazawa 920-1192, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Satoshi","family":"Yamane","sequence":"additional","affiliation":[{"name":"Division of Electrical Engineering and Computer Science, Kanazawa University, Kanazawa 920-1192, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,13]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Zenati, N., and Zerhouni, N. (2007, January 24\u201327). Dense Stereo Matching with Application to Augmented Reality. Proceedings of the IEEE International Conference on Signal Processing and Communications, Dubai, United Arab Emirates.","DOI":"10.1109\/ICSPC.2007.4728616"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Noh, Z., S, M., and Pan, Z. (2009, January 9\u201311). A Review on Augmented Reality for Virtual Heritage System. Proceedings of the Springer International Conference on Technologies for E-Learning and Digital Entertainment, Banff, AB, Canada.","DOI":"10.1007\/978-3-642-03364-3_7"},{"key":"ref_3","unstructured":"H\u00e4ne, C., Sattler, T., and Pollefeys, M. (October, January 28). Obstacle detection for self-driving cars using only monocular cameras and wheel odometry. Proceedings of the IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Hamburg, Germany."},{"key":"ref_4","unstructured":"Nalpantidis, L., Sirakoulis, G.C., and Gasteratos, A. (2007, January 21\u201323). Review of stereo matching algorithms for 3D vision. Proceedings of the 16th International Symposium on Measurement and Control in Robotic, Warsaw, Poland."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Samadi, M., and Othman, M.F. (2013, January 9\u201310). A New Fast and Robust Stereo Matching Algorithm for Robotic Systems. Proceedings of the 9th International Conference on Computing and InformationTechnology (IC2IT2013), Bangkok, Thailand.","DOI":"10.1007\/978-3-642-37371-8_31"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zbontar, J., and LeCun, Y. (2015, January 7\u201312). Computing the stereo matching cost with a convolutional neural network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298767"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1023\/A:1014573219977","article-title":"A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms","volume":"47","author":"Scharstein","year":"2002","journal-title":"Int. J. Comput. Vis."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"331","DOI":"10.1109\/TPAMI.2007.36","article-title":"Estimating optimal parameters for MRF stereo from a single image pair","volume":"29","author":"Zhang","year":"2007","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell. (TPAMI)"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Scharstein, D., and Pal, C. (2007, January 17\u201322). Learning Conditional Random Fields for Stereo. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383191"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Haeusler, R., Nair, R., and Kondermann, D. (2013, January 23\u201328). Ensemble Learning for Confidence Measures in Stereo Vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Portland, OR, USA.","DOI":"10.1109\/CVPR.2013.46"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Kendall, A., Martirosyan, H., Dasgupta, S., Henry, P., Kennedy, R., Bachrach, A., and Bry, A. (2017, January 22\u201329). End-to-end learning of geometry and context for deep stereo regression. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy.","DOI":"10.1109\/ICCV.2017.17"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Chang, J.-R., and Chen, Y.-S. (2018, January 18\u201322). Pyramid Stereo Matching Network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00567"},{"key":"ref_13","unstructured":"Iandola, F.N., Moskewicz, M.W., Ashraf, K., Han, S., Dally, W.J., and Keutzer, K. (2016). SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size. arXiv."},{"key":"ref_14","unstructured":"Krizhevsky, A., Sutskever, I., and Hinton, G. (2012, January 3\u20136). ImageNet Classification with Deep Convolutional Neural Networks. Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), Lake Tahoe, NV, USA."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017, January 22\u201325). Xception: Deep Learning With Depthwise Separable Convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_16","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. (2018, January 18\u201322). Mobilenetv2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2018, January 18\u201322). Shufflenet: An extremely efficient convolutional neural network for mobile devices. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ma, N., Zhang, X., Zheng, H.T., and Sun, J. (2018, January 8\u201314). ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Desig. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201322). Squeeze-and-Excitation Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, W., Hu, X., and Yang, J. (2019, January 15\u201320). Selective Kernel Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00060"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Wang, Q., Wu, B., Zhu, P., Li, P., Zuo, W., and Hu, Q. (2020, January 13\u201319). ECA-Net: Efficient Channel Attention for Deep Convolutional Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01155"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"8562323","DOI":"10.1155\/2020\/8562323","article-title":"Review of Stereo Matching Algorithms Based on Deep Learning","volume":"2020","author":"Zhou","year":"2020","journal-title":"Computational Intelligence and Neuroscience"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Guo, X., Yang, K., Yang, W., Wang, X., and Li, H. (2019, January 15\u201320). Group-wise correlation stereo network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00339"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Khamis, S., Fanello, S., Rhemann, C., Kowdle, A., Valentin, J., and Izadi, S. (2018, January 8\u201314). StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_35"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhang, F., Prisacariu, V., Yang, R., and Torr, P.H. (2019, January 15\u201320). GA-Net: Guided Aggregation Net for End-To-End Stereo Matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00027"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Zhu, C., and Chang, Y. (2019). Hierarchical Guided-Image-Filtering for Efficient Stereo Matching. Appl. Sci., 15.","DOI":"10.3390\/app9153122"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Chen, Y., Bai, X., Yu, S., Yu, K., Li, Z., and Yang, K. (2020, January 7\u201312). Adaptive Unimodal Cost Volume Filtering for Deep Stereo Matching. Proceedings of the AAAI Conference on Artificial Intelligence, New York, NY, USA.","DOI":"10.1609\/aaai.v34i07.6991"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Duggal, S., Wang, S., Ma, W., Hu, R., and Urtasun, R. (2019, January 27\u201328). DeepPruner: Learning Efficient Stereo Matching via Differentiable PatchMatch. Proceedings of the IEEE International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00448"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"24","DOI":"10.1145\/1531326.1531330","article-title":"PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing","volume":"28","author":"Barnes","year":"2009","journal-title":"Acm Trans. Graph. (Proc. Siggraph)"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Xu, H., and Zhang, J. (2020, January 13\u201319). AANet: Adaptive Aggregation Network for Efficient Stereo Matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00203"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Gu, X., Fan, Z., Zhu, S., Dai, Z., Tan, F., and Tan, P. (2020, January 13\u201319). Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00257"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Shen, Z., Dai, Y., and Rao, Z. (2021, January 19\u201325). CFNet: Cascade and Fused Cost Volume for Robust Stereo Matching. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Virtual.","DOI":"10.1109\/CVPR46437.2021.01369"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wei, M., Zhu, M., Wu, Y., Sun, J., Wang, J., and Liu, C. (2021). A Fast Stereo Matching Network with Multi-Cross Attention. Sensors, 18.","DOI":"10.3390\/s21186016"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Huang, Z., Gu, J., Li, J., and Yu, X. (2021). A stereo matching algorithm based on the improved PSMNet. PLoS ONE, 25.","DOI":"10.1371\/journal.pone.0251657"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R., Doll\u00e1r, P., Tu, Z., and He, K. (2017, January 21\u201326). Aggregated Residual Transformations for Deep Neural Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_37","unstructured":"Cheng, L., Papandreou, G., Schroff, F., and Adam, H. (1706). Rethinking Atrous Convolution for Semantic Image Segmentation. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Jia, X., Chen, W., Liang, Z., Luo, X., Wu, M., Li, C., He, Y., Tan, Y., and Huang, L. (2021). A Joint 2D-3D Complementary Network for Stereo Matching. Sensors, 21.","DOI":"10.3390\/s21041430"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Vasileiadis, M., Bouganis, C.-S., Stavropoulos, G., and Tzovaras, D. (2019, January 9\u201312). Optimising 3D-CNN Design towards Human Pose Estimation on Low Power Devices. Proceedings of the British Machine Vision Conference (BMVC), Cardiff, UK.","DOI":"10.1016\/j.cviu.2019.04.011"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_41","unstructured":"Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J. (2017, January 24\u201326). Pruning Convolutional Neural Networks for Resource Efficient Inference. Proceedings of the International Conference on Learning Representations (ICLR), Toulon, France."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Mayer, N., Ilg, E., Hausser, P., Fischer, P., Cremers, D., Dosovitskiy, A., and Brox, T. (2016, January 27\u201330). A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.438"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Menze, M., and Geiger, A. (2015, January 7\u201312). Object scene flow for autonomous vehicles. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298925"},{"key":"ref_44","unstructured":"Ng, A.Y. (1997, January 8\u201312). Preventing \u201coverfitting\u201d of cross-validation data. Machine Learning. Proceedings of the International Conference on Machine Learning (ICML), Nashville, TN, USA."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Tosi, F., Liao, Y., Schmitt, C., and Geiger, A. (2021, January 19\u201325). SMD-Nets: Stereo Mixture Density Networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), virtual.","DOI":"10.1109\/CVPR46437.2021.00883"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Yee, K., and Chakrabarti, A. (2020, January 1\u20135). Fast Deep Stereo with 2D Convolutional Processing of Cost Signatures. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Snowmass Village, CO, USA.","DOI":"10.1109\/WACV45572.2020.9093273"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Qin, Z., Zhang, Z., Li, D., Zhang, Y., and Peng, Y. (2018, January 8\u201313). Diagonalwise Refactorization: An Efficient Training Method for Depthwise Convolutions. Proceedings of the 2018 International Joint Conference on Neural Networks (IJCNN), Rio de Janeiro, Brazil.","DOI":"10.1109\/IJCNN.2018.8489312"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1109\/TPDS.2021.3084813","article-title":"Optimizing Depthwise Separable Convolution Operations on GPUs","volume":"33","author":"Lu","year":"2021","journal-title":"IEEE Trans. Parallel Distrib. Syst."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/20\/6808\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:12:44Z","timestamp":1760166764000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/20\/6808"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,13]]},"references-count":48,"journal-issue":{"issue":"20","published-online":{"date-parts":[[2021,10]]}},"alternative-id":["s21206808"],"URL":"https:\/\/doi.org\/10.3390\/s21206808","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,10,13]]}}}