{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T14:03:38Z","timestamp":1781273018477,"version":"3.54.1"},"reference-count":69,"publisher":"Springer Science and Business Media LLC","issue":"5","license":[{"start":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T00:00:00Z","timestamp":1773619200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T00:00:00Z","timestamp":1781222400000},"content-version":"vor","delay-in-days":88,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"publisher","award":["ZR2022MF274"],"award-info":[{"award-number":["ZR2022MF274"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100007129","name":"Natural Science Foundation of Shandong Province","doi-asserted-by":"publisher","award":["ZR2025MS1098"],"award-info":[{"award-number":["ZR2025MS1098"]}],"id":[{"id":"10.13039\/501100007129","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J. King Saud Univ. Comput. Inf. Sci."],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    2D convolutional neural networks (CNNs) achieve remarkable accuracy across diverse computer vision tasks, yet their inference efficiency remains suboptimal. As a fast convolution algorithm, Winograd convolution can significantly accelerate convolution operations, which constitute the primary performance bottleneck in CNNs. However, existing GPU-based implementations are restricted to accelerating standard convolution with unit dilation, while lacking support for dilated convolution. We propose TC-DWC, a GPU-based 2D Winograd convolution implementation tailored for dilated convolution, which effectively accelerates dilated convolution tasks across various configurations. TC-DWC employs a two-step tile reorganization scheme that reconciles the incompatibility between Winograd convolution and dilated convolution. Moreover, it leverages specialized high-throughput tensor cores (TCs), replacing conventional vector units to accelerate the computationally dominant matrix multiplications. To alleviate TC-DWC\u2019s substantial memory footprint, we further develop a multi-stage kernel fusion strategy that eliminates intermediate array allocations, yielding Fused TC-DWC. Experimental results demonstrate that, for single convolutional layers, TC-DWC and Fused TC-DWC deliver average speedups of\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$1.56\\times $$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mn>1.56<\/mml:mn>\n                            <mml:mo>\u00d7<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    (up to\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$2.45\\times $$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mn>2.45<\/mml:mn>\n                            <mml:mo>\u00d7<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    ) and\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$1.29\\times $$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mn>1.29<\/mml:mn>\n                            <mml:mo>\u00d7<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    (up to\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$1.80\\times $$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mn>1.80<\/mml:mn>\n                            <mml:mo>\u00d7<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    ) over cuDNN optimal implementation, respectively, with Fused TC-DWC reducing memory footprint to 35% of TC-DWC and 55% of cuDNN. For end-to-end inference of CSRNet and DLinkNet34, TC-DWC and Fused TC-DWC deliver average speedups of\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$1.38\\times $$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mn>1.38<\/mml:mn>\n                            <mml:mo>\u00d7<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    (up to\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$1.41\\times $$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mn>1.41<\/mml:mn>\n                            <mml:mo>\u00d7<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    ) and\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$1.20\\times $$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mn>1.20<\/mml:mn>\n                            <mml:mo>\u00d7<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    (up to\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:tex-math>$$1.22\\times $$<\/jats:tex-math>\n                        <mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                          <mml:mrow>\n                            <mml:mn>1.22<\/mml:mn>\n                            <mml:mo>\u00d7<\/mml:mo>\n                          <\/mml:mrow>\n                        <\/mml:math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    ) over PyTorch baseline, respectively, with Fused TC-DWC enabling support for larger batch sizes.\n                  <\/jats:p>","DOI":"10.1007\/s44443-025-00452-1","type":"journal-article","created":{"date-parts":[[2026,3,16]],"date-time":"2026-03-16T11:54:59Z","timestamp":1773662099000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Accelerating dilated Winograd convolution with fused GPU kernel using tensor cores"],"prefix":"10.1007","volume":"38","author":[{"given":"Shixiang","family":"Zhang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianguo","family":"Liang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7809-4233","authenticated-orcid":false,"given":"You","family":"Fu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rong","family":"Hua","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianzhi","family":"Yu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qianqian","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,3,16]]},"reference":[{"key":"452_CR1","doi-asserted-by":"publisher","unstructured":"Abdelkhalik H, Arafa Y, Santhi N et al. (2022) Demystifying the nvidia ampere architecture through microbenchmarking and instruction-level analysis. In: 2022 IEEE High Performance Extreme Computing Conference (HPEC), pp 1\u20138, https:\/\/doi.org\/10.1109\/HPEC55821.2022.9926299","DOI":"10.1109\/HPEC55821.2022.9926299"},{"key":"452_CR2","unstructured":"Bai S, Kolter JZ, Koltun V (2018) An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv:1803.01271"},{"key":"452_CR3","doi-asserted-by":"publisher","first-page":"213","DOI":"10.1007\/978-3-030-58452-8_13","volume-title":"Computer Vision - ECCV 2020","author":"N Carion","year":"2020","unstructured":"Carion N, Massa F, Synnaeve G et al (2020) End-to-end object detection with transformers. Computer Vision - ECCV 2020. Springer International Publishing, Cham, pp 213\u2013229"},{"key":"452_CR4","doi-asserted-by":"publisher","unstructured":"Castro RL, Andrade D, Fraguela BB (2021) Opencnn: A winograd minimal filtering algorithm implementation in cuda. Mathematics 9(17). https:\/\/doi.org\/10.3390\/math9172033","DOI":"10.3390\/math9172033"},{"key":"452_CR5","unstructured":"Chellapilla K, Puri S, Simard P (2006) High Performance Convolutional Neural Networks for Document Processing. In: Tenth International Workshop on Frontiers in Handwriting Recognition, Universit\u00e9 de Rennes 1. Suvisoft, La Baule (France). http:\/\/www.suvisoft.com"},{"issue":"4","key":"452_CR6","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","volume":"40","author":"LC Chen","year":"2018","unstructured":"Chen LC, Papandreou G, Kokkinos I et al (2018) Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans Pattern Anal Mach Intell 40(4):834\u2013848. https:\/\/doi.org\/10.1109\/TPAMI.2017.2699184","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"452_CR7","doi-asserted-by":"crossref","unstructured":"Chen T, Xu W, Chen W et al (2023) Towards efficient and accurate winograd convolution via full quantization. In: Advances in Neural Information Processing Systems, vol 36. Curran Associates Inc, pp 20164\u201320178","DOI":"10.52202\/075280-0885"},{"key":"452_CR8","unstructured":"Cho M, Brand D (2017) MEC: Memory-efficient convolution for deep neural network. In: Proceedings of the 34th International Conference on Machine Learning, Proceedings of Machine Learning Research, vol 70. PMLR, pp 815\u2013824"},{"key":"452_CR9","doi-asserted-by":"publisher","unstructured":"Chowdhury R, Silvestri F, Vella F (2020) A computational model for tensor core units. In: Proceedings of the 32nd ACM Symposium on Parallelism in Algorithms and Architectures. Association for Computing Machinery, New York, NY, USA, SPAA \u201920, p 519\u2013521. https:\/\/doi.org\/10.1145\/3350755.3400252","DOI":"10.1145\/3350755.3400252"},{"key":"452_CR10","doi-asserted-by":"crossref","unstructured":"Dakkak A, Li C, Xiong J et al. (2019) Accelerating reduction and scan using tensor core units. In: Proceedings of the ACM International Conference on Supercomputing, pp 46\u201357","DOI":"10.1145\/3330345.3331057"},{"key":"452_CR11","doi-asserted-by":"crossref","unstructured":"Demir I, Koperski K, Lindenbaum D et al. (2018) Deepglobe 2018: A challenge to parse the earth through satellite images. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp 172\u2013181","DOI":"10.1109\/CVPRW.2018.00031"},{"key":"452_CR12","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A et al. (2020) An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929"},{"key":"452_CR13","doi-asserted-by":"publisher","unstructured":"Dukhan M (2020) Indirect deconvolution algorithm. In: 2020 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), pp 922\u2013926. https:\/\/doi.org\/10.1109\/IPDPSW50202.2020.00154","DOI":"10.1109\/IPDPSW50202.2020.00154"},{"key":"452_CR14","doi-asserted-by":"publisher","unstructured":"Fan R, Wang W, Chu X (2024) Dtc-spmm: Bridging the gap in accelerating general sparse matrix multiplication with tensor cores. In: Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3. Association for Computing Machinery, New York, NY, USA, ASPLOS \u201924, p 253\u2013267. https:\/\/doi.org\/10.1145\/3620666.3651378","DOI":"10.1145\/3620666.3651378"},{"key":"452_CR15","doi-asserted-by":"publisher","DOI":"10.7717\/peerj-cs.330","volume":"7","author":"M Fasi","year":"2021","unstructured":"Fasi M, Higham NJ, Mikaitis M et al (2021) Numerical behavior of nvidia tensor cores. PeerJ Computer Science 7:e330","journal-title":"PeerJ Computer Science"},{"key":"452_CR16","doi-asserted-by":"publisher","unstructured":"Feng B, Wang Y, Chen G et al. (2021) Egemm-tc: accelerating scientific computing on tensor cores with extended precision. In: Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. Association for Computing Machinery, New York, NY, USA, PPoPP \u201921, p 278\u2013291. https:\/\/doi.org\/10.1145\/3437801.3441599","DOI":"10.1145\/3437801.3441599"},{"key":"452_CR17","unstructured":"Fu DY, Kumbong H, Nguyen E et al. (2023) Flashfftconv: Efficient convolutions for long sequences with tensor cores. arXiv:2311.05908"},{"key":"452_CR18","doi-asserted-by":"crossref","unstructured":"Girshick R, Donahue J, Darrell T et al. (2014) Rich feature hierarchies for accurate object detection and semantic segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2014.81"},{"key":"452_CR19","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S et al. (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2016.90"},{"issue":"7","key":"452_CR20","doi-asserted-by":"publisher","first-page":"986","DOI":"10.1109\/TC.2020.2973144","volume":"69","author":"L Jia","year":"2020","unstructured":"Jia L, Liang Y, Li X et al (2020) Enabling efficient fast convolution algorithms on gpus via megakernels. IEEE Trans Comput 69(7):986\u2013997. https:\/\/doi.org\/10.1109\/TC.2020.2973144","journal-title":"IEEE Trans Comput"},{"key":"452_CR21","doi-asserted-by":"publisher","unstructured":"Jia Z, Zlateski A, Durand F et al. (2018) Optimizing n-dimensional, winograd-based convolution for manycore cpus. In: Proceedings of the 23rd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. Association for Computing Machinery, New York, NY, USA, PPoPP \u201918, p 109\u2013123. https:\/\/doi.org\/10.1145\/3178487.3178496","DOI":"10.1145\/3178487.3178496"},{"issue":"1","key":"452_CR22","doi-asserted-by":"publisher","first-page":"2473","DOI":"10.1038\/s41467-020-16108-9","volume":"11","author":"V Joshi","year":"2020","unstructured":"Joshi V, Le Gallo M, Haefeli S et al (2020) Accurate deep neural network inference using computational phase-change memory. Nat Commun 11(1):2473","journal-title":"Nat Commun"},{"key":"452_CR23","unstructured":"Judd P, Albericio J, Hetherington T et al. (2015) Reduced-precision strategies for bounded memory in deep neural nets. arXiv preprint arXiv:1511.05236"},{"key":"452_CR24","doi-asserted-by":"publisher","unstructured":"Kim H, Ahn S, Oh Y et al. (2020) Duplo: Lifting redundant memory accesses of deep neural networks for gpu tensor cores. In: 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO), pp 725\u2013737. https:\/\/doi.org\/10.1109\/MICRO50266.2020.00065","DOI":"10.1109\/MICRO50266.2020.00065"},{"key":"452_CR25","doi-asserted-by":"publisher","unstructured":"Kim M, Park C, Kim S et al. (2019) Efficient dilated-winograd convolutional neural networks. In: 2019 IEEE International Conference on Image Processing (ICIP), pp 2711\u20132715. https:\/\/doi.org\/10.1109\/ICIP.2019.8803277","DOI":"10.1109\/ICIP.2019.8803277"},{"key":"452_CR26","unstructured":"Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: Advances in Neural Information Processing Systems, vol 25. Curran Associates Inc"},{"key":"452_CR27","doi-asserted-by":"crossref","unstructured":"Lavin A, Gray S (2016) Fast algorithms for convolutional neural networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2016.435"},{"issue":"7","key":"452_CR28","doi-asserted-by":"publisher","first-page":"1878","DOI":"10.1109\/TPDS.2020.3045828","volume":"32","author":"A Li","year":"2021","unstructured":"Li A, Su S (2021) Accelerating binarized neural networks via bit-tensor-cores in turing gpus. IEEE Trans Parallel Distrib Syst 32(7):1878\u20131891. https:\/\/doi.org\/10.1109\/TPDS.2020.3045828","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"452_CR29","doi-asserted-by":"crossref","unstructured":"Li Y, Zhang X, Chen D (2018) Csrnet: Dilated convolutional neural networks for understanding the highly congested scenes. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2018.00120"},{"key":"452_CR30","doi-asserted-by":"publisher","unstructured":"Liu J, Yang D, Lai J (2021) Optimizing winograd-based convolution with tensor cores. In: Proceedings of the 50th International Conference on Parallel Processing. Association for Computing Machinery, New York, NY, USA, ICPP \u201921. https:\/\/doi.org\/10.1145\/3472456.3472473","DOI":"10.1145\/3472456.3472473"},{"key":"452_CR31","unstructured":"Liu X, Pool J, Han S et al. (2018) Efficient sparse-winograd convolutional neural networks. arXiv preprint arXiv:1802.06367"},{"key":"452_CR32","doi-asserted-by":"publisher","unstructured":"Liu X, Zheng X, Yang H et al. (2024) Tetris: Accelerating sparse convolution by exploiting memory reuse on gpu. In: Proceedings of the 29th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming. Association for Computing Machinery, New York, NY, USA, PPoPP \u201924, p 229\u2013242. https:\/\/doi.org\/10.1145\/3627535.3638471","DOI":"10.1145\/3627535.3638471"},{"issue":"1","key":"452_CR33","doi-asserted-by":"publisher","first-page":"70","DOI":"10.1109\/TPDS.2021.3084813","volume":"33","author":"G Lu","year":"2022","unstructured":"Lu G, Zhang W, Wang Z (2022) Optimizing depthwise separable convolution operations on gpus. IEEE Trans Parallel Distrib Syst 33(1):70\u201387. https:\/\/doi.org\/10.1109\/TPDS.2021.3084813","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"452_CR34","doi-asserted-by":"publisher","unstructured":"M V, Pinto R, (2025) High-performance winograd based accelerator architecture for convolutional neural network. IEEE Comput Archit Lett 24(1):21\u201324. https:\/\/doi.org\/10.1109\/LCA.2025.3525970","DOI":"10.1109\/LCA.2025.3525970"},{"key":"452_CR35","unstructured":"Mathieu M, Henaff M, LeCun Y (2014) Fast training of convolutional networks through ffts. arXiv:1312.5851"},{"key":"452_CR36","doi-asserted-by":"publisher","unstructured":"Mori P, Sampath SB, Frickenstein L et al. (2023) Winotrain: Winograd-aware training for accurate full 8-bit convolution acceleration. In: 2023 60th ACM\/IEEE Design Automation Conference (DAC), pp 1\u20136. https:\/\/doi.org\/10.1109\/DAC56929.2023.10247805","DOI":"10.1109\/DAC56929.2023.10247805"},{"key":"452_CR37","doi-asserted-by":"crossref","unstructured":"Nam H, Han B (2016) Learning multi-domain convolutional neural networks for visual tracking. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2016.465"},{"key":"452_CR38","unstructured":"NVIDIA (2017) Nvidia tesla v100 gpu architecture. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf, accessed on July 29, 2025"},{"key":"452_CR39","unstructured":"NVIDIA (2020) Nvidia a100 tensor core gpu architecture. https:\/\/images.nvidia.com\/aem-dam\/en-zz\/Solutions\/data-center\/nvidia-ampere-architecture-whitepaper.pdf, accessed on July 29, 2025"},{"key":"452_CR40","unstructured":"NVIDIA (2023) Nvidia ada gpu architecture. https:\/\/images.nvidia.cn\/aem-dam\/Solutions\/geforce\/ada\/nvidia-ada-gpu-architecture.pdf, accessed on July 29, 2025"},{"key":"452_CR41","unstructured":"NVIDIA (2025a) Cuda c++ best practices guide. https:\/\/docs.nvidia.com\/cuda\/cuda-c-best-practices-guide\/index.html, accessed on July 29, 2025"},{"key":"452_CR42","unstructured":"NVIDIA (2025b) Cuda c++ programming guide. https:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide\/index.html, accessed on July 29, 2025"},{"key":"452_CR43","unstructured":"NVIDIA (2025c) Nvidia cudnn documentation. https:\/\/docs.nvidia.com\/deeplearning\/cudnn\/latest\/index.html accessed on July 29, 2025"},{"key":"452_CR44","unstructured":"NVIDIA (2025d) Parallel thread execution isa. https:\/\/docs.nvidia.com\/cuda\/parallel-thread-execution\/index.html, accessed on July 29, 2025"},{"key":"452_CR45","doi-asserted-by":"publisher","unstructured":"Okanovic P, Kwasniewski G, Labini PS et al. (2024) High performance unstructured spmm computation using tensor cores. In: SC24: International Conference for High Performance Computing, Networking, Storage and Analysis, pp 1\u201314. https:\/\/doi.org\/10.1109\/SC41406.2024.00060","DOI":"10.1109\/SC41406.2024.00060"},{"key":"452_CR46","unstructured":"van den Oord A, Dieleman S, Zen H et al. (2016) Wavenet: A generative model for raw audio. arXiv:1609.03499"},{"issue":"4","key":"452_CR47","doi-asserted-by":"publisher","first-page":"475","DOI":"10.1177\/10943420221090256","volume":"36","author":"H Ootomo","year":"2022","unstructured":"Ootomo H, Yokota R (2022) Recovering single precision accuracy from tensor cores while surpassing the fp32 theoretical peak performance. The International Journal of High Performance Computing Applications 36(4):475\u2013491. https:\/\/doi.org\/10.1177\/10943420221090256","journal-title":"The International Journal of High Performance Computing Applications"},{"issue":"4","key":"452_CR48","doi-asserted-by":"publisher","first-page":"297","DOI":"10.1177\/10943420241239588","volume":"38","author":"H Ootomo","year":"2024","unstructured":"Ootomo H, Ozaki K, Yokota R (2024) Dgemm on integer matrix multiplication unit. The International Journal of High Performance Computing Applications 38(4):297\u2013313. https:\/\/doi.org\/10.1177\/10943420241239588","journal-title":"The International Journal of High Performance Computing Applications"},{"key":"452_CR49","unstructured":"Park C, Yoon MK, Park M et al. (2024) Winograd structured pruning"},{"key":"452_CR50","unstructured":"Pybind11 (2025) Pybind11 documentation. https:\/\/pybind11.readthedocs.io\/en\/stable\/index.html, accessed on July 29, 2025"},{"key":"452_CR51","unstructured":"PyTorch (2025) Pytorch documentation. https:\/\/pytorch.org\/docs\/stable\/torch.html, accessed on July 29, 2025"},{"key":"452_CR52","unstructured":"Qin Z, Lin M, Lin W (2023) Low-rank winograd transformation for 3d convolutional neural networks. arXiv:2301.11180"},{"key":"452_CR53","doi-asserted-by":"publisher","unstructured":"Raihan MA, Goli N, Aamodt TM (2019) Modeling deep learning accelerator enabled gpus. In: 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), pp 79\u201392. https:\/\/doi.org\/10.1109\/ISPASS.2019.00016","DOI":"10.1109\/ISPASS.2019.00016"},{"key":"452_CR54","unstructured":"Salimans T, Karpathy A, Chen X et al. (2017) Pixelcnn++: Improving the pixelcnn with discretized logistic mixture likelihood and other modifications. arXiv:1701.05517"},{"issue":"7","key":"452_CR55","doi-asserted-by":"publisher","first-page":"1442","DOI":"10.1109\/TCAD.2019.2912894","volume":"39","author":"J Shen","year":"2020","unstructured":"Shen J, Huang Y, Wen M et al (2020) Toward an efficient deep pipelined template-based architecture for accelerating the entire 2-d and 3-d cnns on fpga. IEEE Trans Comput Aided Des Integr Circuits Syst 39(7):1442\u20131455. https:\/\/doi.org\/10.1109\/TCAD.2019.2912894","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"452_CR56","doi-asserted-by":"publisher","unstructured":"Shi J, Li S, Xu Y et al. (2025) Flashsparse: Minimizing computation redundancy for fast sparse matrix multiplications on tensor cores. In: Proceedings of the 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming. Association for Computing Machinery, New York, NY, USA, PPoPP \u201925, p 312\u2013325. https:\/\/doi.org\/10.1145\/3710848.3710858","DOI":"10.1145\/3710848.3710858"},{"key":"452_CR57","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556"},{"key":"452_CR58","doi-asserted-by":"crossref","unstructured":"Tong G, Yan R, Yang L et al (2022) Optimizing winograd convolution on gpus via partial kernel fusion. Network and Parallel Computing. Springer Nature Switzerland, Cham, pp 17\u201329","DOI":"10.1007\/978-3-031-21395-3_2"},{"issue":"11","key":"452_CR59","doi-asserted-by":"publisher","first-page":"4290","DOI":"10.1109\/TCAD.2020.3012323","volume":"39","author":"X Wang","year":"2020","unstructured":"Wang X, Wang C, Cao J et al (2020) Winonn: Optimizing fpga-based convolutional neural network accelerators using sparse winograd algorithm. IEEE Trans Comput Aided Des Integr Circuits Syst 39(11):4290\u20134302. https:\/\/doi.org\/10.1109\/TCAD.2020.3012323","journal-title":"IEEE Trans Comput Aided Des Integr Circuits Syst"},{"key":"452_CR60","doi-asserted-by":"publisher","unstructured":"Xie K, Lu Y, He X et al. (2024) Winols: A large-tiling sparse winograd cnn accelerator on fpgas. ACM Trans Archit Code Optim 21(2). https:\/\/doi.org\/10.1145\/3643682","DOI":"10.1145\/3643682"},{"key":"452_CR61","doi-asserted-by":"publisher","unstructured":"Yan D, Wang W, Chu X (2020) Optimizing batched winograd convolution on gpus. In: Proceedings of the 25th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming. Association for Computing Machinery, New York, NY, USA, PPoPP \u201920, p 32\u201344. https:\/\/doi.org\/10.1145\/3332466.3374520","DOI":"10.1145\/3332466.3374520"},{"key":"452_CR62","doi-asserted-by":"publisher","unstructured":"Yan J, Jiang W, He D et al. (2025) Rt-gnn: Accelerating sparse graph neural networks by tensor-cuda kernel fusion. ACM Trans Archit Code Optim 22(1). https:\/\/doi.org\/10.1145\/3702001","DOI":"10.1145\/3702001"},{"key":"452_CR63","doi-asserted-by":"publisher","unstructured":"Yang C, Meng Y, Xi J et al. (2024) Wra-ss: A high-performance accelerator integrating winograd with structured sparsity for convolutional neural networks. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 32(1), 164\u2013177. https:\/\/doi.org\/10.1109\/TVLSI.2023.3330993","DOI":"10.1109\/TVLSI.2023.3330993"},{"key":"452_CR64","doi-asserted-by":"publisher","unstructured":"Yepez J, Ko SB (2020) Stride 2 1-d, 2-d, and 3-d winograd for convolutional neural networks. IEEE Transactions on Very Large Scale Integration (VLSI) Systems 28(4), 853\u2013863. https:\/\/doi.org\/10.1109\/TVLSI.2019.2961602","DOI":"10.1109\/TVLSI.2019.2961602"},{"key":"452_CR65","unstructured":"Yu F, Koltun V (2016) Multi-scale context aggregation by dilated convolutions. arXiv:1511.07122"},{"key":"452_CR66","doi-asserted-by":"crossref","unstructured":"Zhang Y, Zhou D, Chen S et al. (2016) Single-image crowd counting via multi-column convolutional neural network. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 589\u2013597","DOI":"10.1109\/CVPR.2016.70"},{"key":"452_CR67","doi-asserted-by":"publisher","unstructured":"Zhang Z, Zhang P, Xu Z et al. (2024) Im2col-winograd: An efficient and flexible fused-winograd convolution for nhwc format on gpus. In: Proceedings of the 53rd International Conference on Parallel Processing. Association for Computing Machinery, New York, NY, USA, ICPP \u201924, p 1072\u20131081, https:\/\/doi.org\/10.1145\/3673038.3673039","DOI":"10.1145\/3673038.3673039"},{"key":"452_CR68","doi-asserted-by":"publisher","unstructured":"Zhao H, Li S, Wang J et al. (2025) Acc-spmm: Accelerating general-purpose sparse matrix-matrix multiplication with gpu tensor cores. In: Proceedings of the 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming. Association for Computing Machinery, New York, NY, USA, PPoPP \u201925, p 326\u2013338. https:\/\/doi.org\/10.1145\/3710848.3710888","DOI":"10.1145\/3710848.3710888"},{"key":"452_CR69","doi-asserted-by":"crossref","unstructured":"Zhou L, Zhang C, Wu M (2018) D-linknet: Linknet with pretrained encoder and dilated convolution for high resolution satellite imagery road extraction. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pp 182\u2013186","DOI":"10.1109\/CVPRW.2018.00034"}],"container-title":["Journal of King Saud University Computer and Information Sciences"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s44443-025-00452-1","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44443-025-00452-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s44443-025-00452-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,12]],"date-time":"2026-06-12T13:48:22Z","timestamp":1781272102000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s44443-025-00452-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,16]]},"references-count":69,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["452"],"URL":"https:\/\/doi.org\/10.1007\/s44443-025-00452-1","relation":{},"ISSN":["1319-1578","2213-1248"],"issn-type":[{"value":"1319-1578","type":"print"},{"value":"2213-1248","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,16]]},"assertion":[{"value":"22 September 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 December 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}},{"value":"This study does not involve human participants or animals and thus does not require ethical approval.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}}],"article-number":"225"}}