{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T05:55:02Z","timestamp":1783749302504,"version":"3.55.0"},"reference-count":181,"publisher":"Springer Science and Business Media LLC","issue":"12","license":[{"start":{"date-parts":[[2023,5,1]],"date-time":"2023-05-01T00:00:00Z","timestamp":1682899200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2023,5,1]],"date-time":"2023-05-01T00:00:00Z","timestamp":1682899200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Artif Intell Rev"],"published-print":{"date-parts":[[2023,12]]},"DOI":"10.1007\/s10462-023-10489-1","type":"journal-article","created":{"date-parts":[[2023,5,1]],"date-time":"2023-05-01T11:11:23Z","timestamp":1682939483000},"page":"14257-14295","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":23,"title":["Dimensionality reduced training by pruning and freezing parts of a deep neural network: a survey"],"prefix":"10.1007","volume":"56","author":[{"given":"Paul","family":"Wimmer","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jens","family":"Mehnert","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alexandru Paul","family":"Condurache","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,5,1]]},"reference":[{"key":"10489_CR1","unstructured":"Aladago MM, Torresani L (2021) Slot machines: discovering winning combinations of random weights in neural networks. In: Proceedings of the 38th international conference on machine learning"},{"key":"10489_CR2","unstructured":"Alizadeh M, Tailor SA, Zintgraf LM et al (2022) Prospect pruning: finding trainable weights at initialization using meta-gradients. In: 10th International conference on learning representations"},{"key":"10489_CR3","unstructured":"Amodei D, Hernandez D, Sastry G et al (2018) AI and compute. OpenAI Blog. https:\/\/openai.com\/blog\/ai-and-compute\/. accessed 16 Nov 2022"},{"issue":"3","key":"10489_CR4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3005348","volume":"13","author":"S Anwar","year":"2017","unstructured":"Anwar S, Hwang K, Sung W (2017) Structured pruning of deep convolutional neural networks. ACM J Emerg Technol Comput Syst 13(3):1\u201318","journal-title":"ACM J Emerg Technol Comput Syst"},{"key":"10489_CR5","unstructured":"Arora S, Ge R, Neyshabur B et al (2018) Stronger generalization bounds for deep nets via a compression approach. In: Proceedings of the 35th international conference on machine learning, 2018"},{"key":"10489_CR6","unstructured":"Arora S, Du SS, Hu W et al (2019) On exact computation with an infinitely wide neural net. In: Advances in neural information processing systems, 2019, vol 32"},{"key":"10489_CR7","unstructured":"Bai Y, Wang H, Tao Z et al (2022) Dual lottery ticket hypothesis. In: 10th International conference on learning representations"},{"key":"10489_CR8","unstructured":"Barsbey M, Sefidgaran M, Erdogdu MA et al (2021) Heavy tails in SGD and compressibility of overparametrized neural networks. In: Advances in neural information processing systems, 2021, vol 34"},{"key":"10489_CR9","unstructured":"Bartoldson B, Morcos A, Barbu A et al (2020) The generalization-stability tradeoff in neural network pruning. In: Advances in neural information processing systems, 2020, vol 33"},{"key":"10489_CR10","unstructured":"Bellec G, Kappel D, Maass W et al (2018) Deep rewiring: training very sparse deep networks. In: 6th International conference on learning representations"},{"key":"10489_CR11","unstructured":"Bengio Y, L\u00e9onard N, Courville AC (2013) Estimating or propagating gradients through stochastic neurons for conditional computation. CoRR abs\/1308.3432. arXiv:1308.3432. Accessed 31 Oct 2022"},{"key":"10489_CR12","unstructured":"Blalock DW, Ortiz JJG, Frankle J et al (2020) What is the state of neural network pruning? In: Proceedings of machine learning and systems, 2020, vol 2"},{"key":"10489_CR13","unstructured":"Brutzkus A, Globerson A, Malach E et al (2018) SGD learns over-parameterized networks that provably generalize on linearly separable data. In: 6th International conference on learning representations"},{"key":"10489_CR14","unstructured":"Burkholz R, Laha N, Mukherjee R et al (2022) On the existence of universal lottery tickets. In: 10th International conference on learning representations"},{"key":"10489_CR15","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1016\/j.neuroscience.2017.06.005","volume":"357","author":"AR Chambers","year":"2017","unstructured":"Chambers AR, Rumpel S (2017) A stable brain from unstable components: emerging concepts and implications for neural computation. Neuroscience 357:172\u2013184","journal-title":"Neuroscience"},{"key":"10489_CR16","unstructured":"Chen W, Wilson J, Tyree S et al (2015) Compressing neural networks with the hashing trick. In: Proceedings of the 32nd international conference on machine learning"},{"key":"10489_CR17","unstructured":"Chen J, Chen S, Pan SJ (2020a) Storage efficient and dynamic flexible runtime channel pruning via deep reinforcement learning. In: Advances in neural information processing systems, 2020, vol 33"},{"key":"10489_CR18","unstructured":"Chen T, Frankle J, Chang S et al (2020b) The lottery ticket hypothesis for pre-trained BERT networks. In: Advances in neural information processing systems, 2020, vol 33"},{"key":"10489_CR19","doi-asserted-by":"crossref","unstructured":"Chen T, Frankle J, Chang S et al (2021a) The lottery tickets hypothesis for supervised and self-supervised pre-training in computer vision models. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR46437.2021.01604"},{"key":"10489_CR20","unstructured":"Chen X, Chen T, Zhang Z et al (2021b) You are caught stealing my winning lottery ticket! Making a lottery ticket claim its ownership. In: Advances in neural information processing systems, 2021, vol 34"},{"key":"10489_CR21","unstructured":"Chen X, Zhang J, Wang Z (2022) Peek-a-boo: what (more) is disguised in a randomly weighted neural network, and how to find it efficiently. In: 10th International conference on learning representations"},{"key":"10489_CR22","unstructured":"Chijiwa D, Yamaguchi S, Ida Y et al (2021) Pruning randomly initialized neural networks with iterative randomization. In: Advances in neural information processing systems, 2021, vol 34"},{"key":"10489_CR23","unstructured":"Courbariaux M, Bengio Y, David JP (2015) Binaryconnect: training deep neural networks with binary weights during propagations. In: Advances in neural information processing systems, 2015, vol 28"},{"key":"10489_CR24","unstructured":"Da Cunha A, Natale E, Viennot L (2022) Proving the lottery ticket hypothesis for convolutional neural networks. In: 10th International conference on learning representations"},{"key":"10489_CR25","unstructured":"De Jorge P, Sanyal A, Behl H et al (2021) Progressive skeletonization: trimming more fat from a network at initialization. In: 9th International conference on learning representations"},{"key":"10489_CR26","doi-asserted-by":"crossref","unstructured":"Deng J, Dong W, Socher R et al (2009) ImageNet: a large-scale hierarchical image database. In: Proceedings of the IEEE conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"10489_CR27","unstructured":"Denton E, Zaremba W, Bruna J et al (2014) Exploiting linear structure within convolutional networks for efficient evaluation. In: Advances in neural information processing systems, 2014, vol 27"},{"key":"10489_CR28","unstructured":"Dettmers T, Zettlemoyer L (2019) Sparse networks from scratch: faster training without losing performance. CoRR abs\/1907.04840v2. arXiv:1907.04840v2. Accessed 2 Oct 2022"},{"key":"10489_CR29","unstructured":"Diffenderfer J, Kailkhura B (2021) Multi-prize lottery ticket hypothesis: finding accurate binary neural networks by pruning a randomly weighted network. In: 9th International conference on learning representations"},{"key":"10489_CR30","unstructured":"Diffenderfer J, Bartoldson BR, Chaganti S et al (2021) A winning hand: compressing deep networks can improve out-of-distribution robustness. In: Advances in neural information processing systems, 2021, vol 34"},{"key":"10489_CR31","unstructured":"Ding X, Ding G, Zhou X et al (2019) Global sparse momentum SGD for pruning very deep neural networks. In: Advances in neural information processing systems, 2019, vol 32"},{"key":"10489_CR32","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A et al (2021) An image is worth 16 \u00d7 words: transformers for image recognition at scale. In: 9th International conference on learning representations"},{"key":"10489_CR33","unstructured":"Du SS, Zhai X, Poczos B et al (2019) Gradient descent provably optimizes over-parameterized neural networks. In: 7th International conference on learning representations"},{"issue":"61","key":"10489_CR34","first-page":"2121","volume":"12","author":"J Duchi","year":"2011","unstructured":"Duchi J, Hazan E, Singer Y (2011) Adaptive subgradient methods for online learning and stochastic optimization. J Mach Learn Res 12(61):2121\u20132159","journal-title":"J Mach Learn Res"},{"key":"10489_CR35","unstructured":"Elesedy B, Kanade V, Teh YW (2021) Lottery tickets in linear models: an analysis of iterative magnitude pruning. In: Sparsity in neural networks workshop"},{"key":"10489_CR36","doi-asserted-by":"crossref","unstructured":"Elsen E, Dukhan M, Gale T et al (2020) Fast sparse ConvNets. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR42600.2020.01464"},{"key":"10489_CR37","doi-asserted-by":"publisher","first-page":"290","DOI":"10.5486\/PMD.1959.6.3-4.12","volume":"6","author":"P Erd\u0151s","year":"1959","unstructured":"Erd\u0151s P, R\u00e9nyi A (1959) On random graphs I. Publ Math Debr 6:290\u2013297","journal-title":"Publ Math Debr"},{"key":"10489_CR38","unstructured":"Evci U, Gale T, Menick J et al (2020) Rigging the lottery: making all tickets winners. In: Proceedings of the 37th international conference on machine learning"},{"key":"10489_CR39","doi-asserted-by":"crossref","unstructured":"Evci U, Dauphin Y, Ioannou Y et al (2022) Gradient flow in sparse neural networks and how lottery tickets win. In: Proceedings of the AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v36i6.20611"},{"key":"10489_CR40","unstructured":"Fischer J, Burkholz R (2022) Plant \u2018n\u2019 seek: can you find the winning ticket? In: 10th International conference on learning representations"},{"key":"10489_CR41","unstructured":"Frankle J, Carbin M (2018) The lottery ticket hypothesis: finding sparse, trainable neural networks. In: 6th International conference on learning representations"},{"key":"10489_CR42","unstructured":"Frankle J, Dziugaite GK, Roy D et al (2020a) Linear mode connectivity and the lottery ticket hypothesis. In: Proceedings of the 37th international conference on machine learning"},{"key":"10489_CR43","unstructured":"Frankle J, Schwab DJ, Morcos AS (2020b) The early phase of neural network training. In: 8th International conference on learning representations"},{"key":"10489_CR44","unstructured":"Frankle J, Dziugaite GK, Roy D et al (2021a) Pruning neural networks at initialization: why are we missing the mark? In: 9th International conference on learning representations"},{"key":"10489_CR45","unstructured":"Frankle J, Schwab DJ, Morcos AS (2021b) Training batchnorm and only batchnorm: on the expressive power of random features in CNNs. In: 9th International conference on learning representations"},{"key":"10489_CR46","unstructured":"Gale T, Elsen E, Hooker S (2019) The state of sparsity in deep neural networks. In: 36th International conference on machine learning joint workshop on on-device machine learning and compact deep neural network representations (ODML-CDNNR)"},{"key":"10489_CR47","doi-asserted-by":"crossref","unstructured":"Gale T, Zaharia M, Young C et al (2020) Sparse GPU kernels for deep learning. In: Proceedings of the international conference for high performance computing, networking, storage and analysis","DOI":"10.1109\/SC41405.2020.00021"},{"key":"10489_CR48","unstructured":"Gebhart T, Saxena U, Schrater P (2021) A unified paths perspective for pruning at initialization. CoRR abs\/2101.10552. arXiv:2101.10552. Accessed 19 Nov 2022"},{"key":"10489_CR49","doi-asserted-by":"crossref","unstructured":"Girish S, Maiya SR, Gupta K et al (2021) The lottery ticket hypothesis for object recognition. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR46437.2021.00082"},{"issue":"13","key":"10489_CR50","doi-asserted-by":"publisher","first-page":"3444","DOI":"10.1109\/TSP.2016.2546221","volume":"64","author":"R Giryes","year":"2016","unstructured":"Giryes R, Sapiro G, Bronstein AM (2016) Deep neural networks with random Gaussian weights: a universal classification strategy? IEEE Trans Signal Process 64(13):3444\u20133457","journal-title":"IEEE Trans Signal Process"},{"key":"10489_CR51","unstructured":"Glorot X, Bengio Y (2010) Understanding the difficulty of training deep feedforward neural networks. In: Proceedings of the 13th international conference on artificial intelligence and statistics"},{"key":"10489_CR52","unstructured":"Guo Y, Yao A, Chen Y (2016) Dynamic network surgery for efficient DNNs. In: Advances in neural information processing systems, 2016, vol 29"},{"key":"10489_CR53","unstructured":"Gustafson JL (2011) Moore\u2019s law. In: Encyclopedia of Parallel Computing, pp 1177\u20131184"},{"key":"10489_CR54","unstructured":"Han S, Pool J, Tran J et al (2015) Learning both weights and connections for efficient neural network. In: Advances in neural information processing systems, 2015, vol 28"},{"issue":"3","key":"10489_CR55","doi-asserted-by":"publisher","first-page":"243","DOI":"10.1145\/3007787.3001163","volume":"44","author":"S Han","year":"2016","unstructured":"Han S, Liu X, Mao H et al (2016) EIE: efficient inference engine on compressed deep neural network. ACM SIGARCH Comput Archit News 44(3):243\u2013254","journal-title":"ACM SIGARCH Comput Archit News"},{"key":"10489_CR56","unstructured":"Hanin B, Rolnick D (2018) How to start training: the effect of initialization and architecture. In: Advances in neural information processing systems, vol 31"},{"key":"10489_CR57","unstructured":"Hayou S, Ton JF, Doucet A et al (2021) Robust pruning at initialization. In: 9th International conference on learning representations"},{"key":"10489_CR58","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S et al (2015) Delving deep into rectifiers: surpassing human-level performance on ImageNet classification. In: IEEE international conference on computer vision","DOI":"10.1109\/ICCV.2015.123"},{"key":"10489_CR59","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S et al (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR.2016.90"},{"key":"10489_CR60","unstructured":"Hoffer E, Hubara I, Soudry D (2018) Fix your classifier: the marginal value of training the last weight layer. In: 6th International conference on learning representations"},{"key":"10489_CR61","unstructured":"Holmes C, Zhang M, He Y et al (2021) NxMTransformer: semi-structured sparsification for natural language understanding via ADMM. In: Advances in neural information processing systems, vol 34"},{"key":"10489_CR62","unstructured":"Huang GB, Zhu QY, Siew CK (2004) Extreme learning machine: a new learning scheme of feedforward neural networks. In: IEEE international joint conference on neural networks, vol 2"},{"issue":"2","key":"10489_CR63","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1007\/s13042-011-0019-y","volume":"2","author":"GB Huang","year":"2011","unstructured":"Huang GB, Wang DH, Lan Y (2011) Extreme learning machines: a survey. Int J Mach Learn Cybern 2(2):107\u2013122","journal-title":"Int J Mach Learn Cybern"},{"key":"10489_CR64","doi-asserted-by":"crossref","unstructured":"Huang Z, Wang N (2018) Data-driven sparse structure selection for deep neural networks. In: Proceedings of the European conference on computer vision","DOI":"10.1007\/978-3-030-01270-0_19"},{"key":"10489_CR65","unstructured":"Hubara I, Chmiel B, Island M et al (2021) Accelerated sparse neural training: a provable and efficient method to find n:m transposable masks. In: Advances in neural information processing systems, vol 34"},{"key":"10489_CR66","unstructured":"Ioffe S, Szegedy C (2015) Batch normalization: accelerating deep network training by reducing internal covariate shift. In: Proceedings of the 32nd international conference on machine learning"},{"key":"10489_CR67","doi-asserted-by":"crossref","unstructured":"Jacob B, Kligys S, Chen B et al (2018) Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: Proceedings of the IEEE conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR.2018.00286"},{"key":"10489_CR68","unstructured":"Jacot A, Hongler C, Gabriel F (2018) Neural tangent kernel: convergence and generalization in neural networks. In: Advances in neural information processing systems, vol 31"},{"key":"10489_CR69","doi-asserted-by":"publisher","first-page":"6600","DOI":"10.1103\/PhysRevA.39.6600","volume":"39","author":"SA Janowsky","year":"1989","unstructured":"Janowsky SA (1989) Pruning versus clipping in neural networks. Phys Rev A 39:6600\u20136603","journal-title":"Phys Rev A"},{"key":"10489_CR70","unstructured":"Jayakumar S, Pascanu R, Rae J et al (2020) Top-KAST: top-k always sparse training. In: Advances in neural information processing systems, vol 33"},{"key":"10489_CR71","doi-asserted-by":"crossref","unstructured":"Joo D, Yi E, Baek S et al (2021) Linearly replaceable filters for deep network channel pruning. In: Proceedings of the AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v35i9.16978"},{"issue":"2","key":"10489_CR72","doi-asserted-by":"publisher","first-page":"239","DOI":"10.1109\/72.80236","volume":"1","author":"ED Karnin","year":"1990","unstructured":"Karnin ED (1990) A simple procedure for pruning back-propagation trained neural networks. IEEE Trans Neural Netw 1(2):239\u2013242","journal-title":"IEEE Trans Neural Netw"},{"key":"10489_CR73","unstructured":"Kingma DP, Ba J (2015) Adam: a method for stochastic optimization. In: 3rd International conference on learning representations"},{"key":"10489_CR74","doi-asserted-by":"crossref","unstructured":"Kolesnikov A, Beyer L, Zhai X et al (2020) Big transfer (bit): general visual representation learning. In: Proceedings of the European conference on computer vision","DOI":"10.1007\/978-3-030-58558-7_29"},{"key":"10489_CR75","unstructured":"Koster N, Grothe O, Rettinger A (2022) Signing the supermask: keep, hide, invert. In: 10th International conference on learning representations"},{"key":"10489_CR76","unstructured":"Krizhevsky A (2012) Learning multiple layers of features from tiny images. University of Toronto. http:\/\/www.cs.toronto.edu\/~kriz\/cifar.html. Accessed 13 May 2022"},{"key":"10489_CR77","unstructured":"Kusupati A, Ramanujan V, Somani R et al (2020) Soft threshold weight reparametrization for learnable sparsity. In: Proceedings of the 37th international conference on machine learning"},{"key":"10489_CR78","unstructured":"Le DH, Hua BS (2021) Network pruning that matters: a case study on retraining variants. In: 9th International conference on learning representations"},{"key":"10489_CR79","unstructured":"Lebedev V, Ganin Y, Rakhuba M et al (2015) Speeding-up convolutional neural networks using fine-tuned CP-decomposition. In: 3rd International conference on learning representations"},{"key":"10489_CR80","unstructured":"LeCun Y, Denker JS, Solla SA (1990) Optimal brain damage. In: Advances in neural information processing systems, vol 2"},{"issue":"11","key":"10489_CR81","doi-asserted-by":"publisher","first-page":"2278","DOI":"10.1109\/5.726791","volume":"86","author":"Y LeCun","year":"1998","unstructured":"LeCun Y, Bottou L, Bengio Y et al (1998) Gradient-based learning applied to document recognition. Proc IEEE 86(11):2278\u20132324","journal-title":"Proc IEEE"},{"key":"10489_CR82","doi-asserted-by":"crossref","unstructured":"Lee J, Xiao L, Schoenholz S et al (2019a) Wide neural networks of any depth evolve as linear models under gradient descent. In: Advances in neural information processing systems, vol 32","DOI":"10.1088\/1742-5468\/abc62b"},{"key":"10489_CR83","unstructured":"Lee N, Ajanthan T, Torr PH (2019b) SNIP: single-shot network pruning based on connection sensitivity. In: 7th International conference on learning representations"},{"key":"10489_CR84","unstructured":"Lee N, Ajanthan T, Gould S et al (2020) A signal propagation perspective for pruning neural networks at initialization. In: 8th International conference on learning representations"},{"key":"10489_CR85","unstructured":"Lee J, Park S, Mo S et al (2021) Layer-adaptive sparsity for the magnitude-based pruning. In: 9th International conference on learning representations"},{"key":"10489_CR86","unstructured":"Li Y, Liang Y (2018) Learning overparameterized neural networks via stochastic gradient descent on structured data. In: Advances in neural information processing systems, vol 31"},{"key":"10489_CR87","unstructured":"Liu B, Wang M, Foroosh H et al (2015) Sparse convolutional neural networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition"},{"key":"10489_CR88","unstructured":"Li H, Kadav A, Durdanovic I et al (2017) Pruning filters for efficient ConvNets. In: 5th International conference on learning representations"},{"key":"10489_CR89","unstructured":"Li C, Farkhoor H, Liu R et al (2018) Measuring the intrinsic dimension of objective landscapes. In: 6th International conference on learning representations"},{"key":"10489_CR90","doi-asserted-by":"crossref","unstructured":"Li R, Wang Y, Liang F et al (2019) Fully quantized network for object detection. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR.2019.00292"},{"key":"10489_CR91","unstructured":"Liu T, Zenke F (2020) Finding trainable sparse networks through neural tangent transfer. In: Proceedings of the 37th international conference on machine learning"},{"key":"10489_CR92","unstructured":"Liu Z, Sun M, Zhou T et al (2019) Rethinking the value of network pruning. In: 7th International conference on learning representations"},{"key":"10489_CR93","unstructured":"Liu J, Xu Z, Shi R et al (2020) Dynamic sparse training: find efficient sparse network from scratch with trainable masked layers. In: 8th International conference on learning representations"},{"issue":"7","key":"10489_CR94","doi-asserted-by":"publisher","first-page":"2589","DOI":"10.1007\/s00521-020-05136-7","volume":"33","author":"S Liu","year":"2021","unstructured":"Liu S, Mocanu DC, Matavalam ARR et al (2021) Sparse evolutionary deep learning with over one million artificial neurons on commodity hardware. Neural Comput Appl 33(7):2589\u20132604","journal-title":"Neural Comput Appl"},{"key":"10489_CR95","unstructured":"Liu S, Yin L, Mocanu DC et al (2021b) Do we actually need dense over-parameterization? In-time over-parameterization in sparse training. In: Proceedings of the 38th international conference on machine learning"},{"key":"10489_CR96","unstructured":"Liu S, Chen T, Chen X et al (2022) The unreasonable effectiveness of random pruning: return of the most Naive baseline for sparse training. In: 10th International conference on learning representations"},{"key":"10489_CR97","unstructured":"Lubana ES, Dick R (2021) A gradient flow framework for analyzing network pruning. In: 9th International conference on learning representations"},{"key":"10489_CR98","doi-asserted-by":"crossref","unstructured":"Mahajan D, Girshick R, Ramanathan V et al (2018) Exploring the limits of weakly supervised pretraining. In: Proceedings of the European conference on computer vision","DOI":"10.1007\/978-3-030-01216-8_12"},{"key":"10489_CR99","unstructured":"Malach E, Yehudai G, Shalev-Schwartz S et al (2020) Proving the lottery ticket hypothesis: pruning is all you need. In: Proceedings of the 37th international conference on machine learning"},{"key":"10489_CR100","doi-asserted-by":"crossref","unstructured":"Mallya A, Davis D, Lazebnik S (2018) Piggyback: adapting a single network to multiple tasks by learning to mask weights. In: Proceedings of the European conference on computer vision","DOI":"10.1007\/978-3-030-01225-0_5"},{"key":"10489_CR101","doi-asserted-by":"crossref","unstructured":"Mao H, Han S, Pool J et al (2017) Exploring the granularity of sparsity in convolutional neural networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition workshops","DOI":"10.1109\/CVPRW.2017.241"},{"key":"10489_CR102","unstructured":"Martens J (2010) Deep learning via Hessian-free optimization. In: Proceedings of the 27th international conference on machine learning"},{"issue":"1","key":"10489_CR103","doi-asserted-by":"publisher","first-page":"309","DOI":"10.1007\/s11071-005-2824-x","volume":"41","author":"I Mezi\u0107","year":"2005","unstructured":"Mezi\u0107 I (2005) Spectral properties of dynamical systems, model reduction and decompositions. Nonlinear Dyn 41(1):309\u2013325","journal-title":"Nonlinear Dyn"},{"issue":"1","key":"10489_CR104","doi-asserted-by":"publisher","first-page":"2383","DOI":"10.1038\/s41467-018-04316-3","volume":"9","author":"D Mocanu","year":"2018","unstructured":"Mocanu D, Mocanu E, Stone P et al (2018) Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science. Nat Commun 9(1):2383","journal-title":"Nat Commun"},{"key":"10489_CR105","unstructured":"Morcos A, Yu H, Paganini M et al (2019) One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers. In: Advances in neural information processing systems, vol 32"},{"key":"10489_CR106","unstructured":"Mostafa H, Wang X (2019) Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization. In: Proceedings of the 36th international conference on machine learning, 2019"},{"key":"10489_CR107","doi-asserted-by":"crossref","unstructured":"Mozer MC, Smolensky P (1989) Skeletonization: a technique for trimming the fat from a network via relevance assessment. In: Advances in neural information processing systems, vol 1","DOI":"10.1080\/09540098908915626"},{"key":"10489_CR108","unstructured":"Neyshabur B, Salakhutdinov R, Srebro N (2015a) Path-SGD: path-normalized optimization in deep neural networks. In: Advances in neural information processing systems, vol 28"},{"key":"10489_CR109","unstructured":"Neyshabur B, Tomioka R, Srebro N (2015b) Norm-based capacity control in neural networks. In: Proceedings of the 28th conference on learning theory"},{"key":"10489_CR110","unstructured":"Novikov A, Podoprikhin D, Osokin A et al (2015) Tensorizing neural networks. In: Advances in neural information processing systems, vol 28"},{"issue":"4","key":"10489_CR111","doi-asserted-by":"publisher","first-page":"473","DOI":"10.1162\/neco.1992.4.4.473","volume":"4","author":"SJ Nowlan","year":"1992","unstructured":"Nowlan SJ, Hinton GE (1992) Simplifying neural networks by soft weight-sharing. Neural Comput 4(4):473\u2013493","journal-title":"Neural Comput"},{"key":"10489_CR112","unstructured":"NVIDIA (2020) NVIDIA a 100 tensor core GPU architecture. https:\/\/images.nvidia.com\/aem-dam\/en-zz\/Solutions\/data-center\/nvidia-ampere-architecture-whitepaper.pdf. Accessed 31 Oct 2022"},{"key":"10489_CR113","unstructured":"Orseau L, Hutter M, Rivasplata O (2020) Logarithmic pruning is all you need. In: Advances in neural information processing systems, vol 33"},{"issue":"5","key":"10489_CR114","doi-asserted-by":"publisher","first-page":"76","DOI":"10.1109\/2.144401","volume":"25","author":"YH Pao","year":"1992","unstructured":"Pao YH, Takefuji Y (1992) Functional-link net computing: theory, system architecture, and functionalities. Computer 25(5):76\u201379","journal-title":"Computer"},{"issue":"2","key":"10489_CR115","doi-asserted-by":"publisher","first-page":"163","DOI":"10.1016\/0925-2312(94)90053-1","volume":"6","author":"YH Pao","year":"1994","unstructured":"Pao YH, Park GH, Sobajic DJ (1994) Learning and generalization characteristics of the random vector functional-link net. Neurocomputing 6(2):163\u2013180","journal-title":"Neurocomputing"},{"key":"10489_CR116","doi-asserted-by":"crossref","unstructured":"Parashar A, Rhu M, Mukkara A et al (2017) SCNN. In: Proceedings of the 44th annual international symposium on computer architecture. ACM","DOI":"10.1145\/3079856.3080254"},{"key":"10489_CR117","unstructured":"Park J, Li SR, Wen W et al (2017) Faster CNNs with direct sparse convolutions and guided pruning. In: 5th International conference on learning representations"},{"key":"10489_CR118","doi-asserted-by":"crossref","unstructured":"Park DS, Zhang Y, Chiu C et al (2020) Specaugment on large scale datasets. In: IEEE international conference on acoustics, speech and signal processing","DOI":"10.1109\/ICASSP40776.2020.9053205"},{"key":"10489_CR119","unstructured":"Patil SM, Dovrolis C (2021) PHEW: constructing sparse networks that learn fast and generalize well without training data. In: Proceedings of the 38th international conference on machine learning"},{"key":"10489_CR120","unstructured":"Pensia A, Rajput S, Nagle A et al (2020) Optimal lottery tickets via subset sum: logarithmic over-parameterization is sufficient. In: Advances in neural information processing systems, vol 33"},{"key":"10489_CR121","unstructured":"Peste A, Iofinova E, Vladu A et al (2021) AC\/DC: alternating compressed\/decompressed training of deep neural networks. In: Advances in neural information processing systems, vol 34"},{"key":"10489_CR122","doi-asserted-by":"crossref","unstructured":"Peters ME, Neumann M, Iyyer M et al (2018) Deep contextualized word representations. In: Proceedings of the 2018 conference of the North American Chapter of the Association for Computational Linguistics: human language technologies","DOI":"10.18653\/v1\/N18-1202"},{"key":"10489_CR123","doi-asserted-by":"crossref","unstructured":"Pham H, Dai Z, Xie Q et al (2021) Meta pseudo labels. In: Proceedings of the IEEE conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR46437.2021.01139"},{"key":"10489_CR124","unstructured":"Pool J, Yu C (2021) Channel permutations for n:m sparsity. In: Advances in neural information processing systems, vol 34"},{"key":"10489_CR125","unstructured":"Poole B, Lahiri S, Raghu M et al (2016) Exponential expressivity in deep neural networks through transient chaos. In: Advances in neural information processing systems, vol 29"},{"key":"10489_CR126","unstructured":"Price I, Tanner J (2021) Dense for the price of sparse: improved performance of sparsely initialized networks via a subspace offset. In: Proceedings of the 38th international conference on machine learning"},{"key":"10489_CR127","unstructured":"Qian X, Klabjan D (2021) A probabilistic approach to neural network pruning. In: Proceedings of the 38th international conference on machine learning"},{"key":"10489_CR128","doi-asserted-by":"publisher","first-page":"426","DOI":"10.1016\/j.neucom.2020.06.110","volume":"412","author":"Y Qing","year":"2020","unstructured":"Qing Y, Zeng Y, Li Y et al (2020) Deep and wide feature based extreme learning machine for image classification. Neurocomputing 412:426\u2013436","journal-title":"Neurocomputing"},{"key":"10489_CR129","doi-asserted-by":"crossref","unstructured":"Ramanujan V, Wortsman M, Kembhavi A et al (2020) What\u2019s hidden in a randomly weighted neural network? In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR42600.2020.01191"},{"key":"10489_CR130","unstructured":"Redman WT, Fonoberova M, Mohr R et al (2022) An operator theoretic view on pruning deep neural networks. In: 10th International conference on learning representations"},{"key":"10489_CR131","unstructured":"Renda A, Frankle J, Carbin M (2020) Comparing rewinding and fine-tuning in neural network pruning. In: 8th International conference on learning representations"},{"issue":"3","key":"10489_CR132","doi-asserted-by":"publisher","first-page":"400","DOI":"10.1214\/aoms\/1177729586","volume":"22","author":"H Robbins","year":"1951","unstructured":"Robbins H, Monro S (1951) A stochastic approximation method. Ann Math Stat 22(3):400\u2013407","journal-title":"Ann Math Stat"},{"key":"10489_CR133","doi-asserted-by":"crossref","unstructured":"Rosenfeld A, Tsotsos JK (2019) Intriguing properties of randomly weighted networks: generalizing while learning next to nothing. In: Conference on computer and robot vision","DOI":"10.1109\/CRV.2019.00010"},{"key":"10489_CR134","unstructured":"Rosenfeld JS, Frankle J, Carbin M et al (2021) On the predictability of pruning across scales. In: Proceedings of the 38th international conference on machine learning"},{"key":"10489_CR135","doi-asserted-by":"crossref","unstructured":"Sainath T, Kingsbury B, Sindhwani V et al (2013) Low-rank matrix factorization for deep neural network training with high-dimensional output targets. In: IEEE international conference on acoustics, speech and signal processing","DOI":"10.1109\/ICASSP.2013.6638949"},{"key":"10489_CR136","unstructured":"Sanh V, Wolf T, Rush AM (2020) Movement pruning: adaptive sparsity by fine-tuning. In: Advances in neural information processing systems, vol 33"},{"key":"10489_CR137","unstructured":"Saxe A, Koh PW, Chen Z et al (2011) On random weights and unsupervised feature learning. In: Proceedings of the 28th international conference on machine learning"},{"key":"10489_CR138","unstructured":"Saxe AM, McClelland JL, Ganguli S (2014) Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. In: 2nd International conference on learning representations"},{"key":"10489_CR139","unstructured":"Schoenholz SS, Gilmer J, Ganguli S et al (2017) Deep information propagation. In: 5th International conference on learning representations"},{"issue":"12","key":"10489_CR140","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1145\/3381831","volume":"63","author":"R Schwartz","year":"2020","unstructured":"Schwartz R, Dodge J, Smith NA et al (2020) Green AI. Commun ACM 63(12):54\u201363","journal-title":"Commun ACM"},{"key":"10489_CR141","unstructured":"Schwarz J, Jayakumar S, Pascanu R et al (2021) Powerpropagation: a sparsity inducing weight reparameterisation. In: Advances in neural information processing systems, vol 34"},{"key":"10489_CR181","doi-asserted-by":"crossref","unstructured":"Shen X, Kong Z, Qin M, et al (2022) The lottery ticket hypothesis for vision transformers. CoRR abs\/2211.01484. Accessed 23 Apr 2023","DOI":"10.24963\/ijcai.2023\/153"},{"key":"10489_CR142","doi-asserted-by":"crossref","unstructured":"Soelen RV, Sheppard JW (2019) Using winning lottery tickets in transfer learning for convolutional neural networks. In: International joint conference on neural networks","DOI":"10.1109\/IJCNN.2019.8852405"},{"key":"10489_CR143","doi-asserted-by":"crossref","unstructured":"Strubell E, Ganesh A, McCallum A (2019) Energy and policy considerations for deep learning in NLP. In: Proceedings of the 57th conference of the Association for Computational Linguistics","DOI":"10.18653\/v1\/P19-1355"},{"key":"10489_CR144","doi-asserted-by":"crossref","unstructured":"Strubell E, Ganesh A, McCallum A (2020) Energy and policy considerations for modern deep learning research. In: Proceedings of the AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v34i09.7123"},{"key":"10489_CR145","unstructured":"Su J, Chen Y, Cai T et al (2020) Sanity-checking pruning methods: random tickets can win the jackpot. In: Advances in neural information processing systems, vol 33"},{"key":"10489_CR146","unstructured":"Sun W, Zhou A, Stuijk S et al (2021) Dominosearch: find layer-wise fine-grained n:m sparse schemes from dense neural networks. In: Advances in neural information processing systems, vol 34"},{"key":"10489_CR147","unstructured":"Sung YL, Nair V, Raffel C (2021) Training neural networks with fixed sparse masks. In: Advances in neural information processing systems, vol 34"},{"key":"10489_CR148","unstructured":"Sutskever I, Martens J, Dahl G et al (2013) On the importance of initialization and momentum in deep learning. In: Proceedings of the 30th international conference on machine learning"},{"key":"10489_CR149","unstructured":"Tanaka H, Kunin D, Yamins DL et al (2020) Pruning neural networks without any data by iteratively conserving synaptic flow. In: Advances in neural information processing systems, vol 33"},{"issue":"11","key":"10489_CR150","doi-asserted-by":"publisher","first-page":"1801","DOI":"10.1109\/PROC.1967.6011","volume":"55","author":"W Tinney","year":"1967","unstructured":"Tinney W, Walker J (1967) Direct solutions of sparse network equations by optimally ordered triangular factorization. Proc IEEE 55(11):1801\u20131809","journal-title":"Proc IEEE"},{"key":"10489_CR151","unstructured":"Ullrich K, Meeds E, Welling M (2017) Soft weight-sharing for neural network compression. In: 5th International conference on learning representations"},{"key":"10489_CR152","unstructured":"Verdenius S, Stol M, Forr\u00e9 P (2020) Pruning via iterative ranking of sensitivity statistics. CoRR abs\/2006.00896v2. arXiv:2006.00896v2. Accessed 19 Sep 2022"},{"key":"10489_CR153","unstructured":"Vischer M, Lange RT, Sprekeler H (2022) On lottery tickets and minimal task representations in deep reinforcement learning. In: 10th International conference on learning representations"},{"key":"10489_CR154","doi-asserted-by":"crossref","unstructured":"Wang Z (2020) SparSERT: accelerating unstructured sparsity on GPUs for deep learning inference. In: Proceedings of the ACM international conference on parallel architectures and compilation techniques","DOI":"10.1145\/3410463.3414654"},{"key":"10489_CR155","unstructured":"Wang C, Zhang G, Grosse R (2020a) Picking winning tickets before training by preserving gradient flow. In: 8th International conference on learning representations"},{"key":"10489_CR156","doi-asserted-by":"crossref","unstructured":"Wang CY, Bochkovskiy A, Liao HYM (2021a) Scaled-YOLOV4: scaling cross stage partial network. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR46437.2021.01283"},{"key":"10489_CR157","unstructured":"Wang H, Qin C, Zhang Y et al (2021b) Emerging paradigms of neural network pruning. CoRR abs\/2103.06460v2. arXiv:2103.06460v2. Accessed 5 Oct 2022"},{"key":"10489_CR158","doi-asserted-by":"crossref","unstructured":"Wang Y, Zhang X, Hu X et al (2020b) Dynamic network pruning with interpretable layerwise channel selection. In: Proceedings of the AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v34i04.6098"},{"key":"10489_CR159","doi-asserted-by":"crossref","unstructured":"Wang Y, Zhang X, Xie L et al (2020c) Pruning from scratch. In: Proceedings of the AAAI conference on artificial intelligence","DOI":"10.1609\/aaai.v34i07.6910"},{"key":"10489_CR160","doi-asserted-by":"crossref","unstructured":"Wimmer P, Mehnert J, Condurache AP (2020) FreezeNet: full performance by reduced storage costs. In: Proceedings of the Asian conference on computer vision","DOI":"10.1007\/978-3-030-69544-6_41"},{"key":"10489_CR161","unstructured":"Wimmer P, Mehnert J, Condurache AP (2021) COPS: controlled pruning before training starts. In: International joint conference on neural networks"},{"key":"10489_CR162","doi-asserted-by":"crossref","unstructured":"Wimmer P, Mehnert J, Condurache AP (2022) Interspace pruning: using adaptive filter representations to improve training of sparse CNNs. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR52688.2022.01220"},{"key":"10489_CR163","doi-asserted-by":"crossref","unstructured":"Wu J, Leng C, Wang Y et al (2016) Quantized convolutional neural networks for mobile devices. In: Proceedings of the IEEE conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR.2016.521"},{"key":"10489_CR164","unstructured":"Xiao L, Bahri Y, Sohl-Dickstein J et al (2018) Dynamical isometry and a mean field theory of CNNs: how to train 10,000-layer vanilla convolutional neural networks. In: Proceedings of the 35th international conference on machine learning"},{"key":"10489_CR165","doi-asserted-by":"crossref","unstructured":"Xue J, Li J, Gong Y (2013) Restructuring of deep neural network acoustic models with singular value decomposition. In: INTERSPEECH","DOI":"10.21437\/Interspeech.2013-552"},{"key":"10489_CR166","unstructured":"Yang Z, Dai Z, Yang Y et al (2019) XLNet: generalized autoregressive pretraining for language understanding. In: Advances in neural information processing systems, vol 32"},{"key":"10489_CR167","unstructured":"You H, Li C, Xu P et al (2020) Drawing early-bird tickets: toward more efficient training of deep networks. In: 8th International conference on learning representations"},{"key":"10489_CR168","unstructured":"Yu H, Edunov S, Tian Y et al (2020) Playing the lottery with rewards and multiple languages: lottery tickets in RL and NLP. In: 8th International conference on learning representations"},{"key":"10489_CR169","doi-asserted-by":"crossref","unstructured":"Zagoruyko S, Komodakis N (2016) Wide residual networks. In: Proceedings of the British machine vision conference","DOI":"10.5244\/C.30.87"},{"key":"10489_CR170","doi-asserted-by":"crossref","unstructured":"Zhang D, Yang J, Ye D et al (2018) LQ-Nets: learned quantization for highly accurate and compact deep neural networks. In: Proceedings of the European conference on computer vision","DOI":"10.1007\/978-3-030-01237-3_23"},{"key":"10489_CR171","unstructured":"Zhang S, Stadie BC (2020) One-shot pruning of recurrent neural networks by Jacobian spectrum evaluation. In: 8th International conference on learning representations"},{"key":"10489_CR172","unstructured":"Zhang Z, Chen X, Chen T et al (2021a) Efficient lottery ticket finding: less data is more. In: Proceedings of the 38th international conference on machine learning"},{"key":"10489_CR173","unstructured":"Zhang Z, Jin J, Zhang Z et al (2021b) Validating the lottery ticket hypothesis with inertial manifold theory. In: Advances in neural information processing systems, vol 34"},{"key":"10489_CR174","unstructured":"Zhou A, Yao A, Guo Y et al (2017) Incremental network quantization: towards lossless CNNs with low-precision weights. In: 5th International conference on learning representations"},{"key":"10489_CR175","unstructured":"Zhou H, Lan J, Liu R et al (2019) Deconstructing lottery tickets: zeros, signs, and the supermask. In: Advances in neural information processing systems, vol 32"},{"key":"10489_CR176","unstructured":"Zhou A, Ma Y, Zhu J et al (2021a) Learning n:m fine-grained structured sparse neural networks from scratch. In: 9th International conference on learning representations"},{"key":"10489_CR177","unstructured":"Zhou X, Zhang W, Chen Z et al (2021b) Efficient neural network training via forward and backward propagation sparsification. In: Advances in neural information processing systems, vol 34"},{"key":"10489_CR178","doi-asserted-by":"crossref","unstructured":"Zhou X, Zhang W, Xu H et al (2021c) Effective sparsification of neural networks with global sparsity constraint. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition","DOI":"10.1109\/CVPR46437.2021.00360"},{"key":"10489_CR179","unstructured":"Zhuang Z, Tan M, Zhuang B et al (2018) Discrimination-aware channel pruning for deep neural networks. In: Advances in neural information processing systems, vol 31"},{"key":"10489_CR180","unstructured":"Zhuang T, Zhang Z, Huang Y et al (2020) Neuron-level structured pruning using polarization regularizer. In: Advances in neural information processing systems, vol 33"}],"container-title":["Artificial Intelligence Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-023-10489-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10462-023-10489-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10462-023-10489-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,19]],"date-time":"2024-10-19T10:58:33Z","timestamp":1729335513000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10462-023-10489-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,1]]},"references-count":181,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2023,12]]}},"alternative-id":["10489"],"URL":"https:\/\/doi.org\/10.1007\/s10462-023-10489-1","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-2458016\/v1","asserted-by":"object"}]},"ISSN":["0269-2821","1573-7462"],"issn-type":[{"value":"0269-2821","type":"print"},{"value":"1573-7462","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,1]]},"assertion":[{"value":"1 May 2023","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"During preparation of this manuscript, all authors were employed by the Robert Bosch GmbH. Also, Alexandru Paul Condurache and Paul Wimmer were part of the Institute for Signal Processing of the University of L\u00fcbeck.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}