{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T14:39:59Z","timestamp":1777559999258,"version":"3.51.4"},"reference-count":33,"publisher":"SAGE Publications","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["AIC"],"published-print":{"date-parts":[[2022,7,18]]},"abstract":"<jats:p>We propose a novel pruning method which uses the oscillations around 0, i.e. sign flips, that a weight has undergone during training in order to determine its saliency. Our method can perform pruning before the network has converged, requires little tuning effort due to having good default values for its hyperparameters, and can directly target the level of sparsity desired by the user. Our experiments, performed on a variety of object classification architectures, show that it is competitive with existing methods and achieves state-of-the-art performance for levels of sparsity of 99.6 % and above for 2 out of 3 of the architectures tested. Moreover, we demonstrate that our method is compatible with quantization, another model compression technique. For reproducibility, we release our code at https:\/\/github.com\/AndreiXYZ\/flipout.<\/jats:p>","DOI":"10.3233\/aic-210127","type":"journal-article","created":{"date-parts":[[2021,11,5]],"date-time":"2021-11-05T14:45:57Z","timestamp":1636123557000},"page":"65-85","source":"Crossref","is-referenced-by-count":0,"title":["Pruning by leveraging training dynamics"],"prefix":"10.1177","volume":"35","author":[{"given":"Andrei C.","family":"Apostol","sequence":"first","affiliation":[{"name":"Informatics Institute, University of Amsterdam, Amsterdam, The Netherlands"},{"name":"BrainCreators B.V., Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maarten C.","family":"Stol","sequence":"additional","affiliation":[{"name":"BrainCreators B.V., Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Patrick","family":"Forr\u00e9","sequence":"additional","affiliation":[{"name":"Informatics Institute, University of Amsterdam, Amsterdam, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","reference":[{"key":"10.3233\/AIC-210127_ref1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-76640-5_2"},{"key":"10.3233\/AIC-210127_ref2","unstructured":"G.\u00a0Bellec, D.\u00a0Kappel, W.\u00a0Maass and R.\u00a0Legenstein, Deep rewiring: Training very sparse deep networks, in: International Conference on Learning Representations, 2018."},{"key":"10.3233\/AIC-210127_ref3","doi-asserted-by":"crossref","unstructured":"L.\u00a0Bottou, Online algorithms and stochastic approximations, in: Online Learning and Neural Networks, Cambridge University Press, Cambridge, UK, 1998, revised, Oct. 2012.","DOI":"10.1017\/CBO9780511569920.003"},{"key":"10.3233\/AIC-210127_ref4","unstructured":"A.\u00a0Brock, J.\u00a0Donahue and K.\u00a0Simonyan, Large scale GAN training for high fidelity natural image synthesis, in: International Conference on Learning Representations, 2019."},{"key":"10.3233\/AIC-210127_ref5","unstructured":"X.\u00a0Ding, X.\u00a0Zhou, Y.\u00a0Guo, J.\u00a0Han, J.\u00a0Liu et al., Global sparse momentum sgd for pruning very deep neural networks, in: Advances in Neural Information Processing Systems, 2019."},{"key":"10.3233\/AIC-210127_ref6","unstructured":"J.\u00a0Frankle and M.\u00a0Carbin, The lottery ticket hypothesis: Finding sparse, trainable neural networks, in: International Conference on Learning Representations, 2019."},{"key":"10.3233\/AIC-210127_ref7","unstructured":"I.\u00a0Goodfellow, J.\u00a0Pouget-Abadie, M.\u00a0Mirza, B.\u00a0Xu, D.\u00a0Warde-Farley, S.\u00a0Ozair, A.\u00a0Courville and Y.\u00a0Bengio, Generative adversarial nets, in: Advances in Neural Information Processing Systems, 2014, pp.\u00a02672\u20132680."},{"key":"10.3233\/AIC-210127_ref8","unstructured":"S.\u00a0Han, J.\u00a0Pool, J.\u00a0Tran and W.\u00a0Dally, Learning both weights and connections for efficient neural network, in: Advances in Neural Information Processing Systems, 2015."},{"key":"10.3233\/AIC-210127_ref9","unstructured":"B.\u00a0Hassibi and D.G.\u00a0Stork, Second order derivatives for network pruning: Optimal brain surgeon, in: Advances in Neural Information Processing Systems, 1993."},{"key":"10.3233\/AIC-210127_ref11","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"10.3233\/AIC-210127_ref12","doi-asserted-by":"crossref","unstructured":"Y.\u00a0He, X.\u00a0Zhang and J.\u00a0Sun, Channel pruning for accelerating very deep neural networks, in: Proceedings of the IEEE International Conference on Computer Vision, 2017, pp.\u00a01389\u20131397.","DOI":"10.1109\/ICCV.2017.155"},{"issue":"8","key":"10.3233\/AIC-210127_ref13","doi-asserted-by":"publisher","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Computation"},{"key":"10.3233\/AIC-210127_ref15","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"10.3233\/AIC-210127_ref16","unstructured":"S.\u00a0Ioffe and C.\u00a0Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift, in: Proceedings of the 32nd International Conference on Machine Learning, Vol.\u00a037, ICML\u201915, JMLR.org, 2015."},{"key":"10.3233\/AIC-210127_ref17","doi-asserted-by":"crossref","unstructured":"B.\u00a0Jacob, S.\u00a0Kligys, B.\u00a0Chen, M.\u00a0Zhu, M.\u00a0Tang, A.\u00a0Howard, H.\u00a0Adam and D.\u00a0Kalenichenko, Quantization and training of neural networks for efficient integer-arithmetic-only inference, in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp.\u00a02704\u20132713.","DOI":"10.1109\/CVPR.2018.00286"},{"key":"10.3233\/AIC-210127_ref18","unstructured":"D.\u00a0Kingma and J.\u00a0Ba, Adam: A method for stochastic optimization, in: International Conference on Learning Representations, 2015."},{"key":"10.3233\/AIC-210127_ref19","unstructured":"D.P.\u00a0Kingma and M.\u00a0Welling, Auto-encoding variational Bayes, in: Proceedings of the 2nd International Conference on Learning Representations, 2014."},{"key":"10.3233\/AIC-210127_ref20","doi-asserted-by":"publisher","DOI":"10.1109\/BigData47090.2019.9005692"},{"key":"10.3233\/AIC-210127_ref23","unstructured":"Y.\u00a0LeCun, J.S.\u00a0Denker and S.A.\u00a0Solla, Optimal brain damage, in: Advances in Neural Information Processing Systems, Vol.\u00a02, 1990."},{"key":"10.3233\/AIC-210127_ref24","unstructured":"N.\u00a0Lee, T.\u00a0Ajanthan and P.\u00a0Torr, SNIP: Single-shot network pruning based on connection sensitivity, in: International Conference on Learning Representations, 2019."},{"key":"10.3233\/AIC-210127_ref25","unstructured":"H.\u00a0Li, A.\u00a0Kadav, I.\u00a0Durdanovic, H.\u00a0Samet and H.P.\u00a0Graf, Pruning filters for efficient ConvNets, in: International Conference on Learning Representations, 2017."},{"key":"10.3233\/AIC-210127_ref26","unstructured":"J.\u00a0Liu, Z.\u00a0Xu, R.\u00a0Shi, R.C.C.\u00a0Cheung and H.K.H.\u00a0So, Dynamic sparse training: Find efficient sparse network from scratch with trainable masked layers, in: International Conference on Learning Representations, 2020."},{"key":"10.3233\/AIC-210127_ref28","unstructured":"J.\u00a0Martens, Deep learning via Hessian-free optimization, in: ICML, Vol.\u00a027, 2010, pp.\u00a0735\u2013742."},{"key":"10.3233\/AIC-210127_ref30","unstructured":"D.\u00a0Molchanov, A.\u00a0Ashukha and D.\u00a0Vetrov, Variational dropout sparsifies deep neural networks, in: International Conference on Machine Learning, 2017."},{"key":"10.3233\/AIC-210127_ref31","unstructured":"V.\u00a0Nair and G.E.\u00a0Hinton, Rectified linear units improve restricted Boltzmann machines, in: Proceedings of the 27th International Conference on Machine Learning (ICML-10), 2010."},{"key":"10.3233\/AIC-210127_ref34","unstructured":"A.\u00a0Paszke, S.\u00a0Gross, F.\u00a0Massa, A.\u00a0Lerer, J.\u00a0Bradbury, G.\u00a0Chanan, T.\u00a0Killeen, Z.\u00a0Lin, N.\u00a0Gimelshein, L.\u00a0Antiga, A.\u00a0Desmaison, A.\u00a0Kopf, E.\u00a0Yang, Z.\u00a0DeVito, M.\u00a0Raison, A.\u00a0Tejani, S.\u00a0Chilamkurthy, B.\u00a0Steiner, L.\u00a0Fang, J.\u00a0Bai and S.\u00a0Chintala, PyTorch: An imperative style, high-performance deep learning library, in: Advances in Neural Information Processing Systems, Vol.\u00a032, Curran Associates, Inc., 2019. http:\/\/papers.neurips.cc\/paper\/9015-pytorch-an-imperative-style-high-performance-deep-learning-library.pdf."},{"key":"10.3233\/AIC-210127_ref36","doi-asserted-by":"crossref","unstructured":"Y.\u00a0Saad, Iterative Methods for Sparse Linear Systems, SIAM, 2003.","DOI":"10.1137\/1.9780898718003"},{"key":"10.3233\/AIC-210127_ref37","unstructured":"K.\u00a0Simonyan and A.\u00a0Zisserman, Very deep convolutional networks for large-scale image recognition, in: International Conference on Learning Representations, 2015."},{"key":"10.3233\/AIC-210127_ref38","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P19-1355"},{"key":"10.3233\/AIC-210127_ref42","unstructured":"A.\u00a0Vaswani, N.\u00a0Shazeer, N.\u00a0Parmar, J.\u00a0Uszkoreit, L.\u00a0Jones, A.N.\u00a0Gomez, \u0141.\u00a0Kaiser and I.\u00a0Polosukhin, Attention is all you need, in: Advances in Neural Information Processing Systems, 2017."},{"key":"10.3233\/AIC-210127_ref44","unstructured":"M.\u00a0Welling and Y.W.\u00a0Teh, Bayesian learning via stochastic gradient Langevin dynamics, in: International Conference on Machine Learning, 2011."},{"key":"10.3233\/AIC-210127_ref45","unstructured":"H.\u00a0Yang, W.\u00a0Wen and H.\u00a0Li, DeepHoyer: Learning sparser neural network with differentiable scale-invariant sparsity measures, in: International Conference on Learning Representations, 2020."},{"key":"10.3233\/AIC-210127_ref46","unstructured":"H.\u00a0Zhou, J.\u00a0Lan, R.\u00a0Liu and J.\u00a0Yosinski, Deconstructing lottery tickets: Zeros, signs, and the supermask, in: Advances in Neural Information Processing Systems, 2019."}],"container-title":["AI Communications"],"original-title":[],"link":[{"URL":"https:\/\/content.iospress.com\/download?id=10.3233\/AIC-210127","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T18:28:01Z","timestamp":1777400881000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/full\/10.3233\/AIC-210127"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,7,18]]},"references-count":33,"journal-issue":{"issue":"2"},"URL":"https:\/\/doi.org\/10.3233\/aic-210127","relation":{},"ISSN":["1875-8452","0921-7126"],"issn-type":[{"value":"1875-8452","type":"electronic"},{"value":"0921-7126","type":"print"}],"subject":[],"published":{"date-parts":[[2022,7,18]]}}}