{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T09:20:42Z","timestamp":1775640042943,"version":"3.50.1"},"reference-count":63,"publisher":"Oxford University Press (OUP)","issue":"7","license":[{"start":{"date-parts":[[2019,9,27]],"date-time":"2019-09-27T00:00:00Z","timestamp":1569542400000},"content-version":"vor","delay-in-days":1,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,7,17]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The successful application of deep learning has led to increasing expectations of their use in embedded systems. This, in turn, has created the need to find ways of reducing the size of neural networks. Decreasing the size of a neural network requires deciding which weights should be removed without compromising accuracy, which is analogous to the kind of problems addressed by multi-armed bandits (MABs). Hence, this paper explores the use of MABs for reducing the number of parameters of a neural network. Different MAB algorithms, namely $\\epsilon $-greedy, win-stay, lose-shift, UCB1, KL-UCB, BayesUCB, UGapEb, successive rejects and Thompson sampling are evaluated and their performance compared to existing approaches. The results show that MAB pruning methods, especially those based on UCB, outperform other pruning methods.<\/jats:p>","DOI":"10.1093\/comjnl\/bxz078","type":"journal-article","created":{"date-parts":[[2019,7,9]],"date-time":"2019-07-09T19:09:11Z","timestamp":1562699351000},"page":"1099-1108","source":"Crossref","is-referenced-by-count":6,"title":["Pruning Neural Networks Using Multi-Armed Bandits"],"prefix":"10.1093","volume":"63","author":[{"given":"Salem","family":"Ameen","sequence":"first","affiliation":[{"name":"School of Computing, Science and Engineering, University of Salford, Manchester, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sunil","family":"Vadera","sequence":"additional","affiliation":[{"name":"School of Computing, Science and Engineering, University of Salford, Manchester, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2019,9,26]]},"reference":[{"key":"2020071706564273000_ref1","first-page":"1106","article-title":"ImageNet classification with deep convolutional neural networks","volume-title":"Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems","author":"Krizhevsky","year":"2012"},{"key":"2020071706564273000_ref2","doi-asserted-by":"crossref","first-page":"818","DOI":"10.1007\/978-3-319-10590-1_53","article-title":"Visualizing and understanding convolutional networks","volume-title":"Computer Vision\u2014ECCV 2014","author":"Zeiler","year":"2014"},{"key":"2020071706564273000_ref3","first-page":"1","article-title":"Going deeper with convolutions","volume-title":"IEEE Conf. Computer Vision and Pattern Recognition, CVPR 2015","author":"Szegedy","year":"2015"},{"key":"2020071706564273000_ref4","volume-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan","year":"2015"},{"key":"2020071706564273000_ref5","first-page":"770","volume-title":"Deep residual learning for image recognition","author":"He","year":"2016"},{"key":"2020071706564273000_ref6","first-page":"2924","article-title":"On the number of linear regions of deep neural networks","volume-title":"Proc. 27th Int. Conf. Neural Information Processing Systems","author":"Mont\u00fafar","year":"2014"},{"key":"2020071706564273000_ref7","doi-asserted-by":"crossref","DOI":"10.1111\/exsy.12197","article-title":"A convolutional neural network to classify american sign language fingerspelling from depth and colour images","volume":"34","author":"Ameen","year":"2017","journal-title":"Expert Syst."},{"key":"2020071706564273000_ref8","first-page":"243","volume-title":"EIE: efficient inference engine on compressed deep neural network","author":"Han","year":"2016"},{"key":"2020071706564273000_ref9","doi-asserted-by":"crossref","first-page":"740","DOI":"10.1109\/72.248452","article-title":"Pruning algorithms\u2014a survey","volume":"4","author":"Reed","year":"1993","journal-title":"IEEE Trans. Neural Netw."},{"key":"2020071706564273000_ref10","first-page":"519","article-title":"A back-propagation algorithm with optimal use of hidden units","volume-title":"Advances in Neural Information Processing Systems","author":"Chauvin","year":"1989"},{"key":"2020071706564273000_ref11","first-page":"105","article-title":"Back-propagation, weight-elimination and time series prediction","author":"Weigend","year":"1991","journal-title":"Proc. 1990 Connectionist Models Summer School"},{"key":"2020071706564273000_ref12","first-page":"837","article-title":"Generalization by weight-elimination applied to currency exchange rate prediction","volume-title":"Proc. IEEE Int. Joint Conf. Neural Networks (IJCNN-91)","author":"Weigend","year":"1991"},{"key":"2020071706564273000_ref13","first-page":"875","article-title":"Generalization by weight-elimination with application to forecasting","volume-title":"Proc. 1990 Conf. Advances in Neural Information Processing Systems","author":"Weigend","year":"1990"},{"key":"2020071706564273000_ref14","first-page":"177","article-title":"Comparing biases for minimal network construction with back-propagation","volume-title":"Proc. 1st Int. Conf. Neural Information Processing Systems","author":"Hanson","year":"1988"},{"key":"2020071706564273000_ref15","volume-title":"The Handbook of Brain Theory and Neural Networks","author":"Arbib","year":"1995"},{"key":"2020071706564273000_ref16","volume-title":"Adaptive Learning of Polynomial Networks: Genetic Programming, Backpropagation and Bayesian Methods","author":"Nikolaev","year":"2006"},{"key":"2020071706564273000_ref17","doi-asserted-by":"crossref","first-page":"771","DOI":"10.1016\/S0893-6080(05)80122-4","article-title":"Improving model selection by nonconvergent methods","volume":"6","author":"Finnoff","year":"1993","journal-title":"Neural Netw."},{"key":"2020071706564273000_ref18","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1007\/3-540-49430-8_18","article-title":"How to train neural networks","volume-title":"Neural networks: tricks of the trade","author":"Neuneier","year":"1998"},{"key":"2020071706564273000_ref19","first-page":"598","article-title":"Optimal brain damage","volume-title":"Advances in Neural Information Processing Systems","author":"Le Cun","year":"1990"},{"key":"2020071706564273000_ref20","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1109\/ICNN.1993.298572","article-title":"Optimal brain surgeon and general network pruning","volume-title":"IEEE Int. Conf. Neural Networks","author":"Hassibi","year":"1993"},{"key":"2020071706564273000_ref21","first-page":"263","article-title":"Optimal brain surgeon: extensions and performance comparisons","volume-title":"Proc. 6th Int. Conf. Neural Information Processing Systems","author":"Hassibi","year":"1993"},{"key":"2020071706564273000_ref22","first-page":"164","article-title":"Second order derivatives for network pruning: optimal brain surgeon","volume-title":"Advances in Neural Information Processing Systems","author":"Hassibi","year":"1993"},{"key":"2020071706564273000_ref23","volume-title":"Memory bounded deep convolutional networks","author":"Collins","year":"2014"},{"key":"2020071706564273000_ref24","volume-title":"Static and Dynamic Neural Networks: From Fundamentals to Advanced Theory","author":"Gupta","year":"2004"},{"key":"2020071706564273000_ref25","doi-asserted-by":"crossref","DOI":"10.5244\/C.29.31","volume-title":"Data-free parameter pruning for deep neural networks","author":"Srinivas","year":"2015"},{"key":"2020071706564273000_ref26","doi-asserted-by":"crossref","first-page":"285","DOI":"10.1093\/biomet\/25.3-4.285","article-title":"On the likelihood that one unknown probability exceeds another in view of the evidence of two samples","volume":"25","author":"Thompson","year":"1933","journal-title":"Biometrika"},{"key":"2020071706564273000_ref27","doi-asserted-by":"crossref","first-page":"450","DOI":"10.2307\/2371219","article-title":"On the theory of apportionment","volume":"57","author":"Thompson","year":"1935","journal-title":"Amer. J. Math."},{"key":"2020071706564273000_ref28","doi-asserted-by":"crossref","first-page":"527","DOI":"10.1090\/S0002-9904-1952-09620-8","article-title":"Some aspects of the sequential design of experiments","volume":"58","author":"Robbins","year":"1952","journal-title":"Bull. Amer. Math. Soc."},{"key":"2020071706564273000_ref29","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1145\/1566374.1566386","article-title":"Characterizing truthful multi-armed bandit mechanisms: extended abstract","volume-title":"Proc. 10th ACM Conf. Electronic Commerce","author":"Babaioff","year":"2009"},{"key":"2020071706564273000_ref30","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1145\/1566374.1566388","article-title":"The price of truthfulness for pay-per-click auctions","volume-title":"Proc. 10th ACM Conf. Electronic Commerce","author":"Devanur","year":"2009"},{"key":"2020071706564273000_ref31","article-title":"Exploration exploitation in Go: UCT for Monte-Carlo Go","volume-title":"NIPS: Neural Information Processing Systems Conference On-line trading of Exploration and Exploitation Workshop","author":"Gelly","year":"2006"},{"key":"2020071706564273000_ref32","first-page":"697","article-title":"Nearly tight bounds for the continuum-armed bandit problem","volume-title":"Proc. 17th Int. Conf. Neural Information Processing Systems","author":"Kleinberg","year":"2004"},{"key":"2020071706564273000_ref33","doi-asserted-by":"crossref","first-page":"681","DOI":"10.1145\/1374376.1374475","article-title":"Multi-armed bandits in metric spaces","volume-title":"Proc. Fortieth Annual ACM Symposium on Theory of Computing","author":"Kleinberg","year":"2008"},{"key":"2020071706564273000_ref34","doi-asserted-by":"crossref","first-page":"672","DOI":"10.1145\/1390156.1390241","article-title":"Empirical Bernstein stopping","volume-title":"Proc. 25th Int. Conf. Machine Learning","author":"Mnih","year":"2008"},{"key":"2020071706564273000_ref35","first-page":"201","article-title":"Online optimization in x-armed bandits","volume-title":"Advances in Neural Information Processing Systems","author":"Bubeck","year":"2008"},{"key":"2020071706564273000_ref36","first-page":"143","article-title":"Fast boosting using adversarial bandits","volume-title":"27th Int. Conf. Machine Learning (ICML 2010)","author":"Busa-Fekete"},{"key":"2020071706564273000_ref37","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton","year":"1998"},{"key":"2020071706564273000_ref38","volume-title":"Playing Atari with deep reinforcement learning","author":"Mnih","year":"2013"},{"key":"2020071706564273000_ref39","doi-asserted-by":"crossref","first-page":"235","DOI":"10.1023\/A:1013689704352","article-title":"Finite-time analysis of the multiarmed bandit problem","volume":"47","author":"Auer","year":"2002","journal-title":"Mach. Learn."},{"key":"2020071706564273000_ref40","first-page":"13","article-title":"Best arm identification in multi-armed bandits","volume-title":"COLT\u201423th Conf. Learning Theory\u20142010","author":"Audibert","year":"2010"},{"key":"2020071706564273000_ref41","doi-asserted-by":"crossref","first-page":"1054","DOI":"10.2307\/1427934","article-title":"Sample mean based index policies by o (log n) regret for the multi-armed bandit problem","volume":"27","author":"Agrawal","year":"1995","journal-title":"Adv. in Appl. Probab."},{"key":"2020071706564273000_ref42","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1137\/S0097539701398375","article-title":"The nonstochastic multiarmed bandit problem","volume":"32","author":"Auer","year":"2002","journal-title":"SIAM J. Comput."},{"key":"2020071706564273000_ref43","first-page":"1015","article-title":"Gaussian process optimization in the bandit setting: no regret and experimental design","volume-title":"Proc. 27th Int. Conf. Machine Learning","author":"Srinivas","year":"2010"},{"key":"2020071706564273000_ref44","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1007\/978-3-642-34106-9_18","article-title":"Thompson sampling: an asymptotically optimal finite-time analysis","volume-title":"Algorithmic Learning Theory","author":"Kaufmann","year":"2012"},{"key":"2020071706564273000_ref45","volume-title":"A survey of online experiment design with the stochastic multi-armed bandit","author":"Burtini","year":"2015"},{"key":"2020071706564273000_ref46","volume-title":"Bandit Algorithms for Website Optimization","author":"White","year":"2012"},{"key":"2020071706564273000_ref47","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1038\/364056a0","article-title":"A strategy of win-stay, lose-shift that outperforms tit-for-tat in the prisoner\u2019s dilemma game","volume":"364","author":"Nowak","year":"1993","journal-title":"Nature"},{"key":"2020071706564273000_ref48","volume-title":"Win Stay\u2014Lose Shift: An Elementary Learning Rule for Normal Form Games","author":"Posch","year":"1997"},{"key":"2020071706564273000_ref49","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1016\/0196-8858(85)90002-8","article-title":"Asymptotically efficient adaptive allocation rules","volume":"6","author":"Lai","year":"1985","journal-title":"Adv. Appl. Math."},{"key":"2020071706564273000_ref50","first-page":"497","article-title":"Finite-time analysis of multi-armed bandits problems with Kullback\u2013Leibler divergences","volume-title":"Proc. 24th Annual Conf. Learning Theory","author":"Maillard","year":"2011"},{"key":"2020071706564273000_ref51","first-page":"359","article-title":"The KL-UCB algorithm for bounded stochastic bandits and beyond","volume-title":"Proc. 24th Annual Conf. Learning Theory, Budapest, Hungary, 09\u201311 June, Proceedings of Machine Learning Research","author":"Garivier","year":"2011"},{"key":"2020071706564273000_ref52","first-page":"99","article-title":"Further optimal regret bounds for thompson sampling","volume-title":"Sixteenth Int. Conf. Artificial Intelligence and Statistics (AISTATS)","author":"Agrawal","year":"2013"},{"key":"2020071706564273000_ref53","first-page":"3212","article-title":"Best arm identification: a unified approach to fixed budget and fixed confidence","volume-title":"Proc. 25th Int. Conf. Neural Information Processing Systems","author":"Gabillon","year":"2012"},{"key":"2020071706564273000_ref54","first-page":"592","article-title":"On Bayesian upper confidence bounds for bandit problems","volume-title":"Proc. Fifteenth Int. Conf. Artificial Intelligence and Statistics","author":"Kaufmann","year":"2012"},{"key":"2020071706564273000_ref55","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1076\/mcmd.8.1.1.8342","article-title":"Nnsysid\u2014toolbox for system identification with neural networks","volume":"8","author":"Norgaard","year":"2002","journal-title":"Math. Comput. Model. Dyn. Syst."},{"key":"2020071706564273000_ref56","volume-title":"UCI machine learning repository","author":"Bache","year":"2013"},{"key":"2020071706564273000_ref57","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"2020071706564273000_ref58","first-page":"1135","article-title":"Learning both weights and connections for efficient neural networks","volume-title":"Proc. 28th Int. Conf. Neural Information Processing Systems","author":"Han","year":"2015"},{"key":"2020071706564273000_ref59","doi-asserted-by":"crossref","first-page":"770","DOI":"10.1093\/bioinformatics\/btx638","article-title":"Machine learning accelerates MD-based binding pose prediction between ligands and proteins","volume":"34","author":"Terayama","year":"2018","journal-title":"Bioinformatics"},{"key":"2020071706564273000_ref60","first-page":"1","article-title":"Statistical comparisons of classifiers over multiple data sets","volume":"7","author":"Dem\u0161ar","year":"2006","journal-title":"J. Mach. Learn. Res."},{"key":"2020071706564273000_ref61","doi-asserted-by":"crossref","first-page":"675","DOI":"10.1080\/01621459.1937.10503522","article-title":"The use of ranks to avoid the assumption of normality implicit in the analysis of variance","volume":"32","author":"Friedman","year":"1937","journal-title":"J. Amer. Statist. Assoc."},{"key":"2020071706564273000_ref62","article-title":"Distribution-free multiple comparisons","author":"Nemenyi","year":"1963"},{"key":"2020071706564273000_ref63","article-title":"On the efficiency of Bayesian bandit algorithms from a frequentist point of view","volume-title":"Neural Information Processing Systems (NIPS)","author":"Kaufmann","year":"2011"}],"container-title":["The Computer Journal"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/comjnl\/article-pdf\/63\/7\/1099\/33505995\/bxz078.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"http:\/\/academic.oup.com\/comjnl\/article-pdf\/63\/7\/1099\/33505995\/bxz078.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,1,10]],"date-time":"2021-01-10T22:05:18Z","timestamp":1610316318000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/comjnl\/article\/63\/7\/1099\/5574718"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,9,26]]},"references-count":63,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2019,9,26]]},"published-print":{"date-parts":[[2020,7,17]]}},"URL":"https:\/\/doi.org\/10.1093\/comjnl\/bxz078","relation":{},"ISSN":["0010-4620","1460-2067"],"issn-type":[{"value":"0010-4620","type":"print"},{"value":"1460-2067","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2020,7]]},"published":{"date-parts":[[2019,9,26]]}}}