{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,15]],"date-time":"2026-01-15T00:33:32Z","timestamp":1768437212507,"version":"3.49.0"},"reference-count":32,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T00:00:00Z","timestamp":1674518400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"European Commission (INTREPID, Intelligent Toolkit for Reconnaissance and assessmEnt in Perilous Incidents)","award":["883345"],"award-info":[{"award-number":["883345"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>The undeniable computational power of artificial neural networks has granted the scientific community the ability to exploit the available data in ways previously inconceivable. However, deep neural networks require an overwhelming quantity of data in order to interpret the underlying connections between them, and therefore, be able to complete the specific task that they have been assigned to. Feeding a deep neural network with vast amounts of data usually ensures efficiency, but may, however, harm the network\u2019s ability to generalize. To tackle this, numerous regularization techniques have been proposed, with dropout being one of the most dominant. This paper proposes a selective gradient dropout method, which, instead of relying on dropping random weights, learns to freeze the training process of specific connections, thereby increasing the overall network\u2019s sparsity in an adaptive manner, by driving it to utilize more salient weights. The experimental results show that the produced sparse network outperforms the baseline on numerous image classification datasets, and additionally, the yielded results occurred after significantly less training epochs.<\/jats:p>","DOI":"10.3390\/s23031325","type":"journal-article","created":{"date-parts":[[2023,1,25]],"date-time":"2023-01-25T03:23:49Z","timestamp":1674617029000},"page":"1325","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Less Is More: Adaptive Trainable Gradient Dropout for Deep Neural Networks"],"prefix":"10.3390","volume":"23","author":[{"given":"Christos","family":"Avgerinos","sequence":"first","affiliation":[{"name":"Information Technologies Institute (ITI), Centre for Research and Technology Hellas (CERTH), 57001 Thessaloniki, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3604-9685","authenticated-orcid":false,"given":"Nicholas","family":"Vretos","sequence":"additional","affiliation":[{"name":"Information Technologies Institute (ITI), Centre for Research and Technology Hellas (CERTH), 57001 Thessaloniki, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3814-6710","authenticated-orcid":false,"given":"Petros","family":"Daras","sequence":"additional","affiliation":[{"name":"Information Technologies Institute (ITI), Centre for Research and Technology Hellas (CERTH), 57001 Thessaloniki, Greece"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,1,24]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"125","DOI":"10.1080\/00401706.1974.10489157","article-title":"The Relationship between Variable Selection and Data Agumentation and a Method for Prediction","volume":"16","author":"Allen","year":"1974","journal-title":"Technometrics"},{"key":"ref_2","unstructured":"Kohavi, R. (1995). 14th International Joint Conference on Artificial Intelligence\u2014Volume 2, Morgan Kaufmann Publishers Inc.. IJCAI\u201995."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"256","DOI":"10.1006\/inco.1995.1136","article-title":"Boosting a Weak Learning Algorithm by Majority","volume":"121","author":"Freund","year":"1995","journal-title":"Inf. Comput."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"123","DOI":"10.1007\/BF00058655","article-title":"Bagging predictors","volume":"24","author":"Breiman","year":"1996","journal-title":"Mach. Learn."},{"key":"ref_5","unstructured":"Perez, L., and Wang, J. (2017). The Effectiveness of Data Augmentation in Image Classification using Deep Learning. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Cubuk, E.D., Zoph, B., Man\u00e9, D., Vasudevan, V., and Le, Q.V. (2018). AutoAugment: Learning Augmentation Policies from Data. arXiv.","DOI":"10.1109\/CVPR.2019.00020"},{"key":"ref_7","unstructured":"Ohashi, H., Al-Naser, M., Ahmed, S., Akiyama, T., Sato, T., Nguyen, P., Nakamura, K., and Dengel, A. (2017, January 6\u201311). Augmenting Wearable Sensor Data with Physical Constraint for DNN-Based Human-Action Recognition. Proceedings of the ICML 2017 Times Series Workshop, PMLR, Sydney, Australia."},{"key":"ref_8","unstructured":"Prechelt, L. (1998). Neural Networks: Tricks of the Trade, Springer. This Book Is an Outgrowth of a 1996 NIPS Workshop."},{"key":"ref_9","unstructured":"Krogh, A., and Hertz, J.A. A Simple Weight Decay Can Improve Generalization. Proceedings of the 4th International Conference on Neural Information Processing Systems, NIPS\u201991."},{"key":"ref_10","first-page":"1929","article-title":"Dropout: A Simple Way to Prevent Neural Networks from Overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_11","unstructured":"Wan, L., Zeiler, M.D., Zhang, S., LeCun, Y., and Fergus, R. (2013, January 16\u201321). Regularization of Neural Networks using DropConnect. Proceedings of the International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. (2016). Deep Networks with Stochastic Depth. arXiv.","DOI":"10.1007\/978-3-319-46493-0_39"},{"key":"ref_13","unstructured":"Ghiasi, G., Lin, T.Y., and Le, Q.V. (2018). DropBlock: A regularization method for convolutional networks. arXiv."},{"key":"ref_14","unstructured":"DeVries, T., and Taylor, G.W. (2017). Improved Regularization of Convolutional Neural Networks with Cutout. arXiv."},{"key":"ref_15","unstructured":"Larsson, G., Maire, M., and Shakhnarovich, G. (2017). FractalNet: Ultra-Deep Neural Networks without Residuals. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zoph, B., Vasudevan, V., Shlens, J., and Le, Q.V. (2018). Learning Transferable Architectures for Scalable Image Recognition. arXiv.","DOI":"10.1109\/CVPR.2018.00907"},{"key":"ref_17","unstructured":"Gastaldi, X. (2017). Shake-Shake regularization. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"186126","DOI":"10.1109\/ACCESS.2019.2960566","article-title":"Shakedrop Regularization for Deep Residual Learning","volume":"7","author":"Yamada","year":"2019","journal-title":"IEEE Access"},{"key":"ref_19","first-page":"1319","article-title":"Maxout Networks","volume":"Volume 28","author":"Dasgupta","year":"2013","journal-title":"Proceedings of the 30th International Conference on Machine Learning"},{"key":"ref_20","unstructured":"Tseng, H.Y., Chen, Y.W., Tsai, Y.H., Liu, S., Lin, Y.Y., and Yang, M.H. (2020). Regularizing Meta-Learning via Gradient Dropout. arXiv."},{"key":"ref_21","unstructured":"Burges, C.J.C., Bottou, L., Welling, M., Ghahramani, Z., and Weinberger, K.Q. (2013). Proceedings of the Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_22","unstructured":"Gomez, A.N., Zhang, I., Kamalakara, S.R., Madaan, D., Swersky, K., Gal, Y., and Hinton, G.E. (2019). Learning Sparse Networks Using Targeted Dropout. arXiv."},{"key":"ref_23","unstructured":"Lin, H., Zeng, W., Ding, X., Huang, Y., Huang, C., and Paisley, J. (2019). Learning Rate Dropout. arXiv."},{"key":"ref_24","unstructured":"Krizhevsky, A., Nair, V., and Hinton, G. (2019). Learning Multiple Layers of Features from Tiny Images, University of Toronto."},{"key":"ref_25","unstructured":"Seewald, A.K. (2005). Digits\u2014A Dataset for Handwritten Digit Recognition, Institute for Artificial Intelligence."},{"key":"ref_26","unstructured":"Xiao, H., Rasul, K., and Vollgraf, R. (2017). Fashion-MNIST: A Novel Image Dataset for Benchmarking Machine Learning Algorithms. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"141","DOI":"10.1109\/MSP.2012.2211477","article-title":"The Mnist Database of Handwritten Digit Images for Machine Learning Research","volume":"29","author":"Deng","year":"2012","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_28","unstructured":"Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A.Y. (2011). Reading Digits in Natural Images with Unsupervised Feature Learning NIPS Workshop on Deep Learning and Unsupervised Feature Learning, Springer."},{"key":"ref_29","unstructured":"Coates, A., Ng, A., and Lee, H. (,  2011). An Analysis of Single Layer Networks in Unsupervised Feature Learning. Proceedings of the Artificial Intelligence and Statistics AISTATS, Ft. Lauderdale, FL, USA. Available online: https:\/\/cs.stanford.edu\/~acoates\/papers\/coatesleeng_aistats_2011.pdf."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_31","unstructured":"Kingma, D.P., and Welling, M. (2014). Auto-Encoding Variational Bayes. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1671","DOI":"10.4208\/cicp.OA-2020-0165","article-title":"Dying ReLU and Initialization: Theory and Numerical Examples","volume":"28","author":"Lu","year":"2020","journal-title":"Commun. Comput. Phys."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1325\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:14:50Z","timestamp":1760120090000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1325"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,24]]},"references-count":32,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,2]]}},"alternative-id":["s23031325"],"URL":"https:\/\/doi.org\/10.3390\/s23031325","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,24]]}}}