{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:13:19Z","timestamp":1760242399013,"version":"build-2065373602"},"reference-count":29,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2017,6,30]],"date-time":"2017-06-30T00:00:00Z","timestamp":1498780800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Chinese 863 Program","award":["2015AA015403"],"award-info":[{"award-number":["2015AA015403"]}]},{"name":"the Key Project of Tianjin Natural Science Foundation","award":["15JCZDJC31100"],"award-info":[{"award-number":["15JCZDJC31100"]}]},{"name":"the Tianjin Younger Natural Science Foundation","award":["14JCQNJC00400"],"award-info":[{"award-number":["14JCQNJC00400"]}]},{"name":"the Major Project of Chinese National Social Science Fund","award":["14ZDB153"],"award-info":[{"award-number":["14ZDB153"]}]},{"name":"MSCA-ITN-ETN - European Training Networks Project","award":["721321, QUARTZ"],"award-info":[{"award-number":["721321, QUARTZ"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Regularization of neural networks can alleviate overfitting in the training phase. Current regularization methods, such as Dropout and DropConnect, randomly drop neural nodes or connections based on a uniform prior. Such a data-independent strategy does not take into consideration of the quality of individual unit or connection. In this paper, we aim to develop a data-dependent approach to regularizing neural network in the framework of Information Geometry. A measurement for the quality of connections is proposed, namely confidence. Specifically, the confidence of a connection is derived from its contribution to the Fisher information distance. The network is adjusted by retaining the confident connections and discarding the less confident ones. The adjusted network, named as ConfNet, would carry the majority of variations in the sample data. The relationships among confidence estimation, Maximum Likelihood Estimation and classical model selection criteria (like Akaike information criterion) is investigated and discussed theoretically. Furthermore, a Stochastic ConfNet is designed by adding a self-adaptive probabilistic sampling strategy. The proposed data-dependent regularization methods achieve promising experimental results on three data collections including MNIST, CIFAR-10 and CIFAR-100.<\/jats:p>","DOI":"10.3390\/e19070313","type":"journal-article","created":{"date-parts":[[2017,6,30]],"date-time":"2017-06-30T10:04:58Z","timestamp":1498817098000},"page":"313","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Regularizing Neural Networks via Retaining Confident Connections"],"prefix":"10.3390","volume":"19","author":[{"given":"Shengnan","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuexian","family":"Hou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"},{"name":"Department of Computing, The Hong Kong Polytechnic University, Hung Hom, Kowloon, Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1501-9914","authenticated-orcid":false,"given":"Benyou","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dawei","family":"Song","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"},{"name":"Department of Computing and Communications, The Open University, Milton Keynes MK76AA, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2017,6,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","article-title":"Deep Neural Networks for Acoustic Modeling in Speech Recognition","volume":"29","author":"Hinton","year":"2012","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Mikolov, T., Deoras, A., Povey, D., and Burget, L. (2011, January 11\u201315). Strategies for training large scale neural network language models. Proceedings of the 2011 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU), Waikoloa, HI, USA.","DOI":"10.1109\/ASRU.2011.6163930"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Sainath, T.N., Mohamed, A.R., Kingsbury, B., and Ramabhadran, B. (2013, January 26\u201331). Deep convolutional neural networks for LVCSR. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6639347"},{"key":"ref_4","first-page":"2012","article-title":"ImageNet Classification with Deep Convolutional Neural Networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1915","DOI":"10.1109\/TPAMI.2012.231","article-title":"Learning Hierarchical Features for Scene Labeling","volume":"35","author":"Farabet","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_6","unstructured":"Tompson, J., Jain, A., Lecun, Y., and Bregler, C. (2014, January 8\u201313). Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation. Proceedings of the 27th International Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Collobert, R., and Weston, J. (2008, January 5\u20139). A unified architecture for natural language processing: Deep neural networks with multitask learning. Proceedings of the 25th International Conference on Machine learning, Helsinki, Finland.","DOI":"10.1145\/1390156.1390177"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"504","DOI":"10.1126\/science.1127647","article-title":"Reducing the dimensionality of data with neural networks","volume":"313","author":"Hinton","year":"2006","journal-title":"Science"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"528","DOI":"10.1080\/01621459.1987.10478458","article-title":"The Calculation of Posterior Distributions by Data Augmentation","volume":"82","author":"Tanner","year":"1987","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_10","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_11","unstructured":"Wan, L., Zeiler, M., Zhang, S., Cun, Y.L., and Fergus, R. (2013, January 16\u201321). Regularization of neural networks using dropconnect. Proceedings of the 30th International Conference on International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1016\/0893-6080(90)90049-Q","article-title":"Probabilistic Neural Networks","volume":"3","author":"Specht","year":"1990","journal-title":"Neural Netw."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Zhao, X., Hou, Y., Song, D., and Li, W. (2017). A Confident Information First Principle for Parameter Reduction and Model Selection of Boltzmann Machines. IEEE Trans. Neural Netw. Learn. Syst.","DOI":"10.1109\/TNNLS.2017.2664100"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1109\/72.125867","article-title":"Information geometry of Boltzmann machines","volume":"3","author":"Amari","year":"1992","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"1527","DOI":"10.1162\/neco.2006.18.7.1527","article-title":"A fast learning algorithm for deep belief nets","volume":"18","author":"Hinton","year":"2006","journal-title":"Neural Comput."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1145\/2493175.2493177","article-title":"Mining pure high-order word associations via information geometry for information retrieval","volume":"31","author":"Hou","year":"2013","journal-title":"ACM Trans. Inf. Syst."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Amari, S., and Nagaoka, H. (2007). Methods of Information Geometry (Translations of Mathematical Monographs), American Mathematical Society.","DOI":"10.1090\/mmono\/191"},{"key":"ref_18","unstructured":"Rao, C.R. (1945). Information and the Accuracy Attainable in the Estimation of Statistical Parameters, Springer."},{"key":"ref_19","first-page":"188","article-title":"The Geometry of Asymptotic Inference","volume":"4","author":"Kass","year":"1989","journal-title":"Stat. Sci."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"3670","DOI":"10.3390\/e16073670","article-title":"Extending the extreme physical information to universal cognitive models via a confident information first principle","volume":"16","author":"Zhao","year":"2014","journal-title":"Entropy"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Gibilisco, P., Riccomagno, E., Rogantin, M.P., and Wynn, H.P. (2010). Algebraic and Geometric Methods in Statistics, Cambridge University Press.","DOI":"10.1017\/CBO9780511642401"},{"key":"ref_22","unstructured":"Cencov, N.N. (1982). Statistical Decision Rules and Optimal Inference, American Mathematical Society."},{"key":"ref_23","unstructured":"Hinton, G.E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R.R. (arXiv, 2012). Improving neural networks by preventing co-adaptation of feature detectors, arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"1701","DOI":"10.1109\/18.930911","article-title":"Information geometry on hierarchy of probability distributions","volume":"47","author":"Amari","year":"2001","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"716","DOI":"10.1109\/TAC.1974.1100705","article-title":"IEEE Xplore Abstract\u2014A new look at the statistical model identification","volume":"19","author":"Akaike","year":"1974","journal-title":"IEEE Trans. Autom. Control"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"345","DOI":"10.1007\/BF02294361","article-title":"Model selection and Akaike\u2019s Information Criterion (AIC): The general theory and its analytical extensions","volume":"52","author":"Bozdogan","year":"1987","journal-title":"Psychometrika"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"2269","DOI":"10.1162\/08997660260293238","article-title":"Information-geometric measure for neural spikes","volume":"14","author":"Nakahara","year":"2002","journal-title":"Neural Comput."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"Lecun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_29","unstructured":"Krizhevsky, A., and Hinton, G. (2017, June 26). Learning Multiple Layers of Features from Tiny Images. Available online: http:\/\/www.cs.utoronto.ca\/kriz\/learning-features-2009-TR.pdf."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/19\/7\/313\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T18:40:55Z","timestamp":1760208055000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/19\/7\/313"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,6,30]]},"references-count":29,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2017,7]]}},"alternative-id":["e19070313"],"URL":"https:\/\/doi.org\/10.3390\/e19070313","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2017,6,30]]}}}