{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,25]],"date-time":"2026-02-25T01:26:23Z","timestamp":1771982783311,"version":"3.50.1"},"reference-count":56,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2022,1,17]],"date-time":"2022-01-17T00:00:00Z","timestamp":1642377600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-2002908"],"award-info":[{"award-number":["CNS-2002908"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"publisher","award":["N00014-19-1-2621"],"award-info":[{"award-number":["N00014-19-1-2621"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61807021"],"award-info":[{"award-number":["61807021"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Shenzhen Science and Technology Program","award":["KQTD20170810150821146"],"award-info":[{"award-number":["KQTD20170810150821146"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>With the unprecedented performance achieved by deep learning, it is commonly believed that deep neural networks (DNNs) attempt to extract informative features for learning tasks. To formalize this intuition, we apply the local information geometric analysis and establish an information-theoretic framework for feature selection, which demonstrates the information-theoretic optimality of DNN features. Moreover, we conduct a quantitative analysis to characterize the impact of network structure on the feature extraction process of DNNs. Our investigation naturally leads to a performance metric for evaluating the effectiveness of extracted features, called the H-score, which illustrates the connection between the practical training process of DNNs and the information-theoretic framework. Finally, we validate our theoretical results by experimental designs on synthesized data and the ImageNet dataset.<\/jats:p>","DOI":"10.3390\/e24010135","type":"journal-article","created":{"date-parts":[[2022,1,17]],"date-time":"2022-01-17T08:20:42Z","timestamp":1642407642000},"page":"135","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":14,"title":["An Information Theoretic Interpretation to Deep Neural Networks"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4178-0934","authenticated-orcid":false,"given":"Xiangxiang","family":"Xu","sequence":"first","affiliation":[{"name":"Data Science and Information Technology Research Center, Tsinghua\u2013Berkeley Shenzhen Institute, Shenzhen 518055, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2827-4022","authenticated-orcid":false,"given":"Shao-Lun","family":"Huang","sequence":"additional","affiliation":[{"name":"Data Science and Information Technology Research Center, Tsinghua\u2013Berkeley Shenzhen Institute, Shenzhen 518055, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lizhong","family":"Zheng","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA 02139, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9166-4758","authenticated-orcid":false,"given":"Gregory W.","family":"Wornell","sequence":"additional","affiliation":[{"name":"Department of Electrical Engineering and Computer Science, Massachusetts Institute of Technology, Cambridge, MA 02139, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,1,17]]},"reference":[{"key":"ref_1","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_2","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (2019, January 3\u20135). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA. Volume 1 (Long and Short Papers)."},{"key":"ref_3","first-page":"1877","article-title":"Language Models are Few-Shot Learners","volume":"Volume 33","author":"Larochelle","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"484","DOI":"10.1038\/nature16961","article-title":"Mastering the game of Go with deep neural networks and tree search","volume":"529","author":"Silver","year":"2016","journal-title":"Nature"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Arulkumaran, K., Cully, A., and Togelius, J. (2019, January 13\u201317). Alphastar: An evolutionary computation perspective. Proceedings of the Genetic and Evolutionary Computation Conference Companion, Prague, Czech Republic.","DOI":"10.1145\/3319619.3321894"},{"key":"ref_6","unstructured":"MacKay, D.J.C. (2003). Information Theory, Inference, and Learning Algorithms, Cambridge University Press."},{"key":"ref_7","unstructured":"Zintgraf, L.M., Cohen, T.S., Adel, T., and Welling, M. (2017, January 24\u201326). Visualizing Deep Neural Network Decisions: Prediction Difference Analysis. Proceedings of the 5th International Conference on Learning Representations, ICLR 2017, Toulon, France."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"24652","DOI":"10.1073\/pnas.2015509117","article-title":"Prevalence of neural collapse during the terminal phase of deep learning training","volume":"117","author":"Papyan","year":"2020","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_9","unstructured":"Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (2018). Sanity Checks for Saliency Maps. Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3236009","article-title":"A survey of methods for explaining black box models","volume":"51","author":"Guidotti","year":"2018","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_11","unstructured":"Bengio, S., Wallach, H., Larochelle, H., Grauman, K., Cesa-Bianchi, N., and Garnett, R. (2018). Neural Tangent Kernel: Convergence and Generalization in Neural Networks. Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"E7665","DOI":"10.1073\/pnas.1806579115","article-title":"A mean field view of the landscape of two-layer neural networks","volume":"115","author":"Mei","year":"2018","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_13","unstructured":"Arora, S., Du, S., Hu, W., Li, Z., and Wang, R. (2019, January 9\u201315). Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks. Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA."},{"key":"ref_14","unstructured":"Cover, T.M., and Thomas, J.A. (2012). Elements of Information Theory, John Wiley & Sons."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Huang, S.L., Xu, X., Zheng, L., and Wornell, G.W. (2019, January 7\u201312). An information theoretic interpretation to deep neural networks. Proceedings of the 2019 IEEE International Symposium on Information Theory (ISIT), Paris, France.","DOI":"10.1109\/ISIT.2019.8849720"},{"key":"ref_16","unstructured":"Tishby, N., and Zaslavsky, N. (May, January 26). Deep learning and the information bottleneck principle. Proceedings of the Information Theory Workshop (ITW), Jerusalem, Israel."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"19","DOI":"10.1109\/JSAIT.2020.2991561","article-title":"The information bottleneck problem and its applications in machine learning","volume":"1","author":"Goldfeld","year":"2020","journal-title":"IEEE J. Sel. Areas Inf. Theory"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Huang, S.L., Makur, A., Zheng, L., and Wornell, G.W. (2017, January 25\u201330). An information-theoretic approach to universal feature selection in high-dimensional inference. Proceedings of the 2017 IEEE International Symposium on Information Theory (ISIT), Aachen, Germany.","DOI":"10.1109\/ISIT.2017.8006746"},{"key":"ref_19","unstructured":"Arjovsky, M., Chintala, S., and Bottou, L. (2017, January 6\u201311). Wasserstein generative adversarial networks. Proceedings of the International Conference on Machine Learning, PMLR, Sydney, Australia."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"124020","DOI":"10.1088\/1742-5468\/ab3985","article-title":"On the information bottleneck theory of deep learning","volume":"2019","author":"Saxe","year":"2019","journal-title":"J. Stat. Mech. Theory Exp."},{"key":"ref_21","unstructured":"Goodfellow, I., Bengio, J., and Courville, A. (2017). Deep Learning, MIT Press."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"ImageNet Large Scale Visual Recognition Challenge","volume":"115","author":"Olga","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Huang, S.L., and Zheng, L. (2012, January 1\u20136). Linear information coupling problems. Proceedings of the 2012 IEEE International Symposium on Information Theory Proceedings, Cambridge, MA, USA.","DOI":"10.1109\/ISIT.2012.6283007"},{"key":"ref_24","unstructured":"Huang, S.L., Makur, A., Wornell, G.W., and Zheng, L. (2019). On universal features for high-dimensional learning and inference. arXiv."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"520","DOI":"10.1017\/S0305004100013517","article-title":"A connection between correlation and contingency","volume":"31","author":"Hirschfeld","year":"1935","journal-title":"Proc. Camb. Phil. Soc."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"364","DOI":"10.1002\/zamm.19410210604","article-title":"Das statistische problem der Korrelation als variations-und Eigenwertproblem und sein Zusammenhang mit der Ausgleichungsrechnung","volume":"21","author":"Gebelein","year":"1941","journal-title":"Z. Angew. Math. Mech."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"441","DOI":"10.1007\/BF02024507","article-title":"On Measures of Dependence","volume":"10","year":"1959","journal-title":"Acta Math. Acad. Sci. Hung."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"5011","DOI":"10.1109\/TIT.2017.2700857","article-title":"Principal inertia components and applications","volume":"63","author":"Makhdoumi","year":"2017","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Hsu, H., Asoodeh, S., Salamatian, S., and Calmon, F.P. (2018, January 17\u201322). Generalizing bottleneck problems. Proceedings of the 2018 IEEE International Symposium on Information Theory (ISIT), Vail, CO, USA.","DOI":"10.1109\/ISIT.2018.8437632"},{"key":"ref_30","unstructured":"Hsu, H., Salamatian, S., and Calmon, F.P. (2019, January 16\u201318). Correspondence analysis using neural networks. Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, PMLR, Okinawa, Japan,."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Anantharam, V., Gohari, A., Kamath, S., and Nair, C. (July, January 29). On hypercontractivity and a data processing inequality. Proceedings of the 2014 IEEE International Symposium on Information Theory, Honolulu, HI, USA,.","DOI":"10.1109\/ISIT.2014.6875389"},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"3355","DOI":"10.1109\/TIT.2016.2549542","article-title":"Strong data processing inequalities and \u03a6-Sobolev inequalities for discrete channels","volume":"62","author":"Raginsky","year":"2016","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Polyanskiy, Y., and Wu, Y. (2017). Strong data-processing inequalities for channels and Bayesian networks. Convexity and Concentration, Springer.","DOI":"10.1007\/978-1-4939-7005-6_7"},{"key":"ref_34","unstructured":"Greenacre, M.J. (1984). Theory and Applications Of Correspondence Analysis, Academic Press."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"8025","DOI":"10.1109\/TIT.2019.2934414","article-title":"Privacy with estimation guarantees","volume":"65","author":"Wang","year":"2019","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_36","first-page":"614","article-title":"Estimating Optimal Transformations for Multiple Regression and Correlation","volume":"80","author":"Breiman","year":"1985","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_37","unstructured":"Sutskever, I., Vinyals, O., and Le, Q.V. (2014, January 8\u201313). Sequence to sequence learning with neural networks. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Hastie, T., Tibshirani, R., and Friedman, J. (2009). Neural Networks. The Elements of Statistical Learning: Data Mining, Inference, and Prediction, Springer.","DOI":"10.1007\/978-0-387-84858-7"},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/BF02551274","article-title":"Approximation by superpositions of a sigmoidal function","volume":"2","author":"Cybenko","year":"1989","journal-title":"Math. Control. Signals Syst."},{"key":"ref_41","unstructured":"Stoer, J., and Bulirsch, R. (2013). Introduction to Numerical Analysis, Springer Science & Business Media."},{"key":"ref_42","unstructured":"Alain, G., and Bengio, Y. (2017, January 24\u201326). Understanding intermediate layers using linear classifier probes. Proceedings of the 5th International Conference on Learning Representations, ICLR 2017, Toulon, France."},{"key":"ref_43","unstructured":"Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013, January 17\u201319). On the importance of initialization and momentum in deep learning. Proceedings of the International Conference on Machine Learning, PMLR, Atlanta, GA, USA."},{"key":"ref_44","unstructured":"Bengio, Y., and LeCun, Y. (2015, January 7\u20139). Very Deep Convolutional Networks for Large-Scale Image Recognition. Proceedings of the 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA. Conference Track Proceedings."},{"key":"ref_45","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Weinberger, K.Q., and van der Maaten, L. (2017, January 21\u201329). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Chollet, F. (2017, January 21\u201329). Xception: Deep learning with depthwise separable convolutions. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.195"},{"key":"ref_48","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (July, January 26). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA,."},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A.A. (2017, January 4\u20139). Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning. Proceedings of the AAAI, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Xu, X., Huang, S.L., Zheng, L., and Zhang, L. (2018, January 25\u201329). The geometric structure of generalized softmax learning. Proceedings of the 2018 IEEE Information Theory Workshop (ITW), Guangzhou, China.","DOI":"10.1109\/ITW.2018.8613303"},{"key":"ref_51","first-page":"2074","article-title":"Learning structured sparsity in deep neural networks","volume":"29","author":"Wen","year":"2016","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_52","unstructured":"Wang, L., Wu, J., Huang, S.L., Zheng, L., Xu, X., Zhang, L., and Huang, J. (February, January 27). An efficient approach to informative feature extraction from multimodal data. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_53","first-page":"4370","article-title":"Learning new tricks from old dogs: Multi-source transfer learning from pre-trained networks","volume":"32","author":"Lee","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Dembo, A., and Zeitouni, O. (2010). Large Deviations Techniques and Applications, Springer. Stochastic Modelling and Applied Probability.","DOI":"10.1007\/978-3-642-03311-7"},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"5973","DOI":"10.1109\/TIT.2016.2603151","article-title":"f-divergence Inequalities","volume":"62","author":"Sason","year":"2016","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/BF02288367","article-title":"The approximation of one matrix by another of lower rank","volume":"1","author":"Eckart","year":"1936","journal-title":"Psychometrika"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/1\/135\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:02:34Z","timestamp":1760133754000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/1\/135"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,17]]},"references-count":56,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,1]]}},"alternative-id":["e24010135"],"URL":"https:\/\/doi.org\/10.3390\/e24010135","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,17]]}}}