{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:28:02Z","timestamp":1760243282607,"version":"build-2065373602"},"reference-count":23,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2014,7,1]],"date-time":"2014-07-01T00:00:00Z","timestamp":1404172800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/3.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>The principle of extreme physical information (EPI) can be used to derive many known laws and distributions in theoretical physics by extremizing the physical information loss K, i.e., the difference between the observed Fisher information I and the intrinsic information bound J of the physical phenomenon being measured. However, for complex cognitive systems of high dimensionality (e.g., human language processing and image recognition), the information bound J could be excessively larger than I (J \u226b I), due to insufficient observation, which would lead to serious over-fitting problems in the derivation of cognitive models. Moreover, there is a lack of an established exact invariance principle that gives rise to the bound information in universal cognitive systems. This limits the direct application of EPI. To narrow down the gap between I and J, in this paper, we propose a confident-information-first (CIF) principle to lower the information bound J by preserving confident parameters and ruling out unreliable or noisy parameters in the probability density function being measured. The confidence of each parameter can be assessed by its contribution to the expected Fisher information distance between the physical phenomenon and its observations. In addition, given a specific parametric representation, this contribution can often be directly assessed by the Fisher information, which establishes a connection with the inverse variance of any unbiased estimate for the parameter via the Cram\u00e9r\u2013Rao bound. We then consider the dimensionality reduction in the parameter spaces of binary multivariate distributions. We show that the single-layer Boltzmann machine without hidden units (SBM) can be derived using the CIF principle. An illustrative experiment is conducted to show how the CIF principle improves the density estimation performance.<\/jats:p>","DOI":"10.3390\/e16073670","type":"journal-article","created":{"date-parts":[[2014,7,1]],"date-time":"2014-07-01T12:08:07Z","timestamp":1404216487000},"page":"3670-3688","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Extending the Extreme Physical Information to Universal Cognitive Models via a Confident Information First Principle"],"prefix":"10.3390","volume":"16","author":[{"given":"Xiaozhao","family":"Zhao","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuexian","family":"Hou","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"},{"name":"Department of Computing, The Hong Kong Polytechnic University, Hung Hom, Kowloon,Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dawei","family":"Song","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"},{"name":"Department of Computing and Communications, The Open University, Milton Keynes MK76AA, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenjie","family":"Li","sequence":"additional","affiliation":[{"name":"Department of Computing, The Hong Kong Polytechnic University, Hung Hom, Kowloon,Hong Kong, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2014,7,1]]},"reference":[{"key":"ref_1","unstructured":"Wheeler, J.A. (1994). Time Today, Cambridge University Press."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Frieden, B.R. (2004). Science from Fisher Information: A Unification, Cambridge University Press.","DOI":"10.1017\/CBO9780511616907"},{"key":"ref_3","first-page":"81","article-title":"Information and the accuracy attainable in the estimation of statistical parameters","volume":"37","author":"Rao","year":"1945","journal-title":"Bull. Calcutta Math. Soc"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"042144","DOI":"10.1103\/PhysRevE.88.042144","article-title":"Principle of maximum Fisher information from Hardy\u2019s axioms applied to statistical systems","volume":"88","author":"Frieden","year":"2013","journal-title":"Phys. Rev. E"},{"key":"ref_5","unstructured":"Burnham, K.P., and Anderson, D.R. (2002). Model Selection and Multimodel Inference: A Practical Information\u2014Theoretic Approach, Springer."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"975","DOI":"10.1103\/PhysRevE.51.975","article-title":"Interpretation of the extreme physical information principle in terms of shift information","volume":"51","author":"Vstovsky","year":"1995","journal-title":"Phys. Rev. E"},{"key":"ref_7","unstructured":"Amari, S., and Nagaoka, H. (1993). Methods of Information Geometry; Translations of Mathematical Monographs;, Oxford University Press."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"260","DOI":"10.1109\/72.125867","article-title":"Information geometry of Boltzmann machines","volume":"3","author":"Amari","year":"1992","journal-title":"IEEE Trans. Neural Netw"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"12:1","DOI":"10.1145\/2493175.2493177","article-title":"Mining pure high-order word associations via information geometry for information retrieval","volume":"31","author":"Hou","year":"2013","journal-title":"ACM Trans. Inf. Syst"},{"key":"ref_10","first-page":"188","article-title":"The geometry of asymptotic inference","volume":"4","author":"Kass","year":"1989","journal-title":"Stat. Sci"},{"key":"ref_11","first-page":"147","article-title":"A learning algorithm for Boltzmann machines","volume":"9","author":"Ackley","year":"1985","journal-title":"Cogn. Sci"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"504","DOI":"10.1126\/science.1127647","article-title":"Reducing the Dimensionality of Data with Neural Networks","volume":"313","author":"Hinton","year":"2006","journal-title":"Science"},{"key":"ref_13","unstructured":"Rasch, G. (1960). Probabilistic Models for Some Intelligence and Attainment Tests, Danish Institute for Educational Research."},{"key":"ref_14","unstructured":"Bond, T., and Fox, C. (2013). Applying the Rasch Model: Fundamental Measurement in the Human Sciences, Psychology Press."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Gibilisco, P. (2010). Algebraic and Geometric Methods in Statistics, Cambridge University Press.","DOI":"10.1017\/CBO9780511642401"},{"key":"ref_16","unstructured":"\u010cencov, N.N. (1982). Statistical Decision Rules and Optimal Inference, American Mathematical Society."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1214\/aoms\/1177729694","article-title":"On Information and Sufficiency","volume":"22","author":"Kullback","year":"1951","journal-title":"Ann. Math. Stat"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Buhlmann, P., and van de Geer, S. (2011). Statistics for High-Dimensional Data: Methods, Theory And Applications, Springer.","DOI":"10.1007\/978-3-642-20192-9"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1421","DOI":"10.1214\/aos\/1176350602","article-title":"Some classes of global Cram\u00e9r-Rao bounds","volume":"15","author":"Bobrovsky","year":"1987","journal-title":"Ann. Stat"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1771","DOI":"10.1162\/089976602760128018","article-title":"Training products of experts by minimizing contrastive divergence","volume":"14","author":"Hinton","year":"2002","journal-title":"Neural Comput"},{"key":"ref_21","unstructured":"Carreira-Perpinan, M.A., and Hinton, G.E. (2005, January 6\u20138). On contrastive divergence learning.. Key West, FL, USA."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Gilks, W.R., Richardson, S., and Spiegelhalter, D. (1996). Markov Chain Monte Carlo in Practice, Chapman and Hall\/CRC.","DOI":"10.1201\/b14835"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"2269","DOI":"10.1162\/08997660260293238","article-title":"Information geometric measure for neural spikes","volume":"14","author":"Nakahara","year":"2002","journal-title":"Neural Comput"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/16\/7\/3670\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T21:13:09Z","timestamp":1760217189000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/16\/7\/3670"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2014,7,1]]},"references-count":23,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2014,7]]}},"alternative-id":["e16073670"],"URL":"https:\/\/doi.org\/10.3390\/e16073670","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2014,7,1]]}}}