{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:28:58Z","timestamp":1760236138773,"version":"build-2065373602"},"reference-count":25,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T00:00:00Z","timestamp":1635379200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Modern computational models in supervised machine learning are often highly parameterized universal approximators. As such, the value of the parameters is unimportant, and only the out of sample performance is considered. On the other hand much of the literature on model estimation assumes that the parameters themselves have intrinsic value, and thus is concerned with bias and variance of parameter estimates, which may not have any simple relationship to out of sample model performance. Therefore, within supervised machine learning, heavy use is made of ridge regression (i.e., L2 regularization), which requires the the estimation of hyperparameters and can be rendered ineffective by certain model parameterizations. We introduce an objective function which we refer to as Information-Corrected Estimation (ICE) that reduces KL divergence based generalization error for supervised machine learning. ICE attempts to directly maximize a corrected likelihood function as an estimator of the KL divergence. Such an approach is proven, theoretically, to be effective for a wide class of models, with only mild regularity restrictions. Under finite sample sizes, this corrected estimation procedure is shown experimentally to lead to significant reduction in generalization error compared to maximum likelihood estimation and L2 regularization.<\/jats:p>","DOI":"10.3390\/e23111419","type":"journal-article","created":{"date-parts":[[2021,10,28]],"date-time":"2021-10-28T23:50:28Z","timestamp":1635465028000},"page":"1419","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Information-Corrected Estimation: A Generalization Error Reducing Parameter Estimation Method"],"prefix":"10.3390","volume":"23","author":[{"given":"Matthew","family":"Dixon","sequence":"first","affiliation":[{"name":"Department of Applied Mathematics, Illinois Institute of Technology, Chicago, IL 60616, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5656-1155","authenticated-orcid":false,"given":"Tyler","family":"Ward","sequence":"additional","affiliation":[{"name":"Department of Financial Engineering, NYU Tandon School of Engineering, New York, NY 11201, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,28]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1214\/aoms\/1177729694","article-title":"On Information and Sufficiency","volume":"22","author":"Kullback","year":"1951","journal-title":"Ann. Math. Stat."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1093","DOI":"10.3982\/ECTA12609","article-title":"Berk\u2013Nash Equilibrium: A Framework for Modeling Agents With Misspecified Models","volume":"84","author":"Esponda","year":"2016","journal-title":"Econometrica"},{"key":"ref_3","first-page":"1711","article-title":"Approximate Bayesian Computation with Kullback-Leibler Divergence as Data Discrepancy","volume":"Volume 84","author":"Storkey","year":"2018","journal-title":"Proceedings of the Twenty-First International Conference on Artificial Intelligence and Statistics, Playa Blanca, Spain, 9\u201311 April 2018"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"5847","DOI":"10.1109\/TIT.2010.2068870","article-title":"Estimating Divergence Functionals and the Likelihood Ratio by Convex Risk Minimization","volume":"56","author":"Nguyen","year":"2010","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"P\u00e9rez-Cruz, F. (2008, January 6\u201311). Kullback-Leibler divergence estimation of continuous distributions. Proceedings of the 2008 IEEE International Symposium on Information Theory, Toronto, ON, Canada.","DOI":"10.1109\/ISIT.2008.4595271"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"875","DOI":"10.1093\/biomet\/83.4.875","article-title":"Generalised Information Criteria in Model Selection","volume":"83","author":"Konishi","year":"1996","journal-title":"Biometrika"},{"key":"ref_7","unstructured":"Roelofs, R. (2019). Measuring Generalization and Overfitting in Machine Learning. [Ph.D. Thesis, University of California]."},{"key":"ref_8","first-page":"1","article-title":"The Jackknife\u2014A Review","volume":"61","author":"Miller","year":"1974","journal-title":"Biometrika"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1093\/biomet\/80.1.27","article-title":"Bias Reduction of Maximum Likelihood Estimates","volume":"80","author":"Firth","year":"1993","journal-title":"Biometrika"},{"key":"ref_10","unstructured":"Kosmidis, I. (2007). Bias Reduction in Exponential Family Nonlinear Models. [Ph.D. Thesis, University of Warwick]."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1097","DOI":"10.1214\/10-EJS579","article-title":"A generic algorithm for reducing bias in parametric estimation","volume":"4","author":"Kosmidis","year":"2010","journal-title":"Electron. J. Stat."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"923","DOI":"10.1093\/biomet\/asx046","article-title":"Median bias reduction of maximum likelihood estimates","volume":"104","author":"Salvan","year":"2017","journal-title":"Biometrika"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1002\/wics.1296","article-title":"Bias in parametric estimation: Reduction and useful side-effects","volume":"6","author":"Kosmidis","year":"2014","journal-title":"Wiley Interdiscip. Rev. Comput. Stat."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"1425","DOI":"10.1162\/089976698300017232","article-title":"Bias\/Variance Decompositions for Likelihood-Based Estimators","volume":"10","author":"Heskes","year":"1998","journal-title":"Neural Comput."},{"key":"ref_15","unstructured":"Petrov, B.N., and Csaki, F. (1971, January 2\u20138). Information Theory and an Extension of the Maximum Likelihood Principle. Proceedings of the 2nd International Symposium on Information Theory, Tsahkadsor, Armenia."},{"key":"ref_16","first-page":"12","article-title":"Distribution of information statistics and validity criteria of models","volume":"153","author":"Takeuchi","year":"1976","journal-title":"Math. Sci."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1111\/j.2517-6161.1977.tb01603.x","article-title":"An Asymptotic Equivalence of Choice of Model by Cross-Validation and Akaike\u2019s Criterion","volume":"39","author":"Stone","year":"1977","journal-title":"J. R. Stat. Soc. Ser. B Methodol."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Hastie, T., Tibshirani, R., and Friedman, J. (2009). The Elements of Statistical Learning, Springer.","DOI":"10.1007\/978-0-387-84858-7"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ross, A., and Doshi-Velez, F. (2018, January 2\u20137). Improving the Adversarial Robustness and Interpretability of Deep Neural Networks by Regularizing Their Input Gradients. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.11504"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"215","DOI":"10.1080\/00401706.1979.10489751","article-title":"Generalized Cross-Validation as a Method for Choosing a Good Ridge Parameter","volume":"21","author":"Golub","year":"1979","journal-title":"Technometrics"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1965","DOI":"10.1016\/j.jmva.2005.10.009","article-title":"Bias correction of cross-validation criterion based on Kullback\u2013Leibler information under a general condition","volume":"97","author":"Yanagihara","year":"2006","journal-title":"J. Multivar. Anal."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1","DOI":"10.2307\/1912526","article-title":"Maximum Likelihood Estimation of Misspecified Models","volume":"50","author":"White","year":"1982","journal-title":"Econometrica"},{"key":"ref_23","unstructured":"Ward, T. (2021, October 02). Information Corrected Estimation: A Bias Correcting Parameter Estimation Method, Supplementary Materials. Available online: https:\/\/figshare.com\/articles\/software\/Information_Corrected_Estimation_A_Bias_Correcting_Parameter_Estimation_Method_Supplementary_Materials\/14312852\/1."},{"key":"ref_24","first-page":"1","article-title":"Multivariate adaptive regression splines","volume":"19","author":"Friedman","year":"1991","journal-title":"Ann. Stat."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"577","DOI":"10.1090\/S0025-5718-1965-0198670-6","article-title":"A Class of Methods for Solving Nonlinear Simultaneous Equations","volume":"19","author":"Broyden","year":"1965","journal-title":"Math. Comput."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/11\/1419\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:22:02Z","timestamp":1760167322000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/11\/1419"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,28]]},"references-count":25,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2021,11]]}},"alternative-id":["e23111419"],"URL":"https:\/\/doi.org\/10.3390\/e23111419","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2021,10,28]]}}}