{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,30]],"date-time":"2026-01-30T04:26:51Z","timestamp":1769747211435,"version":"3.49.0"},"reference-count":38,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2025,6,25]],"date-time":"2025-06-25T00:00:00Z","timestamp":1750809600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Fundamental Research Funds for the Universities in Heilongjiang Province","award":["YWF10236220242"],"award-info":[{"award-number":["YWF10236220242"]}]},{"name":"Fundamental Research Funds for the Universities in Heilongjiang Province","award":["YWF10236240126"],"award-info":[{"award-number":["YWF10236240126"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Hyperparameter optimization (HPO), which is also called hyperparameter tuning, is a vital component of developing machine learning models. These parameters, which regulate the behavior of the machine learning algorithm and cannot be directly learned from the given training data, can significantly affect the performance of the model. In the context of relevance vector machine hyperparameter optimization, we have used zero-mean Gaussian weight priors to derive iterative equations through evidence function maximization. For a general Gaussian weight prior and Bayesian linear regression, we similarly derive iterative reestimation equations for hyperparameters through evidence function maximization. Subsequently, after using relative entropy and Bayesian optimization, the aforementioned non-closed-form reestimation equations can be partitioned into E and M steps, providing a clear mathematical and statistical explanation for the iterative reestimation equations of hyperparameters. The experimental result shows the effectiveness of the EM algorithm of hyperparameter optimization, and the algorithm also has the merit of fast convergence, except that the covariance of the posterior distribution is a singular matrix, which affects the increase in the likelihood.<\/jats:p>","DOI":"10.3390\/e27070678","type":"journal-article","created":{"date-parts":[[2025,6,26]],"date-time":"2025-06-26T06:55:13Z","timestamp":1750920913000},"page":"678","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Hyperparameter Optimization EM Algorithm via Bayesian Optimization and Relative Entropy"],"prefix":"10.3390","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2728-0121","authenticated-orcid":false,"given":"Dawei","family":"Zou","sequence":"first","affiliation":[{"name":"School of Information Engineering, Suihua University, Suihua 152061, China"},{"name":"Engineering Technology Research Center of Artificial Intelligence Innovation Application, Suihua University, Suihua 152061, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chunhua","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Suihua University, Suihua 152061, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peng","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Suihua University, Suihua 152061, China"},{"name":"Engineering Technology Research Center of Artificial Intelligence Innovation Application, Suihua University, Suihua 152061, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanqiu","family":"Geng","sequence":"additional","affiliation":[{"name":"School of Information Engineering, Suihua University, Suihua 152061, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,6,25]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Alibrahim, H., and Ludwig, S.A. (July, January 28). Hyperparameter Optimization: Comparing Genetic Algorithm against Grid Search and Bayesian Optimization. Proceedings of the 2021 IEEE Congress on Evolutionary Computation (CEC), Krakow, Poland.","DOI":"10.1109\/CEC45853.2021.9504761"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"249","DOI":"10.3390\/telecom4020015","article-title":"Deep Learning Optimisation of Static Malware Detection with Grid Search and Covering Arrays","volume":"4","author":"Algorain","year":"2023","journal-title":"Telecom"},{"key":"ref_3","unstructured":"Claesen, M., Simm, J., Popovic, D., Moreau, Y., and Moor, B.D. (2014). Easy Hyperparameter Search Using Optunity. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1502","DOI":"10.12928\/telkomnika.v14i4.3956","article-title":"SVM parameter optimization using grid search and genetic algorithm to improve classification performance","volume":"14","author":"Syarif","year":"2016","journal-title":"TELKOMNIKA (Telecommun. Comput. Electron. Control.)"},{"key":"ref_5","unstructured":"Liu, B. (2018). A Very Brief and Critical Discussion on AutoML. arXiv."},{"key":"ref_6","first-page":"281","article-title":"Random Search for Hyper-Parameter Optimization","volume":"13","author":"Bergstra","year":"2012","journal-title":"J. Mach. Learn. Res."},{"key":"ref_7","first-page":"1","article-title":"Practical Bayesian Optimization of Machine Learning Algorithms","volume":"4","author":"Snoek","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1016\/j.is.2018.01.003","article-title":"Genetic Algorithms for Hyperparameter Optimization in Predictive Business Process Monitoring","volume":"74","author":"Dumas","year":"2018","journal-title":"Inf. Syst."},{"key":"ref_9","first-page":"1997","article-title":"Neural architecture search: A survey","volume":"20","author":"Elsken","year":"2019","journal-title":"J. Mach. Learn. Res."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1016\/j.neunet.2019.02.001","article-title":"Gradient based hyperparameter optimization in Echo State Networks","volume":"115","author":"Thiede","year":"2019","journal-title":"Neural Netw."},{"key":"ref_11","unstructured":"Montavon, G., Orr, G.B., and M\u00fcller, K.-R. (2012). Practical Recommendations for Gradient-Based Training of Deep Architectures. Neural Networks: Tricks of the Trade, Springer. [2nd ed.]."},{"key":"ref_12","unstructured":"Ruder, S. (2016). An overview of gradient descent optimization algorithms. arXiv."},{"key":"ref_13","unstructured":"Bergstra, J., Bardenet, R., K\u00e9gl, B., and Bengio, Y. (2011, January 12\u201315). Algorithms for Hyper-Parameter Optimization. Proceedings of the 25th International Conference on Neural Information Processing Systems, Granada, Spain."},{"key":"ref_14","unstructured":"Snoek, J., Rippel, O., Swersky, K., Kiros, R., Satish, N., Sundaram, N., Patwary, M., Prabhat, M., and Adams, R. (2015). Scalable Bayesian Optimization Using Deep Neural Networks. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1109\/JPROC.2015.2494218","article-title":"Taking the Human Out of the Loop: A Review of Bayesian Optimization","volume":"104","author":"Shahriari","year":"2016","journal-title":"Proc. IEEE"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"144","DOI":"10.1016\/j.neucom.2017.03.058","article-title":"Pre-training the deep generative models with adaptive hyperparameter optimization","volume":"247","author":"Yao","year":"2017","journal-title":"Neurocomputing"},{"key":"ref_17","first-page":"2944","article-title":"Efficient and robust automated machine learning","volume":"28","author":"Feurer","year":"2015","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Kaul, A., Maheshwary, S., and Pudi, V. (2017, January 18\u201321). AutoLearn\u2014Automated Feature Generation and Selection. Proceedings of the 2017 IEEE International Conference on Data Mining (ICDM), Orleans, LA, USA.","DOI":"10.1109\/ICDM.2017.31"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Jin, H., Song, Q., and Hu, X. (2019, January 4\u20138). Auto-Keras: An Efficient Neural Architecture Search System. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Anchorage, AK, USA.","DOI":"10.1145\/3292500.3330648"},{"key":"ref_20","first-page":"52","article-title":"AutoML: A systematic review on automated machine learning with neural architecture search","volume":"2","author":"Salehin","year":"2024","journal-title":"J. Inf. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Chartcharnchai, P., Jewajinda, Y., and Praditwong, K. (2024, January 6\u20138). A Categorical Particle Swarm Optimization for Hyperparameter Optimization in Low-Resource Transformer-Based Machine Translation. Proceedings of the 28th International Computer Science and Engineering Conference (ICSEC), Khon Kaen, Thailand.","DOI":"10.1109\/ICSEC62781.2024.10770725"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Indrawati, A., and Wahyuni, I.N. (2023, January 4\u20135). Enhancing Machine Learning Models through Hyperparameter Optimization with Particle Swarm Optimization. Proceedings of the International Conference on Computer, Control, Informatics and Its Applications (IC3INA), Bandung, Indonesia.","DOI":"10.1109\/IC3INA60834.2023.10285736"},{"key":"ref_23","unstructured":"Marchisio, A., Ghillino, E., Curri, V., Carena, A., and Bardella, P. (February, January 27). Particle swarm optimization hyperparameters tuning for physical-model fitting of VCSEL measurements. Proceedings of the SPIE OPTO, San Francisco, CA, USA."},{"key":"ref_24","unstructured":"Bishop, M. (2006). Pattern Recognition and Machine Learning, Springer."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Golub, G.H., and Van Loan, C.F. (2013). Matrix Computations, JHU Press.","DOI":"10.56021\/9781421407944"},{"key":"ref_26","unstructured":"L\u00fctkepohl, H. (1997). Handbook of Matrices, John Wiley & Sons."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Aggarwal, C. (2020). Linear Algebra and Optimization for Machine Learning, Springer.","DOI":"10.1007\/978-3-030-40344-7"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Simovici, D. (2018). Mathematical Analysis for Machine Learning and Data Mining, World Scientific.","DOI":"10.1142\/10702"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"2548310","DOI":"10.1155\/2020\/2548310","article-title":"A Logical Framework of the Evidence Function Approximation Associated with Relevance Vector Machine","volume":"2020","author":"Zou","year":"2020","journal-title":"Math. Probl. Eng."},{"key":"ref_30","unstructured":"Bernardo, M., and Smith, A.F. (2009). Bayesian Theory, John Wiley & Sons."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Gelman, A., Carlin, J.B., Stern, H.S., and Rubin, D.B. (1995). Bayesian Data Analysis, Chapman and Hall\/CRC.","DOI":"10.1201\/9780429258411"},{"key":"ref_32","unstructured":"Berger, J.O. (2013). Statistical Decision Theory and Bayesian Analysis, Springer Science & Business Media."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1378","DOI":"10.1214\/aos\/1176349743","article-title":"A comparison of GCV and GML for choosing the smoothing parameter in the generalized spline smoothing problem","volume":"13","author":"Wahba","year":"1985","journal-title":"Ann. Stat."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"415","DOI":"10.1162\/neco.1992.4.3.415","article-title":"Bayesian Interpolation","volume":"4","author":"MacKay","year":"1992","journal-title":"Neural Comput."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hastie, T., Tibshirani, R., and Friedman, J. (2009). The Elements of Statistical Learning: Data Mining, Inference, and Prediction, Springer. [2nd ed.].","DOI":"10.1007\/978-0-387-84858-7"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"McLachlan, G.J., and Krishnan, T. (2008). The EM Algorithm and Extensions, Wiley-Interscience. [2nd ed.].","DOI":"10.1002\/9780470191613"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","article-title":"Maximum likelihood from incomplete data via the EM algorithm","volume":"39","author":"Dempster","year":"1977","journal-title":"J. R. Stat. Soc."},{"key":"ref_38","unstructured":"Quinonero-Candela, J. (2004). Sparse probabilistic linear models and the RVM. Learning with Uncertainty: Gaussian Processes and Relevance Vector Machines, Technical University of Denmark."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/7\/678\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T17:58:46Z","timestamp":1760032726000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/7\/678"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,25]]},"references-count":38,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2025,7]]}},"alternative-id":["e27070678"],"URL":"https:\/\/doi.org\/10.3390\/e27070678","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,25]]}}}