{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T20:02:59Z","timestamp":1784318579650,"version":"3.55.0"},"reference-count":53,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2025,2,5]],"date-time":"2025-02-05T00:00:00Z","timestamp":1738713600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>We define an evolving in-time Bayesian neural network called a Hidden Markov Neural Network, which addresses the crucial challenge in time-series forecasting and continual learning: striking a balance between adapting to new data and appropriately forgetting outdated information. This is achieved by modelling the weights of a neural network as the hidden states of a Hidden Markov model, with the observed process defined by the available data. A filtering algorithm is employed to learn a variational approximation of the evolving-in-time posterior distribution over the weights. By leveraging a sequential variant of Bayes by Backprop, enriched with a stronger regularization technique called variational DropConnect, Hidden Markov Neural Networks achieve robust regularization and scalable inference. Experiments on MNIST, dynamic classification tasks, and next-frame forecasting in videos demonstrate that Hidden Markov Neural Networks provide strong predictive performance while enabling effective uncertainty quantification.<\/jats:p>","DOI":"10.3390\/e27020168","type":"journal-article","created":{"date-parts":[[2025,2,5]],"date-time":"2025-02-05T10:09:52Z","timestamp":1738750192000},"page":"168","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Hidden Markov Neural Networks"],"prefix":"10.3390","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8846-1482","authenticated-orcid":false,"given":"Lorenzo","family":"Rimella","sequence":"first","affiliation":[{"name":"Dipartimento di Scienze Economico-Sociali e Matematico-Statistiche, University of Torino, 10124 Torino, Italy"},{"name":"Statistics Initiative, Collegio Carlo Alberto, 10124 Torino, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8337-626X","authenticated-orcid":false,"given":"Nick","family":"Whiteley","sequence":"additional","affiliation":[{"name":"School of Mathematics, University of Bristol, Fry Building, Woodland Road, Bristol BS8 1UG, UK"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,2,5]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/MASSP.1986.1165342","article-title":"An introduction to hidden Markov models","volume":"3","author":"Rabiner","year":"1986","journal-title":"IEEE ASSP Mag."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"567","DOI":"10.1006\/jmbi.2000.4315","article-title":"Predicting transmembrane protein topology with a hidden Markov model: Application to complete genomes","volume":"305","author":"Krogh","year":"2001","journal-title":"J. Mol. Biol."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1023\/A:1007425814087","article-title":"Factorial Hidden Markov Models","volume":"29","author":"Ghahramani","year":"1997","journal-title":"Mach. Learn."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"859","DOI":"10.1080\/01621459.2017.1285773","article-title":"Variational inference: A review for statisticians","volume":"112","author":"Blei","year":"2017","journal-title":"J. Am. Stat. Assoc."},{"key":"ref_5","unstructured":"Graves, A. (2011, January 12\u201315). Practical variational inference for neural networks. Proceedings of the Advances in Neural Information Processing Systems, Granada, Spain."},{"key":"ref_6","unstructured":"Kingma, D.P., and Welling, M. (2013). Auto-encoding variational bayes. arXiv."},{"key":"ref_7","unstructured":"Blundell, C., Cornebise, J., Kavukcuoglu, K., and Wierstra, D. (2015, January 6\u201311). Weight uncertainty in neural network. Proceedings of the International Conference on Machine Learning, PMLR, Lile, France."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"5362","DOI":"10.1109\/TPAMI.2024.3367329","article-title":"A comprehensive survey of continual learning: Theory, method and application","volume":"46","author":"Wang","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_9","unstructured":"Kurle, R., Cseke, B., Klushyn, A., van der Smagt, P., and G\u00fcnnemann, S. (2019, January 6\u20139). Continual Learning with Bayesian Neural Networks for Non-Stationary Data. Proceedings of the International Conference on Learning Representations, New Orleans, LA, USA."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"3521","DOI":"10.1073\/pnas.1611835114","article-title":"Overcoming catastrophic forgetting in neural networks","volume":"114","author":"Kirkpatrick","year":"2017","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"ref_11","unstructured":"Nguyen, C.V., Li, Y., Bui, T.D., and Turner, R.E. (2017). Variational continual learning. arXiv."},{"key":"ref_12","unstructured":"Ritter, H., Botev, A., and Barber, D. (2018, January 3\u20138). Online structured laplace approximations for overcoming catastrophic forgetting. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Chopin, N., and Papaspiliopoulos, O. (2020). An Introduction to Sequential Monte Carlo, Springer.","DOI":"10.1007\/978-3-030-47845-2"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"2809","DOI":"10.1214\/14-AAP1061","article-title":"Can local particle filters beat the curse of dimensionality?","volume":"25","author":"Rebeschini","year":"2015","journal-title":"Ann. Appl. Probab."},{"key":"ref_15","first-page":"1","article-title":"Exploiting locality in high-dimensional Factorial hidden Markov models","volume":"23","author":"Rimella","year":"2022","journal-title":"J. Mach. Learn. Res."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1262","DOI":"10.1093\/jrsssc\/qlae035","article-title":"A state-space perspective on modelling and inference for online skill rating","volume":"73","author":"Duffield","year":"2024","journal-title":"J. R. Stat. Soc. Ser. C Appl. Stat."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1080\/00207176808905650","article-title":"Non-linear filtering by approximation of the a posteriori density","volume":"8","author":"Sorenson","year":"1968","journal-title":"Int. J. Control"},{"key":"ref_18","unstructured":"Minka, T.P. (2001). A Family of Algorithms for Approximate Bayesian Inference. [Ph.D. Thesis, Massachusetts Institute of Technology]."},{"key":"ref_19","unstructured":"Wan, L., Zeiler, M., Zhang, S., Le Cun, Y., and Fergus, R. (2013, January 16\u201321). Regularization of neural networks using dropconnect. Proceedings of the International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_20","unstructured":"Franzini, M., Lee, K.F., and Waibel, A. (1990, January 3\u20136). Connectionist Viterbi training: A new hybrid method for continuous speech recognition. Proceedings of the International Conference on Acoustics, Speech, and Signal Processing, Albuquerque, NM, USA."},{"key":"ref_21","unstructured":"Bengio, Y., Cardin, R., De Mori, R., and Normandin, Y. (1990, January 3\u20136). A hybrid coder for hidden Markov models using a recurrent neural networks. Proceedings of the International Conference on Acoustics, Speech, and Signal Processing, Albuquerque, NM, USA."},{"key":"ref_22","unstructured":"Bengio, Y., De Mori, R., Flammia, G., and Kompe, R. (1991, January 8\u201312). Global optimization of a neural network-hidden Markov model hybrid. Proceedings of the IJCNN-91-Seattle International Joint Conference on Neural Networks, Seattle, WA, USA."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"541","DOI":"10.1162\/089976699300016764","article-title":"Hidden neural networks","volume":"11","author":"Krogh","year":"1999","journal-title":"Neural Comput."},{"key":"ref_24","unstructured":"Johnson, M.J., Duvenaud, D., Wiltschko, A.B., Datta, S.R., and Adams, R.P. (2016). Composing graphical models with neural networks for structured representations and fast inference. arXiv."},{"key":"ref_25","unstructured":"Karl, M., Soelch, M., Bayer, J., and Van der Smagt, P. (2016). Deep variational bayes filters: Unsupervised learning of state space models from raw data. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Krishnan, R., Shalit, U., and Sontag, D. (2017, January 4\u20139). Structured inference networks for nonlinear state space models. Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.10779"},{"key":"ref_27","unstructured":"Aitchison, L., Pouget, A., and Latham, P.E. (2014). Probabilistic synapses. arXiv."},{"key":"ref_28","unstructured":"Chang, P.G., Dur\u00e1n-Mart\u00edn, G., Shestopaloff, A.Y., Jones, M., and Murphy, K. (2023). Low-rank extended Kalman filtering for online learning of neural networks from streaming data. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Jones, M., Scott, T.R., Ren, M., ElSayed, G., Hermann, K., Mayo, D., and Mozer, M. (2023, January 1\u20135). Learning in Temporally Structured Environments. Proceedings of the International Conference on Learning Representations, Kigali, Rwanda.","DOI":"10.1609\/aaaiss.v3i1.31273"},{"key":"ref_30","first-page":"18633","article-title":"Online variational filtering and parameter learning","volume":"34","author":"Campbell","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Mobiny, A., Yuan, P., Moulik, S.K., Garg, N., Wu, C.C., and Van Nguyen, H. (2021). Dropconnect is effective in modeling uncertainty of bayesian deep networks. Sci. Rep., 11.","DOI":"10.1038\/s41598-021-84854-x"},{"key":"ref_32","first-page":"1929","article-title":"Dropout: A simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Salehin, I., and Kang, D.K. (2023). A review on dropout regularization approaches for deep neural networks within the scholarly domain. Electronics, 12.","DOI":"10.3390\/electronics12143106"},{"key":"ref_34","unstructured":"Kingma, D.P., Salimans, T., and Welling, M. (2015, January 7\u201312). Variational dropout and the local reparameterization trick. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_35","unstructured":"Gal, Y., and Ghahramani, Z. (2016, January 19\u201324). Dropout as a bayesian approximation: Representing model uncertainty in deep learning. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1109\/72.279191","article-title":"Neurocontrol of nonlinear dynamical systems with Kalman filter trained recurrent networks","volume":"5","author":"Puskorius","year":"1994","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Puskorius, G.V., and Feldkamp, L.A. (2001). Parameter-based Kalman filter training: Theory and implementation. Kalman Filtering and Neural Networks, Wiley.","DOI":"10.1002\/0471221546.ch2"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"683","DOI":"10.1016\/S0893-6080(03)00127-8","article-title":"Simple and conditioned adaptive behavior from Kalman filter trained recurrent networks","volume":"16","author":"Feldkamp","year":"2003","journal-title":"Neural Netw."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"2930","DOI":"10.1214\/18-EJS1468","article-title":"Online natural gradient as a Kalman filter","volume":"12","author":"Ollivier","year":"2018","journal-title":"Electron. J. Stat."},{"key":"ref_40","unstructured":"Aitchison, L. (2018). Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods. arXiv."},{"key":"ref_41","first-page":"19757","article-title":"Knowledge-Adaptation Priors","volume":"Volume 34","author":"Ranzato","year":"2021","journal-title":"Proceedings of the Advances in Neural Information Processing Systems"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/TPAMI.2023.3322743","article-title":"A bayesian federated learning framework with online laplace approximation","volume":"46","author":"Liu","year":"2023","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_43","unstructured":"Sliwa, J., Schneider, F., Bosch, N., Kristiadi, A., and Hennig, P. (2024). Efficient Weight-Space Laplace-Gaussian Filtering and Smoothing for Sequential Deep Learning. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"2278","DOI":"10.1109\/5.726791","article-title":"Gradient-based learning applied to document recognition","volume":"86","author":"LeCun","year":"1998","journal-title":"Proc. IEEE"},{"key":"ref_45","first-page":"2825","article-title":"Scikit-learn: Machine learning in Python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Chan, A.B., and Vasconcelos, N. (2007, January 17\u201322). Classifying video with kernel dynamic textures. Proceedings of the 2007 IEEE Conference on Computer Vision and Pattern Recognition, Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.382996"},{"key":"ref_47","unstructured":"Boots, B., Gordon, G.J., and Siddiqi, S.M. (2008, January 8\u201311). A constraint generation approach to learning stable linear dynamical systems. Proceedings of the Advances in Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Basharat, A., and Shah, M. (October, January 29). Time series prediction by chaotic modeling of nonlinear dynamical systems. Proceedings of the 2009 IEEE 12th International Conference on Computer Vision, Kyoto, Japan.","DOI":"10.1109\/ICCV.2009.5459429"},{"key":"ref_49","unstructured":"Nair, V., and Hinton, G.E. (2010, January 21\u201324). Rectified linear units improve restricted boltzmann machines. Proceedings of the 27th International Conference on Machine Learning (ICML-10), Haifa, Israel."},{"key":"ref_50","unstructured":"Glorot, X., Bordes, A., and Bengio, Y. (2011, January 11\u201313). Deep sparse rectifier neural networks. Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, Fort Lauderdale, FL, USA."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Venkatraman, A., Hebert, M., and Bagnell, J.A. (2015, January 25\u201330). Improving multi-step prediction of learned time series models. Proceedings of the Twenty-Ninth AAAI Conference on Artificial Intelligence, Austin, TX, USA.","DOI":"10.1609\/aaai.v29i1.9590"},{"key":"ref_52","unstructured":"Yang, Y., Pati, D., and Bhattacharya, A. (2017). \u03b1-Variational Inference with Statistical Guarantees. arXiv."},{"key":"ref_53","unstructured":"Ch\u00e9rief-Abdellatif, B.E. (2019). Convergence Rates of Variational Inference in Sparse Deep Learning. arXiv."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/2\/168\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T16:27:24Z","timestamp":1760027244000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/2\/168"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,5]]},"references-count":53,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["e27020168"],"URL":"https:\/\/doi.org\/10.3390\/e27020168","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,5]]}}}