{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,10]],"date-time":"2026-07-10T16:24:03Z","timestamp":1783700643206,"version":"3.55.0"},"reference-count":176,"publisher":"Emerald","issue":"1-2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,1,1]]},"abstract":"<jats:p>Variational autoencoders (VAEs) are powerful deep generative models widely used to represent high-dimensional complex data through a low-dimensional latent space learned in an unsupervised manner. In the original VAE model, the input data vectors are processed independently. Recently, a series of papers have presented different extensions of the VAE to process sequential data, which model not only the latent space but also the temporal dependencies within a sequence of data vectors and corresponding latent vectors, relying on recurrent neural networks or state-space models. In this monograph, we perform a literature review of these models. We introduce and discuss a general class of models, called dynamical variational autoencoders (DVAEs), which encompasses a large subset of these temporal VAE extensions. Then, we present in detail seven recently proposed DVAE models, with an aim to homogenize the notations and presentation lines, as well as to relate these models with existing classical temporal models. We have reimplemented those seven DVAE models and present the results of an experimental benchmark conducted on the speech analysis-resynthesis task (the PyTorch code is made publicly available). The monograph concludes with a discussion on important issues concerning the DVAE class of models and future research guidelines.<\/jats:p>","DOI":"10.1561\/2200000089","type":"journal-article","created":{"date-parts":[[2021,12,2]],"date-time":"2021-12-02T04:01:59Z","timestamp":1638417719000},"page":"1-175","source":"Crossref","is-referenced-by-count":142,"title":["Dynamical Variational Autoencoders: A Comprehensive Review"],"prefix":"10.1108","volume":"15","author":[{"given":"Laurent","family":"Girin","sequence":"first","affiliation":[{"name":"Univ. Grenoble Alpes , CNRS, Grenoble-INP, GIPSA-lab"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Simon","family":"Leglaive","sequence":"additional","affiliation":[{"name":"CentraleSup\u00e9lec , IETR"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaoyu","family":"Bie","sequence":"additional","affiliation":[{"name":"Inria, Univ. Grenoble Alpes , CNRS, LJK"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Julien","family":"Diard","sequence":"additional","affiliation":[{"name":"Univ. Grenoble Alpes , CNRS, LPNC"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas","family":"Hueber","sequence":"additional","affiliation":[{"name":"Univ. Grenoble Alpes , CNRS, Grenoble-INP, GIPSA-lab"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xavier","family":"Alameda-Pineda","sequence":"additional","affiliation":[{"name":"Inria, Univ. Grenoble Alpes , CNRS, LJK"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2022,1,1]]},"reference":[{"key":"2026033012252967000_ref001","volume-title":"Tensorflow: Large-scale machine learning on heterogeneous distributed systems","author":"Abadi","year":"2016"},{"key":"2026033012252967000_ref002","volume-title":"STCN: Stochastic temporal convo-lutional networks","author":"Aksan","year":"2019"},{"key":"2026033012252967000_ref003","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Andrychowicz","year":"2016"},{"key":"2026033012252967000_ref004","volume-title":"Black box variational inference for state space models","author":"Archer","year":"2015"},{"key":"2026033012252967000_ref005","volume-title":"Stochastic variational video prediction","author":"Babaeizadeh","year":"2018"},{"key":"2026033012252967000_ref006","volume-title":"An empirical evaluation of generic convolutional and recurrent networks for sequence modeling","author":"Bai","year":"2018"},{"key":"2026033012252967000_ref007","doi-asserted-by":"crossref","DOI":"10.1109\/ICASSP.2018.8461530","volume-title":"Statistical speech enhancement based on probabilistic integration of variational autoencoder and non-negative matrix factorization","author":"Bando","year":"2018"},{"key":"2026033012252967000_ref008","volume-title":"Learning stochastic recurrent networks","author":"Bayer","year":"2014"},{"key":"2026033012252967000_ref009","volume-title":"Scheduled sampling for sequence prediction with recurrent neural networks","author":"Bengio","year":"2015"},{"key":"2026033012252967000_ref010","volume-title":"Unsu-pervised speech enhancement using dynamical variational autoencoders","author":"Bie","year":"2021"},{"key":"2026033012252967000_ref011","volume-title":"Pattern Recognition and Machine Learning","author":"Bishop","year":"2006"},{"key":"2026033012252967000_ref012","volume-title":"Neural granular sound synthesis","author":"Bitton","year":"2020"},{"key":"2026033012252967000_ref013","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2016-1183","volume-title":"Modeling and transforming speech using variational autoencoders","author":"Blaauw","year":"2016"},{"issue":"518","key":"2026033012252967000_ref014","doi-asserted-by":"crossref","first-page":"859","DOI":"10.1080\/01621459.2017.1285773","article-title":"Variational inference: A review for statisticians","volume":"112","author":"Blei","year":"2017","journal-title":"Journal of the American Statistical Association"},{"key":"2026033012252967000_ref015","volume-title":"Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image","author":"Bogo","year":"2016"},{"key":"2026033012252967000_ref016","doi-asserted-by":"crossref","first-page":"146","DOI":"10.1007\/978-3-540-28650-9_7","article-title":"Stochastic Learning","volume-title":"Advanced Lectures on Machine Learning: ML Summer Schools 2003, Canberra, Australia","author":"Bottou","year":"2004"},{"key":"2026033012252967000_ref017","volume-title":"Multi-level variational autoencoder: Learning disentangled representations from grouped observations","author":"Bouchacourt","year":"2018"},{"key":"2026033012252967000_ref018","volume-title":"Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription","author":"Boulanger-Lewandowski","year":"2012"},{"key":"2026033012252967000_ref019","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/K16-1002","article-title":"Generating sentences from a continuous space","author":"Bowman","year":"2016"},{"key":"2026033012252967000_ref020","volume-title":"Isolating sources of disentanglement in variational autoencoders","author":"Chen","year":"2018"},{"key":"2026033012252967000_ref021","volume-title":"Variational lossy autoencoder","author":"Chen","year":"2017"},{"key":"2026033012252967000_ref022","doi-asserted-by":"crossref","DOI":"10.3115\/v1\/D14-1179","volume-title":"Learning phrase representations using RNN encoder-decoder for statistical machine translation","author":"Cho","year":"2014"},{"key":"2026033012252967000_ref023","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Chung","year":"2015"},{"key":"2026033012252967000_ref024","volume-title":"Inference suboptimality in variational autoencoders","author":"Cremer","year":"2018"},{"key":"2026033012252967000_ref025","volume-title":"The usual suspects? Reassessing blame for VAE posterior collapse","author":"Dai","year":"2020"},{"issue":"8","key":"2026033012252967000_ref026","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1109\/MAES.2005.1499276","article-title":"Nonlinear filters: Beyond the Kalman filter","volume":"20","author":"Daum","year":"2005","journal-title":"IEEE Aerospace and Electronic Systems Magazine"},{"issue":"5","key":"2026033012252967000_ref027","doi-asserted-by":"crossref","first-page":"889","DOI":"10.1162\/neco.1995.7.5.889","article-title":"The Helmholtz machine","volume":"7","author":"Dayan","year":"1995","journal-title":"Neural Computation"},{"issue":"1","key":"2026033012252967000_ref028","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1111\/j.2517-6161.1977.tb01600.x","article-title":"Maximum likelihood from incomplete data via the EM algorithm","volume":"39","author":"Dempster","year":"1977","journal-title":"Journal of the Royal Statistical Society. Series B (Methodological)"},{"key":"2026033012252967000_ref029","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.637","volume-title":"Factorized variational autoencoders for modeling audience reactions to movies","author":"Deng","year":"2017"},{"key":"2026033012252967000_ref030","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Denton","year":"2017"},{"key":"2026033012252967000_ref031","volume-title":"Stochastic video generation with a learned prior","author":"Denton","year":"2018"},{"issue":"2","key":"2026033012252967000_ref032","doi-asserted-by":"crossref","first-page":"193","DOI":"10.1111\/j.2517-6161.1984.tb01290.x","article-title":"Monte Carlo methods of inference for implicit statistical models","volume":"46","author":"Diggle","year":"1984","journal-title":"Journal of the Royal Statistical Society: Series B (Methodological)"},{"key":"2026033012252967000_ref033","doi-asserted-by":"crossref","DOI":"10.1093\/acprof:oso\/9780199641178.001.0001","volume-title":"Time series analysis by state space methods","author":"Durbin","year":"2012"},{"key":"2026033012252967000_ref034","volume-title":"Bridging audio analysis, perception and synthesis with perceptually-regularized variational timbre space","author":"Esling","year":"2018"},{"key":"2026033012252967000_ref035","volume-title":"Variational recurrent auto-encoders","author":"Fabius","year":"2014"},{"key":"2026033012252967000_ref036","article-title":"Unsupervised learning for physical interaction through video prediction","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Finn","year":"2016"},{"key":"2026033012252967000_ref037","volume-title":"GP-VAE: Deep probabilistic time series imputation","author":"Fortuin","year":"2020"},{"issue":"4","key":"2026033012252967000_ref038","doi-asserted-by":"crossref","first-page":"1569","DOI":"10.1109\/TSP.2010.2102756","article-title":"Bayesian nonparametric inference of switching dynamic linear models","volume":"59","author":"Fox","year":"2011","journal-title":"IEEE Transactions on Signal Processing"},{"key":"2026033012252967000_ref039","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Fraccaro","year":"2016"},{"key":"2026033012252967000_ref040","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Fraccaro","year":"2017"},{"key":"2026033012252967000_ref041","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/3348.001.0001","volume-title":"Graphical models for machine learning and digital communication","author":"Frey","year":"1998"},{"key":"2026033012252967000_ref042","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Gan","year":"2015"},{"key":"2026033012252967000_ref043","doi-asserted-by":"crossref","DOI":"10.1109\/ICASSP.2019.8683277","volume-title":"Low bit-rate speech coding with VQ-VAE and a WaveNet decoder","author":"G\u00e2rbacea","year":"2019"},{"key":"2026033012252967000_ref044","article-title":"CSR-I (WSJ0) Sennheiser LDC93S6B. https:\/\/catalog.ldc.upenn.edu\/LDC93S6B","volume-title":"Philadelphia: Linguistic Data Consortium","author":"Garofolo","year":"1993"},{"issue":"5","key":"2026033012252967000_ref045","doi-asserted-by":"crossref","first-page":"507","DOI":"10.1002\/net.3230200504","article-title":"Identifying independence in Bayesian networks","volume":"20","author":"Geiger","year":"1990","journal-title":"Networks"},{"key":"2026033012252967000_ref046","volume-title":"Vector quantization and signal compression","author":"Gersho","year":"2012"},{"key":"2026033012252967000_ref047","volume-title":"Notes on the use of variational autoencoders for speech and audio spectrogram modeling","author":"Girin","year":"2019"},{"key":"2026033012252967000_ref048","volume-title":"Deep Learning","author":"Goodfellow","year":"2016"},{"key":"2026033012252967000_ref049","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Goodfellow","year":"2014"},{"key":"2026033012252967000_ref050","volume-title":"NIPS 2016 tutorial: Generative adversarial networks","author":"Goodfellow","year":"2016"},{"key":"2026033012252967000_ref051","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Goyal","year":"2017"},{"key":"2026033012252967000_ref052","volume-title":"Generating sequences with recurrent neural networks","author":"Graves","year":"2013"},{"key":"2026033012252967000_ref053","doi-asserted-by":"crossref","DOI":"10.1109\/ICASSP.2013.6638947","volume-title":"Speech recognition with deep recurrent neural networks","author":"Graves","year":"2013"},{"key":"2026033012252967000_ref054","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Gregor","year":"2016"},{"key":"2026033012252967000_ref055","volume-title":"DRAW: A recurrent neural network for image generation","author":"Gregor","year":"2015"},{"key":"2026033012252967000_ref056","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Gu","year":"2015"},{"key":"2026033012252967000_ref057","volume-title":"PixelVAE: A latent variable model for natural images","author":"Gulrajani","year":"2016"},{"key":"2026033012252967000_ref058","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV.2019.00713","volume-title":"Video compression with rate-distortion autoencoders","author":"Habibian","year":"2019"},{"key":"2026033012252967000_ref059","volume-title":"Learning latent dynamics for planning from pixels","author":"Hafner","year":"2018"},{"key":"2026033012252967000_ref060","doi-asserted-by":"crossref","DOI":"10.2307\/j.ctv14jx6sm","volume-title":"Time series analysis","author":"Hamilton","year":"2020"},{"key":"2026033012252967000_ref061","volume-title":"Kalman filtering and neural networks","author":"Haykin","year":"2004"},{"key":"2026033012252967000_ref062","volume-title":"Lagging inference networks and posterior collapse in variational autoencoders","author":"He","year":"2018"},{"key":"2026033012252967000_ref063","volume-title":"\u03b2-VAE: learning basic visual concepts with a constrained variational framework","author":"Higgins","year":"2017"},{"issue":"5786","key":"2026033012252967000_ref064","doi-asserted-by":"crossref","first-page":"504","DOI":"10.1126\/science.1127647","article-title":"Reducing the dimensionality of data with neural networks","volume":"313","author":"Hinton","year":"2006","journal-title":"Science"},{"issue":"5214","key":"2026033012252967000_ref065","doi-asserted-by":"crossref","first-page":"1158","DOI":"10.1126\/science.7761831","article-title":"The \u201cWake-Sleep\u201d algorithm for unsupervised neural networks","volume":"268","author":"Hinton","year":"1995","journal-title":"Science"},{"issue":"8","key":"2026033012252967000_ref066","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Computation"},{"issue":"1","key":"2026033012252967000_ref067","first-page":"1303","article-title":"Stochastic variational inference","volume":"14","author":"Hoffman","year":"2013","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012252967000_ref068","first-page":"3235","article-title":"Approximate Riemannian conjugate gradient learning for fixed-form variational Bayes","volume":"11","author":"Honkela","year":"2010","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012252967000_ref069","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2017-349","volume-title":"Learning latent representations for speech generation and transformation","author":"Hsu","year":"2017"},{"key":"2026033012252967000_ref070","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Hsu","year":"2017"},{"key":"2026033012252967000_ref071","volume-title":"Toward controlled generation of text","author":"Hu","year":"2017"},{"issue":"19","key":"2026033012252967000_ref072","doi-asserted-by":"crossref","first-page":"5052","DOI":"10.1109\/TSP.2016.2576427","article-title":"A flexible and efficient algorithmic framework for constrained matrix and tensor factorization","volume":"64","author":"Huang","year":"2016","journal-title":"IEEE Transactions on Signal Processing"},{"issue":"7","key":"2026033012252967000_ref073","doi-asserted-by":"crossref","first-page":"1325","DOI":"10.1109\/TPAMI.2013.248","article-title":"Hu-man3.6M: Large scale datasets and predictive methods for 3D human sensing in natural environments","volume":"36","author":"Ionescu","year":"2014","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026033012252967000_ref074","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1016\/j.ins.2019.03.066","article-title":"Recurrent neural network-based semantic variational autoencoder for sequence-to-sequence learning","volume":"490","author":"Jang","year":"2019","journal-title":"Information Sciences"},{"key":"2026033012252967000_ref075","volume-title":"Transformer VAE: A hierarchical model for structure-aware and interpretable music representation learning","author":"Jiang","year":"2020"},{"key":"2026033012252967000_ref076","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Johnson","year":"2016"},{"key":"2026033012252967000_ref077","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1023\/A:1007665907178","article-title":"An introduction to variational methods for graphical models","volume":"37","author":"Jordan","year":"1999","journal-title":"Machine Learning"},{"key":"2026033012252967000_ref078","volume-title":"Semi-blind source separation with multichannel variational autoencoder","author":"Kameoka","year":"2018"},{"key":"2026033012252967000_ref079","volume-title":"Deep variational Bayes filters: Unsupervised learning of state space models from raw data","author":"Karl","year":"2017"},{"key":"2026033012252967000_ref080","volume-title":"Disentangling by factorising","author":"Kim","year":"2018"},{"key":"2026033012252967000_ref081","volume-title":"Semi-amortized variational autoencoders","author":"Kim","year":"2018"},{"key":"2026033012252967000_ref082","volume-title":"Adam: A method for stochastic optimization","author":"Kingma","year":"2014"},{"key":"2026033012252967000_ref083","volume-title":"Auto-encoding variational Bayes","author":"Kingma","year":"2014"},{"issue":"4","key":"2026033012252967000_ref084","doi-asserted-by":"crossref","first-page":"307","DOI":"10.1561\/2200000056","article-title":"An introduction to variational autoencoders","volume":"12","author":"Kingma","year":"2019","journal-title":"Foundations and Trends in Machine Learning"},{"key":"2026033012252967000_ref085","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Kingma","year":"2016"},{"key":"2026033012252967000_ref086","volume-title":"Probabilistic graphical models: Principles and techniques","author":"Koller","year":"2009"},{"key":"2026033012252967000_ref087","volume-title":"Deep Kalman filters","author":"Krishnan","year":"2015"},{"key":"2026033012252967000_ref088","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v31i1.10779","volume-title":"Structured inference networks for nonlinear state space models","author":"Krishnan","year":"2017"},{"key":"2026033012252967000_ref089","volume-title":"On the challenges of learning with inference networks on sparse, high-dimensional data","author":"Krishnan","year":"2018"},{"key":"2026033012252967000_ref090","first-page":"507","article-title":"Tensor factorization via matrix factorization","volume-title":"Artificial Intelligence and Statistics","author":"Kuleshov","year":"2015"},{"key":"2026033012252967000_ref091","volume-title":"Auto-encoding sequential Monte Carlo","author":"Le","year":"2018"},{"key":"2026033012252967000_ref092","article-title":"SDR: Half-baked or well done?","author":"Le Roux","year":"2019"},{"key":"2026033012252967000_ref093","volume-title":"Temporal convolutional networks: A unified approach to action segmentation","author":"Lea","year":"2016"},{"key":"2026033012252967000_ref094","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2018-1598","volume-title":"Acoustic modeling using adversarially trained variational recurrent neural network for speech synthesis","author":"Lee","year":"2018"},{"key":"2026033012252967000_ref095","doi-asserted-by":"crossref","DOI":"10.1109\/MLSP.2018.8516711","volume-title":"A variance modeling framework based on variational autoencoders for speech enhancement","author":"Leglaive","year":"2018"},{"key":"2026033012252967000_ref096","doi-asserted-by":"crossref","DOI":"10.1109\/ICASSP.2019.8683704","volume-title":"Semi-supervised multichannel speech enhancement with variational autoencoders and non-negative matrix factorization","author":"Leglaive","year":"2019"},{"key":"2026033012252967000_ref097","doi-asserted-by":"crossref","DOI":"10.1109\/ICASSP40776.2020.9053164","volume-title":"A recurrent variational autoencoder for speech enhancement","author":"Leglaive","year":"2020"},{"key":"2026033012252967000_ref098","doi-asserted-by":"crossref","DOI":"10.1109\/WASPAA.2017.8169994","volume-title":"An EM algorithm for audio source separation based on the convolutive transfer function","author":"Li","year":"2017"},{"key":"2026033012252967000_ref099","volume-title":"Disentangled sequential autoencoder","author":"Li","year":"2018"},{"key":"2026033012252967000_ref100","volume-title":"Recurrent switching linear dynamical systems","author":"Linderman","year":"2016"},{"key":"2026033012252967000_ref101","doi-asserted-by":"crossref","DOI":"10.1109\/IJCNN.2019.8852155","volume-title":"A Transformer-based variational autoencoder for sentence generation","author":"Liu","year":"2019"},{"key":"2026033012252967000_ref102","first-page":"1","article-title":"A sober look at the unsupervised learning of disentangled representations and their evaluation","volume":"21","author":"Locatello","year":"2020","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012252967000_ref103","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Lombardo","year":"2019"},{"key":"2026033012252967000_ref104","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Lucas","year":"2019"},{"key":"2026033012252967000_ref105","volume-title":"Auxiliary guided autoregressive variational autoencoders","author":"Lucas","year":"2018"},{"issue":"8","key":"2026033012252967000_ref106","doi-asserted-by":"crossref","first-page":"1256","DOI":"10.1109\/TASLP.2019.2915167","article-title":"Conv-Tasnet: Surpassing ideal time-frequency magnitude masking for speech separation","volume":"27","author":"Luo","year":"2019","journal-title":"IEEE\/ACM Transactions on Audio, Speech, and Language Processing"},{"key":"2026033012252967000_ref107","volume-title":"Auxiliary deep generative models","author":"Maal\u00f8e","year":"2016"},{"key":"2026033012252967000_ref108","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Maddison","year":"2017"},{"key":"2026033012252967000_ref109","volume-title":"History repeats itself: Human motion prediction via motion attention","author":"Mao","year":"2020"},{"key":"2026033012252967000_ref110","volume-title":"Iterative amortized inference","author":"Marino","year":"2018"},{"key":"2026033012252967000_ref111","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Marino","year":"2018"},{"key":"2026033012252967000_ref112","doi-asserted-by":"crossref","DOI":"10.1109\/CVPR.2017.497","volume-title":"On human motion prediction using recurrent neural networks","author":"Martinez","year":"2017"},{"key":"2026033012252967000_ref113","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Mathieu","year":"2016"},{"key":"2026033012252967000_ref114","volume-title":"Neural variational inference for text processing","author":"Miao","year":"2016"},{"key":"2026033012252967000_ref115","volume-title":"Disentangled state space representations","author":"Miladinovic","year":"2019"},{"key":"2026033012252967000_ref116","volume-title":"Neural variational inference and learning in belief networks","author":"Mnih","year":"2014"},{"key":"2026033012252967000_ref117","volume-title":"Kalman filter: Recent advances and applications","author":"Moreno","year":"2009"},{"key":"2026033012252967000_ref118","volume-title":"Switching Kalman filters","author":"Murphy","year":"1998"},{"key":"2026033012252967000_ref119","volume-title":"Machine learning: A probabilistic perspective","author":"Murphy","year":"2012"},{"key":"2026033012252967000_ref120","volume-title":"Variational sequential Monte Carlo","author":"Naesseth","year":"2018"},{"key":"2026033012252967000_ref121","doi-asserted-by":"crossref","first-page":"355","DOI":"10.1007\/978-94-011-5014-9_12","volume-title":"Learning in graphical models","author":"Neal","year":"1998"},{"issue":"4","key":"2026033012252967000_ref122","doi-asserted-by":"crossref","first-page":"1293","DOI":"10.1109\/18.243446","article-title":"Proper complex random processes with applications to information theory","volume":"39","author":"Neeser","year":"1993","journal-title":"IEEE Transactions on Information Theory"},{"key":"2026033012252967000_ref123","volume-title":"Signal analysis","author":"Papoulis","year":"1977"},{"key":"2026033012252967000_ref124","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2019-1398","volume-title":"A statistically principled and computationally efficient approach to speech enhancement using variational autoencoders","author":"Pariente","year":"2019"},{"key":"2026033012252967000_ref125","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Paszke","year":"2019"},{"key":"2026033012252967000_ref126","doi-asserted-by":"crossref","DOI":"10.1109\/ICMLA.2018.00207","volume-title":"Unsupervised anomaly detection in energy time series data using variational recurrent autoencoders with attention","author":"Pereira","year":"2018"},{"key":"2026033012252967000_ref127","doi-asserted-by":"crossref","DOI":"10.1109\/ICCV48922.2021.01080","volume-title":"Action-conditioned 3D human motion synthesis with Transformer VAE","author":"Petrovich","year":"2021"},{"issue":"16-18","key":"2026033012252967000_ref128","doi-asserted-by":"crossref","first-page":"3704","DOI":"10.1016\/j.neucom.2009.06.009","article-title":"Variational Bayesian learning of nonlinear hidden state-space models for model predictive control","volume":"72","author":"Raiko","year":"2009","journal-title":"Neurocomputing"},{"key":"2026033012252967000_ref129","volume-title":"Hierarchical variational models","author":"Ranganath","year":"2016"},{"key":"2026033012252967000_ref130","volume-title":"Preventing posterior collapse with delta-VAEs","author":"Razavi","year":"2019"},{"key":"2026033012252967000_ref131","volume-title":"Stochastic back-propagation and approximate inference in deep generative models","author":"Rezende","year":"2014"},{"key":"2026033012252967000_ref132","volume-title":"Variational inference with normalizing flows","author":"Rezende","year":"2015"},{"key":"2026033012252967000_ref133","volume-title":"Perceptual evaluation of speech quality (PESQ): A new method for speech quality assessment of telephone networks and codecs","author":"Rix","year":"2001"},{"key":"2026033012252967000_ref134","first-page":"400","article-title":"A stochastic approximation method","volume-title":"The Annals of Mathematical Statistics:","author":"Robbins","year":"1951"},{"key":"2026033012252967000_ref135","volume-title":"A hierarchical latent vector model for learning long-term structure in music","author":"Roberts","year":"2018"},{"key":"2026033012252967000_ref136","volume-title":"Autoencoders for music sound synthesis: A comparison of linear, shallow, deep and variational models","author":"Roche","year":"2019"},{"key":"2026033012252967000_ref137","volume-title":"A structured variational auto-encoder for learning deep hierarchies of sparse features","author":"Salimans","year":"2016"},{"key":"2026033012252967000_ref138","volume-title":"Markov chain Monte Carlo and variational inference: Bridging the gap","author":"Salimans","year":"2015"},{"issue":"4","key":"2026033012252967000_ref139","doi-asserted-by":"crossref","first-page":"837","DOI":"10.1214\/13-BA858","article-title":"Fixed-form variational posterior approximation through stochastic linear regression","volume":"8","author":"Salimans","year":"2013","journal-title":"Bayesian Analysis"},{"key":"2026033012252967000_ref140","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Saul","year":"1996"},{"key":"2026033012252967000_ref141","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/D17-1066","volume-title":"A hybrid convolutional variational autoencoder for text generation","author":"Semeniuta","year":"2017"},{"key":"2026033012252967000_ref142","volume-title":"Piecewise latent variables for neural variational text processing","author":"Serban","year":"2016"},{"key":"2026033012252967000_ref143","volume-title":"A hierarchical latent variable encoderdecoder model for generating dialogues","author":"Serban","year":"2017"},{"key":"2026033012252967000_ref144","doi-asserted-by":"crossref","DOI":"10.1109\/WACV.2018.00136","volume-title":"Channel-recurrent autoencoding for image modeling","author":"Shang","year":"2018"},{"key":"2026033012252967000_ref145","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Siddharth","year":"2017"},{"key":"2026033012252967000_ref146","volume-title":"The variational Bayes method in signal processing","author":"\u0160m\u00eddl","year":"2006"},{"key":"2026033012252967000_ref147","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Sohn","year":"2015"},{"key":"2026033012252967000_ref148","volume-title":"How to train deep variational autoencoders and probabilistic ladder networks","author":"S\u00f8nderby","year":"2016"},{"key":"2026033012252967000_ref149","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Sonderby","year":"2016"},{"key":"2026033012252967000_ref150","doi-asserted-by":"crossref","DOI":"10.1609\/aaai.v32i1.11985","volume-title":"Variational recurrent neural machine translation","author":"Su","year":"2018"},{"key":"2026033012252967000_ref151","volume-title":"Training recurrent neural networks","author":"Sutskever","year":"2013"},{"issue":"7","key":"2026033012252967000_ref152","doi-asserted-by":"crossref","first-page":"2125","DOI":"10.1109\/TASL.2011.2114881","article-title":"An algorithm for intelligibility prediction of time-frequency weighted noisy speech","volume":"19","author":"Taal","year":"2011","journal-title":"IEEE Transactions on Audio, Speech, and Language Processing"},{"issue":"3","key":"2026033012252967000_ref153","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1111\/1467-9868.00196","article-title":"Probabilistic principal component analysis","volume":"61","author":"Tipping","year":"1999","journal-title":"Journal of the Royal Statistical Society: Series B (Statistical Methodology)"},{"key":"2026033012252967000_ref154","volume-title":"MoCoGan: Decomposing motion and content for video generation","author":"Tulyakov","year":"2018"},{"key":"2026033012252967000_ref155","article-title":"NVAE: A deep hierarchical variational autoencoder","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Vahdat","year":"2020"},{"key":"2026033012252967000_ref156","volume-title":"Wavenet: A generative model for raw audio","author":"van den Oord","year":"2016"},{"key":"2026033012252967000_ref157","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"van den Oord","year":"2016"},{"key":"2026033012252967000_ref158","article-title":"Pixel recurrent neural networks","author":"van den Oord","year":"2016"},{"key":"2026033012252967000_ref159","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"van den Oord","year":"2017"},{"key":"2026033012252967000_ref160","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Vaswani","year":"2017"},{"key":"2026033012252967000_ref161","volume-title":"Decomposing motion and content for natural video sequence prediction","author":"Villegas","year":"2017"},{"key":"2026033012252967000_ref162","first-page":"3371","article-title":"Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion","volume":"11","author":"Vincent","year":"2010","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012252967000_ref163","doi-asserted-by":"crossref","DOI":"10.1109\/ASSPCC.2000.882463","volume-title":"The unscented Kalman filter for nonlinear estimation","author":"Wan","year":"2000"},{"key":"2026033012252967000_ref164","volume-title":"T-CVAE: Transformer-based conditioned variational autoencoder for story completion","author":"Wang","year":"2019"},{"key":"2026033012252967000_ref165","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Watter","year":"2015"},{"issue":"411","key":"2026033012252967000_ref166","doi-asserted-by":"crossref","first-page":"699","DOI":"10.1080\/01621459.1990.10474930","article-title":"A Monte Carlo implementation of the EM algorithm and the poor man\u2019s data augmentation algorithms","volume":"85","author":"Wei","year":"1990","journal-title":"Journal of the American Statistical Association"},{"key":"2026033012252967000_ref167","volume-title":"Gaussian processes for machine learning","author":"Williams","year":"2006"},{"issue":"2","key":"2026033012252967000_ref168","doi-asserted-by":"crossref","first-page":"270","DOI":"10.1162\/neco.1989.1.2.270","article-title":"A learning algorithm for continually running fully recurrent neural networks","volume":"1","author":"Williams","year":"1989","journal-title":"Neural Computation"},{"key":"2026033012252967000_ref169","first-page":"661","article-title":"Variational message passing","volume":"6","author":"Winn","year":"2005","journal-title":"Journal of Machine Learning Research"},{"key":"2026033012252967000_ref170","doi-asserted-by":"crossref","DOI":"10.1109\/ICASSP40776.2020.9054074","volume-title":"Feedback recurrent autoencoder","author":"Yang","year":"2020"},{"key":"2026033012252967000_ref171","volume-title":"Improved variational autoencoders for text modeling using dilated convolutions","author":"Yang","year":"2017"},{"key":"2026033012252967000_ref172","volume-title":"Tackling over-pruning in variational autoencoders","author":"Yeung","year":"2017"},{"key":"2026033012252967000_ref173","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/P18-1101","volume-title":"Unsupervised discrete sentence representation learning for interpretable neural dialog generation","author":"Zhao","year":"2018"},{"key":"2026033012252967000_ref174","doi-asserted-by":"crossref","DOI":"10.18653\/v1\/P17-1061","volume-title":"Learning discourse-level diversity for neural dialog models using conditional variational autoencoders","author":"Zhao","year":"2017"},{"key":"2026033012252967000_ref175","volume-title":"Streaming variational Monte Carlo","author":"Zhao","year":"2019"},{"key":"2026033012252967000_ref176","volume-title":"Variational online learning of neural dynamics","author":"Zhao","year":"2017"}],"container-title":["Foundations and Trends\u00ae in Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftmal\/article-pdf\/15\/1-2\/1\/11134975\/2200000089en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftmal\/article-pdf\/15\/1-2\/1\/11134975\/2200000089en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T18:10:54Z","timestamp":1777486254000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftmal\/article\/15\/1-2\/1\/1331294\/Dynamical-Variational-Autoencoders-A-Comprehensive"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,1]]},"references-count":176,"journal-issue":{"issue":"1-2","published-print":{"date-parts":[[2022,1,1]]}},"URL":"https:\/\/doi.org\/10.1561\/2200000089","relation":{},"ISSN":["1935-8237","1935-8245"],"issn-type":[{"value":"1935-8237","type":"print"},{"value":"1935-8245","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,1,1]]}}}