{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,27]],"date-time":"2026-06-27T08:05:32Z","timestamp":1782547532488,"version":"3.54.5"},"reference-count":161,"publisher":"Emerald","issue":"3-4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018,8,14]]},"abstract":"<jats:p>This monograph aims at providing an introduction to key concepts, algorithms, and theoretical results in machine learning. The treatment concentrates on probabilistic models for supervised and unsupervised learning problems. It introduces fundamental concepts and algorithms by building on first principles, while also exposing the reader to more advanced topics with extensive pointers to the literature, within a unified notation and mathematical framework. The material is organized according to clearly defined categories, such as discriminative and generative models, frequentist and Bayesian approaches, exact and approximate inference, as well as directed and undirected models. This monograph is meant as an entry point for researchers with an engineering background in probability and linear algebra.<\/jats:p>","DOI":"10.1561\/2000000102","type":"journal-article","created":{"date-parts":[[2018,8,29]],"date-time":"2018-08-29T06:37:23Z","timestamp":1535524643000},"page":"200-431","source":"Crossref","is-referenced-by-count":83,"title":["A Brief Introduction to Machine Learning for Engineers"],"prefix":"10.1108","volume":"12","author":[{"given":"Osvaldo","family":"Simeone","sequence":"first","affiliation":[{"name":"Department of Informatics,King\u2019s College London"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2018,8,14]]},"reference":[{"key":"2026040313220152500_ref001","article-title":"On the Protection of Private Information in Machine Learning Systems: Two Recent Approaches","author":"Abadi","year":"2017","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref002","volume":"4","author":"Abu-Mostafa","year":"2012","journal-title":"Learning from data"},{"key":"2026040313220152500_ref003","volume-title":"Variational Information Maximization in Stochastic Environments (PhD thesis)","author":"Agakov","year":"2005"},{"key":"2026040313220152500_ref004","article-title":"An Information-Theoretic Analysis of Deep Latent-Variable Models","author":"Alemi","year":"2017","journal-title":"ArXiv e-prints"},{"issue":"2","key":"2026040313220152500_ref005","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1162\/089976698300017746","article-title":"Natural gradient works efficiently in learning","volume":"10","author":"Amari","year":"1998","journal-title":"Neural computation"},{"key":"2026040313220152500_ref006","doi-asserted-by":"crossref","DOI":"10.1007\/978-4-431-55978-8","volume-title":"Information geometry and its applications","author":"Amari","year":"2016"},{"issue":"2-3","key":"2026040313220152500_ref007","doi-asserted-by":"crossref","first-page":"119","DOI":"10.1561\/2200000052","article-title":"Patterns of scalable Bayesian inference","volume":"9","author":"Angelino","year":"2016","journal-title":"Foundations and Trends R in Machine Learning"},{"key":"2026040313220152500_ref008","article-title":"Wasserstein GAN","author":"Arjovsky","year":"2017","journal-title":"arXiv preprint arXiv:1701.07875"},{"issue":"6","key":"2026040313220152500_ref009","doi-asserted-by":"publisher","first-page":"26","DOI":"10.1109\/MSP.2017.2743240","article-title":"Deep Reinforcement Learning: A Brief Survey","volume":"34","author":"Arulkumaran","year":"2017","journal-title":"IEEE Signal Processing Magazine"},{"issue":"3","key":"2026040313220152500_ref010","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1023\/A:1010896012157","article-title":"Relative loss bounds for on-line density estimation with the exponential family of distributions","volume":"43","author":"Azoury","year":"2001","journal-title":"Machine Learning"},{"key":"2026040313220152500_ref011","article-title":"Training Probabilistic Spiking Neural Networks with First-to-spike Decoding","author":"Bagheri","year":"2017","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref012","article-title":"Learning in the machine: Random backpropagation and the learning channel","author":"Baldi","year":"2016","journal-title":"arXiv preprint arXiv:1612.02734"},{"key":"2026040313220152500_ref013","article-title":"Perturbative Black Box Variational Inference","author":"Bamler","year":"2017","journal-title":"ArXiv e-prints"},{"issue":"4","key":"2026040313220152500_ref014","doi-asserted-by":"crossref","first-page":"118","DOI":"10.1109\/MSP.2007.4286571","article-title":"Compressive sensing [lecture notes]","volume":"24","author":"Baraniuk","year":"2007","journal-title":"IEEE signal processing magazine"},{"key":"2026040313220152500_ref015","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511804779","volume-title":"Bayesian reasoning and machine learning","author":"Barber","year":"2012"},{"key":"2026040313220152500_ref016","volume-title":"Variational algorithms for approximate Bayesian inference","author":"Beal","year":"2003"},{"key":"2026040313220152500_ref017","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9781139042918","volume-title":"Scaling up machine learning: Parallel and distributed approaches","author":"Bekkerman","year":"2011"},{"key":"2026040313220152500_ref018","article-title":"MINE: Mutual Information Neural Estimation","author":"Belghazi","year":"2018","journal-title":"arXiv preprint arXiv:1801.04062"},{"key":"2026040313220152500_ref019","first-page":"17","article-title":"Deep learning of representations for unsupervised and transfer learning","author":"Bengio","year":"2012","journal-title":"Proceedings of ICML Workshop on Unsupervised and Transfer Learning"},{"issue":"8","key":"2026040313220152500_ref020","doi-asserted-by":"crossref","first-page":"1798","DOI":"10.1109\/TPAMI.2013.50","article-title":"Representation learning: A review and new perspectives","volume":"35","author":"Bengio","year":"2013","journal-title":"IEEE transactions on pattern analysis and machine intelligence"},{"issue":"3","key":"2026040313220152500_ref021","doi-asserted-by":"crossref","first-page":"580","DOI":"10.1109\/TSP.2015.2477805","article-title":"Empirically estimable classification bounds based on a nonparametric divergence measure","volume":"64","author":"Berisha","year":"2016","journal-title":"IEEE Transactions on Signal Processing"},{"issue":"1-38","key":"2026040313220152500_ref022","first-page":"3","article-title":"Incremental gradient, subgradient, and proximal methods for convex optimization: A survey","volume":"2010","author":"Bertsekas","year":"2011","journal-title":"Optimization for Machine Learning"},{"key":"2026040313220152500_ref023","volume-title":"Pattern recognition and machine learning","author":"Bishop","year":"2006"},{"key":"2026040313220152500_ref024","doi-asserted-by":"crossref","DOI":"10.1080\/01621459.2017.1285773","article-title":"Variational inference: A review for statisticians","author":"Blei","year":"2017","journal-title":"Journal of the American Statistical Association"},{"key":"2026040313220152500_ref025","article-title":"Variational Inference: Foundations and Modern Methods","author":"Blei"},{"key":"2026040313220152500_ref026","first-page":"129","article-title":"Learning bounds for domain adaptation","author":"Blitzer","year":"2008","journal-title":"Advances in neural information processing systems"},{"key":"2026040313220152500_ref027","article-title":"Weight uncertainty in neural networks","author":"Blundell","year":"2015","journal-title":"arXiv preprint arXiv:1505.05424"},{"key":"2026040313220152500_ref028","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511804441","volume-title":"Convex optimization","author":"Boyd","year":"2004"},{"key":"2026040313220152500_ref029","article-title":"Learning Independent Features with Adversarial Nets for Non-linear ICA","author":"Brakel","year":"2017","journal-title":"ArXiv e-prints"},{"issue":"4","key":"2026040313220152500_ref030","doi-asserted-by":"crossref","first-page":"18","DOI":"10.1109\/MSP.2017.2693418","article-title":"Geometric deep learning: going beyond euclidean data","volume":"34","author":"Bronstein","year":"2017","journal-title":"IEEE Signal Processing Magazine"},{"issue":"6370","key":"2026040313220152500_ref031","doi-asserted-by":"crossref","first-page":"1530","DOI":"10.1126\/science.aap8062","article-title":"What can machine learning do? Workforce implications","volume":"358","author":"Brynjolfsson","year":"2017","journal-title":"Science"},{"key":"2026040313220152500_ref032","article-title":"Importance weighted autoencoders","author":"Burda","year":"2015","journal-title":"arXiv preprint arXiv:1509.00519"},{"issue":"5","key":"2026040313220152500_ref033","doi-asserted-by":"crossref","first-page":"32","DOI":"10.1109\/MSP.2014.2329397","article-title":"Convex optimization for big data: Scalable, randomized, and parallel algorithms for big data analytics","volume":"31","author":"Cevher","year":"2014","journal-title":"IEEE Signal Processing Magazine"},{"key":"2026040313220152500_ref034","first-page":"1002","article-title":"In Defense of Probability.","volume":"85","author":"Cheeseman","year":"1985","journal-title":"IJCAI"},{"issue":"2","key":"2026040313220152500_ref035","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1109\/MSP.2013.2297439","article-title":"Tensor decompositions for signal processing applications: From two-way to multiway component analysis","volume":"32","author":"Cichocki","year":"2015","journal-title":"IEEE Signal Processing Magazine"},{"key":"2026040313220152500_ref036","first-page":"617","article-title":"A generalization of principal components analysis to the exponential family","author":"Collins","year":"2002","journal-title":"Advances in neural information processing systems"},{"issue":"3","key":"2026040313220152500_ref037","doi-asserted-by":"crossref","first-page":"273","DOI":"10.1023\/A:1022627411411","article-title":"Support-vector networks","volume":"20","author":"Cortes","year":"1995","journal-title":"Machine learning"},{"key":"2026040313220152500_ref038","volume-title":"Elements of information theory","author":"Cover","year":"2012"},{"key":"2026040313220152500_ref039","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511801389","volume-title":"An introduction to support vector machines and other kernel-based learning methods","author":"Cristianini","year":"2000"},{"issue":"4","key":"2026040313220152500_ref040","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1561\/0100000004","article-title":"Information theory and statistics: A tutorial","volume":"1","author":"Csisz\u00e1r","year":"2004","journal-title":"Foundations and Trends R in Communications and Information Theory"},{"key":"2026040313220152500_ref041","article-title":"Probabilistic Programming & Bayesian Methods for Hackers","author":"Davidson-Pilon","year":"2015"},{"issue":"5","key":"2026040313220152500_ref042","doi-asserted-by":"crossref","first-page":"889","DOI":"10.1162\/neco.1995.7.5.889","article-title":"The helmholtz machine","volume":"7","author":"Dayan","year":"1995","journal-title":"Neural computation"},{"key":"2026040313220152500_ref043","article-title":"Variance Reduction for Distributed Stochastic Gradient Descent","author":"De","year":"2015","journal-title":"arXiv preprint arXiv:1512.01708"},{"issue":"2","key":"2026040313220152500_ref044","doi-asserted-by":"crossref","first-page":"120","DOI":"10.1109\/TSIPN.2016.2524588","article-title":"Next: In-network nonconvex optimization","volume":"2","author":"Di Lorenzo","year":"2016","journal-title":"IEEE Transactions on Signal and Information Processing over Networks"},{"key":"2026040313220152500_ref045","article-title":"Lecture Notes for Statistics 311\/Electrical Engineering 377","author":"Duchi","year":"2016"},{"key":"2026040313220152500_ref046","article-title":"Information Measures, Experiments, Multi-category Hypothesis Tests, and Surrogate Losses","author":"Duchi","year":"2016","journal-title":"arXiv preprint arXiv:1603.00126"},{"key":"2026040313220152500_ref047","article-title":"Adversarially learned inference","author":"Dumoulin","year":"2016","journal-title":"arXiv preprint arXiv:1606.00704"},{"key":"2026040313220152500_ref048","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9781316576533","volume-title":"Computer Age Statistical Inference","author":"Efron","year":"2016"},{"key":"2026040313220152500_ref049","article-title":"Many Paths to Equilibrium: GANs Do Not Need to Decrease a Divergence At Every Step","author":"Fedus","year":"2017","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref050","article-title":"Learning Anonymized Representations with Adversarial Neural Networks","author":"Feutry","year":"2018","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref051","volume-title":"The elements of statistical learning","author":"Friedman","year":"2001"},{"key":"2026040313220152500_ref052","article-title":"Recent advances in zero-shot recognition","author":"Fu","year":"2017","journal-title":"arXiv preprint arXiv:1710.04837"},{"key":"2026040313220152500_ref053","article-title":"Uncertainty in Deep Learning","author":"Gal","year":"2016","journal-title":"PhD thesis"},{"key":"2026040313220152500_ref054","volume":"159","author":"Gersho","year":"2012","journal-title":"Vector quantization and signal compression"},{"key":"2026040313220152500_ref055","article-title":"Explaining and harnessing adversarial examples","author":"Goodfellow","year":"2014","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref056","volume-title":"Deep learning","author":"Goodfellow","year":"2016"},{"key":"2026040313220152500_ref057","first-page":"2672","article-title":"Generative adversarial nets","author":"Goodfellow","year":"2014","journal-title":"Advances in neural information processing systems"},{"key":"2026040313220152500_ref058","article-title":"cvx users\u2019 guide","author":"Grant","year":"2009"},{"key":"2026040313220152500_ref059","article-title":"Backpropagation through the Void: Optimizing control variates for black-box gradient estimation","author":"Grathwohl","year":"2017","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref060","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/4643.001.0001","volume-title":"The minimum description length principle","author":"Grunwald","year":"2007"},{"issue":"4","key":"2026040313220152500_ref061","doi-asserted-by":"crossref","first-page":"1367","DOI":"10.1214\/009053604000000553","article-title":"Game theory, maximum entropy, minimum discrepancy and robust Bayesian decision theory","volume":"32","author":"Gr\u00fcnwald","year":"2004","journal-title":"The Annals of Statistics"},{"issue":"4","key":"2026040313220152500_ref062","doi-asserted-by":"crossref","first-page":"784","DOI":"10.1109\/TKDE.2003.1208999","article-title":"Topic-sensitive pagerank: A context-sensitive ranking algorithm for web search","volume":"15","author":"Haveliwala","year":"2003","journal-title":"IEEE transactions on knowledge and data engineering"},{"key":"2026040313220152500_ref063","article-title":"Neural Networks for Machine Learning (online course)","author":"Hinton","year":"2016"},{"issue":"5214","key":"2026040313220152500_ref064","doi-asserted-by":"crossref","first-page":"1158","DOI":"10.1126\/science.7761831","article-title":"The \u201cwake-sleep\u201d algorithm for unsupervised neural networks","volume":"268","author":"Hinton","year":"1995","journal-title":"Science"},{"issue":"1","key":"2026040313220152500_ref065","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1162\/neco.1997.9.1.1","article-title":"Flat minima","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Computation"},{"issue":"1","key":"2026040313220152500_ref066","doi-asserted-by":"crossref","first-page":"489","DOI":"10.1016\/j.neucom.2005.12.126","article-title":"Extreme learning machine: theory and applications","volume":"70","author":"Huang","year":"2006","journal-title":"Neurocomputing"},{"key":"2026040313220152500_ref067","unstructured":"Husz\u00e1r, F.\n           \u201cEverything that Works Works Because it\u2019s Bayesian: Why Deep Nets Generalize?\u201d url: http:\/\/www.inference.vc\/."},{"key":"2026040313220152500_ref068","unstructured":"Husz\u00e1r, F.\n          \n          2017a. \u201cChoice of Recognition Models in VAEs: a regularisation view\u201d. url: http:\/\/www.inference.vc\/."},{"key":"2026040313220152500_ref069","unstructured":"Husz\u00e1r, F.\n          \n          2017b. \u201cIs Maximum Likelihood Useful for Representation Learning?\u201d url: http:\/\/www.inference.vc\/."},{"key":"2026040313220152500_ref070","article-title":"Variational Inference using Implicit Distributions","author":"Husz\u00e1r","year":"2017","journal-title":"arXiv preprint arXiv:1702.08235"},{"key":"2026040313220152500_ref071","doi-asserted-by":"crossref","unstructured":"Jain, P. and P.Kar. 2017. \u201cNon-convex Optimization for Machine Learning\u201d. Foundations and Trends R in Machine Learning. 10(3-4): 142\u2013336. issn: 1935-8237. url: 10.1561\/2200000058.","DOI":"10.1561\/2200000058"},{"key":"2026040313220152500_ref072","article-title":"Categorical reparameterization with gumbel-softmax","author":"Jang","year":"2016","journal-title":"arXiv preprint arXiv:1611.01144"},{"issue":"12","key":"2026040313220152500_ref073","doi-asserted-by":"crossref","first-page":"7616","DOI":"10.1109\/TIT.2014.2360184","article-title":"Information measures: the curious case of the binary alphabet","volume":"60","author":"Jiao","year":"2014","journal-title":"IEEE Transactions on Information Theory"},{"issue":"10","key":"2026040313220152500_ref074","doi-asserted-by":"crossref","first-page":"5357","DOI":"10.1109\/TIT.2015.2462848","article-title":"Justification of logarithmic loss via the benefit of side information","volume":"61","author":"Jiao","year":"2015","journal-title":"IEEE Transactions on Information Theory"},{"key":"2026040313220152500_ref075","first-page":"315","article-title":"Accelerating stochastic gradient descent using predictive variance reduction","author":"Johnson","year":"2013","journal-title":"Advances in neural information processing systems"},{"key":"2026040313220152500_ref076","unstructured":"Karpathy, A.\n           \u201cDeep Reinforcement Learning: Pong from Pixels\u201d. url: http:\/\/karpathy.github.io\/2016\/05\/31\/rl\/."},{"key":"2026040313220152500_ref077","article-title":"Generalization in Deep Learning","author":"Kawaguchi","year":"2017","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref078","article-title":"On large-batch training for deep learning: Generalization gap and sharp minima","author":"Keskar","year":"2016","journal-title":"arXiv preprint arXiv:1609.04836"},{"key":"2026040313220152500_ref079","article-title":"Auto-encoding variational bayes","author":"Kingma","year":"2013","journal-title":"arXiv preprint arXiv:1312.6114"},{"key":"2026040313220152500_ref080","volume-title":"Probabilistic graphical models: principles and techniques","author":"Koller","year":"2009"},{"key":"2026040313220152500_ref081","article-title":"Benchmarking quantum hardware for training of fully visible boltzmann machines","author":"Korenkevych","year":"2016","journal-title":"arXiv preprint arXiv:1611.04528"},{"key":"2026040313220152500_ref082","article-title":"A tutorial on energy-based learning","volume":"1","author":"LeCun","year":"2006","journal-title":"Predicting structured data"},{"key":"2026040313220152500_ref083","doi-asserted-by":"crossref","DOI":"10.3389\/fnins.2016.00508","article-title":"Training deep spiking neural networks using backpropagation","volume":"10","author":"Lee","year":"2016","journal-title":"Frontiers in neuroscience"},{"issue":"2","key":"2026040313220152500_ref084","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1162\/089976699300016719","article-title":"Independent component analysis using an extended infomax algorithm for mixed subgaussian and supergaussian sources","volume":"11","author":"Lee","year":"1999","journal-title":"Neural computation"},{"key":"2026040313220152500_ref085","volume-title":"Theory of point estimation","author":"Lehmann","year":"2006"},{"key":"2026040313220152500_ref086","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/10909.001.0001","volume-title":"Common Sense, the Turing Test, and the Quest for Real AI","author":"Levesque","year":"2017"},{"key":"2026040313220152500_ref087","unstructured":"Levin, S.\n          \n          2016. A beauty contest was judged by AI and the robots didn\u2019t like dark skin. url: https:\/\/www.theguardian.com\/technology\/2016\/sep\/08\/artificial-intelligence-beauty-contestdoesnt-like-black-people."},{"key":"2026040313220152500_ref088","unstructured":"Levine, S.\n          \n          2017. Deep Reinforcement Learning. url: http:\/\/rll.berkeley.edu\/deeprlcourse\/#lecture-videos."},{"key":"2026040313220152500_ref089","unstructured":"Li, Y.\n           \u201cTopics in Approximate Inference\u201d. url: http:\/\/yingzhenli.net\/home\/pdf\/topics_approx_infer.pdf."},{"key":"2026040313220152500_ref090","first-page":"1073","article-title":"R\u00e9nyi divergence variational inference","author":"Li","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"11","key":"2026040313220152500_ref091","doi-asserted-by":"crossref","first-page":"617","DOI":"10.3390\/e19110617","article-title":"On Lower Bounds for Statistical Learning Theory","volume":"19","author":"Loh","year":"2017","journal-title":"Entropy"},{"key":"2026040313220152500_ref092","doi-asserted-by":"crossref","DOI":"10.1201\/b13613","volume-title":"The BUGS book: A practical introduction to Bayesian analysis","author":"Lunn","year":"2012"},{"key":"2026040313220152500_ref093","first-page":"2579","article-title":"Visualizing data using t-SNE","author":"Maaten","year":"2008","journal-title":"Journal of Machine Learning Research"},{"key":"2026040313220152500_ref094","volume-title":"Information theory, inference and learning algorithms","author":"MacKay","year":"2003"},{"key":"2026040313220152500_ref095","article-title":"The concrete distribution: A continuous relaxation of discrete random variables","author":"Maddison","year":"2016","journal-title":"arXiv preprint arXiv:1611.00712"},{"key":"2026040313220152500_ref096","article-title":"Divergence measures and message passing","author":"Minka","year":"2005","journal-title":"Tech. rep."},{"key":"2026040313220152500_ref097","article-title":"Perceptrons.","author":"Minsky","year":"1969"},{"key":"2026040313220152500_ref098","article-title":"Neural variational inference and learning in belief networks","author":"Mnih","year":"2014","journal-title":"arXiv preprint arXiv:1402.0030"},{"key":"2026040313220152500_ref099","article-title":"Playing atari with deep reinforcement learning","author":"Mnih","year":"2013","journal-title":"arXiv preprint arXiv:1312.5602"},{"key":"2026040313220152500_ref100","article-title":"Learning in implicit generative models","author":"Mohamed","year":"2016","journal-title":"arXiv preprint arXiv:1610.03483"},{"key":"2026040313220152500_ref101","article-title":"First-Order Adaptive Sample Size Methods to Reduce Complexity of Empirical Risk Minimization","author":"Mokhtari","year":"2017","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref102","article-title":"Methods for interpreting and understanding deep neural networks","author":"Montavon","year":"2017","journal-title":"arXiv preprint arXiv:1706.07979"},{"issue":"7676","key":"2026040313220152500_ref103","doi-asserted-by":"crossref","first-page":"375","DOI":"10.1038\/nature24047","article-title":"Solving a Higgs optimization problem with quantum annealing for machine learning","volume":"550","author":"Mott","year":"2017","journal-title":"Nature"},{"key":"2026040313220152500_ref104","volume-title":"Machine learning: a probabilistic perspective","author":"Murphy","year":"2012"},{"issue":"11","key":"2026040313220152500_ref105","doi-asserted-by":"crossref","first-page":"5847","DOI":"10.1109\/TIT.2010.2068870","article-title":"Estimating divergence functionals and the likelihood ratio by convex risk minimization","volume":"56","author":"Nguyen","year":"2010","journal-title":"IEEE Transactions on Information Theory"},{"key":"2026040313220152500_ref106","article-title":"Chernoff information of exponential families","author":"Nielsen","year":"2011","journal-title":"arXiv preprint arXiv:1102.2684"},{"key":"2026040313220152500_ref107","first-page":"271","article-title":"f-GAN: Training generative neural samplers using variational divergence minimization","author":"Nowozin","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026040313220152500_ref108","article-title":"Conditional image synthesis with auxiliary classifier gans","author":"Odena","year":"2016","journal-title":"arXiv preprint arXiv:1610.09585"},{"key":"2026040313220152500_ref109","volume-title":"Weapons of Math Destruction","author":"O\u2019Neil","year":"2016"},{"key":"2026040313220152500_ref110","article-title":"The PageRank citation ranking: Bringing order to the web.","author":"Page","year":"1999","journal-title":"Tech. rep."},{"key":"2026040313220152500_ref111","article-title":"Practical Black-Box Attacks against Machine Learning","author":"Papernot","year":"2016","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref112","article-title":"Theoretical Impediments to Machine Learning With Seven Sparks from the Causal Revolution","author":"Pearl","year":"2018","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref113","volume-title":"Causal inference in statistics: a primer","author":"Pearl","year":"2016"},{"issue":"2","key":"2026040313220152500_ref114","doi-asserted-by":"crossref","first-page":"224","DOI":"10.1109\/JSTSP.2015.2496908","article-title":"A survey of stochastic simulation and optimization methods in signal processing","volume":"10","author":"Pereyra","year":"2016","journal-title":"IEEE Journal of Selected Topics in Signal Processing"},{"key":"2026040313220152500_ref115","volume-title":"Elements of Causal Inference: Foundations and Learning Algorithms","author":"Peters","year":"2017"},{"key":"2026040313220152500_ref116","volume-title":"How the Mind Works","author":"Pinker","year":"1997"},{"issue":"1","key":"2026040313220152500_ref117","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/MASSP.1986.1165342","article-title":"An introduction to hidden Markov models","volume":"3","author":"Rabiner","year":"1986","journal-title":"IEEE ASSP magazine"},{"key":"2026040313220152500_ref118","doi-asserted-by":"crossref","first-page":"958","DOI":"10.1109\/Allerton.2011.6120270","article-title":"Directed information and Pearl\u2019s causal calculus","author":"Raginsky","year":"2011","journal-title":"Communication, Control, and Computing (Allerton), 2011 49th Annual Allerton Conference on"},{"key":"2026040313220152500_ref119","doi-asserted-by":"crossref","first-page":"26","DOI":"10.1109\/ITW.2016.7606789","article-title":"Information-theoretic analysis of stability and bias of learning algorithms","author":"Raginsky","year":"2016","journal-title":"Information Theory Workshop (ITW), 2016 IEEE"},{"key":"2026040313220152500_ref120","first-page":"814","article-title":"Black box variational inference","author":"Ranganath","year":"2014","journal-title":"Artificial Intelligence and Statistics"},{"key":"2026040313220152500_ref121","first-page":"762","article-title":"Deep exponential families","author":"Ranganath","year":"2015","journal-title":"Artificial Intelligence and Statistics"},{"key":"2026040313220152500_ref122","article-title":"Stochastic backpropagation and approximate inference in deep generative models","author":"Rezende","year":"2014","journal-title":"arXiv preprint arXiv:1401.4082"},{"key":"2026040313220152500_ref123","article-title":"Stabilizing Training of Generative Adversarial Networks through Regularization","author":"Roth","year":"2017","journal-title":"arXiv preprint arXiv:1705.09367"},{"key":"2026040313220152500_ref124","article-title":"Structured Embedding Models for Grouped Data","author":"Rudolph","year":"2017","journal-title":"ArXiv e-prints"},{"issue":"3","key":"2026040313220152500_ref125","first-page":"1","article-title":"Learning representations by back-propagating errors","volume":"5","author":"Rumelhart","year":"1988","journal-title":"Cognitive modeling"},{"key":"2026040313220152500_ref126","volume-title":"Artificial Intelligence: A Modern Approach","author":"Russel","year":"2009"},{"key":"2026040313220152500_ref127","doi-asserted-by":"crossref","first-page":"791","DOI":"10.1145\/1273496.1273596","article-title":"Restricted Boltzmann machines for collaborative filtering","author":"Salakhutdinov","year":"2007","journal-title":"Proceedings of the 24th international conference on Machine learning"},{"key":"2026040313220152500_ref128","article-title":"Evolution strategies as a scalable alternative to reinforcement learning","author":"Salimans","year":"2017","journal-title":"arXiv preprint arXiv:1703.03864"},{"issue":"3","key":"2026040313220152500_ref129","doi-asserted-by":"crossref","first-page":"578","DOI":"10.1162\/NECO_a_00929","article-title":"Deep Learning with Dynamic Spiking Neurons and Fixed Feedback Weights","volume":"29","author":"Samadi","year":"2017","journal-title":"Neural Computation"},{"key":"2026040313220152500_ref130","article-title":"Distributed methods for constrained nonconvex multi-agent optimization-part I: theory","author":"Scutari","year":"2014"},{"key":"2026040313220152500_ref131","article-title":"Bayesian Dirichlet Bayesian Network Scores and the Maximum Entropy Principle","author":"Scutari","year":"2017","journal-title":"ArXiv e-prints"},{"issue":"1","key":"2026040313220152500_ref132","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1109\/JPROC.2015.2494218","article-title":"Taking the human out of the loop: A review of bayesian optimization","volume":"104","author":"Shahriari","year":"2016","journal-title":"Proceedings of the IEEE"},{"key":"2026040313220152500_ref133","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9781107298019","volume-title":"Understanding machine learning: From theory to algorithms","author":"Shalev-Shwartz","year":"2014"},{"issue":"3","key":"2026040313220152500_ref134","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","article-title":"A mathematical theory of communication","volume":"27","author":"Shannon","year":"1948","journal-title":"The Bell System Technical Journal"},{"key":"2026040313220152500_ref135","unstructured":"Silver, D.\n          \n          2015. Course on reinforcement learning. url: http:\/\/www0.cs.ucl.ac.uk\/staff\/d.silver\/web\/Teaching.html."},{"key":"2026040313220152500_ref136","article-title":"Don\u2019t Decay the Learning Rate, Increase the Batch Size","author":"Smith","year":"2017","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref137","unstructured":"Spectrum, I.\n          \n          Will the Future of AI Learning Depend More on Nature or Nurture? url: https:\/\/spectrum.ieee.org\/tech-talk\/robotics\/artificial-intelligence\/ai-and-psychology-researchersdebate-the-future-of-deep-learning."},{"key":"2026040313220152500_ref138","doi-asserted-by":"crossref","DOI":"10.4159\/9780674970199","volume-title":"The seven pillars of statistical wisdom","author":"Stigler","year":"2016"},{"key":"2026040313220152500_ref139","first-page":"187","article-title":"Online outlier detection in sensor data using non-parametric models","author":"Subramaniam","year":"2006","journal-title":"Proceedings of the 32nd international conference on Very large data bases"},{"key":"2026040313220152500_ref140","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9781139035613","volume-title":"Density ratio estimation in machine learning","author":"Sugiyama","year":"2012"},{"issue":"3","key":"2026040313220152500_ref141","doi-asserted-by":"crossref","first-page":"794","DOI":"10.1109\/TSP.2016.2601299","article-title":"Majorizationminimization algorithms in signal processing, communications, and machine learning","volume":"65","author":"Sun","year":"2017","journal-title":"IEEE Transactions on Signal Processing"},{"key":"2026040313220152500_ref142","volume-title":"Life 3.0: Being Human in the Age of Artificial Intelligence","author":"Tegmark","year":"2017"},{"key":"2026040313220152500_ref143","first-page":"640","article-title":"Is learning the n-th thing any easier than learning the first?","author":"Thrun","year":"1996","journal-title":"Advances in neural information processing systems"},{"key":"2026040313220152500_ref144","unstructured":"Times, T. N. Y.\n          \n          1958. NEW NAVY DEVICE LEARNS BY DOING; Psychologist Shows Embryo of Computer Designed to Read and Grow Wiser. url: http:\/\/www.nytimes.com\/1958\/07\/08\/archives\/new-navy-device-learns-by-doing-psychologistshows-embryo-of.html."},{"key":"2026040313220152500_ref145","article-title":"The information bottleneck method","author":"Tishby","year":"2000","journal-title":"arXiv preprint physics\/0004057"},{"key":"2026040313220152500_ref146","doi-asserted-by":"crossref","DOI":"10.1007\/b13794","article-title":"Introduction to nonparametric estimation","author":"Tsybakov","year":"2009"},{"key":"2026040313220152500_ref147","first-page":"115","article-title":"Two problems with variational expectation maximisation for time-series models","author":"Turner","year":"2011","journal-title":"Bayesian Time series models"},{"key":"2026040313220152500_ref148","unstructured":"Uber\n          . Pyro: Deep universal probabilistic programming. url: http:\/\/pyro.ai\/."},{"issue":"6","key":"2026040313220152500_ref149","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1109\/MSP.2017.2740460","article-title":"Deep-Learning Systems for Domain Adaptation in Computer Vision: Learning Transferable Feature Representations","volume":"34","author":"Venkateswara","year":"2017","journal-title":"IEEE Signal Processing Magazine"},{"key":"2026040313220152500_ref150","first-page":"3371","article-title":"Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion","author":"Vincent","year":"2010","journal-title":"Journal of Machine Learning Research"},{"issue":"1\u20132","key":"2026040313220152500_ref151","first-page":"1","article-title":"Graphical models, exponential families, and variational inference","volume":"1","author":"Wainwright","year":"2008","journal-title":"Foundations and Trends R in Machine Learning"},{"key":"2026040313220152500_ref152","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9781316402276","volume-title":"Machine Learning Refined: Foundations, Algorithms, and Applications","author":"Watt","year":"2016"},{"key":"2026040313220152500_ref153","first-page":"1481","article-title":"Exponential family harmoniums with an application to information retrieval","author":"Welling","year":"2005","journal-title":"Advances in neural information processing systems"},{"key":"2026040313220152500_ref154","unstructured":"Wikipedia\n          . AI Winter. url: https:\/\/en.wikipedia.org\/wiki\/AI_winter."},{"key":"2026040313220152500_ref155","unstructured":"Wikipedia\n          . Conjugate priors. url: https:\/\/en.wikipedia.org\/wiki\/Conjugate_prior."},{"key":"2026040313220152500_ref156","unstructured":"Wikipedia\n          . Exponential family. url: https:\/\/en.wikipedia.org\/wiki\/Exponential_family."},{"key":"2026040313220152500_ref157","article-title":"The Marginal Value of Adaptive Gradient Methods in Machine Learning","author":"Wilson","year":"2017","journal-title":"arXiv preprint arXiv:1705.08292"},{"key":"2026040313220152500_ref158","volume-title":"Data Mining: Practical machine learning tools and techniques","author":"Witten","year":"2016"},{"key":"2026040313220152500_ref159","article-title":"Advances in Variational Inference","author":"Zhang","year":"2017","journal-title":"ArXiv e-prints"},{"key":"2026040313220152500_ref160","first-page":"2328","article-title":"Information-theoretic lower bounds for distributed statistical estimation with communication constraints","author":"Zhang","year":"2013","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"4","key":"2026040313220152500_ref161","doi-asserted-by":"crossref","first-page":"811","DOI":"10.1109\/TSP.2014.2385046","article-title":"Asynchronous adaptation and learning over networks\u2014Part I: Modeling and stability analysis","volume":"63","author":"Zhao","year":"2015","journal-title":"IEEE Transactions on Signal Processing"}],"container-title":["Foundations and Trends\u00ae in Signal Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftsig\/article-pdf\/12\/3-4\/200\/11135923\/2000000102en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftsig\/article-pdf\/12\/3-4\/200\/11135923\/2000000102en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T18:55:20Z","timestamp":1777488920000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftsig\/article\/12\/3-4\/200\/1331350\/A-Brief-Introduction-to-Machine-Learning-for"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,8,14]]},"references-count":161,"journal-issue":{"issue":"3-4","published-print":{"date-parts":[[2018,8,14]]}},"URL":"https:\/\/doi.org\/10.1561\/2000000102","relation":{},"ISSN":["1932-8346","1932-8354"],"issn-type":[{"value":"1932-8346","type":"print"},{"value":"1932-8354","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,8,14]]}}}