{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,7]],"date-time":"2026-07-07T18:07:51Z","timestamp":1783447671525,"version":"3.55.0"},"reference-count":53,"publisher":"MIT Press","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Neural Computation"],"published-print":{"date-parts":[[2010,4]]},"abstract":"<jats:p>We present a novel algorithm for efficient learning and feature selection in high-dimensional regression problems. We arrive at this model through a modification of the standard regression model, enabling us to derive a probabilistic version of the well-known statistical regression technique of backfitting. Using the expectation-maximization algorithm, along with variational approximation methods to overcome intractability, we extend our algorithm to include automatic relevance detection of the input features. This variational Bayesian least squares (VBLS) approach retains its simplicity as a linear model, but offers a novel statistically robust black-box approach to generalized linear regression with high-dimensional inputs. It can be easily extended to nonlinear regression and classification problems. In particular, we derive the framework of sparse Bayesian learning, the relevance vector machine, with VBLS at its core, offering significant computational and robustness advantages for this class of methods. The iterative nature of VBLS makes it most suitable for real-time incremental learning, which is crucial especially in the application domain of robotics, brain-machine interfaces, and neural prosthetics, where real-time learning of models for control is needed. We evaluate our algorithm on synthetic and neurophysiological data sets, as well as on standard regression and classification benchmark data sets, comparing it with other competitive statistical approaches and demonstrating its suitability as a drop-in replacement for other generalized linear regression techniques.<\/jats:p>","DOI":"10.1162\/neco.2009.02-08-702","type":"journal-article","created":{"date-parts":[[2009,12,22]],"date-time":"2009-12-22T20:18:13Z","timestamp":1261513093000},"page":"831-886","source":"Crossref","is-referenced-by-count":24,"title":["Efficient Learning and Feature Selection in High-Dimensional Regression"],"prefix":"10.1162","volume":"22","author":[{"given":"Jo-Anne","family":"Ting","sequence":"first","affiliation":[{"name":"University of Edinburgh, Edinburgh EH8 9AB, U.K."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Aaron","family":"D'Souza","sequence":"additional","affiliation":[{"name":"Google Inc., Mountain View, CA 94043, U.S.A."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sethu","family":"Vijayakumar","sequence":"additional","affiliation":[{"name":"University of Edinburgh, Edinburgh EH8 9AB, U.K."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stefan","family":"Schaal","sequence":"additional","affiliation":[{"name":"University of Southern California, Los Angeles, CA 90089, U.S.A., and ATR Computational Neuroscience Laboratories, Kyoto, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.1002\/0471725153"},{"key":"B2","first-page":"46","volume-title":"Proceedings of the 16th Conference in Uncertainty in Artificial Intelligence","author":"Bishop C. M.","year":"2000"},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.1007\/BF00994018"},{"key":"B4","doi-asserted-by":"publisher","DOI":"10.1111\/j.2044-8317.1992.tb00992.x"},{"key":"B5","doi-asserted-by":"publisher","DOI":"10.1214\/009053604000000067"},{"key":"B6","doi-asserted-by":"publisher","DOI":"10.1007\/978-94-009-5564-6"},{"key":"B7","volume-title":"Advances in neural information processing systems","volume":"2","author":"Fahlman S. E.","year":"1989"},{"key":"B8","doi-asserted-by":"publisher","DOI":"10.1109\/29.1552"},{"key":"B9","doi-asserted-by":"publisher","DOI":"10.2307\/1269656"},{"key":"B10","doi-asserted-by":"publisher","DOI":"10.1145\/355744.355745"},{"key":"B11","doi-asserted-by":"publisher","DOI":"10.1214\/07-AOAS131"},{"key":"B12","volume-title":"Kernel dimensionality reduction in regression","author":"Fukumizu K.","year":"2006"},{"key":"B13","volume-title":"Bayesian data analysis","author":"Gelman A.","year":"2000"},{"key":"B14","volume-title":"Advanced mean field methods: Theory and practice","author":"Ghahramani Z.","year":"2000"},{"key":"B15","volume-title":"Advances in neural information processing systems","volume":"12","author":"Ghahramani Z.","year":"2000"},{"key":"B16","volume-title":"The EM algorithm for mixtures of factor analyzers","author":"Ghahramani Z.","year":"1997"},{"key":"B17","volume-title":"Proceedings of the Annual Meeting of the Association for Computational Linguistics","author":"Goodman J.","year":"2004"},{"key":"B18","volume-title":"Advances in neural information processing systems","volume":"13","author":"Gray A. G.","year":"2001"},{"key":"B19","volume-title":"Generalized additive models","author":"Hastie T.","year":"1990"},{"key":"B20","doi-asserted-by":"publisher","DOI":"10.1214\/ss\/1009212815"},{"key":"B21","doi-asserted-by":"publisher","DOI":"10.1023\/A:1008932416310"},{"key":"B22","first-page":"453","volume":"186","author":"Jeffreys H.","year":"1946","journal-title":"Journal of the Royal Statistical Society. Series A"},{"key":"B23","doi-asserted-by":"publisher","DOI":"10.1126\/science.285.5436.2136"},{"key":"B24","first-page":"1493","volume":"7","author":"Keerthi S. S.","year":"2006","journal-title":"Journal of Machine Learning Research"},{"key":"B25","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2007.910971"},{"key":"B26","volume-title":"Proceedings of the 17th International Conference on Machine Learning","author":"Komarek P.","year":"2000"},{"key":"B27","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2005.127"},{"key":"B28","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"B29","volume-title":"Proceedings of the 21st National Conference on Artificial Intelligence","author":"Lee S.","year":"2006"},{"key":"B30","volume-title":"The LASSO and generalized linear models","author":"Lokhorst J.","year":"1999"},{"key":"B31","doi-asserted-by":"publisher","DOI":"10.1162\/089976699300016331"},{"key":"B32","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.1965.10480787"},{"key":"B33","first-page":"397","volume-title":"Proceedings of the 16th Conference in Uncertainty in Artificial Intelligence","author":"Moore A.","year":"2000"},{"key":"B34","doi-asserted-by":"publisher","DOI":"10.1613\/jair.453"},{"key":"B36","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015435"},{"key":"B37","volume-title":"Advances in neural information processing systems","volume":"3","author":"Omohundro S. M.","year":"1990"},{"key":"B38","volume-title":"Statistical field theory","author":"Parisi G.","year":"1988"},{"key":"B39","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007618119488"},{"key":"B40","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511812651"},{"key":"B41","volume-title":"Variational methods in statistics","author":"Rustagi J.","year":"1976"},{"key":"B42","volume-title":"Advances in neural information processing systems","volume":"10","author":"Schaal S.","year":"1998"},{"key":"B43","doi-asserted-by":"publisher","DOI":"10.1152\/jn.1998.80.3.1577"},{"key":"B44","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btg308"},{"issue":"1","key":"B45","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1111\/j.2517-6161.1996.tb02080.x","volume":"58","author":"Tibshirani R.","year":"1996","journal-title":"Journal of Royal Statistical Society, Series B"},{"key":"B46","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143962"},{"key":"B47","volume-title":"Advances in neural information processing systems","volume":"18","author":"Ting J.","year":"2005"},{"key":"B48","doi-asserted-by":"publisher","DOI":"10.1162\/15324430152748236"},{"key":"B49","doi-asserted-by":"publisher","DOI":"10.1162\/089976699300016728"},{"key":"B50","volume-title":"Proceedings of the 9th International Workshop on Artificial Intelligence and Statistics","author":"Tipping M. E.","year":"2003"},{"key":"B51","first-page":"1079","volume-title":"Proceedings of the 17th International Conference on Machine Learning","author":"Vijayakumar S.","year":"2000"},{"key":"B52","first-page":"589","volume":"16","author":"Wang L.","year":"2006","journal-title":"Statistica Sinica"},{"key":"B53","volume-title":"Advances in neural information processing systems","volume":"8","author":"Williams C. K. I.","year":"1996"},{"key":"B54","volume-title":"Perspectives in probability and statistics: Papers in honor of M.S. Bartlett","author":"Wold H.","year":"1975"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mitpressjournals.org\/doi\/pdf\/10.1162\/neco.2009.02-08-702","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,15]],"date-time":"2025-02-15T21:19:27Z","timestamp":1739654367000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/22\/4\/831-886\/7527"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2010,4]]},"references-count":53,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2010,4]]}},"alternative-id":["10.1162\/neco.2009.02-08-702"],"URL":"https:\/\/doi.org\/10.1162\/neco.2009.02-08-702","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2010,4]]}}}