{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T18:33:25Z","timestamp":1787337205515,"version":"build-2736575974"},"reference-count":50,"publisher":"Society for Industrial & Applied Mathematics (SIAM)","issue":"1","funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CCF-2009030"],"award-info":[{"award-number":["CCF-2009030"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["SIAM Journal on Mathematics of Data Science"],"published-print":{"date-parts":[[2022,3]]},"abstract":"<jats:p>Deep neural networks generalize well despite being exceedingly overparameterized and being trained without explicit regularization. This curious phenomenon has inspired extensive research activity in establishing its statistical principles: Under what conditions is it observed? How do these depend on the data and on the training algorithm? When does regularization benefit generalization? While such questions remain wide open for deep neural nets, recent works have attempted gaining insights by studying simpler, often linear, models. Our paper contributes to this growing line of work by examining binary linear classification under a generative Gaussian mixture model in which the feature vectors take the form ${{\\it x}}=\\pm{{\\eta}}+{{\\it q}}$, where for a mean vector $\\eta$ and feature noise ${{\\it q}} \\sim \\mathcal{N}(0,{{\\Sigma}})$. Motivated by recent results on the implicit bias of gradient descent, we study both max-margin support vector machine (SVM) classifiers (corresponding to logistic loss) and min-norm interpolating classifiers (corresponding to least-squares loss). First, we leverage an idea introduced in [V. Muthukumar et al., arXiv:2005.08054, 2020a] to relate the SVM solution to the min-norm interpolating solution. Second, we derive novel nonasymptotic bounds on the classification error of the latter. Combining the two, we present novel sufficient conditions on the covariance spectrum and on the signal-to-noise ratio (SNR) $SNR={||{{\\eta}}||_2^4}\/{{\\eta}}^T{{\\Sigma\\eta}}$ under which interpolating estimators achieve asymptotically optimal performance as overparameterization increases. Interestingly, our results extend to a noisy model with constant probability noise flips. Contrary to previously studied discriminative data models, our results emphasize the crucial role of the SNR and its interplay with the data covariance. Finally, via a combination of analytical arguments and numerical demonstrations we identify conditions under which the interpolating estimator performs better than corresponding regularized estimates.<\/jats:p>","DOI":"10.1137\/21m1415121","type":"journal-article","created":{"date-parts":[[2022,3,3]],"date-time":"2022-03-03T10:56:12Z","timestamp":1646304972000},"page":"260-284","source":"Crossref","is-referenced-by-count":7,"title":["Binary Classification of Gaussian Mixtures: Abundance of Support Vectors, Benign Overfitting, and Regularization"],"prefix":"10.1137","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4501-5102","authenticated-orcid":true,"given":"Ke","family":"Wang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christos","family":"Thrampoulidis","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"351","published-online":{"date-parts":[[2022,3,3]]},"reference":[{"key":"atypb1","unstructured":"J. Ba, M. Erdogdu, T. Suzuki, D. Wu, and T. Zhang (2019),\n                      Generalization of two-layer neural networks: An asymptotic viewpoint\n                      , in Proceedings of the International Conference on Learning Representations."},{"key":"atypb2","doi-asserted-by":"crossref","first-page":"138","DOI":"10.1198\/016214505000000907","volume":"101","author":"Bartlett P. L.","year":"2006","journal-title":"J. Amer. Statist. Assoc."},{"key":"atypb3","doi-asserted-by":"crossref","first-page":"30063","DOI":"10.1073\/pnas.1907378117","volume":"117","author":"Bartlett P. L.","year":"2020","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"atypb4","doi-asserted-by":"crossref","unstructured":"M. Belkin, D. Hsu, S. Ma, and S. Mandal (2018a),\n                      Reconciling Modern Machine Learning and the Bias-Variance Trade-Off\n                      , arXiv:1812.11118.","DOI":"10.1073\/pnas.1903070116"},{"key":"atypb5","unstructured":"M. Belkin, D. J. Hsu, and P. Mitra (2018b),\n                      Overfitting or perfect fitting? Risk bounds for classification and regression rules that interpolate\n                      , in Advances in Neural Information Processing Systems, pp. 2300-2311."},{"key":"atypb6","unstructured":"M. Belkin, S. Ma, and S. Mandal (2018c),\n                      To Understand Deep Learning We Need to Understand Kernel Learning\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1802.01396 arXiv:1802.01396."},{"key":"atypb7","unstructured":"M. Belkin, D. Hsu, and J. Xu (2019),\n                      Two Models of Double Descent for Weak Features\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1903.07571 arXiv:1903.07571."},{"key":"atypb8","unstructured":"Y. Cao, Q. Gu, and M. Belkin (2021),\n                      Risk Bounds for Over-Parameterized Maximum Margin Classification on Sub-Gaussian Mixtures\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2104.13628 arXiv:2104.13628."},{"key":"atypb9","unstructured":"N. S. Chatterji and P. M. Long (2020),\n                      Finite-Sample Analysis of Interpolating Linear Classifiers in the Overparameterized Regime\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2004.12019 arXiv:2004.12019."},{"key":"atypb10","unstructured":"N. S. Chatterji, P. M. Long, and P. L. Bartlett (2020),\n                      When Does Gradient Descent with Logistic Loss Find Interpolating Two-Layer Networks?\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2012.02409 arXiv:2012.02409."},{"key":"atypb11","unstructured":"T. H. Cormen, C. E. Leiserson, R. L. Rivest, and C. Stein (2009),\n                      Introduction to Algorithms\n                      , MIT Press, Cambridge, MA."},{"key":"atypb12","unstructured":"Z. Deng, A. Kammoun, and C. Thrampoulidis (2019),\n                      A Model of Double Descent for High-Dimensional Binary Linear Classification\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1911.05822 arXiv:1911.05822."},{"key":"atypb13","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/ICPR.2000.906006","volume":"2","author":"Duin R. P.","year":"2000","journal-title":"Proceedings of the 15th International Conference on Pattern Recognition"},{"key":"atypb14","unstructured":"I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio (2016),\n                      Deep learning\n                      , Vol. 1, MIT Press, Cambridge MA."},{"key":"atypb15","doi-asserted-by":"crossref","unstructured":"T. Hastie, R. Tibshirani, and J. Friedman (2009),\n                      The Elements of Statistical Learning: Data Mining, Inference, and Prediction\n                      , Springer, New York.","DOI":"10.1007\/978-0-387-84858-7"},{"key":"atypb16","unstructured":"T. Hastie, A. Montanari, S. Rosset, and R. J. Tibshirani (2019),\n                      Surprises in High-Dimensional Ridgeless Least Squares Interpolation\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1903.08560 arXiv:1903.08560."},{"key":"atypb17","unstructured":"R. A. Horn and C. R. Johnson (2012),\n                      Matrix Analysis\n                      , Cambridge University Press, Cambridge, UK."},{"key":"atypb18","unstructured":"D. Hsu, V. Muthukumar, and J. Xu (2020),\n                      On the Proliferation of Support Vectors in High Dimensions\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2009.10670 arXiv:2009.10670."},{"key":"atypb19","unstructured":"Z. Ji and M. Telgarsky (2019),\n                      The Implicit Bias of Gradient Descent on Nonseparable Data\n                      , in Proceedings of the Conference on Learning Theory, pp. 1772-1798."},{"key":"atypb20","unstructured":"A. Kammoun and M.S. Alouini (2020),\n                      On the Precise Error Analysis of Support Vector Machines\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2003.12972 arXiv:2003.12972."},{"key":"atypb21","doi-asserted-by":"crossref","unstructured":"G. Kini and C. Thrampoulidis (2020),\n                      Analytic Study of Double Descent in Binary Classification: The Impact of Loss\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2001.11572 arXiv:2001.11572.","DOI":"10.1109\/ISIT44484.2020.9174344"},{"key":"atypb22","unstructured":"D. Kobak, J. Lomond, and B. Sanchez (2018),\n                      Optimal Ridge Penalty for Real-World High-Dimensional Data Can Be Zero or Negative Due to the Implicit Ridge Regularization\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1805.10939 arXiv:1805.10939."},{"key":"atypb23","first-page":"1097","volume":"25","author":"Krizhevsky A.","year":"2012","journal-title":"Advances in Neural Information Processing Systems"},{"key":"atypb24","unstructured":"T. Liang, A. Rakhlin, and X. Zhai (2019),\n                      On the Risk of Minimum-Norm Interpolants and Restricted Lower Isometry of Kernels\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1908.10292 arXiv:1908.10292."},{"key":"atypb25","doi-asserted-by":"crossref","unstructured":"Z. Liao, R. Couillet, and M. W. Mahoney (2020),\n                      A Random Matrix Analysis of Random Fourier Features: Beyond the Gaussian Kernel, a Precise Phase Transition, and the Corresponding Double Descent\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2006.05013 arXiv:2006.05013.","DOI":"10.1088\/1742-5468\/ac3a77"},{"key":"atypb26","doi-asserted-by":"crossref","first-page":"10625","DOI":"10.1073\/pnas.2001875117","volume":"117","author":"Loog M.","year":"2020","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"atypb27","unstructured":"S. Mei and A. Montanari (2019),\n                      The Generalization Error of Random Features Regression: Precise Asymptotics and Double Descent Curve\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1908.05355 arXiv:1908.05355."},{"key":"atypb28","unstructured":"F. Mignacco, F. Krzakala, Y. M. Lu, and L. Zdeborov\u00e1 (2020),\n                      The Role of Regularization in Classification of High-Dimensional Noisy Gaussian Mixture\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2002.11544 arXiv:2002.11544."},{"key":"atypb29","unstructured":"A. Montanari, F. Ruan, Y. Sohn, and J. Yan (2019),\n                      The Generalization Error of max-Margin Linear Classifiers: High-Dimensional Asymptotics in the Overparametrized Regime\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1911.01544 arXiv:1911.01544."},{"key":"atypb30","unstructured":"G. F. Montufar, R. Pascanu, K. Cho, and Y. Bengio (2014),\n                      On the number of linear regions of deep neural networks\n                      , in Advances in Neural Information Processing Systems, pp. 2924-2932."},{"key":"atypb31","unstructured":"V. Muthukumar, A. Narang, V. Subramanian, M. Belkin, D. Hsu, and A. Sahai (2020a),\n                      Classification vs Regression in Overparameterized Regimes: Does the Loss Function Matter?\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2005.08054 arXiv:2005.08054."},{"key":"atypb32","doi-asserted-by":"crossref","unstructured":"V. Muthukumar, K. Vodrahalli, V. Subramanian, and A. Sahai (2020b),\n                      Harmless interpolation of noisy data in regression\n                      , IEEE J. Selected Areas Information Theory, pp. 67-83.","DOI":"10.1109\/JSAIT.2020.2984716"},{"key":"atypb33","unstructured":"P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever (2019),\n                      Deep Double Descent: Where Bigger Models and More Data Hurt\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1912.02292 arXiv:1912.02292."},{"key":"atypb34","doi-asserted-by":"crossref","first-page":"L581","DOI":"10.1088\/0305-4470\/23\/11\/012","volume":"23","author":"Opper M.","year":"1990","journal-title":"J. Phys. A"},{"key":"atypb35","doi-asserted-by":"crossref","first-page":"503","DOI":"10.1007\/s11633-017-1054-2","volume":"14","author":"Poggio T.","year":"2017","journal-title":"Internat. J. Automation Comput."},{"key":"atypb36","unstructured":"S. Rosset, J. Zhu, and T. Hastie (2003),\n                      Margin maximizing loss functions\n                      , in Proceedings of NIPS, pp. 1237-1244."},{"key":"atypb37","unstructured":"F. Salehi, E. Abbasi, and B. Hassibi (2019),\n                      The Impact of Regularization on High-Dimensional Logistic Regression\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1906.03761 arXiv:1906.03761."},{"key":"atypb38","unstructured":"F. Salehi, E. Abbasi, and B. Hassibi (2020),\n                      The Performance Analysis of Generalized Margin Maximizer (GMM) on Separable Data\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2010.15379 arXiv:2010.15379."},{"key":"atypb39","first-page":"2822","volume":"19","author":"Soudry D.","year":"2018","journal-title":"J. Mach. Learn. Res."},{"key":"atypb40","unstructured":"M. Stojnic (2013),\n                      A Framework to Characterize Performance of Lasso Algorithms\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1303.7291 arXiv:1303.7291."},{"key":"atypb41","doi-asserted-by":"crossref","first-page":"14516","DOI":"10.1073\/pnas.1810420116","volume":"116","author":"Sur P.","year":"2019","journal-title":"Proc. Natl. Acad. Sci."},{"key":"atypb42","first-page":"1683","volume":"40","author":"Thrampoulidis C.","year":"2015","journal-title":"Proc. Mach. Learn. Res."},{"key":"atypb43","doi-asserted-by":"crossref","first-page":"5592","DOI":"10.1109\/TIT.2018.2840720","volume":"64","author":"Thrampoulidis C.","year":"2018","journal-title":"IEEE Trans. Inform. Theory"},{"key":"atypb44","unstructured":"A. Tsigler and P. L. Bartlett (2020),\n                      Benign Overfitting in Ridge Regression\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/2009.14286 arXiv:2009.14286."},{"key":"atypb45","doi-asserted-by":"crossref","first-page":"315","DOI":"10.1209\/0295-5075\/9\/4\/003","volume":"9","author":"Vallet F.","year":"1989","journal-title":"Europhys. Lett."},{"key":"atypb46","doi-asserted-by":"crossref","unstructured":"R. Vershynin (2018),\n                      High-Dimensional Probability: An Introduction with Applications in Data Science\n                      , Cambridge Ser. Statist. Probab. Math. 47, Cambridge University Press, Cambridge, UK,","DOI":"10.1017\/9781108231596"},{"key":"atypb47","doi-asserted-by":"crossref","unstructured":"M. J. Wainwright (2019),\n                      High-Dimensional Statistics: A Non-Asymptotic Viewpoint\n                      , Cambridge Ser. Statist. Probab. Math. 48. Cambridge University Press, Cambridge, UK,","DOI":"10.1017\/9781108627771"},{"key":"atypb48","doi-asserted-by":"crossref","unstructured":"K. Wang and C. Thrampoulidis (2021),\n                      Benign overfitting in binary classification of gaussian mixtures\n                      , in Proceedings of the 2021 IEEE International Conference on Acoustics, Speech and Signal Processing.","DOI":"10.1109\/ICASSP39728.2021.9413946"},{"key":"atypb49","unstructured":"Z. Yang, Y. Yu, C. You, J. Steinhardt, and Y. Ma (2020),\n                      Rethinking bias-variance trade-off for generalization of neural networks\n                      , in Proceedings of the International Conference on Machine Learning, pp. 10767-10777."},{"key":"atypb50","unstructured":"C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals (2016),\n                      Understanding Deep Learning Requires Rethinking Generalization\n                      , preprint, \\hrefhttps:\/\/arxiv.org\/abs\/1611.03530 arXiv:1611.03530."}],"container-title":["SIAM Journal on Mathematics of Data Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/epubs.siam.org\/doi\/pdf\/10.1137\/21M1415121","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,21]],"date-time":"2026-08-21T17:49:09Z","timestamp":1787334549000},"score":1,"resource":{"primary":{"URL":"https:\/\/epubs.siam.org\/doi\/10.1137\/21M1415121"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3]]},"references-count":50,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,3]]}},"alternative-id":["10.1137\/21M1415121"],"URL":"https:\/\/doi.org\/10.1137\/21m1415121","relation":{},"ISSN":["2577-0187"],"issn-type":[{"value":"2577-0187","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3]]}}}