{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T07:29:36Z","timestamp":1786606176858,"version":"build-2736575974"},"reference-count":88,"publisher":"Society for Industrial & Applied Mathematics (SIAM)","issue":"3","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["12150410304"],"award-info":[{"award-number":["12150410304"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["W2532005"],"award-info":[{"award-number":["W2532005"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100010877","name":"Shenzhen Science and Technology Innovation Commission","doi-asserted-by":"publisher","award":["RCYX20221008093033010"],"award-info":[{"award-number":["RCYX20221008093033010"]}],"id":[{"id":"10.13039\/501100010877","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Guangdong Provincial Key Laboratory of Mathematical Foundations for Artificial Intelligence","award":["2023B1212010001"],"award-info":[{"award-number":["2023B1212010001"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["SIAM J. Optim."],"published-print":{"date-parts":[[2026,9,30]]},"abstract":"<jats:p>Abstract.<\/jats:p>\n                  <jats:p>The proximal stochastic gradient method ([Formula: see text]) is one of the state-of-the-art approaches for stochastic composite-type problems.\u00a0In contrast to its deterministic counterpart, [Formula: see text] has been found to have difficulties with the correct identification of underlying substructures (such as supports, low rank patterns, or active constraints) and it does not possess a finite-time manifold identification property. Existing solutions rely on convexity assumptions or the additional usage of variance reduction techniques.\u00a0In this paper, we address these limitations and present a simple variant of [Formula: see text] based on Robinson\u2019s normal map.\u00a0The proposed normal map-based proximal stochastic gradient method ([Formula: see text]) is shown to converge globally; i.e., accumulation points of the generated iterates correspond to stationary points almost surely. In addition, we establish complexity bounds for [Formula: see text] that match the known results for [Formula: see text], and we prove that [Formula: see text] can almost surely identify active manifolds in finite time in a general nonconvex setting. Our derivations are built on almost sure iterate convergence guarantees and utilize analysis techniques based on the Kurdyka\u2013\u0141ojasiewicz inequality.<\/jats:p>","DOI":"10.1137\/25m1757332","type":"journal-article","created":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T07:00:43Z","timestamp":1786604443000},"page":"1773-1804","source":"Crossref","is-referenced-by-count":0,"title":["A Normal Map-Based Proximal Stochastic Gradient Method: Convergence and Identification Properties"],"prefix":"10.1137","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-2251-1951","authenticated-orcid":true,"given":"Junwen","family":"Qiu","sequence":"first","affiliation":[{"name":"School of Information Management and Engineering, Shanghai University of Finance and Economics, Shanghai 200433, China; Department of Industrial Systems Engineering and Management, National University of Singapore, Singapore 119077, Singapore."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Li","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Guangdong 518172, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6784-5417","authenticated-orcid":true,"given":"Andre","family":"Milzarek","sequence":"additional","affiliation":[{"name":"School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Guangdong 518172, China."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"351","published-online":{"date-parts":[[2026,8,13]]},"reference":[{"key":"ref1","doi-asserted-by":"publisher","DOI":"10.1137\/040605266"},{"key":"ref2","first-page":"1","volume":"18","author":"Atchad\u00e9 Y. F.","year":"2017","journal-title":"J. Mach. Learn. Res."},{"key":"ref3","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-007-0133-5"},{"key":"ref4","doi-asserted-by":"publisher","DOI":"10.1287\/moor.1100.0449"},{"key":"ref5","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4419-9467-7"},{"key":"ref6","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611974997"},{"key":"ref7","doi-asserted-by":"publisher","DOI":"10.1007\/BFb0096509"},{"key":"ref8","series-title":"Inf. Sci. Stat.","volume-title":"Pattern Recognition and Machine Learning","author":"Bishop C. M.","year":"2006"},{"key":"ref9","doi-asserted-by":"publisher","DOI":"10.1137\/050644641"},{"key":"ref10","doi-asserted-by":"publisher","DOI":"10.1137\/060670080"},{"key":"ref11","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-013-0701-9"},{"key":"ref12","series-title":"Texts Read. Math.\u00a048","volume-title":"Stochastic Approximation: A Dynamical Systems Viewpoint","author":"Borkar V. S.","year":"2023","edition":"2"},{"key":"ref13","doi-asserted-by":"publisher","DOI":"10.1137\/21M1465470"},{"key":"ref14","doi-asserted-by":"publisher","DOI":"10.1137\/16M1080173"},{"key":"ref15","doi-asserted-by":"publisher","DOI":"10.1137\/22M1515550"},{"key":"ref16","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611971286"},{"key":"ref17","doi-asserted-by":"crossref","unstructured":"D. L. Burkholder, B. J. Davis, and R. F. Gundy, Integral inequalities for convex functions of operators on martingales, in Proceedings of the Sixth Berkeley Symposium on Mathematical Statistics and Probability, Vol. II: Probability Theory, University of California Press,\u00a0Berkeley, CA, 1972, pp. 223\u2013240.","DOI":"10.1525\/9780520423671-018"},{"key":"ref18","doi-asserted-by":"publisher","DOI":"10.1145\/1970392.1970395"},{"key":"ref19","unstructured":"E. Chouzenoux, J.B. Fest, and A. Repetti, A Kurdyka-\u0141ojasiewicz Property for Stochastic Optimization Algorithms in a Non-Convex Setting, preprint, arXiv:2302.06447, 2023."},{"key":"ref20","doi-asserted-by":"publisher","DOI":"10.1137\/140971233"},{"key":"ref21","unstructured":"M. Coste, An Introduction to O-Minimal Geometry, hal-05413940, 1999,\u00a0https:\/\/univ-rennes.hal.science\/hal-05413940v1."},{"key":"ref22","unstructured":"A. Cutkosky and F. Orabona, Momentum-based variance reduction in non-convex SGD, in Proceedings of the 33rd International Conference on Neural Information Processing Systems, Adv. Neural Inf. Process. Syst. 32, Curran Associates Inc., Red Hook, NY, 2019, pp.\u00a015236\u201315245."},{"key":"ref23","unstructured":"Y. Dai, G. Wang, F. E. Curtis, and D. P. Robinson, A variance-reduced and stabilized proximal stochastic gradient method with support identification guarantees for structured optimization, in Proceedings of the 26th Int. Conf. Artif. Intell. Stat., 2023, pp. 5107\u20135133."},{"key":"ref24","doi-asserted-by":"publisher","DOI":"10.1137\/18M1178244"},{"key":"ref25","doi-asserted-by":"publisher","DOI":"10.1007\/s10208-018-09409-5"},{"key":"ref26","doi-asserted-by":"publisher","DOI":"10.4208\/jml.240109"},{"key":"ref27","doi-asserted-by":"publisher","DOI":"10.1137\/20M1387213"},{"key":"ref28","doi-asserted-by":"publisher","DOI":"10.1287\/moor.2017.0889"},{"key":"ref29","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-018-1311-3"},{"key":"ref30","first-page":"2899","volume":"10","author":"Duchi J.","year":"2009","journal-title":"J. Mach. Learn. Res."},{"key":"ref31","doi-asserted-by":"publisher","DOI":"10.1137\/17M1135086"},{"key":"ref32","doi-asserted-by":"publisher","DOI":"10.1214\/19-AOS1831"},{"key":"ref33","series-title":"Springer Ser. Oper. Res.","volume-title":"Finite-Dimensional Variational Inequalities and Complementarity Problems","author":"Facchinei F.","year":"2003"},{"key":"ref34","doi-asserted-by":"publisher","DOI":"10.1007\/s10957-014-0642-3"},{"key":"ref35","doi-asserted-by":"publisher","DOI":"10.1007\/s002220050066"},{"key":"ref36","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-014-0846-1"},{"key":"ref37","doi-asserted-by":"publisher","DOI":"10.2307\/1967124"},{"key":"ref38","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-21606-5"},{"key":"ref39","unstructured":"Z.S. Huang and C.P. Lee, Training structured neural networks through manifold identification and variance reduction, in Proceedings of the Tenth International Conference on Learning Representations (ICLR), virtual, 2022, https:\/\/openreview.net\/forum?id=mdUYT5QV0O."},{"key":"ref40","unstructured":"S. J. Reddi, S. Sra, B. Poczos, and A. J. Smola, Proximal stochastic methods for nonsmooth nonconvex finite-sum optimization, in NIPS\u201916: Proceedings of the 30th International Conference on Neural Information Processing Systems, Adv. Neural Inf. Process. Syst. 29, Curran Associates, Inc., Red Hook, NY, 2016, pp.\u00a01153\u20131161, https:\/\/dl.acm.org\/doi\/abs\/10.5555\/3157096.3157225."},{"key":"ref41","doi-asserted-by":"publisher","DOI":"10.1137\/100817206"},{"key":"ref42","doi-asserted-by":"publisher","DOI":"10.1137\/23M1545720"},{"key":"ref43","unstructured":"C. Josz, L. Lai, and X. Li, Proximal Random Reshuffling under Local Lipschitz Continuity, preprint, arXiv:2408.07182, 2024."},{"key":"ref44","doi-asserted-by":"publisher","DOI":"10.5802\/aif.1638"},{"key":"ref45","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4684-9352-8"},{"key":"ref46","unstructured":"R. M. Larsen, PROPACK\u2014Software for Large and Sparse SVD Calculations, http:\/\/sun.stanford.edu\/\u223crmunk\/PROPACK\/, 2012."},{"key":"ref47","doi-asserted-by":"publisher","DOI":"10.1137\/21M140376X"},{"key":"ref48","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-020-01599-7"},{"key":"ref49","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"ref50","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-022-01916-2"},{"key":"ref51","first-page":"1705","volume":"13","author":"Lee S.","year":"2012","journal-title":"J. Mach. Learn. Res."},{"key":"ref52","doi-asserted-by":"publisher","DOI":"10.1137\/S1052623401387623"},{"key":"ref53","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-015-0943-9"},{"key":"ref54","doi-asserted-by":"publisher","DOI":"10.1137\/110852103"},{"key":"ref55","doi-asserted-by":"publisher","DOI":"10.1007\/s10208-017-9366-8"},{"key":"ref56","doi-asserted-by":"crossref","unstructured":"X. Li and A. Milzarek, A unified convergence theorem for stochastic optimization methods, in NIPS\u201922: Proceedings of the 36th International Conference on Neural Information Processing Systems, Adv. Neural Inf. Process. Syst. 35, Curran Associates, Inc., Red Hook, NY, 2022, pp. 33107\u201333119, https:\/\/dl.acm.org\/doi\/10.5555\/3600270.3602669.","DOI":"10.52202\/068431-2399"},{"key":"ref57","doi-asserted-by":"publisher","DOI":"10.1137\/21M1468048"},{"key":"ref58","doi-asserted-by":"publisher","DOI":"10.1137\/16M106340X"},{"key":"ref59","doi-asserted-by":"publisher","DOI":"10.1080\/02331934.2023.2230976"},{"key":"ref60","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.1977.1101561"},{"key":"ref61","series-title":"Colloq. Internat. CNRS,\u00a0No. 117 [International Colloquia of the CNRS]","first-page":"87","volume-title":"Les \u00c9quations aux D\u00e9riv\u00e9es Partielles (Paris, 1962)","author":"\u0141ojasiewicz S.","year":"1963"},{"key":"ref62","unstructured":"S. Majewski, B. Miasojedow, and E. Moulines, Analysis of Nonsmooth Stochastic Approximation: The Differential Inclusion Approach, preprint, https:\/\/arxiv.org\/abs\/1805.01916v1, 2018."},{"key":"ref63","unstructured":"A. Milzarek, Numerical Methods and Second Order Theory for Nonsmooth Problems, Ph.D. thesis, Technische Universit\u00e4t M\u00fcnchen, 2016."},{"key":"ref64","doi-asserted-by":"publisher","DOI":"10.1137\/18M1181249"},{"key":"ref65","doi-asserted-by":"publisher","DOI":"10.24033\/bsmf.1625"},{"key":"ref66","doi-asserted-by":"publisher","DOI":"10.1137\/070704277"},{"key":"ref67","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-007-0149-x"},{"key":"ref68","unstructured":"A. Nitanda, Stochastic proximal gradient descent with acceleration techniques, in NIPS\u201914: Proceedings of the 28th International Conference on Neural Information Processing Systems, Adv. Neural Inf. Process. Syst. 27, MIT Press, Cambridge, MA,\u00a02014, pp. 1574\u20131582, https:\/\/proceedings.neurips.cc\/paper_files\/paper\/2014\/hash\/2d6cd90d4f3fa50e6d9bdbc81a2e3712-Abstract.html."},{"key":"ref69","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-024-02110-2"},{"key":"ref70","doi-asserted-by":"publisher","DOI":"10.1007\/s11590-021-01702-7"},{"key":"ref71","unstructured":"C. Poon, J. Liang, and C.B. Sch\u00f6nlieb, Local convergence properties of SAGA\/Prox-SVRG and acceleration, in Proceedings of the 35th International Conference on Machine Learning,\u00a0PMLR 80, 2018, pp. 4124\u20134132."},{"key":"ref72","unstructured":"Z. Qin and J. Liang, Partial Smoothness, Subdifferentials and Set-Valued Operators, preprint, https:\/\/arxiv.org\/abs\/2501.15540, 2025."},{"key":"ref73","first-page":"1","volume":"26","author":"Qiu J.","year":"2025","journal-title":"J. Mach. Learn. Res."},{"key":"ref74","unstructured":"J. Qiu, B. Ma, and A. Milzarek, Convergence of SGD with Momentum in the Nonconvex Case: A Novel Time Window-Based Analysis, preprint, arXiv:2405.16954, 2024."},{"key":"ref75","doi-asserted-by":"publisher","DOI":"10.1214\/aoms\/1177729586"},{"key":"ref76","doi-asserted-by":"publisher","DOI":"10.1287\/moor.17.3.691"},{"key":"ref77","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-02431-3"},{"key":"ref78","doi-asserted-by":"publisher","DOI":"10.1007\/s00245-019-09617-7"},{"key":"ref79","first-page":"1865","volume":"12","author":"Shalev-Shwartz S.","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref80","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611973433"},{"key":"ref81","volume-title":"Probability Theory","author":"Stroock D. W.","year":"2011","edition":"2"},{"key":"ref82","unstructured":"Y. Sun, H. Jeong, J. Nutini, and M. Schmidt, Are we there yet? Manifold identification of gradient-related proximal methods, in Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics, PMLR 89, 2019, pp. 1110\u20131119."},{"key":"ref83","doi-asserted-by":"publisher","DOI":"10.1016\/j.spa.2014.11.001"},{"key":"ref84","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2017.2713822"},{"key":"ref85","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511525919"},{"key":"ref86","doi-asserted-by":"publisher","DOI":"10.1137\/15M1053141"},{"key":"ref87","first-page":"2543","volume":"11","author":"Xiao L.","year":"2010","journal-title":"J. Mach. Learn. Res."},{"key":"ref88","doi-asserted-by":"publisher","DOI":"10.1137\/140961791"}],"container-title":["SIAM Journal on Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/epubs.siam.org\/doi\/pdf\/10.1137\/25M1757332","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,13]],"date-time":"2026-08-13T07:01:07Z","timestamp":1786604467000},"score":1,"resource":{"primary":{"URL":"https:\/\/epubs.siam.org\/doi\/10.1137\/25M1757332"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,8,13]]},"references-count":88,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,9,30]]}},"alternative-id":["10.1137\/25M1757332"],"URL":"https:\/\/doi.org\/10.1137\/25m1757332","relation":{},"ISSN":["1052-6234","1095-7189"],"issn-type":[{"value":"1052-6234","type":"print"},{"value":"1095-7189","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,8,13]]}}}