{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,30]],"date-time":"2026-07-30T17:44:18Z","timestamp":1785433458912,"version":"3.56.0"},"reference-count":23,"publisher":"Cambridge University Press (CUP)","issue":"2","license":[{"start":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T00:00:00Z","timestamp":1770681600000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["J. Appl. Probab."],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    In recent works on the theory of machine learning, it has been observed that heavy tail properties of stochastic gradient descent (SGD) can be studied in the probabilistic framework of stochastic recursions. In particular, G\u00fcrb\u00fczbalaban et al. (2021) considered a setup corresponding to linear regression for which iterations of SGD can be modelled by a multivariate affine stochastic recursion\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:inline-graphic xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" mime-subtype=\"png\" content-type=\"simple\" xlink:href=\"S0021900225100363_inline1.png\">\n                          <jats:alt-text content-type=\"machine-generated\">upper X Subscript n Baseline equals upper A Subscript n Baseline upper X Subscript n minus 1 Baseline plus upper B Subscript n<\/jats:alt-text>\n                        <\/jats:inline-graphic>\n                        <mml:math xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xmlns:mnf=\"http:\/\/cambridge.org\/core\/manifest\" xmlns:cup=\"http:\/\/contentservices.cambridge.org\" xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" xmlns:m=\"http:\/\/cambridge.org\/core\/metadata\" xmlns:core=\"http:\/\/cambridge.org\/core\" xmlns:c=\"http:\/\/cambridge.org\/core\/content\">\n                          <mml:msub>\n                            <mml:mi>X<\/mml:mi>\n                            <mml:mi>n<\/mml:mi>\n                          <\/mml:msub>\n                          <mml:mo>=<\/mml:mo>\n                          <mml:msub>\n                            <mml:mi>A<\/mml:mi>\n                            <mml:mi>n<\/mml:mi>\n                          <\/mml:msub>\n                          <mml:msub>\n                            <mml:mi>X<\/mml:mi>\n                            <mml:mrow>\n                              <mml:mi>n<\/mml:mi>\n                              <mml:mo>\u2212<\/mml:mo>\n                              <mml:mn>1<\/mml:mn>\n                            <\/mml:mrow>\n                          <\/mml:msub>\n                          <mml:mo>+<\/mml:mo>\n                          <mml:msub>\n                            <mml:mi>B<\/mml:mi>\n                            <mml:mi>n<\/mml:mi>\n                          <\/mml:msub>\n                        <\/mml:math>\n                        <jats:tex-math>$X_n=A_nX_{n-1}+B_n$<\/jats:tex-math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    for independent and identically distributed pairs\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:inline-graphic xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" mime-subtype=\"png\" content-type=\"simple\" xlink:href=\"S0021900225100363_inline2.png\">\n                          <jats:alt-text content-type=\"machine-generated\">left parenthesis upper A Subscript n Baseline comma upper B Subscript n Baseline right parenthesis<\/jats:alt-text>\n                        <\/jats:inline-graphic>\n                        <mml:math xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xmlns:mnf=\"http:\/\/cambridge.org\/core\/manifest\" xmlns:cup=\"http:\/\/contentservices.cambridge.org\" xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" xmlns:m=\"http:\/\/cambridge.org\/core\/metadata\" xmlns:core=\"http:\/\/cambridge.org\/core\" xmlns:c=\"http:\/\/cambridge.org\/core\/content\">\n                          <mml:mo stretchy=\"false\">(<\/mml:mo>\n                          <mml:msub>\n                            <mml:mi>A<\/mml:mi>\n                            <mml:mi>n<\/mml:mi>\n                          <\/mml:msub>\n                          <mml:mo>,<\/mml:mo>\n                          <mml:msub>\n                            <mml:mi>B<\/mml:mi>\n                            <mml:mi>n<\/mml:mi>\n                          <\/mml:msub>\n                          <mml:mo stretchy=\"false\">)<\/mml:mo>\n                        <\/mml:math>\n                        <jats:tex-math>$(A_n,B_n)$<\/jats:tex-math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    , where\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:inline-graphic xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" mime-subtype=\"png\" content-type=\"simple\" xlink:href=\"S0021900225100363_inline3.png\">\n                          <jats:alt-text content-type=\"machine-generated\">upper A Subscript n<\/jats:alt-text>\n                        <\/jats:inline-graphic>\n                        <mml:math xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xmlns:mnf=\"http:\/\/cambridge.org\/core\/manifest\" xmlns:cup=\"http:\/\/contentservices.cambridge.org\" xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" xmlns:m=\"http:\/\/cambridge.org\/core\/metadata\" xmlns:core=\"http:\/\/cambridge.org\/core\" xmlns:c=\"http:\/\/cambridge.org\/core\/content\">\n                          <mml:msub>\n                            <mml:mi>A<\/mml:mi>\n                            <mml:mi>n<\/mml:mi>\n                          <\/mml:msub>\n                        <\/mml:math>\n                        <jats:tex-math>$A_n$<\/jats:tex-math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    is a random symmetric matrix and\n                    <jats:inline-formula>\n                      <jats:alternatives>\n                        <jats:inline-graphic xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" mime-subtype=\"png\" content-type=\"simple\" xlink:href=\"S0021900225100363_inline4.png\">\n                          <jats:alt-text content-type=\"machine-generated\">upper B Subscript n<\/jats:alt-text>\n                        <\/jats:inline-graphic>\n                        <mml:math xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xmlns:mnf=\"http:\/\/cambridge.org\/core\/manifest\" xmlns:cup=\"http:\/\/contentservices.cambridge.org\" xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\" xmlns:m=\"http:\/\/cambridge.org\/core\/metadata\" xmlns:core=\"http:\/\/cambridge.org\/core\" xmlns:c=\"http:\/\/cambridge.org\/core\/content\">\n                          <mml:msub>\n                            <mml:mi>B<\/mml:mi>\n                            <mml:mi>n<\/mml:mi>\n                          <\/mml:msub>\n                        <\/mml:math>\n                        <jats:tex-math>$B_n$<\/jats:tex-math>\n                      <\/jats:alternatives>\n                    <\/jats:inline-formula>\n                    is a random vector. However, their approach is not completely correct and, in the present paper, the problem is put into the right framework by applying the theory of irreducible-proximal matrices.\n                  <\/jats:p>","DOI":"10.1017\/jpr.2025.10036","type":"journal-article","created":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T06:52:03Z","timestamp":1770706323000},"page":"459-483","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":0,"title":["Analysing heavy-tail properties of stochastic gradient descent by means of stochastic recurrence equations"],"prefix":"10.1017","volume":"63","author":[{"given":"Ewa","family":"Damek","sequence":"first","affiliation":[{"id":[{"id":"https:\/\/ror.org\/00yae6e25","id-type":"ROR","asserted-by":"publisher"}],"name":"University of Wroc\u0142aw"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sebastian","family":"Mentemeier","sequence":"additional","affiliation":[{"id":[{"id":"https:\/\/ror.org\/02f9det96","id-type":"ROR","asserted-by":"publisher"}],"name":"University of Hildesheim"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"56","published-online":{"date-parts":[[2026,2,10]]},"reference":[{"key":"S0021900225100363_ref15","doi-asserted-by":"publisher","DOI":"10.1007\/BF02392040"},{"key":"S0021900225100363_ref7","doi-asserted-by":"publisher","DOI":"10.1214\/13-AIHP566"},{"key":"S0021900225100363_ref23","unstructured":"[23] Wang, H. , Gurbuzbalaban, M. , Zhu, L. , Simsekli, U. and Erdogdu, M. A. (2021). Convergence rates of stochastic gradient descent under infinite noise variance. Adv. Neural Inf. Proc. Syst. 34, 18866\u201318877."},{"key":"S0021900225100363_ref12","unstructured":"[12] G\u00fcrb\u00fczbalaban, M. , Simsekli, U. and Zhu, L. (2021). The heavy-tail phenomenon in SGD. Proc. Mach. Learn. Res. 139, 3964\u20133975."},{"key":"S0021900225100363_ref21","doi-asserted-by":"publisher","DOI":"10.2307\/3213932"},{"key":"S0021900225100363_ref11","unstructured":"[11] G\u00fcrb\u00fczbalaban, M. , Hu, Y. , Simsekli, U. and Zhu, L. (2023). Cyclic and randomized stepsizes invoke heavier tails in SGD. Trans. Mach. Learn. Res. 2023."},{"key":"S0021900225100363_ref17","volume-title":"Probabilistic Machine Learning: An Introduction","author":"Murphy","year":"2022"},{"key":"S0021900225100363_ref8","doi-asserted-by":"publisher","DOI":"10.1090\/S0002-9947-05-03664-0"},{"key":"S0021900225100363_ref4","doi-asserted-by":"publisher","DOI":"10.1007\/s00440-008-0172-8"},{"key":"S0021900225100363_ref9","doi-asserted-by":"publisher","DOI":"10.1093\/oso\/9780198572237.001.0001"},{"key":"S0021900225100363_ref18","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9781107298019"},{"key":"S0021900225100363_ref22","doi-asserted-by":"publisher","DOI":"10.2307\/1426858"},{"key":"S0021900225100363_ref6","doi-asserted-by":"publisher","DOI":"10.1080\/10236198.2016.1234613"},{"key":"S0021900225100363_ref3","doi-asserted-by":"crossref","unstructured":"[3] Bougerol, P. and Lacroix, J. (1985). Products of Random Matrices with Applications to Schr\u00f6dinger Operators (Progress Prob. Statist. 8). Birkh\u00e4user, Boston.","DOI":"10.1007\/978-1-4684-9172-2"},{"key":"S0021900225100363_ref1","doi-asserted-by":"publisher","DOI":"10.1080\/10236198.2011.571383"},{"key":"S0021900225100363_ref14","unstructured":"[14] Hodgkinson, L. and Mahoney, M. (2021). Multiplicative noise and heavy tails in stochastic optimization. Proc. Mach. Learn. Res. 139, 4262\u20134274."},{"key":"S0021900225100363_ref16","first-page":"354","article-title":"A variational analysis of stochastic gradient algorithms","volume":"48","author":"Mandt","year":"2016","journal-title":"Proc, Mach. Learn. Res."},{"key":"S0021900225100363_ref5","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-29679-1"},{"key":"S0021900225100363_ref2","unstructured":"[2] Blanchet, J. , Mijatovi\u0107, A. and Yang, W. (2024). Limit theorems for stochastic gradient descent with infinite variance. Preprint, arXiv:2410.16340."},{"key":"S0021900225100363_ref13","first-page":"3964","article-title":"The heavy-tail phenomenon in SGD. Supplementary document","volume":"139","author":"G\u00fcrb\u00fczbalaban","year":"2021","journal-title":"Proc. Mach. Learn. Res."},{"key":"S0021900225100363_ref19","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2017.58"},{"key":"S0021900225100363_ref10","doi-asserted-by":"publisher","DOI":"10.1214\/15-AIHP668"},{"key":"S0021900225100363_ref20","unstructured":"[20] Smith, L. N. (2018). A disciplined approach to neural network hyper-parameters: Part 1 \u2013 learning rate, batch size, momentum, and weight decay. Preprint, arXiv:1803.09820."}],"container-title":["Journal of Applied Probability"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S0021900225100363","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T08:38:33Z","timestamp":1782290313000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S0021900225100363\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,10]]},"references-count":23,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["S0021900225100363"],"URL":"https:\/\/doi.org\/10.1017\/jpr.2025.10036","relation":{},"ISSN":["0021-9002","1475-6072"],"issn-type":[{"value":"0021-9002","type":"print"},{"value":"1475-6072","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,10]]},"assertion":[{"value":"\u00a9 The Author(s), 2026. Published by Cambridge University Press on behalf of Applied Probability Trust","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (https:\/\/creativecommons.org\/licenses\/by\/4.0\/), which permits unrestricted re-use, distribution, and reproduction in any medium, provided the original work is properly cited.","name":"license","label":"License","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This content has been made available to all.","name":"free","label":"Free to read"}]}}