{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T19:03:07Z","timestamp":1757617387520,"version":"3.44.0"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2025,1,9]],"date-time":"2025-01-09T00:00:00Z","timestamp":1736380800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,1,9]],"date-time":"2025-01-09T00:00:00Z","timestamp":1736380800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"publisher","award":["450330162"],"award-info":[{"award-number":["450330162"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Comput Stat"],"published-print":{"date-parts":[[2025,9]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Due to the popularity of deep learning models there have recently been many attempts to translate generalized additive models to neural nets. Generalized additive models are usually regularized by a penalty in the loss function and the magnitude of penalization is controlled by one or more smoothing parameters. In the statistical literature these smoothing parameters are estimated by criteria such as generalized cross-validation or restricted maximum likelihood. While the estimation of the primary regression coefficients is well calibrated and investigated for neural net based additive models, the estimation of smoothing parameters is often either based on testing data (and grid search), implicitly estimated or completely neglected. In this paper, we address the issue of explicit smoothing parameter estimation in neural net-based additive models fitted via gradient-based methods, such as the well-known Adam algorithm. We therefore investigate the data-driven smoothing parameter selection via gradient-based optimization of generalized cross-validation and restricted maximum likelihood. Thus we do not need to calculate Hessian information of the smoothing parameters. As an additive model structure, we use a translation of P-splines to neural nets, so-called neural P-splines. The fitting process of neural P-splines as well as the gradient-based smoothing parameter selection are investigated in a simulation study and an application.<\/jats:p>","DOI":"10.1007\/s00180-024-01593-z","type":"journal-article","created":{"date-parts":[[2025,1,9]],"date-time":"2025-01-09T06:18:21Z","timestamp":1736403501000},"page":"3645-3663","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Gradient-based smoothing parameter estimation for neural P-splines"],"prefix":"10.1007","volume":"40","author":[{"given":"Lea M.","family":"Dammann","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0920-2535","authenticated-orcid":false,"given":"Marei","family":"Freitag","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anton","family":"Thielmann","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Benjamin","family":"S\u00e4fken","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,1,9]]},"reference":[{"key":"1593_CR1","unstructured":"Abadi Mart\u00edn, Agarwal Ashish, Barham Paul, Brevdo Eugene, Chen Zhifeng, Citro Craig, Corrado Greg\u00a0S, Davis Andy, Dean Jeffrey, Devin Matthieu, Ghemawat Sanjay, Goodfellow Ian, Harp Andrew, Irving Geoffrey, Isard Michael, Jia Yangqing, Jozefowicz Rafal, Kaiser Lukasz, Kudlur Manjunath, Levenberg Josh, Man\u00e9 Dandelion, Monga Rajat, Moore Sherry, Murray Derek, Olah Chris, Schuster Mike, Shlens Jonathon, Steiner Benoit, Sutskever Ilya, Talwar Kunal, Tucker Paul, Vanhoucke Vincent, Vasudevan Vijay, Vi\u00e9gas Fernanda, Vinyals Oriol, Warden Pete, Wattenberg Martin, Wicke Martin, Yu Yuan, Zheng Xiaoqiang (2015) TensorFlow: Large-scale machine learning on heterogeneous systems, URL https:\/\/www.tensorflow.org\/. Software available from tensorflow.org"},{"key":"1593_CR2","unstructured":"Agarwal Rishabh, Melnick Levi, Frosst Nicholas, Zhang Xuezhou, Lengerich Ben, Caruana Rich, Hinton Geoffrey (2021) Neural additive models: Interpretable machine learning with neural nets. In A.\u00a0Beygelzimer, Y.\u00a0Dauphin, P.\u00a0Liang, and J.\u00a0Wortman Vaughan, editors, Advances in Neural Information Processing Systems, URL https:\/\/openreview.net\/forum?id=wHkKTW2wrmm"},{"key":"1593_CR3","unstructured":"Chang Chun-Hao, Caruana Rich, Goldenberg Anna (2021) Node-gam: Neural generalized additive model for interpretable deep learning. arXiv preprint arXiv:2106.01613"},{"issue":"4","key":"1593_CR4","doi-asserted-by":"publisher","first-page":"377","DOI":"10.1007\/BF01404567","volume":"31","author":"P Craven","year":"1979","unstructured":"Craven P, Wahba G (1979) Smoothing noisy data with spline functions. Numer Math 31(4):377\u2013403 (ISSN 0029-599X)","journal-title":"Numer Math"},{"key":"1593_CR5","unstructured":"Dubey Abhimanyu, Radenovic Filip, Mahajan Dhruv (2022) Scalable interpretability via polynomials. arXiv preprint arXiv:2205.14108"},{"key":"1593_CR6","unstructured":"Duchi John, Hazan Elad, Singer Yoram (2011) Adaptive subgradient methods for online learning and stochastic optimization. Journal of machine learning research, 12 (7)"},{"issue":"2","key":"1593_CR7","doi-asserted-by":"publisher","first-page":"89","DOI":"10.1214\/ss\/1038425655","volume":"11","author":"HC Eilers Paul","year":"1996","unstructured":"Eilers Paul HC, Marx Brian D (1996) Flexible smoothing with B -splines and penalties. Stat Sci 11(2):89\u2013121. https:\/\/doi.org\/10.1214\/ss\/1038425655. (ISSN 0883-4237)","journal-title":"Stat Sci"},{"key":"1593_CR8","unstructured":"Enouen James, Liu Yan (2022) Sparse interaction additive networks via feature interaction detection and sparse selection. arXiv preprint arXiv:2209.09326"},{"key":"1593_CR9","doi-asserted-by":"crossref","unstructured":"Fahrmeir Ludwig, Kneib Thomas, Lang Stefan, Marx Brian\u00a0D (2013) Regression Models, Methods and Applications. Springer, 2 edition","DOI":"10.1007\/978-3-642-34333-9"},{"issue":"535","key":"1593_CR10","doi-asserted-by":"publisher","first-page":"1402","DOI":"10.1080\/01621459.2020.1725521","volume":"116","author":"M Fasiolo","year":"2021","unstructured":"Fasiolo M, Wood SN, Zaffran M, Nedellec R, Goude Y (2021) Fast calibrated additive quantile regression. J Am Stat Assoc 116(535):1402\u20131412","journal-title":"J Am Stat Assoc"},{"issue":"12","key":"1593_CR11","doi-asserted-by":"publisher","first-page":"1683","DOI":"10.1016\/j.cad.2011.07.010","volume":"43","author":"Akemi G\u00e1lvez","year":"2011","unstructured":"G\u00e1lvez Akemi, Iglesias Andr\u00e9s (2011) Efficient particle swarm optimization approach for data fitting with free knot b-splines. Comput Aided Des 43(12):1683\u20131692","journal-title":"Comput Aided Des"},{"key":"1593_CR12","unstructured":"Goodfellow Ian, Bengio Yoshua, Courville Aaron (2016) Deep Learning. MIT Press, http:\/\/www.deeplearningbook.org"},{"volume-title":"Automated Machine Learning - Methods Systems, Challenges","year":"2019","key":"1593_CR13","unstructured":"Frank Hutter, Lars Kotthoff, Joaquin Vanschoren (eds) (2019) Automated Machine Learning - Methods Systems, Challenges. Springer, Berlin Heidelberg"},{"key":"1593_CR14","unstructured":"Karakida Ryo, Akaho Shotaro, Amari Shun-ichi (Apr 2019) Universal statistics of fisher information in deep neural networks: Mean field approach. In Kamalika Chaudhuri and Masashi Sugiyama, editors, Proceedings of the Twenty-Second International Conference on Artificial Intelligence and Statistics, volume\u00a089 of Proceedings of Machine Learning Research, pages 1032\u20131041. PMLR, 16\u201318. URL https:\/\/proceedings.mlr.press\/v89\/karakida19a.html"},{"key":"1593_CR15","unstructured":"Kim Minkyu, Choi Hyun-Soo, Kim Jinho (2022) Higher-order neural additive models: An interpretable machine learning model with feature interactions. arXiv preprint arXiv:2209.15409"},{"key":"1593_CR16","unstructured":"Kingma Diederik\u00a0P, Ba Jimmy (2015) Adam: A method for stochastic optimization, 2014. URL http:\/\/arxiv.org\/abs\/1412.6980. cite arxiv:1412.6980Comment: Published as a conference paper at the 3rd International Conference for Learning Representations, San Diego"},{"key":"1593_CR17","unstructured":"Kneib Thomas (Februar 2006) Mixed model based inference in structured additive regression. URL http:\/\/nbn-resolving.de\/urn:nbn:de:bvb:19-50112"},{"issue":"4","key":"1593_CR18","doi-asserted-by":"publisher","first-page":"725","DOI":"10.1111\/rssb.12010","volume":"75","author":"Tatyana Krivobokova","year":"2013","unstructured":"Krivobokova Tatyana (2013) Smoothing parameter selection in two frameworks for penalized splines. J R Stat Soc\u202f: Series B (Statistical Methodology) 75(4):725\u2013741","journal-title":"J R Stat Soc : Series B (Statistical Methodology)"},{"key":"1593_CR19","unstructured":"Kuka\u010dka Jan, Golkov Vladimir, Cremers Daniel (2017) Regularization for deep learning: A taxonomy. arXiv preprint arXiv:1710.10686"},{"key":"1593_CR20","doi-asserted-by":"crossref","unstructured":"LeCun Yann, Bottou Leon, Orr Genevieve\u00a0B, M\u00fcller Klaus\u00a0Robert (1998) Efficient BackProp, Springer Berlin Heidelberg, p 9\u201350 doi 10.1007\/3-540-49430-8_2","DOI":"10.1007\/3-540-49430-8_2"},{"issue":"7553","key":"1593_CR21","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Yann LeCun","year":"2015","unstructured":"LeCun Yann, Bengio Yoshua, Hinton Geoffrey (2015) Deep learning. Nature 521(7553):436\u2013444. https:\/\/doi.org\/10.1038\/nature14539","journal-title":"Nature"},{"key":"1593_CR22","unstructured":"Luber Mattias, Thielmann Anton, S\u00e4fken Benjamin (2023) Structural neural additive models: Enhanced interpretable machine learning. arXiv preprint arXiv:2302.09275"},{"key":"1593_CR23","doi-asserted-by":"publisher","first-page":"40","DOI":"10.1016\/j.jmp.2017.05.006","volume":"80","author":"Ly Alexander","year":"2017","unstructured":"Alexander Ly, Maarten Marsman, Josine Verhagen, Grasman Raoul PPP, Eric-Jan Wagenmakers (2017) A tutorial on Fisher information. J Math Psychol 80:40\u201355","journal-title":"J Math Psychol"},{"issue":"1","key":"1593_CR24","doi-asserted-by":"publisher","first-page":"155","DOI":"10.1007\/s00180-020-01022-x","volume":"36","author":"SD Mohanty","year":"2021","unstructured":"Mohanty SD, Fahnestock E (2021) Adaptive spline fitting with particle swarm optimization. Comput Statistics 36(1):155\u2013191","journal-title":"Comput Statistics"},{"key":"1593_CR25","unstructured":"Radenovic Filip, Dubey Abhimanyu, Mahajan Dhruv (2022) Neural basis models for interpretability. arXiv preprint arXiv:2205.14120"},{"key":"1593_CR26","doi-asserted-by":"publisher","first-page":"505","DOI":"10.1111\/j.1467-9868.2008.00695.x","volume":"71","author":"PT Reiss","year":"2009","unstructured":"Reiss PT, Ogden R (2009) Smoothing parameter selection for a class of semiparametric linear models. J R Stat Soc Series B (Methodological) 71:505\u2013523","journal-title":"J R Stat Soc Series B (Methodological)"},{"key":"1593_CR27","unstructured":"Reuter Arik, Thielmann Anton, S\u00e4fken Benjamin (2024) Neural additive image model: Interpretation through interpolation, URL https:\/\/arxiv.org\/abs\/2405.02295"},{"issue":"3","key":"1593_CR28","doi-asserted-by":"publisher","first-page":"507","DOI":"10.1111\/j.1467-9876.2005.00510.x","volume":"54","author":"RA Rigby","year":"2005","unstructured":"Rigby RA, Stasinopoulos DM (2005) Generalized additive models for location, scale and shape. J R Stat Soc\u202f: Series C (Applied Statistics) 54(3):507\u2013554","journal-title":"J R Stat Soc : Series C (Applied Statistics)"},{"key":"1593_CR29","unstructured":"R\u00fcgamer David (2023) mixdistreg: An r package for fitting mixture of experts distributional regression with adaptive first-order methods. arXiv preprint arXiv:2302.02043"},{"issue":"1","key":"1593_CR30","first-page":"1","volume":"78","author":"D R\u00fcgamer","year":"2023","unstructured":"R\u00fcgamer D, Kolb C, Klein N (2023) Semi-structured distributional regression. Am Stat 78(1):1\u201312","journal-title":"Am Stat"},{"issue":"6088","key":"1593_CR31","doi-asserted-by":"publisher","first-page":"533","DOI":"10.1038\/323533a0","volume":"323","author":"E Rumelhart David","year":"1986","unstructured":"Rumelhart David E, Hinton Geoffrey E, Williams Ronald J (1986) Learning representations by back-propagating errors. Nature 323(6088):533\u2013536","journal-title":"Nature"},{"key":"1593_CR32","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/asae052","author":"B S\u00e4fken","year":"2024","unstructured":"S\u00e4fken B, Kneib T, Wood SN (2024) On the degrees of freedom of the smoothing parameter. Biometrika. https:\/\/doi.org\/10.1093\/biomet\/asae052","journal-title":"Biometrika"},{"key":"1593_CR33","doi-asserted-by":"crossref","unstructured":"Seifert Quentin\u00a0Edward, Thielmann Anton, Bergherr Elisabeth, S\u00e4fken Benjamin, Zierk Jakob, Rauh Manfred, Hepp Tobias (2022) Penalized regression splines in mixture density networks","DOI":"10.21203\/rs.3.rs-2398185\/v1"},{"key":"1593_CR34","first-page":"5708","volume":"34","author":"Soen Alexander","year":"2021","unstructured":"Alexander Soen, Ke Sun (2021) On the variance of the fisher information for deep learning. Adv Neural Inf Process Syst 34:5708\u20135719","journal-title":"Adv Neural Inf Process Syst"},{"issue":"1","key":"1593_CR35","first-page":"1929","volume":"15","author":"Nitish Srivastava","year":"2014","unstructured":"Srivastava Nitish, Hinton Geoffrey, Krizhevsky Alex, Sutskever Ilya, Salakhutdinov Ruslan (2014) Dropout: a simple way to prevent neural networks from overfitting. J Mach Learn Res 15(1):1929\u20131958","journal-title":"J Mach Learn Res"},{"key":"1593_CR36","unstructured":"Thielmann Anton\u00a0Frederik, Kruse Ren\u00e9-Marcel, Kneib Thomas,\u00e4fken Benjamin S (2024a) Neural additive models for location scale and shape: A framework for interpretable neural regression beyond the mean. In International Conference on Artificial Intelligence and Statistics, p 1783\u20131791. PMLR"},{"key":"1593_CR37","unstructured":"Thielmann AF, Kumar M, Weisser C, Reuter A, S\u00e4fken B, Samiee S  (2024b) Mambular: A sequential model for tabular deep learning. URL https:\/\/arxiv.org\/abs\/2408.06291"},{"key":"1593_CR38","unstructured":"Thielmann AF, Reuter A, Kneib T, R\u00fcgamer D, S\u00e4fken B (2024) Interpretable additive tabular transformer networks, Transactions on machine learning research"},{"key":"1593_CR39","doi-asserted-by":"crossref","unstructured":"Umlauf Nikolaus, Seiler Johannes, Wetscher Mattias, Simon Thorsten, Lang Stefan, Klein Nadja (2023) Scalable estimation for structured additive distributional regression. arXiv preprint arXiv:2301.05593","DOI":"10.1080\/10618600.2024.2388604"},{"key":"1593_CR40","doi-asserted-by":"publisher","first-page":"1378","DOI":"10.1214\/aos\/1176349743","volume":"13","author":"G Wahba","year":"1985","unstructured":"Wahba G (1985) A comparison of GCV and GML for choosing the smoothing parameter in the generalized spline smoothing problem. Ann Statist 13:1378\u20131402","journal-title":"Ann Statist"},{"key":"1593_CR41","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1111\/j.1467-9868.2010.00749.x","volume":"73","author":"SN Wood","year":"2011","unstructured":"Wood SN (2011) Fast stable restricted maximum likelihood and marginal likelihood estimation of semiparametric generalized linear models. J R Stat Soc\u202f: Series B (Statistical Methodology) 73:3\u201336 (ISSN 1369-7412)","journal-title":"J R Stat Soc : Series B (Statistical Methodology)"},{"issue":"519","key":"1593_CR42","doi-asserted-by":"publisher","first-page":"1199","DOI":"10.1080\/01621459.2016.1195744","volume":"112","author":"Simon\u00a0N Wood","year":"2017","unstructured":"Wood Simon\u00a0N, Li Zheyuan, Shaddick Gavin, Augustin Nicole\u00a0H (2017) Generalized additive models for gigadata: Modeling the u.k. black smoke network daily data. J Am Stat Assoc 112(519):1199\u20131210. https:\/\/doi.org\/10.1080\/01621459.2016.1195744","journal-title":"J Am Stat Assoc"},{"key":"1593_CR43","doi-asserted-by":"crossref","unstructured":"Wood S.N (2017) Generalized Additive Models: An Introduction with R. Chapman and Hall\/CRC, 2 edition","DOI":"10.1201\/9781315370279"}],"container-title":["Computational Statistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00180-024-01593-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00180-024-01593-z\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00180-024-01593-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,6]],"date-time":"2025-09-06T03:14:58Z","timestamp":1757128498000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00180-024-01593-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,9]]},"references-count":43,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2025,9]]}},"alternative-id":["1593"],"URL":"https:\/\/doi.org\/10.1007\/s00180-024-01593-z","relation":{},"ISSN":["0943-4062","1613-9658"],"issn-type":[{"type":"print","value":"0943-4062"},{"type":"electronic","value":"1613-9658"}],"subject":[],"published":{"date-parts":[[2025,1,9]]},"assertion":[{"value":"6 March 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"14 December 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 January 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 February 2025","order":4,"name":"change_date","label":"Change Date","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Update","order":5,"name":"change_type","label":"Change Type","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"In this article, the year in the funding statement was incorrect. This has been corrected.","order":6,"name":"change_details","label":"Change Details","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors certify that they have no affiliations with or involvement in any organization or entity with any financial or non-financial interest that are directly or indirectly related to the work discussed in this manuscript.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interests"}}]}}