{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,25]],"date-time":"2026-05-25T15:04:30Z","timestamp":1779721470677,"version":"3.53.1"},"reference-count":50,"publisher":"Oxford University Press (OUP)","issue":"3","license":[{"start":{"date-parts":[[2026,5,25]],"date-time":"2026-05-25T00:00:00Z","timestamp":1779667200000},"content-version":"vor","delay-in-days":24,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100002341","name":"Academy of Finland","doi-asserted-by":"publisher","award":["326291"],"award-info":[{"award-number":["326291"]}],"id":[{"id":"10.13039\/501100002341","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,5,4]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>High-dimensional genomic datasets contain complex patterns shaped by substantial biological noise, which pose major challenges for predictive modeling in genetics and breeding. Residual neural networks (ResNets) provide a powerful framework for capturing nonlinear genomic effects, but often overfit in settings where marker numbers greatly exceed sample sizes. As a solution, a range of regularization methods have been proposed. One promising approach relies on the proximal mapping technique, which is computationally efficient since it can be directly incorporated into the optimization algorithm. However, the performance of ResNets with various convex or non-convex proximal regularizers remains under-explored on high-dimensional data. In this study, we propose an extended stochastic adaptive proximal gradient ResNet method that can handle both convex and non-convex regularizers that range from $L_{0}$ to $L_{\\infty }$ and give more analysis of the convergence guarantee for the convex and non-convex regularizers. Moreover, we evaluate the prediction performance in a supervised regression setting on four real high-dimensional genomic datasets from mice, pig, wheat, and loblolly pine. For comparison, we also implement and evaluate traditional sparse linear proximal methods with the same regularizers, as well as LightGBM. Experimental results demonstrate that an 18-layer ResNet with $L_{\\frac{1}{2}}$ regularization outperforms other configurations on both mice and pig datasets. For the wheat and loblolly pine data, the 15-layer ResNet $L_{\\frac{1}{2}}$ configuration achieves the lowest test mean squared errors and the highest distance correlation (dCor). These findings highlight the effectiveness of the regularized adaptive proximal gradient ResNet method and its potential for prediction tasks on high-dimensional genomic data.<\/jats:p>","DOI":"10.1093\/bib\/bbag246","type":"journal-article","created":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T12:14:20Z","timestamp":1777637660000},"source":"Crossref","is-referenced-by-count":0,"title":["Proximal regularization of deep residual neural networks applied to high-dimensional genomic data"],"prefix":"10.1093","volume":"27","author":[{"given":"Yuhua","family":"Fan","sequence":"first","affiliation":[{"name":"Research Unit of Mathematical Sciences, University of Oulu , Pentti Kaiteran katu 1, 90570 Oulu,","place":["Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ilkka","family":"Launonen","sequence":"additional","affiliation":[{"name":"Research Unit of Mathematical Sciences, University of Oulu , Pentti Kaiteran katu 1, 90570 Oulu,","place":["Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mikko J","family":"Sillanp\u00e4\u00e4","sequence":"additional","affiliation":[{"name":"Research Unit of Mathematical Sciences, University of Oulu , Pentti Kaiteran katu 1, 90570 Oulu,","place":["Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Patrik","family":"Waldmann","sequence":"additional","affiliation":[{"name":"Research Unit of Mathematical Sciences, University of Oulu , Pentti Kaiteran katu 1, 90570 Oulu,","place":["Finland"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2026,5,25]]},"reference":[{"key":"2026052510355932300_ref1","doi-asserted-by":"crossref","first-page":"756","DOI":"10.1016\/j.tplants.2024.12.009","article-title":"Expanding genomic prediction in plant breeding: harnessing big data, machine learning, and advanced software","volume":"30","author":"Crossa","year":"2025","journal-title":"Trends Plant Sci"},{"key":"2026052510355932300_ref2","first-page":"770","article-title":"Deep residual learning for image recognition","volume-title":"Proceedings of the IEEE conference on Computer Vision and Pattern Recognition","author":"He","year":"2016"},{"key":"2026052510355932300_ref3","author":"Bondarenko","year":"2021"},{"key":"2026052510355932300_ref4","doi-asserted-by":"crossref","first-page":"bbab530","DOI":"10.1093\/bib\/bbab530","article-title":"Prediction of disease-associated nsSNPs by integrating multi-scale ResNet models with deep feature fusion","volume":"23","author":"Ge","year":"2022","journal-title":"Brief Bioinform"},{"key":"2026052510355932300_ref5","author":"Zagoruyko","year":"2016"},{"key":"2026052510355932300_ref6","first-page":"1492","article-title":"Aggregated residual transformations for deep neural networks","volume-title":"Proceedings of the IEEE conference on Computer Vision and Pattern Recognition","author":"Xie","year":"2017"},{"key":"2026052510355932300_ref7","first-page":"5927","article-title":"Deep pyramidal residual networks","volume-title":"Proceedings of the IEEE conference on Computer Vision and Pattern Recognition","author":"Han","year":"2017"},{"key":"2026052510355932300_ref8","doi-asserted-by":"crossref","first-page":"1979","DOI":"10.1109\/TPAMI.2018.2858821","article-title":"Virtual adversarial training: a regularization method for supervised and semi-supervised learning","volume":"41","author":"Miyato","year":"2018","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"2026052510355932300_ref9","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3510413","article-title":"Avoiding overfitting: a survey on regularization methods for convolutional neural networks","volume":"54","author":"Santos","year":"2022","journal-title":"ACM Comput Surv"},{"key":"2026052510355932300_ref10","article-title":"Learning structured sparsity in deep neural networks","volume-title":"30th Conference on Neural Information Processing Systems, 29","author":"Wei","year":"2016"},{"key":"2026052510355932300_ref11","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1186\/s40537-019-0197-0","article-title":"A survey on image data augmentation for deep learning","volume":"6","author":"Shorten","year":"2019","journal-title":"J Big Data"},{"key":"2026052510355932300_ref12","doi-asserted-by":"crossref","first-page":"53","DOI":"10.1007\/978-3-642-35289-8_5","article-title":"Early stopping\u2014but when?","volume-title":"Neural Networks: Tricks of the Trade: Second Edition","author":"Prechelt","year":"2012"},{"key":"2026052510355932300_ref13","article-title":"When does label smoothing help?","author":"M\u00fcller","journal-title":"33th Conference on Neural Information Processing Systems"},{"key":"2026052510355932300_ref14","first-page":"1929","article-title":"Dropout: a simple way to prevent neural networks from overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J Mach Learn Res"},{"key":"2026052510355932300_ref15","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1145\/3446776","article-title":"Understanding deep learning (still) requires rethinking generalization","volume":"64","author":"Zhang","year":"2021","journal-title":"Commun ACM"},{"key":"2026052510355932300_ref16","article-title":"A simple weight decay can improve generalization","author":"Krogh","journal-title":"5th Conference of Neural Information Processing Systems"},{"key":"2026052510355932300_ref17","volume-title":"32nd International Conference on Machine Learning","author":"Ioffe"},{"key":"2026052510355932300_ref18","article-title":"How does batch normalization help optimization?","volume-title":"32nd International Conference on Neural Information Processing System","author":"Santurkar"},{"key":"2026052510355932300_ref19","doi-asserted-by":"crossref","first-page":"499","DOI":"10.1038\/nrg3012","article-title":"Genome-wide genetic marker discovery and genotyping using next-generation sequencing","volume":"12","author":"Davey","year":"2011","journal-title":"Nat Rev Genet"},{"key":"2026052510355932300_ref20","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1534\/genetics.112.143313","article-title":"Whole-genome regression and prediction methods applied to plant and animal breeding","volume":"193","author":"de Los","year":"2013","journal-title":"Genetics"},{"key":"2026052510355932300_ref21","doi-asserted-by":"crossref","first-page":"iyae161","DOI":"10.1093\/genetics\/iyae161","article-title":"A review of multimodal deep learning methods for genomic-enabled prediction in plant breeding","volume":"228","author":"Montesinos-Lopez","year":"2024","journal-title":"Genetics"},{"key":"2026052510355932300_ref22","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1017\/S0016672310000662","article-title":"Prediction of body mass index in mice using dense molecular markers and a regularized neural network","volume":"93","author":"Okut","year":"2011","journal-title":"Genet Res (Camb)"},{"key":"2026052510355932300_ref23","article-title":"LightGBM: a highly efficient gradient boosting decision tree","volume-title":"31st International Conference on Neural Information Processing Systems","author":"Ke"},{"key":"2026052510355932300_ref24","article-title":"Adaptive proximal gradient methods for structured neural networks","volume-title":"35th International Conference on Neural Information Processing System","author":"Yun"},{"key":"2026052510355932300_ref25","doi-asserted-by":"crossref","first-page":"52665","DOI":"10.1109\/ACCESS.2024.3386093","article-title":"Evaluation of sparse proximal multi-task learning for genome-wide prediction","volume":"12","author":"Fan","year":"2024","journal-title":"IEEE Access"},{"key":"2026052510355932300_ref26","author":"Hinton","year":"2012"},{"key":"2026052510355932300_ref27","article-title":"L1 and L2 regularized deep residual network model for automated detection of myocardial infarction (heart attack) using electrocardiogram signals","volume-title":"CIKM Workshops","author":"Ukil","year":"2021"},{"key":"2026052510355932300_ref28","first-page":"499","article-title":"Stability and generalization","volume":"2","author":"Bousquet","year":"2002","journal-title":"J Mach Learn Res"},{"key":"2026052510355932300_ref29","first-page":"8757","article-title":"A proximal gradient method for regularized deep neural networks","volume-title":"42nd Chinese Control Conference (CCC)","author":"Liang","year":"2023"},{"key":"2026052510355932300_ref30","first-page":"2121","article-title":"Adaptive subgradient methods for online learning and stochastic optimization","volume":"12","author":"Duchi","year":"2011","journal-title":"J Mach Learn Res"},{"key":"2026052510355932300_ref31","article-title":"RMSprop: divide the gradient by a running average of its recent magnitude. Coursera: Neural networks for machine learning","volume":"17","author":"Tieleman","year":"2012","journal-title":"COURSERA Neural Networks Mach Learn"},{"key":"2026052510355932300_ref32","article-title":"Adam: a method for stochastic optimization","volume-title":"3rd International Conference on Learning Representations","author":"Kingma"},{"key":"2026052510355932300_ref33","first-page":"1762","article-title":"Diagonal preconditioning for first order primal-dual algorithms in convex optimization","volume-title":"International Conference on Computer Vision","author":"Pock","year":"2011"},{"key":"2026052510355932300_ref34","doi-asserted-by":"crossref","first-page":"1399","DOI":"10.1016\/j.mri.2013.05.010","article-title":"Smoothly clipped absolute deviation (SCAD) regularization for compressed sensing MRI using an augmented Lagrangian scheme","volume":"31","author":"Mehranian","year":"2013","journal-title":"Magn Reson Imaging"},{"key":"2026052510355932300_ref35","doi-asserted-by":"crossref","first-page":"894","DOI":"10.1214\/09-AOS729","article-title":"Nearly unbiased variable selection under minimax concave penalty","volume":"38","author":"Zhang","year":"2010","journal-title":"Ann Stat"},{"key":"2026052510355932300_ref36","doi-asserted-by":"crossref","first-page":"483","DOI":"10.1534\/genetics.114.164442","article-title":"Genome-wide regression and prediction with the BGLR statistical package","volume":"198","author":"P\u00e9rez","year":"2014","journal-title":"Genetics"},{"key":"2026052510355932300_ref37","doi-asserted-by":"crossref","first-page":"611","DOI":"10.1534\/genetics.108.088575","article-title":"Performance of genomic selection in mice","volume":"180","author":"Legarra","year":"2008","journal-title":"Genetics"},{"key":"2026052510355932300_ref38","doi-asserted-by":"crossref","first-page":"429","DOI":"10.1534\/g3.111.001453","article-title":"A common dataset for genomic analysis of livestock populations","volume":"2","author":"Cleveland","year":"2012","journal-title":"G3: Genes\u2014 Genomes\u2014 Genetics"},{"key":"2026052510355932300_ref39","doi-asserted-by":"crossref","first-page":"1503","DOI":"10.1534\/genetics.111.137026","article-title":"Accuracy of genomic selection methods in a standard data set of loblolly pine (Pinus taeda L.)","volume":"190","author":"Resende","year":"2012","journal-title":"Genetics"},{"key":"2026052510355932300_ref40","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1139\/x06-203","article-title":"Genetic analysis of early field growth of loblolly pine clones and seedlings from the same full-sib families","volume":"37","author":"Baltunis","year":"2006","journal-title":"Can J For Res"},{"key":"2026052510355932300_ref41","doi-asserted-by":"crossref","first-page":"969","DOI":"10.1534\/genetics.110.115543","article-title":"Patterns of population structure and environmental associations to aridity across the range of loblolly pine (P. taeda L., Pinaceae)","volume":"185","author":"Eckert","year":"2010","journal-title":"Genetics"},{"key":"2026052510355932300_ref42","doi-asserted-by":"crossref","first-page":"322","DOI":"10.1186\/s12859-024-05940-1","article-title":"Tabular deep learning: a comparative study applied to multi-task genome-wide prediction","volume":"25","author":"Fan","year":"2024","journal-title":"BMC Bioinformatics"},{"key":"2026052510355932300_ref43","article-title":"Why do tree-based models still outperform deep learning on tabular data?","volume-title":"36th International Conference on Neural Information Processing Systems","author":"L\u00e9o","year":"2022"},{"key":"2026052510355932300_ref44","doi-asserted-by":"crossref","first-page":"S11","DOI":"10.1186\/1753-6561-5-S3-S11","article-title":"A comparison of random forests, boosting and support vector machines for genomic selection","volume":"5","author":"Ogutu","year":"2011","journal-title":"BMC Proc"},{"key":"2026052510355932300_ref45","article-title":"Bayesian optimization","volume-title":"Recent advances in optimization and modeling of contemporary problems","author":"Peter","year":"2018"},{"key":"2026052510355932300_ref46","doi-asserted-by":"crossref","first-page":"20","DOI":"10.25080\/Majora-8b375195-004","article-title":"Hyperopt: a python library for optimizing the hyperparameters of machine learning algorithms","volume":"13","author":"Bergstra","year":"2013","journal-title":"SciPy"},{"key":"2026052510355932300_ref47","doi-asserted-by":"crossref","first-page":"80","DOI":"10.2307\/3001968","article-title":"Individual comparisons by ranking methods","volume":"1","author":"Wilcoxon","year":"1945","journal-title":"Biom Bull"},{"key":"2026052510355932300_ref48","first-page":"1","article-title":"Statistical comparisons of classifiers over multiple data sets","volume":"7","author":"Dem\u0161ar","year":"2006","journal-title":"J Mach Learn Res"},{"key":"2026052510355932300_ref49","doi-asserted-by":"crossref","DOI":"10.1007\/978-0-387-84858-7","volume-title":"The Elements of Statistical Learning: Data Mining, Inference, and Prediction 2nd edn","author":"Hastie","year":"2009"},{"key":"2026052510355932300_ref50","volume-title":"Deep Learning","author":"Goodfellow","year":"2016"}],"container-title":["Briefings in Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bib\/article-pdf\/27\/3\/bbag246\/68381879\/bbag246.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bib\/article-pdf\/27\/3\/bbag246\/68381879\/bbag246.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,25]],"date-time":"2026-05-25T14:36:10Z","timestamp":1779719770000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bib\/article\/doi\/10.1093\/bib\/bbag246\/8692743"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5]]},"references-count":50,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2026,5,4]]}},"URL":"https:\/\/doi.org\/10.1093\/bib\/bbag246","relation":{},"ISSN":["1467-5463","1477-4054"],"issn-type":[{"value":"1467-5463","type":"print"},{"value":"1477-4054","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026,5]]},"published":{"date-parts":[[2026,5]]},"article-number":"bbag246"}}