{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T13:45:16Z","timestamp":1782913516836,"version":"3.54.5"},"reference-count":43,"publisher":"Oxford University Press (OUP)","issue":"6","license":[{"start":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T00:00:00Z","timestamp":1781654400000},"content-version":"vor","delay-in-days":16,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["P01CA196569"],"award-info":[{"award-number":["P01CA196569"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["U01CA261339"],"award-info":[{"award-number":["U01CA261339"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["5P30CA014089-47"],"award-info":[{"award-number":["5P30CA014089-47"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2026,6,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Motivation<\/jats:title>\n                    <jats:p>High-dimensional omics data are typically measured on limited sample sizes, which challenges model-based clustering methods such as Gaussian mixture models (GMMs), often leading to instability and poor generalization under complex mixture structures. To address these limitations, we developed Praxis-BGM, a natural-gradient variational inference framework for GMMs. Praxis-BGM enables semi-supervised transfer learning by incorporating an informative prior GMM estimated from large-scale reference data with robust cluster structures. The prior model can encode cluster-specific means, covariance structures, and structural connectivity patterns, and is updated using the target data with variational inference to improve clustering in small-sample settings.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>Using the Variational Online Newton (VON) algorithm, we derived natural-gradient updates for the standard parameters of GMMs. Implemented in the Python library JAX for accelerator-oriented computation, Praxis-BGM is computationally efficient and scalable. Across extensive simulations and two real-world applications\u2014breast cancer bulk transcriptomics for subtype recovery and single-cell transcriptomics for cross-platform cell-type label transfer\u2014Praxis-BGM improves posterior clustering performance, stability, and biological interpretability, even when priors are partially mismatched.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Availability and implementation<\/jats:title>\n                    <jats:p>Praxis-BGM is freely available at https:\/\/github.com\/ContiLab-usc\/Praxis-BGM, and an archival version is available on Zenodo at https:\/\/doi.org\/10.5281\/zenodo.19657680.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btag395","type":"journal-article","created":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T12:32:23Z","timestamp":1781613143000},"source":"Crossref","is-referenced-by-count":0,"title":["Praxis-BGM: clustering of omics data using semi-supervised transfer learning for Gaussian mixture models via natural-gradient variational inference"],"prefix":"10.1093","volume":"42","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0790-5967","authenticated-orcid":false,"given":"Qiran","family":"Jia","sequence":"first","affiliation":[{"name":"Division of Biostatistics and Health Data Science, Department of Population and Public Health Sciences, Keck School of Medicine, University of Southern California , Los Angeles, CA 90033,","place":["United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6615-0472","authenticated-orcid":false,"given":"Jesse A","family":"Goodrich","sequence":"additional","affiliation":[{"name":"Division of Environmental Health, Department of Population and Public Health Sciences, Keck School of Medicine, University of Southern California , Los Angeles, CA 90033,","place":["United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2941-7833","authenticated-orcid":false,"given":"David V","family":"Conti","sequence":"additional","affiliation":[{"name":"Division of Biostatistics and Health Data Science, Department of Population and Public Health Sciences, Keck School of Medicine, University of Southern California , Los Angeles, CA 90033,","place":["United States"]},{"name":"Department of Biostatistics and Informatics, Colorado School of Public Health , Aurora, CO 80045,","place":["United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2026,6,17]]},"reference":[{"key":"2026063023013928900_btag395-B1","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1162\/089976698300017746","article-title":"Natural gradient works efficiently in learning","volume":"10","author":"Amari","year":"1998","journal-title":"Neural Comput"},{"key":"2026063023013928900_btag395-B2","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1038\/75556","article-title":"Gene ontology: tool for the unification of biology","volume":"25","author":"Ashburner","year":"2000","journal-title":"Nat Genet"},{"key":"2026063023013928900_btag395-B3","doi-asserted-by":"publisher","first-page":"1122","DOI":"10.7150\/thno.11543","article-title":"Micrornas: new biomarkers for diagnosis, prognosis, therapy prediction and therapeutic tools for breast cancer","volume":"5","author":"Bertoli","year":"2015","journal-title":"Theranostics"},{"key":"2026063023013928900_btag395-B4","doi-asserted-by":"publisher","first-page":"859","DOI":"10.1080\/01621459.2017.1285773","article-title":"Variational inference: a review for statisticians","volume":"112","author":"Blei","year":"2017","journal-title":"J Am Stat Assoc"},{"key":"2026063023013928900_btag395-B5","volume-title":"JAX: Composable Transformations of Python Numpy Programs","author":"Bradbury"},{"key":"2026063023013928900_btag395-B6","doi-asserted-by":"publisher","first-page":"401","DOI":"10.1158\/2159-8290.CD-12-0095","article-title":"The CBIO cancer genomics portal: an open platform for exploring multidimensional cancer genomics data","volume":"2","author":"Cerami","year":"2012","journal-title":"Cancer Discov"},{"key":"2026063023013928900_btag395-B7","doi-asserted-by":"publisher","author":"Chen","year":"2016","DOI":"10.1145\/2939672.2939785"},{"key":"2026063023013928900_btag395-B8","doi-asserted-by":"publisher","first-page":"D1779","DOI":"10.1093\/nar\/gkaf1292","article-title":"The gene ontology knowledgebase in 2026","volume":"54","author":"Consortium","year":"2026","journal-title":"Nucleic Acids Res"},{"key":"2026063023013928900_btag395-B9","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1038\/nature10983","article-title":"The genomic and transcriptomic architecture of 2,000 breast tumours reveals novel subgroups","volume":"486","author":"Curtis","year":"2012","journal-title":"Nature"},{"key":"2026063023013928900_btag395-B10","doi-asserted-by":"publisher","first-page":"3861","DOI":"10.1158\/0008-5472.CAN-23-0816","article-title":"Analysis and visualization of longitudinal genomic and clinical data from the AACR project GENIE biopharma collaborative in cBioPortal","volume":"83","author":"de Bruijn","year":"2023","journal-title":"Cancer Res"},{"key":"2026063023013928900_btag395-B11","doi-asserted-by":"publisher","first-page":"btac757","DOI":"10.1093\/bioinformatics\/btac757","article-title":"Gseapy: a comprehensive package for performing gene set enrichment analysis in python","volume":"39","author":"Fang","year":"2023","journal-title":"Bioinformatics"},{"key":"2026063023013928900_btag395-B12","doi-asserted-by":"publisher","first-page":"611","DOI":"10.1198\/016214502760047131","article-title":"Model-based clustering, discriminant analysis, and density estimation","volume":"97","author":"Fraley","year":"2002","journal-title":"J Am Stat Assoc"},{"key":"2026063023013928900_btag395-B13","doi-asserted-by":"publisher","first-page":"pl1","DOI":"10.1126\/scisignal.2004088","article-title":"Integrative analysis of complex cancer genomics and clinical profiles using the cBioPortal","volume":"6","author":"Gao","year":"2013","journal-title":"Sci Signal"},{"key":"2026063023013928900_btag395-B14","doi-asserted-by":"publisher","first-page":"108930","DOI":"10.1016\/j.envint.2024.108930","article-title":"Integrating multi-omics with environmental data for precision health: a novel analytic framework and case study on prenatal mercury induced childhood fatty liver disease","volume":"190","author":"Goodrich","year":"2024","journal-title":"Environ Int"},{"key":"2026063023013928900_btag395-B15","doi-asserted-by":"publisher","first-page":"3573","DOI":"10.1016\/j.cell.2021.04.048","article-title":"Integrated analysis of multimodal single-cell data","volume":"184","author":"Hao","year":"2021","journal-title":"Cell"},{"key":"2026063023013928900_btag395-B16","doi-asserted-by":"publisher","first-page":"2283","DOI":"10.1038\/s41596-024-00991-3","article-title":"Scanorama: integrating large and diverse single-cell transcriptomic datasets","volume":"19","author":"Hie","year":"2024","journal-title":"Nat Protoc"},{"key":"2026063023013928900_btag395-B17"},{"key":"2026063023013928900_btag395-B18","doi-asserted-by":"publisher","author":"Khan","year":"2018","DOI":"10.23919\/ISITA.2018.8664326"},{"key":"2026063023013928900_btag395-B19","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1111\/rssb.12479","article-title":"Transfer learning for high-dimensional linear regression: prediction, estimation and minimax optimality","volume":"84","author":"Li","year":"2022","journal-title":"J R Stat Soc Series B Stat Methodol"},{"key":"2026063023013928900_btag395-B20","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1016\/j.cels.2015.12.004","article-title":"The molecular signatures database (MSIGDB) hallmark gene set collection","volume":"1","author":"Liberzon","year":"2015","journal-title":"Cell Syst"},{"key":"2026063023013928900_btag395-B21","first-page":"3992","author":"Lin","year":"2019"},{"key":"2026063023013928900_btag395-B22","doi-asserted-by":"publisher","first-page":"1053","DOI":"10.1038\/s41592-018-0229-2","article-title":"Deep generative modeling for single-cell transcriptomics","volume":"15","author":"Lopez","year":"2018","journal-title":"Nat Methods"},{"key":"2026063023013928900_btag395-B23","doi-asserted-by":"publisher","first-page":"41","DOI":"10.1038\/s41592-021-01336-8","article-title":"Benchmarking atlas-level data integration in single-cell genomics","volume":"19","author":"Luecken","year":"2022","journal-title":"Nat Methods"},{"key":"2026063023013928900_btag395-B24","author":"Mahdisoltani","year":"2021"},{"key":"2026063023013928900_btag395-B25","doi-asserted-by":"crossref","first-page":"4245","DOI":"10.1073\/pnas.1208949110","article-title":"Pattern discovery and cancer gene identification in integrated cancer genomic data","volume":"110","author":"Mo","year":"2013","journal-title":"Proc Natl Acad Sci USA"},{"key":"2026063023013928900_btag395-B26","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1038\/nature11412","article-title":"Comprehensive molecular portraits of human breast tumours","volume":"490","author":"Network","year":"2012","journal-title":"Nature"},{"key":"2026063023013928900_btag395-B27","doi-asserted-by":"publisher","first-page":"1345","DOI":"10.1109\/TKDE.2009.191","article-title":"A survey on transfer learning","volume":"22","author":"Pan","year":"2010","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"2026063023013928900_btag395-B28","doi-asserted-by":"crossref","first-page":"1160","DOI":"10.1200\/JCO.2008.18.1370","article-title":"Supervised risk predictor of breast cancer based on intrinsic subtypes","volume":"27","author":"Parker","year":"2009","journal-title":"J Clin Oncol"},{"key":"2026063023013928900_btag395-B29","doi-asserted-by":"publisher","first-page":"961","DOI":"10.1016\/j.csbj.2021.01.015","article-title":"Automated methods for cell type annotation on scrna-seq data","volume":"19","author":"Pasquini","year":"2021","journal-title":"Comput Struct Biotechnol J"},{"key":"2026063023013928900_btag395-B30","first-page":"2825","article-title":"Scikit-learn: machine learning in python","volume":"12","author":"Pedregosa","year":"2011","journal-title":"J Mach Learn Res"},{"key":"2026063023013928900_btag395-B31","doi-asserted-by":"publisher","first-page":"747","DOI":"10.1038\/35021093","article-title":"Molecular portraits of human breast tumours","volume":"406","author":"Perou","year":"2000","journal-title":"Nature"},{"key":"2026063023013928900_btag395-B32","doi-asserted-by":"publisher","first-page":"958","DOI":"10.1200\/CCI.19.00119","article-title":"Multiomic integration of public oncology databases in bioconductor","volume":"4","author":"Ramos","year":"2020","journal-title":"JCO Clin Cancer Inform"},{"key":"2026063023013928900_btag395-B33","doi-asserted-by":"crossref","first-page":"10546","DOI":"10.1093\/nar\/gky889","article-title":"Multi-omic and multi-view clustering algorithms: review and cancer benchmark","volume":"46","author":"Rappoport","year":"2018","journal-title":"Nucleic Acids Res"},{"key":"2026063023013928900_btag395-B34","doi-asserted-by":"publisher","first-page":"477","DOI":"10.1214\/25-STS987","article-title":"Bayesian transfer learning","volume":"40","author":"Suder","year":"2025","journal-title":"Stat Sci"},{"key":"2026063023013928900_btag395-B35","doi-asserted-by":"publisher","first-page":"21","DOI":"10.1186\/1479-7364-4-1-21","article-title":"Use of pathway information in molecular epidemiology","volume":"4","author":"Thomas","year":"2009","journal-title":"Hum Genomics"},{"key":"2026063023013928900_btag395-B36","doi-asserted-by":"publisher","volume-title":"J Am Stat Assoc","DOI":"10.1080\/01621459.2026.2670031"},{"key":"2026063023013928900_btag395-B37","doi-asserted-by":"publisher","first-page":"20220149","DOI":"10.1098\/rsta.2022.0149","article-title":"Bayesian cluster analysis","volume":"381","author":"Wade","year":"2023","journal-title":"Philos Trans A Math Phys Eng Sci"},{"key":"2026063023013928900_btag395-B38","doi-asserted-by":"publisher","first-page":"7058","DOI":"10.1109\/TCYB.2022.3177242","article-title":"Transfer-learning-based gaussian mixture model for distributed clustering","volume":"53","author":"Wang","year":"2023","journal-title":"IEEE Trans Cybern"},{"key":"2026063023013928900_btag395-B39","author":"Wycoff","year":"2025"},{"key":"2026063023013928900_btag395-B40","doi-asserted-by":"publisher","first-page":"e9620","DOI":"10.15252\/msb.20209620","article-title":"Probabilistic harmonization and annotation of single-cell transcriptomics data with deep generative models","volume":"17","author":"Xu","year":"2021","journal-title":"Mol Syst Biol"},{"key":"2026063023013928900_btag395-B41","doi-asserted-by":"publisher","first-page":"116615","DOI":"10.1016\/j.biopha.2024.116615","article-title":"Targeting the crosstalk between estrogen receptors and membrane growth factor receptors in breast cancer treatment: advances and opportunities","volume":"175","author":"Yan","year":"2024","journal-title":"Biomed Pharmacother"},{"key":"2026063023013928900_btag395-B42","doi-asserted-by":"publisher","first-page":"lqaa078","DOI":"10.1093\/nargab\/lqaa078","article-title":"Combat-seq: batch effect adjustment for RNA-seq count data","volume":"2","author":"Zhang","year":"2020","journal-title":"NAR Genom Bioinform"},{"key":"2026063023013928900_btag395-B43","doi-asserted-by":"publisher","first-page":"4","DOI":"10.32614\/RJ-2024-012","article-title":"Lucidus: an R package for implementing latent unknown clustering by integrating multi-omics data (lucid) with phenotypic traits","volume":"16","author":"Zhao","year":"2025","journal-title":"R J"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btag395\/68542719\/btag395.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/42\/6\/btag395\/68542719\/btag395.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/42\/6\/btag395\/68542719\/btag395.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T12:53:08Z","timestamp":1782910388000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/doi\/10.1093\/bioinformatics\/btag395\/8709948"}},"subtitle":[],"editor":[{"given":"Jonathan","family":"Wren","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2026,6,1]]},"references-count":43,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2026,6,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btag395","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2026,6]]},"published":{"date-parts":[[2026,6,1]]},"article-number":"btag395"}}