{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T01:10:33Z","timestamp":1781053833737,"version":"3.54.1"},"reference-count":24,"publisher":"Oxford University Press (OUP)","issue":"Supplement_1","license":[{"start":{"date-parts":[[2024,6,28]],"date-time":"2024-06-28T00:00:00Z","timestamp":1719532800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000002","name":"National Institutes of Health","doi-asserted-by":"publisher","award":["R01 HG009937"],"award-info":[{"award-number":["R01 HG009937"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CCF-1750472"],"award-info":[{"award-number":["CCF-1750472"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["CNS-1763680"],"award-info":[{"award-number":["CNS-1763680"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2022-252586"],"award-info":[{"award-number":["2022-252586"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Chan Zuckerberg Initiative DAF"},{"DOI":"10.13039\/100000923","name":"Silicon Valley Community Foundation","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000923","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,6,28]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Motivation<\/jats:title>\n                    <jats:p>Short-read single-cell RNA-sequencing (scRNA-seq) has been used to study cellular heterogeneity, cellular fate, and transcriptional dynamics. Modeling splicing dynamics in scRNA-seq data is challenging, with inherent difficulty in even the seemingly straightforward task of elucidating the splicing status of the molecules from which sequenced fragments are drawn. This difficulty arises, in part, from the limited read length and positional biases, which substantially reduce the specificity of the sequenced fragments. As a result, the splicing status of many reads in scRNA-seq is ambiguous because of a lack of definitive evidence. We are therefore in need of methods that can recover the splicing status of ambiguous reads which, in turn, can lead to more accuracy and confidence in downstream analyses.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>We develop Forseti, a predictive model to probabilistically assign a splicing status to scRNA-seq reads. Our model has two key components. First, we train a binding affinity model to assign a probability that a given transcriptomic site is used in fragment generation. Second, we fit a robust fragment length distribution model that generalizes well across datasets deriving from different species and tissue types. Forseti combines these two trained models to predict the splicing status of the molecule of origin of reads by scoring putative fragments that associate each alignment of sequenced reads with proximate potential priming sites. Using both simulated and experimental data, we show that our model can precisely predict the splicing status of many reads and identify the true gene origin of multi-gene mapped reads.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Availability and implementation<\/jats:title>\n                    <jats:p>Forseti and the code used for producing the results are available at https:\/\/github.com\/COMBINE-lab\/forseti under a BSD 3-clause license.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btae207","type":"journal-article","created":{"date-parts":[[2024,4,12]],"date-time":"2024-04-12T10:54:02Z","timestamp":1712919242000},"page":"i297-i306","source":"Crossref","is-referenced-by-count":2,"title":["<tt>Forseti<\/tt>\n                    : a mechanistic and predictive model of the splicing status of scRNA-seq reads"],"prefix":"10.1093","volume":"40","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8259-7434","authenticated-orcid":false,"given":"Dongze","family":"He","sequence":"first","affiliation":[{"name":"Center for Bioinformatics and Computational Biology, University of Maryland , College Park, MD 20742, United States"},{"name":"Program in Computational Biology, Bioinformatics and Genomices, University of Maryland , College Park, MD 20742, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7517-588X","authenticated-orcid":false,"given":"Yuan","family":"Gao","sequence":"additional","affiliation":[{"name":"Center for Bioinformatics and Computational Biology, University of Maryland , College Park, MD 20742, United States"},{"name":"Program in Computational Biology, Bioinformatics and Genomices, University of Maryland , College Park, MD 20742, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7061-3695","authenticated-orcid":false,"given":"Spencer Skylar","family":"Chan","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Maryland , College Park, MD 20742, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Natalia","family":"Quintana-Parrilla","sequence":"additional","affiliation":[{"name":"Department of Biology, University of Puerto Rico, Mayag\u00fcez Campus , Mayag\u00fcez 00682, Puerto Rico"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8463-1675","authenticated-orcid":false,"given":"Rob","family":"Patro","sequence":"additional","affiliation":[{"name":"Center for Bioinformatics and Computational Biology, University of Maryland , College Park, MD 20742, United States"},{"name":"Department of Computer Science, University of Maryland , College Park, MD 20742, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2024,6,28]]},"reference":[{"key":"2024071814131716600_btae207-B1","author":"10x Genomics","year":"2018"},{"key":"2024071814131716600_btae207-B2","author":"10x Genomics","year":"2021"},{"key":"2024071814131716600_btae207-B3","author":"10x Genomics","year":"2022"},{"key":"2024071814131716600_btae207-B4","author":"10x Genomics","year":"2022"},{"key":"2024071814131716600_btae207-B5","doi-asserted-by":"crossref","first-page":"1408","DOI":"10.1038\/s41587-020-0591-3","article-title":"Generalizing RNA velocity to transient cell states through dynamical modeling","volume":"38","author":"Bergen","year":"2020","journal-title":"Nat Biotechnol"},{"key":"2024071814131716600_btae207-B6","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1101\/gr.278253.123","article-title":"Differences in molecular sampling and data processing explain variation among single-cell and single-nucleus RNA-seq experiments","volume":"34","author":"Chamberlin","year":"2024","journal-title":"Genome Res"},{"key":"2024071814131716600_btae207-B7","unstructured":"Chen X, Roelli P, Here\u00f1\u00fa D \u00a0et al \u00a02023. Teichlab\/scg_lib_structs: Release October 26, 2023. https:\/\/zenodo.org\/doi\/10.5281\/zenodo.10042390"},{"key":"2024071814131716600_btae207-B8","author":"Eldj\u00e1rn Hj\u00f6rleifsson","year":"2022"},{"key":"2024071814131716600_btae207-B9","doi-asserted-by":"crossref","first-page":"114","DOI":"10.1007\/s11538-023-01213-9","article-title":"Assessing Markovian and delay models for single-nucleus RNA sequencing","volume":"85","author":"Gorin","year":"2023","journal-title":"Bull Math Biol"},{"key":"2024071814131716600_btae207-B10","doi-asserted-by":"crossref","first-page":"521","DOI":"10.1093\/bioinformatics\/bty630","article-title":"Simulating illumina metagenomic data with insilicoseq","volume":"35","author":"Gourl\u00e9","year":"2019","journal-title":"Bioinformatics"},{"key":"2024071814131716600_btae207-B11","doi-asserted-by":"crossref","DOI":"10.1093\/bioinformatics\/btad614","article-title":"simpleaf: a simple, flexible, and scalable framework for single-cell data processing using alevin-fry","volume":"39","author":"He","year":"2023","journal-title":"Bioinformatics"},{"key":"2024071814131716600_btae207-B12","author":"He","year":"2024"},{"key":"2024071814131716600_btae207-B13","author":"He","year":"2023"},{"key":"2024071814131716600_btae207-B14","doi-asserted-by":"crossref","first-page":"316","DOI":"10.1038\/s41592-022-01408-3","article-title":"Alevin-fry unlocks rapid, accurate and memory-frugal quantification of single-cell RNA-seq data","volume":"19","author":"He","year":"2022","journal-title":"Nat Methods"},{"key":"2024071814131716600_btae207-B15","author":"Kaminow","year":"2021"},{"key":"2024071814131716600_btae207-B16","doi-asserted-by":"crossref","first-page":"494","DOI":"10.1038\/s41586-018-0414-6","article-title":"RNA velocity of single cells","volume":"560","author":"La Manno","year":"2018","journal-title":"Nature"},{"key":"2024071814131716600_btae207-B17","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1038\/s41587-023-01728-5","article-title":"A relay velocity model infers cell-dependent RNA velocity","volume":"42","author":"Li","year":"2023","journal-title":"Nat Biotechnol"},{"key":"2024071814131716600_btae207-B18","doi-asserted-by":"crossref","first-page":"813","DOI":"10.1038\/s41587-021-00870-2","article-title":"Modular, efficient and constant-memory single-cell RNA-seq preprocessing","volume":"39","author":"Melsted","year":"2021","journal-title":"Nat Biotechnol"},{"key":"2024071814131716600_btae207-B19","doi-asserted-by":"crossref","first-page":"6152","DOI":"10.1073\/pnas.092140899","article-title":"Oligo(dT) primer generates a high frequency of truncated cDNAs through internal poly(A) priming during reverse transcription","volume":"99","author":"Nam","year":"2002","journal-title":"Proc Natl Acad Sci USA"},{"key":"2024071814131716600_btae207-B20","doi-asserted-by":"crossref","first-page":"1506","DOI":"10.1038\/s41592-023-02003-w","article-title":"Recovery of missing single-cell RNA-sequencing data with optimized transcriptomic references","volume":"20","author":"Pool","year":"2023","journal-title":"Nat Methods"},{"key":"2024071814131716600_btae207-B21","doi-asserted-by":"crossref","first-page":"841","DOI":"10.1093\/bioinformatics\/btq033","article-title":"BEDTools: a flexible suite of utilities for comparing genomic features","volume":"26","author":"Quinlan","year":"2010","journal-title":"Bioinformatics"},{"key":"2024071814131716600_btae207-B22","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1186\/s13059-019-1670-y","article-title":"Alevin efficiently estimates accurate gene abundances from dscRNA-seq data","volume":"20","author":"Srivastava","year":"2019","journal-title":"Genome Biol"},{"key":"2024071814131716600_btae207-B23","doi-asserted-by":"crossref","first-page":"631","DOI":"10.1038\/s41576-019-0150-2","article-title":"RNA sequencing: the teenage years","volume":"20","author":"Stark","year":"2019","journal-title":"Nat Rev Genet"},{"key":"2024071814131716600_btae207-B24","doi-asserted-by":"crossref","first-page":"lqac035","DOI":"10.1093\/nargab\/lqac035","article-title":"Internal oligo(dT) priming introduces systematic bias in bulk and single-cell RNA sequencing count data","volume":"4","author":"Svoboda","year":"2022","journal-title":"NAR Genom Bioinform"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/40\/Supplement_1\/i297\/58585828\/btae207.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/40\/Supplement_1\/i297\/58585828\/btae207.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,18]],"date-time":"2024-07-18T11:36:19Z","timestamp":1721302579000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/40\/Supplement_1\/i297\/7700855"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,6,28]]},"references-count":24,"journal-issue":{"issue":"Supplement_1","published-print":{"date-parts":[[2024,6,28]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btae207","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2024.02.01.577813","asserted-by":"object"}]},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024,7]]},"published":{"date-parts":[[2024,6,28]]}}}