{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T18:04:25Z","timestamp":1787508265166,"version":"build-2736575974"},"update-to":[{"DOI":"10.1371\/journal.pcbi.1011734","type":"new_version","label":"New version","source":"publisher","updated":{"date-parts":[[2024,1,5]],"date-time":"2024-01-05T00:00:00Z","timestamp":1704412800000}}],"reference-count":36,"publisher":"Public Library of Science (PLoS)","issue":"12","license":[{"start":{"date-parts":[[2023,12,21]],"date-time":"2023-12-21T00:00:00Z","timestamp":1703116800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["DBI-2019797"],"award-info":[{"award-number":["DBI-2019797"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["DBI-214517"],"award-info":[{"award-number":["DBI-214517"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000051","name":"National Human Genome Research Institute","doi-asserted-by":"publisher","award":["R01HG011065"],"award-info":[{"award-number":["R01HG011065"]}],"id":[{"id":"10.13039\/100000051","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["www.ploscompbiol.org"],"crossmark-restriction":false},"short-container-title":["PLoS Comput Biol"],"abstract":"<jats:p>\n                    Transcript annotations play a critical role in gene expression analysis as they serve as a reference for quantifying isoform-level expression. The two main sources of annotations are RefSeq and Ensembl\/GENCODE, but discrepancies between their methodologies and information resources can lead to significant differences. It has been demonstrated that the choice of annotation can have a significant impact on gene expression analysis. Furthermore, transcript assembly is closely linked to annotations, as assembling large-scale available RNA-seq data is an effective data-driven way to construct annotations, and annotations are often served as benchmarks to evaluate the accuracy of assembly methods. However, the influence of different annotations on transcript assembly is not yet fully understood. We investigate the impact of annotations on transcript assembly. Surprisingly, we observe that opposite conclusions can arise when evaluating assemblers with different annotations. To understand this striking phenomenon, we compare the structural similarity of annotations at various levels and find that the primary structural difference across annotations occurs at the intron-chain level. Next, we examine the biotypes of annotated and assembled transcripts and uncover a significant bias towards annotating and assembling transcripts with intron retentions, which explains above the contradictory conclusions. We develop a standalone tool, available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/Shao-Group\/irtool\" xlink:type=\"simple\">https:\/\/github.com\/Shao-Group\/irtool<\/jats:ext-link>\n                    , that can be combined with an assembler to generate an assembly without intron retentions. We evaluate the performance of such a pipeline and offer guidance to select appropriate assembling tools for different application scenarios.\n                  <\/jats:p>","DOI":"10.1371\/journal.pcbi.1011734","type":"journal-article","created":{"date-parts":[[2023,12,21]],"date-time":"2023-12-21T14:00:57Z","timestamp":1703167257000},"page":"e1011734","update-policy":"https:\/\/doi.org\/10.1371\/journal.pcbi.corrections_policy","source":"Crossref","is-referenced-by-count":9,"title":["Transcript assembly and annotations: Bias and adjustment"],"prefix":"10.1371","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5990-1765","authenticated-orcid":true,"given":"Qimin","family":"Zhang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6112-5139","authenticated-orcid":true,"given":"Mingfu","family":"Shao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"340","published-online":{"date-parts":[[2023,12,21]]},"reference":[{"key":"pcbi.1011734.ref001","doi-asserted-by":"crossref","first-page":"323","DOI":"10.1186\/1471-2105-12-323","article-title":"RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome","volume":"12","author":"B Li","year":"2011","journal-title":"BMC Bioinformatics"},{"key":"pcbi.1011734.ref002","doi-asserted-by":"crossref","first-page":"525","DOI":"10.1038\/nbt.3519","article-title":"Near-optimal probabilistic RNA-seq quantification","volume":"34","author":"NL Bray","year":"2016","journal-title":"Nat Biotechnol"},{"key":"pcbi.1011734.ref003","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1038\/nmeth.4197","article-title":"Salmon provides fast and bias-aware quantification of transcript expression","volume":"14","author":"R Patro","year":"2017","journal-title":"Nat Methods"},{"issue":"1","key":"pcbi.1011734.ref004","doi-asserted-by":"crossref","first-page":"139","DOI":"10.1093\/bioinformatics\/btp616","article-title":"edgeR: a Bioconductor package for differential expression analysis of digital gene expression data","volume":"26","author":"MD Robinson","year":"2010","journal-title":"Bioinformatics"},{"key":"pcbi.1011734.ref005","doi-asserted-by":"crossref","first-page":"550","DOI":"10.1186\/s13059-014-0550-8","article-title":"Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2","volume":"15","author":"MI Love","year":"2014","journal-title":"Genome Biol"},{"issue":"5","key":"pcbi.1011734.ref006","doi-asserted-by":"crossref","first-page":"462","DOI":"10.1038\/nbt.2862","article-title":"Sailfish enables alignment-free isoform quantification from RNA-seq reads using lightweight algorithms","volume":"32","author":"R Patro","year":"2014","journal-title":"Nat Biotechnol"},{"issue":"51","key":"pcbi.1011734.ref007","doi-asserted-by":"crossref","first-page":"E5593","DOI":"10.1073\/pnas.1419161111","article-title":"rMATS: robust and flexible detection of differential alternative splicing from replicate RNA-Seq data","volume":"111","author":"S Shen","year":"2014","journal-title":"Proc Natl Acad Sci USA"},{"issue":"3","key":"pcbi.1011734.ref008","doi-asserted-by":"crossref","first-page":"eabq5072","DOI":"10.1126\/sciadv.abq5072","article-title":"ESPRESSO: Robust discovery and quantification of transcript isoforms from error-prone long-read RNA-seq data","volume":"9","author":"Y Gao","year":"2023","journal-title":"Sci Adv"},{"key":"pcbi.1011734.ref009","doi-asserted-by":"crossref","first-page":"D733","DOI":"10.1093\/nar\/gkv1189","article-title":"Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation","volume":"44","author":"NA O\u2019Leary","year":"2016","journal-title":"Nucleic Acids Res"},{"key":"pcbi.1011734.ref010","doi-asserted-by":"crossref","first-page":"D916","DOI":"10.1093\/nar\/gkaa1087","article-title":"GENCODE 2021","volume":"49","author":"A Frankish","year":"2021","journal-title":"Nucleic Acids Res"},{"issue":"Suppl 8","key":"pcbi.1011734.ref011","doi-asserted-by":"crossref","first-page":"S2","DOI":"10.1186\/1471-2164-16-S8-S2","article-title":"Comparison of GENCODE and RefSeq gene annotation and the impact of reference geneset on variant effect prediction","volume":"16","author":"A Frankish","year":"2015","journal-title":"BMC Genomics"},{"key":"pcbi.1011734.ref012","doi-asserted-by":"crossref","first-page":"310","DOI":"10.1038\/s41586-022-04558-8","article-title":"A joint NCBI and EMBL-EBI transcript set for clinical genomics and research","volume":"604","author":"J Morales","year":"2022","journal-title":"Nature"},{"key":"pcbi.1011734.ref013","doi-asserted-by":"crossref","first-page":"699","DOI":"10.1038\/s41586-020-2493-4","article-title":"Expanded encyclopaedias of DNA elements in the human and mouse genomes","volume":"583","author":"TEP Consortium","year":"2020","journal-title":"Nature"},{"key":"pcbi.1011734.ref014","doi-asserted-by":"crossref","first-page":"434","DOI":"10.1038\/s41586-020-2308-7","article-title":"The mutational constraint spectrum quantified from variation in 141,456 humans","volume":"581","author":"KJ Karczewski","year":"2020","journal-title":"Nature"},{"key":"pcbi.1011734.ref015","doi-asserted-by":"crossref","first-page":"1318","DOI":"10.1126\/science.aaz1776","article-title":"The GTEx Consortium atlas of genetic regulatory effects across human tissues","volume":"369","author":"G Consortium","year":"2020","journal-title":"Science"},{"key":"pcbi.1011734.ref016","first-page":"1","article-title":"Assessing the impact of human genome annotation choice on RNA-seq expression estimates","volume":"14","author":"P Wu","year":"2013","journal-title":"BMC Bioinformatics"},{"issue":"7","key":"pcbi.1011734.ref017","doi-asserted-by":"crossref","first-page":"e101374","DOI":"10.1371\/journal.pone.0101374","article-title":"Assessment of the impact of using a reference transcriptome in mapping short RNA-seq reads","volume":"9","author":"S Zhao","year":"2014","journal-title":"PLoS One"},{"issue":"97","key":"pcbi.1011734.ref018","article-title":"A comprehensive evaluation of ensembl, RefSeq, and UCSC annotations in the context of RNA-seq read mapping and gene quantification","volume":"16","author":"S Zhao","year":"2015","journal-title":"BMC Genomics"},{"issue":"4","key":"pcbi.1011734.ref019","doi-asserted-by":"crossref","first-page":"479","DOI":"10.1261\/rna.037473.112","article-title":"Incorporating the human gene annotations in different databases significantly improved transcriptomic and genetic analyses","volume":"19","author":"G Chen","year":"2013","journal-title":"RNA"},{"key":"pcbi.1011734.ref020","doi-asserted-by":"crossref","first-page":"658","DOI":"10.1109\/TCBB.2017.2779509","article-title":"Theory and a heuristic for the minimum path flow decomposition problem","volume":"16","author":"M Shao","year":"2017","journal-title":"IEEE\/ACM Trans Comput Biol Bioinform"},{"issue":"11","key":"pcbi.1011734.ref021","doi-asserted-by":"crossref","first-page":"1252","DOI":"10.1089\/cmb.2022.0257","article-title":"Efficient minimum flow decomposition via integer linear programming","volume":"29","author":"FHC Dias","year":"2022","journal-title":"J Comput Biol"},{"issue":"1","key":"pcbi.1011734.ref022","doi-asserted-by":"crossref","first-page":"360","DOI":"10.1109\/TCBB.2022.3147697","article-title":"Flow decomposition with subpath constraints","volume":"20","author":"L Williams","year":"2023","journal-title":"IEEE\/ACM Trans Comput Biol Bioinform"},{"issue":"5","key":"pcbi.1011734.ref023","doi-asserted-by":"crossref","first-page":"511","DOI":"10.1038\/nbt.1621","article-title":"Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform switching during cell differentiation","volume":"28","author":"C Trapnell","year":"2010","journal-title":"Nat Biotechnol"},{"issue":"10","key":"pcbi.1011734.ref024","doi-asserted-by":"crossref","first-page":"e98","DOI":"10.1093\/nar\/gkw158","article-title":"CLASS2: accurate and efficient splice variant annotation from RNA-seq reads","volume":"44","author":"L Song","year":"2016","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"pcbi.1011734.ref025","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1186\/s13059-016-1074-1","article-title":"TransComb: genome-guided transcriptome assembly via combing junctions in splicing graphs","volume":"17","author":"J Liu","year":"2016","journal-title":"Genome Biol"},{"key":"pcbi.1011734.ref026","doi-asserted-by":"crossref","first-page":"1438","DOI":"10.1038\/s41467-020-15171-6","article-title":"Full-length transcript characterization of SF3B1 mutation in chronic lymphocytic leukemia reveals downregulation of retained introns","volume":"11","author":"AD Tang","year":"2020","journal-title":"Nat Commun"},{"issue":"3","key":"pcbi.1011734.ref027","doi-asserted-by":"crossref","first-page":"290","DOI":"10.1038\/nbt.3122","article-title":"StringTie enables improved reconstruction of a transcriptome from RNA-seq reads","volume":"33","author":"M Pertea","year":"2015","journal-title":"Nat Biotechnol"},{"issue":"12","key":"pcbi.1011734.ref028","doi-asserted-by":"crossref","first-page":"1167","DOI":"10.1038\/nbt.4020","article-title":"Accurate assembly of transcripts through phase-preserving graph decomposition","volume":"35","author":"M Shao","year":"2017","journal-title":"Nat Biotechnol"},{"key":"pcbi.1011734.ref029","doi-asserted-by":"crossref","first-page":"278","DOI":"10.1186\/s13059-019-1910-1","article-title":"Transcriptome assembly from long-read RNA-seq alignments with StringTie2","volume":"20","author":"S Kovaka","year":"2019","journal-title":"Genome Biol"},{"key":"pcbi.1011734.ref030","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1038\/s43588-022-00216-1","article-title":"Accurate assembly of multi-end RNA-seq data with Scallop2","volume":"2","author":"Q Zhang","year":"2022","journal-title":"Nat Comput Sci"},{"key":"pcbi.1011734.ref031","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1186\/s13059-018-1590-2","article-title":"CHESS: a new human gene catalog curated from thousands of large-scale RNA sequencing experiments reveals extensive transcriptional noise","volume":"19","author":"M Pertea","year":"2018","journal-title":"Genome Biol"},{"key":"pcbi.1011734.ref032","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1126\/science.abj6987","article-title":"The complete sequence of a human genome","volume":"376","author":"S Nurk","year":"2022","journal-title":"Science"},{"issue":"1","key":"pcbi.1011734.ref033","doi-asserted-by":"crossref","first-page":"15","DOI":"10.1093\/bioinformatics\/bts635","article-title":"STAR: ultrafast universal RNA-seq aligner","volume":"29","author":"A Dobin","year":"2013","journal-title":"Bioinformatics"},{"key":"pcbi.1011734.ref034","doi-asserted-by":"crossref","first-page":"907","DOI":"10.1038\/s41587-019-0201-4","article-title":"Graph-based genome alignment and genotyping with HISAT2 and HISAT-genotype","volume":"37","author":"D Kim","year":"2019","journal-title":"Nat Biotechnol"},{"key":"pcbi.1011734.ref035","doi-asserted-by":"crossref","first-page":"304","DOI":"10.12688\/f1000research.23297.1","article-title":"GFF utilities: GffRead and GffCompare","volume":"9","author":"G Pertea","year":"2020","journal-title":"F1000Res"},{"key":"pcbi.1011734.ref036","article-title":"Systematic assessment of long-read RNA-seq methods for transcript identification and quantification","author":"FJ Pardo-Palacios","year":"2023","journal-title":"bioRxiv"}],"updated-by":[{"DOI":"10.1371\/journal.pcbi.1011734","type":"new_version","label":"New version","source":"publisher","updated":{"date-parts":[[2024,1,5]],"date-time":"2024-01-05T00:00:00Z","timestamp":1704412800000}}],"container-title":["PLOS Computational Biology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dx.plos.org\/10.1371\/journal.pcbi.1011734","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,6]],"date-time":"2024-11-06T12:08:23Z","timestamp":1730894903000},"score":1,"resource":{"primary":{"URL":"https:\/\/dx.plos.org\/10.1371\/journal.pcbi.1011734"}},"subtitle":[],"editor":[{"given":"Marc","family":"Robinson-Rechavi","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2023,12,21]]},"references-count":36,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2023,12,21]]}},"URL":"https:\/\/doi.org\/10.1371\/journal.pcbi.1011734","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/2023.04.20.537700","asserted-by":"object"}]},"ISSN":["1553-7358"],"issn-type":[{"value":"1553-7358","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,12,21]]}}}