{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,25]],"date-time":"2025-11-25T14:03:14Z","timestamp":1764079394916,"version":"3.37.3"},"reference-count":16,"publisher":"Oxford University Press (OUP)","issue":"4","license":[{"start":{"date-parts":[[2018,7,23]],"date-time":"2018-07-23T00:00:00Z","timestamp":1532304000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"DOI":"10.13039\/100000002","name":"NIH","doi-asserted-by":"publisher","award":["R01-HG007196","R01-HL129239"],"award-info":[{"award-number":["R01-HG007196","R01-HL129239"]}],"id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,2,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>DNA sequencing archives have grown to enormous scales in recent years, and thousands of human genomes have already been sequenced. The size of these data sets has made searching the raw read data infeasible without high-performance data-query technology. Additionally, it is challenging to search a repository of short-read data using relational logic and to apply that logic across samples from multiple whole-genome sequencing samples.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>We have built a compact, efficiently-indexed database that contains the raw read data for over 250 human genomes, encompassing trillions of bases of DNA, and that allows users to search these data in real-time. The Terabase Search Engine enables retrieval from this database of all the reads for any genomic location in a matter of seconds. Users can search using a range of positions or a specific sequence that is aligned to the genome on the fly.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>Public access to the Terabase Search Engine database is available at http:\/\/tse.idies.jhu.edu.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information<\/jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/bty657","type":"journal-article","created":{"date-parts":[[2018,7,20]],"date-time":"2018-07-20T20:48:45Z","timestamp":1532119725000},"page":"665-670","source":"Crossref","is-referenced-by-count":8,"title":["The Terabase Search Engine: a large-scale relational database of short-read sequences"],"prefix":"10.1093","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1263-5532","authenticated-orcid":false,"given":"Richard","family":"Wilton","sequence":"first","affiliation":[{"name":"Department of Physics and Astronomy, Johns Hopkins University, Baltimore, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sarah J","family":"Wheelan","sequence":"additional","affiliation":[{"name":"Department of Oncology, Johns Hopkins University School of Medicine, Baltimore, MD, USA"},{"name":"Department of Biostatistics, Johns Hopkins Bloomberg School of Public Health, Baltimore, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alexander S","family":"Szalay","sequence":"additional","affiliation":[{"name":"Department of Physics and Astronomy, Johns Hopkins University, Baltimore, MD, USA"},{"name":"Department of Computer Science, Johns Hopkins University School of Medicine, Baltimore, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Steven L","family":"Salzberg","sequence":"additional","affiliation":[{"name":"Department of Biostatistics, Johns Hopkins Bloomberg School of Public Health, Baltimore, MD, USA"},{"name":"Department of Computer Science, Johns Hopkins University School of Medicine, Baltimore, MD, USA"},{"name":"Department of Biomedical Engineering, Johns Hopkins University School of Medicine, Baltimore, MD, USA"},{"name":"Center for Computational Biology, McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University School of Medicine, Baltimore, MD, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2018,7,23]]},"reference":[{"key":"2023051511010661700_bty657-B1","doi-asserted-by":"crossref","first-page":"1498","DOI":"10.1101\/gr.123638.111","article-title":"Accurate and comprehensive sequencing of personal genomes","volume":"21","author":"Ajay","year":"2011","journal-title":"Genome Res"},{"key":"2023051511010661700_bty657-B2","doi-asserted-by":"crossref","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","article-title":"Gapped BLAST and PSI-BLAST: a new generation of protein database search programs","volume":"25","author":"Altschul","year":"1997","journal-title":"Nucleic Acids Res"},{"key":"2023051511010661700_bty657-B3","doi-asserted-by":"crossref","first-page":"e25452.","DOI":"10.1371\/journal.pone.0025452","article-title":"Mucin variable number tandem repeat polymorphisms and severity of cystic fibrosis lung disease: significant association with MUC5AC","volume":"6","author":"Guo","year":"2011","journal-title":"PLoS One"},{"key":"2023051511010661700_bty657-B4","doi-asserted-by":"crossref","first-page":"734","DOI":"10.1101\/gr.114819.110","article-title":"Efficient storage of high throughput DNA sequencing data using reference-based compression","volume":"21","author":"Hsi-Yang Fritz","year":"2011","journal-title":"Genome Res"},{"year":"2012","key":"2023051511010661700_bty657-B5"},{"key":"2023051511010661700_bty657-B6","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1038\/nmeth.1923","article-title":"Fast gapped-read alignment with Bowtie 2","volume":"9","author":"Langmead","year":"2012","journal-title":"Nat. Methods"},{"key":"2023051511010661700_bty657-B7","doi-asserted-by":"crossref","first-page":"1754","DOI":"10.1093\/bioinformatics\/btp324","article-title":"Fast and accurate short read alignment with Burrows-Wheeler transform","volume":"25","author":"Li","year":"2009","journal-title":"Bioinformatics"},{"key":"2023051511010661700_bty657-B8","doi-asserted-by":"crossref","first-page":"201","DOI":"10.1038\/nature18964","article-title":"The Simons Genome Diversity Project: 300 genomes from 142 diverse populations","volume":"538","author":"Mallick","year":"2016","journal-title":"Nature"},{"key":"2023051511010661700_bty657-B9","doi-asserted-by":"crossref","first-page":"849","DOI":"10.1101\/gr.213611.116","article-title":"Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly","volume":"27","author":"Schneider","year":"2017","journal-title":"Genome Res"},{"key":"2023051511010661700_bty657-B10","doi-asserted-by":"crossref","first-page":"631","DOI":"10.1016\/j.ajhg.2015.09.010","article-title":"Privacy risks from genomic data-sharing beacons","volume":"97","author":"Shringarpure","year":"2015","journal-title":"Am. J. Hum. Genet"},{"key":"2023051511010661700_bty657-B11","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1038\/nature15393","article-title":"A global reference for human genetic variation","volume":"526","year":"2015","journal-title":"Nature"},{"key":"2023051511010661700_bty657-B12","doi-asserted-by":"crossref","first-page":"178","DOI":"10.1093\/bib\/bbs017","article-title":"Integrative Genomics Viewer (IGV): high-performance genomics data visualization and exploration","volume":"14","author":"Thorvaldsd\u00f3ttir","year":"2013","journal-title":"Brief. Bioinform"},{"volume-title":"Presented at the NVidia GPU Technology Conference","year":"2017","author":"Wilton","key":"2023051511010661700_bty657-B13"},{"key":"2023051511010661700_bty657-B14","doi-asserted-by":"crossref","first-page":"e808.","DOI":"10.7717\/peerj.808","article-title":"Arioc: high-throughput read alignment with GPU-accelerated exploration of the seed-and-extend search space","volume":"3","author":"Wilton","year":"2015","journal-title":"PeerJ"},{"key":"2023051511010661700_bty657-B15","doi-asserted-by":"crossref","first-page":"240","DOI":"10.1038\/nbt.3170","article-title":"Quality score compression improves genotyping accuracy","volume":"33","author":"Yu","year":"2015","journal-title":"Nat. Biotechnol"},{"first-page":"500","year":"2014","author":"Zola","key":"2023051511010661700_bty657-B16"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/4\/665\/50321287\/bioinformatics_35_4_665.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/4\/665\/50321287\/bioinformatics_35_4_665.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,5,15]],"date-time":"2023-05-15T11:02:37Z","timestamp":1684148557000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/35\/4\/665\/5057158"}},"subtitle":[],"editor":[{"given":"Jonathan","family":"Wren","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2018,7,23]]},"references-count":16,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2019,2,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bty657","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"type":"print","value":"1367-4803"},{"type":"electronic","value":"1367-4811"}],"subject":[],"published-other":{"date-parts":[[2019,2,15]]},"published":{"date-parts":[[2018,7,23]]}}}