{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T04:23:57Z","timestamp":1776140637831,"version":"3.50.1"},"reference-count":36,"publisher":"Oxford University Press (OUP)","issue":"9","license":[{"start":{"date-parts":[[2017,12,15]],"date-time":"2017-12-15T00:00:00Z","timestamp":1513296000000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100004807","name":"DFG","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100004807","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2018,5,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec><jats:title>Motivation<\/jats:title><jats:p>The increasing amount of next-generation sequencing data poses a fundamental challenge on large scale genomic analytics. Existing tools use different distributed computational platforms to scale-out bioinformatics workloads. However, the scalability of these tools is not efficient. Moreover, they have heavy run time overheads when pre-processing large amounts of data. To address these limitations, we have developed Sparkhit: a distributed bioinformatics framework built on top of the Apache Spark platform.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results<\/jats:title><jats:p>Sparkhit integrates a variety of analytical methods. It is implemented in the Spark extended MapReduce model. It runs 92\u2013157 times faster than MetaSpark on metagenomic fragment recruitment and 18\u201332 times faster than Crossbow on data pre-processing. We analyzed 100 terabytes of data across four genomic projects in the cloud in 21\u2009h, which includes the run times of cluster deployment and data downloading. Furthermore, our application on the entire Human Microbiome Project shotgun sequencing data was completed in 2\u2009h, presenting an approach to easily associate large amounts of public datasets with reference data.<\/jats:p><\/jats:sec><jats:sec><jats:title>Availability and implementation<\/jats:title><jats:p>Sparkhit is freely available at: https:\/\/rhinempi.github.io\/sparkhit\/.<\/jats:p><\/jats:sec><jats:sec><jats:title>Supplementary information<\/jats:title><jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p><\/jats:sec>","DOI":"10.1093\/bioinformatics\/btx808","type":"journal-article","created":{"date-parts":[[2017,12,14]],"date-time":"2017-12-14T20:11:53Z","timestamp":1513282313000},"page":"1457-1465","source":"Crossref","is-referenced-by-count":13,"title":["Analyzing large scale genomic data on the cloud with Sparkhit"],"prefix":"10.1093","volume":"34","author":[{"given":"Liren","family":"Huang","sequence":"first","affiliation":[{"name":"Faculty of Technology, Bielefeld University, Bielefeld, Germany"},{"name":"Center for Biotechnology \u2013 CeBiTec, Bielefeld University, Bielefeld, Germany"},{"name":"Computational Methods for the Analysis of the Diversity and Dynamics of Genomes, Bielefeld University, Bielefeld, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jan","family":"Kr\u00fcger","sequence":"additional","affiliation":[{"name":"Faculty of Technology, Bielefeld University, Bielefeld, Germany"},{"name":"Center for Biotechnology \u2013 CeBiTec, Bielefeld University, Bielefeld, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4405-3847","authenticated-orcid":false,"given":"Alexander","family":"Sczyrba","sequence":"additional","affiliation":[{"name":"Faculty of Technology, Bielefeld University, Bielefeld, Germany"},{"name":"Center for Biotechnology \u2013 CeBiTec, Bielefeld University, Bielefeld, Germany"},{"name":"Computational Methods for the Analysis of the Diversity and Dynamics of Genomes, Bielefeld University, Bielefeld, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2017,12,15]]},"reference":[{"key":"2023012713023706000_btx808-B1","doi-asserted-by":"crossref","first-page":"e0155461.","DOI":"10.1371\/journal.pone.0155461","article-title":"Sparkbwa: speeding up the alignment of high-throughput dna sequencing data","volume":"11","author":"Abuin","year":"2016","journal-title":"PLoS One"},{"key":"2023012713023706000_btx808-B2","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1038\/nature15393","article-title":"A global reference for human genetic variation","volume":"526","author":"Auton","year":"2015","journal-title":"Nature"},{"key":"2023012713023706000_btx808-B3","first-page":"67","article-title":"Review of current methods, applications, and data management for the bioinformatics analysis of whole exome sequencing","volume":"13","author":"Bao","year":"2014","journal-title":"Cancer Inf"},{"key":"2023012713023706000_btx808-B4","doi-asserted-by":"crossref","first-page":"1519","DOI":"10.1089\/cmb.2009.0238","article-title":"Ray: simultaneous assembly of reads from a mix of high-throughput sequencing technologies","volume":"17","author":"Boisvert","year":"2010","journal-title":"J. Comput. Biol"},{"key":"2023012713023706000_btx808-B5","doi-asserted-by":"crossref","first-page":"888","DOI":"10.1038\/nbt0816-888d","article-title":"Near-optimal probabilistic RNA-seq quantification","volume":"34","author":"Bray","year":"2016","journal-title":"Nat. Biotechnol"},{"key":"2023012713023706000_btx808-B6","author":"Chen","year":"2015"},{"key":"2023012713023706000_btx808-B7","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1145\/1327452.1327492","article-title":"Mapreduce: simplified data processing on large clusters","volume":"51","author":"Dean","year":"2008","journal-title":"Commun. ACM"},{"key":"2023012713023706000_btx808-B8","doi-asserted-by":"crossref","first-page":"2482","DOI":"10.1093\/bioinformatics\/btv179","article-title":"Halvade: scalable sequence analysis with mapreduce","volume":"31","author":"Decap","year":"2015","journal-title":"Bioinformatics"},{"key":"2023012713023706000_btx808-B9","doi-asserted-by":"crossref","first-page":"1267","DOI":"10.1093\/bioinformatics\/btv698","article-title":"qsubsec: a lightweight template system for defining sun grid engine workflows","volume":"32","author":"Droop","year":"2016","journal-title":"Bioinformatics"},{"key":"2023012713023706000_btx808-B10","doi-asserted-by":"crossref","first-page":"10476","DOI":"10.1038\/ncomms10476","article-title":"Global metagenomic survey reveals a new bacterial candidate phylum in geothermal springs","volume":"7","author":"Eloe-Fadrosh","year":"2016","journal-title":"Nat. Commun"},{"key":"2023012713023706000_btx808-B11","doi-asserted-by":"crossref","first-page":"789","DOI":"10.1016\/0167-8191(96)00024-5","article-title":"A high-performance, portable implementation of the mpi message passing interface standard","volume":"22","author":"Gropp","year":"1996","journal-title":"Parallel Comput"},{"key":"2023012713023706000_btx808-B12","doi-asserted-by":"crossref","DOI":"10.1002\/0471250953.bi1107s32","article-title":"Aligning short sequencing reads with bowtie","author":"Langmead","year":"2010","journal-title":"Curr. Protoc. Bioinf"},{"key":"2023012713023706000_btx808-B13","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1038\/nmeth.1923","article-title":"Fast gapped-read alignment with bowtie 2","volume":"9","author":"Langmead","year":"2012","journal-title":"Nat. Methods"},{"key":"2023012713023706000_btx808-B14","doi-asserted-by":"crossref","first-page":"R134.","DOI":"10.1186\/gb-2009-10-11-r134","article-title":"Searching for snps with cloud computing","volume":"10","author":"Langmead","year":"2009","journal-title":"Genome Biol"},{"key":"2023012713023706000_btx808-B15","doi-asserted-by":"crossref","first-page":"R83.","DOI":"10.1186\/gb-2010-11-8-r83","article-title":"Cloud-scale RNA-sequencing differential expression analysis with myrna","volume":"11","author":"Langmead","year":"2010","journal-title":"Genome Biol"},{"key":"2023012713023706000_btx808-B16","doi-asserted-by":"crossref","first-page":"1754","DOI":"10.1093\/bioinformatics\/btp324","article-title":"Fast and accurate short read alignment with burrows-wheeler transform","volume":"25","author":"Li","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012713023706000_btx808-B17","doi-asserted-by":"crossref","first-page":"2078","DOI":"10.1093\/bioinformatics\/btp352","article-title":"The sequence alignment\/map format and samtools","volume":"25","author":"Li","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012713023706000_btx808-B18","doi-asserted-by":"crossref","first-page":"713","DOI":"10.1093\/bioinformatics\/btn025","article-title":"Soap: short oligonucleotide alignment program","volume":"24","author":"Li","year":"2008","journal-title":"Bioinformatics"},{"key":"2023012713023706000_btx808-B19","doi-asserted-by":"crossref","first-page":"1297","DOI":"10.1101\/gr.107524.110","article-title":"The genome analysis toolkit: a mapreduce framework for analyzing next-generation dna sequencing data","volume":"20","author":"McKenna","year":"2010","journal-title":"Genome Res"},{"key":"2023012713023706000_btx808-B20","first-page":"2","article-title":"Docker: lightweight linux containers for consistent development and deployment","volume":"2014","author":"Merkel","year":"2014","journal-title":"Linux J"},{"key":"2023012713023706000_btx808-B21","doi-asserted-by":"crossref","first-page":"1704","DOI":"10.1093\/bioinformatics\/btr252","article-title":"Fr-hit, a very fast program to recruit metagenomic reads to homologous reference genomes","volume":"27","author":"Niu","year":"2011","journal-title":"Bioinformatics"},{"key":"2023012713023706000_btx808-B22","doi-asserted-by":"crossref","first-page":"2317","DOI":"10.1101\/gr.096651.109","article-title":"The NIH human microbiome project","volume":"19","author":"Peterson","year":"2009","journal-title":"Genome Res"},{"key":"2023012713023706000_btx808-B23","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1186\/2047-217X-3-7","article-title":"The 3,000 rice genomes project","volume":"3","author":"R Genomes Project","year":"2014","journal-title":"Gigascience"},{"key":"2023012713023706000_btx808-B24","doi-asserted-by":"crossref","first-page":"709","DOI":"10.1056\/NEJMoa1106920","article-title":"Origins of the E. coli strain causing an outbreak of hemolytic-uremic syndrome in germany","volume":"365","author":"Rasko","year":"2011","journal-title":"N. Engl. J. Med"},{"key":"2023012713023706000_btx808-B25","doi-asserted-by":"crossref","first-page":"431","DOI":"10.1038\/nature12352","article-title":"Insights into the phylogeny and coding potential of microbial dark matter","volume":"499","author":"Rinke","year":"2013","journal-title":"Nature"},{"key":"2023012713023706000_btx808-B26","doi-asserted-by":"crossref","first-page":"e77","DOI":"10.1371\/journal.pbio.0050077","article-title":"The sorcerer ii global ocean sampling expedition: northwest atlantic through eastern tropical pacific","volume":"5","author":"Rusch","year":"2007","journal-title":"PLoS Biol"},{"key":"2023012713023706000_btx808-B27","doi-asserted-by":"crossref","first-page":"647","DOI":"10.1038\/nrg2857","article-title":"Computational solutions to large-scale data management and analysis","volume":"11","author":"Schadt","year":"2010","journal-title":"Nat. Rev. Genet"},{"key":"2023012713023706000_btx808-B28","doi-asserted-by":"crossref","first-page":"1363","DOI":"10.1093\/bioinformatics\/btp236","article-title":"Cloudburst: highly sensitive read mapping with mapreduce","volume":"25","author":"Schatz","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012713023706000_btx808-B29","author":"Shvachko","year":"2010"},{"key":"2023012713023706000_btx808-B30","doi-asserted-by":"crossref","first-page":"1117","DOI":"10.1101\/gr.089532.108","article-title":"Abyss: a parallel assembler for short read sequence data","volume":"19","author":"Simpson","year":"2009","journal-title":"Genome Res"},{"key":"2023012713023706000_btx808-B31","doi-asserted-by":"crossref","first-page":"R46.","DOI":"10.1186\/gb-2014-15-3-r46","article-title":"Kraken: ultrafast metagenomic sequence classification using exact alignments","volume":"15","author":"Wood","year":"2014","journal-title":"Genome Biol"},{"key":"2023012713023706000_btx808-B32","doi-asserted-by":"crossref","first-page":"426.","DOI":"10.1186\/s13059-014-0426-y","article-title":"Heterogeneity in the inter-tumor transcriptome of high risk prostate cancer","volume":"15","author":"Wyatt","year":"2014","journal-title":"Genome Biol"},{"key":"2023012713023706000_btx808-B33","first-page":"15","author":"Zaharia","year":"2012"},{"key":"2023012713023706000_btx808-B34","first-page":"845","author":"Zhao","year":"2015"},{"key":"2023012713023706000_btx808-B35","doi-asserted-by":"crossref","first-page":"1090","DOI":"10.1093\/bioinformatics\/btw750","article-title":"Metaspark: a spark-based distributed processing tool to recruit metagenomic reads to reference genomes","volume":"33","author":"Zhou","year":"2017","journal-title":"Bioinformatics"},{"key":"2023012713023706000_btx808-B36","first-page":"246","article-title":"Integrating human sequence data sets provides a resource of benchmark snp and indel genotype calls","volume-title":"Nat. Biotechnol","author":"Zook","year":"2014"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/34\/9\/1457\/48916245\/bioinformatics_34_9_1457.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/34\/9\/1457\/48916245\/bioinformatics_34_9_1457.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,8,29]],"date-time":"2023-08-29T22:52:22Z","timestamp":1693349542000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/34\/9\/1457\/4747885"}},"subtitle":[],"editor":[{"given":"Inanc","family":"Birol","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2017,12,15]]},"references-count":36,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2018,5,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btx808","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2018,5,1]]},"published":{"date-parts":[[2017,12,15]]}}}