{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T07:41:04Z","timestamp":1781854864803,"version":"3.54.5"},"reference-count":8,"publisher":"Oxford University Press (OUP)","issue":"2","license":[{"start":{"date-parts":[[2016,10,2]],"date-time":"2016-10-02T00:00:00Z","timestamp":1475366400000},"content-version":"vor","delay-in-days":1058,"URL":"http:\/\/creativecommons.org\/licenses\/by\/3.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2014,1,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation:\u2003Nucleotide sequence data are being produced at an ever increasing rate. Clustering such sequences by similarity is often an essential first step in their analysis\u2014intended to reduce redundancy, define gene families or suggest taxonomic units. Exact clustering algorithms, such as hierarchical clustering, scale relatively poorly in terms of run time and memory usage, yet they are desirable because heuristic shortcuts taken during clustering might have unintended consequences in later analysis steps.<\/jats:p>\n               <jats:p>Results:\u2003Here we present HPC-CLUST, a highly optimized software pipeline that can cluster large numbers of pre-aligned DNA sequences by running on distributed computing hardware. It allocates both memory and computing resources efficiently, and can process more than a million sequences in a few hours on a small cluster.<\/jats:p>\n               <jats:p>Availability and implementation:\u2003Source code and binaries are freely available at http:\/\/meringlab.org\/software\/hpc-clust\/; the pipeline is implemented in C++ and uses the Message Passing Interface (MPI) standard for distributed computing.<\/jats:p>\n               <jats:p>Contact:\u2003 \u00a0mering@imls.uzh.ch<\/jats:p>\n               <jats:p>Supplementary Information:\u2003 \u00a0Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btt657","type":"journal-article","created":{"date-parts":[[2013,11,10]],"date-time":"2013-11-10T01:09:20Z","timestamp":1384045760000},"page":"287-288","source":"Crossref","is-referenced-by-count":53,"title":["HPC-CLUST: distributed hierarchical clustering for large sets of nucleotide sequences"],"prefix":"10.1093","volume":"30","author":[{"given":"Jo\u00e3o F.","family":"Matias Rodrigues","sequence":"first","affiliation":[{"name":"Institute of Molecular Life Sciences and Swiss Institute of Bioinformatics, University of Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christian","family":"von Mering","sequence":"additional","affiliation":[{"name":"Institute of Molecular Life Sciences and Swiss Institute of Bioinformatics, University of Zurich, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2013,11,9]]},"reference":[{"key":"2023012710385851200_btt657-B1","doi-asserted-by":"crossref","first-page":"D141","DOI":"10.1093\/nar\/gkn879","article-title":"The Ribosomal Database Project: improved alignments and new tools for rRNA analysis","volume":"37","author":"Cole","year":"2009","journal-title":"Nucleic Acids Res."},{"key":"2023012710385851200_btt657-B2","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1007\/BF01890115","article-title":"Efficient algorithms for agglomerative hierarchical clustering methods","volume":"1","author":"Day","year":"1984","journal-title":"J. Classif."},{"key":"2023012710385851200_btt657-B3","doi-asserted-by":"crossref","first-page":"2460","DOI":"10.1093\/bioinformatics\/btq461","article-title":"Search and clustering orders of magnitude faster than BLAST","volume":"26","author":"Edgar","year":"2010","journal-title":"Bioinformatics"},{"key":"2023012710385851200_btt657-B4","doi-asserted-by":"crossref","first-page":"1658","DOI":"10.1093\/bioinformatics\/btl158","article-title":"CD-HIT: a fast program for clustering and comparing large sets of protein or nucleotide sequences","volume":"22","author":"Li","year":"2006","journal-title":"Bioinformatics"},{"key":"2023012710385851200_btt657-B5","doi-asserted-by":"crossref","first-page":"1335","DOI":"10.1093\/bioinformatics\/btp157","article-title":"Infernal 1.0: inference of RNA alignments","volume":"25","author":"Nawrocki","year":"2009","journal-title":"Bioinformatics"},{"key":"2023012710385851200_btt657-B6","doi-asserted-by":"crossref","first-page":"7537","DOI":"10.1128\/AEM.01541-09","article-title":"Introducing MOTHUR: open-source, platform-independent, community-supported software for describing and comparing microbial communities","volume":"75","author":"Schloss","year":"2009","journal-title":"Appl. Environ. Microbiol."},{"key":"2023012710385851200_btt657-B7","doi-asserted-by":"crossref","first-page":"e76","DOI":"10.1093\/nar\/gkp285","article-title":"ESPRIT: estimating species richness using large collections of 16S rRNA pyrosequences","volume":"37","author":"Sun","year":"2009","journal-title":"Nucleic Acids Res."},{"key":"2023012710385851200_btt657-B8","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1093\/bib\/bbr009","article-title":"A large-scale benchmark study of existing algorithms for taxonomy-independent microbial community analysis","volume":"13","author":"Sun","year":"2012","journal-title":"Brief. Bioinform."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/2\/287\/48914553\/bioinformatics_30_2_287.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/30\/2\/287\/48914553\/bioinformatics_30_2_287.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T10:39:17Z","timestamp":1674815957000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/30\/2\/287\/223574"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,11,9]]},"references-count":8,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2014,1,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btt657","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2014,1,15]]},"published":{"date-parts":[[2013,11,9]]}}}