{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T02:53:00Z","timestamp":1782442380469,"version":"3.54.5"},"reference-count":19,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2008,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>During the last decade, the use of microarrays to assess the transcriptome of many biological systems has generated an enormous amount of data. A common technique used to organize and analyze microarray data is to perform cluster analysis. While many clustering algorithms have been developed, they all suffer a significant decrease in computational performance as the size of the dataset being analyzed becomes very large. For example, clustering 10000 genes from an experiment containing 200 microarrays can be quite time consuming and challenging on a desktop PC. One solution to the scalability problem of clustering algorithms is to distribute or parallelize the algorithm across multiple computers.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>The software described in this paper is a high performance multithreaded application that implements a parallelized version of the K-means Clustering algorithm. Most parallel processing applications are not accessible to the general public and require specialized software libraries (e.g. MPI) and specialized hardware configurations. The parallel nature of the application comes from the use of a web service to perform the distance calculations and cluster assignments. Here we show our parallel implementation provides significant performance gains over a wide range of datasets using as little as seven nodes. The software was written in C# and was designed in a modular fashion to provide both deployment flexibility as well as flexibility in the user interface.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>ParaKMeans was designed to provide the general scientific community with an easy and manageable client-server application that can be installed on a wide variety of Windows operating systems.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-9-200","type":"journal-article","created":{"date-parts":[[2008,4,16]],"date-time":"2008-04-16T18:13:30Z","timestamp":1208369610000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["ParaKMeans: Implementation of a parallelized K-means algorithm suitable for general laboratory use"],"prefix":"10.1186","volume":"9","author":[{"given":"Piotr","family":"Kraj","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ashok","family":"Sharma","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nikhil","family":"Garge","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Robert","family":"Podolsky","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Richard A","family":"McIndoe","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2008,4,16]]},"reference":[{"key":"2185_CR1","volume-title":"Proc 1997 IEEE International Conference on Parallel and Distributed Systems,","author":"H Tsai","year":"1997","unstructured":"Tsai H, Horng S, Tsai S, Lee S, Kao T, Chen C: Parallel clustering algorithms on a reconfigurable array of processors with wider bus networks. Proc 1997 IEEE International Conference on Parallel and Distributed Systems, 1997."},{"key":"2185_CR2","first-page":"243","volume":"1","author":"S Kantabutra","year":"2000","unstructured":"Kantabutra S, Couch AL: Parallel K-means clustering algorithm on NOWs. NOCTEC Technical Journal 2000, 1: 243\u2013248.","journal-title":"NOCTEC Technical Journal"},{"key":"2185_CR3","doi-asserted-by":"publisher","first-page":"245","DOI":"10.1007\/3-540-46502-2_13","volume":"1759","author":"IS Dhillon","year":"1999","unstructured":"Dhillon IS, Modha DS: A data-clustering algorithm on distributed memory multiprocessors. Large-Scale Parallel Data Mining, 1999, 1759: 245\u2013260.","journal-title":"Large-Scale Parallel Data Mining,"},{"key":"2185_CR4","volume-title":"A Hybrid Parallel Web Document Clustering Algorithm and Its Performance Study","author":"S Xu","year":"2003","unstructured":"Xu S, Zhang J: A Hybrid Parallel Web Document Clustering Algorithm and Its Performance Study.2003. [http:\/\/www.cs.uky.edu\/~jzhang\/pub\/MINING\/ppddp.ps.gz]"},{"key":"2185_CR5","unstructured":"Gursoy A: Data Decomposition for Parallel K-means: 2003. Parallel Processing and Applied Mathematics. Edited by: Wyrzykowski R, Dongarra J, Paprzycki M and Wasniewski J. Edited by: Wyrzykowski R. Czestochowa, Poland, Springer-Verlag, Berlin Heidelberg; 2004:241\u2013248."},{"key":"2185_CR6","volume-title":"Parallel K-Means Data Clustering","author":"W Liao","year":"2005","unstructured":"Liao W: Parallel K-Means Data Clustering.2005. [http:\/\/www.ece.northwestern.edu\/~wkliao\/Kmeans\/index.html]"},{"key":"2185_CR7","doi-asserted-by":"publisher","first-page":"19","DOI":"10.1007\/s11227-006-0002-7","volume":"39","author":"L Yanjun","year":"2007","unstructured":"Yanjun L, Soon C: Parallel bisecting k-means with prediction clustering algorithm. The Journal of Supercomputing 2007, 39: 19\u201337.","journal-title":"The Journal of Supercomputing"},{"key":"2185_CR8","unstructured":"de Souza PSL, Britto AS, Sabourin R, de Souza SRS, Borges DL: K-Means VQ Algorithm using a Low-Cost Parallel Cluster Computing: 2004\/2\/19. Parallel and Distributed Computing and Networks 2004. Edited by: Hamza MH. Edited by: Hamza MH. Innsbruck, Austria; 2004:420\u2013124."},{"key":"2185_CR9","volume-title":"Proc of the 4th European Conference on Principles of Data Mining and Knowledge Discovery (PKDD 2000) Lyon, France","author":"B Zhang","year":"2000","unstructured":"Zhang B, Hsu M, Forman G: Accurate recasting of parameter estimation algorithms using sufficient statistics for efficient parallel speed-up: Demonstrated for center-based data clustering algorithms. Proc of the 4th European Conference on Principles of Data Mining and Knowledge Discovery (PKDD 2000) Lyon, France 2000."},{"key":"2185_CR10","volume-title":"Parallel K-Means Algorithm on Distributed Memory Multiprocessors","author":"J M.N.","year":"2003","unstructured":"M.N. J: Parallel K-Means Algorithm on Distributed Memory Multiprocessors.2003. [http:\/\/www-users.cs.umn.edu\/~mnjoshi\/PKMeans.pdf]"},{"key":"2185_CR11","doi-asserted-by":"publisher","first-page":"258","DOI":"10.1016\/j.vph.2006.08.003","volume":"45","author":"CD Collins","year":"2006","unstructured":"Collins CD, Purohit S, Podolsky RH, Zhao HS, Schatz D, Eckenrode SE, Yang P, Hopkins D, Muir A, Hoffman M, McIndoe RA, Rewers M, She JX: The application of genomic and proteomic technologies in predictive, preventive and personalized medicine. Vascul Pharmacol 2006, 45: 258\u2013267.","journal-title":"Vascul Pharmacol"},{"key":"2185_CR12","doi-asserted-by":"publisher","first-page":"14863","DOI":"10.1073\/pnas.95.25.14863","volume":"95","author":"MB Eisen","year":"1998","unstructured":"Eisen MB, Spellman PT, Brown PO, Botstein D: Cluster analysis and display of genome-wide expression patterns. Proc Natl Acad Sci U S A 1998, 95: 14863\u201314868.","journal-title":"Proc Natl Acad Sci U S A"},{"key":"2185_CR13","doi-asserted-by":"publisher","first-page":"441","DOI":"10.1207\/s15327906mbr2104_5","volume":"21","author":"GW Milligan","year":"1986","unstructured":"Milligan GW, Cooper MC: A study of the comparability of external criteria for heirarchical cluster analysis. Multivar Behav Res 1986, 21: 441\u2013458.","journal-title":"Multivar Behav Res"},{"key":"2185_CR14","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1007\/BF01908075","volume":"2","author":"L Hubert","year":"1985","unstructured":"Hubert L, Arabie P: Comparing Partitions. J Classification 1985, 2: 193\u2013218.","journal-title":"J Classification"},{"key":"2185_CR15","doi-asserted-by":"publisher","first-page":"846","DOI":"10.1080\/01621459.1971.10482356","volume":"66","author":"WM Rand","year":"1971","unstructured":"Rand WM: Objective criteria for the evaluation of clustering methods. J Am Stat Assoc 1971, 66: 846\u2013850.","journal-title":"J Am Stat Assoc"},{"key":"2185_CR16","doi-asserted-by":"publisher","first-page":"R34","DOI":"10.1186\/gb-2003-4-5-r34","volume":"4","author":"KY Yeung","year":"2003","unstructured":"Yeung KY, Medvedovic M, Bumgarner RE: Clustering gene-expression data with repeated measurements. Genome Biol 2003, 4: R34.","journal-title":"Genome Biol"},{"key":"2185_CR17","doi-asserted-by":"publisher","first-page":"2405","DOI":"10.1093\/bioinformatics\/btl406","volume":"22","author":"A Thalamuthu","year":"2006","unstructured":"Thalamuthu A, Mukhopadhyay I, Zheng X, Tseng GC: Evaluation and comparison of gene clustering methods in microarray analysis. Bioinformatics 2006, 22: 2405\u20132412.","journal-title":"Bioinformatics"},{"key":"2185_CR18","unstructured":"ParaKMeans Help2008. [http:\/\/bioanalysis.genomics.mcg.edu\/parakmeans\/help\/webframe.html]"},{"key":"2185_CR19","unstructured":"ParaKMeans Public Use2008. [http:\/\/bioanalysis.genomics.mcg.edu\/parakmeans]"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-9-200.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T11:00:10Z","timestamp":1630494010000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-9-200"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2008,4,16]]},"references-count":19,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2008,12]]}},"alternative-id":["2185"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-9-200","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2008,4,16]]},"assertion":[{"value":"7 September 2007","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 April 2008","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 April 2008","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"200"}}