{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,2]],"date-time":"2025-08-02T17:25:18Z","timestamp":1754155518173,"version":"3.41.2"},"reference-count":26,"publisher":"Emerald","issue":"4","license":[{"start":{"date-parts":[[2013,4,19]],"date-time":"2013-04-19T00:00:00Z","timestamp":1366329600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.emerald.com\/insight\/site-policies"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2013,4,19]]},"abstract":"<jats:sec><jats:title content-type=\"abstract-heading\">Purpose<\/jats:title><jats:p>The<jats:italic>K<\/jats:italic>\u2010means clustering algorithm has been intensely researched owing to its simplicity of implementation and usefulness in the clustering task. However, there have also been criticisms on its performance, in particular, for demanding the value of<jats:italic>K<\/jats:italic>before the actual clustering task. It is evident from previous researches that providing the number of clusters a priori does not in any way assist in the production of good quality clusters. The authors' investigations in this paper also confirm this finding. The purpose of this paper is to investigate further, the usefulness of the<jats:italic>K<\/jats:italic>\u2010means clustering in the clustering of high and multi\u2010dimensional data by applying it to biological sequence data.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Design\/methodology\/approach<\/jats:title><jats:p>The authors suggest a scheme which maps the high dimensional data into low dimensions, then show that the<jats:italic>K<\/jats:italic>\u2010means algorithm with pre\u2010processor produces good quality, compact and well\u2010separated clusters of the biological data mapped in low dimensions. For the purpose of clustering, a character\u2010to\u2010numeric conversion was conducted to transform the nucleic\/amino acids symbols to numeric values.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Findings<\/jats:title><jats:p>A preprocessing technique has been suggested.<\/jats:p><\/jats:sec><jats:sec><jats:title content-type=\"abstract-heading\">Originality\/value<\/jats:title><jats:p>Conceptually this is a new paper with new results.<\/jats:p><\/jats:sec>","DOI":"10.1108\/k-02-2013-0028","type":"journal-article","created":{"date-parts":[[2013,6,20]],"date-time":"2013-06-20T09:34:19Z","timestamp":1371720859000},"page":"614-627","source":"Crossref","is-referenced-by-count":2,"title":["An investigation of<i>K<\/i>\u2010means clustering to high and multi\u2010dimensional biological data"],"prefix":"10.1108","volume":"42","author":[{"given":"Barile\u00e9 B.","family":"Baridam","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"M.","family":"Montaz Ali","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","reference":[{"key":"key2022031220201278000_b24","doi-asserted-by":"crossref","unstructured":"Andreopoulos, B., An, A. and Wang, X. (2006), \u201cBi\u2010level clustering of mixed categorical and numerical biomedical data\u201d, International Journal of Data Mining and Bioinformatics, Vol. 1 No. 1, pp. 19\u201056.","DOI":"10.1504\/IJDMB.2006.009920"},{"key":"key2022031220201278000_b21","doi-asserted-by":"crossref","unstructured":"Azuaje, F. (2002), \u201cCluster validity framework for genome expression data\u201d, Bioinformatics, Vol. 18 No. 2.","DOI":"10.1093\/bioinformatics\/18.2.319"},{"key":"key2022031220201278000_b25","unstructured":"Baridam, B.B. and Owolabi, O. (2010), \u201cConceptual clustering of RNA sequences with the codon usage model\u201d, Global Journal of Computer Science and Technology, Vol. 10 No. 8, pp. 41\u201045."},{"key":"key2022031220201278000_b1","unstructured":"Berkhin, P. (2002), \u201cSurvey of clustering data mining techniques\u201d, Technical Report No. 4, Accrue Software, Inc., San Jose, CA."},{"key":"key2022031220201278000_b10","doi-asserted-by":"crossref","unstructured":"Bezdek, J.C. (1980), \u201cA convergence theorem for the fuzzy ISODATA clustering algorithms\u201d, IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol. 2, pp. 1\u20108.","DOI":"10.1109\/TPAMI.1980.4766964"},{"key":"key2022031220201278000_b2","unstructured":"Binder, D.A. (1977), \u201cCluster analysis under parametric models\u201d, PhD thesis, University of London, London."},{"key":"key2022031220201278000_b23","doi-asserted-by":"crossref","unstructured":"Bolshakova, N. and Azuaje, F. (2003), \u201cCluster validation techniques for genome expression data\u201d, Signal Processing, Vol. 83, pp. 825\u2010833.","DOI":"10.1016\/S0165-1684(02)00475-9"},{"key":"key2022031220201278000_b12","doi-asserted-by":"crossref","unstructured":"Bourne, P.E. and Weissig, H. (Eds) (2003), Structural Bioinformatics, Wiley\u2010Liss, Inc., Hoboken, NJ, pp. 35\u201049.","DOI":"10.1002\/0471721204"},{"key":"key2022031220201278000_b22","unstructured":"Gonz\u00e1lez Teledo, M.D. (2005), \u201cA comparison in cluster validation techniques\u201d, MSc thesis, Department of Mathematics (Statistics), University of Puerto Rico."},{"key":"key2022031220201278000_b17","doi-asserted-by":"crossref","unstructured":"Gupta, S.R., Rao, K.S. and Bhatnagar, V. (1999), \u201cK\u2010means clustering algorithm for categorical attributes\u201d, Proceedings of 1st International Conference on Data Warehousing and Knowledge Discovery, Florence, Italy, pp. 203\u2010208.","DOI":"10.1007\/3-540-48298-9_22"},{"key":"key2022031220201278000_b7","unstructured":"Huang, K. (2002), \u201cA synergistic automatic clustering technique (SYNERACT) for multispectral image analysis\u201d, Photogrammetric Engineering and Remote Sensing, Vol. 1 No. 1, pp. 33\u201040."},{"key":"key2022031220201278000_b19","doi-asserted-by":"crossref","unstructured":"Kaufman, L. and Rousseuw, P. (1990), Finding Groups in Data: An Introduction to Cluster Analysis, Wiley, New York, NY.","DOI":"10.1002\/9780470316801"},{"key":"key2022031220201278000_b4","unstructured":"MacQueen, J.B. (1967), \u201cSome methods for classification and analysis of multivariate observations\u201d, Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability, University of California Press, Berkeley, CA, pp. 281\u2010297."},{"key":"key2022031220201278000_b18","unstructured":"MATLAB (2004), \u201cThe language of technical computing\u201d, Version 7.0, The MathWorks, Inc., Natick, MA."},{"key":"key2022031220201278000_b13","unstructured":"National Human Genome Research Institute (2007), \u201cThe structure of ribonucleic and deoxyribonucleic acids\u201d, National Institutes of Health, Division of Intramural Research, available at: www.nhgri.gov."},{"key":"key2022031220201278000_b11","unstructured":"Omran, M.G.H. (2004), \u201cParticle swarm optimization methods for pattern recognition and image processing\u201d, PhD thesis, Department of Computer Science, Faculty of Engineering, Built Environment and Information Technology, University of Pretoria."},{"key":"key2022031220201278000_b14","doi-asserted-by":"crossref","unstructured":"Ramoni, M.F., Sebastiani, P. and Kohane, I.I. (2002), \u201cCluster analysis of gene expression dynamics\u201d, Proceedings of National Academy of Science, Vol. 99, July, pp. 9121\u20109126.","DOI":"10.1073\/pnas.132656399"},{"key":"key2022031220201278000_b9","unstructured":"Rosenberger, C. and Chehdi, K. (2000), \u201cUnsupervised clustering method with optimal estimation of the number of clusters: application to image segmentation\u201d, Proceedings of the International Conference on Pattern Recognition (ICPR'00), pp. 1656\u20101659."},{"key":"key2022031220201278000_b20","doi-asserted-by":"crossref","unstructured":"Rousseuw, P. (1987), \u201cSilhouettes: a practical aid to the interpretation and validation of cluster analysis\u201d, Computational and Applied Mathematics, Vol. 20.","DOI":"10.1016\/0377-0427(87)90125-7"},{"key":"key2022031220201278000_b5","doi-asserted-by":"crossref","unstructured":"Smet, F.D., Mathys, J., Marchal, K., Thijs, G., Moor, B.D. and Moreau, Y. (2002), \u201cAdaptive quality\u2010based clustering of gene expression profiles\u201d, Bioinformatics, Vol. 18 No. 6, pp. 735\u2010748.","DOI":"10.1093\/bioinformatics\/18.5.735"},{"key":"key2022031220201278000_b26","doi-asserted-by":"crossref","unstructured":"Sonstegard, T., Capuco, A.V., White, J., Van Tastell, C.P., Connor, E.E., Cho, J., Sultana, R., Shade, L., Wray, J.E., Wells, K.D. and Quackenbush, J. (2002), \u201cAnalysis of bovine mammary gland EST and functional annotation of the Bos Taurus gene index\u201d, Mammary Genome, Vol. 13 No. 7, pp. 373\u2010379.","DOI":"10.1007\/s00335-001-2145-4"},{"key":"key2022031220201278000_b8","doi-asserted-by":"crossref","unstructured":"Tou, J. (1979), \u201cDYNOC \u2013 a dynamic optimal cluster\u2010seeking technique\u201d, International Journal of Computer and Information Sciences, Vol. 8 No. 6, pp. 541\u2010547.","DOI":"10.1007\/BF00995502"},{"key":"key2022031220201278000_b3","unstructured":"Tou, J. and Gonz\u00e1lez, R. (1974), Pattern Recognition Principles, Addison\u2010Wesley, Boston, MA."},{"key":"key2022031220201278000_b6","unstructured":"Turi, R.H. (2001), \u201cClustering\u2010based colour image segmentation\u201d, PhD thesis, Monash University."},{"key":"key2022031220201278000_b16","unstructured":"Xu, R. and Wunsch, D. II (2005), \u201cSurvey of clustering algorithms\u201d, International Journal of Intelligent Computing and Cybernetics, Vol. 16 No. 3, pp. 601\u2010614."},{"key":"key2022031220201278000_b15","doi-asserted-by":"crossref","unstructured":"Xu, Y., Olman, V. and Xu, D. (2002), \u201cClustering gene expression data using a graph theoretic approach: an application of minimum spanning trees\u201d, Bioinformatics, Vol. 18 No. 4, pp. 536\u2010545.","DOI":"10.1093\/bioinformatics\/18.4.536"}],"container-title":["Kybernetes"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.emeraldinsight.com\/doi\/full-xml\/10.1108\/K-02-2013-0028","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/K-02-2013-0028\/full\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/K-02-2013-0028\/full\/html","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,24]],"date-time":"2025-07-24T21:46:43Z","timestamp":1753393603000},"score":1,"resource":{"primary":{"URL":"http:\/\/www.emerald.com\/k\/article\/42\/4\/614-627\/270699"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,4,19]]},"references-count":26,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2013,4,19]]}},"alternative-id":["10.1108\/K-02-2013-0028"],"URL":"https:\/\/doi.org\/10.1108\/k-02-2013-0028","relation":{},"ISSN":["0368-492X"],"issn-type":[{"type":"print","value":"0368-492X"}],"subject":[],"published":{"date-parts":[[2013,4,19]]}}}