{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T00:27:21Z","timestamp":1773275241975,"version":"3.50.1"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2006,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>Various statistical and machine learning methods have been successfully applied to the classification of DNA microarray data. Simple instance-based classifiers such as nearest neighbor (NN) approaches perform remarkably well in comparison to more complex models, and are currently experiencing a renaissance in the analysis of data sets from biology and biotechnology. While binary classification of microarray data has been extensively investigated, studies involving multiclass data are rare. The question remains open whether there exists a significant difference in performance between NN approaches and more complex multiclass methods. Comparative studies in this field commonly assess different models based on their classification accuracy only; however, this approach lacks the rigor needed to draw reliable conclusions and is inadequate for testing the null hypothesis of equal performance. Comparing novel classification models to existing approaches requires focusing on the significance of differences in performance.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results<\/jats:title>\n            <jats:p>We investigated the performance of instance-based classifiers, including a NN classifier able to assign a degree of class membership to each sample. This model alleviates a major problem of conventional instance-based learners, namely the lack of confidence values for predictions. The model translates the distances to the nearest neighbors into 'confidence scores'; the higher the confidence score, the closer is the considered instance to a pre-defined class. We applied the models to three real gene expression data sets and compared them with state-of-the-art methods for classifying microarray data of multiple classes, assessing performance using a statistical significance test that took into account the data resampling strategy. Simple NN classifiers performed as well as, or significantly better than, their more intricate competitors.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>Given its highly intuitive underlying principles \u2013 simplicity, ease-of-use, and robustness \u2013 the <jats:italic>k<\/jats:italic>-NN classifier complemented by a suitable distance-weighting regime constitutes an excellent alternative to more complex models for multiclass microarray data sets. Instance-based classifiers using weighted distances are not limited to microarray data sets, but are likely to perform competitively in classifications of high-dimensional biological data sets such as those generated by high-throughput mass spectrometry.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-7-73","type":"journal-article","created":{"date-parts":[[2006,2,17]],"date-time":"2006-02-17T07:45:46Z","timestamp":1140162346000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":16,"title":["Instance-based concept learning from multiclass DNA microarray data"],"prefix":"10.1186","volume":"7","author":[{"given":"Daniel","family":"Berrar","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ian","family":"Bradbury","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Werner","family":"Dubitzky","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2006,2,16]]},"reference":[{"issue":"3","key":"812_CR1","doi-asserted-by":"publisher","first-page":"227","DOI":"10.1038\/73432","volume":"24","author":"DT Ross","year":"2000","unstructured":"Ross DT, Scherf U, Eisen MB, Perou CM, Rees C, Spellman P, Iyer V, Jeffrey SS, van de Rijn M, Waltham M, Pergamenschikov A, Lee JC, Lashkari D, Shalon D, Myers TG, Weinstein JN, Botstein D, Brown PO: Systematic variation in gene expression patterns in human cancer cell lines. Nat Gen 2000, 24(3):227\u2013235.","journal-title":"Nat Gen"},{"issue":"26","key":"812_CR2","doi-asserted-by":"publisher","first-page":"15149","DOI":"10.1073\/pnas.211566398","volume":"98","author":"S Ramaswamy","year":"2001","unstructured":"Ramaswamy S, Tamayo P, Rifkin R, Mukherjee S, Yeang CH, Angelo MLC, Reich M, Latulippe E, Mesirov JP, Poggio T, Gerald W, Loda M, Lander ES, Golub TR: Multiclass cancer diagnosis using tumor gene expression signatures. Proc Natl Acad Sci USA 2001, 98(26):15149\u201315154.","journal-title":"Proc Natl Acad Sci USA"},{"key":"812_CR3","doi-asserted-by":"publisher","first-page":"133","DOI":"10.1016\/S1535-6108(02)00032-6","volume":"1","author":"EJ Yeoh","year":"2002","unstructured":"Yeoh EJ, Ross ME, Shurtleff SA, Williams WK, Patel D, Mahfouz R, Behm FG, Raimondi SC, Relling MV, Patel A, Cheng C, Campana D, Wilkins D, Zhou X, Li J, Liu H, Pui CH, Evans WE, Naeve C, Wong L, Downing JR: Classification, subtype discovery, and prediction of outcome in pediatric acute lymphoblastic leukemia by gene expression profiling. Cancer Cell 2002, 1: 133\u2013143.","journal-title":"Cancer Cell"},{"issue":"12","key":"812_CR4","doi-asserted-by":"publisher","first-page":"1484","DOI":"10.1093\/bioinformatics\/btg182","volume":"19","author":"RL Somorjai","year":"2003","unstructured":"Somorjai RL, Dolenko B, Baumgartner R: Class prediction and discovery using gene microarray and proteomics mass spectroscopy data: curses, caveats, cautions. Bioinformatics 2003, 19(12):1484\u20131491.","journal-title":"Bioinformatics"},{"key":"812_CR5","first-page":"131","volume-title":"A Practical Approach to Microarray Data Analysis","author":"S Dudoit","year":"2002","unstructured":"Dudoit S, Fridlyand J: Introduction to classification in microarray experiments. In A Practical Approach to Microarray Data Analysis. Edited by: Berrar D, Dubitzky W, Granzow M. Boston: Kluwer Academic Publishers; 2002:131\u2013151."},{"issue":"2","key":"812_CR6","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1145\/980972.980981","volume":"5","author":"S Dudoit","year":"2003","unstructured":"Dudoit S, van der Laan MJ, Kele\u015f S, Molinaro AM, Sinisi SE, Teng SL: Loss-based estimation with cross-validation: applications to microarray data. SIGKDD Explorations 2003, 5(2):56\u201368.","journal-title":"SIGKDD Explorations"},{"key":"812_CR7","doi-asserted-by":"publisher","first-page":"6562","DOI":"10.1073\/pnas.102102699","volume":"98","author":"C Ambroise","year":"2002","unstructured":"Ambroise C, McLachlan GJ: Selection bias in gene extraction on th basis of microarray gene expression data. Proc Natl Acad Sci USA 2002, 98: 6562\u20136566.","journal-title":"Proc Natl Acad Sci USA"},{"issue":"2","key":"812_CR8","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1145\/980972.980978","volume":"5","author":"R Simon","year":"2003","unstructured":"Simon R: Supervised analysis when the number of candidate features ( p ) greatly exceeds the number of cases ( n ). SIGKDD Explorations 2003, 5(2):31\u201336.","journal-title":"SIGKDD Explorations"},{"key":"812_CR9","doi-asserted-by":"publisher","first-page":"559","DOI":"10.1089\/106652700750050943","volume":"7","author":"A Ben-Dor","year":"2000","unstructured":"Ben-Dor A, Bruhn L, Friedman N, Nachman I, Schummer M, Yakhini Z: Tissue classification with gene expression profiles. J Comp Biol 2000, 7: 559\u2013583.","journal-title":"J Comp Biol"},{"issue":"2\u20133","key":"812_CR10","doi-asserted-by":"publisher","first-page":"227","DOI":"10.1089\/1066527041410463","volume":"11","author":"B Krishnapuram","year":"2004","unstructured":"Krishnapuram B, Carin L, Hartemink A: Joint classifier and feature optimization for comprehensive cancer diagnosis using gene expression data. J Comp Bio 2004, 11(2\u20133):227\u2013242.","journal-title":"J Comp Bio"},{"issue":"15","key":"812_CR11","doi-asserted-by":"publisher","first-page":"2429","DOI":"10.1093\/bioinformatics\/bth267","volume":"20","author":"T Li","year":"2004","unstructured":"Li T, Zhang C, Ogihara M: A comparative study of feature selection and multiclass classification methods for tissue classification based on gene expression. Bioinformatics 2004, 20(15):2429\u20132437.","journal-title":"Bioinformatics"},{"issue":"1","key":"812_CR12","doi-asserted-by":"publisher","first-page":"S316","DOI":"10.1093\/bioinformatics\/17.suppl_1.S316","volume":"17","author":"CH Yeang","year":"2001","unstructured":"Yeang CH, Ramaswamy S, Tamayo P, Mukherjee S, Rifkin RM, Angelo M, Reich M, Lander E, Mesirov J, Golub T: Molecular classification of multiple tumor types. Bioinformatics 2001, 17(1):S316-S322.","journal-title":"Bioinformatics"},{"issue":"7","key":"812_CR13","doi-asserted-by":"publisher","first-page":"1895","DOI":"10.1162\/089976698300017197","volume":"10","author":"T Dietterich","year":"1998","unstructured":"Dietterich T: Approximate statistical tests for comparing supervised classification learning algorithms. Neural Comp 1998, 10(7):1895\u20131924.","journal-title":"Neural Comp"},{"key":"812_CR14","doi-asserted-by":"publisher","first-page":"77","DOI":"10.1198\/016214502753479248","volume":"97","author":"S Dudoit","year":"2002","unstructured":"Dudoit S, Fridlyand J, Speed TP: Comparison of discrimination methods for the classification of tumors using gene expression data. J Am Stat Assoc 2002, 97: 77\u201387.","journal-title":"J Am Stat Assoc"},{"issue":"24","key":"812_CR15","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/415436a","volume":"415","author":"SL Pomeroy","year":"2002","unstructured":"Pomeroy SL, Tamayo P, Gaasenbeek M, Sturla LM, Angelo M, McLaughlin ME, Kim JY, Goumnerova LC, Black PM, Lau C, Allen JC, Zagzag D, Olson J, Curran T, Wetmore C, Biegel JA, Poggio T, Mukherjee S, Rifkin R, Califano A, Stolovitzky G, Louis DN, Mesirov JP, Lander ES, Golub TR: Prediction of central nervous system embryonal tumour outcome based on gene expression. Nature 2002, 415(24):436\u2013442.","journal-title":"Nature"},{"key":"812_CR16","first-page":"427","volume-title":"The elements of statistical learning \u2013 Data mining, inference, and prediction","author":"T Hastie","year":"2002","unstructured":"Hastie T, Tibshirani R, Friedman J: The elements of statistical learning \u2013 Data mining, inference, and prediction. New York\/Berlin\/Heidelberg: Springer Series in Statistics; 2002:427\u2013433."},{"key":"812_CR17","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511812651","volume-title":"Pattern recognition and neural networks","author":"BD Ripley","year":"1996","unstructured":"Ripley BD: Pattern recognition and neural networks. Cambridge: University Press; 1996."},{"issue":"12","key":"812_CR18","doi-asserted-by":"publisher","first-page":"1131","DOI":"10.1093\/bioinformatics\/17.12.1131","volume":"17","author":"L Li","year":"2001","unstructured":"Li L, Weinberg CR, Darden TA, Pedersen LG: Gene selection for sample classification based on gene expression data: study of sensitivity to choice of parameters of the GA\/KNN method. Bioinformatics 2001, 17(12):1131\u20131142.","journal-title":"Bioinformatics"},{"issue":"1","key":"812_CR19","doi-asserted-by":"publisher","first-page":"79","DOI":"10.1016\/j.mbs.2004.07.002","volume":"193","author":"CA Tsai","year":"2005","unstructured":"Tsai CA, Lee TC, Ho IC, Yang UC, Chen CH, Chen JJ: Multi-class clustering and prediction in the analysis of microarray data. Math Biosci 2005, 193(1):79\u2013100.","journal-title":"Math Biosci"},{"key":"812_CR20","first-page":"216","volume-title":"A Practical Approach to Microarray Data Analysis","author":"L Li","year":"2002","unstructured":"Li L, Weinberg CR: Gene selection and sample classification using a genetic algorithm and k-nearest neighbor method. In A Practical Approach to Microarray Data Analysis. Edited by: Berrar D, Dubitzky W, Granzow M. Boston: Kluwer Academic Publishers; 2002:216\u2013229."},{"key":"812_CR21","doi-asserted-by":"publisher","first-page":"60","DOI":"10.1186\/1471-2105-4-60","volume":"4","author":"J Wang","year":"2003","unstructured":"Wang J, Bo TH, Jonassen I, Myklebost O, Hovig E: Tumor classification and marker gene prediction by feature selection and fuzzy c-means clustering using microarray data. BMC Bioinformatics 2003, 4: 60.","journal-title":"BMC Bioinformatics"},{"issue":"5","key":"812_CR22","doi-asserted-by":"publisher","first-page":"644","DOI":"10.1093\/bioinformatics\/bti036","volume":"21","author":"MH Asyali","year":"2005","unstructured":"Asyali MH, Alci M: Reliability analysis of microarray data using fuzzy c-means and normal mixture modeling based classification methods. Bioinformatics 2005, 21(5):644\u2013649.","journal-title":"Bioinformatics"},{"key":"812_CR23","doi-asserted-by":"publisher","first-page":"239","DOI":"10.1023\/A:1024068626366","volume":"52","author":"C Nadeau","year":"2003","unstructured":"Nadeau C, Bengio Y: Inference for generalization error. Machine Learning 2003, 52: 239\u2013281.","journal-title":"Machine Learning"},{"key":"812_CR24","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1007\/978-3-540-24775-3_3","volume-title":"Proceedings of the Eighth Pacific-Asia Conference on Knowledge Discovery and Data Mining: 26\u201328 May 2004, Sydney, Australia","author":"R Bouckaert","year":"2004","unstructured":"Bouckaert R, Frank E: Evaluating the replicability of significance tests for comparing learning algorithms. In Proceedings of the Eighth Pacific-Asia Conference on Knowledge Discovery and Data Mining: 26\u201328 May 2004, Sydney, Australia. Edited by: Dai H, Srikant R, Zhang C. Sydney, Australia: Springer; 2004:3\u201312."},{"issue":"3","key":"812_CR25","doi-asserted-by":"publisher","first-page":"236","DOI":"10.1038\/73439","volume":"24","author":"U Scherf","year":"2000","unstructured":"Scherf U, Ross D, Waltham M, Smith L, Lee J, Tanabe L, Kohn K, Reinhold W, Myers T, Andrews D, Scudiero D, Eisen M, Sausville E, Pommier Y, Botstein D, Brown P, Weinstein J: A gene expression database for the molecular pharmacology of cancer. Nat Gen 2000, 24(3):236\u2013244.","journal-title":"Nat Gen"},{"key":"812_CR26","volume-title":"Multivariate Analysis","author":"KV Mardia","year":"1980","unstructured":"Mardia KV, Kent JT, J Bibby M: Multivariate Analysis. Academic Press: London; 1980."},{"key":"812_CR27","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1145\/332306.332564","volume-title":"Proceedings of the Fourth Annual International Conference on Computational Molecular Biology: 8\u201311 April 2000; Tokyo, Japan","author":"D Slonim","year":"2000","unstructured":"Slonim D, Tamayo P, Mesirov J, Golub T, Lander E: Class prediction and discovery using gene expression data. In Proceedings of the Fourth Annual International Conference on Computational Molecular Biology: 8\u201311 April 2000; Tokyo, Japan. Edited by: Shamir R, Miyano S, Istrail S, Pevzner P, Waterman M. Universal Academy Press; 2000:263\u2013272."},{"issue":"18","key":"812_CR28","doi-asserted-by":"publisher","first-page":"10101","DOI":"10.1073\/pnas.97.18.10101","volume":"97","author":"O Alter","year":"2000","unstructured":"Alter O, Brown PO, Botstein D: Singular-value decomposition for genome-wide expression data processing and modeling. Proc Natl Acad Sci USA 2000, 97(18):10101\u201310106.","journal-title":"Proc Natl Acad Sci USA"},{"issue":"3","key":"812_CR29","doi-asserted-by":"publisher","first-page":"505","DOI":"10.1089\/106652702760138592","volume":"9","author":"MD Radmacher","year":"2002","unstructured":"Radmacher MD, McShane LM, Simon R: A paradigm for class prediction using gene expression profiles. J Comp Biol 2002, 9(3):505\u2013511.","journal-title":"J Comp Biol"},{"key":"812_CR30","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1111\/j.1432-1033.1997.00225.x","volume":"248","author":"H Rechreche","year":"1997","unstructured":"Rechreche H, Mallo GV, Montalto G, Dagorn JC, Iovanna JL: Cloning and expression of the mRNA of human galectin-4, an S-type lectin down-regulated in colorectal cancer. Europ J Biochem 1997, 248: 225\u2013230.","journal-title":"Europ J Biochem"},{"key":"812_CR31","first-page":"420","volume-title":"Proceedings of the Eighth International Conference on Database Theory (ICDT): 4\u20136 January 2001, London, UK","author":"CC Aggarwal","year":"2001","unstructured":"Aggarwal CC, Hinneburg A, Keim DA: On the surprising behavior of distance metrics in high dimensional space. In Proceedings of the Eighth International Conference on Database Theory (ICDT): 4\u20136 January 2001, London, UK. Edited by: Van den Bussche J, Vianu V. Springer; 2001:420\u2013434."},{"key":"812_CR32","volume-title":"Statistical Learning Theory","author":"V Vapnik","year":"1998","unstructured":"Vapnik V: Statistical Learning Theory. New York: John Wiley & Sons; 1998."},{"key":"812_CR33","unstructured":"Cawley GC:Support Vector Machine Toolbox (v0.55b). University of East Anglia, School of Information Systems, Norwich, Norfolk, UK, NR4 7TJ; [http:\/\/theoval.sys.uea.ac.uk\/~gcc\/svm\/toolbox\/]"},{"key":"812_CR34","first-page":"547","volume-title":"Advances in Neural Information Processing Systems","author":"J Platt","year":"2000","unstructured":"Platt J, Christianini N, Shawe-Taylor J: Large margin DAGs for multiclass classification. In Advances in Neural Information Processing Systems. Volume 12. Edited by: Solla SA, Leen TK, Mueller KR. Cambridge, MA: MIT Press; 2000:547\u2013553."},{"key":"812_CR35","first-page":"166","volume-title":"A Practical Approach to Microarray Data Analysis","author":"S Mukherjee","year":"2002","unstructured":"Mukherjee S: Classifying microarray data using support vector machines. In A Practical Approach to Microarray Data Analysis. Edited by: Berrar D, Dubitzky W, Granzow M. Boston: Kluwer Academic Publishers; 2002:166\u2013185."},{"issue":"12","key":"812_CR36","doi-asserted-by":"publisher","first-page":"6730","DOI":"10.1073\/pnas.111153698","volume":"98","author":"H Zhang","year":"2001","unstructured":"Zhang H, Yu CH, Singer B, Xiong M: Recursive partitioning for tumor classification with gene expression microarray data. Proc Natl Acad Sci USA 2001, 98(12):6730\u20136735.","journal-title":"Proc Natl Acad Sci USA"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-7-73.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T03:12:41Z","timestamp":1630465961000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-7-73"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,2,16]]},"references-count":36,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2006,12]]}},"alternative-id":["812"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-7-73","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2006,2,16]]},"assertion":[{"value":"8 April 2005","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 February 2006","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 February 2006","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"73"}}