{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T13:57:52Z","timestamp":1780063072770,"version":"3.54.0"},"reference-count":26,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2004,10,26]],"date-time":"2004-10-26T00:00:00Z","timestamp":1098748800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0"},{"start":{"date-parts":[[2004,10,26]],"date-time":"2004-10-26T00:00:00Z","timestamp":1098748800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/2.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                        <jats:title>Background<\/jats:title>\n                        <jats:p>The imputation of missing values is necessary for the efficient use of DNA microarray data, because many clustering algorithms and some statistical analysis require a complete data set. A few imputation methods for DNA microarray data have been introduced, but the efficiency of the methods was low and the validity of imputed values in these methods had not been fully checked.<\/jats:p>\n                     <\/jats:sec><jats:sec>\n                        <jats:title>Results<\/jats:title>\n                        <jats:p>We developed a new cluster-based imputation method called sequential K-nearest neighbor (SKNN) method. This imputes the missing values sequentially from the gene having least missing values, and uses the imputed values for the later imputation. Although it uses the imputed values, the efficiency of this new method is greatly improved in its accuracy and computational complexity over the conventional KNN-based method and other methods based on maximum likelihood estimation. The performance of SKNN was in particular higher than other imputation methods for the data with high missing rates and large number of experiments.<\/jats:p>\n                        <jats:p>Application of Expectation Maximization (EM) to the SKNN method improved the accuracy, but increased computational time proportional to the number of iterations. The Multiple Imputation (MI) method, which is well known but not applied previously to microarray data, showed a similarly high accuracy as the SKNN method, with slightly higher dependency on the types of data sets.<\/jats:p>\n                     <\/jats:sec><jats:sec>\n                        <jats:title>Conclusions<\/jats:title>\n                        <jats:p>Sequential reuse of imputed data in KNN-based imputation greatly increases the efficiency of imputation. The SKNN method should be practically useful to save the data of some microarray experiments which have high amounts of missing entries. The SKNN method generates reliable imputed values which can be used for further cluster-based analysis of microarray data.<\/jats:p>\n                     <\/jats:sec>","DOI":"10.1186\/1471-2105-5-160","type":"journal-article","created":{"date-parts":[[2004,10,28]],"date-time":"2004-10-28T16:38:47Z","timestamp":1098981527000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":151,"title":["Reuse of imputed data in microarray analysis increases imputation efficiency"],"prefix":"10.1186","volume":"5","author":[{"given":"Ki-Yeol","family":"Kim","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Byoung-Jin","family":"Kim","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gwan-Su","family":"Yi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2004,10,26]]},"reference":[{"key":"276_CR1","first-page":"111","volume":"12","author":"S Dudoit","year":"2002","unstructured":"Dudoit S, Yang YH, Callow MJ, Speed TP: Statistical methods for identifying differentially expressed genes in replicated cDNA microarray experiments.\n                           Statistica Sinica 2002, 12: 111\u2013139.","journal-title":"Statistica Sinica"},{"issue":"3","key":"276_CR2","doi-asserted-by":"publisher","first-page":"503","DOI":"10.1038\/35000501","volume":"403","author":"AA Alizadeh","year":"2000","unstructured":"Alizadeh AA, Eisen MB, Davis RE, Ma C, Lossos IS, Rosenwald A, Boldrick JC, Sabet H, Tran T, Yu X, Powell JI, Yang L, Marti GE, Moore T, Jr JH, Lu L, Lewis DB, Tibshirani R, Sherlock G, Chan WC, Greiner TC, Weisenburger DD, Armitage JO, Warnke R, Levy R, Wilson W, Grever MR, Byrd JC, Botstein D, Brown PO, Staudt LM: Distinct types of diffuse large B-cell lymphoma identified by gene expression profiling.\n                           Nature 2000, 403(3):503\u2013511.","journal-title":"Nature"},{"key":"276_CR3","doi-asserted-by":"publisher","first-page":"R34","DOI":"10.1186\/gb-2003-4-5-r34","volume":"4","author":"KY Yeung","year":"2003","unstructured":"Yeung KY, Medvedovic M, Bumgarner RE: Clustering gene-expression data with repeated measurements.\n                           Genome Biol 2003, 4: R34. 10.1186\/gb-2003-4-5-r34","journal-title":"Genome Biol"},{"key":"276_CR4","first-page":"14863","volume-title":"Proc Natl Acad Sci U S A","author":"MB Eisen","year":"1998","unstructured":"Eisen MB, Spellman PT, Brown PO, Botstein D: Cluster analysis and display of genome-wide expression patterns.\n                           Proc Natl Acad Sci U S A 1998, 14863\u201314868. 10.1073\/pnas.95.25.14863"},{"issue":"6","key":"276_CR5","doi-asserted-by":"publisher","first-page":"418","DOI":"10.1038\/35076576","volume":"2","author":"J Quackenbush","year":"2001","unstructured":"Quackenbush J: Computational analysis of microarray data.\n                           Nat Rev Genet 2001, 2(6):418\u2013427. 10.1038\/35076576","journal-title":"Nat Rev Genet"},{"key":"276_CR6","volume-title":"Tech rep","author":"TM Beasley","year":"1998","unstructured":"Beasley TM: Comments on the Analysis of Data with Missing Values.\n                           Tech rep 1998."},{"key":"276_CR7","volume-title":"Statistical Analysis With Missing Data","author":"R Little","year":"1987","unstructured":"Little R, Rubin D: Statistical Analysis With Missing Data. Wiley, New Work; 1987."},{"issue":"3","key":"276_CR8","doi-asserted-by":"publisher","first-page":"380","DOI":"10.1076\/clin.15.3.380.10266","volume":"15","author":"V Narhi","year":"2001","unstructured":"Narhi V, Laassonen S, Hietala R, Ahonen T, Lyyti H: Treating missing data in a clinical neuropsychological dataset-data imputation.\n                           The Clinical Neuropsychologist 2001, 15(3):380\u2013392. 10.1076\/clin.15.3.380.10266","journal-title":"The Clinical Neuropsychologist"},{"key":"276_CR9","doi-asserted-by":"publisher","first-page":"4241","DOI":"10.1091\/mbc.11.12.4241","volume":"11","author":"AP Gasch","year":"2000","unstructured":"Gasch AP, Spellman PT, Kao CM, Carmel-Harel O, Eisen MB, Storz G, Botstein D, Brown PO: Genomic expression programs in the response of yeast cells to environmental changes.\n                           Molecular Biology of the Cell 2000, 11: 4241\u20134257.","journal-title":"Molecular Biology of the Cell"},{"issue":"24","key":"276_CR10","doi-asserted-by":"publisher","first-page":"13784","DOI":"10.1073\/pnas.241500798","volume":"98","author":"ME Garber","year":"2001","unstructured":"Garber ME, Troyanskaya OG, Schluens K, Petersen S, Thaesler Z, Pacyna-Gengelbach M, van de Rijn M, Rosen GD, Perou CM, Whyte RI, Altman RB, Brown PO, Botstein D, Petersen I: Diversity of gene expression in adenocarcinoma of the lung.\n                           Proc Natl Acad Sci U S A 2001, 98(24):13784\u201313789. 10.1073\/pnas.241500798","journal-title":"Proc Natl Acad Sci U S A"},{"issue":"4","key":"276_CR11","doi-asserted-by":"publisher","first-page":"1926","DOI":"10.1073\/pnas.0437875100","volume":"100","author":"SP Bohen","year":"2003","unstructured":"Bohen SP, Troyanskaya OG, Alter O, Warnke R, Botstein D, Brown PO, Levy R: Variation in gene expression patterns in follicular lymphoma and the response to rituximab.\n                           Proc Natl Acad Sci U S A 2003, 100(4):1926\u20131930. 10.1073\/pnas.0437875100","journal-title":"Proc Natl Acad Sci U S A"},{"key":"276_CR12","first-page":"18","volume-title":"In Conference of European Statistics","author":"SS Kuzin","year":"2000","unstructured":"Kuzin SS: Data imputation based on regression models with variayions of entropy.\n                           In Conference of European Statistics 2000, 18\u201320."},{"key":"276_CR13","first-page":"580","volume-title":"In Proceeding of the Survey Research Methods Section, American Statistical Asssociation","author":"JF Beaumont","year":"2000","unstructured":"Beaumont JF: On regression imputation in the presence of nonignorable nonresponse.\n                           In Proceeding of the Survey Research Methods Section, American Statistical Asssociation 2000, 580\u2013585."},{"key":"276_CR14","first-page":"1491","volume-title":"In Proceedings of the International Conference on Spoken Language Precessing","author":"B Raj","year":"1998","unstructured":"Raj B, Singh R, Stern RM: Inference of missing spectrographic features for robust speech recognition.\n                           In Proceedings of the International Conference on Spoken Language Precessing 1998, 1491\u20131494."},{"key":"276_CR15","volume-title":"Classification and Regressions Trees","author":"L Breiman","year":"1984","unstructured":"Breiman L, Friedman JH, Olshen RA, Stone CJ: Classification and Regressions Trees. Chapman & Hall Inc; 1984."},{"key":"276_CR16","volume-title":"In PKDD 99, 3rd European Conference of Principles and Practice of Knowledge Discovery in Databases","author":"A Feelders","year":"1999","unstructured":"Feelders A: Cluster analysis and display of genome-wide expression patterns.\n                           In PKDD 99, 3rd European Conference of Principles and Practice of Knowledge Discovery in Databases 1999."},{"key":"276_CR17","doi-asserted-by":"publisher","DOI":"10.1002\/9780470316696","volume-title":"Multiple Imputation for Nonresponse in Surveys","author":"D Rubin","year":"1987","unstructured":"Rubin D: Multiple Imputation for Nonresponse in Surveys. Wiley, New York; 1987."},{"key":"276_CR18","doi-asserted-by":"publisher","DOI":"10.1201\/9781439821862","volume-title":"Analysis of Incomplete Multivariate Data","author":"J Schafer","year":"1997","unstructured":"Schafer J: Analysis of Incomplete Multivariate Data. Chapman & Hall, New York; 1997."},{"key":"276_CR19","volume-title":"Technical report, Division of Biostatistics, Stanford University","author":"T Hastie","year":"1999","unstructured":"Hastie T, Tibshirani R, Sherlock G, Eisen M, Brown P, Botsein D: Imputing Missing Data for Gene Expression Arrays.\n                           Technical report, Division of Biostatistics, Stanford University 1999."},{"issue":"6","key":"276_CR20","doi-asserted-by":"publisher","first-page":"520","DOI":"10.1093\/bioinformatics\/17.6.520","volume":"7","author":"O Troyanskaya","year":"2001","unstructured":"Troyanskaya O, Cantor M, Sherlock G, Brown P, Hastie T, Tibshirani R, Botstein D, Altman RB: Missing value estimation methods for DNA microarrays.\n                           Bioinformatics 2001, 7(6):520\u2013525. 10.1093\/bioinformatics\/17.6.520","journal-title":"Bioinformatics"},{"issue":"453","key":"276_CR21","doi-asserted-by":"publisher","first-page":"260","DOI":"10.1198\/016214501750332839","volume":"96","author":"JS Chen","year":"2001","unstructured":"Chen JS: Jackknife Variance Estimation for Nearest Neighbor Imputation.\n                           Journal of Statistical Association 2001, 96(453):260\u2013269. 10.1198\/016214501750332839","journal-title":"Journal of Statistical Association"},{"issue":"12","key":"276_CR22","doi-asserted-by":"publisher","first-page":"3273","DOI":"10.1091\/mbc.9.12.3273","volume":"9","author":"PT Spellman","year":"1998","unstructured":"Spellman PT, Sherlock G, Zhang MQ, Iyer VR, Anders K, Eisen MB, Brown PO, Botstein D, Futcher B: Comprehensive Identification of Cell-regulated Genes of the Yeast Saccharomyces cerevisiae by Microarray Hybridzation.\n                           Molecular biology of the Cell 1998, 9(12):3273\u20133297.","journal-title":"Molecular biology of the Cell"},{"issue":"34","key":"276_CR23","doi-asserted-by":"publisher","first-page":"31079","DOI":"10.1074\/jbc.M202718200","volume":"277","author":"H Yoshimoto","year":"2002","unstructured":"Yoshimoto H, Saltsman K, Gasch AP, Li HX, Ogawa N, D B, Brown PO, Cyert MS: Genome-wide Analysis of Gene Expression Regulated by the Calcineurin\/Crzlp Signaling Pathway in Saccharomyces cerevisiae.\n                           The Journal of Biological Chemistry 2002, 277(34):31079\u201331088. 10.1074\/jbc.M202718200","journal-title":"The Journal of Biological Chemistry"},{"key":"276_CR24","unstructured":"Helix group at Stanford University[http:\/\/smi-web.stanford.edu\/projects\/helix\/pubs\/impute]"},{"key":"276_CR25","volume-title":"AI and Statistics","author":"R Caruana","year":"2001","unstructured":"Caruana R: A Non-Parametric EM-Style Algorithm for Imputing Missing Values.\n                           AI and Statistics 2001."},{"key":"276_CR26","unstructured":"R: A language and environment for statistical computing[http:\/\/www.R-project.org]"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-5-160.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/1471-2105-5-160\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-5-160.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,7]],"date-time":"2024-10-07T12:19:45Z","timestamp":1728303585000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-5-160"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,10,26]]},"references-count":26,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2004,12]]}},"alternative-id":["276"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-5-160","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2004,10,26]]},"assertion":[{"value":"19 May 2004","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 October 2004","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 October 2004","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"160"}}