{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T16:24:19Z","timestamp":1780590259882,"version":"3.54.1"},"reference-count":18,"publisher":"Oxford University Press (OUP)","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2005,1,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Gene expression data often contain missing expression values. Effective missing value estimation methods are needed since many algorithms for gene expression data analysis require a complete matrix of gene array values. In this paper, imputation methods based on the least squares formulation are proposed to estimate missing values in the gene expression data, which exploit local similarity structures in the data as well as least squares optimization process.<\/jats:p>\n               <jats:p>Results: The proposed local least squares imputation method (LLSimpute) represents a target gene that has missing values as a linear combination of similar genes. The similar genes are chosen by k-nearest neighbors or k coherent genes that have large absolute values of Pearson correlation coefficients. Non-parametric missing values estimation method of LLSimpute are designed by introducing an automatic k-value estimator. In our experiments, the proposed LLSimpute method shows competitive results when compared with other imputation methods for missing value estimation on various datasets and percentages of missing values in the data.<\/jats:p>\n               <jats:p>Availability: The software is available at http:\/\/www.cs.umn.edu\/~hskim\/tools.html<\/jats:p>\n               <jats:p>Contact: \u00a0hpark@cs.umn.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/bth499","type":"journal-article","created":{"date-parts":[[2004,8,28]],"date-time":"2004-08-28T01:15:02Z","timestamp":1093655702000},"page":"187-198","source":"Crossref","is-referenced-by-count":370,"title":["Missing value estimation for DNA microarray gene expression data: local least squares imputation"],"prefix":"10.1093","volume":"21","author":[{"given":"Hyunsoo","family":"Kim","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gene H.","family":"Golub","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haesun","family":"Park","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2004,8,27]]},"reference":[{"key":"2023013107194201500_B1","doi-asserted-by":"crossref","unstructured":"Alter, O., Brown, P.O., Botstein, D. 2000Singular value decomposition for genome-wide expression data processing and modeling. Proc. Natl Acad. Sci. USA9710101\u201310106","DOI":"10.1073\/pnas.97.18.10101"},{"key":"2023013107194201500_B2","doi-asserted-by":"crossref","unstructured":"Alter, O., Brown, P.O., Botstein, D. 2003Generalized singular value decomposition for comparative analysis of genome-scale expression datasets of two different organisms. Proc. Natl Acad. Sci. USA1003351\u20133356","DOI":"10.1073\/pnas.0530258100"},{"key":"2023013107194201500_B3","doi-asserted-by":"crossref","unstructured":"B\u00f8, T.H., Dysvik, B., Jonassen, I. 2004LSimpute: accurate estimation of missing values in microarray data with least squares methods. Nucleic Acids Res.32e34","DOI":"10.1093\/nar\/gnh026"},{"key":"2023013107194201500_B4","unstructured":"Cho, J.H., Lee, D., Park, J.H., Lee, I.B. 2003New gene selection method for classification of cancer subtypes considering within-class variation. FEBS Lett.5513\u20137"},{"key":"2023013107194201500_B5","unstructured":"Institute for Mathematics and its Applications Preprint Series. Friedland, S., Niknejad, A., Chihara, L. 2003A simultaneous reconstruction of missing data in DNA microarrays.   No. 1948"},{"key":"2023013107194201500_B6","doi-asserted-by":"crossref","unstructured":"Gasch, A.P., Huang, M., Metzner, S., Botstein, D., Elledge, S.J., Brown, P.O. 2001Genomic expression responses to DNA-damaging agents and the regulatory role of the yeast ATR homolog Mec1p. Mol. Biol. Cell122987\u20133003","DOI":"10.1091\/mbc.12.10.2987"},{"key":"2023013107194201500_B7","unstructured":"Golub, G.H. and van Loan, C.F. Matrix Computations1996 3rd edn , Baltimore, CA  Johns Hopkins University Press"},{"key":"2023013107194201500_B8","doi-asserted-by":"crossref","unstructured":"Golub, T.R., Slonim, D.K., Tamayo, P., Huard, C., Gaasenbeek, M., Mesirov, J.P., Coller, H., Loh, M.L., Downing, J.R., Caligiuri, M.A., Bloomfield, C.D., Lander, E.S. 1999Molecular classification of cancer: class discovery and class prediction by gene expression monitoring. Science286,  pp. 531\u2013537","DOI":"10.1126\/science.286.5439.531"},{"key":"2023013107194201500_B9","doi-asserted-by":"crossref","unstructured":"Oba, S., Sato, M., Takemasa, I., Monden, M., Matsubara, K., Ishii, S. 2003A Bayesian missing value estimation method for gene expression profile data. Bioinformatics192088\u20132096","DOI":"10.1093\/bioinformatics\/btg287"},{"key":"2023013107194201500_B10","unstructured":"Pearson, K. 1894Contributions to the mathematical theory of evolution. Phil. Trans. R. Soc. London18571\u2013110"},{"key":"2023013107194201500_B11","unstructured":"Perou, C.M., Sorlie, T., Eisen, M.B., van de Rijn, M., Jeffrey, S.S., Rees, C.A., Pollack, J.R., Ross, D.T., Johnsen, H., Akslen, L.A., et al. 2000Molecular portraits of human breast tumors. Nature406747\u2013752"},{"key":"2023013107194201500_B12","unstructured":"Sherlock, G., Hernandez-Boussard, T., Kasarskis, A., Binkley, G., Matese, J.C., Dwight, S.S., Kaloper, M., Weng, S., Jin, H., Ball, C.A., et al. 2001The stanford microarray database. Nucleic Acids Res.29152\u2013155"},{"key":"2023013107194201500_B13","doi-asserted-by":"crossref","unstructured":"Shipp, M.A., Ross, K.N., Tamayo, P., Weng, A.P., Kutok, J.L., Aguiar, R.C., Gaasenbeek, M., Angelo, M., Reich, M., Pinkus, G.S., et al. 2002Diffuse large B-cell lymphoma outcome prediction by gene-expression profiling and supervised machine learning. Nat. Med.868\u201374","DOI":"10.1038\/nm0102-68"},{"key":"2023013107194201500_B14","doi-asserted-by":"crossref","unstructured":"Spellman, P.T., Sherlock, G., Zhang, M.Q., Iyer, V.R., Anders, K., Eisen, M.B., Brown, P.O., Botstein, D., Futcher, B. 1998Comprehensive identification of cell cycle-regulated genes of the yeast Saccharomyces cerevisiae by microarray hybridization. Mol. Biol. Cell93273\u20133297","DOI":"10.1091\/mbc.9.12.3273"},{"key":"2023013107194201500_B15","doi-asserted-by":"crossref","unstructured":"Takemasa, I., Higuchi, H., Yamamoto, H., Sekimoto, M., Tomita, N., Nakamori, S., Matoba, R., Monden, M., Matsubara, K. 2001Construction of preferential cDNA microarray specialized for human colorectal carcinoma: molecular sketch of colorectal cancer. Biochem. Biophys. Res. Commun.2851244\u20131249","DOI":"10.1006\/bbrc.2001.5277"},{"key":"2023013107194201500_B16","doi-asserted-by":"crossref","unstructured":"Troyanskaya, O., Cantor, M., Sherlock, G., Brown, P., Hastie, T., Tibshirani, R., Botstein, D., Altman, R.B. 2001Missing value estimation methods for DNA microarray. Bioinformatics17520\u2013525","DOI":"10.1093\/bioinformatics\/17.6.520"},{"key":"2023013107194201500_B17","doi-asserted-by":"crossref","unstructured":"van't Veer, L.J., Dai, H., van de Vijver, M.J., He, Y.D., Hart, A.A., Mao, M., Peterse, H.L., van der Kooy, K., Marton, M.J., Witteveen, A.T., et al. 2002Gene expression profiling predicts clinical outcome of breast cancer. Nature415530\u2013536","DOI":"10.1038\/415530a"},{"key":"2023013107194201500_B18","doi-asserted-by":"crossref","unstructured":"Vapnik, V. The Nature of Statistical Learning Theory1995, New York  Springer-Verlag","DOI":"10.1007\/978-1-4757-2440-0"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/2\/187\/48962001\/bioinformatics_21_2_187.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/2\/187\/48962001\/bioinformatics_21_2_187.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,31]],"date-time":"2023-01-31T09:57:59Z","timestamp":1675159079000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/21\/2\/187\/187450"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,8,27]]},"references-count":18,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2005,1,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bth499","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2005,1,15]]},"published":{"date-parts":[[2004,8,27]]}}}