{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,7]],"date-time":"2026-02-07T12:55:20Z","timestamp":1770468920144,"version":"3.49.0"},"reference-count":14,"publisher":"Oxford University Press (OUP)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2005,1,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: The standard paradigm for a classifier design is to obtain a sample of feature-label pairs and then to apply a classification rule to derive a classifier from the sample data. Typically in laboratory situations the sample size is limited by cost, time or availability of sample material. Thus, an investigator may wish to consider a sequential approach in which there is a sufficient number of patients to train a classifier in order to make a sound decision for diagnosis while at the same time keeping the number of patients as small as possible to make the studies affordable.<\/jats:p><jats:p>Results: A sequential classification procedure is studied via the martingale central limit theorem. It updates the classification rule at each step and provides stopping criteria to ensure with a certain confidence that at stopping a future subject will have misclassification probability smaller than a predetermined threshold. Simulation studies and applications to microarray data analysis are provided. The procedure possesses several attractive properties: (1) it updates the classification rule sequentially and thus does not rely on distributions of primary measurements from other studies; (2) it assesses the stopping criteria at each sequential step and thus can substantially reduce cost via early stopping; and (3) it is not restricted to any particular classification rule and therefore applies to any parametric or non-parametric method, including feature selection or extraction.<\/jats:p><jats:p>Availability: R-code for the sequential stopping rule is available at http:\/\/stat.tamu.edu\/~wfu\/microarray\/sequential\/R-code.html<\/jats:p><jats:p>Contact: \u00a0wfu@stat.tamu.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/bth461","type":"journal-article","created":{"date-parts":[[2004,8,6]],"date-time":"2004-08-06T00:13:20Z","timestamp":1091751200000},"page":"63-70","source":"Crossref","is-referenced-by-count":23,"title":["How many samples are needed to build a classifier: a general sequential approach"],"prefix":"10.1093","volume":"21","author":[{"given":"Wenjiang J.","family":"Fu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Edward R.","family":"Dougherty","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bani","family":"Mallick","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Raymond J.","family":"Carroll","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2004,8,5]]},"reference":[{"key":"2023013107192769500_B1","doi-asserted-by":"crossref","unstructured":"Ambroise, C and McLachlan, G.J. 2002Selection bias in gene extraction on the basis of microarray gene-expression data. Proc. Natl. Acad. Sci., USA996562\u20136566","DOI":"10.1073\/pnas.102102699"},{"key":"2023013107192769500_B2","unstructured":"Braga-Neto, U.M. and Dougherty, E.R. 2000Is cross-validation valid for small-sample microarray classification?. Bioinformatics20374\u2013380"},{"key":"2023013107192769500_B3","doi-asserted-by":"crossref","unstructured":"Emerson, S.S. and Fleming, T.R. 1990Parameter estimation following group sequential hypothesis testing. Biometrika77875\u2013892","DOI":"10.2307\/2337110"},{"key":"2023013107192769500_B4","unstructured":"Hall, P. and Heyde, C.C. Martingale Limit Theory and Its Application1980, New York Academic Press Inc"},{"key":"2023013107192769500_B5","doi-asserted-by":"crossref","unstructured":"Knight, K. Mathematical Statistics2000, New York Chapman and Hall\/CRC","DOI":"10.1201\/9781584888567"},{"key":"2023013107192769500_B6","unstructured":"Lai, T.L. 1997On optimal stopping problems in sequential hypothesis testing. Statist. Sinica7, pp. 33\u201351"},{"key":"2023013107192769500_B7","unstructured":"Liu, A. and Hall, W.J. 1999Unbiased estimation following a group sequential test. Biometrika8671\u201378"},{"key":"2023013107192769500_B8","unstructured":"Perou, C.M., Sorlie, T., Eisen, M.B., van de Rijn, M., Jeffrey, S.S., Rees, C.A., Pollack, J.R., Ross, D.T., Johnsen, H., Akslen, L.A., et al. 2000Molecular portraits of Human Breast Tumours. Nature406747\u2013752"},{"key":"2023013107192769500_B9","doi-asserted-by":"crossref","unstructured":"Pinheiro, J.C. and DeMets, D.L. 1997Estimating and reducing bias in group sequential designs with Gaussian independent increment structure. Biometrika84831\u2013845","DOI":"10.1093\/biomet\/84.4.831"},{"key":"2023013107192769500_B10","unstructured":"Shorack, G.R. Probability for Statisticians2000, New York Springer"},{"key":"2023013107192769500_B11","doi-asserted-by":"crossref","unstructured":"Todd, S. and Whitehead, J. 1996Point and interval estimation following a sequential clinical trial. Biometrika83, pp. 453\u2013461","DOI":"10.1093\/biomet\/83.2.453"},{"key":"2023013107192769500_B12","unstructured":"van de Vijver, M.J., He, Y.D., van't Veer, L.J., Dai, H., Hart, A.A.M., Voskuil, D.W., Schreiber, G.J., Peterse, J.L., Roberts, C., Marton, M.J., et al. 2002A gene-expression signature as a predictor of survival in breast cancer. N. Engl. J. Med.3471999\u20132009"},{"key":"2023013107192769500_B13","doi-asserted-by":"crossref","unstructured":"van't Veer, L.J., Dai, H., van de Vijver, M.J., He, Y.D., Hart, A.A.M., Mao, M., Peterse, H.L., van der Kooy, K., Marton, M.J., Witteveen, A.T., et al. 2002Gene expression profiling predicts clinical outcome of breast cancer. Nature415530\u2013536","DOI":"10.1038\/415530a"},{"key":"2023013107192769500_B14","doi-asserted-by":"crossref","unstructured":"Whitehead, J. 1986On the bias of maximum likelihood estimation following a sequential test. Biometrika73573\u2013581","DOI":"10.2307\/2336521"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/1\/63\/48961852\/bioinformatics_21_1_63.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/21\/1\/63\/48961852\/bioinformatics_21_1_63.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,12,18]],"date-time":"2024-12-18T04:24:33Z","timestamp":1734495873000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/21\/1\/63\/212378"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,8,5]]},"references-count":14,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2005,1,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bth461","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2005,1,1]]},"published":{"date-parts":[[2004,8,5]]}}}