{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T00:10:53Z","timestamp":1773274253921,"version":"3.50.1"},"reference-count":26,"publisher":"Oxford University Press (OUP)","issue":"7","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2006,4,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Given a large set of potential features, such as the set of all gene-expression values from a microarray, it is necessary to find a small subset with which to classify. The task of finding an optimal feature set of a given size is inherently combinatoric because to assure optimality all feature sets of a given size must be checked. Thus, numerous suboptimal feature-selection algorithms have been proposed. There are strong impediments to evaluate feature-selection algorithms using real data when data are limited, a common situation in genetic classification. The difficulty is compound. First, there are no class-conditional distributions from which to draw data points, only a single small labeled sample. Second, there are no test data with which to estimate the feature-set errors, and one must depend on a training-data-based error estimator. Finally, there is no optimal feature set with which to compare the feature sets found by the algorithms.<\/jats:p>\n               <jats:p>Results: This paper describes a genetic test bed for the evaluation of feature-selection algorithms. It begins with a large biological feature-label dataset that is used as an empirical distribution and, using massively parallel computation, finds the top feature sets of various sizes based on a given sample size and classification rule. The user can draw random samples from the data, apply a proposed algorithm, and evaluate the proficiency of the proposed algorithm via three different measures (code provided). A key feature of the test bed is that, once a dataset is input, a single command creates the entire test bed relative to the dataset. The particular dataset used for the first version of the test bed comes from a microarray-based classification study that analyzes a large number of microarrays, prepared with RNA from breast tumor samples from each of 295 patients.<\/jats:p>\n               <jats:p>Availability: The software and supplementary material are available at<\/jats:p>\n               <jats:p>Contact: \u00a0edward@ece.tamu.edu<\/jats:p>","DOI":"10.1093\/bioinformatics\/btl008","type":"journal-article","created":{"date-parts":[[2006,1,21]],"date-time":"2006-01-21T01:32:41Z","timestamp":1137807161000},"page":"837-842","source":"Crossref","is-referenced-by-count":15,"title":["Genetic test bed for feature selection"],"prefix":"10.1093","volume":"22","author":[{"given":"Ashish","family":"Choudhary","sequence":"first","affiliation":[{"name":"Department of Electrical and Computer Engineering, Texas A&M University 1 \u00a0 1 \u00a0 \u00a0 College Station, TX 77843, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marcel","family":"Brun","sequence":"additional","affiliation":[{"name":"TGen 2 \u00a0 2 \u00a0 \u00a0 445 North Fifth Street, Phoenix, AZ 85004, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianping","family":"Hua","sequence":"additional","affiliation":[{"name":"TGen 2 \u00a0 2 \u00a0 \u00a0 445 North Fifth Street, Phoenix, AZ 85004, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James","family":"Lowey","sequence":"additional","affiliation":[{"name":"TGen 2 \u00a0 2 \u00a0 \u00a0 445 North Fifth Street, Phoenix, AZ 85004, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ed","family":"Suh","sequence":"additional","affiliation":[{"name":"TGen 2 \u00a0 2 \u00a0 \u00a0 445 North Fifth Street, Phoenix, AZ 85004, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Edward R.","family":"Dougherty","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, Texas A&M University 1 \u00a0 1 \u00a0 \u00a0 College Station, TX 77843, USA"},{"name":"TGen 2 \u00a0 2 \u00a0 \u00a0 445 North Fifth Street, Phoenix, AZ 85004, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2006,1,20]]},"reference":[{"key":"2023012409110274200_b1","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1038\/ng765","article-title":"MLL Translocations specify a distinct gene expression profile that distinguishes a unique leukemia","volume":"30","author":"Armstrong","year":"2002","journal-title":"Nat. Genet."},{"key":"2023012409110274200_b2","doi-asserted-by":"crossref","first-page":"536","DOI":"10.1038\/35020115","article-title":"Molecular classification of cutaneous malignant melanoma by gene expression profiling","volume":"406","author":"Bittner","year":"2000","journal-title":"Nature"},{"key":"2023012409110274200_b3","doi-asserted-by":"crossref","first-page":"1926","DOI":"10.1073\/pnas.0437875100","article-title":"Variation in gene expression patterns in follicular lymphoma and the response to rituximab","volume":"100","author":"Bohen","year":"2003","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"2023012409110274200_b4","doi-asserted-by":"crossref","first-page":"374","DOI":"10.1093\/bioinformatics\/btg419","article-title":"Is cross-validation valid for small-sample microarray classification?","volume":"20","author":"Braga-Neto","year":"2004","journal-title":"Bioinformatics"},{"key":"2023012409110274200_b5","doi-asserted-by":"crossref","first-page":"1267","DOI":"10.1016\/j.patcog.2003.08.017","article-title":"Bolstered error estimation","volume":"37","author":"Braga-Neto","year":"2004","journal-title":"Pattern Recognit."},{"key":"2023012409110274200_b6","doi-asserted-by":"crossref","first-page":"657","DOI":"10.1109\/TSMC.1977.4309803","article-title":"On the possible orderings in the measurement selection problem","volume":"7","author":"Cover","year":"1977","journal-title":"IEEE Trans. Syst. Man Cybern."},{"key":"2023012409110274200_b7","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4612-0711-5","volume-title":"A Probabilistic Theory of Pattern Recognition","author":"Devroye","year":"1996"},{"key":"2023012409110274200_b8","doi-asserted-by":"crossref","first-page":"316","DOI":"10.1080\/01621459.1983.10477973","article-title":"Estimating the error rate of a prediction rule: improvement on cross-validation","volume":"78","author":"Efron","year":"1983","journal-title":"J. Am. Stat. Soc."},{"key":"2023012409110274200_b9","doi-asserted-by":"crossref","first-page":"531","DOI":"10.1126\/science.286.5439.531","article-title":"Molecular classification of cancer: class discovery and class prediction by gene expression monitoring","volume":"286","author":"Golub","year":"1999","journal-title":"Science"},{"key":"2023012409110274200_b10","doi-asserted-by":"crossref","first-page":"1509","DOI":"10.1093\/bioinformatics\/bti171","article-title":"Optimal number of features as a function of sample size for various classification rules","volume":"21","author":"Hua","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012409110274200_b11","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1016\/j.patcog.2004.08.007","article-title":"Determination of the optimal number of features for quadratic discriminant analysis via the normal approximation to discriminant distribution","volume":"38","author":"Hua","year":"2005","journal-title":"Pattern Recognit."},{"key":"2023012409110274200_b12","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1109\/TIT.1968.1054102","article-title":"On the mean accuracy of the statistical pattern recognizers","volume":"14","author":"Hughes","year":"1968","journal-title":"IEEE Trans. Inform. Theory"},{"key":"2023012409110274200_b13","doi-asserted-by":"crossref","first-page":"835","DOI":"10.1016\/S0169-7161(82)02042-2","article-title":"Dimensionality and sample size consideration in pattern recognition practice","volume-title":"Classification, Pattern Recognition and Reduction of Dimensionality. Handbook of Statistics","author":"Jain","year":"1982"},{"key":"2023012409110274200_b14","doi-asserted-by":"crossref","first-page":"153","DOI":"10.1109\/34.574797","article-title":"Feature selection\u2014evaluation, application, and small sample performance","volume":"19","author":"Jain","year":"1997","journal-title":"IEEE Trans. Pattern Anal. Mach. Intelli."},{"key":"2023012409110274200_b15","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1016\/0031-3203(71)90013-6","article-title":"On dimensionality and sample size in statistical pattern classification","volume":"3","author":"Kanal","year":"1971","journal-title":"Pattern Recognit."},{"key":"2023012409110274200_b16","first-page":"1229","article-title":"Identification of combination gene sets for glioma classification","volume":"1","author":"Kim","year":"2002","journal-title":"Mol. Cancer Ther."},{"key":"2023012409110274200_b17","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1016\/S0031-3203(99)00041-2","article-title":"Comparison of algorithms that select features for pattern classifiers","volume":"33","author":"Kudo","year":"2000","journal-title":"Pattern Recognit."},{"key":"2023012409110274200_b18","doi-asserted-by":"crossref","first-page":"1131","DOI":"10.1093\/bioinformatics\/17.12.1131","article-title":"Gene selection for sample classification based on gene expression data: study of sensitivity to choice of parameters of the GA\/KNN method","volume":"17","author":"Li","year":"2001","journal-title":"Bioinformatics"},{"key":"2023012409110274200_b19","doi-asserted-by":"crossref","first-page":"1119","DOI":"10.1016\/0167-8655(94)90127-9","article-title":"Floating search methods in feature selection","volume":"15","author":"Pudil","year":"1994","journal-title":"Pattern Recognit. Lett."},{"key":"2023012409110274200_b20","doi-asserted-by":"crossref","first-page":"4376","DOI":"10.1091\/mbc.e03-05-0279","article-title":"Gene expression patterns in ovarian carcinomas","volume":"14","author":"Schaner","year":"2003","journal-title":"Mol. Biol. Cell"},{"key":"2023012409110274200_b21","doi-asserted-by":"crossref","first-page":"1046","DOI":"10.1093\/bioinformatics\/bti081","article-title":"Superior feature-set ranking for small samples using bolstered error estimation","volume":"21","author":"Sima","year":"2005","journal-title":"Bioinformatics"},{"key":"2023012409110274200_b22","doi-asserted-by":"crossref","first-page":"2472","DOI":"10.1016\/j.patcog.2005.03.026","article-title":"Impact of error estimation on feature-selection algorithms","volume":"38","author":"Sima","year":"2005","journal-title":"Pattern Recognit."},{"key":"2023012409110274200_b23","first-page":"2578","article-title":"Tumor classification based on gene expression profiling shows that uveal melanomas with and without monosomy 3 represent two distinct entities","volume":"63","author":"Tschentscher","year":"2003","journal-title":"Cancer Res."},{"key":"2023012409110274200_b24","doi-asserted-by":"crossref","first-page":"1999","DOI":"10.1056\/NEJMoa021967","article-title":"A gene-expression signature as a predictor of survival in breast cancer","volume":"374","author":"van de Vijver","year":"2002","journal-title":"N. Engl. J. Med."},{"key":"2023012409110274200_b25","doi-asserted-by":"crossref","first-page":"530","DOI":"10.1038\/415530a","article-title":"Gene expression profiling predicts clinical outcome of breast cancer","volume":"415","author":"van 't Veer","year":"2002","journal-title":"Nature"},{"key":"2023012409110274200_b26","doi-asserted-by":"crossref","first-page":"11462","DOI":"10.1073\/pnas.201162998","article-title":"Predicting the clinical status of human breast cancer by using gene expression profiles","volume":"98","author":"West","year":"2001","journal-title":"Proc. Natl Acad. Sci. USA"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/7\/837\/48839631\/bioinformatics_22_7_837.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/22\/7\/837\/48839631\/bioinformatics_22_7_837.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T09:45:34Z","timestamp":1674553534000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/22\/7\/837\/202450"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,1,20]]},"references-count":26,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2006,4,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btl008","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2006,4,1]]},"published":{"date-parts":[[2006,1,20]]}}}