{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,5]],"date-time":"2026-05-05T01:24:39Z","timestamp":1777944279329,"version":"3.51.4"},"reference-count":17,"publisher":"Oxford University Press (OUP)","issue":"20","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2007,10,15]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Motivation: Protein expression profiling for differences indicative of early cancer holds promise for improving diagnostics. Due to their high dimensionality, statistical analysis of proteomic data from mass spectrometers is challenging in many aspects such as dimension reduction, feature subset selection as well as construction of classification rules. Search of an optimal feature subset, commonly known as the feature subset selection (FSS) problem, is an important step towards disease classification\/diagnostics with biomarkers.<\/jats:p><jats:p>Methods: We develop a parsimonious threshold-independent feature selection (PTIFS) method based on the concept of area under the curve (AUC) of the receiver operating characteristic (ROC). To reduce computational complexity to a manageable level, we use a sigmoid approximation to the empirical AUC as the criterion function. Starting from an anchor feature, the PTIFS method selects a feature subset through an iterative updating algorithm. Highly correlated features that have similar discriminating power are precluded from being selected simultaneously. The classification rule is then determined from the resulting feature subset.<\/jats:p><jats:p>Results: The performance of the proposed approach is investigated by extensive simulation studies, and by applying the method to two mass spectrometry data sets of prostate cancer and of liver cancer. We compare the new approach with the threshold gradient descent regularization (TGDR) method. The results show that our method can achieve comparable performance to that of the TGDR method in terms of disease classification, but with fewer features selected.<\/jats:p><jats:p>Availability: Supplementary Material and the PTIFS implementations are available at http:\/\/staff.ustc.edu.cn\/~ynyang\/PTIFS<\/jats:p><jats:p>Contact: \u00a0ynyang@ustc.edu.cn or czzhuliang@126.com<\/jats:p><jats:p>Supplementary information: Supplementary data are available at Bioinformatics online.<\/jats:p>","DOI":"10.1093\/bioinformatics\/btm442","type":"journal-article","created":{"date-parts":[[2007,9,19]],"date-time":"2007-09-19T00:25:25Z","timestamp":1190161525000},"page":"2788-2794","source":"Crossref","is-referenced-by-count":38,"title":["A parsimonious threshold-independent protein feature selection method through the area under receiver operating characteristic curve"],"prefix":"10.1093","volume":"23","author":[{"given":"Zhanfeng","family":"Wang","sequence":"first","affiliation":[{"name":"1 Department of Statistics and Finance, University of Science and Technology of China, Hefei, 230026, China, 2Institute of Statistical Science, Academia Sinica, Taipei 11529, Taiwan, 3Department of Statistics, Columbia University, New York 10027, USA and 4Department of Gastroenterology, Changzheng Hospital, Second Military Medical University, Shanghai 200003, China."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuan-chin I.","family":"Chang","sequence":"additional","affiliation":[{"name":"1 Department of Statistics and Finance, University of Science and Technology of China, Hefei, 230026, China, 2Institute of Statistical Science, Academia Sinica, Taipei 11529, Taiwan, 3Department of Statistics, Columbia University, New York 10027, USA and 4Department of Gastroenterology, Changzheng Hospital, Second Military Medical University, Shanghai 200003, China."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhiliang","family":"Ying","sequence":"additional","affiliation":[{"name":"1 Department of Statistics and Finance, University of Science and Technology of China, Hefei, 230026, China, 2Institute of Statistical Science, Academia Sinica, Taipei 11529, Taiwan, 3Department of Statistics, Columbia University, New York 10027, USA and 4Department of Gastroenterology, Changzheng Hospital, Second Military Medical University, Shanghai 200003, China."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liang","family":"Zhu","sequence":"additional","affiliation":[{"name":"1 Department of Statistics and Finance, University of Science and Technology of China, Hefei, 230026, China, 2Institute of Statistical Science, Academia Sinica, Taipei 11529, Taiwan, 3Department of Statistics, Columbia University, New York 10027, USA and 4Department of Gastroenterology, Changzheng Hospital, Second Military Medical University, Shanghai 200003, China."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yaning","family":"Yang","sequence":"additional","affiliation":[{"name":"1 Department of Statistics and Finance, University of Science and Technology of China, Hefei, 230026, China, 2Institute of Statistical Science, Academia Sinica, Taipei 11529, Taiwan, 3Department of Statistics, Columbia University, New York 10027, USA and 4Department of Gastroenterology, Changzheng Hospital, Second Military Medical University, Shanghai 200003, China."}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2007,9,18]]},"reference":[{"key":"2023041106214180800_","first-page":"3609","article-title":"Serum protein fingerprinting coupled with a pattern-matching algorithm distinguishes prostate cancer from benign prostate hyperplasia and healthy men","volume":"62","author":"Adam","year":"2002","journal-title":"Cancer Res"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"407","DOI":"10.1214\/009053604000000067","article-title":"Last angle regression","volume":"32","author":"Efron","year":"2004","journal-title":"Ann. Stat"},{"key":"2023041106214180800_","article-title":"Gradient directed regularization for linear regression and classification","volume-title":"Technical report","author":"Friedman","year":"2004"},{"key":"2023041106214180800_","volume-title":"Computational Learning and Probabilistic Reasoning","author":"Gammerman","year":"1996"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"185","DOI":"10.1155\/2004\/546293","article-title":"Serum protein expression profiling for cancer detection: validation of a SELDI-based approach for prostate cancer","volume":"19","author":"Grizzle","year":"2003","journal-title":"Dis. Markers"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1186\/1471-2105-6-68","article-title":"Feature selection and nearest centroid classification for protein mass spectrometry","volume":"6","author":"Levner","year":"2005","journal-title":"BMC Bioinformatics"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"37","DOI":"10.1002\/sim.1922","article-title":"On linear combinations of biomarkers to improve diagnostic accuracy","volume":"24","author":"Liu","year":"2005","journal-title":"Stat. Med"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"4356","DOI":"10.1093\/bioinformatics\/bti724","article-title":"Regularized ROC method for disease classification and biomarker selection with microarray data","volume":"21","author":"Ma","year":"2005","journal-title":"Bioinformatrics"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1016\/S0001-2998(78)80014-2","article-title":"Basic principles of ROC analysis","volume":"8","author":"Metz","year":"1978","journal-title":"Semin. Nucl. Med"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"432","DOI":"10.1007\/978-94-009-6045-9_25","article-title":"A new approach for testing the significance of differences between the ROC curves measured from correlated data. In","volume-title":"Information Processing in Medical imaging VIII","author":"Metz","year":"1984"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"1054","DOI":"10.1093\/jnci\/93.14.1054","article-title":"Phases of biomarker development for early detection of cancer","volume":"93","author":"Pepe","year":"2001","journal-title":"J Natl Cancer Inst"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"1835","DOI":"10.1093\/clinchem\/48.10.1835","article-title":"Boosted decision tree analysis of surface-enhanced laser desorption\/ionization mass spectral serum profiles discriminates prostate cancer from noncancer patients","volume":"48","author":"Qu","year":"2002","journal-title":"Clin. Chem"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"1350","DOI":"10.1080\/01621459.1993.10476417","article-title":"Linear combinations of multiple diagnostic markers","volume":"88","author":"Su","year":"1993","journal-title":"J. Am. Stat. Ass"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"1285","DOI":"10.1126\/science.3287615","article-title":"Measuring the accuracy of diagnostic systems","volume":"240","author":"Swets","year":"1988","journal-title":"Science"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"449","DOI":"10.1093\/biostatistics\/4.3.449","article-title":"A data-analytic strategy for protein biomarker discovery: profiling of high-dimensional proteomic data for cancer detection","volume":"4","author":"Yasui","year":"2003","journal-title":"Biostatistics"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","first-page":"i487","DOI":"10.1093\/bioinformatics\/bti1030","article-title":"Bayesian neural network approaches to ovarian cancer identification from high-resolution mass spectrometry data","volume":"21","author":"Yu","year":"2005","journal-title":"Bioinformatics"},{"key":"2023041106214180800_","doi-asserted-by":"crossref","DOI":"10.1002\/9780470317082","volume-title":"Statistical Methods in Diagnostic Medicine","author":"Zhou","year":"2002"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/20\/2788\/49818512\/bioinformatics_23_20_2788.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/23\/20\/2788\/49818512\/bioinformatics_23_20_2788.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,21]],"date-time":"2025-01-21T02:45:54Z","timestamp":1737427554000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/23\/20\/2788\/231045"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,9,18]]},"references-count":17,"journal-issue":{"issue":"20","published-print":{"date-parts":[[2007,10,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btm442","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2007,10,15]]},"published":{"date-parts":[[2007,9,18]]}}}