{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,18]],"date-time":"2026-06-18T14:17:02Z","timestamp":1781792222858,"version":"3.54.5"},"reference-count":19,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2004,12,1]],"date-time":"2004-12-01T00:00:00Z","timestamp":1101859200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Transactions on Asian Language Information Processing"],"published-print":{"date-parts":[[2004,12]]},"abstract":"<jats:p>\n            <jats:italic>k<\/jats:italic>\n            is the most important parameter in a text categorization system based on the\n            <jats:italic>k<\/jats:italic>\n            -nearest neighbor algorithm (\n            <jats:italic>k<\/jats:italic>\n            NN). To classify a new document, the\n            <jats:italic>k<\/jats:italic>\n            -nearest documents in the training set are determined first. The prediction of categories for this document can then be made according to the category distribution among the\n            <jats:italic>k<\/jats:italic>\n            nearest neighbors. Generally speaking, the class distribution in a training set is not even; some classes may have more samples than others. The system's performance is very sensitive to the choice of the parameter\n            <jats:italic>k<\/jats:italic>\n            . And it is very likely that a fixed\n            <jats:italic>k<\/jats:italic>\n            value will result in a bias for large categories, and will not make full use of the information in the training set. To deal with these problems, an improved kNN strategy, in which different numbers of nearest neighbors for different categories are used instead of a fixed number across all categories, is proposed in this article. More samples (nearest neighbors) will be used to decide whether a test document should be classified in a category that has more samples in the training set. The numbers of nearest neighbors selected for different categories are adaptive to their sample size in the training set. Experiments on two different datasets show that our methods are less sensitive to the parameter\n            <jats:italic>k<\/jats:italic>\n            than the traditional ones, and can properly classify documents belonging to smaller classes with a large\n            <jats:italic>k<\/jats:italic>\n            . The strategy is especially applicable and promising for cases where estimating the parameter\n            <jats:italic>k<\/jats:italic>\n            via cross-validation is not possible and the class distribution of a training set is skewed.\n          <\/jats:p>","DOI":"10.1145\/1039621.1039623","type":"journal-article","created":{"date-parts":[[2005,1,26]],"date-time":"2005-01-26T16:35:53Z","timestamp":1106757353000},"page":"215-226","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":78,"title":["An adaptive\n            <i>k<\/i>\n            -nearest neighbor text categorization strategy"],"prefix":"10.1145","volume":"3","author":[{"given":"Li","family":"Baoli","sequence":"first","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lu","family":"Qin","sequence":"additional","affiliation":[{"name":"The Hong Kong Polytechnic University, Kowloon, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Shiwen","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2004,12]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Topic Detection and Tracking: Event-based Information Organization","author":"Allan J.","unstructured":"Allan , J. 2002. Topic Detection and Tracking: Event-based Information Organization . Kluwer Academic Boston , MA .]] Allan, J. 2002. Topic Detection and Tracking: Event-based Information Organization. Kluwer Academic Boston, MA.]]"},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the 10th International Symposium on String Processing and Information Retrieval","author":"Cardoso-Cachopo A.","year":"2003","unstructured":"Cardoso-Cachopo , A. , and Olivera , A. L . 2003. An empirical comparison of text categorization methods . In Proceedings of the 10th International Symposium on String Processing and Information Retrieval ( Manaus, Brazil, Oct.8--10 , 2003 ). M.A. Nasciente et al. eds. Springer--Verlag, Heidelberg. 183--196.]] Cardoso-Cachopo, A., and Olivera, A. L. 2003. An empirical comparison of text categorization methods. In Proceedings of the 10th International Symposium on String Processing and Information Retrieval (Manaus, Brazil, Oct.8--10, 2003). M.A. Nasciente et al. eds. Springer--Verlag, Heidelberg. 183--196.]]"},{"key":"e_1_2_1_3_1","volume-title":"Nearest Neighbor (NN) Norms: NN Pattern Classification Techniques","author":"Dasarathy B.V.","unstructured":"Dasarathy , B.V. 1991. Nearest Neighbor (NN) Norms: NN Pattern Classification Techniques . IEEE Computer Society Press , Las Alamitos, CA .]] Dasarathy, B.V. 1991. Nearest Neighbor (NN) Norms: NN Pattern Classification Techniques. IEEE Computer Society Press, Las Alamitos, CA.]]"},{"key":"e_1_2_1_4_1","volume-title":"Proceedings of the 5th Pacific-Asia Conference on Knowledge Discovery and Data Mining","author":"Han E. H.","year":"2001","unstructured":"Han , E. H. , Karypis , G. , and Kumar , V . 2001. Text categorization using weight adjusted k-nearest neighbor classification . In Proceedings of the 5th Pacific-Asia Conference on Knowledge Discovery and Data Mining ( Hong Kong, April 16--18 , 2001 ). D. Cheung, et al. eds. Springer-Verlag, Heidelberg. 53--65.]] Han, E. H., Karypis, G., and Kumar, V. 2001. Text categorization using weight adjusted k-nearest neighbor classification. In Proceedings of the 5th Pacific-Asia Conference on Knowledge Discovery and Data Mining (Hong Kong, April 16--18, 2001). D. Cheung, et al. eds. Springer-Verlag, Heidelberg. 53--65.]]"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the ACL'2000 2nd Workshop on Chinese Language Processing","author":"He J.","year":"2000","unstructured":"He , J. , Tan , A.H. , and Tan , C. L . 2000. Machine learning methods for Chinese web page categorization . In Proceedings of the ACL'2000 2nd Workshop on Chinese Language Processing ( Hong Kong , Oct. 2000 ). 93--100.]] 10.3115\/1117769.1117785 He, J., Tan, A.H., and Tan, C. L. 2000. Machine learning methods for Chinese web page categorization. In Proceedings of the ACL'2000 2nd Workshop on Chinese Language Processing (Hong Kong, Oct. 2000). 93--100.]] 10.3115\/1117769.1117785"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the 10th European Conference on Machine Learning","author":"Joachims T.","year":"1998","unstructured":"Joachims , T. 1998 . Text categorization with support vector machines: learning with many relevant features . In Proceedings of the 10th European Conference on Machine Learning ( Chemnitz, Germany, April 21--24 , 1998). 137--142.]] Joachims, T. 1998. Text categorization with support vector machines: learning with many relevant features. In Proceedings of the 10th European Conference on Machine Learning (Chemnitz, Germany, April 21--24, 1998). 137--142.]]"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the Twelfth International Conference on Machine Learning","author":"Lang K.","year":"1995","unstructured":"Lang , K. 1995 . Newsweeder: learning to filter netnews . In Proceedings of the Twelfth International Conference on Machine Learning ( Tahoe City, CA, July 9--12 , 1995). A. Prieditis et al. eds. Morgan Kaufmann. 331--339.]] Lang, K. 1995. Newsweeder: learning to filter netnews. In Proceedings of the Twelfth International Conference on Machine Learning (Tahoe City, CA, July 9--12, 1995). A. Prieditis et al. eds. Morgan Kaufmann. 331--339.]]"},{"key":"e_1_2_1_8_1","unstructured":"Manning C. D. and Schutze H. 1999. Foundations of Statistical Natural Language Processing. MIT Press Cambridge MA.]]   Manning C. D. and Schutze H. 1999. Foundations of Statistical Natural Language Processing. MIT Press Cambridge MA.]]"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of 15th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Copenhagen, June 21--24","author":"Masand B.","year":"1992","unstructured":"Masand , B. , Linoff , G. , and Waltz , D . 1992. Classifying news stories using memory based reasoning . In Proceedings of 15th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Copenhagen, June 21--24 , 1992 ). N. J. Belkin et al. eds. ACM Press, New York. 59--64.]] 10.1145\/133160.133177 Masand, B., Linoff, G., and Waltz, D. 1992. Classifying news stories using memory based reasoning. In Proceedings of 15th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Copenhagen, June 21--24, 1992). N. J. Belkin et al. eds. ACM Press, New York. 59--64.]] 10.1145\/133160.133177"},{"key":"e_1_2_1_10_1","volume-title":"Machine Learning","author":"Mitchell T.","unstructured":"Mitchell , T. 1997. Machine Learning . McGraw Hill , New York .]] Mitchell, T. 1997. Machine Learning. McGraw Hill, New York.]]"},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","first-page":"130","DOI":"10.1108\/eb046814","article-title":"An algorithm for suffix stripping","volume":"14","author":"Porter M. F.","year":"1980","unstructured":"Porter , M. F. 1980 . An algorithm for suffix stripping . Program 14 , 3 (1980), 130 -- 137 .]] Porter, M. F. 1980. An algorithm for suffix stripping. Program 14, 3 (1980), 130--137.]]","journal-title":"Program"},{"key":"e_1_2_1_12_1","volume-title":"Automatic Text Processing: The Transformation, Analysis, and Retrieval of Information by Computer","author":"Salton G.","unstructured":"Salton , G. 1989. Automatic Text Processing: The Transformation, Analysis, and Retrieval of Information by Computer . Addison-Wesley Longman Publishing , Boston, MA .]] Salton, G. 1989. Automatic Text Processing: The Transformation, Analysis, and Retrieval of Information by Computer. Addison-Wesley Longman Publishing, Boston, MA.]]"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/505282.505283"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 17th International Conference on Research and Development in Information Retrieval (SIGIR'94","author":"Yang Y.","year":"1994","unstructured":"Yang , Y. 1994 . Expert network: effective and efficient learning from human decisions in text categorization and retrieval . In Proceedings of the 17th International Conference on Research and Development in Information Retrieval (SIGIR'94 , Dublin, July 3--6 , 1994). W.B. Croft et al. eds. ACM\/Springer. 13--22.]] Yang, Y. 1994. Expert network: effective and efficient learning from human decisions in text categorization and retrieval. In Proceedings of the 17th International Conference on Research and Development in Information Retrieval (SIGIR'94, Dublin, July 3--6, 1994). W.B. Croft et al. eds. ACM\/Springer. 13--22.]]"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009982220290"},{"key":"e_1_2_1_16_1","volume-title":"Proceedings of the 23rd International Conference on Research and Development in Information Retrieval (SIGIR-2000","author":"Yang Y.","year":"2000","unstructured":"Yang , Y. , Ault , T. , Pierce , T. , and Lattimer , C.W . 2000. Improving text categorization methods for event tracking . In Proceedings of the 23rd International Conference on Research and Development in Information Retrieval (SIGIR-2000 , Athens, July 24--28 , 2000 ). N. J. Belkin et al. eds. ACM Press, New York, 65--72.]] 10.1145\/345508.345550 Yang, Y., Ault, T., Pierce, T., and Lattimer, C.W. 2000. Improving text categorization methods for event tracking. In Proceedings of the 23rd International Conference on Research and Development in Information Retrieval (SIGIR-2000, Athens, July 24--28, 2000). N. J. Belkin et al. eds. ACM Press, New York, 65--72.]] 10.1145\/345508.345550"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/183422.183424"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Yang Y.","year":"1999","unstructured":"Yang , Y. and Liu , X . 1999. A re-examination of text categorization methods . In Proceedings of 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval ( Berkeley, CA, Aug. 15--19 , 1999 ). ACM Press, New York, 42--49.]] 10.1145\/312624.312647 Yang, Y. and Liu, X. 1999. A re-examination of text categorization methods. In Proceedings of 22nd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (Berkeley, CA, Aug. 15--19, 1999). ACM Press, New York, 42--49.]] 10.1145\/312624.312647"},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of Fourteenth International Conference on Machine Learning","author":"Yang Y.","year":"1997","unstructured":"Yang , Y. and Pedersen , J. O . 1997. A comparative study on feature selection in text categorization . In Proceedings of Fourteenth International Conference on Machine Learning ( Nashville, TN, July 8--12 , 1997 ). D. H. Fisher, ed. Morgan Kaufmann, 412--420.]] Yang, Y. and Pedersen, J. O. 1997. A comparative study on feature selection in text categorization. In Proceedings of Fourteenth International Conference on Machine Learning (Nashville, TN, July 8--12, 1997). D. H. Fisher, ed. Morgan Kaufmann, 412--420.]]"}],"container-title":["ACM Transactions on Asian Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1039621.1039623","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1039621.1039623","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T16:31:45Z","timestamp":1750264305000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1039621.1039623"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2004,12]]},"references-count":19,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2004,12]]}},"alternative-id":["10.1145\/1039621.1039623"],"URL":"https:\/\/doi.org\/10.1145\/1039621.1039623","relation":{},"ISSN":["1530-0226","1558-3430"],"issn-type":[{"value":"1530-0226","type":"print"},{"value":"1558-3430","type":"electronic"}],"subject":[],"published":{"date-parts":[[2004,12]]},"assertion":[{"value":"2004-12-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}