{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T19:23:49Z","timestamp":1773775429316,"version":"3.50.1"},"reference-count":21,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2007,10,1]],"date-time":"2007-10-01T00:00:00Z","timestamp":1191196800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Inf. Syst."],"published-print":{"date-parts":[[2007,10]]},"abstract":"<jats:p>Stemmers attempt to reduce a word to its stem or root form and are used widely in information retrieval tasks to increase the recall rate. Most popular stemmers encode a large number of language-specific rules built over a length of time. Such stemmers with comprehensive rules are available only for a few languages. In the absence of extensive linguistic resources for certain languages, statistical language processing tools have been successfully used to improve the performance of IR systems. In this article, we describe a clustering-based approach to discover equivalence classes of root words and their morphological variants. A set of string distance measures are defined, and the lexicon for a given text collection is clustered using the distance measures to identify these equivalence classes. The proposed approach is compared with Porter's and Lovin's stemmers on the AP and WSJ subcollections of the Tipster dataset using 200 queries. Its performance is comparable to that of Porter's and Lovin's stemmers, both in terms of average precision and the total number of relevant documents retrieved. The proposed stemming algorithm also provides consistent improvements in retrieval performance for French and Bengali, which are currently resource-poor.<\/jats:p>","DOI":"10.1145\/1281485.1281489","type":"journal-article","created":{"date-parts":[[2007,10,12]],"date-time":"2007-10-12T15:47:29Z","timestamp":1192204049000},"page":"18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":95,"title":["YASS"],"prefix":"10.1145","volume":"25","author":[{"given":"Prasenjit","family":"Majumder","sequence":"first","affiliation":[{"name":"Indian Statistical Institute, Kolkata, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mandar","family":"Mitra","sequence":"additional","affiliation":[{"name":"Indian Statistical Institute, Kolkata, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Swapan K.","family":"Parui","sequence":"additional","affiliation":[{"name":"Indian Statistical Institute, Kolkata, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gobinda","family":"Kole","sequence":"additional","affiliation":[{"name":"Indian Statistical Institute, Kolkata, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pabitra","family":"Mitra","sequence":"additional","affiliation":[{"name":"Indian Institute of Technology, Kharagpur, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kalyankumar","family":"Datta","sequence":"additional","affiliation":[{"name":"Jadavpur University, Calcutta, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2007,10]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1016\/0020-0271(74)90020-5"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2004.04.006"},{"key":"e_1_2_1_3_1","volume-title":"the 5th Text Retrieval Conference.","author":"Buckley C.","unstructured":"Buckley , C. , Singhal , A. , and Mitra , M . 1996. Using query zoning and correlation within SMART: TREC 5 . In the 5th Text Retrieval Conference. Buckley, C., Singhal, A., and Mitra, M. 1996. Using query zoning and correlation within SMART: TREC 5. In the 5th Text Retrieval Conference."},{"key":"e_1_2_1_4_1","volume-title":"the 4th Text Retrieval Conference.","author":"Buckley C.","unstructured":"Buckley , C. , Singhal , A. , and Mitra , M . 1995. New retrieval approaches using SMART: TREC 4 . In the 4th Text Retrieval Conference. Buckley, C., Singhal, A., and Mitra, M. 1995. New retrieval approaches using SMART: TREC 4. In the 4th Text Retrieval Conference."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/792550.792564"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1162\/089120101750300490"},{"key":"e_1_2_1_7_1","volume-title":"Proceedings of the Workshop on Cross-Language Evaluation Forum (CLEF), 273--284","author":"Goldsmith J. A.","unstructured":"Goldsmith , J. A. , Higgins , D. , and Soglasnova , S . 2000. Automatic language-specific stemming in information retrieval . In Proceedings of the Workshop on Cross-Language Evaluation Forum (CLEF), 273--284 . Goldsmith, J. A., Higgins, D., and Soglasnova, S. 2000. Automatic language-specific stemming in information retrieval. In Proceedings of the Workshop on Cross-Language Evaluation Forum (CLEF), 273--284."},{"key":"e_1_2_1_8_1","volume-title":"Multiple Comparisons: Theory and Methods","author":"Hsu J.","year":"1986","unstructured":"Hsu , J. 1986 . Multiple Comparisons: Theory and Methods . Chapman and Hall . Hsu, J. 1986. Multiple Comparisons: Theory and Methods. Chapman and Hall."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/331499.331504"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(99)00101-0"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/564376.564425"},{"key":"e_1_2_1_12_1","first-page":"358","article-title":"Binary codes capable of correcting deletions, insertions and reversals","volume":"27","author":"Levenstein V. I.","year":"1966","unstructured":"Levenstein , V. I. 1966 . Binary codes capable of correcting deletions, insertions and reversals . Commun. ACM 27 , 4, 358 -- 368 Levenstein, V. I. 1966. Binary codes capable of correcting deletions, insertions and reversals. Commun. ACM 27, 4, 358--368","journal-title":"Commun. ACM"},{"key":"e_1_2_1_13_1","first-page":"22","article-title":"Development of a stemming algorithm","volume":"11","author":"Lovins J.","year":"1968","unstructured":"Lovins , J. 1968 . Development of a stemming algorithm . Mech. Trans. Comput. Linguis. 11 , 22 -- 31 . Lovins, J. 1968. Development of a stemming algorithm. Mech. Trans. Comput. Linguis. 11, 22--31.","journal-title":"Mech. Trans. Comput. Linguis."},{"key":"e_1_2_1_14_1","unstructured":"Majumder P. Mitra M. and Chaudhuri B. 2004. Construction and statistical analysis of an Indic language corpus for applied language research. Computing Science Tech. Rep. TR\/ISI\/CVPR\/01\/2004 CVPR Unit Indian Statistical Institute Kolkata.  Majumder P. Mitra M. and Chaudhuri B. 2004. Construction and statistical analysis of an Indic language corpus for applied language research. Computing Science Tech. Rep. TR\/ISI\/CVPR\/01\/2004 CVPR Unit Indian Statistical Institute Kolkata."},{"key":"e_1_2_1_15_1","volume-title":"Maryland: Statistical stemming and backoff translation. In Revised Papers from the Workshop of Cross-Language Evaluation Forum on Cross-Language Information Retrieval and Evaluation (CLEF)","author":"Oard D. W.","year":"2001","unstructured":"Oard , D. W. , Levow , G.-A. , and Cabezas , C. I . 2001 . CLEF experiments at Maryland: Statistical stemming and backoff translation. In Revised Papers from the Workshop of Cross-Language Evaluation Forum on Cross-Language Information Retrieval and Evaluation (CLEF) , Springer , London , 176--187. Oard, D. W., Levow, G.-A., and Cabezas, C. I. 2001. CLEF experiments at Maryland: Statistical stemming and backoff translation. In Revised Papers from the Workshop of Cross-Language Evaluation Forum on Cross-Language Information Retrieval and Evaluation (CLEF), Springer, London, 176--187."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1108\/eb046814"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 10th Conference of the European Chapter of the Association for Computational Linguistics (EACL), on Computatinal Linguistics for South Asian Languages (Budapest, Apr.) Workshop.","author":"Ramanathan A.","unstructured":"Ramanathan , A. and Rao , D . 2003. A lightweight stemmer for Hindi . In Proceedings of the 10th Conference of the European Chapter of the Association for Computational Linguistics (EACL), on Computatinal Linguistics for South Asian Languages (Budapest, Apr.) Workshop. Ramanathan, A. and Rao, D. 2003. A lightweight stemmer for Hindi. In Proceedings of the 10th Conference of the European Chapter of the Association for Computational Linguistics (EACL), on Computatinal Linguistics for South Asian Languages (Budapest, Apr.) Workshop."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.3115\/1075218.1075244"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.3115\/1075096.1075146"},{"key":"e_1_2_1_20_1","volume-title":"1971. The SMART Retrieval System---Experiments in Automatic Document Retrieval","author":"Salton G.","unstructured":"Salton , G. , Ed. 1971. The SMART Retrieval System---Experiments in Automatic Document Retrieval . Prentice Hall , Englewood Cliffs, NJ . Salton, G., Ed. 1971. The SMART Retrieval System---Experiments in Automatic Document Retrieval. Prentice Hall, Englewood Cliffs, NJ."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/267954.267957"}],"container-title":["ACM Transactions on Information Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1281485.1281489","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1281485.1281489","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T15:13:46Z","timestamp":1750259626000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1281485.1281489"}},"subtitle":["Yet another suffix stripper"],"short-title":[],"issued":{"date-parts":[[2007,10]]},"references-count":21,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2007,10]]}},"alternative-id":["10.1145\/1281485.1281489"],"URL":"https:\/\/doi.org\/10.1145\/1281485.1281489","relation":{},"ISSN":["1046-8188","1558-2868"],"issn-type":[{"value":"1046-8188","type":"print"},{"value":"1558-2868","type":"electronic"}],"subject":[],"published":{"date-parts":[[2007,10]]},"assertion":[{"value":"2007-10-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}