{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,14]],"date-time":"2026-03-14T02:39:30Z","timestamp":1773455970941,"version":"3.50.1"},"reference-count":52,"publisher":"Cambridge University Press (CUP)","issue":"1","license":[{"start":{"date-parts":[[2020,2,18]],"date-time":"2020-02-18T00:00:00Z","timestamp":1581984000000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Efficiently exploiting all sources of information such as labeled instances, classes\u2019 representation, and relations of them has a high impact on the performance of Multi-Label Text Classification (MLTC) systems. Most of the current approaches use labeled documents as the primary source of information for MLTC. We investigate the effectiveness of different sources of information\u2014 such as the labeled training data, textual labels of classes, and taxonomy relations of classes\u2014 for MLTC. More specifically, first, for each document\u2013class pair, different features are extracted using different sources of information. The features reflect the similarity of classes and documents. Then, MLTC is considered to be a ranking problem, and a learning to rank (LTR) approach is used for ranking classes regarding documents and selecting labels of documents. An important characteristic of many MLTC instances is that documents can belong to multiple classes and there are implicit relations between classes. We apply score propagation on top of LTR to incorporate co-occurrence patterns of classes in labeled documents. Our main findings are the following. First, using an LTR approach integrating all features, we observe significantly better performance than previous systems for MLTC. Specifically, we show that simple classification approaches fail when there is a high number of classes. Second, the analysis of feature weights reveals the relative importance of various sources of evidence, also giving insight into the underlying classification problem. Interestingly, the results indicate that the titles of documents are more informative than all other sources of information. Third, a lean-and-mean system using only four features is able to perform at 96% of the large LTR model that we propose in this paper. Fourth, using the co-occurrence information of classes helps in classifying documents more accurately. Our results show that the co-occurrence information is more helpful when the underlying classifier has a poor performance.<\/jats:p>","DOI":"10.1017\/s1351324920000029","type":"journal-article","created":{"date-parts":[[2020,2,18]],"date-time":"2020-02-18T07:17:16Z","timestamp":1582010236000},"page":"89-111","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":28,"title":["Learning to rank for multi-label text classification: Combining different sources of information"],"prefix":"10.1017","volume":"27","author":[{"given":"Hosein","family":"Azarbonyad","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mostafa","family":"Dehghani","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maarten","family":"Marx","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6614-0087","authenticated-orcid":false,"given":"Jaap","family":"Kamps","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2020,2,18]]},"reference":[{"key":"S1351324920000029_ref40","unstructured":"Steinberger, R. , Ebrahim, M. and Turchi, M. (2012). JRC EuroVoc indexer JEX-A freely available multi-label categorisation tool. In Proceedings of the 5th International Conference on Language Resources and Evaluation, LREC."},{"key":"S1351324920000029_ref16","unstructured":"EuroVoc, . (2014). Multilingual thesaurus of the European Union. Available at http:\/\/eurovoc.europa.eu\/"},{"key":"S1351324920000029_ref9","unstructured":"Daudaravicius, V. (2012). Automatic multilingual annotation of EU legislation with Eurovoc descriptors. In Proceedings of Exploring and Exploiting Official Publications Workshop Programme, EEOP2012, pp. 14\u201320."},{"key":"S1351324920000029_ref33","doi-asserted-by":"publisher","DOI":"10.3115\/1667583.1667634"},{"key":"S1351324920000029_ref46","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2015.10.006"},{"key":"S1351324920000029_ref47","doi-asserted-by":"publisher","DOI":"10.1145\/3019612.3019664"},{"key":"S1351324920000029_ref52","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2013.39"},{"key":"S1351324920000029_ref49","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-011-5270-7"},{"key":"S1351324920000029_ref2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-28577-7_11"},{"key":"S1351324920000029_ref7","first-page":"25","article-title":"Learning to rank using an ensemble of lambda gradient models","volume":"14","author":"Burges","year":"2011","journal-title":"Journal of Machine Learning Research: Workshop and Conference Proceedings"},{"key":"S1351324920000029_ref11","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-44564-9_6"},{"key":"S1351324920000029_ref51","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2006.162"},{"key":"S1351324920000029_ref17","doi-asserted-by":"crossref","unstructured":"Fauzan, A. and Khodra, M.L. (2014). Automatic multilabel categorization using learning to rank framework for complaint text on bandung government. In Proceedings of the International Conference of Advanced Informatics: Concept, Theory and Application, ICAICTA, pp. 28\u201333.","DOI":"10.1109\/ICAICTA.2014.7005910"},{"key":"S1351324920000029_ref3","doi-asserted-by":"publisher","DOI":"10.1145\/3018661.3018741"},{"key":"S1351324920000029_ref45","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-30228-5_20"},{"key":"S1351324920000029_ref25","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2016.04.069"},{"key":"S1351324920000029_ref10","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijar.2008.10.006"},{"key":"S1351324920000029_ref8","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-44794-6_4"},{"key":"S1351324920000029_ref34","doi-asserted-by":"crossref","first-page":"876","DOI":"10.1016\/j.patcog.2011.08.007","article-title":"Multilabel classifiers with a probabilistic thresholding strategy","volume":"45","author":"Quevedo","year":"2012","journal-title":"Pattern Recognition"},{"key":"S1351324920000029_ref39","first-page":"1601","article-title":"Kernel-based learning of hierarchical multilabel classification models","volume":"7","author":"Rousu","year":"2006","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324920000029_ref14","volume-title":"Multi Label Eurovoc Classification for Eastern and Southern Eu Languages","author":"Ebrahim","year":"2012"},{"key":"S1351324920000029_ref27","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-36973-5_16"},{"key":"S1351324920000029_ref29","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-12837-0_11"},{"key":"S1351324920000029_ref28","doi-asserted-by":"publisher","DOI":"10.1016\/0306-4573(95)80034-Q"},{"key":"S1351324920000029_ref23","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2016.2608339"},{"key":"S1351324920000029_ref26","doi-asserted-by":"publisher","DOI":"10.1145\/1150402.1150429"},{"key":"S1351324920000029_ref37","doi-asserted-by":"publisher","DOI":"10.1145\/2600428.2609595"},{"key":"S1351324920000029_ref50","doi-asserted-by":"crossref","unstructured":"Yuan, P. , Chen, Y. , Jin, H. and Huang, L. (2008). MSVM-kNN: Combining SVM and k-NN for multi-class text classification. In Proceedings of the IEEE International Workshop on Semantic Computing and Systems, WSCS, pp. 133\u2013140.","DOI":"10.1109\/WSCS.2008.36"},{"key":"S1351324920000029_ref13","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-16354-3_63"},{"key":"S1351324920000029_ref21","first-page":"17","volume-title":"Multilabel Classification : Problem Analysis, Metrics and Techniques","author":"Herrera","year":"2016"},{"key":"S1351324920000029_ref30","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-662-44851-9_28"},{"key":"S1351324920000029_ref15","unstructured":"Elisseeff, A. and Weston, J. (2001). A kernel method for multi-labelled classification. In Proceedings of the 14th International Conference on Neural Information Processing Systems: Natural and Synthetic, NIPS, pp. 681\u2013687."},{"key":"S1351324920000029_ref36","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2016.12.004"},{"key":"S1351324920000029_ref6","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2004.03.009"},{"key":"S1351324920000029_ref22","doi-asserted-by":"publisher","DOI":"10.1145\/2661829.2661989"},{"key":"S1351324920000029_ref35","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-011-5256-5"},{"key":"S1351324920000029_ref42","doi-asserted-by":"crossref","unstructured":"Tang, L. , Rajan, S. and Narayanan, V.K. (2009). Large scale multi-label classification via metalabeler. In Proceedings of the 18th International Conference on World Wide Web, pp. 211\u2013220.","DOI":"10.1145\/1526709.1526738"},{"key":"S1351324920000029_ref12","doi-asserted-by":"crossref","unstructured":"Dehghani, M. , Azarbonyad, H. , Kamps, J. and Marx, M. (2016b). On horizontal and vertical separation in hierarchical text classification. In Proceedings of the 2016 ACM International Conference on the Theory of Information Retrieval, ICTIR, pp. 185\u2013194.","DOI":"10.1145\/2970398.2970408"},{"key":"S1351324920000029_ref20","unstructured":"Hariharan, B. , Zelnik-manor, L. , Vishwanathan, S.V.N.M and Varma, M. (2010). Large scale max-margin multi-label classification with priors. In Proceedings of the 27th International Conference on Machine Learning, ICML, pp. 423\u2013430."},{"key":"S1351324920000029_ref24","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI.2010.65"},{"key":"S1351324920000029_ref31","doi-asserted-by":"publisher","DOI":"10.1109\/FSKD.2009.207"},{"key":"S1351324920000029_ref43","first-page":"667","volume-title":"In Data Mining and Knowledge Discovery Handbook","author":"Tsoumakas","year":"2010"},{"key":"S1351324920000029_ref4","unstructured":"Bi, W. and Kwok, J.T. (2011). Multi-label classification on tree and dag-structured hierarchies. In Proceedings of the 28th International Conference on Machine Learning, ICML, pp. 17\u201324."},{"key":"S1351324920000029_ref19","doi-asserted-by":"publisher","DOI":"10.21236\/ADA440081"},{"key":"S1351324920000029_ref41","unstructured":"Steinberger, R. , Pouliquen, B. , Widiger, A. , Ignat, C. , Erjavec, T. and Tufis, D. (2006). The JRC-Acquis: A multilingual aligned parallel corpus with 20+ languages. In Proceedings of the 5th International Conference on Language Resources and Evaluation, LREC."},{"key":"S1351324920000029_ref1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-56608-5_6"},{"key":"S1351324920000029_ref18","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-008-5064-8"},{"key":"S1351324920000029_ref48","doi-asserted-by":"crossref","unstructured":"Xu, J. and Li, H. (2007). Adarank: A boosting algorithm for information retrieval. In Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR, pp. 391\u2013398.","DOI":"10.1145\/1277741.1277809"},{"key":"S1351324920000029_ref44","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2014.02.006"},{"key":"S1351324920000029_ref38","doi-asserted-by":"crossref","first-page":"217","DOI":"10.1016\/j.ipm.2015.07.004","article-title":"Optimization and label propagation in bipartite heterogeneous networks to improve transductive classification of texts","volume":"52","author":"Rossi","year":"2016","journal-title":"Information Processing and Management"},{"key":"S1351324920000029_ref5","unstructured":"Bi, W. and Kwok, J.T. (2013). Efficient multi-label classification with many labels. In Proceedings of the 30th International Conference on Machine Learning, ICML, pp. 405\u2013413."},{"key":"S1351324920000029_ref32","unstructured":"Pouliquen, B. , Steinberger, R. and Ignat, C. (2003). Automatic annotation of multilingual text collections with a conceptual thesaurus. In Proceedings of the Ontologies and Information Extraction Workshop, EUROLAN."}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324920000029","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,15]],"date-time":"2022-10-15T20:43:20Z","timestamp":1665866600000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324920000029\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,2,18]]},"references-count":52,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["S1351324920000029"],"URL":"https:\/\/doi.org\/10.1017\/s1351324920000029","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,2,18]]},"assertion":[{"value":"\u00a9 Cambridge University Press 2020","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This is an Open Access article, distributed under the terms of the Creative Commons Attribution licence (http:\/\/creativecommons.org\/licenses\/by\/4.0\/), which permits unrestricted re-use, distribution, and reproduction in any medium, provided the original work is properly cited.","name":"license","label":"License","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}},{"value":"This content has been made available to all.","name":"free","label":"Free to read"}]}}