{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,5,13]],"date-time":"2025-05-13T22:00:10Z","timestamp":1747173610259,"version":"3.40.5"},"reference-count":51,"publisher":"Cambridge University Press (CUP)","issue":"1","license":[{"start":{"date-parts":[[2020,6,17]],"date-time":"2020-06-17T00:00:00Z","timestamp":1592352000000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2022,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In real-world applications, text classification models often suffer from a lack of accurately labelled documents. The available labelled documents may also be out of domain, making the trained model not able to perform well in the target domain. In this work, we mitigate the data problem of text classification using a two-stage approach. First, we mine representative keywords from a noisy out-of-domain data set using statistical methods. We then apply a dataless classification method to learn from the automatically selected keywords and unlabelled in-domain data. The proposed approach outperformed various supervised learning and dataless classification baselines by a large margin. We evaluated different keyword selection methods intrinsically and extrinsically by measuring their impact on the dataless classification accuracy. Last but not least, we conducted an in-depth analysis of the behaviour of the classifier and explained why the proposed dataless classification method outperformed supervised learning counterparts.<\/jats:p>","DOI":"10.1017\/s1351324920000340","type":"journal-article","created":{"date-parts":[[2020,6,17]],"date-time":"2020-06-17T09:22:50Z","timestamp":1592385770000},"page":"39-69","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":3,"title":["Learning from noisy out-of-domain corpus using dataless classification"],"prefix":"10.1017","volume":"28","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9599-9181","authenticated-orcid":false,"given":"Yiping","family":"Jin","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dittaya","family":"Wanvarie","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Phu T. V.","family":"Le","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2020,6,17]]},"reference":[{"key":"S1351324920000340_ref33","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.191"},{"key":"S1351324920000340_ref13","doi-asserted-by":"publisher","DOI":"10.1177\/1471082X0700700303"},{"key":"S1351324920000340_ref43","doi-asserted-by":"publisher","DOI":"10.2307\/3315930"},{"key":"S1351324920000340_ref50","unstructured":"Zheng, R. , Tian, T. , Hu, Z. , Iyer, R. , Sycara, K. et al. (2016). Joint embedding of hierarchical categories and entities for concept categorization and dataless classification. In Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers, Osaka, Japanpp. The COLING 2016 Organising Committee, pp. 2678\u20132688."},{"key":"S1351324920000340_ref9","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2013.2292894"},{"key":"S1351324920000340_ref49","first-page":"649","volume-title":"In Advances in Neural Information Processing Systems","author":"Zhang","year":"2015"},{"key":"S1351324920000340_ref24","unstructured":"Li, X. and Yang, B. (2018). A pseudo label based dataless naive bayes algorithm for text classification with seed words. In Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, New Mexico, USA. Association for Computational Linguistics, pp. 1908\u20131917."},{"key":"S1351324920000340_ref1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2010.5596659"},{"key":"S1351324920000340_ref5","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1411"},{"key":"S1351324920000340_ref48","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-018-1280-0"},{"key":"S1351324920000340_ref21","doi-asserted-by":"publisher","DOI":"10.1145\/3238250"},{"key":"S1351324920000340_ref28","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00020"},{"key":"S1351324920000340_ref30","first-page":"5413","volume-title":"In Advances in Neural Information Processing Systems","author":"Nam","year":"2017"},{"key":"S1351324920000340_ref20","unstructured":"Le, Q. and Mikolov, T. (2014). Distributed representations of sentences and documents. In Proceedings of the 31st International Conference on Machine Learning-Volume 32, Beijing, China. International Machine Learning Society, pp. 1188\u20131196."},{"key":"S1351324920000340_ref36","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939778"},{"key":"S1351324920000340_ref25","unstructured":"Maas, A.L. , Daly, R.E. , Pham, P.T. , Huang, D. , Ng, A.Y. and Potts, C. (2011). Learning word vectors for sentiment analysis. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies-Volume 1, Portland, Oregon, USA. Association for Computational Linguistics, pp. 142\u2013150."},{"key":"S1351324920000340_ref44","unstructured":"Wang, S. and Manning, C.D. (2012). Baselines and bigrams: Simple, good sentiment and topic classification. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Short Papers-Volume 2, Jeju Island, Korea. Association for Computational Linguistics, pp. 90\u201394."},{"key":"S1351324920000340_ref22","doi-asserted-by":"publisher","DOI":"10.1145\/2983323.2983721"},{"key":"S1351324920000340_ref32","unstructured":"Nguyen-Hoang, B.D. , Pham-Hong, B.T. , Jin, Y. and Le, P. (2018). Genre-oriented web content extraction with deep convolutional neural networks and statistical methods. In Proceedings of the 32nd Pacific Asia Conference on Language, Information and Computation (PACLIC 32), Hong Kong, China. Association for Computational Linguistics, pp. 452\u2013459."},{"key":"S1351324920000340_ref3","doi-asserted-by":"publisher","DOI":"10.1145\/290941.291025"},{"key":"S1351324920000340_ref45","doi-asserted-by":"publisher","DOI":"10.1145\/2063576.2063726"},{"key":"S1351324920000340_ref23","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1214"},{"key":"S1351324920000340_ref17","unstructured":"King, B. and Abney, S.P. (2013). Labeling the languages of words in mixed-language documents using weakly supervised methods. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Atlanta, USA. Association for Computational Linguistics, pp. 1110\u20131119."},{"key":"S1351324920000340_ref18","unstructured":"Krizhevsky, A. , Sutskever, I. and Hinton, G.E. (2012). Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25, Lake Tahoe, Nevada, USA, pp. 1097\u20131105."},{"key":"S1351324920000340_ref15","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.33.2.25"},{"key":"S1351324920000340_ref19","doi-asserted-by":"publisher","DOI":"10.1016\/B978-1-55860-377-6.50048-7"},{"key":"S1351324920000340_ref10","unstructured":"Gabrilovich, E. , Markovitch, S. et al. (2007). Computing semantic relatedness using wikipedia-based explicit semantic analysis. In Proceedings of the Twentieth International Joint Conference on Artificial Intelligence, Hyderabad, vol. 7. International Joint Conferences on Artificial Intelligence, pp. 1606\u20131611."},{"key":"S1351324920000340_ref7","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007607513941"},{"key":"S1351324920000340_ref41","unstructured":"Song, Y. , Upadhyay, S. , Peng, H. and Roth, D. (2016). Cross-lingual dataless classification for many languages. In Proceedings of the Twenty-Fifth International Joint Conference on Artificial Intelligence, New York, USA. International Joint Conferences on Artificial Intelligence, pp. 2901\u20132907."},{"key":"S1351324920000340_ref26","doi-asserted-by":"publisher","DOI":"10.1145\/3269206.3271737"},{"key":"S1351324920000340_ref8","doi-asserted-by":"publisher","DOI":"10.1145\/1390334.1390436"},{"key":"S1351324920000340_ref34","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00259"},{"key":"S1351324920000340_ref12","first-page":"2096","article-title":"Domain-adversarial training of neural networks","volume":"17","author":"Ganin","year":"2016","journal-title":"Journal of Machine Learning Research"},{"key":"S1351324920000340_ref40","doi-asserted-by":"publisher","DOI":"10.1016\/j.artint.2019.02.002"},{"key":"S1351324920000340_ref42","first-page":"244","article-title":"Identifying and correcting mislabeled training instances","author":"Sun","year":"2007","journal-title":"Korea"},{"key":"S1351324920000340_ref31","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-010-9156-z"},{"key":"S1351324920000340_ref29","doi-asserted-by":"crossref","unstructured":"Nam, J. , Menca, E.L. and F\u00fcrnkranz, J. (2016). All-in text: Learning document, label, and word representations jointly. In Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, Phoenix, Arizona, USA. Association for the Advancement of Artificial Intelligence.","DOI":"10.1609\/aaai.v30i1.10241"},{"key":"S1351324920000340_ref14","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1031"},{"key":"S1351324920000340_ref2","unstructured":"Brodley, C.E. , Friedl, M.A. et al. (1996). Identifying and eliminating mislabeled training instances. In Proceedings of the National Conference on Artificial Intelligence, Portland, Oregon. Association for the Advancement of Artificial Intelligence, pp. 799\u2013805."},{"key":"S1351324920000340_ref47","unstructured":"Yogatama, D. , Dyer, C. , Ling, W. and Blunsom, P. (2017). Generative and discriminative text classification with recurrent neural networks. In Proceedings of the 34th International Conference on Machine Learning (ICML 2017), Sydney, Australia. International Machine Learning Society."},{"key":"S1351324920000340_ref16","unstructured":"Jin, Y. , Wanvarie, D. and Le, P. (2017). Combining lightly-supervised text classification models for accurate contextual advertising. In Proceedings of the Eighth International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Taipei, Taiwan, vol. 1. Asian Federation of Natural Language Processing, pp. 545\u2013554."},{"key":"S1351324920000340_ref27","unstructured":"Merity, S. , Xiong, C. , Bradbury, J. and Socher, R. (2016). Pointer sentinel mixture models. arXiv preprint."},{"key":"S1351324920000340_ref37","unstructured":"Sachan, D , Zaheer, M. and Salakhutdinov, R. (2018). Investigating the working of text classifiers. In Proceedings of the 27th International Conference on Computational Linguistics, Santa Fe, New Mexico, USA. Association for Computational Linguistics, pp. 2120\u20132131."},{"key":"S1351324920000340_ref38","unstructured":"Settles, B. (2011). Closing the loop: Fast, interactive semi-supervised annotation with queries on features and instances. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Edinburgh, Scotland. Association for Computational Linguistics, pp. 1467\u20131478."},{"key":"S1351324920000340_ref39","doi-asserted-by":"crossref","unstructured":"Song, Y. and Roth, D. (2014). On dataless hierarchical text classification. In Proceedings of the Twenty-Eighth AAAI Conference on Artificial Intelligence, Quebec, Canada. Association for the Advancement of Artificial Intelligence Press.","DOI":"10.1609\/aaai.v28i1.8938"},{"key":"S1351324920000340_ref35","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1202"},{"key":"S1351324920000340_ref46","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D19-1404"},{"volume-title":"The Principle of Least Effort: An Introduction to Human Ecology","year":"1949","author":"Zipf","key":"S1351324920000340_ref51"},{"key":"S1351324920000340_ref6","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-2015"},{"key":"S1351324920000340_ref11","unstructured":"Gamberger, D. , Lavrac, N. and Groselj, C. (1999). Experiments with noise filtering in a medical domain. In Proceedings of the Sixteenth International Conference on Machine Learning (ICML 1999), Bled, Slovenia. International Machine Learning Society, pp. 143\u2013151."},{"key":"S1351324920000340_ref4","unstructured":"Chang, M.W. , Ratinov, L.A. , Roth, D. and Srikumar, V. (2008). Importance of semantic representation: Dataless classification. In Proceedings of the Twenty-Third AAAI Conference on Artificial Intelligence, Chicago, Illinois, vol. 2. Association for the Advancement of Artificial Intelligence, pp. 830\u2013835."}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324920000340","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,29]],"date-time":"2022-10-29T00:12:26Z","timestamp":1667002346000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324920000340\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,17]]},"references-count":51,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,1]]}},"alternative-id":["S1351324920000340"],"URL":"https:\/\/doi.org\/10.1017\/s1351324920000340","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"type":"print","value":"1351-3249"},{"type":"electronic","value":"1469-8110"}],"subject":[],"published":{"date-parts":[[2020,6,17]]},"assertion":[{"value":"\u00a9 The Author(s), 2020. Published by Cambridge University Press","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}}]}}