{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:09:48Z","timestamp":1750219788784,"version":"3.41.0"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,4,12]],"date-time":"2023-04-12T00:00:00Z","timestamp":1681257600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2023,4,30]]},"abstract":"<jats:p>Label smoothing has a wide range of applications in the machine learning field. Nonetheless, label smoothing only softens the targets by adding a uniform distribution into a one-hot vector, which cannot truthfully reflect the underlying relations among categories. However, learning category relations is of vital importance in many fields such as emotion taxonomy and open set recognition. In this work, we propose a method to obtain the label distribution for each category (category distribution) to reveal category relations. Furthermore, based on the learned category distribution, we calculate new soft targets to improve the performance of model classification. Compared with existing methods, our algorithm can improve neural network models without any side information or additional neural network module by considering category relations. Extensive experiments have been conducted on four original datasets and 10 constructed noisy datasets with three basic neural network models to validate our algorithm. The results demonstrate the effectiveness of our algorithm on the classification task. In addition, three experiments (arrangement, clustering, and similarity) are also conducted to validate the intrinsic quality of the learned category distribution. The results indicate that the learned category distribution can well express underlying relations among categories.<\/jats:p>","DOI":"10.1145\/3585279","type":"journal-article","created":{"date-parts":[[2023,2,27]],"date-time":"2023-02-27T12:02:49Z","timestamp":1677499369000},"page":"1-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Learning Category Distribution for Text Classification"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8448-5548","authenticated-orcid":false,"given":"Xiangyu","family":"Wang","sequence":"first","affiliation":[{"name":"National Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences, School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9864-3818","authenticated-orcid":false,"given":"Chengqing","family":"Zong","sequence":"additional","affiliation":[{"name":"National Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences, School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,4,12]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/1451983.1451994"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.111"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2487986"},{"key":"e_1_3_2_5_2","first-page":"3304","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Chen Chen","year":"2019","unstructured":"Chen Chen, Haobo Wang, Weiwei Liu, Xingyuan Zhao, Tianlei Hu, and Gang Chen. 2019. Two-stage label embedding via neural factorization machine for multi-label classification. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33(1). 3304\u20133311."},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"523","DOI":"10.21437\/Interspeech.2017-343","article-title":"Towards better decoding and language model integration in sequence to sequence models","author":"Chorowski Jan","year":"2017","unstructured":"Jan Chorowski and Navdeep Jaitly. 2017. Towards better decoding and language model integration in sequence to sequence models. In Proc. Interspeech 2017. 523\u2013527.","journal-title":"Proc. Interspeech 2017"},{"key":"e_1_3_2_7_2","first-page":"4171","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4171\u20134186."},{"key":"e_1_3_2_8_2","volume-title":"The Nature of Emotion: Fundamental Questions.","author":"Ekman Paul Ed","year":"1994","unstructured":"Paul Ed Ekman and Richard J. Davidson. 1994. The Nature of Emotion: Fundamental Questions.Oxford University Press."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1093\/llc\/11.4.163"},{"key":"e_1_3_2_10_2","article-title":"Recent advances in open set recognition: A survey","author":"Geng Chuanxing","year":"2020","unstructured":"Chuanxing Geng, Sheng-jun Huang, and Songcan Chen. 2020. Recent advances in open set recognition: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 10 (2020), 3614\u20133631.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2016.2545658"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2005.06.042"},{"key":"e_1_3_2_13_2","first-page":"12929","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"35","author":"Guo Biyang","year":"2021","unstructured":"Biyang Guo, Songqiao Han, Xiao Han, Hailiang Huang, and Ting Lu. 2021. Label confusion learning to enhance text classification models. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 35. 12929\u201312936."},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.5555\/3305381.3305518"},{"key":"e_1_3_2_15_2","article-title":"Distilling the knowledge in a neural network","author":"Hinton Geoffrey","year":"2015","unstructured":"Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015. Distilling the knowledge in a neural network. arXiv preprint arXiv:1503.02531 (2015).","journal-title":"arXiv preprint arXiv:1503.02531"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.5555\/645326.649721"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/tkde.2006.180"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1181"},{"key":"e_1_3_2_19_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the 2015 Conference on International Conference on Learning Representations (ICLR) 1\u201315."},{"issue":"3","key":"e_1_3_2_20_2","doi-asserted-by":"crossref","first-page":"226","DOI":"10.1177\/1754073919843185","article-title":"Rethinking the principles of emotion taxonomy","volume":"11","author":"Kron Assaf","year":"2019","unstructured":"Assaf Kron. 2019. Rethinking the principles of emotion taxonomy. Emotion Review 11, 3 (2019), 226\u2013233.","journal-title":"Emotion Review"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4614-3223-4_13"},{"key":"e_1_3_2_22_2","first-page":"171","volume-title":"Proceedings of the 26th International Conference on Computational Linguistics: Technical Papers (COLING\u201916)","author":"Ma Yukun","year":"2016","unstructured":"Yukun Ma, Erik Cambria, and Sa Gao. 2016. Label embedding for zero-shot fine-grained named entity typing. In Proceedings of the 26th International Conference on Computational Linguistics: Technical Papers (COLING\u201916). 171\u2013180."},{"key":"e_1_3_2_23_2","doi-asserted-by":"crossref","first-page":"6870","DOI":"10.18653\/v1\/2020.acl-main.615","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Meister Clara","year":"2020","unstructured":"Clara Meister, Elizabeth Salesky, and Ryan Cotterell. 2020. Generalized entropy regularization or: There\u2019s nothing special about label smoothing. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 6870\u20136886."},{"key":"e_1_3_2_24_2","first-page":"4694","article-title":"When does label smoothing help?","volume":"32","author":"M\u00fcller Rafael","year":"2019","unstructured":"Rafael M\u00fcller, Simon Kornblith, and Geoffrey E. Hinton. 2019. When does label smoothing help? Advances in Neural Information Processing Systems 32 (2019), 4694\u20134703.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_25_2","first-page":"61","volume-title":"IJCAI-99 Workshop on Machine Learning for Information Filtering","author":"Nigam Kamal","year":"1999","unstructured":"Kamal Nigam, John Lafferty, and Andrew McCallum. 1999. Using maximum entropy for text classification. In IJCAI-99 Workshop on Machine Learning for Information Filtering, Vol. 1. 61\u201367."},{"key":"e_1_3_2_26_2","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1023\/A:1007692713085","article-title":"Text classification from labeled and unlabeled documents using EM","volume":"39","author":"Nigam Kamal","year":"2000","unstructured":"Kamal Nigam, Andrew Kachites McCallum, Sebastian Thrun, and Tom Mitchell. 2000. Text classification from labeled and unlabeled documents using EM. Machine Learning 39, 2 (2000), 103\u2013134.","journal-title":"Machine Learning"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N18-1202"},{"key":"e_1_3_2_28_2","first-page":"5142","volume-title":"International Conference on Machine Learning","author":"Phuong Mary","year":"2019","unstructured":"Mary Phuong and Christoph Lampert. 2019. Towards understanding knowledge distillation. In International Conference on Machine Learning. 5142\u20135151."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1016\/0092-6566(82)90005-8"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2014.2321392"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/78.650093"},{"key":"e_1_3_2_32_2","article-title":"Thuctc: An efficient Chinese text classifier","author":"Sun Maosong","year":"2016","unstructured":"Maosong Sun, Jingyang Li, Zhipeng Guo, Z. Yu, Y. Zheng, X. Si, and Z. Liu. 2016. Thuctc: An efficient Chinese text classifier. GitHub Repository (2016).","journal-title":"GitHub Repository"},{"key":"e_1_3_2_33_2","doi-asserted-by":"crossref","first-page":"2818","DOI":"10.1109\/CVPR.2016.308","volume-title":"2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916)","author":"Szegedy Christian","year":"2016","unstructured":"Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016. Rethinking the inception architecture for computer vision. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR\u201916). IEEE, 2818\u20132826."},{"key":"e_1_3_2_34_2","article-title":"Learning soft labels via meta learning","author":"Vyas Nidhi","year":"2020","unstructured":"Nidhi Vyas, Shreyas Saxena, and Thomas Voice. 2020. Learning soft labels via meta learning. arXiv preprint arXiv:2009.09496 (2020).","journal-title":"arXiv preprint arXiv:2009.09496"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1007\/0-306-47815-3_5"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"2321","DOI":"10.18653\/v1\/P18-1216","volume-title":"Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang Guoyin","year":"2018","unstructured":"Guoyin Wang, Chunyuan Li, Wenlin Wang, Yizhe Zhang, Dinghan Shen, Xinyuan Zhang, Ricardo Henao, and Lawrence Carin. 2018. Joint embedding of words and labels for text classification. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2321\u20132331."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.5555\/3172077.3172295"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.184"},{"key":"e_1_3_2_39_2","doi-asserted-by":"crossref","first-page":"2120","DOI":"10.1109\/TKDE.2015.2407371","article-title":"Dual sentiment analysis: Considering two sides of one review","volume":"27","author":"Xia Rui","year":"2015","unstructured":"Rui Xia, Feng Xu, Chengqing Zong, Qianmu Li, Yong Qi, and Tao Li. 2015. Dual sentiment analysis: Considering two sides of one review. IEEE Transactions on Knowledge and Data Engineering 27, 8 (2015), 2120\u20132133.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2010.11.023"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009982220290"},{"key":"e_1_3_2_42_2","article-title":"XLNet: Generalized autoregressive pretraining for language understanding","volume":"32","author":"Yang Zhilin","year":"2019","unstructured":"Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R. Salakhutdinov, and Quoc V. Le. 2019. XLNet: Generalized autoregressive pretraining for language understanding. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_43_2","first-page":"7370","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","volume":"33","author":"Yao Liang","year":"2019","unstructured":"Liang Yao, Chengsheng Mao, and Yuan Luo. 2019. Graph convolutional networks for text classification. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 33(1). 7370\u20137377."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3446776"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3089942"},{"key":"e_1_3_2_46_2","doi-asserted-by":"crossref","first-page":"4545","DOI":"10.18653\/v1\/D18-1484","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Zhang Honglun","year":"2018","unstructured":"Honglun Zhang, Liqiang Xiao, Wenqing Chen, Yongkun Wang, and Yaohui Jin. 2018. Multi-task label embedding for text classification. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 4545\u20134553."},{"key":"e_1_3_2_47_2","first-page":"649","article-title":"Character-level convolutional networks for text classification","volume":"28","author":"Zhang Xiang","year":"2015","unstructured":"Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. Advances in Neural Information Processing Systems 28 (2015), 649\u2013657.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13042-010-0001-0"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-981-16-0100-2"}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3585279","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3585279","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:37:56Z","timestamp":1750178276000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3585279"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,12]]},"references-count":48,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,4,30]]}},"alternative-id":["10.1145\/3585279"],"URL":"https:\/\/doi.org\/10.1145\/3585279","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"type":"print","value":"2375-4699"},{"type":"electronic","value":"2375-4702"}],"subject":[],"published":{"date-parts":[[2023,4,12]]},"assertion":[{"value":"2022-03-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-02-06","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-04-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}