{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,9]],"date-time":"2026-05-09T10:56:21Z","timestamp":1778324181996,"version":"3.51.4"},"reference-count":49,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2025,2,14]],"date-time":"2025-02-14T00:00:00Z","timestamp":1739491200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100010909","name":"Excellent Young Scientists Fund of NSFC","doi-asserted-by":"crossref","award":["No.T2322027"],"award-info":[{"award-number":["No.T2322027"]}],"id":[{"id":"10.13039\/501100010909","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Postdoctoral Fellowship Program of CPSF","award":["GZC20232736"],"award-info":[{"award-number":["GZC20232736"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["2023M743565"],"award-info":[{"award-number":["2023M743565"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2025,2,28]]},"abstract":"<jats:p>\n            The objective of topic inference in research proposals aims to obtain the most suitable disciplinary division from the discipline system defined by a funding agency. The agency will subsequently find appropriate peer-review experts from their database based on this division. Automated topic inference can reduce human errors caused by manual topic filling, bridge the knowledge gap between funding agencies and project applicants, and improve system efficiency. Existing methods focus on modeling this as a hierarchical multi-label classification problem, using generative models to iteratively infer the most appropriate topic information. However, these methods overlook the gap in scale between interdisciplinary research proposals and non-interdisciplinary ones, leading to an unjust phenomenon where the automated inference system categorizes interdisciplinary proposals as non-interdisciplinary, causing unfairness during the expert assignment. How can we address this data imbalance issue under a complex discipline system and hence resolve this unfairness? In this article, we implement a topic label inference system based on a Transformer encoder\u2013decoder architecture. Furthermore, we utilize interpolation techniques to create a series of pseudo-interdisciplinary proposals from non-interdisciplinary ones during training based on non-parametric indicators, such as cross-topic probabilities and topic occurrence probabilities. This approach aims to reduce the bias of the system during model training. Finally, we conduct extensive experiments on a real-world dataset to verify the effectiveness of the proposed method. The experimental results demonstrate that our training strategy can significantly mitigate the unfairness generated in the topic inference task. To improve the reproducibility of our research, we have released accompanying code by Dropbox.\n            <jats:xref ref-type=\"fn\">\n              <jats:sup>1<\/jats:sup>\n            <\/jats:xref>\n          <\/jats:p>","DOI":"10.1145\/3671149","type":"journal-article","created":{"date-parts":[[2024,6,8]],"date-time":"2024-06-08T11:46:15Z","timestamp":1717847175000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Interdisciplinary Fairness in Imbalanced Research Proposal Topic Inference: A Hierarchical Transformer-based Method with Selective Interpolation"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5294-5776","authenticated-orcid":false,"given":"Meng","family":"Xiao","sequence":"first","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0977-3600","authenticated-orcid":false,"given":"Min","family":"Wu","sequence":"additional","affiliation":[{"name":"Institute for Infocomm Research, Agency for Science, Technology and Research, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9485-4861","authenticated-orcid":false,"given":"Ziyue","family":"Qiao","sequence":"additional","affiliation":[{"name":"School of Computing and Information Technology, Great Bay University, Dongguan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1767-8024","authenticated-orcid":false,"given":"Yanjie","family":"Fu","sequence":"additional","affiliation":[{"name":"Arizona State University, School of Computing and AI, Tempe, AZ, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4852-0163","authenticated-orcid":false,"given":"Zhiyuan","family":"Ning","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3121-8937","authenticated-orcid":false,"given":"Yi","family":"Du","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2144-1131","authenticated-orcid":false,"given":"Yuanchun","family":"Zhou","sequence":"additional","affiliation":[{"name":"Computer Network Information Center, Chinese Academy of Sciences, Beijing, China and University of Chinese Academy of Sciences, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,2,14]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/72.286891"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","unstructured":"Guillaume P. Archambault Yongyi Mao Hongyu Guo and Richong Zhang. 2019. Mixup as directional adversarial training. arXiv:1906.06875. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1906.06875","DOI":"10.48550\/arXiv.1906.06875"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3454741"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00051"},{"key":"e_1_3_2_6_2","first-page":"956","volume-title":"Proceedings of the IEEE International Conference on Data Mining Workshops (ICDM\u2019 23)","author":"Cai Xunxin","year":"2023","unstructured":"Xunxin Cai, Meng Xiao, Zhiyuan Ning, and Yuanchun Zhou. 2023. Resolving the imbalance issue in hierarchical disciplinary topic inference via LLM-based data augmentation. In Proceedings of the IEEE International Conference on Data Mining Workshops (ICDM\u2019 23), 956\u2013961."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","unstructured":"Jie-Neng Chen Shuyang Sun Ju He Philip Torr Alan Yuille and Song Bai. 2021. TransMix: Attend to mix for vision transformers. arXiv:2111.09833. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2111.09833","DOI":"10.48550\/arXiv.2111.09833"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3543507.3583355"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","unstructured":"Yong Cheng Lu Jiang Wolfgang Macherey and Jacob Eisenstein. 2020. Advaug: Robust adversarial augmentation for neural machine translation. arXiv:2006.11834. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2006.11834","DOI":"10.48550\/arXiv.2006.11834"},{"key":"e_1_3_2_10_2","unstructured":"Wei-Lin Chiang Zhuohan Li Zi Lin Ying Sheng Zhanghao Wu Hao Zhang Lianmin Zheng Siyuan Zhuang Yonghao Zhuang Joseph E. Gonzalez Ion Stoica and Eric P. Xing. 2023. Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. Retrieved from https:\/\/lmsys.org\/blog\/2023-03-30-vicuna\/"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-65414-6_9"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly Jakob Uszkoreit and Neil Houlsby. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2010.11929","DOI":"10.48550\/arXiv.2010.11929"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-87240-3_31"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1002\/widm.1139"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","unstructured":"Kai Han An Xiao Enhua Wu Jianyuan Guo Chunjing Xu and Yunhe Wang. 2021. Transformer in transformer. arXiv:2103.00112. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2103.00112","DOI":"10.48550\/arXiv.2103.00112"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357384.3357885"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-019-0192-5"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P17-1052"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/E17-2068"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","unstructured":"Anubha Kabra Ayush Chopra Nikaash Puri Pinkesh Badjatiya Sukriti Verma Piyush Gupta and K. Balaji. 2020. MixBoost: Synthetic oversampling with boosted mixup for handling extreme imbalance. arXiv:2009.01571. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2009.01571","DOI":"10.48550\/arXiv.2009.01571"},{"key":"e_1_3_2_21_2","article-title":"Imbalanced classification via adversarial minority over-sampling","author":"Kim Jaehyung","year":"2019","unstructured":"Jaehyung Kim, Jongheon Jeong, and Jinwoo Shin. 2019. Imbalanced classification via adversarial minority over-sampling. OpenReview.","journal-title":"OpenReview"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/D14-1181"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/s13748-016-0094-0"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3533048"},{"key":"e_1_3_2_25_2","first-page":"231","article-title":"Cost-sensitive learning and the class imbalance problem","author":"Ling Charles X.","year":"2008","unstructured":"Charles X. Ling and Victor S. Sheng. 2008. Cost-sensitive learning and the class imbalance problem. Encyclopedia of Machine Learning 2011, 231\u2013235.","journal-title":"Encyclopedia of Machine Learning"},{"key":"e_1_3_2_26_2","first-page":"2873","volume-title":"Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI\u201916)","author":"Liu Pengfei","year":"2016","unstructured":"Pengfei Liu, Xipeng Qiu, and Huang Xuanjing. 2016. Recurrent neural network for text classification with multi-task learning. In Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI\u201916). 2873\u20132879. arXiv:1605.05101"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","unstructured":"Yuning Mao Jingjing Tian Jiawei Han and Xiang Ren. 2019. Hierarchical text classification with reinforced label assignment. arXiv:1908.10419. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.1908.10419","DOI":"10.48550\/arXiv.1908.10419"},{"key":"e_1_3_2_28_2","first-page":"3111","article-title":"Distributed representations of words and phrases and their compositionality","volume":"26","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. Advances in Neural Information Processing Systems 26, 3111\u20133119.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00178"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2019.2959991"},{"key":"e_1_3_2_31_2","first-page":"2123","article-title":"Optimizing F-measures by cost-sensitive classification","volume":"27","author":"Parambath Shameem Puthiya","year":"2014","unstructured":"Shameem Puthiya Parambath, Nicolas Usunier, and Yves Grandvalet. 2014. Optimizing F-measures by cost-sensitive classification. Advances in Neural Information Processing Systems 27, 2123\u20132131.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","unstructured":"Hugo Touvron Thibaut Lavril Gautier Izacard Xavier Martinet Marie-Anne Lachaux Timoth\u00e9e Lacroix Baptiste Rozi\u00e8re Naman Goyal Eric Hambro Faisal Azhar Aurelien Rodriguez Armand Joulin Edouard Grave and Guillaume Lample. 2023. Llama: Open and efficient foundation language models. arXiv:2302.13971. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2302.13971","DOI":"10.48550\/arXiv.2302.13971"},{"key":"e_1_3_2_33_2","first-page":"5998","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Vol. 30. 5998\u20136008."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-008-5077-3"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2021.10.008"},{"key":"e_1_3_2_36_2","first-page":"6438","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Verma Vikas","year":"2019","unstructured":"Vikas Verma, Alex Lamb, Christopher Beckham, Amir Najafi, Ioannis Mitliagkas, David Lopez-Paz, and Yoshua Bengio. 2019. Manifold mixup: Better representations by interpolating hidden states. In Proceedings of the International Conference on Machine Learning. PMLR, 6438\u20136447."},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i11.17203"},{"key":"e_1_3_2_38_2","first-page":"5075","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Wehrmann Jonatas","year":"2018","unstructured":"Jonatas Wehrmann, Ricardo Cerri, and Rodrigo Barros. 2018. Hierarchical multi-label classification networks. In Proceedings of the International Conference on Machine Learning. PMLR, 5075\u20135084."},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i10.21401"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","unstructured":"Congying Xia. 2018. Mixup-transformer: Dynamic data augmentation for NLP tasks. arXiv:2010.02394v2. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2010.02394","DOI":"10.48550\/arXiv.2010.02394"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2023.3248608"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM51629.2021.00087"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.findings-acl.285"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00612"},{"key":"e_1_3_2_45_2","first-page":"1","volume-title":"Proceedings of the 6th International Conference on Learning Representations (ICLR\u201918)","author":"Zhang Hongyi","year":"2018","unstructured":"Hongyi Zhang, Moustapha Cisse, Yann N. Dauphin, and David Lopez-Paz. 2018. MixUp: Beyond empirical risk minimization. In Proceedings of the 6th International Conference on Learning Representations (ICLR\u201918). 1\u201313."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","unstructured":"Linjun Zhang Zhun Deng Kenji Kawaguchi Amirata Ghorbani and James Zou. 2020. How does mixup help with robustness and generalization? arXiv:2010.04819. Retrieved from https:\/\/doi.org\/10.48550\/arXiv.2010.04819","DOI":"10.48550\/arXiv.2010.04819"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2020.acl-main.104"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P16-2034"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2019-2250"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3397271.3401177"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3671149","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3671149","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:09:32Z","timestamp":1750295372000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3671149"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,14]]},"references-count":49,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,2,28]]}},"alternative-id":["10.1145\/3671149"],"URL":"https:\/\/doi.org\/10.1145\/3671149","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,2,14]]},"assertion":[{"value":"2023-09-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}