{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,14]],"date-time":"2026-07-14T22:08:55Z","timestamp":1784066935861,"version":"3.55.0"},"reference-count":45,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2025,1,26]],"date-time":"2025-01-26T00:00:00Z","timestamp":1737849600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Hierarchical text classification (HTC) is a challenging task that requires classifiers to solve a series of multi-label subtasks considering hierarchical dependencies among labels. Recent studies have introduced prompt tuning to create closer connections between the language model (LM) and the complex label hierarchy. However, we find that the model\u2019s attention to the prompt gradually decreases as the prompt moves from the input to the output layer, revealing the limitations of previous prompt tuning methods for HTC. Given the success of prefix tuning-based studies in natural language understanding tasks, we introduce Structural entroPy guIded pRefIx Tuning (SPIRIT). Specifically, we extract the essential structure of the label hierarchy via structural entropy minimization and decode the abstractive structural information as the prefix to prompt all intermediate layers in the LM. Additionally, a depth-wise reparameterization strategy is developed to enhance optimization and propagate the prefix throughout the LM layers. Extensive evaluation on four popular datasets demonstrates that SPIRIT achieves a state-of-the-art performance.<\/jats:p>","DOI":"10.3390\/e27020128","type":"journal-article","created":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T03:54:27Z","timestamp":1737950067000},"page":"128","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["SPIRIT: Structural Entropy Guided Prefix Tuning for Hierarchical Text Classification"],"prefix":"10.3390","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-1328-8431","authenticated-orcid":false,"given":"He","family":"Zhu","sequence":"first","affiliation":[{"name":"State Key Laboratory of Software Development Environment, School of Computer Science and Engineering, Beihang University, No. 37 Xue Yuan Road, Hai Dian District, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinxiang","family":"Xia","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Software Development Environment, School of Computer Science and Engineering, Beihang University, No. 37 Xue Yuan Road, Hai Dian District, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6533-6229","authenticated-orcid":false,"given":"Ruomei","family":"Liu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Software Development Environment, School of Computer Science and Engineering, Beihang University, No. 37 Xue Yuan Road, Hai Dian District, Beijing 100191, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Bowen","family":"Deng","sequence":"additional","affiliation":[{"name":"School of Advanced Technology, Xi\u2019an Jiaotong-Liverpool University, Suzhou 215213, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,1,26]]},"reference":[{"key":"ref_1","unstructured":"Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J.R. (2020, January 5\u201310). Hierarchy-Aware Global Model for Hierarchical Text Classification. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online."},{"key":"ref_2","unstructured":"Zong, C., Xia, F., Li, W., and Navigli, R. (2021, January 1\u20136). Hierarchy-aware Label Semantics Matching Network for Hierarchical Text Classification. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Paperspp), ACL\/IJCNLP 2021, Virtual Event."},{"key":"ref_3","unstructured":"Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-T\u00fcr, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (2021, January 6\u201311). HTCInfoMax: A Global Model for Hierarchical Text Classification via Information Maximization. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online."},{"key":"ref_4","unstructured":"Muresan, S., Nakov, P., and Villavicencio, A. (2022, January 22\u201327). Incorporating Hierarchy into Text Encoder: A Contrastive Learning Approach for Hierarchical Text Classification. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland."},{"key":"ref_5","unstructured":"Bouamor, H., Pino, J., and Bali, K. (2023, January 6\u201310). Instances and Labels: Hierarchy-aware Joint Supervised Contrastive Learning for Hierarchical Multi-Label Text Classification. Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore."},{"key":"ref_6","unstructured":"Goldberg, Y., Kozareva, Z., and Zhang, Y. (2022, January 7\u201311). HPT: Hierarchy-aware Prompt Tuning for Hierarchical Text Classification. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates."},{"key":"ref_7","unstructured":"Goldberg, Y., Kozareva, Z., and Zhang, Y. (2022, January 7\u201311). Exploiting Global and Local Hierarchies for Hierarchical Text Classification. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Ji, K., Lian, Y., Gao, J., and Wang, B. (2023, January 9\u201314). Hierarchical Verbalizer for Few-Shot Hierarchical Text Classification. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Toronto, ON, Canada.","DOI":"10.18653\/v1\/2023.acl-long.164"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Li, X.L., and Liang, P. (2021). Prefix-tuning: Optimizing continuous prompts for generation. arXiv.","DOI":"10.18653\/v1\/2021.acl-long.353"},{"key":"ref_10","unstructured":"Toutanova, K., Rumshisky, A., Zettlemoyer, L., Hakkani-T\u00fcr, D., Beltagy, I., Bethard, S., Cotterell, R., Chakraborty, T., and Zhou, Y. (2021, January 6\u201311). Learning How to Ask: Querying LMs with Mixtures of Soft Prompts. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Liu, X., Ji, K., Fu, Y., Tam, W.L., Du, Z., Yang, Z., and Tang, J. (2021). P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv.","DOI":"10.18653\/v1\/2022.acl-short.8"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"3290","DOI":"10.1109\/TIT.2016.2555904","article-title":"Structural Information and Dynamical Complexity of Networks","volume":"62","author":"Li","year":"2016","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_13","unstructured":"Velickovic, P., Cucurull, G., Casanova, A., Romero, A., Li\u00f2, P., and Bengio, Y. (May, January 30). Graph Attention Networks. Proceedings of the 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada."},{"key":"ref_14","unstructured":"Rogers, A., Boyd-Graber, J.L., and Okazaki, N. (2023, January 9\u201314). HiTIN: Hierarchy-aware Tree Isomorphism Network for Hierarchical Text Classification. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, ON, Canada."},{"key":"ref_15","unstructured":"Su, J., Zhu, M., Murtadha, A., Pan, S., Wen, B., and Liu, Y. (2022). ZLPR: A Novel Loss for Multi-label Classification. arXiv."},{"key":"ref_16","unstructured":"Burstein, J., Doran, C., and Solorio, T. (2019, January 2\u20137). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA. Volume 1 (Long and Short Papers)."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Kowsari, K., Brown, D.E., Heidarysafa, M., Meimandi, K., Gerber, M.S., and Barnes, L.E. (2017, January 18\u201321). HDLTex: Hierarchical Deep Learning for Text Classification. Proceedings of the 2017 16th IEEE International Conference on Machine Learning and Applications (ICMLA), Cancun, Mexico.","DOI":"10.1109\/ICMLA.2017.0-134"},{"key":"ref_18","first-page":"361","article-title":"RCV1: A New Benchmark Collection for Text Categorization Research","volume":"5","author":"Lewis","year":"2004","journal-title":"J. Mach. Learn. Res."},{"key":"ref_19","unstructured":"Sandhaus, E. (2008). The New York Times Annotated Corpus, Linguistic Data Consortium."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Gopal, S., and Yang, Y. (2013, January 11\u201314). Recursive regularization for large-scale classification with hierarchical and graphical dependencies. Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Chicago, IL, USA.","DOI":"10.1145\/2487575.2487644"},{"key":"ref_21","unstructured":"Bengio, Y., and LeCun, Y. (2015, January 7\u20139). Adam: A Method for Stochastic Optimization. Proceedings of the 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA. Conference Track Proceedings."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Gu, Y., Han, X., Liu, Z., and Huang, M. (2021). Ppt: Pre-trained prompt tuning for few-shot learning. arXiv.","DOI":"10.18653\/v1\/2022.acl-long.576"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Yu, C., Shen, Y., and Mao, Y. (2022, January 11\u201315). Constrained sequence-to-tree generation for hierarchical text classification. Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval, Madrid, Spain.","DOI":"10.1145\/3477495.3531765"},{"key":"ref_24","unstructured":"Brown, T.B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., and Askell, A. (2020, January 6\u201312). Language Models are Few-Shot Learners. Proceedings of the Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, Virtual."},{"key":"ref_25","unstructured":"Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., and Anadkat, S. (2023). Gpt-4 technical report. arXiv."},{"key":"ref_26","unstructured":"Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (November, January 31). HFT-CNN: Learning Hierarchical Category Structure for Multi-label Short Text Categorization. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium."},{"key":"ref_27","unstructured":"Korhonen, A., Traum, D.R., and M\u00e0rquez, L. (August, January 28). Hierarchical Transfer Learning for Multi-label Text Classification. Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy. Volume 1: Long Papers."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Huang, W., Chen, E., Liu, Q., Chen, Y., Huang, Z., Liu, Y., Zhao, Z., Zhang, D., and Wang, S. (2019, January 3\u20137). Hierarchical Multi-label Text Classification: An Attention-based Recurrent Network Approach. Proceedings of the 28th ACM International Conference on Information and Knowledge Management, Beijing, China.","DOI":"10.1145\/3357384.3357885"},{"key":"ref_29","unstructured":"Alva-Manchego, F., Choi, E., and Khashabi, D. (August, January 28). Hierarchical Multi-label Classification of Text with Capsule Networks. Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy. Volume 2: Student Research Workshop."},{"key":"ref_30","unstructured":"Inui, K., Jiang, J., Ng, V., and Wan, X. (2019, January 3\u20137). Hierarchical Text Classification with Reinforced Label Assignment. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China."},{"key":"ref_31","unstructured":"Inui, K., Jiang, J., Ng, V., and Wan, X. (2019, January 3\u20137). Learning to Learn and Predict: A Meta-Learning Approach for Multi-Label Classification. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China."},{"key":"ref_32","unstructured":"Jurafsky, D., Chai, J., Schluter, N., and Tetreault, J.R. (2020, January 5\u201310). Efficient Strategies for Hierarchical Text Classification: External Knowledge and Auxiliary Tasks. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Lester, B., Al-Rfou, R., and Constant, N. (2021). The power of scale for parameter-efficient prompt tuning. arXiv.","DOI":"10.18653\/v1\/2021.emnlp-main.243"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Zhong, Z., Friedman, D., and Chen, D. (2021). Factual probing is [mask]: Learning vs. learning to recall. arXiv.","DOI":"10.18653\/v1\/2021.naacl-main.398"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hambardzumyan, K., Khachatrian, H., and May, J. (2021). Warp: Word-level adversarial reprogramming. arXiv.","DOI":"10.18653\/v1\/2021.acl-long.381"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1016\/j.aiopen.2023.08.012","article-title":"GPT understands, too","volume":"5","author":"Liu","year":"2024","journal-title":"AI Open"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"379","DOI":"10.1002\/j.1538-7305.1948.tb01338.x","article-title":"A mathematical theory of communication","volume":"27","author":"Shannon","year":"1948","journal-title":"Bell Syst. Tech. J."},{"key":"ref_38","unstructured":"Liu, Y., Liu, J., Zhang, Z., Zhu, L., and Li, A. (2019, January 8\u201314). REM: From Structural Entropy to Community Structure Deception. Proceedings of the Advances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, Vancouver, BC, Canada."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Wu, J., Li, S., Li, J., Pan, Y., and Xu, K. (2022, January 23\u201329). A Simple yet Effective Method for Graph Classification. Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence, IJCAI 2022, Vienna, Austria.","DOI":"10.24963\/ijcai.2022\/497"},{"key":"ref_40","unstructured":"Calzolari, N., Huang, C., Kim, H., Pustejovsky, J., Wanner, L., Choi, K., Ryu, P., Chen, H., Donatelli, L., and Ji, H. (2022, January 12\u201317). Hierarchical Information Matters: Text Classification via Tree Based Graph Neural Network. Proceedings of the 29th International Conference on Computational Linguistics, COLING 2022, Gyeongju, Republic of Korea."},{"key":"ref_41","unstructured":"Wu, J., Chen, X., Xu, K., and Li, S. (2022, January 17\u201323). Structural entropy guided graph hierarchical pooling. Proceedings of the International Conference on Machine Learning, PMLR, Baltimore, MD, USA."},{"key":"ref_42","unstructured":"Chua, T., Lauw, H.W., Si, L., Terzi, E., and Tsaparas, P. (March, January 27). Minimum Entropy Principle Guided Graph Neural Networks. Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining, WSDM 2023, Singapore."},{"key":"ref_43","unstructured":"Wu, J., Chen, X., Shi, B., Li, S., and Xu, K. (2023, January 23\u201329). SEGA: Structural entropy guided anchor view for graph contrastive learning. Proceedings of the International Conference on Machine Learning, PMLR, Honolulu, HI, USA."},{"key":"ref_44","unstructured":"Williams, B., Chen, Y., and Neville, J. (2023, January 7\u201314). USER: Unsupervised Structural Entropy-Based Robust Graph Neural Network. Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Zou, D., Peng, H., Huang, X., Yang, R., Li, J., Wu, J., Liu, C., and Yu, P.S. (May, January 30). Se-gsl: A general and effective graph structure learning framework through structural entropy optimization. Proceedings of the ACM Web Conference 2023, Austin, TX, USA.","DOI":"10.1145\/3543507.3583453"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/2\/128\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,8]],"date-time":"2025-10-08T10:36:18Z","timestamp":1759919778000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/2\/128"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,26]]},"references-count":45,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["e27020128"],"URL":"https:\/\/doi.org\/10.3390\/e27020128","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,1,26]]}}}