{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T02:28:11Z","timestamp":1781663291624,"version":"3.54.5"},"reference-count":55,"publisher":"MIT Press","license":[{"start":{"date-parts":[[2024,10,2]],"date-time":"2024-10-02T00:00:00Z","timestamp":1727827200000},"content-version":"vor","delay-in-days":275,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2024,9,30]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Hierarchical text classification (HTC) is an important task with broad applications, and few-shot HTC has gained increasing interest recently. While in-context learning (ICL) with large language models (LLMs) has achieved significant success in few-shot learning, it is not as effective for HTC because of the expansive hierarchical label sets and extremely ambiguous labels. In this work, we introduce the first ICL-based framework with LLM for few-shot HTC. We exploit a retrieval database to identify relevant demonstrations, and an iterative policy to manage multi-layer hierarchical labels. Particularly, we equip the retrieval database with HTC label-aware representations for the input texts, which is achieved by continual training on a pretrained language model with masked language modeling (MLM), layer-wise classification (CLS, specifically for HTC), and a novel divergent contrastive learning (DCL, mainly for adjacent semantically similar labels) objective. Experimental results on three benchmark datasets demonstrate superior performance of our method, and we can achieve state-of-the-art results in few-shot HTC.<\/jats:p>","DOI":"10.1162\/tacl_a_00697","type":"journal-article","created":{"date-parts":[[2024,10,2]],"date-time":"2024-10-02T17:17:45Z","timestamp":1727889465000},"page":"1214-1231","update-policy":"https:\/\/doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":17,"title":["Retrieval-style In-context Learning for Few-shot Hierarchical Text Classification"],"prefix":"10.1162","volume":"12","author":[{"given":"Huiyao","family":"Chen","sequence":"first","affiliation":[{"name":"Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen), China"},{"name":"Alibaba Group, China. chenhy1018@gmail.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu","family":"Zhao","sequence":"additional","affiliation":[{"name":"College of Intelligence and Computing, Tianjin University, China. zhaoyucs@tju.edu.cn"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zulong","family":"Chen","sequence":"additional","affiliation":[{"name":"Alibaba Group, China. chenzulong198867@163.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mengjia","family":"Wang","sequence":"additional","affiliation":[{"name":"Alibaba Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Liangyue","family":"Li","sequence":"additional","affiliation":[{"name":"Alibaba Group, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meishan","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen), China. mason.zms@gmail.com"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Min","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Computing and Intelligence, Harbin Institute of Technology (Shenzhen), China. zhangmin2021@hit.edu.cn"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","published-online":{"date-parts":[[2024,9,30]]},"reference":[{"key":"2024102514344833000_bib1","doi-asserted-by":"publisher","first-page":"13","DOI":"10.1145\/2488388.2488391","article-title":"Multi-label learning with millions of labels: Recommending advertiser bid phrases for web pages","volume-title":"22nd International World Wide Web Conference, WWW \u201913","author":"Agrawal","year":"2013"},{"key":"2024102514344833000_bib2","doi-asserted-by":"publisher","first-page":"323","DOI":"10.18653\/v1\/P19-2045","article-title":"Hierarchical multi-label classification of text with capsule networks","volume-title":"Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28 \u2013 August 2, 2019, Volume 2: Student Research Workshop","author":"Aly","year":"2019"},{"key":"2024102514344833000_bib3","doi-asserted-by":"publisher","first-page":"1782","DOI":"10.18653\/v1\/2023.acl-short.152","article-title":"A simple and effective framework for strict zero-shot hierarchical classification","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2023, Toronto, Canada, July 9\u201314, 2023","author":"Bhambhoria","year":"2023"},{"key":"2024102514344833000_bib4","doi-asserted-by":"publisher","first-page":"4370","DOI":"10.18653\/v1\/2021.acl-long.337","article-title":"Hierarchy-aware label semantics matching network for hierarchical text classification","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Chen","year":"2021"},{"key":"2024102514344833000_bib5","doi-asserted-by":"publisher","first-page":"10492","DOI":"10.1609\/aaai.v36i10.21292","article-title":"Contrastnet: A contrastive learning framework for few-shot text classification","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Chen","year":"2022"},{"key":"2024102514344833000_bib6","doi-asserted-by":"publisher","first-page":"657","DOI":"10.18653\/v1\/2020.findings-emnlp.58","article-title":"Revisiting pre-trained models for Chinese natural language processing","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings","author":"Cui","year":"2020"},{"key":"2024102514344833000_bib7","article-title":"Pre-training with whole word masking for chinese BERT","author":"Cui","year":"2019","journal-title":"arXiv preprint arXiv:1906.08101"},{"key":"2024102514344833000_bib8","doi-asserted-by":"publisher","first-page":"4005","DOI":"10.18653\/v1\/2023.findings-acl.247","article-title":"Why can GPT learn in-context? Language models secretly perform gradient descent as meta-optimizers","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9\u201314, 2023","author":"Dai","year":"2023"},{"key":"2024102514344833000_bib9","first-page":"4171","article-title":"BERT: pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2\u20137, 2019, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2024102514344833000_bib10","doi-asserted-by":"publisher","first-page":"105","DOI":"10.18653\/v1\/2022.acl-demo.10","article-title":"OpenPrompt: An open-source framework for prompt-learning","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations","author":"Ding","year":"2022"},{"key":"2024102514344833000_bib11","article-title":"Compositional semantic parsing with large language models","volume-title":"The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1\u20135, 2023","author":"Drozdov","year":"2023"},{"key":"2024102514344833000_bib12","doi-asserted-by":"publisher","first-page":"320","DOI":"10.18653\/v1\/2022.acl-long.26","article-title":"Glm: General language model pretraining with autoregressive blank infilling","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Zhengxiao","year":"2022"},{"key":"2024102514344833000_bib13","doi-asserted-by":"publisher","first-page":"14014","DOI":"10.18653\/v1\/2023.acl-long.783","article-title":"Mitigating label biases for in-context learning","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9\u201314, 2023","author":"Fei","year":"2023"},{"key":"2024102514344833000_bib14","doi-asserted-by":"publisher","first-page":"3816","DOI":"10.18653\/v1\/2021.acl-long.295","article-title":"Making pre-trained language models better few-shot learners","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL\/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1\u20136, 2021","author":"Gao","year":"2021"},{"key":"2024102514344833000_bib15","doi-asserted-by":"publisher","first-page":"12933","DOI":"10.1609\/aaai.v37i11.26520","article-title":"Hierarchical text classification as sub-hierarchy sequence generation","volume-title":"Thirty-Seventh AAAI Conference on Artificial Intelligence, AAAI 2023, Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence, IAAI 2023, Thirteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2023, Washington, DC, USA, February 7\u201314, 2023","author":"Im","year":"2023"},{"key":"2024102514344833000_bib16","doi-asserted-by":"publisher","first-page":"2918","DOI":"10.18653\/v1\/2023.acl-long.164","article-title":"Hierarchical verbalizer for few-shot hierarchical text classification","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Ke","year":"2023"},{"key":"2024102514344833000_bib17","doi-asserted-by":"publisher","first-page":"2092","DOI":"10.1145\/3539618.3592005","article-title":"LADER: Log-augmented dense retrieval for biomedical literature search","volume-title":"Proceedings of SIGIR 2023","author":"Jin","year":"2023"},{"key":"2024102514344833000_bib18","article-title":"Adam: A method for stochastic optimization","volume-title":"3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7\u20139, 2015, Conference Track Proceedings","author":"Kingma","year":"2015"},{"key":"2024102514344833000_bib19","first-page":"170","article-title":"Hierarchically classifying documents using very few words","volume-title":"Proceedings of the Fourteenth International Conference on Machine Learning (ICML 1997), Nashville, Tennessee, USA, July 8\u201312, 1997","author":"Koller","year":"1997"},{"key":"2024102514344833000_bib20","doi-asserted-by":"publisher","first-page":"364","DOI":"10.1109\/ICMLA.2017.0-134","article-title":"Hdltex: Hierarchical deep learning for text classification","volume-title":"16th IEEE International Conference on Machine Learning and Applications, ICMLA 2017, Cancun, Mexico, December 18\u201321, 2017","author":"Kowsari","year":"2017"},{"key":"2024102514344833000_bib21","doi-asserted-by":"publisher","first-page":"4644","DOI":"10.18653\/v1\/2023.acl-long.256","article-title":"Unified demonstration retriever for in-context learning","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Li","year":"2023"},{"key":"2024102514344833000_bib22","article-title":"What makes good in-context examples for gpt-3?","author":"Liu","year":"2021","journal-title":"arXiv preprint arXiv:2101.06804"},{"key":"2024102514344833000_bib23","doi-asserted-by":"publisher","first-page":"100","DOI":"10.18653\/v1\/2022.deelio-1.10","article-title":"What makes good in-context examples for gpt-3?","volume-title":"Proceedings of Deep Learning Inside Out: The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures, DeeLIO@ACL 2022, Dublin, Ireland and Online, May 27, 2022","author":"Liu","year":"2022"},{"key":"2024102514344833000_bib24","doi-asserted-by":"publisher","first-page":"445","DOI":"10.18653\/v1\/D19-1042","article-title":"Hierarchical text classification with reinforced label assignment","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3\u20137, 2019","author":"Mao","year":"2019"},{"key":"2024102514344833000_bib25","doi-asserted-by":"publisher","first-page":"11048","DOI":"10.18653\/v1\/2022.emnlp-main.759","article-title":"Rethinking the role of demonstrations: What makes in-context learning work?","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7\u201311, 2022","author":"Min","year":"2022"},{"issue":"12","key":"2024102514344833000_bib26","doi-asserted-by":"publisher","first-page":"70","DOI":"10.1093\/bioinformatics\/btw294","article-title":"DeepMeSH: Deep semantic representation for improving large-scale mesh indexing","volume":"32","author":"Peng","year":"2016","journal-title":"Bioinformatics"},{"key":"2024102514344833000_bib27","article-title":"Web of science","author":"Reuters","year":"2012"},{"key":"2024102514344833000_bib28","doi-asserted-by":"publisher","first-page":"2655","DOI":"10.18653\/v1\/2022.naacl-main.191","article-title":"Learning to retrieve prompts for in-context learning","volume-title":"Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies","author":"Rubin","year":"2022"},{"key":"2024102514344833000_bib29","article-title":"Exnet: Efficient in-context learning for data-less text classification","volume":"abs\/2305.14622","author":"Shome","year":"2023","journal-title":"CoRR"},{"key":"2024102514344833000_bib30","doi-asserted-by":"publisher","first-page":"817","DOI":"10.18653\/v1\/D18-1094","article-title":"A hierarchical neural attention-based text classifier","volume-title":"Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing","author":"Sinha","year":"2018"},{"key":"2024102514344833000_bib31","doi-asserted-by":"publisher","first-page":"3747","DOI":"10.18653\/v1\/2023.acl-long.207","article-title":"Peer-label assisted hierarchical text classification","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9\u201314, 2023","author":"Song","year":"2023"},{"key":"2024102514344833000_bib32","doi-asserted-by":"publisher","first-page":"819","DOI":"10.18653\/v1\/2022.acl-long.60","article-title":"An information-theoretic approach to prompt engineering without ground truth labels","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22\u201327, 2022","author":"Sorensen","year":"2022"},{"key":"2024102514344833000_bib33","doi-asserted-by":"publisher","first-page":"216","DOI":"10.1016\/j.ins.2018.09.001","article-title":"An analysis of hierarchical text classification using word embeddings","volume":"471","author":"Stein","year":"2019","journal-title":"Information Sciences"},{"key":"2024102514344833000_bib34","doi-asserted-by":"publisher","first-page":"102613","DOI":"10.1016\/j.artmed.2023.102613","article-title":"CEHMR: Curriculum learning enhanced hierarchical multi-label classification for medication recommendation","volume":"143","author":"Sun","year":"2023","journal-title":"Artificial Intelligence in Medicine"},{"key":"2024102514344833000_bib35","doi-asserted-by":"publisher","first-page":"1556","DOI":"10.3115\/v1\/P15-1150","article-title":"Improved semantic representations from tree-structured long short-term memory networks","volume-title":"Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing of the Asian Federation of Natural Language Processing, ACL 2015, July 26\u201331, 2015, Beijing, China, Volume 1: Long Papers","author":"Tai","year":"2015"},{"issue":"11","key":"2024102514344833000_bib36","article-title":"Visualizing data using t-sne","volume":"9","author":"Van der Maaten","year":"2008","journal-title":"Journal of machine learning research"},{"key":"2024102514344833000_bib37","article-title":"GPT-NER: Named entity recognition via large language models","volume":"abs\/2304.10428","author":"Wang","year":"2023","journal-title":"CoRR"},{"key":"2024102514344833000_bib38","doi-asserted-by":"publisher","first-page":"7722","DOI":"10.18653\/v1\/2023.findings-acl.489","article-title":"Towards better hierarchical text classification with data generation","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9\u201314, 2023","author":"Wang","year":"2023"},{"key":"2024102514344833000_bib39","doi-asserted-by":"publisher","first-page":"7109","DOI":"10.18653\/v1\/2022.acl-long.491","article-title":"Incorporating hierarchy into text encoder: A contrastive learning approach for hierarchical text classification","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)","author":"Wang","year":"2022"},{"key":"2024102514344833000_bib40","doi-asserted-by":"publisher","first-page":"3740","DOI":"10.18653\/v1\/2022.emnlp-main.246","article-title":"HPT: Hierarchy-aware prompt tuning for hierarchical text classification","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Wang","year":"2022"},{"key":"2024102514344833000_bib41","doi-asserted-by":"publisher","first-page":"4353","DOI":"10.18653\/v1\/D19-1444","article-title":"Learning to learn and predict: A meta-learning approach for multi-label classification","volume-title":"Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3\u20137, 2019","author":"Jiawei","year":"2019"},{"key":"2024102514344833000_bib42","doi-asserted-by":"publisher","first-page":"115","DOI":"10.1016\/j.ins.2022.11.158","article-title":"XRR: Extreme multi-label text classification with candidate retrieving and deep ranking","volume":"622","author":"Xiong","year":"2023","journal-title":"Information Sciences"},{"key":"2024102514344833000_bib43","article-title":"Approximate nearest neighbor negative contrastive learning for dense text retrieval","volume-title":"9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3\u20137, 2021","author":"Xiong","year":"2021"},{"key":"2024102514344833000_bib44","doi-asserted-by":"publisher","first-page":"11782","DOI":"10.18653\/v1\/2023.findings-acl.748","article-title":"Regen: Zero-shot text classification via training data generation with progressive dense retrieval","volume-title":"Findings of the Association for Computational Linguistics: ACL 2023, Toronto, Canada, July 9\u201314, 2023","author":"Yue","year":"2023"},{"key":"2024102514344833000_bib45","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2210.02414","article-title":"Glm-130b: An open bilingual pre-trained model","author":"Zeng","year":"2022","journal-title":"arXiv preprint arXiv:2210.02414"},{"key":"2024102514344833000_bib46","article-title":"TIM: Teaching large language models to translate with comparison","volume":"abs\/2307.04408","author":"Zeng","year":"2023","journal-title":"CoRR"},{"key":"2024102514344833000_bib47","doi-asserted-by":"publisher","first-page":"1342","DOI":"10.18653\/v1\/2022.emnlp-main.87","article-title":"Prompt-based meta-learning for few-shot text classification","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Zhang","year":"2022"},{"key":"2024102514344833000_bib48","doi-asserted-by":"publisher","first-page":"1062","DOI":"10.18653\/v1\/2023.findings-eacl.81","article-title":"Long-tailed extreme multi-label text classification by the retrieval of generated pseudo label descriptions","volume-title":"Findings of the Association for Computational Linguistics: EACL 2023, Dubrovnik, Croatia, May 2\u20136, 2023","author":"Zhang","year":"2023"},{"key":"2024102514344833000_bib49","doi-asserted-by":"publisher","first-page":"115922","DOI":"10.1016\/j.eswa.2021.115922","article-title":"LA-HCN: Label-based attention for hierarchical multi-label text classification neural network","volume":"187","author":"Zhang","year":"2022","journal-title":"Expert Systems with Applications"},{"key":"2024102514344833000_bib50","doi-asserted-by":"publisher","first-page":"9134","DOI":"10.18653\/v1\/2022.emnlp-main.622","article-title":"Active example selection for in-context learning","volume-title":"Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing","author":"Zhang","year":"2022"},{"key":"2024102514344833000_bib51","doi-asserted-by":"publisher","first-page":"2158","DOI":"10.1109\/TASLP.2023.3282099","article-title":"Label-correction capsule network for hierarchical text classification","volume":"31","author":"Zhao","year":"2023","journal-title":"IEEE ACM Transactions on Audio, Speech, and Language Processing"},{"key":"2024102514344833000_bib52","first-page":"12697","article-title":"Calibrate before use: Improving few-shot performance of language models","volume-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18v24 July 2021, Virtual Event","author":"Zhao","year":"2021"},{"key":"2024102514344833000_bib53","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2306.05685","article-title":"Judging llm-as-a-judge with mt-bench and chatbot arena","author":"Zheng","year":"2023"},{"key":"2024102514344833000_bib54","doi-asserted-by":"publisher","first-page":"1106","DOI":"10.18653\/v1\/2020.acl-main.104","article-title":"Hierarchy-aware global model for hierarchical text classification","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5\u201310, 2020","author":"Zhou","year":"2020"},{"key":"2024102514344833000_bib55","article-title":"Large language models are human-level prompt engineers","volume-title":"The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1\u20135, 2023","author":"Zhou","year":"2023"}],"container-title":["Transactions of the Association for Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00697\/2476595\/tacl_a_00697.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/tacl\/article-pdf\/doi\/10.1162\/tacl_a_00697\/2476595\/tacl_a_00697.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,25]],"date-time":"2024-10-25T14:35:02Z","timestamp":1729866902000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/tacl\/article\/doi\/10.1162\/tacl_a_00697\/124630\/Retrieval-style-In-context-Learning-for-Few-shot"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024]]},"references-count":55,"URL":"https:\/\/doi.org\/10.1162\/tacl_a_00697","relation":{},"ISSN":["2307-387X"],"issn-type":[{"value":"2307-387X","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2024]]},"published":{"date-parts":[[2024]]}}}