{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T23:16:34Z","timestamp":1770938194345,"version":"3.50.1"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"2","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62372139"],"award-info":[{"award-number":["62372139"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100003453","name":"Guangdong Natural Science Foundation","doi-asserted-by":"crossref","award":["2024A1515030024"],"award-info":[{"award-number":["2024A1515030024"]}],"id":[{"id":"10.13039\/501100003453","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100016804","name":"Shenzhen Natural Science Foundation","doi-asserted-by":"crossref","award":["JCYJ20220818102414030"],"award-info":[{"award-number":["JCYJ20220818102414030"]}],"id":[{"id":"10.13039\/100016804","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Guangdong Provincial Key Laboratory","award":["2022B1212010005"],"award-info":[{"award-number":["2022B1212010005"]}]},{"name":"Guangdong S&T Program","award":["2024B0101050003"],"award-info":[{"award-number":["2024B0101050003"]}]},{"name":"Shenzhen Science and Technology Program","award":["ZDSYS20230626091203008, KJZD20231023094700001, KQTD2024072910215406, and RCBS20221008093121053"],"award-info":[{"award-number":["ZDSYS20230626091203008, KJZD20231023094700001, KQTD2024072910215406, and RCBS20221008093121053"]}]},{"name":"Shenzhen College Stability Support Plan","award":["GXWD20220811173340003, GXWD20220817123150002"],"award-info":[{"award-number":["GXWD20220811173340003, GXWD20220817123150002"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Asian Low-Resour. Lang. Inf. Process."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>Multi-teacher knowledge distillation transfers knowledge from multiple large teacher models to a small student model and has performed well on many downstream tasks. However, when distilling knowledge from multiple teachers, it always suffers from the severe problems of being time-consuming and storage-extensive for multiple teacher models training and inference. We present MoE-KD, a simple but effective framework that produces supervision for training the student model from one single teacher model, which fixes the above problems and improves effectiveness. In the proposed MoE-KD, multiple trainable prompts are used to extract different views of samples from a single pre-trained language model and only a few parameters (prompts) need to be trained and stored. To guarantee the generated supervision signals with increased robustness and correctness, we introduce an uncertainty-based mechanism and a selector module, which routes the input instance to its corresponding teacher. We have also extended MoE KD to lifelong learning scenarios, proposing a lightweight solution for catastrophic forgetting. We conduct experiments on traditional KD scenarios and lifelong learning scenarios. MoE-KD yields improvements up to 1.1% and 140% in accuracy and efficiency in knowledge distillation and 2.8% improvements on average in lifelong learning, compared with the strong baseline methods.<\/jats:p>","DOI":"10.1145\/3786604","type":"journal-article","created":{"date-parts":[[2026,1,24]],"date-time":"2026-01-24T18:31:24Z","timestamp":1769279484000},"page":"1-18","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Orchestrating Prompt Expertise: Enhancing Knowledge Distillation via Expert-Guided Tuning"],"prefix":"10.1145","volume":"25","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-4631-8203","authenticated-orcid":false,"given":"Xu","family":"Meng","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology Shenzhen","place":["Shenzhen, China"]},{"name":"Guangdong Key Lab of New Security and Intelligence Technology","place":["Shenzhen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5804-1508","authenticated-orcid":false,"given":"Jun","family":"Rao","sequence":"additional","affiliation":[{"name":"School of computer science and technology, Harbin Institute of Technology Shenzhen","place":["Shenzhen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6903-145X","authenticated-orcid":false,"given":"Shuhan","family":"Qi","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology Shenzhen","place":["Shenzhen, China"]},{"name":"Guangdong Key Lab of New Security and Intelligence Technology","place":["Shenzhen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8524-2006","authenticated-orcid":false,"given":"Xuebo","family":"Liu","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology Shenzhen","place":["Shenzhen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2124-0954","authenticated-orcid":false,"given":"Lei","family":"Wang","sequence":"additional","affiliation":[{"name":"Ping An Technology","place":["Shenzhen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3895-5510","authenticated-orcid":false,"given":"Min","family":"Zhang","sequence":"additional","affiliation":[{"name":"Institute of Computing and Intelligence, Harbin Institute of Technology Shenzhen","place":["Shenzhen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3512-0649","authenticated-orcid":false,"given":"Xuan","family":"Wang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology Shenzhen","place":["Shenzhen, China"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,2,12]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"Proceedings of the European Conference on Artificial Intelligence","author":"Asif Umar","year":"2019","unstructured":"Umar Asif, Jianbin Tang, and Stefan Harrer. 2019. Ensemble knowledge distillation for learning improved and efficient networks. In Proceedings of the European Conference on Artificial Intelligence."},{"key":"e_1_3_1_3_2","volume-title":"Proceedings of the Conference on Neural Information Processing Systems","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et\u00a0al. 2020. Language models are few-shot learners. In Proceedings of the Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_4_2","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Chen Wuyang","year":"2023","unstructured":"Wuyang Chen, Yan-Quan Zhou, Nan Du, Yanping Huang, James Laudon, Z. Chen, and Claire Cu. 2023. Lifelong language pretraining with distribution-specialized experts. In Proceedings of the International Conference on Machine Learning."},{"key":"e_1_3_1_5_2","volume-title":"Proceedings of the International Symposium on Neural Networks","author":"Chen Xingjian","year":"2019","unstructured":"Xingjian Chen, Jianbo Su, and Jun Zhang. 2019. A two-teacher framework for knowledge distillation. In Proceedings of the International Symposium on Neural Networks."},{"key":"e_1_3_1_6_2","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Dauphin Yann","year":"2016","unstructured":"Yann Dauphin, Angela Fan, Michael Auli, and David Grangier. 2016. Language modeling with gated convolutional networks. In Proceedings of the International Conference on Machine Learning."},{"key":"e_1_3_1_7_2","volume-title":"Episodic memory in lifelong language learning","author":"d\u2019Autume Cyprien de Masson","year":"2019","unstructured":"Cyprien de Masson d\u2019Autume, Sebastian Ruder, Lingpeng Kong, and Dani Yogatama. 2019. Episodic memory in lifelong language learning. In Proceedings of the 33rd International Conference on Neural Information Processing Systems. https:\/\/dl.acm.org\/doi\/10.5555\/3454287.3455464"},{"key":"e_1_3_1_8_2","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Douillard Arthur","year":"2020","unstructured":"Arthur Douillard, Matthieu Cord, Charles Ollion, Thomas Robert, and Eduardo Valle. 2020. PODNet: Pooled outputs distillation for small-tasks incremental learning. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11063-007-9043-z"},{"key":"e_1_3_1_10_2","article-title":"Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity","author":"Fedus William","year":"2021","unstructured":"William Fedus, Barret Zoph, and Noam M. Shazeer. 2021. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research 23, 120 (2022), 1\u201339. https:\/\/jmlr.org\/papers\/v23\/21-0998.html","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_11_2","volume-title":"Proceedings of the Interspeech","author":"Fukuda Takashi","year":"2017","unstructured":"Takashi Fukuda, Masayuki Suzuki, Gakuto Kurata, Samuel Thomas, Jia Cui, and Bhuvana Ramabhadran. 2017. Efficient knowledge distillation from an ensemble of teachers. In Proceedings of the Interspeech."},{"key":"e_1_3_1_12_2","doi-asserted-by":"crossref","first-page":"1789","DOI":"10.1007\/s11263-021-01453-z","article-title":"Knowledge distillation: A survey","volume":"129","author":"Gou Jianping","year":"2021","unstructured":"Jianping Gou, B. Yu, Stephen J. Maybank, and Dacheng Tao. 2021. Knowledge distillation: A survey. International Journal of Computer Vision 129 (2021), 1789\u20131819. https:\/\/link.springer.com\/article\/10.1007\/s11263-021-01453-z","journal-title":"International Journal of Computer Vision"},{"key":"e_1_3_1_13_2","first-page":"5085","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition","author":"Gross Sam","year":"2017","unstructured":"Sam Gross, Marc\u2019Aurelio Ranzato, and Arthur Szlam. 2017. Hard mixtures of experts for large scale weakly supervised vision. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. 5085\u20135093. DOI:10.1109\/CVPR.2017.540"},{"key":"e_1_3_1_14_2","first-page":"8410","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","author":"Gu Yuxian","year":"2022","unstructured":"Yuxian Gu, Xu Han, Zhiyuan Liu, and Minlie Huang. 2022. PPT: Pre-trained prompt tuning for few-shot learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics. 8410\u20138423."},{"key":"e_1_3_1_15_2","doi-asserted-by":"crossref","first-page":"14078","DOI":"10.18653\/v1\/2023.findings-acl.885","volume-title":"Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023","author":"Gupta Shivanshu","year":"2023","unstructured":"Shivanshu Gupta, Yoshitomo Matsubara, Ankit Chadha, and Alessandro Moschitti. 2023. Cross-lingual knowledge distillation for answer sentence selection in low-resource languages. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023. 14078\u201314092."},{"key":"e_1_3_1_16_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Hendrycks Dan","year":"2016","unstructured":"Dan Hendrycks and Kevin Gimpel. 2016. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_17_2","volume-title":"Proceedings of the Conference on Neural Information Processing Systems","author":"Hinton Geoffrey E.","year":"2015","unstructured":"Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean. 2015. Distilling the knowledge in a neural network. In Proceedings of the Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_18_2","unstructured":"Nithin Holla Pushkar Mishra Helen Yannakoudakis and Ekaterina Shutova. 2020. Meta-Learning with Sparse Experience Replay for Lifelong Language Learning. arXiv:2009.04891. Retrieved from https:\/\/arxiv.org\/abs\/2009.04891. (2020)."},{"key":"e_1_3_1_19_2","first-page":"1","volume-title":"Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Hou Boyu","year":"2023","unstructured":"Boyu Hou, Chengyu Wang, Xiaoqing Chen, Minghui Qiu, Liang Feng, and Jun Huang. 2023. Prompt-distiller: Few-shot knowledge distillation for prompt-based language learners with dual contrastive learning. In Proceedings of the 2023 IEEE International Conference on Acoustics, Speech and Signal Processing. 1\u20135."},{"key":"e_1_3_1_20_2","unstructured":"J. Edward Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang and Weizhu Chen. 2021. LoRA: Low-rank adaptation of large language models. arXiv:2106.09685. Retrieved from https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_3_1_21_2","first-page":"16050","article-title":"Class-incremental learning by knowledge distillation with adaptive feature consolidation","author":"Kang Minsoo","year":"2022","unstructured":"Minsoo Kang, Jaeyoo Park, and Bohyung Han. 2022. Class-incremental learning by knowledge distillation with adaptive feature consolidation. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201922) (2022), 16050\u201316059. https:\/\/ieeexplore.ieee.org\/document\/9879128","journal-title":"IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR\u201922)"},{"key":"e_1_3_1_22_2","first-page":"3167","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Kim Sungnyun","year":"2024","unstructured":"Sungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan, and Stefano Soatto. 2024. DocKD: Knowledge distillation from LLMs for open-world document understanding models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 3167\u20133193."},{"key":"e_1_3_1_23_2","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Lewis Mike","year":"2021","unstructured":"Mike Lewis, Shruti Bhosale, Tim Dettmers, Naman Goyal, and Luke Zettlemoyer. 2021. BASE layers: Simplifying training of large, sparse models. In Proceedings of the International Conference on Machine Learning."},{"key":"e_1_3_1_24_2","first-page":"1","volume-title":"Proceedings of the 2022 International Joint Conference on Neural Networks","author":"Li Hongjia","year":"2022","unstructured":"Hongjia Li, Lingyu Yang, Lei Li, Chengyin Xu, Shu-Tao Xia, and Chun Yuan. 2022. PTS: A prompt-based teacher-student network for weakly supervised aspect detection. In Proceedings of the 2022 International Joint Conference on Neural Networks. 1\u20138. DOI:10.1109\/IJCNN55064.2022.9892147"},{"key":"e_1_3_1_25_2","first-page":"4582","volume-title":"Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics","author":"Li Xiang Lisa","year":"2021","unstructured":"Xiang Lisa Li and Percy Liang. 2021. Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics. 4582\u20134597."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2022.3153264"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3539618.3592043"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_3_1_29_2","first-page":"5871","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Lu Qixi","year":"2024","unstructured":"Qixi Lu, Endong Xun, and Gongbo Tang. 2024. MTA4DPR: Multi-teaching-assistants based iterative knowledge distillation for dense passage retrieval. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 5871\u20135883."},{"key":"e_1_3_1_30_2","volume-title":"Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Ma Jiaqi","year":"2018","unstructured":"Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. 2018. Modeling task relationships in multi-task learning with multi-gate mixture-of-experts. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining."},{"key":"e_1_3_1_31_2","first-page":"218","volume-title":"Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases","author":"Meng Xv","year":"2024","unstructured":"Xv Meng, Jun Rao, Shuhan Qi, Lei Wang, Jing Xiao, and Xuan Wang. 2024. Harnessing the power of prompt experts: Efficient knowledge distillation for enhanced language understanding. In Proceedings of the Joint European Conference on Machine Learning and Knowledge Discovery in Databases. Springer, 218\u2013234."},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"3675","DOI":"10.18653\/v1\/2024.emnlp-main.215","volume-title":"Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing","author":"Palo Flavio Di","year":"2024","unstructured":"Flavio Di Palo, Prateek Singhi, and Bilal H Fadlallah. 2024. Performance-guided LLM knowledge distillation for efficient text classification at scale. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 3675\u20133687."},{"key":"e_1_3_1_33_2","first-page":"3962","volume-title":"Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Park Wonpyo","year":"2019","unstructured":"Wonpyo Park, Dongju Kim, Yan Lu, and Minsu Cho. 2019. Relational knowledge distillation. In Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 3962\u20133971. DOI:10.1109\/CVPR.2019.00409"},{"key":"e_1_3_1_34_2","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Passalis Nikolaos","year":"2018","unstructured":"Nikolaos Passalis and Anastasios Tefas. 2018. Learning deep representations with probabilistic knowledge transfer. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_1_35_2","first-page":"487","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume","author":"Pfeiffer Jonas","year":"2021","unstructured":"Jonas Pfeiffer, Aishwarya Kamath, Andreas R\u00fcckl\u00e9, Kyunghyun Cho, and Iryna Gurevych. 2021. AdapterFusion: Non-destructive task composition for transfer learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 487\u2013503."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2023.103510"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3236837"},{"key":"e_1_3_1_38_2","volume-title":"Proceedings of the ACL(Findings)","author":"Rao Jun","year":"2025","unstructured":"Jun Rao, Zepeng Lin, Xuebo Liu, Lian Lian, Dong Jin, Shengjun Cheng, Jun Yu, and Min Zhang. 2025. APT: Improving specialist LLM performance with weakness case acquisition and iterative preference training. In Proceedings of the ACL(Findings). Association for Computational Linguistics."},{"key":"e_1_3_1_39_2","first-page":"10064","volume-title":"Proceedings of the EMNLP","author":"Rao Jun","year":"2024","unstructured":"Jun Rao, Xuebo Liu, Lian Lian, Shengjun Cheng, Yunjie Liao, and Min Zhang. 2024. CommonIT: Commonality-aware instruction tuning for large language models via data partitions. In Proceedings of the EMNLP, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for Computational Linguistics, Miami, Florida, USA, 10064\u201310083. DOI:10.18653\/v1\/2024.emnlp-main.561"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3321480"},{"key":"e_1_3_1_41_2","volume-title":"Proceedings of the Conference on Computational Natural Language Learning","author":"Sang Erik Tjong Kim","year":"2003","unstructured":"Erik Tjong Kim Sang and Fien De Meulder. 2003. Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition. In Proceedings of the Conference on Computational Natural Language Learning."},{"key":"e_1_3_1_42_2","first-page":"255","volume-title":"Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics","author":"Schick Timo","year":"2021","unstructured":"Timo Schick and Hinrich Sch\u00fctze. 2021. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics. 255\u2013269."},{"key":"e_1_3_1_43_2","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Kim Byung Cheol Song, Seung Hyun Lee, and Dae Ha","year":"2018","unstructured":"Byung Cheol Song, Seung Hyun Lee, and Dae Ha Kim. 2018. Self-supervised knowledge distillation using singular value decomposition. In Proceedings of the European Conference on Computer Vision."},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","unstructured":"Haizhou Shi Zihao Xu Hengyi Wang Weiyi Qin Wenyuan Wang Yibin Wang and Hao Wang. 2024. Continual learning of large language models: A comprehensive survey. ACM Computing Surveys 58 5 Article 120 (2024) 1\u201342. 10.1145\/3735633","DOI":"10.1145\/3735633"},{"key":"e_1_3_1_45_2","first-page":"4222","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Shin Taylor","year":"2020","unstructured":"Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. AutoPrompt: Eliciting knowledge from language models with automatically generated prompts. In Proceedings of the Conference on Empirical Methods in Natural Language Processing. 4222\u20134235."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3209978.320999"},{"key":"e_1_3_1_47_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Sun Fan-Keng","year":"2019","unstructured":"Fan-Keng Sun, Cheng-Hao Ho, and Hung yi Lee. 2019. LAMOL: LAnguage MOdeling for lifelong language learning. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_48_2","volume-title":"Proceedings of the International Conference on Learning Representations","author":"Tian Yonglong","year":"2020","unstructured":"Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020. Contrastive representation distillation. In Proceedings of the International Conference on Learning Representations."},{"key":"e_1_3_1_49_2","volume-title":"Proceedings of the Conference on Neural Information Processing Systems","author":"Wang Alex","year":"2019","unstructured":"Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. 2019. SuperGLUE: A stickier benchmark for general-purpose language understanding systems. In Proceedings of the Conference on Neural Information Processing Systems."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01563"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00024"},{"key":"e_1_3_1_52_2","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","author":"Xu Julian McAuley, Wangchunshu Zhou, and Canwen","year":"2019","unstructured":"Julian McAuley, Wangchunshu Zhou, and Canwen Xu. 2019. BERT learns to teach: Knowledge distillation with meta learning. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics."},{"key":"e_1_3_1_53_2","unstructured":"Chuhan Wu Fangzhao Wu Tao Qi and Yongfeng Huang. 2022. Unified and effective ensemble knowledge distillation. arXiv:2204.00548. Retrieved from https:\/\/arxiv.org\/abs\/2204.00548"},{"key":"e_1_3_1_54_2","first-page":"2202","volume-title":"Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Wu Meng-Chieh","year":"2019","unstructured":"Meng-Chieh Wu, Ching-Te Chiu, and Kun-Hsuan Wu. 2019. Multi-teacher knowledge distillation for compressed video action recognition on deep neural networks. In Proceedings of the 2019 IEEE International Conference on Acoustics, Speech and Signal Processing. 2202\u20132206. DOI:10.1109\/ICASSP.2019.8682450"},{"key":"e_1_3_1_55_2","first-page":"1464","volume-title":"Proceedings of the 6th International Conference on Intelligent Computing and Signal Processing","author":"Xi Chen","year":"2021","unstructured":"Chen Xi, Xing Zhiqiang, and Cheng Yuyang. 2021. Introduction to model compression knowledge distillation. In Proceedings of the 6th International Conference on Intelligent Computing and Signal Processing. 1464\u20131467."},{"key":"e_1_3_1_56_2","volume-title":"Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics","author":"Ji Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, Jie Tang, Xiao Liu, and Kaixuan","year":"2022","unstructured":"Yicheng Fu, Weng Lam Tam, Zhengxiao Du, Zhilin Yang, Jie Tang, Xiao Liu, and Kaixuan Ji. 2022. P-tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics."},{"key":"e_1_3_1_57_2","volume-title":"Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"You Shan","year":"2017","unstructured":"Shan You, Chang Xu, Chao Xu, and Dacheng Tao. 2017. Learning from multiple teacher networks. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining."},{"key":"e_1_3_1_58_2","first-page":"941","volume-title":"Proceedings of the 2022 IEEE International Conference on Data Mining Workshops","author":"Yu Ping","year":"2022","unstructured":"Ping Yu, Wei Wang, Chunyuan Li, Ruiyi Zhang, Zhanpeng Jin, and Changyou Chen. 2022. STT: Soft template tuning for few-shot adaptation. In Proceedings of the 2022 IEEE International Conference on Data Mining Workshops. 941\u2013946. DOI:10.1109\/ICDMW58026.2022.00122"},{"key":"e_1_3_1_59_2","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence","author":"Yuan Fei","year":"2020","unstructured":"Fei Yuan, Linjun Shou, Jian Pei, Wutao Lin, Ming Gong, Yan Fu, and Daxin Jiang. 2020. Reinforced multi-teacher selection for knowledge distillation. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3336191.3371792"},{"key":"e_1_3_1_61_2","volume-title":"Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics","author":"Zhang Songming","year":"2023","unstructured":"Songming Zhang, Yunlong Liang, Shuaibo Wang, Wenjuan Han, Jian Liu, Jinan Xu, and Yufeng Chen. 2023. Towards understanding and improving knowledge distillation for neural machine translation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics."},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.3037821"},{"key":"e_1_3_1_63_2","volume-title":"Proceedings of the 13th ACM Conference on Recommender Systems","author":"Zhao Zhe","year":"2019","unstructured":"Zhe Zhao, Lichan Hong, Li Wei, Jilin Chen, Aniruddh Nath, Shawn Andrews, Aditee Kumthekar, Maheswaran Sathiamoorthy, Xinyang Yi, and Ed H. Chi. 2019. Recommending what video to watch next: A multitask ranking system. In Proceedings of the 13th ACM Conference on Recommender Systems."}],"container-title":["ACM Transactions on Asian and Low-Resource Language Information Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786604","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,12]],"date-time":"2026-02-12T22:20:22Z","timestamp":1770934822000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786604"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,12]]},"references-count":62,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3786604"],"URL":"https:\/\/doi.org\/10.1145\/3786604","relation":{},"ISSN":["2375-4699","2375-4702"],"issn-type":[{"value":"2375-4699","type":"print"},{"value":"2375-4702","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,12]]},"assertion":[{"value":"2025-03-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-22","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-02-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}