{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,8]],"date-time":"2026-08-08T09:19:23Z","timestamp":1786180763117,"version":"3.56.0"},"reference-count":52,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2023,3,2]],"date-time":"2023-03-02T00:00:00Z","timestamp":1677715200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,3,2]],"date-time":"2023-03-02T00:00:00Z","timestamp":1677715200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Nat Mach Intell"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>With the prevalence of pre-trained language models (PLMs) and the pre-training\u2013fine-tuning paradigm, it has been continuously shown that larger models tend to yield better performance. However, as PLMs scale up, fine-tuning and storing all the parameters is prohibitively costly and eventually becomes practically infeasible. This necessitates a new branch of research focusing on the parameter-efficient adaptation of PLMs, which optimizes a small portion of the model parameters while keeping the rest fixed, drastically cutting down computation and storage costs. In general, it demonstrates that large-scale models could be effectively stimulated by the optimization of a few parameters. Despite the various designs, here we discuss and analyse the approaches under a more consistent and accessible term \u2018delta-tuning\u2019, where \u2018delta\u2019 a mathematical notation often used to denote changes, is borrowed to refer to the portion of parameters that are \u2018changed\u2019 during training. We formally describe the problem and propose a unified categorization criterion for existing delta-tuning methods to explore their correlations and differences. We also discuss the theoretical principles underlying the effectiveness of delta-tuning and interpret them from the perspectives of optimization and optimal control. Furthermore, we provide a holistic empirical study on over 100 natural language processing tasks and investigate various aspects of delta-tuning. With comprehensive study and analysis, our research demonstrates the theoretical and practical properties of delta-tuning in the adaptation of PLMs.<\/jats:p>","DOI":"10.1038\/s42256-023-00626-4","type":"journal-article","created":{"date-parts":[[2023,3,2]],"date-time":"2023-03-02T12:03:13Z","timestamp":1677758593000},"page":"220-235","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":975,"title":["Parameter-efficient fine-tuning of large-scale pre-trained language models"],"prefix":"10.1038","volume":"5","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8758-9484","authenticated-orcid":false,"given":"Ning","family":"Ding","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yujia","family":"Qin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guang","family":"Yang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Fuchao","family":"Wei","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zonghan","family":"Yang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yusheng","family":"Su","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shengding","family":"Hu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yulin","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chi-Min","family":"Chan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weize","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jing","family":"Yi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Weilin","family":"Zhao","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xiaozhi","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7709-2543","authenticated-orcid":false,"given":"Zhiyuan","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5128-5649","authenticated-orcid":false,"given":"Hai-Tao","family":"Zheng","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jianfei","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yang","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jie","family":"Tang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Juanzi","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6011-6115","authenticated-orcid":false,"given":"Maosong","family":"Sun","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,3,2]]},"reference":[{"key":"626_CR1","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun, Y., Bengio, Y. & Hinton, G. Deep learning. Nature 521, 436\u2013444 (2015).","journal-title":"Nature"},{"key":"626_CR2","first-page":"1735","volume":"9","author":"S Hochreiter","year":"1997","unstructured":"Hochreiter, S. & Schmidhuber, J. \u00fcrgen. Long short-term memory. Neural Comput. 9, 1735\u20131780 (1997).","journal-title":"Long short-term memory. Neural Comput."},{"key":"626_CR3","unstructured":"Bengio, Y., Ducharme, R. & Vincent, P. A neural probabilistic language model. In Advances in Neural Information Processing Systems. 13 (2000)."},{"key":"626_CR4","unstructured":"Vaswani, A. et al. Attention is all you need. In Advances in Neural Information Processing Systems. 30 (2017)."},{"key":"626_CR5","unstructured":"Devlin, J., Chang, M.-W., Lee, K. & Toutanova, K. BERT: pre-training of deep bidirectional transformers for language understanding. In Proc. the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1, 4171\u20134186 (2019)."},{"key":"626_CR6","unstructured":"Radford, A., Narasimhan, K., Salimans, T. & Sutskever, I. Improving language understanding by generative pre-training. OpenAI Blog. https:\/\/cdn.openai.com\/research-covers\/language-unsupervised\/language_understanding_paper.pdf (2018)."},{"key":"626_CR7","unstructured":"Radford, A. et al. Language models are unsupervised multitask learners. OpenAI Blog. https:\/\/d4mucfpksywv.cloudfront.net\/better-language-models\/language-models.pdf (2019)."},{"key":"626_CR8","first-page":"5485","volume":"21","author":"C Raffel","year":"2020","unstructured":"Raffel, C. et al. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 21, 5485\u20135551 (2020).","journal-title":"J. Mach. Learn. Res."},{"key":"626_CR9","unstructured":"Brown, T. et al. Language models are few-shot learners. In Advances in Neural Information Processing Systems. 33, 1877\u20131901 (2020)."},{"key":"626_CR10","unstructured":"Rae, J. W. et al. Scaling language models: methods, analysis & insights from training Gopher. Preprint at arXiv https:\/\/arxiv.org\/abs\/2112.11446 (2021)."},{"key":"626_CR11","unstructured":"Smith, S. et al. Using deepspeed and megatron to train Megatron-Turing NLG 530b, a large-scale generative language model. Preprint at arXiv https:\/\/arxiv.org\/abs\/2201.11990 (2022)."},{"key":"626_CR12","unstructured":"Chowdhery, A. et al. PaLM: scaling language modeling with pathways. Preprint at arXiv https:\/\/arxiv.org\/abs\/2204.02311 (2022)."},{"key":"626_CR13","unstructured":"Houlsby, N. et al. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning. (eds Chaudhuri, K. & Salakhutdinov, R.) 2790\u20132799 (2019)."},{"key":"626_CR14","doi-asserted-by":"crossref","unstructured":"Zaken, E. B., Ravfogel, S. & Goldberg, Y. Bitfit: simple parameter-efficient fine-tuning for transformer-based masked language-models. In Proc. the 60th Annual Meeting of the Association for Computational Linguistics. 2, 1\u20139 (2022).","DOI":"10.18653\/v1\/2022.acl-short.1"},{"key":"626_CR15","unstructured":"Hu, E. J. et al. LoRA: low-rank adaptation of large language models. In International Conference on Learning Representations (2022)."},{"key":"626_CR16","doi-asserted-by":"crossref","unstructured":"AAghajanyan, A., Gupt, S. & Zettlemoyer, L. Intrinsic dimensionality explains the effectiveness of language model fine-tuning. In Proc. the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. 1, 7319\u20137328 (2021).","DOI":"10.18653\/v1\/2021.acl-long.568"},{"key":"626_CR17","unstructured":"Qin, Y. et al. Exploring low-dimensional intrinsic task subspace via prompt tuning. Preprint at arXiv https:\/\/arxiv.org\/abs\/2110.07867 (2021)."},{"key":"626_CR18","unstructured":"Lhoest, Q. et al. Datasets: a community library for natural language processing. In Proc. the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 175\u2013184 (2021)."},{"key":"626_CR19","doi-asserted-by":"crossref","unstructured":"Lester, B., Al-Rfou, R. & Constant, N. The power of scale for parameter-efficient prompt tuning. In Proc. the 2021 Conference on Empirical Methods in Natural Language Processing. 3045\u20133059 (2021).","DOI":"10.18653\/v1\/2021.emnlp-main.243"},{"key":"626_CR20","unstructured":"Liu, Y. et al. Roberta: a robustly optimized BERT pretraining approach. Preprint at arXiv https:\/\/arxiv.org\/abs\/1907.11692 (2019)."},{"key":"626_CR21","doi-asserted-by":"crossref","unstructured":"Wang, A. et al. GLUE: a multi-task benchmark and analysis platform for natural language understanding. In International Conference on Learning Representations (2019).","DOI":"10.18653\/v1\/W18-5446"},{"key":"626_CR22","doi-asserted-by":"crossref","unstructured":"Schick, T. & Sch\u00fctze, H. Exploiting cloze-questions for few-shot text classification and natural language inference. In Proc. the 16th Conference of the European Chapter of the Association for Computational Linguistics. 255\u2013269 (2021).","DOI":"10.18653\/v1\/2021.eacl-main.20"},{"key":"626_CR23","unstructured":"Socher, R. et al. Recursive deep models for semantic compositionality over a sentiment treebank. In Proc. the 2013 Conference on Empirical Methods in Natural Language Processing. 1631\u20131642 (2013)."},{"key":"626_CR24","doi-asserted-by":"crossref","unstructured":"Su, Y. et al. On transferability of prompt tuning for natural language understanding. In Proc. the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 3949\u20133969 (2022).","DOI":"10.18653\/v1\/2022.naacl-main.290"},{"key":"626_CR25","doi-asserted-by":"crossref","unstructured":"Williams, A., Nangia, N. & Bowman, S. A broad-coverage challenge corpus for sentence understanding through inference. In Proc. the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 1, 1112\u20131122 (2018).","DOI":"10.18653\/v1\/N18-1101"},{"key":"626_CR26","doi-asserted-by":"crossref","unstructured":"Vu, T., Lester, B., Constant, N., Al-Rfou, R. & Cer, D. Spot: better frozen model adaptation through soft prompt transfer. In Proc. the 60th Annual Meeting of the Association for Computational Linguistics. 1, 5039\u20135059 (2022).","DOI":"10.18653\/v1\/2022.acl-long.346"},{"key":"626_CR27","doi-asserted-by":"crossref","unstructured":"Han, X. et al. Pre-trained models: Past, present and future. AI Open 2, 225-250. https:\/\/www.sciencedirect.com\/science\/article\/pii\/S2666651021000231 (2021).","DOI":"10.1016\/j.aiopen.2021.08.002"},{"key":"626_CR28","first-page":"1022","volume":"34","author":"RK Mahabadi","year":"2021","unstructured":"Mahabadi, R. K., Henderson, J. & Ruder, S. Compacter: efficient low-rank hypercomplex adapter layers. In Advances in Neural Information Processing Systems. 34, 1022\u20131035 (2021).","journal-title":"In Advances in Neural Information Processing Systems."},{"key":"626_CR29","unstructured":"Stickland, A. C. & Murray, I. BERT and pals: projected attention layers for efficient adaptation in multi-task learning. In International Conference on Machine Learning. 5986\u20135995 (2019)."},{"key":"626_CR30","unstructured":"Mahabadi, R. K., Ruder, S., Dehghani, M. & Henderson, J. Parameter-efficient multi-task fine-tuning for transformers via shared hypernetworks. In Proc. the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. 1, 565\u2013576 (2021)."},{"key":"626_CR31","doi-asserted-by":"crossref","unstructured":"Pfeiffer, J., Kamath, A., R\u00fcckl\u00e9, A., Cho, K. & Gurevych, I. AdapterFusion: non-destructive task composition for transfer learning. In Proc. the 16th Conference of the European Chapter of the Association for Computational Linguistics. 487\u2013503 (2021).","DOI":"10.18653\/v1\/2021.eacl-main.39"},{"key":"626_CR32","doi-asserted-by":"crossref","unstructured":"R\u00fcckl\u00e9, A. et al. AdapterDrop: in the efficiency of adapters in transformers. In Proc. the 2021 Conference on Empirical Methods in Natural Language Processing. 7930\u20137946 (2021).","DOI":"10.18653\/v1\/2021.emnlp-main.626"},{"key":"626_CR33","doi-asserted-by":"crossref","unstructured":"He, R. et al. On the effectiveness of adapter-based tuning for pretrained language model adaptation. In Proc. the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. 1, 2208\u20132222 (2021).","DOI":"10.18653\/v1\/2021.acl-long.172"},{"key":"626_CR34","doi-asserted-by":"crossref","unstructured":"Han, W., Pang, B. & Wu, Y. N. Robust transfer learning with pretrained language models through adapters. In Proc. the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. 2, 854\u2013861 (2021).","DOI":"10.18653\/v1\/2021.acl-short.108"},{"key":"626_CR35","doi-asserted-by":"crossref","unstructured":"Pfeiffer, J. et al. AdapterHub: a framework for adapting transformers. In Proc. the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations. 46\u201354 (2020).","DOI":"10.18653\/v1\/2020.emnlp-demos.7"},{"key":"626_CR36","doi-asserted-by":"crossref","unstructured":"Gao, T., Fisch, A. & Chen, D. Making pre-trained language models better few-shot learners. In Proc. the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. 1, 3816\u20133830 (2021).","DOI":"10.18653\/v1\/2021.acl-long.295"},{"key":"626_CR37","doi-asserted-by":"crossref","unstructured":"Hu, S. et al. Knowledgeable prompt-tuning: incorporating knowledge into prompt verbalizer for text classification. In Proc. the 60th Annual Meeting of the Association for Computational Linguistics. 1, 2225\u20132240 (2021).","DOI":"10.18653\/v1\/2022.acl-long.158"},{"key":"626_CR38","first-page":"1","volume":"55","author":"P Liu","year":"2023","unstructured":"Liu, P. et al. Pre-train, prompt, and predict: a systematic survey of prompting methods in natural language processing. ACM Comput. Surv. 55, 1\u201335 (2023).","journal-title":"ACM Comput. Surv."},{"key":"626_CR39","doi-asserted-by":"crossref","unstructured":"Ding, N. et al. Openprompt: an open-source framework for prompt-learning. In Proc. the 60th Annual Meeting of the Association for Computational Linguistics: System Demonstrations. 105\u2013113 (2022).","DOI":"10.18653\/v1\/2022.acl-demo.10"},{"key":"626_CR40","doi-asserted-by":"crossref","unstructured":"Li, X. L. & Liang, P. Prefix-tuning: optimizing continuous prompts for generation. In Proc. the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. 1, 4582\u20134597 (2021).","DOI":"10.18653\/v1\/2021.acl-long.353"},{"key":"626_CR41","first-page":"61","volume":"2","author":"X Liu","year":"2022","unstructured":"Liu, X. et al. P-tuning: prompt tuning can be comparable to fine-tuning universally across scales and tasks. In Proc. the 60th Annual Meeting of the Association for Computational Linguistics. 2, 61\u201368 (2022).","journal-title":"In Proc. the 60th Annual Meeting of the Association for Computational Linguistics."},{"key":"626_CR42","first-page":"8410","volume":"1","author":"Y Gu","year":"2022","unstructured":"Gu, Y., Han, X., Liu, S. & Huang, M. Ppt: pre-trained prompt tuning for few-shot learning. In Proc. the 60th Annual Meeting of the Association for Computational Linguistics. 1, 8410\u20138423 (2022).","journal-title":"In Proc. the 60th Annual Meeting of the Association for Computational Linguistics."},{"key":"626_CR43","unstructured":"Lee, J., Tang, R. & Lin, J. What would elsa do? Freezing layers during transformer fine-tuning. Preprint at arXiv https:\/\/arxiv.org\/abs\/1911.03090 (2019)."},{"key":"626_CR44","doi-asserted-by":"crossref","unstructured":"Guo, D., Rush, A. & Kim, Y. Parameter-efficient transfer learning with diff pruning. In Proc. the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. 1, 4884\u20134896 (2021).","DOI":"10.18653\/v1\/2021.acl-long.378"},{"key":"626_CR45","doi-asserted-by":"crossref","unstructured":"Zhao, M., Lin, T., Mi, F., Jaggi, M. & Sch\u00fctze, H. Masking as an efficient alternative to finetuning for pretrained language models. In Proc. the 2020 Conference on Empirical Methods in Natural Language Processing. 2226\u20132241 (2020).","DOI":"10.18653\/v1\/2020.emnlp-main.174"},{"key":"626_CR46","unstructured":"Li, C., Farkhoor, H., Liu, R. & Yosinski, J. Measuring the intrinsic dimension of objective landscapes. In International Conference on Learning Representations (2018)."},{"key":"626_CR47","doi-asserted-by":"publisher","first-page":"585","DOI":"10.4208\/csiam-am.SO-2021-0016","volume":"2","author":"X Liu","year":"2021","unstructured":"Liu, X., Wen, Z. & Yuan, Y.-X. Subspace methods for nonlinear optimization. CSIAM Trans. Appl. Math. 2, 585\u2013651 (2021).","journal-title":"CSIAM Trans. Appl. Math."},{"key":"626_CR48","unstructured":"Yang, Z. & Liu, Y. On robust prefix-tuning for text classification. In International Conference on Learning Representations (2022)."},{"key":"626_CR49","unstructured":"Yang, Z., Yi, X., Li, P., Liu, Y. & Xie, X. Unified detoxifying and debiasing in language generation via inference-time adaptive optimization. Preprint at arXiv https:\/\/arxiv.org\/abs\/2210.04492 (2022)."},{"key":"626_CR50","unstructured":"Boyd, S. P. & Barratt, C. H. Linear Controller Design: Limits of Performance Vol. 7 (Citeseer, 1991)."},{"key":"626_CR51","doi-asserted-by":"publisher","first-page":"559","DOI":"10.1109\/TCST.2005.847331","volume":"13","author":"KH Ang","year":"2005","unstructured":"Ang, K. H., Chong, G. & Li, Y. PID control system analysis, design, and technology. IEEE Trans. Control Syst. Technol. 13, 559\u2013576 (2005).","journal-title":"IEEE Trans. Control Syst. Technol."},{"key":"626_CR52","unstructured":"He, J., Zhou, C., Ma, X., Berg-Kirkpatrick, T. & Neubig, G. Towards a unified view of parameter-efficient transfer learning. In International Conference on Learning Representations (2022)."}],"container-title":["Nature Machine Intelligence"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s42256-023-00626-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s42256-023-00626-4","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s42256-023-00626-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,22]],"date-time":"2023-03-22T20:07:33Z","timestamp":1679515653000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s42256-023-00626-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,2]]},"references-count":52,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["626"],"URL":"https:\/\/doi.org\/10.1038\/s42256-023-00626-4","relation":{"has-preprint":[{"id-type":"doi","id":"10.21203\/rs.3.rs-1553541\/v1","asserted-by":"object"}]},"ISSN":["2522-5839"],"issn-type":[{"value":"2522-5839","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,3,2]]},"assertion":[{"value":"13 April 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 February 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 March 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}