{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T14:26:31Z","timestamp":1783347991891,"version":"3.54.6"},"reference-count":42,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2026,7,1]],"date-time":"2026-07-01T00:00:00Z","timestamp":1782864000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T00:00:00Z","timestamp":1783296000000},"content-version":"vor","delay-in-days":5,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100008769","name":"Julius-Maximilians-Universit\u00e4t W\u00fcrzburg","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100008769","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>In practical use, language models (LM) must efficiently adapt to new tasks and knowledge while avoiding catastrophic forgetting, a requirement that sparked research on continual learning (CL). Despite advances, current CL methods still lack a unified solution that delivers strong knowledge transfer and parameter-efficient capacity management while preventing catastrophic forgetting, which limits effective use of task synergies under tight training and memory budgets. We bridge this gap by introducing Gated Expandable Parameter-Efficient Fine-Tuning (GE-PEFT), a novel approach that shares knowledge of previous tasks through leveraging a single, dynamically expanding PEFT module within LMs while selectively gating irrelevant previous tasks. Our experiments across multiple task-incremental CL benchmarks show that GE-PEFT outperforms existing state-of-the-art CL approaches in both full CL and few-shot settings. Our ablation and parameter sensitivity studies highlight the benefit of each proposed component, demonstrating that GE-PEFT offers a more efficient and adaptive solution for CL in LMs.<\/jats:p>","DOI":"10.1007\/s10994-026-07104-z","type":"journal-article","created":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T14:02:39Z","timestamp":1783346559000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["GE-PEFT: Gated Expandable Parameter-Efficient Fine-Tuning for Continual Learning"],"prefix":"10.1007","volume":"115","author":[{"given":"Janna","family":"Omeliyanenko","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andreas","family":"Hotho","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Schl\u00f6r","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,7,6]]},"reference":[{"key":"7104_CR1","doi-asserted-by":"crossref","unstructured":"Aljundi, R. (2018). Memory aware synapses: Learning what (not) to forget. European Conference on Computer Vision (ECCV) (pp. 139\u2013154)","DOI":"10.1007\/978-3-030-01219-9_9"},{"key":"7104_CR2","unstructured":"Alabi, J. O., Adelani, D. I., Mosbach, M., & Klakow, D. (2022). Adapting pre-trained language models to African languages via multilingual adaptive fine-tuning. Computational Linguistics"},{"key":"7104_CR3","doi-asserted-by":"publisher","first-page":"6993","DOI":"10.1609\/aaai.v35i8.16861","volume":"35","author":"A Chaudhry","year":"2021","unstructured":"Chaudhry, A. (2021). Using hindsight to anchor past knowledge in continual learning. AAAI Conference on Artificial Intelligence, 35, 6993\u20137001.","journal-title":"AAAI Conference on Artificial Intelligence"},{"key":"7104_CR4","unstructured":"Costa-Juss\u00e0, M. R. (2022). No language left behind: Scaling human-centered machine translation arXiv:2207.04672 arXiv preprint."},{"key":"7104_CR5","unstructured":"Chaudhry, A., Ranzato, M., Rohrbach, M., & Elhoseiny, M. (2018). Efficient lifelong learning with a-gem arXiv:1812.00420 arXiv preprint."},{"key":"7104_CR6","doi-asserted-by":"crossref","unstructured":"Du, W. (2024). Unlocking continual learning abilities in language models. Findings of the ACL: EMNLP 2024 (pp. 6503\u20136522)","DOI":"10.18653\/v1\/2024.findings-emnlp.379"},{"key":"7104_CR7","unstructured":"Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). Bert: Pre-training of deep bidirectional transformers for language understanding (p. NAACL)"},{"key":"7104_CR8","unstructured":"Masson\u00a0D\u2019Autume, C., Ruder, S., Kong, L., & Yogatama, D. (2019). Episodic memory in lifelong language learning. Advances in Neural Information Processing Systems, 32"},{"key":"7104_CR9","unstructured":"Evci, U. (2022). Gradmax: Growing neural networks using gradient information arXiv:2201.05125 arXiv preprint."},{"key":"7104_CR10","unstructured":"GenAI, M. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288"},{"key":"7104_CR11","unstructured":"Houlsby, N. (2019). Parameter-efficient transfer learning for nlp. International Conference on Machine Learning (pp. 2790\u20132799) PMLR."},{"key":"7104_CR12","unstructured":"Hu, E. J. (2021). Lora: Low-rank adaptation of large language models arXiv:2106.09685 arXiv preprint."},{"key":"7104_CR13","doi-asserted-by":"crossref","unstructured":"Huang, Y. (2021). Continual learning for text classification with information disentanglement based regularization arXiv:2104.05489 arXiv preprint.","DOI":"10.18653\/v1\/2021.naacl-main.218"},{"key":"7104_CR14","doi-asserted-by":"crossref","unstructured":"Hyder, R. (2022). Incremental task learning with incremental rank updates (pp. 566\u2013582) European Conference on Computer Vision.","DOI":"10.1007\/978-3-031-20050-2_33"},{"key":"7104_CR15","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2015). Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. IEEE Computer Vision.","DOI":"10.1109\/ICCV.2015.123"},{"key":"7104_CR16","doi-asserted-by":"publisher","first-page":"3521","DOI":"10.1073\/pnas.1611835114","volume":"11413","author":"J Kirkpatrick","year":"2017","unstructured":"Kirkpatrick, J. (2017). Overcoming catastrophic forgetting in neural networks. Proceedings of the national academy of sciences, 11413, 3521\u20133526.","journal-title":"Proceedings of the national academy of sciences"},{"key":"7104_CR17","unstructured":"Ke, Z. (2021). Achieving forgetting prevention and knowledge transfer in continual learning. Advances in Neural Information Processing Systems, 34,"},{"key":"7104_CR18","doi-asserted-by":"crossref","unstructured":"Ke, Z. (2022). Continual training of language models for few-shot learning. Proceedings of EMNLP (pp. 10205\u201310216)","DOI":"10.18653\/v1\/2022.emnlp-main.695"},{"key":"7104_CR19","doi-asserted-by":"crossref","unstructured":"Kowsari, K., Brown, E., Heidarysafa, M., Meimandi, K., Gerber, M., & Barnes, L. (2016). Hierarchical deep learning for text classification. IEEE.","DOI":"10.1109\/ICMLA.2017.0-134"},{"key":"7104_CR20","unstructured":"Kilcher, Y., B\u00e9cigneul, G., & Hofmann, T. (2018). Escaping flat areas via function-preserving structural network modifications"},{"key":"7104_CR21","doi-asserted-by":"crossref","unstructured":"Li, X.L., Liang, P. (2021) Prefix-tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190","DOI":"10.18653\/v1\/2021.acl-long.353"},{"key":"7104_CR22","unstructured":"Loshchilov, I. (2017). Decoupled weight decay regularization arXiv:1711.05101 arXiv preprint."},{"key":"7104_CR23","unstructured":"Mirzadeh, S. I. (2020). Linear mode connectivity in multitask and continual learning arXiv:2010.04495 arXiv preprint."},{"key":"7104_CR24","doi-asserted-by":"crossref","unstructured":"Muhammad, S., (2023). AfriSenti: A Twitter sentiment analysis benchmark for African languages. In: EMNLP 2023, pp. 13968\u201313981. ACL, Singapore.","DOI":"10.18653\/v1\/2023.emnlp-main.862"},{"key":"7104_CR25","doi-asserted-by":"crossref","unstructured":"Omeliyanenko, J., Zehe, A., Hotho, A., & Schl\u00f6r, D. (2023). Capskg: Enabling continual knowledge integration in language models for automatic knowledge graph completion (p. ISWC)","DOI":"10.1007\/978-3-031-47240-4_33"},{"key":"7104_CR26","doi-asserted-by":"crossref","unstructured":"Petroni, F. (2019). Language models as knowledge bases? In: EMNLP-IJCNLP (pp. 2463\u20132473)","DOI":"10.18653\/v1\/D19-1250"},{"key":"7104_CR27","unstructured":"Qin, C., & Joty, S. (2022). Lfpt5: A unified framework for lifelong few-shot language learning based on prompt tuning of t5 (p. ICLR)"},{"key":"7104_CR28","first-page":"1","volume":"21140","author":"C Raffel","year":"2020","unstructured":"Raffel, C. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21140, 1\u201367.","journal-title":"Journal of machine learning research"},{"key":"7104_CR29","unstructured":"Razdaibiedina, A. (2023). Progressive prompts: Continual learning for language models arXiv:2301.12314 arXiv preprint."},{"key":"7104_CR30","doi-asserted-by":"crossref","unstructured":"Rebuffi, S.-A., Kolesnikov, A., Sperl, G., & Lampert, C.H. (2017). icarl: Incremental classifier and representation learning. In: CVPR, pp. 5533\u20135542 IEEE","DOI":"10.1109\/CVPR.2017.587"},{"key":"7104_CR31","unstructured":"Rusu, A. A., Rabinowitz, N. C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., & Hadsell, R. (2016). Progressive neural networks,"},{"key":"7104_CR32","unstructured":"Sun, F.-K., Ho, C.-H., & Lee, H.-Y. (2019). Lamol: Language modeling for lifelong language learning. International Conference on Learning Representations."},{"key":"7104_CR33","unstructured":"Serra, J., Suris, D., Miron, M., & Karatzoglou, A. (2018). Overcoming catastrophic forgetting with hard attention to the task. ICML (pp. 4548\u20134557)"},{"key":"7104_CR34","doi-asserted-by":"crossref","unstructured":"Valipour, M., Rezagholizadeh, M., Kobyzev, I., & Ghodsi, A. (2023). Dylora: Parameter-efficient tuning of pre-trained models using dynamic search-free low-rank adaptation. European Chapter of the ACL (pp. 3274\u20133287)","DOI":"10.18653\/v1\/2023.eacl-main.239"},{"key":"7104_CR35","first-page":"15173","volume":"33","author":"M Wortsman","year":"2020","unstructured":"Wortsman, M. (2020). Supermasks in superposition. Advances in Neural Information Processing Systems, 33, 15173\u201315184.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"7104_CR36","doi-asserted-by":"crossref","unstructured":"Wang, X. (2023a). Orthogonal subspace learning for language model continual learning. EMLNP,","DOI":"10.18653\/v1\/2023.findings-emnlp.715"},{"key":"7104_CR37","unstructured":"Wang, Z. (2023b). Rehearsal-free continual language learning via efficient parameter isolation. In: ACL Volume 1: Long Papers, pp. 10933\u201310946"},{"key":"7104_CR38","doi-asserted-by":"crossref","unstructured":"Wang, M. (2024). Rehearsal-free modular and compositional continual learning for language models. In: NAACL (Volume 2: Short Papers), pp. 469\u2013480","DOI":"10.18653\/v1\/2024.naacl-short.39"},{"key":"7104_CR39","doi-asserted-by":"crossref","unstructured":"Wu, Z., Tran, H., Pirsiavash, H., & Kolouri, S. (2023). Is multi-task learning an upper bound for continual learning? ICASSP 2023 (pp. 1\u20135). IEEE.","DOI":"10.1109\/ICASSP49357.2023.10095984"},{"key":"7104_CR40","unstructured":"Wu, L., Wang, D., & Liu, Q. (2019). Splitting steepest descent for growing neural architectures. Advances in neural information processing systems, 32,"},{"key":"7104_CR41","unstructured":"Zhang, Q., Chen, M., Bukharin, A., Karampatziakis, N., He, P., Cheng, Y., ..., & Zhao, T. (2023). Adaptive budget allocation for parameter-efficient fine-tuning. The Eleventh International Conference on Learning Representations."},{"key":"7104_CR42","unstructured":"Zhang, X., Zhao, J., LeCun, Y. (2015). Character-level convolutional networks for text classification. Advances in Neural Information Processing Systems, 28"}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07104-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-026-07104-z","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-026-07104-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,6]],"date-time":"2026-07-06T14:04:01Z","timestamp":1783346641000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-026-07104-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7]]},"references-count":42,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["7104"],"URL":"https:\/\/doi.org\/10.1007\/s10994-026-07104-z","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7]]},"assertion":[{"value":"15 January 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"26 May 2026","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 June 2026","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 July 2026","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing Interests"}}],"article-number":"171"}}