{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,5]],"date-time":"2026-07-05T04:31:36Z","timestamp":1783225896922,"version":"3.54.6"},"reference-count":28,"publisher":"Springer Science and Business Media LLC","issue":"4","license":[{"start":{"date-parts":[[2024,7,5]],"date-time":"2024-07-05T00:00:00Z","timestamp":1720137600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,7,5]],"date-time":"2024-07-05T00:00:00Z","timestamp":1720137600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006407","name":"Natural Science Foundation of Henan Province","doi-asserted-by":"publisher","award":["232300421240"],"award-info":[{"award-number":["232300421240"]}],"id":[{"id":"10.13039\/501100006407","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62171470"],"award-info":[{"award-number":["62171470"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Henan Zhongyuan Science and Technology Innovation Leading Talent Project","award":["234200510019"],"award-info":[{"award-number":["234200510019"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Process Lett"],"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Meta-learning has proven to be a powerful paradigm for transferring knowledge from prior tasks to facilitate the quick learning of new tasks in automatic speech recognition. However, the differences between languages (tasks) lead to variations in task learning directions, causing the harmful competition for model\u2019s limited resources. To address this challenge, we introduce the task-agreement multilingual meta-learning (TAMML), which adopts the gradient agreement algorithm to guide the model parameters towards a direction where tasks exhibit greater consistency. However, the computation and storage cost of TAMML grows dramatically with model\u2019s depth increases. To address this, we further propose a simplification called TAMML-Light which only uses the output layer for gradient calculation. Experiments on three datasets demonstrate that TAMML and TAMML-Light achieve outperform meta-learning approaches, yielding superior results.Furthermore, TAMML-Light can reduce at least 80<jats:inline-formula><jats:alternatives><jats:tex-math>$$\\%$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\"><mml:mo>%<\/mml:mo><\/mml:math><\/jats:alternatives><\/jats:inline-formula>of the relative increased computation expenses compared to TAMML.<\/jats:p>","DOI":"10.1007\/s11063-024-11661-6","type":"journal-article","created":{"date-parts":[[2024,7,5]],"date-time":"2024-07-05T04:01:47Z","timestamp":1720152107000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["A Lightweight Task-Agreement Meta Learning for Low-Resource Speech Recognition"],"prefix":"10.1007","volume":"56","author":[{"given":"Yaqi","family":"Chen","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hao","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenlin","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dan","family":"Qu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xukui","family":"Yang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,7,5]]},"reference":[{"key":"11661_CR1","unstructured":"Baevski A, Zhou H, Mohamed A-r, Auli M (2020) wav2vec 2.0: A framework for self-supervised learning of speech representations. ArXiv arxiv:2006.11477"},{"key":"11661_CR2","doi-asserted-by":"crossref","unstructured":"Pratap V, Sriram A, Tomasello P, Hannun AY, Liptchinsky V, Synnaeve G, Collobert R (2020) Massively multilingual asr: 50 languages, 1 model, 1 billion parameters. ArXiv arxiv:2007.03001","DOI":"10.21437\/Interspeech.2020-2831"},{"key":"11661_CR3","doi-asserted-by":"publisher","first-page":"3451","DOI":"10.1109\/TASLP.2021.3122291","volume":"29","author":"W-N Hsu","year":"2021","unstructured":"Hsu W-N, Bolte B, Tsai Y-HH, Lakhotia K, Salakhutdinov R, Mohamed A-r (2021) Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE\/ACM Trans Audio Speech Lang Proc 29:3451\u20133460","journal-title":"IEEE\/ACM Trans Audio Speech Lang Proc"},{"key":"11661_CR4","doi-asserted-by":"crossref","unstructured":"Luo J, Wang J, Cheng N, Zheng Z, Xiao J (2022) Adaptive activation network for low resource multilingual speech recognition. 2022 Int Jt Conf Neural Netw (IJCNN), 1\u20137","DOI":"10.1109\/IJCNN55064.2022.9892396"},{"key":"11661_CR5","doi-asserted-by":"crossref","unstructured":"Hou W, Dong Y, Zhuang B, Yang L, Shi J, Shinozaki T (2020) Large-scale end-to-end multilingual speech recognition and language identification with multi-task learning. In: Interspeech","DOI":"10.21437\/Interspeech.2020-2164"},{"key":"11661_CR6","unstructured":"Madhavaraj A, Ganesan RA (2022) Data and knowledge-driven approaches for multilingual training to improve the performance of speech recognition systems of indian languages. ArXiv arxiv:2201.09494"},{"key":"11661_CR7","first-page":"5149","volume":"44","author":"TM Hospedales","year":"2020","unstructured":"Hospedales TM, Antoniou A, Micaelli P, Storkey AJ (2020) Meta-learning in neural networks: a survey. IEEE Trans Pattern Anal Mach Intell 44:5149\u20135169","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"11661_CR8","doi-asserted-by":"crossref","unstructured":"Hsu J, Chen Y, Lee H (2020) Meta learning for end-to-end low-resource speech recognition. In: ICASSP, pp. 7844\u20137848","DOI":"10.1109\/ICASSP40776.2020.9053112"},{"key":"11661_CR9","doi-asserted-by":"publisher","first-page":"317","DOI":"10.1109\/TASLP.2021.3138674","volume":"30","author":"W Hou","year":"2021","unstructured":"Hou W, Zhu H, Wang Y, Wang J, Qin T, Xu R, Shinozaki T (2021) Exploiting adapters for cross-lingual low-resource speech recognition. IEEE\/ACM Trans Audio Speech Lang Proc 30:317\u2013329","journal-title":"IEEE\/ACM Trans Audio Speech Lang Proc"},{"key":"11661_CR10","doi-asserted-by":"publisher","unstructured":"Winata GI, Cahyawijaya S, Liu Z, Lin Z, Madotto A, Xu P, Fung P (2020) Learning fast adaptation on cross-accented speech recognition. In: Meng H, Xu B, Zheng TF (eds.) Interspeech, pp. 1276\u20131280. https:\/\/doi.org\/10.21437\/Interspeech.2020-0045","DOI":"10.21437\/Interspeech.2020-0045"},{"key":"11661_CR11","doi-asserted-by":"publisher","unstructured":"Chopra S, Mathur P, Sawhney R, Shah RR (2021) Meta-learning for low-resource speech emotion recognition. In: ICASSP , pp. 6259\u20136263. https:\/\/doi.org\/10.1109\/ICASSP39728.2021.9414373","DOI":"10.1109\/ICASSP39728.2021.9414373"},{"key":"11661_CR12","doi-asserted-by":"publisher","unstructured":"Klejch O, Fainberg J, Bell P, Renals S (2019) Speaker adaptive training using model agnostic meta-learning. In: ASRU, pp. 881\u2013888. https:\/\/doi.org\/10.1109\/ASRU46091.2019.9003751","DOI":"10.1109\/ASRU46091.2019.9003751"},{"key":"11661_CR13","doi-asserted-by":"crossref","unstructured":"Xiao Y, Gong K, Zhou P, et al (2021) Adversarial meta sampling for multilingual low-resource speech recognition. Proceed AAAI Conf Artif Intell 35(16):14112\u201314120","DOI":"10.1609\/aaai.v35i16.17661"},{"key":"11661_CR14","doi-asserted-by":"crossref","unstructured":"Singh S, Wang R, Hou F (2022) Improved meta learning for low resource speech recognition. ICASSP 2022\u20132022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4798\u20134802","DOI":"10.1109\/ICASSP43922.2022.9746899"},{"key":"11661_CR15","doi-asserted-by":"publisher","unstructured":"Hou W, Wang Y, Gao S, Shinozaki T (2021) Meta-adapter: Efficient cross-lingual adaptation with meta-learning. In: ICASSP, pp. 7028\u20137032. https:\/\/doi.org\/10.1109\/ICASSP39728.2021.9414959","DOI":"10.1109\/ICASSP39728.2021.9414959"},{"key":"11661_CR16","unstructured":"Eshratifar AE, Eigen D, Pedram M (2018) Gradient agreement as an optimization objective for meta-learning. CoRR arxiv:1810.08178"},{"key":"11661_CR17","doi-asserted-by":"publisher","first-page":"1227","DOI":"10.1109\/JSTSP.2022.3184480","volume":"16","author":"J Zhao","year":"2022","unstructured":"Zhao J, Zhang W (2022) Improving automatic speech recognition performance for low-resource languages with self-supervised models. IEEE J Sel Top Signal Proc 16:1227\u20131241","journal-title":"IEEE J Sel Top Signal Proc"},{"key":"11661_CR18","doi-asserted-by":"crossref","unstructured":"Kim S, Hori T, Watanabe S (2017) Joint CTC-attention based end-toend speech recognition using multi-task learning. IEEE international conference on acoustics, speech and signal processing (ICASSP). IEEE 2017:4835\u20134839","DOI":"10.1109\/ICASSP.2017.7953075"},{"key":"11661_CR19","doi-asserted-by":"crossref","unstructured":"Sennrich R, Haddow B, Birch A (2016) Neural machine translation of rare words with subword units. In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp 1715\u20131725","DOI":"10.18653\/v1\/P16-1162"},{"key":"11661_CR20","unstructured":"Zhou S, Xu S, Xu B (2018) Multilingual end-to-end speech recognition with a single transformer on low-resource languages. ArXiv arxiv:1806.05059"},{"key":"11661_CR21","unstructured":"Gales MJF, Knill KM, Ragni A, Rath SP (2014) Speech recognition and keyword spotting for low-resource languages: Babel project research at CUED. In: 4th Workshop on Spoken Language Technologies for Under-resourced Languages, SLTU 2014, pp. 16\u201323. St. Petersburg, Russia, May 14-16, 2014"},{"key":"11661_CR22","unstructured":"Ardila R, Branson M, Davis K, Henretty M, Kohler M, Meyer J, Morais R, Saunders L, Tyers FM, Weber G (2019) Common voice: A massively-multilingual speech corpus"},{"key":"11661_CR23","doi-asserted-by":"crossref","unstructured":"Kim S, Hori T, Watanabe S (2016) Joint ctc-attention based end-to-end speech recognition using multi-task learning. 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 4835\u20134839","DOI":"10.1109\/ICASSP.2017.7953075"},{"key":"11661_CR24","unstructured":"Kingma DP, Ba J (2015) Adam: A method for stochastic optimization. In: Bengio Y, LeCun Y (eds.) 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings . arxiv:1412.6980"},{"key":"11661_CR25","doi-asserted-by":"crossref","unstructured":"Gillick L, Cox SJ (1989) Some statistical issues in the comparison of speech recognition algorithms. In: International Conference on Acoustics, Speech, and Signal Processing, pp. 532\u2013535 . IEEE","DOI":"10.1109\/ICASSP.1989.266481"},{"key":"11661_CR26","unstructured":"Pallet DS, Fisher WM, Fiscus JG (1990) Tools for the analysis of benchmark speech recognition tests. In: International Conference on Acoustics, Speech, and Signal Processing, pp. 97\u2013100 . IEEE"},{"key":"11661_CR27","unstructured":"Finn C, Abbeel P, Levine S (2017) Model-agnostic meta-learning for fast adaptation of deep networks. In: Precup D, Teh YW (eds.) ICML , vol. 70, pp. 1126\u20131135"},{"key":"11661_CR28","unstructured":"Raghu A, Raghu M, Bengio S, Vinyals O (2020) Rapid learning or feature reuse? towards understanding the effectiveness of MAML. In: 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30"}],"container-title":["Neural Processing Letters"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11661-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11063-024-11661-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11063-024-11661-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,23]],"date-time":"2024-11-23T13:33:41Z","timestamp":1732368821000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11063-024-11661-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,5]]},"references-count":28,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2024,8]]}},"alternative-id":["11661"],"URL":"https:\/\/doi.org\/10.1007\/s11063-024-11661-6","relation":{},"ISSN":["1573-773X"],"issn-type":[{"value":"1573-773X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,5]]},"assertion":[{"value":"31 May 2024","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 July 2024","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"210"}}