{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T17:39:43Z","timestamp":1785173983153,"version":"3.55.0"},"reference-count":32,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T00:00:00Z","timestamp":1732147200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Neurorobot."],"abstract":"<jats:sec><jats:title>Introduction<\/jats:title><jats:p>In recent years, with the rapid development of artificial intelligence technology, the field of music education has begun to explore new teaching models. Traditional music education research methods have primarily focused on single-modal studies such as note recognition and instrument performance techniques, often overlooking the importance of multimodal data integration and interactive teaching. Existing methods often struggle with handling multimodal data effectively, unable to fully utilize visual, auditory, and textual information for comprehensive analysis, which limits the effectiveness of teaching.<\/jats:p><\/jats:sec><jats:sec><jats:title>Methods<\/jats:title><jats:p>To address these challenges, this project introduces MusicARLtrans Net, a multimodal interactive music education agent system driven by reinforcement learning. The system integrates Speech-to-Text (STT) technology to achieve accurate transcription of user voice commands, utilizes the ALBEF (Align Before Fuse) model for aligning and integrating multimodal data, and applies reinforcement learning to optimize teaching strategies.<\/jats:p><\/jats:sec><jats:sec><jats:title>Results and discussion<\/jats:title><jats:p>This approach provides a personalized and real-time feedback interactive learning experience by effectively combining auditory, visual, and textual information. The system collects and annotates multimodal data related to music education, trains and integrates various modules, and ultimately delivers an efficient and intelligent music education agent. Experimental results demonstrate that MusicARLtrans Net significantly outperforms traditional methods, achieving an accuracy of <jats:bold>96.77%<\/jats:bold> on the LibriSpeech dataset and <jats:bold>97.55%<\/jats:bold> on the MS COCO dataset, with marked improvements in recall, F1 score, and AUC metrics. These results highlight the system's superiority in speech recognition accuracy, multimodal data understanding, and teaching strategy optimization, which together lead to enhanced learning outcomes and user satisfaction. The findings hold substantial academic and practical significance, demonstrating the potential of advanced AI-driven systems in revolutionizing music education.<\/jats:p><\/jats:sec>","DOI":"10.3389\/fnbot.2024.1479694","type":"journal-article","created":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T06:24:52Z","timestamp":1732170292000},"update-policy":"https:\/\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["MusicARLtrans Net: a multimodal agent interactive music education system driven via reinforcement learning"],"prefix":"10.3389","volume":"18","author":[{"given":"Jie","family":"Chang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhenmeng","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chao","family":"Yan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1965","published-online":{"date-parts":[[2024,11,21]]},"reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3297887","article-title":"Exploration of digital transformation of government governance under the information environment","author":"Ai","year":"2023","journal-title":"IEEE Access"},{"key":"B2","doi-asserted-by":"publisher","first-page":"351","DOI":"10.1016\/S0167-6393(02)00087-0","article-title":"Speech\/music segmentation using entropy and dynamism features in a hmm classification framework","volume":"40","author":"Ajmera","year":"2003","journal-title":"Speech Commun"},{"key":"B3","doi-asserted-by":"crossref","unstructured":"\u201cPortable expert system to voice and speech recognition using an open source computer hardware,\u201d\n          \n          564\n          568\n          \n            \n              Betancourt\n              H. E.\n            \n            \n              Armijos\n              D. A.\n            \n            \n              Martinez\n              P. N.\n            \n            \n              Ponce\n              A. E.\n            \n            \n              Ortega-Zamorano\n              F.\n            \n          \n          IEEE\n          2018 2nd European Conference on Electrical Engineering and Computer Science (EECS)\n          \n          2018","DOI":"10.1109\/EECS.2018.00110"},{"key":"B4","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1037\/pmu0000284","article-title":"Song properties and familiarity affect speech recognition in musical noise","volume":"32","author":"Brown","year":"2022","journal-title":"Psychomusicology"},{"key":"B5","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3397499","article-title":"Understanding optical music recognition","volume":"53","author":"Calvo-Zaragoza","year":"2020","journal-title":"ACM Comp. Surv"},{"key":"B6","doi-asserted-by":"publisher","first-page":"12607","DOI":"10.1609\/aaai.v37i11.26484","article-title":"Leveraging modality-specific representations for audio-visual speech recognition via reinforcement learning","volume":"37","author":"Chen","year":"2023","journal-title":"Proc. AAAI Conf. Artif. Intell"},{"key":"B7","doi-asserted-by":"crossref","unstructured":"\u201cContinuous speech separation: dataset and analysis,\u201d\n          \n          7284\n          7288\n          \n            \n              Chen\n              Z.\n            \n            \n              Yoshioka\n              T.\n            \n            \n              Lu\n              L.\n            \n            \n              Zhou\n              T.\n            \n            \n              Meng\n              Z.\n            \n            \n              Luo\n              Y.\n            \n          \n          25455337\n          IEEE\n          ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)\n          \n          2020","DOI":"10.1109\/ICASSP40776.2020.9053426"},{"key":"B8","doi-asserted-by":"publisher","first-page":"28","DOI":"10.1016\/j.asoc.2016.12.024","article-title":"An evaluation of convolutional neural networks for music classification using spectrograms","volume":"52","author":"Costa","year":"2017","journal-title":"Appl. Soft Comput"},{"key":"B9","doi-asserted-by":"crossref","unstructured":"\u201cExpert system for intelligent audio codification based in speech\/music discrimination,\u201d\n          \n          318\n          322\n          \n            \n              Exposito\n              J. M.\n            \n            \n              Galan\n              S. G.\n            \n            \n              Reyes\n              N. R.\n            \n            \n              Candeas\n              P. V.\n            \n            \n              Pena\n              F. R.\n            \n          \n          IEEE\n          2006 International Symposium on Evolving Fuzzy Systems\n          \n          2006","DOI":"10.1109\/ISEFS.2006.251182"},{"key":"B10","doi-asserted-by":"publisher","first-page":"4","DOI":"10.3389\/fnbot.2012.00004","article-title":"Bayesian exploration for intelligent identification of textures","volume":"6","author":"Fishel","year":"2012","journal-title":"Front. Neurorobot"},{"key":"B11","doi-asserted-by":"publisher","first-page":"109492","DOI":"10.1016\/j.apacoust.2023.109492","article-title":"Emotional speech recognition using cnn and deep learning techniques","volume":"211","author":"Hema","year":"2023","journal-title":"Appl. Acoust"},{"key":"B12","doi-asserted-by":"publisher","first-page":"1338104","DOI":"10.3389\/fnbot.2023.1338104","article-title":"Education robot object detection with a brain-inspired approach integrating faster R-CNN, YOLOv3, and semi-supervised learning","volume":"17","author":"Hong","year":"2024","journal-title":"Front. Neurorobot"},{"key":"B13","doi-asserted-by":"publisher","first-page":"107978","DOI":"10.1016\/j.compeleceng.2022.107978","article-title":"An intelligent music genre analysis using feature extraction and classification using deep learning techniques","volume":"100","author":"Hongdan","year":"2022","journal-title":"Comp. Elect. Eng"},{"key":"B14","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/2009\/239892","article-title":"A decision-tree-based algorithm for speech\/music classification and segmentation","volume":"2009","author":"Lavner","year":"2009","journal-title":"EURASIP J. Audio Speech Music Process"},{"key":"B15","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s13636-021-00215-6","article-title":"Adversarial joint training with self-attention mechanism for robust end-to-end speech recognition","volume":"2021","author":"Li","year":"2021","journal-title":"EURASIP J. Audio Speech Music Process"},{"key":"B16","unstructured":"Neural radiance fields convert 2D to 3D texture\n          \n          40\n          44\n          \n            \n              Lin\n              Z.\n            \n            \n              Wang\n              C.\n            \n            \n              Li\n              Z.\n            \n            \n              Wang\n              Z.\n            \n            \n              Liu\n              X.\n            \n            \n              Zhu\n              Y.\n            \n          \n          37478036\n          Appl. Sci. Biotechnol. J. Adv. Res\n          3"},{"key":"B17","unstructured":"Text sentiment detection and classification based on integrated learning algorithm\n          \n          27\n          33\n          \n            \n              Lin\n              Z.\n            \n            \n              Wang\n              Z.\n            \n            \n              Zhu\n              Y.\n            \n            \n              Li\n              Z.\n            \n            \n              Qin\n              H.\n            \n          \n          Appl. Sci. Eng. J. Adv. Res\n          3"},{"key":"B18","unstructured":"\u201cRule-based word pronunciation networks generation for mandarin speech recognition,\u201d\n          \n          35\n          38\n          \n            \n              Liu\n              Y.\n            \n            \n              Fung\n              P.\n            \n          \n          Citeseer\n          International Symposium of Chinese Spoken Language Processing\n          \n          2000"},{"key":"B19","doi-asserted-by":"crossref","unstructured":"\u201cAuto-AVSR: audio-visual speech recognition with automatic labels,\u201d\n          \n          1\n          5\n          \n            \n              Ma\n              P.\n            \n            \n              Haliassos\n              A.\n            \n            \n              Fernandez-Lopez\n              A.\n            \n            \n              Chen\n              H.\n            \n            \n              Petridis\n              S.\n            \n            \n              Pantic\n              M.\n            \n          \n          IEEE\n          ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)\n          \n          2023","DOI":"10.1109\/ICASSP49357.2023.10096889"},{"key":"B20","unstructured":"\u201cA rule-based approach to extracting relations from music tidbits,\u201d\n          \n          \n            \n              Oramas\n              S.\n            \n            \n              Sordo\n              M.\n            \n            \n              Espinosa-Anke\n              L.\n            \n          \n          Proceedings of the 24th International Conference on World Wide Web\n          \n          2015"},{"key":"B21","doi-asserted-by":"publisher","first-page":"108","DOI":"10.3389\/neuro.12.006.2007","article-title":"What is intrinsic motivation? A typology of computational approaches","volume":"1","author":"Oudeyer","year":"2007","journal-title":"Front. Neurorobot"},{"key":"B22","unstructured":"\u201cNon-negative matrix factorization based compensation of music for automatic speech recognition,\u201d\n          \n          \n            \n              Raj\n              B.\n            \n            \n              Virtanen\n              T.\n            \n            \n              Chaudhuri\n              S.\n            \n            \n              Singh\n              R.\n            \n          \n          Interspeech\n          \n          2010"},{"key":"B23","unstructured":"\u201cAllies: a speech corpus for segmentation, speaker diarization, speech recognition and speaker change detection,\u201d\n          \n          \n            \n              Tahon\n              M.\n            \n            \n              Larcher\n              A.\n            \n            \n              Lebourdais\n              M.\n            \n            \n              Bougares\n              F.\n            \n            \n              Silnova\n              A.\n            \n            \n              Gimeno\n              P.\n            \n          \n          Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)\n          \n          2024"},{"key":"B24","doi-asserted-by":"crossref","unstructured":"\u201cRandom forest algorithm for improving the performance of speech\/non-speech detection,\u201d\n          \n          28\n          32\n          \n            \n              Thambi\n              S. V.\n            \n            \n              Sreekumar\n              K.\n            \n            \n              Kumar\n              C. S.\n            \n            \n              Raj\n              P. R.\n            \n          \n          IEEE\n          2014 First International Conference on Computational Systems and Communications (ICCSC)\n          \n          2014","DOI":"10.1109\/COMPSC.2014.7032615"},{"key":"B25","doi-asserted-by":"publisher","first-page":"103830","DOI":"10.1016\/j.jvcir.2023.103830","article-title":"Rethinking pascal-VOC and MS-COCO dataset for small object detection","volume":"93","author":"Tong","year":"2023","journal-title":"J. Vis. Commun. Image Rep"},{"key":"B26","unstructured":"\u201cFrom imagenet to image classification: contextualizing progress on benchmarks,\u201d\n          \n          9625\n          9635\n          \n            \n              Tsipras\n              D.\n            \n            \n              Santurkar\n              S.\n            \n            \n              Engstrom\n              L.\n            \n            \n              Ilyas\n              A.\n            \n            \n              Madry\n              A.\n            \n          \n          PMLR\n          International Conference on Machine Learning\n          \n          2020"},{"key":"B27","doi-asserted-by":"publisher","first-page":"100721","DOI":"10.1016\/j.entcom.2024.100721","article-title":"Personalized recommendation based on improved speech recognition algorithm in music e-learning course simulation","volume":"52","author":"Wang","year":"2025","journal-title":"Entertain. Comput"},{"key":"B28","doi-asserted-by":"publisher","first-page":"036030","DOI":"10.1117\/1.JRS.10.036030","article-title":"Speckle-reducing scale-invariant feature transform match for synthetic aperture radar image registration","volume":"10","author":"Wang","year":"2016","journal-title":"J. Appl. Remote Sens"},{"key":"B29","doi-asserted-by":"publisher","first-page":"12","DOI":"10.1016\/j.jpdc.2019.03.003","article-title":"The security of machine learning in an adversarial setting: a survey","volume":"130","author":"Wang","year":"","journal-title":"J. Parallel Distrib. Comput"},{"key":"B30","doi-asserted-by":"publisher","first-page":"294","DOI":"10.1016\/j.ins.2019.07.023","article-title":"Multilevel similarity model for high-resolution remote sensing image registration","volume":"505","author":"Wang","year":"","journal-title":"Inf. Sci"},{"key":"B31","doi-asserted-by":"publisher","first-page":"118243","DOI":"10.1109\/ACCESS.2022.3220878","article-title":"A sequence-to-sequence framework based on transformer with masked language model for optical music recognition","volume":"10","author":"Wen","year":"2022","journal-title":"IEEE Access"},{"key":"B32","doi-asserted-by":"crossref","unstructured":"\u201cMusic removal by convolutional denoising autoencoder in speech recognition,\u201d\n          \n          338\n          341\n          \n            \n              Zhao\n              M.\n            \n            \n              Wang\n              D.\n            \n            \n              Zhang\n              Z.\n            \n            \n              Zhang\n              X.\n            \n          \n          IEEE\n          2015 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)\n          \n          2015","DOI":"10.1109\/APSIPA.2015.7415289"}],"container-title":["Frontiers in Neurorobotics"],"original-title":[],"link":[{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2024.1479694\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,21]],"date-time":"2024-11-21T06:25:02Z","timestamp":1732170302000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.frontiersin.org\/articles\/10.3389\/fnbot.2024.1479694\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,11,21]]},"references-count":32,"alternative-id":["10.3389\/fnbot.2024.1479694"],"URL":"https:\/\/doi.org\/10.3389\/fnbot.2024.1479694","relation":{},"ISSN":["1662-5218"],"issn-type":[{"value":"1662-5218","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,11,21]]},"article-number":"1479694"}}