{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T16:38:55Z","timestamp":1783183135199,"version":"3.54.6"},"reference-count":46,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2022,5,11]],"date-time":"2022-05-11T00:00:00Z","timestamp":1652227200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62172267"],"award-info":[{"award-number":["62172267"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2019YFE0190500"],"award-info":[{"award-number":["2019YFE0190500"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["20ZR1420400"],"award-info":[{"award-number":["20ZR1420400"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61936001"],"award-info":[{"award-number":["61936001"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["21PJ1404200"],"award-info":[{"award-number":["21PJ1404200"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2021PE0AC02"],"award-info":[{"award-number":["2021PE0AC02"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&amp;D Program of China","doi-asserted-by":"publisher","award":["62172267"],"award-info":[{"award-number":["62172267"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&amp;D Program of China","doi-asserted-by":"publisher","award":["2019YFE0190500"],"award-info":[{"award-number":["2019YFE0190500"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&amp;D Program of China","doi-asserted-by":"publisher","award":["20ZR1420400"],"award-info":[{"award-number":["20ZR1420400"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&amp;D Program of China","doi-asserted-by":"publisher","award":["61936001"],"award-info":[{"award-number":["61936001"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&amp;D Program of China","doi-asserted-by":"publisher","award":["21PJ1404200"],"award-info":[{"award-number":["21PJ1404200"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012166","name":"National Key R&amp;D Program of China","doi-asserted-by":"publisher","award":["2021PE0AC02"],"award-info":[{"award-number":["2021PE0AC02"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007219","name":"Natural Science Foundation of Shanghai, China","doi-asserted-by":"publisher","award":["62172267"],"award-info":[{"award-number":["62172267"]}],"id":[{"id":"10.13039\/100007219","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007219","name":"Natural Science Foundation of Shanghai, China","doi-asserted-by":"publisher","award":["2019YFE0190500"],"award-info":[{"award-number":["2019YFE0190500"]}],"id":[{"id":"10.13039\/100007219","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007219","name":"Natural Science Foundation of Shanghai, China","doi-asserted-by":"publisher","award":["20ZR1420400"],"award-info":[{"award-number":["20ZR1420400"]}],"id":[{"id":"10.13039\/100007219","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007219","name":"Natural Science Foundation of Shanghai, China","doi-asserted-by":"publisher","award":["61936001"],"award-info":[{"award-number":["61936001"]}],"id":[{"id":"10.13039\/100007219","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007219","name":"Natural Science Foundation of Shanghai, China","doi-asserted-by":"publisher","award":["21PJ1404200"],"award-info":[{"award-number":["21PJ1404200"]}],"id":[{"id":"10.13039\/100007219","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100007219","name":"Natural Science Foundation of Shanghai, China","doi-asserted-by":"publisher","award":["2021PE0AC02"],"award-info":[{"award-number":["2021PE0AC02"]}],"id":[{"id":"10.13039\/100007219","id-type":"DOI","asserted-by":"publisher"}]},{"name":"State Key Program of National Natural Science Foundation of China","award":["62172267"],"award-info":[{"award-number":["62172267"]}]},{"name":"State Key Program of National Natural Science Foundation of China","award":["2019YFE0190500"],"award-info":[{"award-number":["2019YFE0190500"]}]},{"name":"State Key Program of National Natural Science Foundation of China","award":["20ZR1420400"],"award-info":[{"award-number":["20ZR1420400"]}]},{"name":"State Key Program of National Natural Science Foundation of China","award":["61936001"],"award-info":[{"award-number":["61936001"]}]},{"name":"State Key Program of National Natural Science Foundation of China","award":["21PJ1404200"],"award-info":[{"award-number":["21PJ1404200"]}]},{"name":"State Key Program of National Natural Science Foundation of China","award":["2021PE0AC02"],"award-info":[{"award-number":["2021PE0AC02"]}]},{"name":"Shanghai Pujiang Program","award":["62172267"],"award-info":[{"award-number":["62172267"]}]},{"name":"Shanghai Pujiang Program","award":["2019YFE0190500"],"award-info":[{"award-number":["2019YFE0190500"]}]},{"name":"Shanghai Pujiang Program","award":["20ZR1420400"],"award-info":[{"award-number":["20ZR1420400"]}]},{"name":"Shanghai Pujiang Program","award":["61936001"],"award-info":[{"award-number":["61936001"]}]},{"name":"Shanghai Pujiang Program","award":["21PJ1404200"],"award-info":[{"award-number":["21PJ1404200"]}]},{"name":"Shanghai Pujiang Program","award":["2021PE0AC02"],"award-info":[{"award-number":["2021PE0AC02"]}]},{"name":"Key Research Project of Zhejiang Laboratory","award":["62172267"],"award-info":[{"award-number":["62172267"]}]},{"name":"Key Research Project of Zhejiang Laboratory","award":["2019YFE0190500"],"award-info":[{"award-number":["2019YFE0190500"]}]},{"name":"Key Research Project of Zhejiang Laboratory","award":["20ZR1420400"],"award-info":[{"award-number":["20ZR1420400"]}]},{"name":"Key Research Project of Zhejiang Laboratory","award":["61936001"],"award-info":[{"award-number":["61936001"]}]},{"name":"Key Research Project of Zhejiang Laboratory","award":["21PJ1404200"],"award-info":[{"award-number":["21PJ1404200"]}]},{"name":"Key Research Project of Zhejiang Laboratory","award":["2021PE0AC02"],"award-info":[{"award-number":["2021PE0AC02"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Large-scale automatic speech recognition model has achieved impressive performance. However, huge computational resources and massive amount of data are required to train an ASR model. Knowledge distillation is a prevalent model compression method which transfers the knowledge from large model to small model. To improve the efficiency of knowledge distillation for end-to-end speech recognition especially in the low-resource setting, a Mixup-based Knowledge Distillation (MKD) method is proposed which combines Mixup, a data-agnostic data augmentation method, with softmax-level knowledge distillation. A loss-level mixture is presented to address the problem caused by the non-linearity of label in the KL-divergence when adopting Mixup to the teacher\u2013student framework. It is mathematically shown that optimizing the mixture of loss function is equivalent to optimize an upper bound of the original knowledge distillation loss. The proposed MKD takes the advantage of Mixup and brings robustness to the model even with a small amount of training data. The experiments on Aishell-1 show that MKD obtains a 15.6% and 3.3% relative improvement on two student models with different parameter scales compared with the existing methods. Experiments on data efficiency demonstrate MKD achieves similar results with only half of the original dataset.<\/jats:p>","DOI":"10.3390\/a15050160","type":"journal-article","created":{"date-parts":[[2022,5,11]],"date-time":"2022-05-11T10:19:46Z","timestamp":1652264386000},"page":"160","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["MKD: Mixup-Based Knowledge Distillation for Mandarin End-to-End Speech Recognition"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5331-022X","authenticated-orcid":false,"given":"Xing","family":"Wu","sequence":"first","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"},{"name":"Shanghai Institute for Advanced Communication and Data Science, Shanghai University, Shanghai 200444, China"},{"name":"Materials Genome Institute, Shanghai University, Shanghai 200444, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yifan","family":"Jin","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1983-1632","authenticated-orcid":false,"given":"Jianjia","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"},{"name":"Shanghai Institute for Advanced Communication and Data Science, Shanghai University, Shanghai 200444, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Quan","family":"Qian","sequence":"additional","affiliation":[{"name":"School of Computer Engineering and Science, Shanghai University, Shanghai 200444, China"},{"name":"Shanghai Institute for Advanced Communication and Data Science, Shanghai University, Shanghai 200444, China"},{"name":"Materials Genome Institute, Shanghai University, Shanghai 200444, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yike","family":"Guo","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Hong Kong Baptist University, Hong Kong 999077, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,5,11]]},"reference":[{"key":"ref_1","unstructured":"Li, J., Wang, X., and Li, Y. (2019, January 12\u201317). The speechtransformer for large-scale mandarin chinese speech recognition. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Sainath, T.N., Mohamed, A.R., Kingsbury, B., and Ramabhadran, B. (2013, January 26\u201331). Deep convolutional neural networks for LVCSR. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6639347"},{"key":"ref_3","unstructured":"Amodei, D., Ananthanarayanan, S., Anubhai, R., Bai, J., Battenberg, E., Case, C., Casper, J., Catanzaro, B., Cheng, Q., and Chen, G. (2016, January 19\u201324). Deep speech 2: End-to-end speech recognition in english and mandarin. Proceedings of the International Conference on Machine Learning, New York, NY, USA."},{"key":"ref_4","first-page":"38","article-title":"Distilling the Knowledge in a Neural Network","volume":"14","author":"Hinton","year":"2015","journal-title":"Comput. Sci."},{"key":"ref_5","unstructured":"Romero, A., Ballas, N., Kahou, S.E., Chassang, A., Gatta, C., and Bengio, Y. (2015, January 7\u20139). FitNets: Hints for Thin Deep Nets. Proceedings of the 3rd International Conference on Learning Representations (ICLR), San Diego, CA, USA."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Fukuda, T., Suzuki, M., Kurata, G., Thomas, S., and Ramabhadran, B. (2017, January 20\u201324). Efficient Knowledge Distillation from an Ensemble of Teachers. Proceedings of the Interspeech, Stockholm, Sweden.","DOI":"10.21437\/Interspeech.2017-614"},{"key":"ref_7","unstructured":"Liang, K.J., Hao, W., Shen, D., Zhou, Y., Chen, W., Chen, C., and Carin, L. (2021, January 3\u20137). MixKD: Towards Efficient Distillation of Large-scale Language Models. Proceedings of the 9th International Conference on Learning Representations (ICLR), Virtual Event, Austria."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Park, D.S., Chan, W., Zhang, Y., Chiu, C.C., Zoph, B., Cubuk, E.D., and Le, Q.V. (2019, January 15\u201319). SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition. Proceedings of the Interspeech, Graz, Austria.","DOI":"10.21437\/Interspeech.2019-2680"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Meng, L., Xu, J., Tan, X., Wang, J., Qin, T., and Xu, B. (2021, January 6\u201311). MixSpeech: Data Augmentation for Low-resource Automatic Speech Recognition. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9414483"},{"key":"ref_10","unstructured":"Zhang, H., Ciss\u00e9, M., Dauphin, Y.N., and Lopez-Paz, D. (May, January 30). mixup: Beyond Empirical Risk Minimization. Proceedings of the 6th International Conference on Learning Representations (ICLR), Vancouver, BC, Canada."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Bu, H., Du, J., Na, X., Wu, B., and Zheng, H. (2017, January 1\u20133). Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline. Proceedings of the 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I\/O Systems and Assessment (O-COCOSDA), Seoul, Korea.","DOI":"10.1109\/ICSDA.2017.8384449"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"142","DOI":"10.1016\/j.ins.2020.05.066","article-title":"Adaptive stock trading strategies with deep reinforcement learning methods","volume":"538","author":"Wu","year":"2020","journal-title":"Inf. Sci."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1016\/j.ins.2019.08.059","article-title":"The assessment of small bowel motility with attentive deformable neural network","volume":"508","author":"Wu","year":"2020","journal-title":"Inf. Sci."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","article-title":"Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups","volume":"29","author":"Hinton","year":"2012","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Bi, M., Qian, Y., and Yu, K. (2015, January 6\u201310). Very deep convolutional neural networks for LVCSR. Proceedings of the Sixteenth Annual Conference of the International Speech Communication Association, Dresden, Germany.","DOI":"10.21437\/Interspeech.2015-656"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Graves, A., Mohamed, A.R., and Hinton, G. (2013, January 26\u201331). Speech recognition with deep recurrent neural networks. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6638947"},{"key":"ref_17","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, U., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Yim, J., Joo, D., Bae, J., and Kim, J. (2017, January 21\u201326). A Gift from Knowledge Distillation: Fast Optimization, Network Minimization and Transfer Learning. Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR), IEEE Computer Society, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.754"},{"key":"ref_19","unstructured":"Ba, J., and Caruana, R. (2014, January 8\u201313). Do Deep Nets Really Need to be Deep?. Proceedings of the Advances in Neural Information Processing Systems 27: Annual Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_20","unstructured":"Komodakis, N., and Zagoruyko, S. (2017, January 24\u201326). Paying more attention to attention: Improving the performance of convolutional neural networks via attention transfer. Proceedings of the ICLR, Palais des Congres Neptune, Toulon, France."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Li, J., Zhao, R., Huang, J.T., and Gong, Y. (2014, January 14\u201318). Learning Small-Size DNN with Output-Distribution-Based Criteria. Proceedings of the Interspeech, Singapore.","DOI":"10.21437\/Interspeech.2014-432"},{"key":"ref_22","unstructured":"Geras, K.J., Mohamed, A.R., Caruana, R., Urban, G., Wang, S., Aslan, O., Philipose, M., Richardson, M., and Sutton, C. (2015). Blending LSTMs into CNNs. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Kurata, G., and Audhkhasi, K. (2018, January 18\u201321). Improved Knowledge Distillation from Bi-Directional to Uni-Directional LSTM CTC for End-to-End Speech Recognition. Proceedings of the IEEE Spoken Language Technology Workshop (SLT), Athens, Greece.","DOI":"10.1109\/SLT.2018.8639629"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Graves, A., Fernndez, S., Gomez, F., and Schmidhuber, J. (2006, January 25\u201329). Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. Proceedings of the 23rd International Conference on Machine Learning, Pittsburgh, PA, USA.","DOI":"10.1145\/1143844.1143891"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Takashima, R., Li, S., and Kawai, H. (2018, January 15\u201320). An Investigation of a Knowledge Distillation Method for CTC Acoustic Models. Proceedings of the International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461995"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wong, J., and Gales, M. (2016, January 8\u201312). Sequence Student-Teacher Training of Deep Neural Networks. Proceedings of the Interspeech, San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-911"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"1725","DOI":"10.1109\/TASLP.2019.2929859","article-title":"General Sequence Teacher\u2013Student Learning","volume":"27","author":"Wong","year":"2019","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Kim, H., Na, H., Lee, H., Lee, J., Kang, T.G., Lee, M., and Choi, Y.S. (2019, January 12\u201317). Knowledge Distillation Using Output Errors for Self-attention End-to-end Models. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8682775"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Rosenberg, A., Zhang, Y., Ramabhadran, B., Jia, Y., and Wu, Z. (2019, January 14\u201318). Speech Recognition with Augmented Synthesized Speech. Proceedings of the IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), Singapore.","DOI":"10.1109\/ASRU46091.2019.9003990"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Kanda, N., Takeda, R., and Obuchi, Y. (2013, January 8\u201312). Elastic spectral distortion for low resource speech recognition with deep neural networks. Proceedings of the IEEE Workshop on Automatic Speech Recognition and Understanding, Olomouc, Czech Republic.","DOI":"10.1109\/ASRU.2013.6707748"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Ko, T., Peddinti, V., Povey, D., and Khudanpur, S. (2015, January 6\u201310). Audio augmentation for speech recognition. Proceedings of the Sixteenth Annual Conference of the International Speech Communication Association, Dresden, Germany.","DOI":"10.21437\/Interspeech.2015-711"},{"key":"ref_32","unstructured":"Jaitly, N., and Hinton, G.E. (2013, January 16). Vocal tract length perturbation (VTLP) improves speech recognition. Proceedings of the ICML Workshop on Deep Learning for Audio, Speech and Language, Atlanta, GA, USA."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Ragni, A., Knill, K.M., Rath, S.P., and Gales, M.J.F. (2014, January 14\u201318). Data augmentation for low resource languages. Proceedings of the INTERSPEECH: 15th Annual Conference of the International Speech Communication Association, Singapore.","DOI":"10.21437\/Interspeech.2014-207"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Cubuk, E.D., Zoph, B., Mane, D., Vasudevan, V., and Le, Q.V. (2019, January 16\u201320). AutoAugment: Learning Augmentation Strategies From Data. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00020"},{"key":"ref_35","unstructured":"Zhang, X., Wang, Q., Zhang, J., and Zhong, Z. (2020, January 26\u201330). Adversarial autoaugment. Proceedings of the International Conference on Learning Representations, Addis Ababa, Ethiopia."},{"key":"ref_36","first-page":"6665","article-title":"Fast autoaugment","volume":"32","author":"Lim","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Cubuk, E.D., Zoph, B., Shlens, J., and Le, Q.V. (2020, January 14\u201319). Randaugment: Practical automated data augmentation with a reduced search space. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00359"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"15191","DOI":"10.1109\/ACCESS.2021.3050758","article-title":"Local Augment: Utilizing Local Bias Property of Convolutional Neural Networks for Data Augmentation","volume":"9","author":"Kim","year":"2021","journal-title":"IEEE Access"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"35313","DOI":"10.1109\/ACCESS.2021.3062187","article-title":"An Efficient Data Augmentation Network for Out-of-Distribution Image Detection","volume":"9","author":"Lin","year":"2021","journal-title":"IEEE Access"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Park, D.S., Zhang, Y., Chiu, C.C., Chen, Y., Li, B., Chan, W., Le, Q.V., and Wu, Y. (2020, January 4\u20138). Specaugment on large scale datasets. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9053205"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Nguyen, T.S., Stueker, S., Niehues, J., and Waibel, A. (2020, January 4\u20138). Improving sequence-to-sequence speech recognition training with on-the-fly data augmentation. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Barcelona, Spain.","DOI":"10.1109\/ICASSP40776.2020.9054130"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Hu, T.Y., Shrivastava, A., Chang, J.H.R., Koppula, H., Braun, S., Hwang, K., Kalinli, O., and Tuzel, O. (2021, January 6\u201311). Sapaugment: Learning a sample adaptive policy for data augmentation. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9413928"},{"key":"ref_43","unstructured":"Povey, D., Ghoshal, A., Boulianne, G., Burget, L., Glembek, O., Goel, N., Hannemann, M., Motlicek, P., Qian, Y., and Schwarz, P. (2011, January 11\u201315). The Kaldi speech recognition toolkit. Proceedings of the IEEE Workshop on Automatic Speech Recognition and Understanding, IEEE Signal Processing Society, Big Island, HI, USA."},{"key":"ref_44","unstructured":"Kingma, D.P., and Ba, J. (2015, January 7\u20139). Adam: A method for stochastic optimization. Proceedings of the International Conference for Learning Representations, San Diego, CA, USA."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Takashima, R., Li, S., and Kawai, H. (2019, January 12\u201317). Investigation of Sequence-level Knowledge Distillation Methods for CTC Acoustic Models. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK.","DOI":"10.1109\/ICASSP.2019.8682671"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"1626","DOI":"10.1109\/TASLP.2021.3071662","article-title":"TutorNet: Towards Flexible Knowledge Distillation for End-to-End Speech Recognition","volume":"29","author":"Yoon","year":"2021","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/5\/160\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T23:09:02Z","timestamp":1760137742000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/15\/5\/160"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,11]]},"references-count":46,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2022,5]]}},"alternative-id":["a15050160"],"URL":"https:\/\/doi.org\/10.3390\/a15050160","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,11]]}}}