{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T17:10:16Z","timestamp":1780765816078,"version":"3.54.1"},"reference-count":48,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2021,4,15]],"date-time":"2021-04-15T00:00:00Z","timestamp":1618444800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61571192, 61771200"],"award-info":[{"award-number":["61571192, 61771200"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Knowledge Distillation (KD), which transfers the knowledge from a teacher to a student network by penalizing their Kullback\u2013Leibler (KL) divergence, is a widely used tool for Deep Neural Network (DNN) compression in intelligent sensor systems. Traditional KD uses pre-trained teacher, while self-KD distills its own knowledge to achieve better performance. The role of the teacher in self-KD is usually played by multi-branch peers or the identical sample with different augmentation. However, the mentioned self-KD methods above have their limitation for widespread use. The former needs to redesign the DNN for different tasks, and the latter relies on the effectiveness of the augmentation method. To avoid the limitation above, we propose a new self-KD method, Memory-replay Knowledge Distillation (MrKD), that uses the historical models as teachers. Firstly, we propose a novel self-KD training method that penalizes the KD loss between the current model\u2019s output distributions and its backup outputs on the training trajectory. This strategy can regularize the model with its historical output distribution space to stabilize the learning. Secondly, a simple Fully Connected Network (FCN) is applied to ensemble the historical teacher\u2019s output for a better guidance. Finally, to ensure the teacher outputs offer the right class as ground truth, we correct the teacher logit output by the Knowledge Adjustment (KA) method. Experiments on the image (dataset CIFAR-100, CIFAR-10, and CINIC-10) and audio (dataset DCASE) classification tasks show that MrKD improves single model training and working efficiently across different datasets. In contrast to the existing fancy self-KD methods with various external knowledge, the effectiveness of MrKD sheds light on the usually abandoned historical models during the training trajectory.<\/jats:p>","DOI":"10.3390\/s21082792","type":"journal-article","created":{"date-parts":[[2021,4,15]],"date-time":"2021-04-15T21:35:13Z","timestamp":1618522513000},"page":"2792","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["Memory-Replay Knowledge Distillation"],"prefix":"10.3390","volume":"21","author":[{"given":"Jiyue","family":"Wang","sequence":"first","affiliation":[{"name":"School of Electronic and Information Engineering, South China University of Technology, Guangzhou 510641, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6372-5653","authenticated-orcid":false,"given":"Pei","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Northwestern Polytechnical University, Xi\u2019an 710072, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanxiong","family":"Li","sequence":"additional","affiliation":[{"name":"School of Electronic and Information Engineering, South China University of Technology, Guangzhou 510641, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2021,4,15]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_2","unstructured":"Huang, G., Liu, Z., Pleiss, G., Van Der Maaten, L., and Weinberger, K. (2019). Convolutional Networks with Dense Connectivity. IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","unstructured":"Chen, Y., Li, J., Xiao, H., Jin, X., Yan, S., and Feng, J. (2017). Dual path networks. Adv. Neural Inf. Process. Syst., 4467\u20134475."},{"key":"ref_4","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.C. (2018, January 18\u201323). Mobilenetv2: Inverted residuals and linear bottlenecks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Howard, A., Sandler, M., Chu, G., Chen, L.C., Chen, B., Tan, M., Wang, W., Zhu, Y., Pang, R., and Vasudevan, V. (2019, January 27\u201328). Searching for MobileNetV3. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00140"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Ma, N., Zhang, X., Zheng, H.T., and Sun, J. (2018, January 8\u201314). Shufflenet v2: Practical guidelines for efficient cnn architecture design. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"ref_8","unstructured":"Liu, H., Simonyan, K., and Yang, Y. (2018). DARTS:Differentiable Architecture Search. arXiv."},{"key":"ref_9","unstructured":"Chaudhuri, K., and Salakhutdinov, R. (2019, January 9\u201315). EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. Proceedings of the 36th International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_10","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the Knowledge in a Neural Network. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Cho, J., and Lee, M. (2019). Building a Compact Convolutional Neural Network for Embedded Intelligent Sensor Systems Using Group Sparsity and Knowledge Distillation. Sensors, 19.","DOI":"10.3390\/s19194307"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Park, S., and Heo, Y.S. (2020). Knowledge Distillation for Semantic Segmentation Using Channel and Spatial Correlations and Adaptive Cross Entropy. Sensors, 20.","DOI":"10.3390\/s20164616"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Choi, E., Chae, S., and Kim, J. (2019). Machine Learning-Based Fast Banknote Serial Number Recognition Using Knowledge Distillation and Bayesian Optimization. Sensors, 19.","DOI":"10.3390\/s19194218"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Chechlinski, L., Siemi\u0105tkowska, B., and Majewski, M. (2019). A System for Weeds and Crops Identification\u2014Reaching over 10 FPS on Raspberry Pi with the Usage of MobileNets, DenseNet and Custom Modifications. Sensors, 19.","DOI":"10.20944\/preprints201907.0115.v1"},{"key":"ref_15","unstructured":"Furlanello, T., Lipton, Z.C., Tschannen, M., Itti, L., and Anandkumar, A. (2018, January 10\u201315). Born Again Neural Networks. Proceedings of the International Conference on Machine Learning, Stockholm Sweden."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Xiang, T., Hospedales, T.M., and Lu, H. (2018, January 18\u201323). Deep Mutual Learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00454"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Gao, L., Lan, X., Mi, H., Feng, D., Xu, K., and Peng, Y. (2019). Multistructure-Based Collaborative Online Distillation. Entropy, 21.","DOI":"10.3390\/e21040357"},{"key":"ref_18","unstructured":"Zhang, L., Song, J., Gao, A., Chen, J., Bao, C., and Ma, K. (November, January 27). Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self Distillation. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Yun, S., Park, J., Lee, K., and Shin, J. (2020, January 14\u201319). Regularizing Class-Wise Predictions via Self-Knowledge Distillation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01389"},{"key":"ref_20","first-page":"5565","article-title":"Data-Distortion Guided Self-Distillation for Deep Neural Networks","volume":"33","author":"Xu","year":"2019","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_21","unstructured":"Luan, Y., Zhao, H., Yang, Z., and Dai, Y. (2019). MSD: Multi-Self-Distillation Learning via Multi-classifiers within Deep Neural Networks. arXiv."},{"key":"ref_22","unstructured":"Hendrycks, D., Mu, N., Cubuk, E.D., Zoph, B., Gilmer, J., and Lakshminarayanan, B. (2019). AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_24","unstructured":"Mandt, S., Hoffman, M.D., and Blei, D.M. (2017). Stochastic Gradient Descent as Approximate Bayesian Inference. arXiv."},{"key":"ref_25","unstructured":"Wen, T., Lai, S., and Qian, X. (2019). Preparing Lessons: Improve Knowledge Distillation with Better Supervision. arXiv."},{"key":"ref_26","unstructured":"Krizhevsky, A., and Hinton, G. (2009). Learning Multiple Layers of Features from Tiny Images, University of Toronto. Technical report."},{"key":"ref_27","unstructured":"Darlow, L.N., Crowley, E.J., Antoniou, A., and Storkey, A.J. (2018). CINIC-10 is not ImageNet or CIFAR-10. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Zagoruyko, S., and Komodakis, N. (2016). Wide Residual Networks. arXiv.","DOI":"10.5244\/C.30.87"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Xie, S., Girshick, R.B., Doll\u00e1r, P., Tu, Z., and He, K. (2017, January 21\u201326). Aggregated Residual Transformations for Deep Neural Networks. Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.634"},{"key":"ref_30","unstructured":"Mesaros, A., Heittola, T., and Virtanen, T. (2018, January 19\u201320). A multi-device dataset for urban acoustic scene classification. Proceedings of the Detection and Classification of Acoustic Scenes and Events 2018 Workshop (DCASE2018), Surrey, UK."},{"key":"ref_31","unstructured":"Heittola, T., Mesaros, A., and Virtanen, T. (2020, January 2\u20133). Acoustic scene classification in DCASE 2020 Challenge: Generalization across devices and low complexity solutions. Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020), Tokyo, Japan."},{"key":"ref_32","unstructured":"Song, G., and Chai, W. (2018). Collaborative learning for deep neural networks. arXiv."},{"key":"ref_33","unstructured":"Lan, X., Zhu, X., and Gong, S. (2018). Knowledge distillation by on-the-fly native ensemble. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Cho, J.H., and Hariharan, B. (2019). On the Efficacy of Knowledge Distillation. arXiv.","DOI":"10.1109\/ICCV.2019.00489"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Mirzadeh, S.I., Farajtabar, M., Li, A., Levine, N., Matsukawa, A., and Ghasemzadeh, H. (2019). Improved Knowledge Distillation via Teacher Assistant. arXiv.","DOI":"10.1609\/aaai.v34i04.5963"},{"key":"ref_36","unstructured":"Jin, X., Peng, B., Wu, Y., Liu, Y., Liu, J., Liang, D., Yan, J., and Hu, X. (November, January 27). Knowledge Distillation via Route Constrained Optimization. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea."},{"key":"ref_37","unstructured":"Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D., and Wilson, A.G. (2018). Averaging Weights Leads to Wider Optima and Better Generalization. arXiv."},{"key":"ref_38","unstructured":"Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., and Garnett, R. (2017). Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results. Advances in Neural Information Processing Systems 30, Curran Associates, Inc."},{"key":"ref_39","unstructured":"Xu, Y., Xu, Y., Qian, Q., Li, H., and Jin, R. (2020). Towards Understanding Label Smoothing. arXiv."},{"key":"ref_40","unstructured":"Kim, K., Ji, B., Yoon, D.Y., and Hwang, S. (2020). Self-Knowledge Distillation: A Simple Way for Better Generalization. arXiv."},{"key":"ref_41","first-page":"3430","article-title":"Online Knowledge Distillation with Diverse Peers","volume":"34","author":"Chen","year":"2020","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_42","unstructured":"Wu, G., and Gong, S. (2020). Peer Collaborative Learning for Online Knowledge Distillation. arXiv."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Koutini, K., Eghbal-zadeh, H., Dorfer, M., and Widmer, G. (2019, January 2\u20136). The Receptive Field as a Regularizer in Deep Convolutional Neural Networks for Acoustic Scene Classification. Proceedings of the European Signal Processing Conference (EUSIPCO), A Coruna, Spain.","DOI":"10.23919\/EUSIPCO.2019.8902732"},{"key":"ref_44","unstructured":"Koutini, K., Henkel, F., Eghbal-Zadeh, H., and Widmer, G. (2020, January 2\u20133). Low-Complexity Models for Acoustic Scene Classification Based on Receptive Field Regularization and Frequency Damping. Proceedings of the Detection and Classification of Acoustic Scenes and Events 2020 Workshop (DCASE2020), Tokyo, Japan."},{"key":"ref_45","unstructured":"Zhang, H., Cisse, M., Dauphin, Y.N., and Lopez-Paz, D. (2018). mixup: Beyond Empirical Risk Minimization. arXiv."},{"key":"ref_46","unstructured":"Romero, A., Ballas, N., Ebrahimi Kahou, S., Chassang, A., Gatta, C., and Bengio, Y. (2014). FitNets: Hints for Thin Deep Nets. arXiv."},{"key":"ref_47","unstructured":"Zagoruyko, S., and Komodakis, N. (2017). Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer. arXiv."},{"key":"ref_48","first-page":"7350","article-title":"Knowledge Distillation from Internal Representations","volume":"34","author":"Aguilar","year":"2020","journal-title":"Proc. AAAI Conf. Artif. Intell."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/8\/2792\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T05:48:28Z","timestamp":1760161708000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/8\/2792"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,15]]},"references-count":48,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2021,4]]}},"alternative-id":["s21082792"],"URL":"https:\/\/doi.org\/10.3390\/s21082792","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,4,15]]}}}