{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,8]],"date-time":"2026-07-08T15:11:21Z","timestamp":1783523481373,"version":"3.55.0"},"reference-count":28,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2026,1,8]],"date-time":"2026-01-08T00:00:00Z","timestamp":1767830400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62176122"],"award-info":[{"award-number":["62176122"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012130","name":"Aeronautical Science Foundation of China","doi-asserted-by":"publisher","award":["2023Z073052003"],"award-info":[{"award-number":["2023Z073052003"]}],"id":[{"id":"10.13039\/501100012130","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MAKE"],"abstract":"<jats:p>With the rapid development of applications such as edge computing, the Internet of Things (IoT), and embodied intelligence, massive multimodal data are continuously generated on end devices in a streaming manner. To maintain model adaptability and robustness in dynamic environments, incremental learning has gradually become the core training paradigm on edge devices. However, edge devices are constrained by limited computational, storage, and communication resources, making it infeasible to retain and process all data samples over time. This necessitates efficient data selection strategies to reduce redundancy and improve training efficiency. Existing sample selection methods primarily focus on overall sample difficulty or gradient contribution, but they overlook the heterogeneity of multimodal data in terms of information content and discriminative power. This often leads to modality imbalance, causing the model to over-rely on a single modality and suffer performance degradation. To address this issue, this paper proposes a multimodal sample selection strategy based on the Modality Balance Score (MBS). The method computes confidence scores at the modality level for each sample and further quantifies the contribution differences across modalities. In the selection process, samples with balanced modality contributions are prioritized, thereby improving training efficiency while alleviating modality bias. Experiments conducted on two benchmark datasets, CREMA-D and AVE, demonstrate that compared with existing approaches, the MBS strategy achieves the most stable performance under medium-to-high selection ratios (0.25\u20130.4), yielding superior results in both accuracy and robustness. These findings validate the effectiveness of the proposed strategy in resource-constrained scenarios, providing both theoretical insights and practical guidance for multimodal sample selection in learning tasks.<\/jats:p>","DOI":"10.3390\/make8010017","type":"journal-article","created":{"date-parts":[[2026,1,8]],"date-time":"2026-01-08T15:33:16Z","timestamp":1767886396000},"page":"17","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["MBS: A Modality-Balanced Strategy for Multimodal Sample Selection"],"prefix":"10.3390","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3757-7482","authenticated-orcid":false,"given":"Yuntao","family":"Xu","sequence":"first","affiliation":[{"name":"College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics (NUAA), Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2863-5441","authenticated-orcid":false,"given":"Bing","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics (NUAA), Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feng","family":"Hu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics (NUAA), Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jiawei","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics (NUAA), Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Changjie","family":"Zhao","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics (NUAA), Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongtao","family":"Wu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics (NUAA), Nanjing 211106, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,1,8]]},"reference":[{"key":"ref_1","first-page":"1","article-title":"Lightweight deep learning for resource-constrained environments: A survey","volume":"56","author":"Liu","year":"2024","journal-title":"ACM Comput. Surv."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"42","DOI":"10.1109\/JPROC.2022.3226481","article-title":"Efficient acceleration of deep learning inference on resource-constrained edge devices: A review","volume":"111","author":"Shuvo","year":"2022","journal-title":"Proc. IEEE"},{"key":"ref_3","unstructured":"Landola, F.N., Han, S., Moskewicz, M.W., Ashraf, K., Dally, W.J., and Keutzer, K. (2017). SqueezeNet: AlexNet-level accuracy with 50\u00d7 fewer parameters and <0.5 MB model size. arXiv."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zhou, X., Lin, M., and Sun, J. (2018, January 18\u201323). Shufflenet: An extremely efficient convolutional neural network for mobile devices. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00716"},{"key":"ref_5","unstructured":"Tan, M., and Le, Q. (2019, January 9\u201315). Efficientnet: Rethinking model scaling for convolutional neural networks. Proceedings of the International Conference on Machine Learning, PMLR, Long Beach, CA, USA."},{"key":"ref_6","unstructured":"Mehta, S., and Rastegari, M. (2021, January 3\u20137). MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer. Proceedings of the International Conference on Learning Representations, ICLR, Vienna, Austria."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Chen, Y., Dai, X., Chen, D., Liu, M., Dong, X., Yuan, L., and Liu, Z. (2022, January 18\u201324). Mobile-former: Bridging mobilenet and transformer. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00520"},{"key":"ref_8","unstructured":"Han, S., Mao, H., and Dally, W.J. (2016, January 2\u20134). Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. Proceedings of the 4th International Conference on Learning Representations, ICLR, San Juan, Puerto Rico."},{"key":"ref_9","unstructured":"Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J. (2017, January 24\u201326). Pruning Convolutional Neural Networks for Resource Efficient Inference. Proceedings of the International Conference on Learning Representations, ICLR, Toulon, France."},{"key":"ref_10","unstructured":"Lee, N., Ajanthan, T., and Torr, P. (2019, January 6\u20139). Snip: Single-Shot Network Pruning Based on Connection Sensitivity. Proceedings of the International Conference on Learning Representations, ICLR, New Orleans, LA, USA."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D. (2018, January 18\u201323). Quantization and training of neural networks for efficient integer-arithmetic-only inference. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00286"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1789","DOI":"10.1007\/s11263-021-01453-z","article-title":"Knowledge distillation: A survey","volume":"129","author":"Gou","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1109\/JSSC.2016.2616357","article-title":"Eyeriss: An energy-efficient reconfigurable accelerator for deep convolutional neural networks","volume":"52","author":"Chen","year":"2016","journal-title":"IEEE J. Solid-State Circuits"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"615","DOI":"10.1145\/3093337.3037698","article-title":"Neurosurgeon: Collaborative intelligence between the cloud and mobile edge","volume":"45","author":"Kang","year":"2017","journal-title":"ACM SIGARCH Comput. Archit. News"},{"key":"ref_15","first-page":"3366","article-title":"A continual learning survey: Defying forgetting in classification tasks","volume":"44","author":"Aljundi","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"5362","DOI":"10.1109\/TPAMI.2024.3367329","article-title":"A comprehensive survey of continual learning: Theory, method and application","volume":"46","author":"Wang","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Rebuffi, S.A., Kolesnikov, A., Sperl, G., and Lampert, C. (2017, January 21\u201326). iCaRL: Incremental Classifier and Representation Learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.587"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Wang, K., Herranz, L., and van de Weijer, J. (2021, January 19\u201325). Continual Learning in Cross-Modal Retrieval. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPR), Nashville, TN, USA.","DOI":"10.1109\/CVPRW53098.2021.00402"},{"key":"ref_19","unstructured":"Wang, Z., Chen, Y., and Xu, C. (2023, January 2\u20133). Uncertainty-aware Sample Selection for Multimodal Continual Learning. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Paris, France."},{"key":"ref_20","unstructured":"Paul, M., Ganguli, S., and Dziugaite, G.K. (2021, January 6\u201314). Deep learning on a data diet: Finding important examples early in training. Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), Online."},{"key":"ref_21","unstructured":"Yang, Z., Yang, H., Majumder, S., Cardoso, J., and Gallego, G. (2024). Data Pruning Can Do More: A Comprehensive Data Pruning Approach for Object Re-identification. arXiv."},{"key":"ref_22","unstructured":"Gong, C., Zheng, Z., Wu, F., Shao, Y., Li, B., and Chen, G. (May, January 30). To store or not? Online data selection for federated learning with limited storage. Proceedings of the ACM Web Conference (WWW), Austin, TX, USA."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Peng, X., Wei, Y., Deng, A., Wang, D., and Hu, D. (2022, January 18\u201324). Balanced multimodal learning via on-the-fly gradient modulation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00806"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Fan, Y., Xu, W., Wang, H., Wang, J., and Guo, S. (2023, January 17\u201324). PMR: Prototypical modal rebalance for multimodal learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(CVPR), Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01918"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Mahmoud, A., Elhoushi, M., Abbas, A., Yang, Y., Ardalan, N., Leather, H., and Morcos, A. (2024, January 16\u201322). Sieve: Multimodal dataset pruning using image captioning models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02116"},{"key":"ref_26","unstructured":"Ye, W., Wu, Q., Lin, W., and Zhou, Y. (March, January 25). Fit and prune: Fast and training-free visual token pruning for multi-modal large language models. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"377","DOI":"10.1109\/TAFFC.2014.2336244","article-title":"CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset","volume":"5","author":"Cao","year":"2014","journal-title":"IEEE Trans. Affect. Comput."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Tian, Y., Shi, J., Li, B., Duan, Z., and Xu, C. (2018, January 8\u201314). Audio-Visual Event Localization in Unconstrained Videos. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01216-8_16"}],"container-title":["Machine Learning and Knowledge Extraction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/1\/17\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,10]],"date-time":"2026-01-10T05:19:08Z","timestamp":1768022348000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2504-4990\/8\/1\/17"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,8]]},"references-count":28,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,1]]}},"alternative-id":["make8010017"],"URL":"https:\/\/doi.org\/10.3390\/make8010017","relation":{},"ISSN":["2504-4990"],"issn-type":[{"value":"2504-4990","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,8]]}}}