{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T01:51:52Z","timestamp":1781574712689,"version":"3.54.5"},"reference-count":43,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T00:00:00Z","timestamp":1781481600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Science and Technology Projects of State Grid Jiangsu Electric Power Company Ltd.","award":["J2024168"],"award-info":[{"award-number":["J2024168"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Different mixtures of multimodal training data significantly impact the performance of multimodal large language models, and manually tuning data mixtures is inefficient, computationally expensive, and frequently suboptimal because of complex, nonlinear inter-modal interactions. How to determine data-mixture hyperparameters in an efficient and principled manner becomes the bottleneck for progress in the field. This study establishes a scalable, learnable framework, DMPredictor, that treats multimodal data-mixture design as a regression-based hyperparameter-optimization problem and automates the selection of effective training data mixtures. DMPredictor is trained on data mixture samples derived from hundreds of small proxy models (2M parameters), each of which is trained on 1B tokens sampled using different data mixtures. The framework incorporates alignment-aware smoothing and quality-reweighting, enabling diverse exploration of the multimodal data mixture space while avoiding distribution collapse. DMPredictor produces accurate performance forecasts and identifies nearly optimal data mixtures. The predicted optimal mixture surpasses human-designed baselines on diverse benchmarks, achieving +2.7% on MMMU, +6.4% on TextVQA, and +195.2 on MME. Moreover, the mixture optimization complexity is largely reduced by small proxies and a small number of tokens. The proposed approach offers a robust, computationally efficient pathway for optimizing mixtures of multimodal training data, addressing the critical challenge of training data heterogeneity.<\/jats:p>","DOI":"10.3390\/a19060482","type":"journal-article","created":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T00:49:01Z","timestamp":1781570941000},"page":"482","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Optimize Multimodal Data Mixture for Pre-Training with Loss Regression"],"prefix":"10.3390","volume":"19","author":[{"given":"Linjiang","family":"Shang","sequence":"first","affiliation":[{"name":"Information and Communication Branch, State Grid Jiangsu Electric Power Company Ltd., Nanjing 210024, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huanyu","family":"Cheng","sequence":"additional","affiliation":[{"name":"Information and Communication Branch, State Grid Jiangsu Electric Power Company Ltd., Nanjing 210024, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-5724-0951","authenticated-orcid":false,"given":"Bo","family":"Zhou","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yin","family":"Zhang","sequence":"additional","affiliation":[{"name":"College of Computer Science and Technology, Zhejiang University, Hangzhou 310027, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2026,6,15]]},"reference":[{"key":"ref_1","unstructured":"Liu, H., Li, C., Wu, Q., and Lee, Y.J. (2023, January 10\u201316). Visual instruction tuning. Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_2","unstructured":"Xie, S.M., Pham, H., Dong, X., Du, N., Liu, H., Lu, Y., Liang, P.S., Le, Q.V., Ma, T., and Yu, A.W. (2023, January 10\u201316). Doremi: Optimizing data mixtures speeds up language model pretraining. Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_3","unstructured":"Ye, J., Liu, P., Sun, T., Zhan, J., Zhou, Y., and Qiu, X. (2025, January 24\u201328). Data mixing laws: Optimizing data mixtures by predicting language modeling performance. Proceedings of the 13th International Conference on Learning Representations, Singapore."},{"key":"ref_4","unstructured":"Liu, Q., Zheng, X., Muennighoff, N., Zeng, G., Dou, L., Pang, T., Jiang, J., and Lin, M. (2025, January 24\u201328). RegMix: Data Mixture as Regression for Language Model Pre-training. Proceedings of the 13th International Conference on Learning Representations, Singapore."},{"key":"ref_5","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the 38th International Conference on Machine Learning, Virtual."},{"key":"ref_6","unstructured":"Alayrac, J., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., and Reynolds, M. (December, January 28). Flamingo: A visual language model for few-shot learning. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_7","unstructured":"Li, J., Li, D., Xiong, C., and Hoi, S. (2022, January 17\u201323). Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. Proceedings of the 39th International Conference on Machine Learning, Baltimore, MD, USA."},{"key":"ref_8","unstructured":"Li, J., Li, D., Savarese, S., and Hoi, S. (2023, January 23\u201329). Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. Proceedings of the 40th International Conference on Machine Learning, Honolulu, HI, USA."},{"key":"ref_9","unstructured":"Chen, X., Wang, X., Changpinyo, S., Piergiovanni, A.J., Padlewski, P., Salz, D., Goodman, S., Grycner, A., Mustafa, B., and Beyer, L. (2023, January 1\u20135). PaLI: A Jointly-Scaled Multilingual Language-Image Model. Proceedings of the 11th International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_10","unstructured":"Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C., and Zhou, J. (2023). Qwen-vl: A frontier large vision-language model with versatile abilities. arXiv."},{"key":"ref_11","unstructured":"Bai, S., Cai, Y., Chen, R., Chen, K., Chen, X., Cheng, Z., Deng, L., Ding, W., Gao, C., and Ge, C. (2025). Qwen3-VL Technical Report. arXiv."},{"key":"ref_12","unstructured":"Tong, S., Brown, E., Wu, P., Woo, S., Middepogu, M., Akula, S.C., Yang, J., Yang, S., Iyer, A., and Pan, X. (2024, January 10\u201315). Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs. Proceedings of the 38th International Conference on Neural Information Processing Systems, Vancouver, BC, Canada."},{"key":"ref_13","unstructured":"Li, F., Zhang, R., Zhang, H., Zhang, Y., Li, B., Li, W., Ma, Z., and Li, C. (2024). LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models. arXiv."},{"key":"ref_14","unstructured":"Schuhmann, C., Beaumont, R., Vencu, R., Gordon, C., Wightman, R., Cherti, M., Coombes, T., Katta, A., Mullis, C., and Wortsman, M. (December, January 28). Laion-5b: An open large-scale dataset for training next generation image-text models. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Changpinyo, S., Sharma, P., Ding, N., and Soricut, R. (2021, January 20\u201325). Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00356"},{"key":"ref_16","unstructured":"Gadre, S.Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., and Zhang, J. (2023, January 10\u201316). Datacomp: In search of the next generation of multimodal datasets. Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_17","unstructured":"Lauren\u00e7on, H., Saulnier, L., Tronchon, L., Bekman, S., Singh, A., Lozhkov, A., Wang, T., Karamcheti, S., Rush, A., and Kiela, D. (2023, January 10\u201316). Obelics: An open web-scale filtered dataset of interleaved image-text documents. Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_18","unstructured":"Zhu, W., Hessel, J., Awadalla, A., Gadre, S.Y., Dodge, J., Fang, A., Yu, Y., Schmidt, L., Wang, W.Y., and Choi, Y. (2023, January 10\u201316). Multimodal c4: An open, billion-scale corpus of images interleaved with text. Proceedings of the 37th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_19","unstructured":"Lauren\u00e7on, H., Marafioti, A., Sanh, V., and Tronchon, L. (2024). Building and better understanding vision-language models: Insights and future directions. arXiv."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Bain, M., Nagrani, A., Varol, G., and Zisserman, A. (2021, January 11\u201317). Frozen in time: A joint video and image encoder for end-to-end retrieval. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00175"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Miech, A., Zhukov, D., Alayrac, J.B., Tapaswi, M., Laptev, I., and Sivic, J. (November, January 27). Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00272"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Lin, J., Yin, H., Ping, W., Molchanov, P., Shoeybi, M., and Han, S. (2024, January 17\u201321). Vila: On pre-training for visual language models. Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02520"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Nagrani, A., Seo, P.H., Seybold, B., Hauth, A., Manen, S., Sun, C., and Schmid, C. (2022). Learning audio-video modalities from image captions. Proceedings of the 17th European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-031-19781-9_24"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Chen, D., Huang, Y., Ma, Z., Chen, H., Pan, X., Ge, C., Gao, D., Xie, Y., Liu, Z., and Gao, J. (2024, January 9\u201315). Data-juicer: A one-stop data processing system for large language models. Proceedings of the Companion of the 2024 International Conference on Management of Data, Santiago, Chile.","DOI":"10.1145\/3626246.3653385"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D. (2024). Sharegpt4v: Improving large multi-modal models with better captions. Proceedings of the 18th European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-031-72643-9_22"},{"key":"ref_26","unstructured":"Zhang, Z., Teng, J., Yang, Z., Cao, T., Wang, C., Gu, X., Tang, J., Guo, D., and Wang, M. (2025). Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model. arXiv."},{"key":"ref_27","unstructured":"Agnolucci, L., Galteri, L., and Bertini, M. (2024). Quality-aware image-text alignment for opinion-unaware image quality assessment. arXiv."},{"key":"ref_28","unstructured":"Kaplan, J., McCandlish, S., Henighan, T., Brown, T.B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020). Scaling laws for neural language models. arXiv."},{"key":"ref_29","unstructured":"Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D.d.L., Hendricks, L.A., Welbl, J., and Clark, A. (December, January 28). Training compute-optimal large language models. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_30","unstructured":"Sorscher, B., Geirhos, R., Shekhar, S., Ganguli, S., and Morcos, A. (December, January 28). Beyond neural scaling laws: Beating power law scaling via data pruning. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Sharma, P., Ding, N., Goodman, S., and Soricut, R. (2018, January 15\u201320). Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Melbourne, Australia.","DOI":"10.18653\/v1\/P18-1238"},{"key":"ref_32","unstructured":"Brain, K. (2025, December 01). COYO-700M: A Large-Scale Image-Text Pair Dataset. Available online: https:\/\/github.com\/kakaobrain\/coyo-dataset."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Srinivasan, K., Raman, K., Chen, J., Bendersky, M., and Najork, M. (2021, January 11\u201315). Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Virtual.","DOI":"10.1145\/3404835.3463257"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, X., Wu, J., Chen, J., Li, L., Wang, Y.F., and Wang, W.Y. (November, January 27). VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea.","DOI":"10.1109\/ICCV.2019.00468"},{"key":"ref_35","unstructured":"Lu, P., Mishra, S., Xia, T., Qiu, L., Chang, K.W., Zhu, S.C., Tafjord, O., Clark, P., and Kalyan, A. (December, January 28). Learn to explain: Multimodal reasoning via thought chains for science question answering. Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Xu, Z., Feng, C., Shao, R., Ashby, T., Shen, Y., Jin, D., Cheng, Y., Wang, Q., and Huang, L. (2024, January 16\u201321). Vision-Flan: Scaling Human-labeled Tasks in Visual Instruction Tuning. Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, Mexico City, Mexico.","DOI":"10.18653\/v1\/2024.findings-acl.905"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Li, C., Yuan, Z., Yuan, H., Dong, G., Lu, K., Wu, J., Tan, C., Wang, X., and Zhou, C. (2024, January 11\u201316). Mugglemath: Assessing the impact of query and response augmentation on math reasoning. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.551"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Hessel, J., Hwang, J.D., Park, J.S., Zellers, R., Bhagavatula, C., Rohrbach, A., Saenko, K., and Choi, Y. (2022). The abduction of sherlock holmes: A dataset for visual abductive reasoning. Proceedings of the 17th European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-031-20059-5_32"},{"key":"ref_39","unstructured":"Han, M., Yang, L., Chang, X., Yao, L., and Wang, H. (2025, January 24\u201328). Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos. Proceedings of the 13th International Conference on Learning Representations, Singapore."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Maaz, M., Rasheed, H., Khan, S., and Khan, F.S. (2024, January 11\u201316). Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand.","DOI":"10.18653\/v1\/2024.acl-long.679"},{"key":"ref_41","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. Proceedings of the 9th International Conference on Learning Representations, Virtual."},{"key":"ref_42","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L.u., and Polosukhin, I. (2017, January 4\u20139). Attention is All you Need. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Yue, X., Ni, Y., Zhang, K., Zheng, T., Liu, R., Zhang, G., Stevens, S., Jiang, D., Ren, W., and Sun, Y. (2024, January 17\u201321). MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI. Proceedings of the 2024 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.00913"}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/6\/482\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T00:54:47Z","timestamp":1781571287000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/19\/6\/482"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,15]]},"references-count":43,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2026,6]]}},"alternative-id":["a19060482"],"URL":"https:\/\/doi.org\/10.3390\/a19060482","relation":{},"ISSN":["1999-4893"],"issn-type":[{"value":"1999-4893","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,15]]}}}