{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T21:10:57Z","timestamp":1775596257855,"version":"3.50.1"},"reference-count":83,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,4,2]]},"abstract":"<jats:p>\n                    State-of-the-art large language and vision models are trained over trillions of tokens that are aggregated from a large variety of sources. As training data collections grow, manually managing the samples becomes time-consuming, tedious, and prone to errors. Yet recent research shows that the data mixture and the order in which samples are visited during training can significantly influence model accuracy. We build and present\n                    <jats:sc>Mixtera<\/jats:sc>\n                    , a data plane for foundation model training that enables users to declaratively express which data samples should be used in which proportion and in which order during training.\n                    <jats:sc>Mixtera<\/jats:sc>\n                    is a centralized, read-only layer that is deployed on top of existing training data collections and can be declaratively queried. It operates independently of the filesystem structure and supports mixtures across arbitrary properties (e.g., language, source dataset) as well as dynamic adjustment of the mixture based on model feedback. We experimentally evaluate\n                    <jats:sc>Mixtera<\/jats:sc>\n                    and show that our implementation does not bottleneck training and scales to 256 GH200 superchips. We demonstrate how\n                    <jats:sc>Mixtera<\/jats:sc>\n                    supports recent advancements in mixing strategies by implementing the Adaptive Data Optimization (ADO) algorithm in the system and evaluating its performance impact. We also show how\n                    <jats:sc>Mixtera<\/jats:sc>\n                    enables exploring the role of mixtures for vision-language models, which is a growing area of research.\n                  <\/jats:p>","DOI":"10.1145\/3786668","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T17:54:13Z","timestamp":1775584453000},"page":"1-28","source":"Crossref","is-referenced-by-count":0,"title":["Mixtera: A Data Plane for Foundation Model Training"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4093-4361","authenticated-orcid":false,"given":"Maximilian","family":"B\u00f6ther","sequence":"first","affiliation":[{"name":"ETH Zurich, Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4661-533X","authenticated-orcid":false,"given":"Xiaozhe","family":"Yao","sequence":"additional","affiliation":[{"name":"ETH Zurich, Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1175-338X","authenticated-orcid":false,"given":"Tolga","family":"Kerimoglu","sequence":"additional","affiliation":[{"name":"ETH Zurich, Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-0682-2422","authenticated-orcid":false,"given":"Dan","family":"Graur","sequence":"additional","affiliation":[{"name":"ETH Zurich, Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6750-5500","authenticated-orcid":false,"given":"Viktor","family":"Gsteiger","sequence":"additional","affiliation":[{"name":"ETH Zurich, Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8559-0529","authenticated-orcid":false,"given":"Ana","family":"Klimovic","sequence":"additional","affiliation":[{"name":"ETH Zurich, Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,7]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.14778\/2824032.2824076"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.52202\/075280-0722"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2312.02406"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","unstructured":"Loubna Ben Allal Anton Lozhkov Elie Bakouch Gabriel Mart\u00edn Bl\u00e1zquez Guilherme Penedo Lewis Tunstall Andr\u00e9s Marafioti Hynek Kydl\u00ed?ek Agust\u00edn Piqueres Lajar\u00edn Vaibhav Srivastav Joshua Lochner Caleb Fahlgren Xuan-Son Nguyen Cl\u00e9mentine Fourrier Ben Burtenshaw Hugo Larcher Haojun Zhao Cyril Zakka Mathieu Morlon Colin Raffel Leandro von Werra and Thomas Wolf. 2025. SmolLM2: When Smol Goes Big - Data-Centric Training of a Small Language Model. arXiv preprint (2025). doi:10.48550\/arXiv.2502.02737","DOI":"10.48550\/arXiv.2502.02737"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the Conference on Machine Learning and Systems (MLSys).","author":"Barham Paul","year":"2022","unstructured":"Paul Barham, Aakanksha Chowdhery, Jeff Dean, Sanjay Ghemawat, Steven Hand, Dan Hurt, Michael Isard, Hyeontaek Lim, Ruoming Pang, Sudip Roy, Brennan Saeta, Parker Schuh, Ryan Sepassi, Laurent El Shafey, Chandramohan A. Thekkath, and Yonghui Wu. 2022. Pathways: Asynchronous Distributed Dataflow for ML. In Proceedings of the Conference on Machine Learning and Systems (MLSys)."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3320060"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i05.6239"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2504.10950"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","unstructured":"Rishi Bommasani Drew A. Hudson Ehsan Adeli Russ Altman Simran Arora Sydney von Arx Michael S. Bernstein Jeannette Bohg Antoine Bosselut Emma Brunskill Erik Brynjolfsson Shyamal Buch Dallas Card Rodrigo Castellon Niladri Chatterji Annie Chen Kathleen Creel Jared Quincy Davis Dora Demszky Chris Donahue Moussa Doumbouya Esin Durmus Stefano Ermon John Etchemendy Kawin Ethayarajh Li Fei-Fei Chelsea Finn Trevor Gale Lauren Gillespie Karan Goel Noah Goodman Shelby Grossman Neel Guha Tatsunori Hashimoto Peter Henderson John Hewitt Daniel E. Ho Jenny Hong Kyle Hsu Jing Huang Thomas Icard Saahil Jain Dan Jurafsky Pratyusha Kalluri Siddharth Karamcheti Geoff Keeling Fereshte Khani Omar Khattab Pang Wei Koh Mark Krass Ranjay Krishna Rohith Kuditipudi Ananya Kumar Faisal Ladhak Mina Lee Tony Lee Jure Leskovec Isabelle Levent Xiang Lisa Li Xuechen Li Tengyu Ma Ali Malik Christopher D. Manning Suvir Mirchandani Eric Mitchell Zanele Munyikwa Suraj Nair Avanika Narayan Deepak Narayanan Ben Newman Allen Nie Juan Carlos Niebles Hamed Nilforoshan Julian Nyarko Giray Ogut Laurel Orr Isabel Papadimitriou Joon Sung Park Chris Piech Eva Portelance Christopher Potts Aditi Raghunathan Rob Reich Hongyu Ren Frieda Rong Yusuf Roohani Camilo Ruiz Jack Ryan Christopher R\u00e9 Dorsa Sadigh Shiori Sagawa Keshav Santhanam Andy Shih Krishnan Srinivasan Alex Tamkin Rohan Taori Armin W. Thomas Florian Tram\u00e8r Rose E. Wang William Wang Bohan Wu Jiajun Wu Yuhuai Wu Sang Michael Xie Michihiro Yasunaga Jiaxuan You Matei Zaharia Michael Zhang Tianyi Zhang Xikun Zhang Yuhui Zhang Lucia Zheng Kaitlyn Zhou and Percy Liang. 2022. On the Opportunities and Risks of Foundation Models. In arXiv preprint. doi:10.48550\/arXiv.2108.07258","DOI":"10.48550\/arXiv.2108.07258"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3709705"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems (NeurIPS).","author":"Brown Tom B.","year":"2020","unstructured":"Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020. Language Models are Few-Shot Learners. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3626246.3653385"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2411.05735"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2307.14430"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3722212.3724454"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1803.05457"},{"key":"e_1_2_1_17_1","unstructured":"Competition and Markets Authority. 2013. AI Foundation Models: Initial Report. Technical Report. UK Government Agency."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3511265.3550446"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2306.13394"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2101.00027"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.10256836"},{"key":"e_1_2_1_22_1","volume-title":"Proceedings of the USENIX Annual Technical Conference (ATC).","author":"Graur Dan","year":"2022","unstructured":"Dan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici, Chandramohan A Thekkath, and Ana Klimovic. 2022. Cachew: Machine Learning Input Data Processing as a Service. In Proceedings of the USENIX Annual Technical Conference (ATC)."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3623490"},{"key":"e_1_2_1_24_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems (NeurIPS).","author":"Huang Yanping","year":"2019","unstructured":"Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Xu Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen. 2019. GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr.2019.00686"},{"key":"e_1_2_1_26_1","volume-title":"Nanotron: Pretraining models made easy. https:\/\/github.com\/huggingface\/nanotron","year":"2025","unstructured":"HuggingFace. 2025. Nanotron: Pretraining models made easy. https:\/\/github.com\/huggingface\/nanotron"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2405.11788"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2410.11820"},{"key":"e_1_2_1_29_1","unstructured":"Siddharth Karamcheti Laurel Orr Jason Bolton Tianyi Zhang Karan Goel Avanika Narayan Rishi Bommasani Deepak Narayanan Tatsunori Hashimoto Dan Jurafsky Christopher D. Manning Christopher Potts Christopher R\u00e9 and Percy Liang. 2021. Mistral - A Journey towards Reproducible Language Model Training."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0981-7"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the Conference on Machine Learning and Systems (MLSys).","author":"Kuchnik Michael","year":"2022","unstructured":"Michael Kuchnik, Ana Klimovic, Jiri Simsa, Virginia Smith, and George Amvrosiadis. 2022. Plumber: Diagnosing and Removing Performance Bottlenecks in Machine Learning Data Pipelines. In Proceedings of the Conference on Machine Learning and Systems (MLSys)."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.emnlp-main.20"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2502.06244"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2410.06511"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr52733.2024.02484"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.naacl-long.179"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the Conference on Neural Information Processing Systems (NeurIPS).","author":"Lu Pan","year":"2022","unstructured":"Pan Lu, Swaroop Mishra, Tony Xia, Liang Qiu, Kai-Wei Chang, Song-Chun Zhu, Oyvind Tafjord, Peter Clark, and Ashwin Kalyan. 2022. Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2407.21783"},{"key":"e_1_2_1_40_1","unstructured":"Meta. 2024b. Llama 3.3 Model Card. https:\/\/github.com\/meta-llama\/llama-models\/blob\/main\/models\/llama3_3\/MODEL_CARD.md. Accessed: 2024-12-18."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/d18-1260"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/icdar.2019.00156"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.14778\/3446095.3446100"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476374"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476209"},{"key":"e_1_2_1_47_1","unstructured":"OpenAI. 2024. GPT-4 Technical Report. In arXiv preprint. doi:10.48550\/arXiv.2303.08774"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/p16-1144"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2406.17557"},{"key":"e_1_2_1_50_1","unstructured":"Guilherme Penedo Hynek Kydl\u00ed?ek Alessandro Cappelli Mario Sasko and Thomas Wolf. 2024b. DataTrove: large scale data processing. https:\/\/github.com\/huggingface\/datatrove"},{"key":"e_1_2_1_51_1","volume-title":"Proceedings of the Conference on Neural Information Processing Systems (NeurIPS).","author":"Qian Shangshu","year":"2021","unstructured":"Shangshu Qian, Hung Viet Pham, Thibaud Lutellier, Zeou Hu, Jungwon Kim, Lin Tan, Yaoliang Yu, Jiahao Chen, and Sameena Shah. 2021. Are My Deep Learning Systems Fair? An Empirical Study of Fixed-Seed Training. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3299869.3320212"},{"key":"e_1_2_1_53_1","unstructured":"Alec Radford Jeff Wu Rewon Child David Luan Dario Amodei and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. In Self-hosted preprint."},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2405.20512"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3749185"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3474381"},{"key":"e_1_2_1_57_1","volume-title":"Hechtman","author":"Shazeer Noam","year":"2018","unstructured":"Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, and Blake A. Hechtman. 2018. Mesh-TensorFlow: Deep Learning for Supercomputers. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_1_58_1","volume-title":"SlimPajama-DC: Understanding Data Combinations for LLM Training. arXiv preprint","author":"Shen Zhiqiang","year":"2024","unstructured":"Zhiqiang Shen, Tianhua Tao, Liqun Ma, Willie Neiswanger, Zhengzhong Liu, Hongyi Wang, Bowen Tan, Joel Hestness, Natalia Vassilieva, Daria Soboleva, and Eric Xing. 2024. SlimPajama-DC: Understanding Data Combinations for LLM Training. arXiv preprint (2024). 10.48550\/arXiv.2309.10818"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.1909.08053"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","unstructured":"Amanpreet Singh Vivek Natarajan Meet Shah Yu Jiang Xinlei Chen Dhruv Batra Devi Parikh and Marcus Rohrbach. 2019. Towards VQA Models That Can Read. In Proceedings of the Conference on Computer Vision and Pattern Recognition (CVPR). doi:10.1109\/cvpr.2019.00851","DOI":"10.1109\/cvpr.2019.00851"},{"key":"e_1_2_1_61_1","unstructured":"Daria Soboleva Faisal Al-Khateeb Robert Myers Jacob R Steeves Joel Hestness and Nolan Dey. 2023. SlimPajama: A 627B token cleaned and deduplicated version of RedPajama. https:\/\/huggingface.co\/datasets\/cerebras\/SlimPajama-627B"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","unstructured":"Luca Soldaini Rodney Kinney Akshita Bhagia Dustin Schwenk David Atkinson Russell Authur Ben Bogin Khyathi Chandu Jennifer Dumas Yanai Elazar Valentin Hofmann Ananya Harsh Jha Sachin Kumar Li Lucy Xinxi Lyu Nathan Lambert Ian Magnusson Jacob Morrison Niklas Muennighoff Aakanksha Naik Crystal Nam Matthew E. Peters Abhilasha Ravichander Kyle Richardson Zejiang Shen Emma Strubell Nishant Subramani Oyvind Tafjord Pete Walsh Luke Zettlemoyer Noah A. Smith Hannaneh Hajishirzi Iz Beltagy Dirk Groeneveld Jesse Dodge and Kyle Lo. 2024. Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. arXiv preprint (2024). doi:10.48550\/arXiv.2402.00159","DOI":"10.48550\/arXiv.2402.00159"},{"key":"e_1_2_1_63_1","unstructured":"The LlaVA Team. 2023. LLaVA Data Documentation. https:\/\/github.com\/haotian-liu\/LLaVA\/blob\/main\/docs\/Data.md"},{"key":"e_1_2_1_64_1","unstructured":"The LlaVA Team. 2024a. TinyLLaVA: Model Zoo. https:\/\/github.com\/TinyLLaVA\/TinyLLaVA_Factory?tab=readme-ov-file#model-zoo"},{"key":"e_1_2_1_65_1","unstructured":"The Mosaic ML Team. 2022. streaming: Fast accurate streaming of training data from cloud storage. https:\/\/github.com\/mosaicml\/streaming\/"},{"key":"e_1_2_1_66_1","volume-title":"Tensorflow: Determinism. https:\/\/www.tensorflow.org\/api_docs\/python\/tf\/config\/experimental\/enable_op_determinism","author":"TensorFlow Team The","year":"2025","unstructured":"The TensorFlow Team. 2025. Tensorflow: Determinism. https:\/\/www.tensorflow.org\/api_docs\/python\/tf\/config\/experimental\/enable_op_determinism"},{"key":"e_1_2_1_67_1","unstructured":"The TinyLLaVA Factory Team. 2024b. TinyLLaVA Factory: Prepare Datasets. https:\/\/tinyllava-factory.readthedocs.io\/en\/latest\/Prepare%20Datasets.html"},{"key":"e_1_2_1_68_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems (NeurIPS).","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2312.01700"},{"key":"e_1_2_1_70_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems (NeurIPS).","author":"Weber Maurice","year":"2024","unstructured":"Maurice Weber, Daniel Y Fu, Quentin Gregory Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Re, Irina Rish, and Ce Zhang. 2024. RedPajama: an Open Dataset for Training Large Language Models. In Proceedings of Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_2_1_71_1","volume-title":"Advances in Neural Information Processing Systems","volume":"36","author":"Xie Sang Michael","year":"2024","unstructured":"Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy S Liang, Quoc V Le, Tengyu Ma, and Adams Wei Yu. 2024. Doremi: Optimizing data mixtures speeds up language model pretraining. Advances in Neural Information Processing Systems, Vol. 36 (2024)."},{"key":"e_1_2_1_72_1","volume-title":"Shweti Mahajan, Julian McAuley, Jennifer Neville, Ahmed Hassan Awadallah, and Nikhil Rao.","author":"Xu Canwen","year":"2024","unstructured":"Canwen Xu, Corby Rosset, Ethan C. Chau, Luciano Del Corro, Shweti Mahajan, Julian McAuley, Jennifer Neville, Ahmed Hassan Awadallah, and Nikhil Rao. 2024b. Automatic Pair Construction for Contrastive Post-training. arXiv preprint (2024). 10.48550\/arXiv.2310.02263"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2309.1167"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/3731569.3764847"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2403.16952"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/cvpr52733.2024.00913"},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934664"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/p19-1472"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2401.02385"},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2504.09844"},{"key":"e_1_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3533044"},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.14778\/3611540.3611569"},{"key":"e_1_2_1_83_1","volume-title":"Proceedings of the Conference on Machine Learning and Systems (MLSys).","author":"Zhuang Donglin","year":"2022","unstructured":"Donglin Zhuang, Xingyao Zhang, Shuaiwen Song, and Sara Hooker. 2022. Randomness in Neural Network Training: Characterizing the Impact of Tooling. In Proceedings of the Conference on Machine Learning and Systems (MLSys)."}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786668","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T20:03:27Z","timestamp":1775592207000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786668"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,2]]},"references-count":83,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4,2]]}},"alternative-id":["10.1145\/3786668"],"URL":"https:\/\/doi.org\/10.1145\/3786668","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,2]]}}}