{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T03:06:10Z","timestamp":1787022370503,"version":"3.56.0"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"5","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2023,1]]},"abstract":"<jats:p>When training a deep learning (DL) model, input data are pre-processed on CPUs and transformed into tensors, which are then fed into GPUs for gradient computations of model training. Expensive GPUs must be fully utilized during training to accelerate the training speed. However, intensive CPU operations for input data preprocessing (input pipeline) often lead to CPU bottlenecks; correspondingly, various DL training jobs suffer from GPU under-utilization.<\/jats:p>\n                  <jats:p>We propose FastFlow, a DL training system that automatically mitigates the CPU bottleneck by offloading (scaling out) input pipelines to remote CPUs. FastFlow carefully decides various offloading decisions based on performance metrics specific to applications and allocated resources, while leveraging both local and remote CPUs to prevent the inefficient use of remote resources and minimize the training time. FastFlow's smart offloading policy and mechanisms are seamlessly integrated with TensorFlow for users to enjoy the smart offloading features without modifying the main logic. Our evaluations on our private DL cloud with diverse workloads on various resource environments show that FastFlow improves the training throughput by 1 ~ 4.34X compared to TensorFlow without offloading, by 1 ~ 4.52X compared to TensorFlow with manual CPU offloading (tf.data.service), and by 0.63 ~ 2.06X compared to GPU offloading (DALI).<\/jats:p>","DOI":"10.14778\/3579075.3579083","type":"journal-article","created":{"date-parts":[[2023,3,6]],"date-time":"2023-03-06T12:10:26Z","timestamp":1678104626000},"page":"1086-1099","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":35,"title":["FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data Pipeline"],"prefix":"10.14778","volume":"16","author":[{"given":"Taegeon","family":"Um","sequence":"first","affiliation":[{"name":"Samsung Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Byungsoo","family":"Oh","sequence":"additional","affiliation":[{"name":"Samsung Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Byeongchan","family":"Seo","sequence":"additional","affiliation":[{"name":"Samsung Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Minhyeok","family":"Kweun","sequence":"additional","affiliation":[{"name":"Samsung Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Goeun","family":"Kim","sequence":"additional","affiliation":[{"name":"Samsung Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Woo-Yeon","family":"Lee","sequence":"additional","affiliation":[{"name":"Samsung Research"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,3,6]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Accessed in January 2023. Add more choices for data augmentation. https:\/\/github.com\/NVIDIA\/DALI\/issues\/1610.  Accessed in January 2023. Add more choices for data augmentation. https:\/\/github.com\/NVIDIA\/DALI\/issues\/1610."},{"key":"e_1_2_1_2_1","unstructured":"Accessed in January 2023. Amazon EC2 P3 Instances. https:\/\/aws.amazon.com\/ec2\/instance-types\/p3\/.  Accessed in January 2023. Amazon EC2 P3 Instances. https:\/\/aws.amazon.com\/ec2\/instance-types\/p3\/."},{"key":"e_1_2_1_3_1","unstructured":"Accessed in January 2023. Amazon S3. https:\/\/aws.amazon.com\/s3.  Accessed in January 2023. Amazon S3. https:\/\/aws.amazon.com\/s3."},{"key":"e_1_2_1_4_1","unstructured":"Accessed in January 2023. Automatic Speech Recognition using CTC. https:\/\/keras.io\/examples\/audio\/ctc_asr\/.  Accessed in January 2023. Automatic Speech Recognition using CTC. https:\/\/keras.io\/examples\/audio\/ctc_asr\/."},{"key":"e_1_2_1_5_1","unstructured":"Accessed in January 2023. Automatic Speech Recognition with Transformer. https:\/\/keras.io\/examples\/audio\/transformer_asr\/.  Accessed in January 2023. Automatic Speech Recognition with Transformer. https:\/\/keras.io\/examples\/audio\/transformer_asr\/."},{"key":"e_1_2_1_6_1","unstructured":"Accessed in January 2023. Data-efficient GANs with Adaptive Discriminator Augmentation. https:\/\/keras.io\/examples\/generative\/gan_ada\/.  Accessed in January 2023. Data-efficient GANs with Adaptive Discriminator Augmentation. https:\/\/keras.io\/examples\/generative\/gan_ada\/."},{"key":"e_1_2_1_7_1","unstructured":"Accessed in January 2023. A distribution strategy for synchronous training on multiple workers. https:\/\/www.tensorflow.org\/api_docs\/python\/tf\/distribute\/experimental\/MultiWorkerMirroredStrategy.  Accessed in January 2023. A distribution strategy for synchronous training on multiple workers. https:\/\/www.tensorflow.org\/api_docs\/python\/tf\/distribute\/experimental\/MultiWorkerMirroredStrategy."},{"key":"e_1_2_1_8_1","unstructured":"Accessed in January 2023. Does DALI support GPU operation equivalent to tf.signal.fft? https:\/\/github.com\/NVIDIA\/DALI\/issues\/4331.  Accessed in January 2023. Does DALI support GPU operation equivalent to tf.signal.fft? https:\/\/github.com\/NVIDIA\/DALI\/issues\/4331."},{"key":"e_1_2_1_9_1","unstructured":"Accessed in January 2023. Google TPU v4. https:\/\/cloud.google.com\/tpu\/docs\/system-architecture-tpu-vm#tpu_v4.  Accessed in January 2023. Google TPU v4. https:\/\/cloud.google.com\/tpu\/docs\/system-architecture-tpu-vm#tpu_v4."},{"key":"e_1_2_1_10_1","unstructured":"Accessed in January 2023. Keras: Deep Learning for humans. https:\/\/github.com\/keras-team\/keras.  Accessed in January 2023. Keras: Deep Learning for humans. https:\/\/github.com\/keras-team\/keras."},{"key":"e_1_2_1_11_1","unstructured":"Accessed in January 2023. Kubernetes: an open source system for managing containerized applications across multiple hosts. https:\/\/github.com\/kubernetes\/kubernetes.  Accessed in January 2023. Kubernetes: an open source system for managing containerized applications across multiple hosts. https:\/\/github.com\/kubernetes\/kubernetes."},{"key":"e_1_2_1_12_1","unstructured":"Accessed in January 2023. Learning to Resize in Computer Vision. https:\/\/keras.io\/examples\/vision\/learnable_resizer\/.  Accessed in January 2023. Learning to Resize in Computer Vision. https:\/\/keras.io\/examples\/vision\/learnable_resizer\/."},{"key":"e_1_2_1_13_1","unstructured":"Accessed in January 2023. MelGAN-based spectrogram inversion using feature matching. https:\/\/keras.io\/examples\/audio\/melgan_spectrogram_inversion\/.  Accessed in January 2023. MelGAN-based spectrogram inversion using feature matching. https:\/\/keras.io\/examples\/audio\/melgan_spectrogram_inversion\/."},{"key":"e_1_2_1_14_1","unstructured":"Accessed in January 2023. Module: tf.data.experimental.service. https:\/\/www.tensorflow.org\/api_docs\/python\/tf\/data\/experimental\/service.  Accessed in January 2023. Module: tf.data.experimental.service. https:\/\/www.tensorflow.org\/api_docs\/python\/tf\/data\/experimental\/service."},{"key":"e_1_2_1_15_1","unstructured":"Accessed in January 2023. NVIDIA Data Loading Library (DALI). https:\/\/developer.nvidia.com\/dali.  Accessed in January 2023. NVIDIA Data Loading Library (DALI). https:\/\/developer.nvidia.com\/dali."},{"key":"e_1_2_1_16_1","unstructured":"Accessed in January 2023. NVIDIA H100 Performance. https:\/\/developer.nvidia.com\/blog\/nvidia-hopper-architecture-in-depth\/.  Accessed in January 2023. NVIDIA H100 Performance. https:\/\/developer.nvidia.com\/blog\/nvidia-hopper-architecture-in-depth\/."},{"key":"e_1_2_1_17_1","unstructured":"Accessed in January 2023. psutil (process and system utilities): a cross-platform library for retrieving information on running processes. https:\/\/github.com\/giampaolo\/psutil.  Accessed in January 2023. psutil (process and system utilities): a cross-platform library for retrieving information on running processes. https:\/\/github.com\/giampaolo\/psutil."},{"key":"e_1_2_1_18_1","unstructured":"Accessed in January 2023. Pytorch Dataloader. https:\/\/pytorch.org\/docs\/stable\/data.html#torch.utils.data.DataLoader.  Accessed in January 2023. Pytorch Dataloader. https:\/\/pytorch.org\/docs\/stable\/data.html#torch.utils.data.DataLoader."},{"key":"e_1_2_1_19_1","unstructured":"Accessed in January 2023. Speaker Recognition. https:\/\/keras.io\/examples\/audio\/speaker_recognition_using_cnn\/.  Accessed in January 2023. Speaker Recognition. https:\/\/keras.io\/examples\/audio\/speaker_recognition_using_cnn\/."},{"key":"e_1_2_1_20_1","unstructured":"Accessed in January 2023. Speaker Recognition Dataset. https:\/\/www.kaggle.com\/datasets\/kongaevans\/speaker-recognition-dataset.  Accessed in January 2023. Speaker Recognition Dataset. https:\/\/www.kaggle.com\/datasets\/kongaevans\/speaker-recognition-dataset."},{"key":"e_1_2_1_21_1","unstructured":"Accessed in January 2023. tf.data.service commit that does not prefer local reads. https:\/\/github.com\/tensorflow\/tensorflow\/commit\/17e7f5e01bbcdde893309bc302f144e25d43bb81.  Accessed in January 2023. tf.data.service commit that does not prefer local reads. https:\/\/github.com\/tensorflow\/tensorflow\/commit\/17e7f5e01bbcdde893309bc302f144e25d43bb81."},{"key":"e_1_2_1_22_1","unstructured":"Martin Abadi Paul Barham Jianmin Chen Zhifeng Chen Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Geoffrey Irving Michael Isard Manjunath Kudlur Josh Levenberg Rajat Monga Sherry Moore Derek G. Murray Benoit Steiner Paul Tucker Vijay Vasudevan Pete Warden Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2016. TensorFlow: A system for large-scale machine learning. In OSDI. 265--283.  Martin Abadi Paul Barham Jianmin Chen Zhifeng Chen Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Geoffrey Irving Michael Isard Manjunath Kudlur Josh Levenberg Rajat Monga Sherry Moore Derek G. Murray Benoit Steiner Paul Tucker Vijay Vasudevan Pete Warden Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2016. TensorFlow: A system for large-scale machine learning. In OSDI. 265--283."},{"key":"e_1_2_1_23_1","volume-title":"Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach.","author":"Black Sid","year":"2022","unstructured":"Sid Black , Stella Biderman , Eric Hallahan , Quentin Anthony , Leo Gao , Laurence Golding , Horace He , Connor Leahy , Kyle McDonell , Jason Phang , Michael Pieler , USVSN Sai Prashanth , Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022 . GPT-NeoX-20B: An Open-Source Autoregressive Language Model . Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022. GPT-NeoX-20B: An Open-Source Autoregressive Language Model."},{"key":"e_1_2_1_24_1","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell etal 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020) 1877--1901.  Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared D Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. Advances in neural information processing systems 33 (2020) 1877--1901."},{"key":"e_1_2_1_25_1","volume-title":"Apache flink: Stream and batch processing in a single engine. Bulletin of the IEEE Computer Society Technical Committee on Data Engineering 36, 4","author":"Carbone Paris","year":"2015","unstructured":"Paris Carbone , Asterios Katsifodimos , Stephan Ewen , Volker Markl , Seif Haridi , and Kostas Tzoumas . 2015. Apache flink: Stream and batch processing in a single engine. Bulletin of the IEEE Computer Society Technical Committee on Data Engineering 36, 4 ( 2015 ). Paris Carbone, Asterios Katsifodimos, Stephan Ewen, Volker Markl, Seif Haridi, and Kostas Tzoumas. 2015. Apache flink: Stream and batch processing in a single engine. Bulletin of the IEEE Computer Society Technical Committee on Data Engineering 36, 4 (2015)."},{"key":"e_1_2_1_26_1","unstructured":"Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Josh Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating Large Language Models Trained on Code. https:\/\/arxiv.org\/abs\/2107.03374  Mark Chen Jerry Tworek Heewoo Jun Qiming Yuan Henrique Ponde de Oliveira Pinto Jared Kaplan Harri Edwards Yuri Burda Nicholas Joseph Greg Brockman Alex Ray Raul Puri Gretchen Krueger Michael Petrov Heidy Khlaaf Girish Sastry Pamela Mishkin Brooke Chan Scott Gray Nick Ryder Mikhail Pavlov Alethea Power Lukasz Kaiser Mohammad Bavarian Clemens Winter Philippe Tillet Felipe Petroski Such Dave Cummings Matthias Plappert Fotios Chantzis Elizabeth Barnes Ariel Herbert-Voss William Hebgen Guss Alex Nichol Alex Paino Nikolas Tezak Jie Tang Igor Babuschkin Suchir Balaji Shantanu Jain William Saunders Christopher Hesse Andrew N. Carr Jan Leike Josh Achiam Vedant Misra Evan Morikawa Alec Radford Matthew Knight Miles Brundage Mira Murati Katie Mayer Peter Welinder Bob McGrew Dario Amodei Sam McCandlish Ilya Sutskever and Wojciech Zaremba. 2021. Evaluating Large Language Models Trained on Code. https:\/\/arxiv.org\/abs\/2107.03374"},{"key":"e_1_2_1_27_1","first-page":"1802","article-title":"Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and Scheduling","volume":"32","author":"Cheng Yang","year":"2021","unstructured":"Yang Cheng , Dan Li , Zhiyuan Guo , Binyao Jiang , Jinkun Geng , Wei Bai , Jianping Wu , and Yongqiang Xiong . 2021 . Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and Scheduling . IEEE TPDS 32 , 7 (2021), 1802 -- 1814 . Yang Cheng, Dan Li, Zhiyuan Guo, Binyao Jiang, Jinkun Geng, Wei Bai, Jianping Wu, and Yongqiang Xiong. 2021. Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and Scheduling. IEEE TPDS 32, 7 (2021), 1802--1814.","journal-title":"IEEE TPDS"},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Joon Son Chung Arsha Nagrani and Andrew Zisserman. 2018. VoxCeleb2: Deep Speaker Recognition. In Interspeech. 1086--1090.  Joon Son Chung Arsha Nagrani and Andrew Zisserman. 2018. VoxCeleb2: Deep Speaker Recognition. In Interspeech. 1086--1090.","DOI":"10.21437\/Interspeech.2018-1929"},{"key":"e_1_2_1_29_1","volume-title":"Le","author":"Cubuk Ekin D.","year":"2019","unstructured":"Ekin D. Cubuk , Barret Zoph , Dandelion Mane , Vijay Vasudevan , and Quoc V . Le . 2019 . AutoAugment: Learning Augmentation Strategies From Data. In CVPR. Ekin D. Cubuk, Barret Zoph, Dandelion Mane, Vijay Vasudevan, and Quoc V. Le. 2019. AutoAugment: Learning Augmentation Strategies From Data. In CVPR."},{"key":"e_1_2_1_30_1","volume-title":"Le","author":"Cubuk Ekin D.","year":"2020","unstructured":"Ekin D. Cubuk , Barret Zoph , Jonathon Shlens , and Quoc V . Le . 2020 . Randaugment : Practical Automated Data Augmentation With a Reduced Search Space. In CVPR. Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V. Le. 2020. Randaugment: Practical Automated Data Augmentation With a Reduced Search Space. In CVPR."},{"key":"e_1_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Jia Deng Wei Dong Richard Socher Li-Jia Li Kai Li and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In CVPR.  Jia Deng Wei Dong Richard Socher Li-Jia Li Kai Li and Li Fei-Fei. 2009. ImageNet: A large-scale hierarchical image database. In CVPR.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_32_1","volume-title":"Asirra: A CAPTCHA that Exploits Interest-Aligned Manual Image Categorization. In ACM CCS.","author":"Elson Jeremy","year":"2007","unstructured":"Jeremy Elson , John (JD) Douceur , Jon Howell , and Jared Saul . 2007 . Asirra: A CAPTCHA that Exploits Interest-Aligned Manual Image Categorization. In ACM CCS. Jeremy Elson, John (JD) Douceur, Jon Howell, and Jared Saul. 2007. Asirra: A CAPTCHA that Exploits Interest-Aligned Manual Image Categorization. In ACM CCS."},{"key":"e_1_2_1_33_1","volume-title":"2022 USENIX Annual Technical Conference (USENIX ATC 22)","author":"Graur Dan","year":"2022","unstructured":"Dan Graur , Damien Aymon , Dan Kluser , Tanguy Albrici , Chandramohan A. Thekkath , and Ana Klimovic . 2022 . Cachew: Machine Learning Input Data Processing as a Service . In 2022 USENIX Annual Technical Conference (USENIX ATC 22) . 689--706. Dan Graur, Damien Aymon, Dan Kluser, Tanguy Albrici, Chandramohan A. Thekkath, and Ana Klimovic. 2022. Cachew: Machine Learning Input Data Processing as a Service. In 2022 USENIX Annual Technical Conference (USENIX ATC 22). 689--706."},{"key":"e_1_2_1_34_1","volume-title":"IEEE international conference on acoustics, speech and signal processing. 6645--6649","author":"Graves Alex","year":"2013","unstructured":"Alex Graves , Abdel-rahman Mohamed, and Geoffrey Hinton . 2013 . Speech recognition with deep recurrent neural networks . In IEEE international conference on acoustics, speech and signal processing. 6645--6649 . Alex Graves, Abdel-rahman Mohamed, and Geoffrey Hinton. 2013. Speech recognition with deep recurrent neural networks. In IEEE international conference on acoustics, speech and signal processing. 6645--6649."},{"key":"e_1_2_1_35_1","doi-asserted-by":"crossref","unstructured":"Yosuke Higuchi Shinji Watanabe Nanxin Chen Tetsuji Ogawa and Tetsunori Kobayashi. 2020. Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict. In Interspeech. 3655--3659.  Yosuke Higuchi Shinji Watanabe Nanxin Chen Tetsuji Ogawa and Tetsunori Kobayashi. 2020. Mask CTC: Non-Autoregressive End-to-End ASR with CTC and Mask Predict. In Interspeech. 3655--3659.","DOI":"10.21437\/Interspeech.2020-2404"},{"key":"e_1_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Alexander Isenko Ruben Mayer Jeffrey Jedele and Hans-Arno Jacobsen. 2022. Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines. In SIGMOD.  Alexander Isenko Ruben Mayer Jeffrey Jedele and Hans-Arno Jacobsen. 2022. Where Is My Training Bottleneck? Hidden Trade-Offs in Deep Learning Preprocessing Pipelines. In SIGMOD.","DOI":"10.1145\/3514221.3517848"},{"key":"e_1_2_1_37_1","unstructured":"Keith Ito and Linda Johnson. 2017. The LJ Speech Dataset. https:\/\/keithito.com\/LJ-Speech-Dataset\/.  Keith Ito and Linda Johnson. 2017. The LJ Speech Dataset. https:\/\/keithito.com\/LJ-Speech-Dataset\/."},{"key":"e_1_2_1_38_1","unstructured":"Tero Karras Miika Aittala Janne Hellsten Samuli Laine Jaakko Lehtinen and Timo Aila. 2020. Training Generative Adversarial Networks with Limited Data. In NeurIPS. 12104--12114.  Tero Karras Miika Aittala Janne Hellsten Samuli Laine Jaakko Lehtinen and Timo Aila. 2020. Training Generative Adversarial Networks with Limited Data. In NeurIPS. 12104--12114."},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476249.3476308"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of Machine Learning and Systems 4","author":"Kuchnik Michael","year":"2022","unstructured":"Michael Kuchnik , Ana Klimovic , Jiri Simsa , Virginia Smith , and George Amvrosiadis . 2022 . Plumber: Diagnosing and Removing Performance Bottlenecks in Machine Learning Data Pipelines . Proceedings of Machine Learning and Systems 4 (2022). Michael Kuchnik, Ana Klimovic, Jiri Simsa, Virginia Smith, and George Amvrosiadis. 2022. Plumber: Diagnosing and Removing Performance Bottlenecks in Machine Learning Data Pipelines. Proceedings of Machine Learning and Systems 4 (2022)."},{"key":"e_1_2_1_41_1","volume-title":"Quiver: An Informed Storage Cache for Deep Learning. In FAST. 283--296.","author":"Kumar Abhishek Vijaya","year":"2020","unstructured":"Abhishek Vijaya Kumar and Muthian Sivathanu . 2020 . Quiver: An Informed Storage Cache for Deep Learning. In FAST. 283--296. Abhishek Vijaya Kumar and Muthian Sivathanu. 2020. Quiver: An Informed Storage Cache for Deep Learning. In FAST. 283--296."},{"key":"e_1_2_1_42_1","volume-title":"Jose Sotelo, Alexandre de Br\u00e9bisson, Yoshua Bengio, and Aaron C Courville.","author":"Kumar Kundan","year":"2019","unstructured":"Kundan Kumar , Rithesh Kumar , Thibault de Boissiere , Lucas Gestin , Wei Zhen Teoh , Jose Sotelo, Alexandre de Br\u00e9bisson, Yoshua Bengio, and Aaron C Courville. 2019 . MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis. In NeurIPS. Kundan Kumar, Rithesh Kumar, Thibault de Boissiere, Lucas Gestin, Wei Zhen Teoh, Jose Sotelo, Alexandre de Br\u00e9bisson, Yoshua Bengio, and Aaron C Courville. 2019. MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis. In NeurIPS."},{"key":"e_1_2_1_43_1","unstructured":"Gyewon Lee Irene Lee Hyeonmin Ha Kyunggeun Lee Hwarim Hyun Ahnjae Shin and Byung-Gon Chun. 2021. Refurbish Your Training Data: Reusing Partially Augmented Samples for Faster Deep Neural Network Training. In ATC. 537--550.  Gyewon Lee Irene Lee Hyeonmin Ha Kyunggeun Lee Hwarim Hyun Ahnjae Shin and Byung-Gon Chun. 2021. Refurbish Your Training Data: Reusing Partially Augmented Samples for Faster Deep Neural Network Training. In ATC. 537--550."},{"key":"e_1_2_1_44_1","volume-title":"Haeju Lee, Joonyoung Kim, Kangwook Lee, and Kee-Eung Kim.","author":"Lee Youngjune","year":"2021","unstructured":"Youngjune Lee , Oh Joon Kwon , Haeju Lee, Joonyoung Kim, Kangwook Lee, and Kee-Eung Kim. 2021 . Augment & Valuate: A Data Enhancement Pipeline for Data-Centric AI. arXiv preprint arXiv:2112.03837 (2021). Youngjune Lee, Oh Joon Kwon, Haeju Lee, Joonyoung Kim, Kangwook Lee, and Kee-Eung Kim. 2021. Augment & Valuate: A Data Enhancement Pipeline for Data-Centric AI. arXiv preprint arXiv:2112.03837 (2021)."},{"key":"e_1_2_1_45_1","volume-title":"Transformers with convolutional context for asr. arXiv preprint arXiv:1904.11660","author":"Mohamed Abdelrahman","year":"2019","unstructured":"Abdelrahman Mohamed , Dmytro Okhonko , and Luke Zettlemoyer . 2019. Transformers with convolutional context for asr. arXiv preprint arXiv:1904.11660 ( 2019 ). Abdelrahman Mohamed, Dmytro Okhonko, and Luke Zettlemoyer. 2019. Transformers with convolutional context for asr. arXiv preprint arXiv:1904.11660 (2019)."},{"key":"e_1_2_1_46_1","unstructured":"Jayashree Mohan Amar Phanishayee Janardhan Kulkarni and Vijay Chidambaram. 2022. Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters. In OSDI. 579--596.  Jayashree Mohan Amar Phanishayee Janardhan Kulkarni and Vijay Chidambaram. 2022. Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters. In OSDI. 579--596."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.14778\/3446095.3446100"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.14778\/3476311.3476374"},{"key":"e_1_2_1_49_1","doi-asserted-by":"crossref","unstructured":"Pyeongsu Park Heetaek Jeong and Jangwoo Kim. 2020. TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Systematically Balancing Operations. In MICRO. 825--838.  Pyeongsu Park Heetaek Jeong and Jangwoo Kim. 2020. TrainBox: An Extreme-Scale Neural Network Training Server Architecture by Systematically Balancing Operations. In MICRO. 825--838.","DOI":"10.1109\/MICRO50266.2020.00072"},{"key":"e_1_2_1_50_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas Kopf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In NeurIPS. Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In NeurIPS."},{"key":"e_1_2_1_51_1","volume-title":"Pyroomacoustics: A Python Package for Audio Room Simulation and Array Processing Algorithms","author":"Scheibler Robin","year":"2018","unstructured":"Robin Scheibler , Eric Bezzam , and Ivan Dokmani\u0107 . 2018 . Pyroomacoustics: A Python Package for Audio Room Simulation and Array Processing Algorithms . In IEEE ICASSP. 351--355. Robin Scheibler, Eric Bezzam, and Ivan Dokmani\u0107. 2018. Pyroomacoustics: A Python Package for Audio Room Simulation and Array Processing Algorithms. In IEEE ICASSP. 351--355."},{"key":"e_1_2_1_52_1","doi-asserted-by":"crossref","unstructured":"Hossein Talebi and Peyman Milanfar. 2021. Learning to resize images for computer vision tasks. In ICCV. 497--506.  Hossein Talebi and Peyman Milanfar. 2021. Learning to resize images for computer vision tasks. In ICCV. 497--506.","DOI":"10.1109\/ICCV48922.2021.00055"},{"key":"e_1_2_1_53_1","unstructured":"Catherine Wah Steve Branson Peter Welinder Pietro Perona and Serge Belongie. 2011. The caltech-ucsd birds-200-2011 dataset. (2011).  Catherine Wah Steve Branson Peter Welinder Pietro Perona and Serge Belongie. 2011. The caltech-ucsd birds-200-2011 dataset. (2011)."},{"key":"e_1_2_1_54_1","first-page":"1173","article-title":"DIESEL+: Accelerating Distributed Deep Learning Tasks on Image Datasets","volume":"33","author":"Wang Lipeng","year":"2022","unstructured":"Lipeng Wang , Qiong Luo , and Shengen Yan . 2022 . DIESEL+: Accelerating Distributed Deep Learning Tasks on Image Datasets . IEEE TPDS 33 (2022), 1173 -- 1184 . Lipeng Wang, Qiong Luo, and Shengen Yan. 2022. DIESEL+: Accelerating Distributed Deep Learning Tasks on Image Datasets. IEEE TPDS 33 (2022), 1173--1184.","journal-title":"IEEE TPDS"},{"key":"e_1_2_1_55_1","unstructured":"Qizhen Weng Wencong Xiao Yinghao Yu Wei Wang Cheng Wang Jian He Yong Li Liping Zhang Wei Lin and Yu Ding. 2022. MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In NSDI. 945--960.  Qizhen Weng Wencong Xiao Yinghao Yu Wei Wang Cheng Wang Jian He Yong Li Liping Zhang Wei Lin and Yu Ding. 2022. MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In NSDI. 945--960."},{"key":"e_1_2_1_56_1","unstructured":"Matei Zaharia Mosharaf Chowdhury Tathagata Das Ankur Dave Justin Ma Murphy McCauly Michael J Franklin Scott Shenker and Ion Stoica. 2012. Resilient Distributed Datasets: A {Fault-Tolerant} Abstraction for {In-Memory} Cluster Computing. In NSDI. 15--28.  Matei Zaharia Mosharaf Chowdhury Tathagata Das Ankur Dave Justin Ma Murphy McCauly Michael J Franklin Scott Shenker and Ion Stoica. 2012. Resilient Distributed Datasets: A {Fault-Tolerant} Abstraction for {In-Memory} Cluster Computing. In NSDI. 15--28."},{"key":"e_1_2_1_57_1","doi-asserted-by":"crossref","unstructured":"Mark Zhao Niket Agarwal Aarti Basant Bu\u011fra Gedik Satadru Pan Mustafa Ozdal Rakesh Komuravelli Jerry Pan Tianshu Bao Haowei Lu Sundaram Narayanan Jack Langman Kevin Wilfong Harsha Rastogi Carole-Jean Wu Christos Kozyrakis and Parik Pol. 2022. Understanding Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training: Industrial Product. In ISCA. 1042--1057.  Mark Zhao Niket Agarwal Aarti Basant Bu\u011fra Gedik Satadru Pan Mustafa Ozdal Rakesh Komuravelli Jerry Pan Tianshu Bao Haowei Lu Sundaram Narayanan Jack Langman Kevin Wilfong Harsha Rastogi Carole-Jean Wu Christos Kozyrakis and Parik Pol. 2022. Understanding Data Storage and Ingestion for Large-Scale Deep Recommendation Model Training: Industrial Product. In ISCA. 1042--1057.","DOI":"10.1145\/3470496.3533044"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3579075.3579083","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,6]],"date-time":"2023-03-06T12:13:16Z","timestamp":1678104796000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3579075.3579083"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1]]},"references-count":57,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2023,1]]}},"alternative-id":["10.14778\/3579075.3579083"],"URL":"https:\/\/doi.org\/10.14778\/3579075.3579083","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2023,1]]},"assertion":[{"value":"2023-03-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}