{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T15:45:17Z","timestamp":1782834317617,"version":"3.54.5"},"reference-count":68,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2022,7]]},"abstract":"<jats:p>Deep neural networks (DNNs) have grown exponentially in size over the past decade, leaving only those who have massive datacenter-based resources with the ability to develop and train such models. One of the main challenges for the long tail of researchers who might have only limited resources (e.g., a single multi-GPU server) is limited GPU memory capacity compared to model size. The problem is so acute that the memory requirement of training massive DNN models can often exceed the aggregate capacity of all available GPUs on a single server; this problem only gets worse with the trend of ever-growing model sizes. Current solutions that rely on virtualizing GPU memory (by swapping to\/from CPU memory) incur excessive swapping overhead. In this paper, we present a new training framework, Harmony, and advocate rethinking how DNN frameworks schedule computation and move data to push the boundaries of training massive models efficiently on a single commodity server. Across various massive DNN models, Harmony is able to reduce swap load by up to two orders of magnitude and obtain a training throughput speedup of up to 7.6x over highly optimized baselines with virtualized memory.<\/jats:p>","DOI":"10.14778\/3551793.3551828","type":"journal-article","created":{"date-parts":[[2022,9,29]],"date-time":"2022-09-29T22:25:03Z","timestamp":1664490303000},"page":"2747-2760","source":"Crossref","is-referenced-by-count":20,"title":["Harmony"],"prefix":"10.14778","volume":"15","author":[{"given":"Youjie","family":"Li","sequence":"first","affiliation":[{"name":"UIUC"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Amar","family":"Phanishayee","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Derek","family":"Murray","sequence":"additional","affiliation":[{"name":"Lacework"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jakub","family":"Tarnawski","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nam Sung","family":"Kim","sequence":"additional","affiliation":[{"name":"UIUC"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,9,29]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Deep learning for computational biology. Molecular systems biology 12, 7","author":"Angermueller Christof","year":"2016","unstructured":"Christof Angermueller , Tanel P\u00e4rnamaa , Leopold Parts , and Oliver Stegle . 2016. Deep learning for computational biology. Molecular systems biology 12, 7 ( 2016 ), 878. Christof Angermueller, Tanel P\u00e4rnamaa, Leopold Parts, and Oliver Stegle. 2016. Deep learning for computational biology. Molecular systems biology 12, 7 (2016), 878."},{"key":"e_1_2_1_2_1","unstructured":"ASUS. 2019. High-density 4U GPU server https:\/\/www.asus.com\/us\/Commercial-Servers-Workstations\/ESC8000-G4.  ASUS. 2019. High-density 4U GPU server https:\/\/www.asus.com\/us\/Commercial-Servers-Workstations\/ESC8000-G4."},{"key":"e_1_2_1_3_1","volume-title":"GPT-2 fine-tuning with ONNX Runtime. Microsoft Open Source Blog","author":"Bhandare Aishwarya","year":"2020","unstructured":"Aishwarya Bhandare , Tianju Xu , and Kshama Pawar . 2020. GPT-2 fine-tuning with ONNX Runtime. Microsoft Open Source Blog ( 2020 ). https:\/\/cloudblogs.microsoft.com\/opensource\/2020\/08\/24\/pytorch-gpt-2-fine-tuning-onnx-runtime-speedup-training-time Aishwarya Bhandare, Tianju Xu, and Kshama Pawar. 2020. GPT-2 fine-tuning with ONNX Runtime. Microsoft Open Source Blog (2020). https:\/\/cloudblogs.microsoft.com\/opensource\/2020\/08\/24\/pytorch-gpt-2-fine-tuning-onnx-runtime-speedup-training-time"},{"key":"e_1_2_1_4_1","unstructured":"Tom B Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell etal 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020).  Tom B Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. arXiv preprint arXiv:2005.14165 (2020)."},{"key":"e_1_2_1_5_1","volume-title":"Training deep nets with sublinear memory cost. arXiv preprint arXiv.1604.06174","author":"Chen Tianqi","year":"2016","unstructured":"Tianqi Chen , Bing Xu , Chiyuan Zhang , and Carlos Guestrin . 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv.1604.06174 ( 2016 ). Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv preprint arXiv.1604.06174 (2016)."},{"key":"e_1_2_1_6_1","unstructured":"Minsik Cho Tung D Le U Finkler Haruiki Imai Yasushi Negishi Taro Sekiyama Saritha Vinod Vladimir Zolotov Kiyokuni Kawachiya David S Kung etal 2018. Large model support for deep learning in caffe and chainer. SysML'18 (Feb. 2018).  Minsik Cho Tung D Le U Finkler Haruiki Imai Yasushi Negishi Taro Sekiyama Saritha Vinod Vladimir Zolotov Kiyokuni Kawachiya David S Kung et al. 2018. Large model support for deep learning in caffe and chainer. SysML'18 (Feb. 2018)."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_8_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv 1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv 1810.04805 (2018)."},{"key":"e_1_2_1_9_1","volume-title":"Annual Technical Conference (ATC'21)","author":"Eliad Saar","year":"2021","unstructured":"Saar Eliad , Ido Hakimi , Alon De Jagger , Mark Silberstein , and Assaf Schuster . 2021 . Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism . In Annual Technical Conference (ATC'21) . 381--396. Saar Eliad, Ido Hakimi, Alon De Jagger, Mark Silberstein, and Assaf Schuster. 2021. Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism. In Annual Technical Conference (ATC'21). 381--396."},{"key":"e_1_2_1_10_1","unstructured":"Facebook. 2020. Distributed Data Parallel in PyTorch https:\/\/pytorch.org\/docs\/master\/notes\/ddp.html.  Facebook. 2020. Distributed Data Parallel in PyTorch https:\/\/pytorch.org\/docs\/master\/notes\/ddp.html."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3437801.3441593"},{"key":"e_1_2_1_12_1","volume-title":"AI and Memory Wall. RiseLab Medium Post","author":"Gholami Amir","year":"2021","unstructured":"Amir Gholami , Zhewei Yao , Sehoon Kim , Michael W Mahoney , and Kurt Keutzer . 2021. AI and Memory Wall. RiseLab Medium Post ( 2021 ). Amir Gholami, Zhewei Yao, Sehoon Kim, Michael W Mahoney, and Kurt Keutzer. 2021. AI and Memory Wall. RiseLab Medium Post (2021)."},{"key":"e_1_2_1_13_1","unstructured":"Google. 2018. TensorFlow code and pre-trained models for BERT https:\/\/github.com\/google-research\/bert.  Google. 2018. TensorFlow code and pre-trained models for BERT https:\/\/github.com\/google-research\/bert."},{"key":"e_1_2_1_14_1","volume-title":"large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677","author":"Goyal Priya","year":"2017","unstructured":"Priya Goyal , Piotr Doll\u00e1r , Ross Girshick , Pieter Noordhuis , Lukasz Wesolowski , Aapo Kyrola , Andrew Tulloch , Yangqing Jia , and Kaiming He. 2017. Accurate , large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677 ( 2017 ). Priya Goyal, Piotr Doll\u00e1r, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677 (2017)."},{"key":"e_1_2_1_15_1","volume-title":"A robot wrote this entire article. Are you scared yet, human? The Guardian","author":"Guardian The","year":"2020","unstructured":"The Guardian . 2020. A robot wrote this entire article. Are you scared yet, human? The Guardian ( 2020 ). The Guardian. 2020. A robot wrote this entire article. Are you scared yet, human? The Guardian (2020)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378465"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS'17)","author":"Hoffer Elad","year":"2017","unstructured":"Elad Hoffer , Itay Hubara , and Daniel Soudry . 2017 . Train Longer, Generalize Better: Closing the Generalization Gap in Large Batch Training of Neural Networks . In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS'17) (Long Beach, California, USA). 1729--1739. Elad Hoffer, Itay Hubara, and Daniel Soudry. 2017. Train Longer, Generalize Better: Closing the Generalization Gap in Large Batch Training of Neural Networks. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NeurIPS'17) (Long Beach, California, USA). 1729--1739."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378530"},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the 33st International Conference on Neural Information Processing Systems (NeurIPS'19)","author":"Huang Yanping","year":"2019","unstructured":"Yanping Huang , Youlong Cheng , Ankur Bapna , Orhan Firat , Mia Xu Chen , Dehao Chen , HyoukJoong Lee , Jiquan Ngiam , Quoc V Le , Yonghui Wu , 2019 . GPipe: Efficient training of giant neural networks using pipeline parallelism . In Proceedings of the 33st International Conference on Neural Information Processing Systems (NeurIPS'19) . Vancouver, Canada, 103--112. Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al. 2019. GPipe: Efficient training of giant neural networks using pipeline parallelism. In Proceedings of the 33st International Conference on Neural Information Processing Systems (NeurIPS'19). Vancouver, Canada, 103--112."},{"key":"e_1_2_1_21_1","unstructured":"Huggingface. 2018. PyTorch Pretrained Bert https:\/\/github.com\/maknotavailable\/pytorch-pretrained-BERT.  Huggingface. 2018. PyTorch Pretrained Bert https:\/\/github.com\/maknotavailable\/pytorch-pretrained-BERT."},{"key":"e_1_2_1_22_1","unstructured":"Huggingface. 2021. Transformer Examples https:\/\/huggingface.co\/transformers\/v2.3.0\/examples.html.  Huggingface. 2021. Transformer Examples https:\/\/huggingface.co\/transformers\/v2.3.0\/examples.html."},{"key":"e_1_2_1_23_1","unstructured":"IBM. 2018. TensorFlow Large-Model-Support https:\/\/github.com\/IBM\/tensorflow-large-model-support.  IBM. 2018. TensorFlow Large-Model-Support https:\/\/github.com\/IBM\/tensorflow-large-model-support."},{"key":"e_1_2_1_24_1","unstructured":"IBM. 2020. PyTorch Large-Model-Support https:\/\/github.com\/IBM\/pytorch-large-model-support.  IBM. 2020. PyTorch Large-Model-Support https:\/\/github.com\/IBM\/pytorch-large-model-support."},{"key":"e_1_2_1_25_1","unstructured":"Intel. 2017. Intel Xeon Processors https:\/\/www.intel.com\/content\/www\/us\/en\/products\/details\/processors\/xeon.html.  Intel. 2017. Intel Xeon Processors https:\/\/www.intel.com\/content\/www\/us\/en\/products\/details\/processors\/xeon.html."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00070"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of Machine Learning and Systems (MLSys'20)","volume":"2","author":"Jain Paras","year":"2020","unstructured":"Paras Jain , Ajay Jain , Aniruddha Nrusimha , Amir Gholami , Pieter Abbeel , Joseph Gonzalez , Kurt Keutzer , and Ion Stoica . 2020 . Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization . In Proceedings of Machine Learning and Systems (MLSys'20) , Vol. 2 . 497--511. Paras Jain, Ajay Jain, Aniruddha Nrusimha, Amir Gholami, Pieter Abbeel, Joseph Gonzalez, Kurt Keutzer, and Ion Stoica. 2020. Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization. In Proceedings of Machine Learning and Systems (MLSys'20), Vol. 2. 497--511."},{"key":"e_1_2_1_28_1","volume-title":"Layer-centric memory reuse and data migration for extreme-scale deep learning on many-core architectures. ACM Transactions on Architecture and Code Optimization (TACO'18)","author":"Jin Hai","year":"2018","unstructured":"Hai Jin , Bo Liu , Wenbin Jiang , Yang Ma , Xuanhua Shi , Bingsheng He , and Shaofeng Zhao . 2018. Layer-centric memory reuse and data migration for extreme-scale deep learning on many-core architectures. ACM Transactions on Architecture and Code Optimization (TACO'18) ( 2018 ), 1--26. Hai Jin, Bo Liu, Wenbin Jiang, Yang Ma, Xuanhua Shi, Bingsheng He, and Shaofeng Zhao. 2018. Layer-centric memory reuse and data migration for extreme-scale deep learning on many-core architectures. ACM Transactions on Architecture and Code Optimization (TACO'18) (2018), 1--26."},{"key":"e_1_2_1_29_1","volume-title":"Identity Mappings in Deep Residual Networks. arXiv preprint arXiv 1603.05027","author":"He Kaiming","year":"2016","unstructured":"Kaiming He and Xiangyu Zhang and Shaoqing Ren and Jian Sun . 2016. Identity Mappings in Deep Residual Networks. arXiv preprint arXiv 1603.05027 ( 2016 ). Kaiming He and Xiangyu Zhang and Shaoqing Ren and Jian Sun. 2016. Identity Mappings in Deep Residual Networks. arXiv preprint arXiv 1603.05027 (2016)."},{"key":"e_1_2_1_30_1","volume-title":"Scaling laws for neural language models. arXiv preprint arXiv","author":"Kaplan Jared","year":"2001","unstructured":"Jared Kaplan , Sam McCandlish , Tom Henighan , Tom B Brown , Benjamin Chess , Rewon Child , Scott Gray , Alec Radford , Jeffrey Wu , and Dario Amodei . 2020. Scaling laws for neural language models. arXiv preprint arXiv 2001 .08361 (2020). Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv preprint arXiv 2001.08361 (2020)."},{"key":"e_1_2_1_31_1","volume-title":"One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv 1404.5997","author":"Krizhevsky Alex","year":"2014","unstructured":"Alex Krizhevsky . 2014. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv 1404.5997 ( 2014 ). Alex Krizhevsky. 2014. One weird trick for parallelizing convolutional neural networks. arXiv preprint arXiv 1404.5997 (2014)."},{"key":"e_1_2_1_32_1","volume-title":"Proceedings of the 25th International Conference on Neural Information Processing Systems (NeurIPS'12)","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E Hinton . 2012 . Imagenet Classification with Deep Convolutional Neural Networks . In Proceedings of the 25th International Conference on Neural Information Processing Systems (NeurIPS'12) . Lake Tahoe, NV, 1097--1105. Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Imagenet Classification with Deep Convolutional Neural Networks. In Proceedings of the 25th International Conference on Neural Information Processing Systems (NeurIPS'12). Lake Tahoe, NV, 1097--1105."},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR'21)","author":"Lepikhin Dmitry","year":"2021","unstructured":"Dmitry Lepikhin , HyoukJoong Lee , Yuanzhong Xu , Dehao Chen , Orhan Firat , Yanping Huang , Maxim Krikun , Noam Shazeer , and Zhifeng Chen . 2021 . GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding . In Proceedings of the International Conference on Learning Representations (ICLR'21) . Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2021. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. In Proceedings of the International Conference on Learning Representations (ICLR'21)."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2019.2928289"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.5555\/2685048.2685095"},{"key":"e_1_2_1_36_1","unstructured":"Shen Li Yanli Zhao Rohan Varma Omkar Salpekar Pieter Noordhuis Teng Li Adam Paszke Jeff Smith Brian Vaughan Pritam Damania etal 2020. Pytorch distributed: Experiences on accelerating data parallel training. arXiv preprint arXiv 2006.15704 (2020).  Shen Li Yanli Zhao Rohan Varma Omkar Salpekar Pieter Noordhuis Teng Li Adam Paszke Jeff Smith Brian Vaughan Pritam Damania et al. 2020. Pytorch distributed: Experiences on accelerating data parallel training. arXiv preprint arXiv 2006.15704 (2020)."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/3415478.3415530"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458336.3465289"},{"key":"e_1_2_1_39_1","volume-title":"Harmony: Overcoming the Hurdles of GPU Memory Capacity to Train Massive DNN Models on Commodity Servers. arXiv preprint arXiv.2202.01306","author":"Li Youjie","year":"2022","unstructured":"Youjie Li , Amar Phanishayee , Derek Murray , Jakub Tarnawski , and Nam Sung Kim . 2022 . Harmony: Overcoming the Hurdles of GPU Memory Capacity to Train Massive DNN Models on Commodity Servers. arXiv preprint arXiv.2202.01306 (2022). https:\/\/arxiv.org\/abs\/2202.01306 Youjie Li, Amar Phanishayee, Derek Murray, Jakub Tarnawski, and Nam Sung Kim. 2022. Harmony: Overcoming the Hurdles of GPU Memory Capacity to Train Massive DNN Models on Commodity Servers. arXiv preprint arXiv.2202.01306 (2022). https:\/\/arxiv.org\/abs\/2202.01306"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS'18)","author":"Li Youjie","year":"2018","unstructured":"Youjie Li , Mingchao Yu , Songze Li , Salman Avestimehr , Nam Sung Kim , and Alexander Schwing . 2018 . Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training . In Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS'18) . Montreal, Canada, 8056--8067. Youjie Li, Mingchao Yu, Songze Li, Salman Avestimehr, Nam Sung Kim, and Alexander Schwing. 2018. Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training. In Proceedings of the 32nd Conference on Neural Information Processing Systems (NeurIPS'18). Montreal, Canada, 8056--8067."},{"key":"e_1_2_1_41_1","volume-title":"How Do You Know a Human Wrote This? The New York Times","author":"Manjoo Farhad","year":"2020","unstructured":"Farhad Manjoo . 2020. How Do You Know a Human Wrote This? The New York Times ( 2020 ). Farhad Manjoo. 2020. How Do You Know a Human Wrote This? The New York Times (2020)."},{"key":"e_1_2_1_42_1","volume-title":"Pointer sentinel mixture models. arXiv preprint arXiv 1609.07843","author":"Merity Stephen","year":"2016","unstructured":"Stephen Merity , Caiming Xiong , James Bradbury , and Richard Socher . 2016. Pointer sentinel mixture models. arXiv preprint arXiv 1609.07843 ( 2016 ). Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2016. Pointer sentinel mixture models. arXiv preprint arXiv 1609.07843 (2016)."},{"key":"e_1_2_1_43_1","unstructured":"Microsoft. 2020. GPT-2 fine-tuning with ONNX Runtime https:\/\/cloudblogs.microsoft.com\/opensource\/2020\/08\/24\/pytorch-gpt-2-fine-tuning-onnx-runtime-speedup-training-time.  Microsoft. 2020. GPT-2 fine-tuning with ONNX Runtime https:\/\/cloudblogs.microsoft.com\/opensource\/2020\/08\/24\/pytorch-gpt-2-fine-tuning-onnx-runtime-speedup-training-time."},{"key":"e_1_2_1_44_1","volume-title":"Riedmiller","author":"Mnih Volodymyr","year":"2013","unstructured":"Volodymyr Mnih , Koray Kavukcuoglu , David Silver , Alex Graves , Ioannis Antonoglou , Daan Wierstra , and Martin A . Riedmiller . 2013 . Playing Atari with Deep Reinforcement Learning . arXiv preprint arXiv 1312.5602 (2013). Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller. 2013. Playing Atari with Deep Reinforcement Learning. arXiv preprint arXiv 1312.5602 (2013)."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_2_1_46_1","volume-title":"International Conference on Machine Learning(ICML'21)","author":"Narayanan Deepak","year":"2021","unstructured":"Deepak Narayanan , Amar Phanishayee , Kaiyu Shi , Xie Chen , and Matei Zaharia . 2021 . Memory-efficient pipeline-parallel DNN training . In International Conference on Machine Learning(ICML'21) . 7937--7947. Deepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen, and Matei Zaharia. 2021. Memory-efficient pipeline-parallel DNN training. In International Conference on Machine Learning(ICML'21). 7937--7947."},{"key":"e_1_2_1_47_1","volume-title":"Efficient Large-Scale Language Model Training on GPU Clusters. arXiv preprint arXiv:2104.04473","author":"Narayanan Deepak","year":"2021","unstructured":"Deepak Narayanan , Mohammad Shoeybi , Jared Casper , Patrick LeGresley , Mostofa Patwary , Vijay Korthikanti , Dmitri Vainbrand , Prethvi Kashinkunti , Julie Bernauer , Bryan Catanzaro , Amar Phanishayee , and Matei Zaharia . 2021. Efficient Large-Scale Language Model Training on GPU Clusters. arXiv preprint arXiv:2104.04473 ( 2021 ). Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021. Efficient Large-Scale Language Model Training on GPU Clusters. arXiv preprint arXiv:2104.04473 (2021)."},{"key":"e_1_2_1_48_1","unstructured":"NVIDIA. 2016. NVLINK AND NVSWITCH https:\/\/www.nvidia.com\/en-us\/data-center\/nvlink\/.  NVIDIA. 2016. NVLINK AND NVSWITCH https:\/\/www.nvidia.com\/en-us\/data-center\/nvlink\/."},{"key":"e_1_2_1_49_1","unstructured":"NVIDIA. 2017. GEFORCE GTX GPU https:\/\/www.nvidia.com\/en-in\/geforce\/products\/10series\/geforce-gtx-1080-ti\/.  NVIDIA. 2017. GEFORCE GTX GPU https:\/\/www.nvidia.com\/en-in\/geforce\/products\/10series\/geforce-gtx-1080-ti\/."},{"key":"e_1_2_1_50_1","unstructured":"NVIDIA. 2017. NVIDIA DGX-1 System Architecture White Paper https:\/\/www.azken.com\/images\/dgx1_images\/dgx1-system-architecture-whitepaper1.pdf.  NVIDIA. 2017. NVIDIA DGX-1 System Architecture White Paper https:\/\/www.azken.com\/images\/dgx1_images\/dgx1-system-architecture-whitepaper1.pdf."},{"key":"e_1_2_1_51_1","unstructured":"NVIDIA. 2017. Unified Memory https:\/\/developer.nvidia.com\/blog\/unified-memory-cuda-beginners\/.  NVIDIA. 2017. Unified Memory https:\/\/developer.nvidia.com\/blog\/unified-memory-cuda-beginners\/."},{"key":"e_1_2_1_52_1","unstructured":"NVIDIA. 2018. NVIDIA DGX-2H The World's Most Powerful System for The Most Complex AI Challenges https:\/\/www.nvidia.com\/content\/dam\/en-zz\/es_em\/Solutions\/Data-Center\/dgx-2\/dgx-2h-datasheet-us-nvidia-841283-r6-web.pdf.  NVIDIA. 2018. NVIDIA DGX-2H The World's Most Powerful System for The Most Complex AI Challenges https:\/\/www.nvidia.com\/content\/dam\/en-zz\/es_em\/Solutions\/Data-Center\/dgx-2\/dgx-2h-datasheet-us-nvidia-841283-r6-web.pdf."},{"key":"e_1_2_1_53_1","unstructured":"NVIDIA. 2021. NVIDIA TESLA GPUs https:\/\/en.wikipedia.org\/wiki\/Nvidia_Tesla.  NVIDIA. 2021. NVIDIA TESLA GPUs https:\/\/en.wikipedia.org\/wiki\/Nvidia_Tesla."},{"key":"e_1_2_1_54_1","volume-title":"Proceedings of the 25th International Conference on Neural Information Processing Systems (NeurIPS'19)","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas Kopf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019 . PyTorch: An Imperative Style, High-Performance Deep Learning Library . In Proceedings of the 25th International Conference on Neural Information Processing Systems (NeurIPS'19) . Vancouver, Canada, 8024--8035. Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Proceedings of the 25th International Conference on Neural Information Processing Systems (NeurIPS'19). Vancouver, Canada, 8024--8035."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378505"},{"key":"e_1_2_1_56_1","unstructured":"PNY. 2021. Single Root Complex Purley 4U GPU Server for Deep Learning Applications https:\/\/www.pny.eu\/en\/consumer\/explore-all-products\/pny-gpu-servers\/\\983-single-root-complex-purley-4u-gpu\\-server-for-deep-learning-applications.  PNY. 2021. Single Root Complex Purley 4U GPU Server for Deep Learning Applications https:\/\/www.pny.eu\/en\/consumer\/explore-all-products\/pny-gpu-servers\/\\983-single-root-complex-purley-4u-gpu\\-server-for-deep-learning-applications."},{"key":"e_1_2_1_57_1","volume-title":"OpenAI","author":"Radford Alec","year":"2019","unstructured":"Alec Radford , Jeff Wu , Rewon Child , David Luan , Dario Amodei , and Ilya Sutskever . 2019. Language Models are Unsupervised Multitask Learners. Technical report , OpenAI ( 2019 ). Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. Technical report, OpenAI (2019)."},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00024"},{"key":"e_1_2_1_59_1","volume-title":"ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning. arXiv preprint arXiv:2104.07857","author":"Rajbhandari Samyam","year":"2021","unstructured":"Samyam Rajbhandari , Olatunji Ruwase , Jeff Rasley , Shaden Smith , and Yuxiong He. 2021. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning. arXiv preprint arXiv:2104.07857 ( 2021 ). Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021. ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning. arXiv preprint arXiv:2104.07857 (2021)."},{"key":"e_1_2_1_60_1","volume-title":"Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He.","author":"Ren Jie","year":"2021","unstructured":"Jie Ren , Samyam Rajbhandari , Reza Yazdani Aminabadi , Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021 . ZeRO-Offload: Democratizing Billion-Scale Model Training . arXiv preprint arXiv.2101.06840 (2021). Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021. ZeRO-Offload: Democratizing Billion-Scale Model Training. arXiv preprint arXiv.2101.06840 (2021)."},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2016.7783721"},{"key":"e_1_2_1_62_1","volume-title":"The Largest Model So Far. Analytics India Magazine","author":"Sagar Ram","year":"2020","unstructured":"Ram Sagar . 2020. OpenAI Releases GPT-3 , The Largest Model So Far. Analytics India Magazine ( 2020 ). Ram Sagar. 2020. OpenAI Releases GPT-3, The Largest Model So Far. Analytics India Magazine (2020)."},{"key":"e_1_2_1_63_1","volume-title":"Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv","author":"Shoeybi Mohammad","year":"2019","unstructured":"Mohammad Shoeybi , Mostofa Patwary , Raul Puri , Patrick LeGresley , Jared Casper , and Bryan Catanzaro . 2019 . Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv 1909.08053 (2019). Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv preprint arXiv 1909.08053 (2019)."},{"key":"e_1_2_1_64_1","volume-title":"High-resolution representations for labeling pixels and regions. arXiv preprint arXiv.1904.04514","author":"Sun Ke","year":"2019","unstructured":"Ke Sun , Yang Zhao , Borui Jiang , Tianheng Cheng , Bin Xiao , Dong Liu , Yadong Mu , Xinggang Wang , Wenyu Liu , and Jingdong Wang . 2019. High-resolution representations for labeling pixels and regions. arXiv preprint arXiv.1904.04514 ( 2019 ). Ke Sun, Yang Zhao, Borui Jiang, Tianheng Cheng, Bin Xiao, Dong Liu, Yadong Mu, Xinggang Wang, Wenyu Liu, and Jingdong Wang. 2019. High-resolution representations for labeling pixels and regions. arXiv preprint arXiv.1904.04514 (2019)."},{"key":"e_1_2_1_65_1","volume-title":"GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv","author":"Wang Alex","year":"2018","unstructured":"Alex Wang , Amanpreet Singh , Julian Michael , Felix Hill , Omer Levy , and Samuel R Bowman . 2018 . GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv 1804.07461 (2018). Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman. 2018. GLUE: A multi-task benchmark and analysis platform for natural language understanding. arXiv preprint arXiv 1804.07461 (2018)."},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178491"},{"key":"e_1_2_1_67_1","unstructured":"Yonghui Wu Mike Schuster Zhifeng Chen Quoc V Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey etal 2016. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv: 1609.08144 (2016).  Yonghui Wu Mike Schuster Zhifeng Chen Quoc V Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey et al. 2016. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv preprint arXiv: 1609.08144 (2016)."},{"key":"e_1_2_1_68_1","volume-title":"Large Batch Training of Convolutional Networks. arXiv preprint arXiv 1708.03888","author":"You Yang","year":"2017","unstructured":"Yang You , Igor Gitman , and Boris Ginsburg . 2017. Large Batch Training of Convolutional Networks. arXiv preprint arXiv 1708.03888 ( 2017 ). Yang You, Igor Gitman, and Boris Ginsburg. 2017. Large Batch Training of Convolutional Networks. arXiv preprint arXiv 1708.03888 (2017)."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3551793.3551828","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:37:50Z","timestamp":1672223870000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3551793.3551828"}},"subtitle":["overcoming the hurdles of GPU memory capacity to train massive DNN models on commodity servers"],"short-title":[],"issued":{"date-parts":[[2022,7]]},"references-count":68,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2022,7]]}},"alternative-id":["10.14778\/3551793.3551828"],"URL":"https:\/\/doi.org\/10.14778\/3551793.3551828","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2022,7]]}}}