{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,26]],"date-time":"2025-06-26T04:51:50Z","timestamp":1750913510557,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":68,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,6,1]],"date-time":"2021-06-01T00:00:00Z","timestamp":1622505600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,6]]},"DOI":"10.1145\/3458336.3465289","type":"proceedings-article","created":{"date-parts":[[2021,6,4]],"date-time":"2021-06-04T02:03:56Z","timestamp":1622772236000},"page":"119-127","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Doing more with less"],"prefix":"10.1145","author":[{"given":"Youjie","family":"Li","sequence":"first","affiliation":[{"name":"University of Illinois at Urbana-Champaign"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Amar","family":"Phanishayee","sequence":"additional","affiliation":[{"name":"Microsoft Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Derek","family":"Murray","sequence":"additional","affiliation":[{"name":"Microsoft"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nam Sung","family":"Kim","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana-Champaign"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,6,3]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Deep learning for computational biology. Molecular systems biology 12, 7","author":"Angermueller Christof","year":"2016","unstructured":"Christof Angermueller , Tanel P\u00e4rnamaa , Leopold Parts , and Oliver Stegle . 2016. Deep learning for computational biology. Molecular systems biology 12, 7 ( 2016 ), 878. Christof Angermueller, Tanel P\u00e4rnamaa, Leopold Parts, and Oliver Stegle. 2016. Deep learning for computational biology. Molecular systems biology 12, 7 (2016), 878."},{"unstructured":"ASUS. 2019. High-density 4U GPU server https:\/\/www.asus.com\/us\/Commercial-Servers-Workstations\/ESC8000-G4\/HelpDesk_Manual.  ASUS. 2019. High-density 4U GPU server https:\/\/www.asus.com\/us\/Commercial-Servers-Workstations\/ESC8000-G4\/HelpDesk_Manual.","key":"e_1_3_2_1_2_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_3_1","DOI":"10.1145\/3320060"},{"key":"e_1_3_2_1_4_1","volume-title":"GPT-2 fine-tuning with ONNX Runtime. Microsoft Open Source Blog","author":"Bhandare Aishwarya","year":"2020","unstructured":"Aishwarya Bhandare , Tianju Xu , and Kshama Pawar . 2020. GPT-2 fine-tuning with ONNX Runtime. Microsoft Open Source Blog ( 2020 ). https:\/\/cloudblogs.microsoft.com\/opensource\/2020\/08\/24\/pytorch-gpt-2-fine-tuning-onnx-runtime-speedup-training-time Aishwarya Bhandare, Tianju Xu, and Kshama Pawar. 2020. GPT-2 fine-tuning with ONNX Runtime. Microsoft Open Source Blog (2020). https:\/\/cloudblogs.microsoft.com\/opensource\/2020\/08\/24\/pytorch-gpt-2-fine-tuning-onnx-runtime-speedup-training-time"},{"unstructured":"Tom B Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell etal 2020. Language models are few-shot learners. arXiv arXiv\/2005.14165 (2020).  Tom B Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell et al. 2020. Language models are few-shot learners. arXiv arXiv\/2005.14165 (2020).","key":"e_1_3_2_1_5_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_6_1","DOI":"10.5555\/1841497"},{"key":"e_1_3_2_1_7_1","volume-title":"Training deep nets with sublinear memory cost. arXiv arXiv\/1604.06174","author":"Chen Tianqi","year":"2016","unstructured":"Tianqi Chen , Bing Xu , Chiyuan Zhang , and Carlos Guestrin . 2016. Training deep nets with sublinear memory cost. arXiv arXiv\/1604.06174 ( 2016 ). Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin. 2016. Training deep nets with sublinear memory cost. arXiv arXiv\/1604.06174 (2016)."},{"key":"e_1_3_2_1_8_1","volume-title":"Thirteenth Annual Conference of the International Speech Communication Association (INTERSPEECH'12)","author":"Chen Xie","year":"2012","unstructured":"Xie Chen , Adam Eversole , Gang Li , Dong Yu , and Frank Seide . 2012 . Pipelined back-propagation for context-dependent deep neural networks . In Thirteenth Annual Conference of the International Speech Communication Association (INTERSPEECH'12) . Portland, USA. Xie Chen, Adam Eversole, Gang Li, Dong Yu, and Frank Seide. 2012. Pipelined back-propagation for context-dependent deep neural networks. In Thirteenth Annual Conference of the International Speech Communication Association (INTERSPEECH'12). Portland, USA."},{"unstructured":"Minsik Cho Tung D Le U Finkler Haruiki Imai Yasushi Negishi Taro Sekiyama Saritha Vinod Vladimir Zolotov Kiyokuni Kawachiya David S Kung etal 2018. Large model support for deep learning in caffe and chainer. SysML'18 (Feb. 2018).  Minsik Cho Tung D Le U Finkler Haruiki Imai Yasushi Negishi Taro Sekiyama Saritha Vinod Vladimir Zolotov Kiyokuni Kawachiya David S Kung et al. 2018. Large model support for deep learning in caffe and chainer. SysML'18 (Feb. 2018).","key":"e_1_3_2_1_9_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_10_1","DOI":"10.1145\/1327452.1327492"},{"key":"e_1_3_2_1_11_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv arXiv\/1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv arXiv\/1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv arXiv\/1810.04805 (2018)."},{"unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly etal 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv arXiv\/2010.11929 (2020).  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv arXiv\/2010.11929 (2020).","key":"e_1_3_2_1_12_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_13_1","DOI":"10.5555\/3026877.3026886"},{"key":"e_1_3_2_1_14_1","volume-title":"Generative adversarial networks. arXiv arXiv\/1406.2661","author":"Goodfellow Ian J","year":"2014","unstructured":"Ian J Goodfellow , Jean Pouget-Abadie , Mehdi Mirza , Bing Xu , David Warde-Farley , Sherjil Ozair , Aaron Courville , and Yoshua Bengio . 2014. Generative adversarial networks. arXiv arXiv\/1406.2661 ( 2014 ). Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative adversarial networks. arXiv arXiv\/1406.2661 (2014)."},{"key":"e_1_3_2_1_15_1","volume-title":"A robot wrote this entire article. Are you scared yet, human? The Guardian","author":"Guardian The","year":"2020","unstructured":"The Guardian . 2020. A robot wrote this entire article. Are you scared yet, human? The Guardian ( 2020 ). The Guardian. 2020. A robot wrote this entire article. Are you scared yet, human? The Guardian (2020)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_16_1","DOI":"10.1109\/CVPR.2016.90"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_17_1","DOI":"10.5555\/567110"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_18_1","DOI":"10.1109\/CVPR.2018.00745"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_19_1","DOI":"10.1145\/3373376.3378530"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_20_1","DOI":"10.5555\/3454287.3454297"},{"unstructured":"Hugging Face. 2021. Transformer Examples https:\/\/huggingface.co\/transformers\/v2.3.0\/examples.html.  Hugging Face. 2021. Transformer Examples https:\/\/huggingface.co\/transformers\/v2.3.0\/examples.html.","key":"e_1_3_2_1_21_1"},{"key":"e_1_3_2_1_22_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning (ICML'18)","author":"Huo Zhouyuan","year":"2018","unstructured":"Zhouyuan Huo , Bin Gu , Heng Huang , 2018 . Decoupled parallel backpropagation with convergence guarantee . In Proceedings of the 35th International Conference on Machine Learning (ICML'18) . Stockholm, Sweden. Zhouyuan Huo, Bin Gu, Heng Huang, et al. 2018. Decoupled parallel backpropagation with convergence guarantee. In Proceedings of the 35th International Conference on Machine Learning (ICML'18). Stockholm, Sweden."},{"unstructured":"IBM. 2018. TensorFlow Large-Model-Support https:\/\/github.com\/IBM\/tensorflow-large-model-support.  IBM. 2018. TensorFlow Large-Model-Support https:\/\/github.com\/IBM\/tensorflow-large-model-support.","key":"e_1_3_2_1_23_1"},{"unstructured":"IBM. 2020. PyTorch Large-Model-Support https:\/\/github.com\/IBM\/pytorch-large-model-support.  IBM. 2020. PyTorch Large-Model-Support https:\/\/github.com\/IBM\/pytorch-large-model-support.","key":"e_1_3_2_1_24_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_25_1","DOI":"10.1145\/1272998.1273005"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_26_1","DOI":"10.1145\/1629575.1629601"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_27_1","DOI":"10.1109\/ISCA.2018.00070"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_28_1","DOI":"10.1145\/3243904"},{"key":"e_1_3_2_1_29_1","volume-title":"Scaling laws for neural language models. arXiv arXiv\/2001.08361","author":"Kaplan Jared","year":"2020","unstructured":"Jared Kaplan , Sam McCandlish , Tom Henighan , Tom B Brown , Benjamin Chess , Rewon Child , Scott Gray , Alec Radford , Jeffrey Wu , and Dario Amodei . 2020. Scaling laws for neural language models. arXiv arXiv\/2001.08361 ( 2020 ). Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. arXiv arXiv\/2001.08361 (2020)."},{"key":"e_1_3_2_1_30_1","volume-title":"Adam: A method for stochastic optimization. arXiv arXiv\/1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba . 2014 . Adam: A method for stochastic optimization. arXiv arXiv\/1412.6980 (2014). Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv arXiv\/1412.6980 (2014)."},{"key":"e_1_3_2_1_31_1","volume-title":"One weird trick for parallelizing convolutional neural networks. arXiv arXiv\/1404.5997","author":"Krizhevsky Alex","year":"2014","unstructured":"Alex Krizhevsky . 2014. One weird trick for parallelizing convolutional neural networks. arXiv arXiv\/1404.5997 ( 2014 ). Alex Krizhevsky. 2014. One weird trick for parallelizing convolutional neural networks. arXiv arXiv\/1404.5997 (2014)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_32_1","DOI":"10.5555\/2999134.2999257"},{"key":"e_1_3_2_1_33_1","first-page":"e190177","article-title":"The importance of image resolution in building deep learning models for medical imaging. Radiology","volume":"2","author":"Lakhani Paras","year":"2020","unstructured":"Paras Lakhani . 2020 . The importance of image resolution in building deep learning models for medical imaging. Radiology : Artificial Intelligence 2 , 1 (2020), e190177 . Paras Lakhani. 2020. The importance of image resolution in building deep learning models for medical imaging. Radiology: Artificial Intelligence 2, 1 (2020), e190177.","journal-title":"Artificial Intelligence"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_34_1","DOI":"10.1109\/5.726791"},{"key":"e_1_3_2_1_35_1","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR'21)","author":"Lepikhin Dmitry","year":"2021","unstructured":"Dmitry Lepikhin , HyoukJoong Lee , Yuanzhong Xu , Dehao Chen , Orhan Firat , Yanping Huang , Maxim Krikun , Noam Shazeer , and Zhifeng Chen . 2021 . GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding . In Proceedings of the International Conference on Learning Representations (ICLR'21) . Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2021. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. In Proceedings of the International Conference on Learning Representations (ICLR'21)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_36_1","DOI":"10.1109\/TPDS.2019.2928289"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_37_1","DOI":"10.14778\/3415478.3415530"},{"key":"e_1_3_2_1_38_1","volume-title":"How Do You Know a Human Wrote This? The New York Times","author":"Manjoo Farhad","year":"2020","unstructured":"Farhad Manjoo . 2020. How Do You Know a Human Wrote This? The New York Times ( 2020 ). Farhad Manjoo. 2020. How Do You Know a Human Wrote This? The New York Times (2020)."},{"key":"e_1_3_2_1_39_1","volume-title":"Nitish Shirish Keskar, and Richard Socher","author":"Merity Stephen","year":"2017","unstructured":"Stephen Merity , Nitish Shirish Keskar, and Richard Socher . 2017 . Regularizing and optimizing LSTM language models. arXiv arXiv\/1708.02182 (2017). Stephen Merity, Nitish Shirish Keskar, and Richard Socher. 2017. Regularizing and optimizing LSTM language models. arXiv arXiv\/1708.02182 (2017)."},{"key":"e_1_3_2_1_40_1","volume-title":"Riedmiller","author":"Mnih Volodymyr","year":"2013","unstructured":"Volodymyr Mnih , Koray Kavukcuoglu , David Silver , Alex Graves , Ioannis Antonoglou , Daan Wierstra , and Martin A . Riedmiller . 2013 . Playing Atari with Deep Reinforcement Learning . arXiv arXiv\/1312.5602 (2013). http:\/\/arxiv.org\/abs\/1312.5602 Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin A. Riedmiller. 2013. Playing Atari with Deep Reinforcement Learning. arXiv arXiv\/1312.5602 (2013). http:\/\/arxiv.org\/abs\/1312.5602"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_41_1","DOI":"10.5555\/3291168.3291210"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_42_1","DOI":"10.1145\/3341301.3359646"},{"key":"e_1_3_2_1_43_1","volume-title":"Memory-efficient pipeline-parallel DNN training. arXiv arXiv\/2006.09503","author":"Narayanan Deepak","year":"2020","unstructured":"Deepak Narayanan , Amar Phanishayee , Kaiyu Shi , Xie Chen , and Matei Zaharia . 2020. Memory-efficient pipeline-parallel DNN training. arXiv arXiv\/2006.09503 ( 2020 ). Deepak Narayanan, Amar Phanishayee, Kaiyu Shi, Xie Chen, and Matei Zaharia. 2020. Memory-efficient pipeline-parallel DNN training. arXiv arXiv\/2006.09503 (2020)."},{"key":"e_1_3_2_1_44_1","volume-title":"Efficient Large-Scale Language Model Training on GPU Clusters. arXiv arXiv\/2104.04473","author":"Narayanan Deepak","year":"2021","unstructured":"Deepak Narayanan , Mohammad Shoeybi , Jared Casper , Patrick LeGresley , Mostofa Patwary , Vijay Korthikanti , Dmitri Vainbrand , Prethvi Kashinkunti , Julie Bernauer , Bryan Catanzaro , Amar Phanishayee , and Matei Zaharia . 2021. Efficient Large-Scale Language Model Training on GPU Clusters. arXiv arXiv\/2104.04473 ( 2021 ). Deepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley, Mostofa Patwary, Vijay Korthikanti, Dmitri Vainbrand, Prethvi Kashinkunti, Julie Bernauer, Bryan Catanzaro, Amar Phanishayee, and Matei Zaharia. 2021. Efficient Large-Scale Language Model Training on GPU Clusters. arXiv arXiv\/2104.04473 (2021)."},{"unstructured":"NVIDIA. 2017. NVIDIA DGX-1 System Architecture White Paper https:\/\/www.azken.com\/images\/dgx1_images\/dgx1-system-architecture-whitepaper1.pdf.  NVIDIA. 2017. NVIDIA DGX-1 System Architecture White Paper https:\/\/www.azken.com\/images\/dgx1_images\/dgx1-system-architecture-whitepaper1.pdf.","key":"e_1_3_2_1_45_1"},{"unstructured":"NVIDIA. 2017. Unified Memory https:\/\/developer.nvidia.com\/blog\/unified-memory-cuda-beginners\/.  NVIDIA. 2017. Unified Memory https:\/\/developer.nvidia.com\/blog\/unified-memory-cuda-beginners\/.","key":"e_1_3_2_1_46_1"},{"unstructured":"NVIDIA. 2018. NVIDIA DGX-2H The World's Most Powerful System for The Most Complex AI Challenges https:\/\/www.nvidia.com\/content\/dam\/en-zz\/es_em\/Solutions\/Data-Center\/dgx-2\/dgx-2h-datasheet-us-nvidia-841283-r6-web.pdf.  NVIDIA. 2018. NVIDIA DGX-2H The World's Most Powerful System for The Most Complex AI Challenges https:\/\/www.nvidia.com\/content\/dam\/en-zz\/es_em\/Solutions\/Data-Center\/dgx-2\/dgx-2h-datasheet-us-nvidia-841283-r6-web.pdf.","key":"e_1_3_2_1_47_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_48_1","DOI":"10.1145\/2517349.2522716"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_49_1","DOI":"10.5555\/3454287.3455008"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_50_1","DOI":"10.1145\/3373376.3378505"},{"unstructured":"PNY. 2021. Single Root Complex Purley 4U GPU Server for Deep Learning Applications https:\/\/www.pny.eu\/en\/consumer\/explore-all-products\/pny-gpu-servers\/983-single-root-complex-purley-4u-gpu-server-for-deep-learning-applications  PNY. 2021. Single Root Complex Purley 4U GPU Server for Deep Learning Applications https:\/\/www.pny.eu\/en\/consumer\/explore-all-products\/pny-gpu-servers\/983-single-root-complex-purley-4u-gpu-server-for-deep-learning-applications","key":"e_1_3_2_1_51_1"},{"key":"e_1_3_2_1_52_1","volume-title":"OpenAi","author":"Radford Alec","year":"2019","unstructured":"Alec Radford , Jeff Wu , Rewon Child , David Luan , Dario Amodei , and Ilya Sutskever . 2019. Language Models are Unsupervised Multitask Learners. Technical report , OpenAi ( 2019 ). Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language Models are Unsupervised Multitask Learners. Technical report, OpenAi (2019)."},{"key":"e_1_3_2_1_53_1","volume-title":"Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv arXiv\/1910.10683","author":"Raffel Colin","year":"2019","unstructured":"Colin Raffel , Noam Shazeer , Adam Roberts , Katherine Lee , Sharan Narang , Michael Matena , Yanqi Zhou , Wei Li , and Peter J Liu . 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv arXiv\/1910.10683 ( 2019 ). Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2019. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv arXiv\/1910.10683 (2019)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_54_1","DOI":"10.5555\/3433701.3433727"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_55_1","DOI":"10.1609\/aaai.v33i01.33014780"},{"key":"e_1_3_2_1_56_1","volume-title":"Proceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA'21)","author":"Ren Jie","year":"2021","unstructured":"Jie Ren , Jiaolin Luo , Kai Wu , Minjia Zhang , and Dong Li . 2021 . Sentinel: Runtime Data Management on Heterogeneous Main MemorySystems for Deep Learning . In Proceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA'21) . Jie Ren, Jiaolin Luo, Kai Wu, Minjia Zhang, and Dong Li. 2021. Sentinel: Runtime Data Management on Heterogeneous Main MemorySystems for Deep Learning. In Proceedings of the IEEE International Symposium on High-Performance Computer Architecture (HPCA'21)."},{"key":"e_1_3_2_1_57_1","volume-title":"Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He.","author":"Ren Jie","year":"2021","unstructured":"Jie Ren , Samyam Rajbhandari , Reza Yazdani Aminabadi , Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021 . ZeRO-Offload: Democratizing Billion-Scale Model Training . arXiv arXiv\/2101.06840 (2021). Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021. ZeRO-Offload: Democratizing Billion-Scale Model Training. arXiv arXiv\/2101.06840 (2021)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_58_1","DOI":"10.5555\/3195638.3195660"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_59_1","DOI":"10.1038\/323533a0"},{"key":"e_1_3_2_1_60_1","first-page":"e190015","article-title":"The effect of image resolution on deep learning in radiography. Radiology","volume":"2","author":"Sabottke Carl F","year":"2020","unstructured":"Carl F Sabottke and Bradley M Spieler . 2020 . The effect of image resolution on deep learning in radiography. Radiology : Artificial Intelligence 2 , 1 (2020), e190015 . Carl F Sabottke and Bradley M Spieler. 2020. The effect of image resolution on deep learning in radiography. Radiology: Artificial Intelligence 2, 1 (2020), e190015.","journal-title":"Artificial Intelligence"},{"key":"e_1_3_2_1_61_1","volume-title":"The Largest Model So Far. Analytics India Magazine","author":"Sagar Ram","year":"2020","unstructured":"Ram Sagar . 2020. OpenAI Releases GPT-3 , The Largest Model So Far. Analytics India Magazine ( 2020 ). Ram Sagar. 2020. OpenAI Releases GPT-3, The Largest Model So Far. Analytics India Magazine (2020)."},{"key":"e_1_3_2_1_62_1","volume-title":"Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv arXiv\/1909.08053","author":"Shoeybi Mohammad","year":"2019","unstructured":"Mohammad Shoeybi , Mostofa Patwary , Raul Puri , Patrick LeGresley , Jared Casper , and Bryan Catanzaro . 2019 . Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv arXiv\/1909.08053 (2019). Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019. Megatron-lm: Training multi-billion parameter language models using model parallelism. arXiv arXiv\/1909.08053 (2019)."},{"key":"e_1_3_2_1_63_1","volume-title":"High-resolution representations for labeling pixels and regions. arXiv arXiv\/1904.04514","author":"Sun Ke","year":"2019","unstructured":"Ke Sun , Yang Zhao , Borui Jiang , Tianheng Cheng , Bin Xiao , Dong Liu , Yadong Mu , Xinggang Wang , Wenyu Liu , and Jingdong Wang . 2019. High-resolution representations for labeling pixels and regions. arXiv arXiv\/1904.04514 ( 2019 ). Ke Sun, Yang Zhao, Borui Jiang, Tianheng Cheng, Bin Xiao, Dong Liu, Yadong Mu, Xinggang Wang, Wenyu Liu, and Jingdong Wang. 2019. High-resolution representations for labeling pixels and regions. arXiv arXiv\/1904.04514 (2019)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_64_1","DOI":"10.1145\/3200691.3178491"},{"unstructured":"Yonghui Wu Mike Schuster Zhifeng Chen Quoc V Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey etal 2016. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv arXiv\/1609.08144 (2016).  Yonghui Wu Mike Schuster Zhifeng Chen Quoc V Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey et al. 2016. Google's neural machine translation system: Bridging the gap between human and machine translation. arXiv arXiv\/1609.08144 (2016).","key":"e_1_3_2_1_65_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_66_1","DOI":"10.1145\/1755913.1755940"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_67_1","DOI":"10.5555\/1863103.1863113"},{"key":"e_1_3_2_1_68_1","volume-title":"Yao Shu, Bingsheng He, and Wei Wang.","author":"Zhang Junzhe","year":"2019","unstructured":"Junzhe Zhang , Sai Ho Yeung , Yao Shu, Bingsheng He, and Wei Wang. 2019 . Efficient memory management for gpu-based deep learning systems. arXiv arXiv\/1903.06631 (2019). Junzhe Zhang, Sai Ho Yeung, Yao Shu, Bingsheng He, and Wei Wang. 2019. Efficient memory management for gpu-based deep learning systems. arXiv arXiv\/1903.06631 (2019)."}],"event":{"sponsor":["SIGOPS ACM Special Interest Group on Operating Systems"],"acronym":"HotOS '21","name":"HotOS '21: Workshop on Hot Topics in Operating Systems","location":"Ann Arbor Michigan"},"container-title":["Proceedings of the Workshop on Hot Topics in Operating Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3458336.3465289","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3458336.3465289","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:28:19Z","timestamp":1750195699000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3458336.3465289"}},"subtitle":["training large DNN models on commodity servers for the masses"],"short-title":[],"issued":{"date-parts":[[2021,6]]},"references-count":68,"alternative-id":["10.1145\/3458336.3465289","10.1145\/3458336"],"URL":"https:\/\/doi.org\/10.1145\/3458336.3465289","relation":{},"subject":[],"published":{"date-parts":[[2021,6]]},"assertion":[{"value":"2021-06-03","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}