{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T18:02:16Z","timestamp":1783620136482,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":108,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,1,27]],"date-time":"2023-01-27T00:00:00Z","timestamp":1674777600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,1,27]]},"DOI":"10.1145\/3575693.3575705","type":"proceedings-article","created":{"date-parts":[[2023,1,30]],"date-time":"2023-01-30T22:56:55Z","timestamp":1675119415000},"page":"457-472","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":43,"title":["Lucid: A Non-intrusive, Scalable and Interpretable Scheduler for Deep Learning Training Jobs"],"prefix":"10.1145","author":[{"given":"Qinghao","family":"Hu","sequence":"first","affiliation":[{"name":"Nanyang Technological University, Singapore \/ Shanghai AI Laboratory, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Meng","family":"Zhang","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Peng","family":"Sun","sequence":"additional","affiliation":[{"name":"SenseTime, China \/ Shanghai AI Laboratory, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yonggang","family":"Wen","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tianwei","family":"Zhang","sequence":"additional","affiliation":[{"name":"Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,1,30]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2023. gRPC: An RPC library and framework. https:\/\/grpc.io\/ \t\t\t\t  2023. gRPC: An RPC library and framework. https:\/\/grpc.io\/"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.7275326"},{"key":"e_1_3_2_1_3_1","unstructured":"2023. NVIDIA Data Center GPU Manager. https:\/\/developer.nvidia.com\/dcgm \t\t\t\t  2023. NVIDIA Data Center GPU Manager. https:\/\/developer.nvidia.com\/dcgm"},{"key":"e_1_3_2_1_4_1","unstructured":"2023. NVIDIA Multi-Instance GPU. https:\/\/www.nvidia.com\/en-us\/technologies\/multi-instance-gpu\/ \t\t\t\t  2023. NVIDIA Multi-Instance GPU. https:\/\/www.nvidia.com\/en-us\/technologies\/multi-instance-gpu\/"},{"key":"e_1_3_2_1_5_1","unstructured":"2023. NVIDIA Multi-Process Service. https:\/\/docs.nvidia.com\/deploy\/mps\/index.html \t\t\t\t  2023. NVIDIA Multi-Process Service. https:\/\/docs.nvidia.com\/deploy\/mps\/index.html"},{"key":"e_1_3_2_1_6_1","unstructured":"2023. NVIDIA-smi. https:\/\/developer.nvidia.com\/nvidia-system-management-interface \t\t\t\t  2023. NVIDIA-smi. https:\/\/developer.nvidia.com\/nvidia-system-management-interface"},{"key":"e_1_3_2_1_7_1","unstructured":"2023. ONNX: Open Neural Network Exchange. https:\/\/github.com\/onnx\/onnx \t\t\t\t  2023. ONNX: Open Neural Network Exchange. https:\/\/github.com\/onnx\/onnx"},{"key":"e_1_3_2_1_8_1","volume-title":"An Empirical Distribution Function for Sampling with Incomplete Information. The Annals of Mathematical Statistics, 26","author":"Ayer Miriam","year":"1955","unstructured":"Miriam Ayer , H. D. Brunk , G. M. Ewing , W. T. Reid , and Edward Silverman . 1955. An Empirical Distribution Function for Sampling with Incomplete Information. The Annals of Mathematical Statistics, 26 ( 1955 ). Miriam Ayer, H. D. Brunk, G. M. Ewing, W. T. Reid, and Edward Silverman. 1955. An Empirical Distribution Function for Sampling with Incomplete Information. The Annals of Mathematical Statistics, 26 (1955)."},{"key":"e_1_3_2_1_9_1","volume-title":"3rd International Conference on Learning Representations (ICLR \u201915)","author":"Bahdanau Dzmitry","year":"2015","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2015 . Neural Machine Translation by Jointly Learning to Align and Translate . In 3rd International Conference on Learning Representations (ICLR \u201915) . Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015. Neural Machine Translation by Jointly Learning to Align and Translate. In 3rd International Conference on Learning Representations (ICLR \u201915)."},{"key":"e_1_3_2_1_10_1","volume-title":"Deep Learning-based Job Placement in Distributed Machine Learning Clusters. In IEEE Conference on Computer Communications (INFOCOM \u201919)","author":"Bao Yixin","year":"2019","unstructured":"Yixin Bao , Yanghua Peng , and Chuan Wu . 2019 . Deep Learning-based Job Placement in Distributed Machine Learning Clusters. In IEEE Conference on Computer Communications (INFOCOM \u201919) . Yixin Bao, Yanghua Peng, and Chuan Wu. 2019. Deep Learning-based Job Placement in Distributed Machine Learning Clusters. In IEEE Conference on Computer Communications (INFOCOM \u201919)."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3480859"},{"key":"e_1_3_2_1_12_1","volume-title":"Apollo: Scalable and Coordinated Scheduling for Cloud-Scale Computing. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201914)","author":"Boutin Eric","year":"2014","unstructured":"Eric Boutin , Jaliya Ekanayake , Wei Lin , Bing Shi , Jingren Zhou , Zhengping Qian , Ming Wu , and Lidong Zhou . 2014 . Apollo: Scalable and Coordinated Scheduling for Cloud-Scale Computing. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201914) . Eric Boutin, Jaliya Ekanayake, Wei Lin, Bing Shi, Jingren Zhou, Zhengping Qian, Ming Wu, and Lidong Zhou. 2014. Apollo: Scalable and Coordinated Scheduling for Cloud-Scale Computing. In 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201914)."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Leo Breiman. 2001. Random Forests. Machine learning 5\u201332. \t\t\t\t  Leo Breiman. 2001. Random Forests. Machine learning 5\u201332.","DOI":"10.1023\/A:1010933404324"},{"key":"e_1_3_2_1_14_1","unstructured":"Leo Breiman Jerome H Friedman Richard A Olshen and Charles J Stone. 1984. Classification and regression trees. Wadsworth. \t\t\t\t  Leo Breiman Jerome H Friedman Richard A Olshen and Charles J Stone. 1984. Classification and regression trees. Wadsworth."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2898442.2898444"},{"key":"e_1_3_2_1_16_1","volume-title":"ShapeNet: An Information-Rich 3D Model Repository. CoRR, abs\/1512.03012","author":"Chang Angel X.","year":"2015","unstructured":"Angel X. Chang , Thomas Funkhouser , Leonidas Guibas , Pat Hanrahan , Qixing Huang , Zimo Li , Silvio Savarese , Manolis Savva , Shuran Song , Hao Su , Jianxiong Xiao , Li Yi , and Fisher Yu. 2015. ShapeNet: An Information-Rich 3D Model Repository. CoRR, abs\/1512.03012 ( 2015 ). Angel X. Chang, Thomas Funkhouser, Leonidas Guibas, Pat Hanrahan, Qixing Huang, Zimo Li, Silvio Savarese, Manolis Savva, Shuran Song, Hao Su, Jianxiong Xiao, Li Yi, and Fisher Yu. 2015. ShapeNet: An Information-Rich 3D Model Repository. CoRR, abs\/1512.03012 (2015)."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3342195.3387555"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_2_1_19_1","volume-title":"Poster Abstract: Deep Learning Workloads Scheduling with Reinforcement Learning on GPU Clusters. In IEEE Conference on Computer Communications Workshops (INFOCOM \u201919)","author":"Chen Zhaoyun","year":"2019","unstructured":"Zhaoyun Chen , Lei Luo , Wei Quan , Mei Wen , and Chunyuan Zhang . 2019 . Poster Abstract: Deep Learning Workloads Scheduling with Reinforcement Learning on GPU Clusters. In IEEE Conference on Computer Communications Workshops (INFOCOM \u201919) . Zhaoyun Chen, Lei Luo, Wei Quan, Mei Wen, and Chunyuan Zhang. 2019. Poster Abstract: Deep Learning Workloads Scheduling with Reinforcement Learning on GPU Clusters. In IEEE Conference on Computer Communications Workshops (INFOCOM \u201919)."},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3132747.3132772"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2670979.2670981"},{"key":"e_1_3_2_1_22_1","volume-title":"Hawk: Hybrid Datacenter Scheduling. In 2015 USENIX Annual Technical Conference (USENIX ATC \u201915)","author":"Delgado Pamela","year":"2015","unstructured":"Pamela Delgado , Florin Dinu , Anne-Marie Kermarrec , and Willy Zwaenepoel . 2015 . Hawk: Hybrid Datacenter Scheduling. In 2015 USENIX Annual Technical Conference (USENIX ATC \u201915) . Pamela Delgado, Florin Dinu, Anne-Marie Kermarrec, and Willy Zwaenepoel. 2015. Hawk: Hybrid Datacenter Scheduling. In 2015 USENIX Annual Technical Conference (USENIX ATC \u201915)."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_1_24_1","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL \u201919)","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL \u201919) . Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL \u201919)."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W16-3210"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2168836.2168847"},{"key":"e_1_3_2_1_27_1","volume-title":"Frey and Delbert Dueck","author":"Brendan","year":"2007","unstructured":"Brendan J. Frey and Delbert Dueck . 2007 . Clustering by Passing Messages Between Data Points. Science . Brendan J. Frey and Delbert Dueck. 2007. Clustering by Passing Messages Between Data Points. Science."},{"key":"e_1_3_2_1_28_1","volume-title":"Proceedings of the 35th International Conference on Machine Learning (ICML \u201918)","author":"Fujimoto Scott","year":"2018","unstructured":"Scott Fujimoto , Herke van Hoof , and David Meger . 2018 . Addressing Function Approximation Error in Actor-Critic Methods . In Proceedings of the 35th International Conference on Machine Learning (ICML \u201918) . Scott Fujimoto, Herke van Hoof, and David Meger. 2018. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning (ICML \u201918)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446700"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3472883.3486978"},{"key":"e_1_3_2_1_31_1","volume-title":"Tiresias: A GPU Cluster Manager for Distributed Deep Learning. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201919)","author":"Gu Juncheng","year":"2019","unstructured":"Juncheng Gu , Mosharaf Chowdhury , Kang G. Shin , Yibo Zhu , Myeongjae Jeon , Junjie Qian , Hongqiang Liu , and Chuanxiong Guo . 2019 . Tiresias: A GPU Cluster Manager for Distributed Deep Learning. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201919) . Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo. 2019. Tiresias: A GPU Cluster Manager for Distributed Deep Learning. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201919)."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3138825"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3243734.3243792"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3460120.3484589"},{"key":"e_1_3_2_1_35_1","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201920)","author":"Hao Mingzhe","unstructured":"Mingzhe Hao , Levent Toksoz , Nanqinqin Li , Edward Edberg Halim , Henry Hoffmann , and Haryadi S. Gunawi . 2020. LinnOS: Predictability on Unpredictable Flash Storage with a Light Neural Network . In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201920) . Mingzhe Hao, Levent Toksoz, Nanqinqin Li, Edward Edberg Halim, Henry Hoffmann, and Haryadi S. Gunawi. 2020. LinnOS: Predictability on Unpredictable Flash Storage with a Light Neural Network. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201920)."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/2827872"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3038912.3052569"},{"key":"e_1_3_2_1_39_1","volume-title":"Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center. In 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201911)","author":"Hindman Benjamin","year":"2011","unstructured":"Benjamin Hindman , Andy Konwinski , Matei Zaharia , Ali Ghodsi , Anthony D. Joseph , Randy Katz , Scott Shenker , and Ion Stoica . 2011 . Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center. In 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201911) . Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, Anthony D. Joseph, Randy Katz, Scott Shenker, and Ion Stoica. 2011. Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center. In 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201911)."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_3_2_1_41_1","volume-title":"Primo: Practical Learning-Augmented Systems with Interpretable Models. In 2022 USENIX Annual Technical Conference (USENIX ATC \u201922)","author":"Hu Qinghao","year":"2022","unstructured":"Qinghao Hu , Harsha Nori , Peng Sun , Yonggang Wen , and Tianwei Zhang . 2022 . Primo: Practical Learning-Augmented Systems with Interpretable Models. In 2022 USENIX Annual Technical Conference (USENIX ATC \u201922) . Qinghao Hu, Harsha Nori, Peng Sun, Yonggang Wen, and Tianwei Zhang. 2022. Primo: Practical Learning-Augmented Systems with Interpretable Models. In 2022 USENIX Annual Technical Conference (USENIX ATC \u201922)."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476223"},{"key":"e_1_3_2_1_43_1","volume-title":"NVIDIA GTC 2022 KEYNOTE. https:\/\/www.nvidia.com\/gtc\/keynote\/","author":"Huang Jensen","year":"2023","unstructured":"Jensen Huang . 2023 . NVIDIA GTC 2022 KEYNOTE. https:\/\/www.nvidia.com\/gtc\/keynote\/ Jensen Huang. 2023. NVIDIA GTC 2022 KEYNOTE. https:\/\/www.nvidia.com\/gtc\/keynote\/"},{"key":"e_1_3_2_1_44_1","volume-title":"Elastic Resource Sharing for Distributed Deep Learning. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201921)","author":"Hwang Changho","year":"2021","unstructured":"Changho Hwang , Taehyun Kim , Sunghyun Kim , Jinwoo Shin , and KyoungSoo Park . 2021 . Elastic Resource Sharing for Distributed Deep Learning. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201921) . Changho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin, and KyoungSoo Park. 2021. Elastic Resource Sharing for Distributed Deep Learning. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201921)."},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3492321.3519575"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/1272996.1273005"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2785956.2787488"},{"key":"e_1_3_2_1_48_1","volume-title":"Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. In 2019 USENIX Annual Technical Conference (USENIX ATC \u201919)","author":"Jeon Myeongjae","year":"2019","unstructured":"Myeongjae Jeon , Shivaram Venkataraman , Amar Phanishayee , Junjie Qian , Wencong Xiao , and Fan Yang . 2019 . Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. In 2019 USENIX Annual Technical Conference (USENIX ATC \u201919) . Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, and Fan Yang. 2019. Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. In 2019 USENIX Annual Technical Conference (USENIX ATC \u201919)."},{"key":"e_1_3_2_1_49_1","volume-title":"Morpheus: Towards Automated SLOs for Enterprise Clusters. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201916)","author":"Jyothi Sangeetha Abdu","year":"2016","unstructured":"Sangeetha Abdu Jyothi , Carlo Curino , Ishai Menache , Shravan Matthur Narayanamurthy , Alexey Tumanov , Jonathan Yaniv , Ruslan Mavlyutov , Inigo Goiri , Subru Krishnan , Janardhan Kulkarni , and Sriram Rao . 2016 . Morpheus: Towards Automated SLOs for Enterprise Clusters. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201916) . Sangeetha Abdu Jyothi, Carlo Curino, Ishai Menache, Shravan Matthur Narayanamurthy, Alexey Tumanov, Jonathan Yaniv, Ruslan Mavlyutov, Inigo Goiri, Subru Krishnan, Janardhan Kulkarni, and Sriram Rao. 2016. Morpheus: Towards Automated SLOs for Enterprise Clusters. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201916)."},{"key":"e_1_3_2_1_50_1","unstructured":"Guolin Ke Qi Meng Thomas Finley Taifeng Wang Wei Chen Weidong Ma Qiwei Ye and Tie-Yan Liu. 2017. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Advances in Neural Information Processing Systems (NeurIPS \u201917). \t\t\t\t  Guolin Ke Qi Meng Thomas Finley Taifeng Wang Wei Chen Weidong Ma Qiwei Ye and Tie-Yan Liu. 2017. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Advances in Neural Information Processing Systems (NeurIPS \u201917)."},{"key":"e_1_3_2_1_51_1","volume-title":"On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. In International Conference on Learning Representations (ICLR \u201917)","author":"Keskar Nitish Shirish","year":"2017","unstructured":"Nitish Shirish Keskar , Dheevatsa Mudigere , Jorge Nocedal , Mikhail Smelyanskiy , and Ping Tak Peter Tang . 2017 . On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. In International Conference on Learning Representations (ICLR \u201917) . Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2017. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. In International Conference on Learning Representations (ICLR \u201917)."},{"key":"e_1_3_2_1_52_1","volume-title":"Co-scheML: Interference-aware Container Co-scheduling Scheme Using Machine Learning Application Profiles for GPU Clusters. In 2020 IEEE International Conference on Cluster Computing (CLUSTER \u201920)","author":"Kim Sejin","year":"2020","unstructured":"Sejin Kim and Yoonhee Kim . 2020 . Co-scheML: Interference-aware Container Co-scheduling Scheme Using Machine Learning Application Profiles for GPU Clusters. In 2020 IEEE International Conference on Cluster Computing (CLUSTER \u201920) . Sejin Kim and Yoonhee Kim. 2020. Co-scheML: Interference-aware Container Co-scheduling Scheme Using Machine Learning Application Profiles for GPU Clusters. In 2020 IEEE International Conference on Cluster Computing (CLUSTER \u201920)."},{"key":"e_1_3_2_1_53_1","unstructured":"Alex Krizhevsky. 2023. The CIFAR10 Dataset. https:\/\/www.cs.toronto.edu\/ kriz\/cifar.html \t\t\t\t  Alex Krizhevsky. 2023. The CIFAR10 Dataset. https:\/\/www.cs.toronto.edu\/ kriz\/cifar.html"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304028"},{"key":"e_1_3_2_1_55_1","volume-title":"Proceedings of the 38th International Conference on Machine Learning (ICML \u201921)","author":"Landajuela Mikel","year":"2021","unstructured":"Mikel Landajuela , Brenden K Petersen , Sookyung Kim , Claudio P Santiago , Ruben Glatt , Nathan Mundhenk , Jacob F Pettit , and Daniel Faissol . 2021 . Discovering symbolic policies with deep reinforcement learning . In Proceedings of the 38th International Conference on Machine Learning (ICML \u201921) . Mikel Landajuela, Brenden K Petersen, Sookyung Kim, Claudio P Santiago, Ruben Glatt, Nathan Mundhenk, Jacob F Pettit, and Daniel Faissol. 2021. Discovering symbolic policies with deep reinforcement learning. In Proceedings of the 38th International Conference on Machine Learning (ICML \u201921)."},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1038\/nature14539"},{"key":"e_1_3_2_1_57_1","volume-title":"Aryl: An Elastic Cluster Scheduler for Deep Learning. CoRR, abs\/2202.07896","author":"Li Jiamin","year":"2022","unstructured":"Jiamin Li , Hong Xu , Yibo Zhu , Zherui Liu , Chuanxiong Guo , and Cong Wang . 2022 . Aryl: An Elastic Cluster Scheduler for Deep Learning. CoRR, abs\/2202.07896 (2022). Jiamin Li, Hong Xu, Yibo Zhu, Zherui Liu, Chuanxiong Guo, and Cong Wang. 2022. Aryl: An Elastic Cluster Scheduler for Deep Learning. CoRR, abs\/2202.07896 (2022)."},{"key":"e_1_3_2_1_58_1","volume-title":"Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS \u201921)","author":"Li Rui","unstructured":"Rui Li , Yufan Xu , Aravind Sukumaran-Rajam , Atanas Rountev , and P. Sadayappan . 2021. Analytical Characterization and Design Space Exploration for Optimization of CNNs . In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS \u201921) . Rui Li, Yufan Xu, Aravind Sukumaran-Rajam, Atanas Rountev, and P. Sadayappan. 2021. Analytical Characterization and Design Space Exploration for Optimization of CNNs. In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS \u201921)."},{"key":"e_1_3_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/2487575.2487579"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/GLOBECOM38437.2019.9014110"},{"key":"e_1_3_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378499"},{"key":"e_1_3_2_1_62_1","volume-title":"Themis: Fair and Efficient GPU Cluster Scheduling. In 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201920)","author":"Mahajan Kshiteej","year":"2020","unstructured":"Kshiteej Mahajan , Arjun Balasubramanian , Arjun Singhvi , Shivaram Venkataraman , Aditya Akella , Amar Phanishayee , and Shuchi Chawla . 2020 . Themis: Fair and Efficient GPU Cluster Scheduling. In 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201920) . Kshiteej Mahajan, Arjun Balasubramanian, Arjun Singhvi, Shivaram Venkataraman, Aditya Akella, Amar Phanishayee, and Shuchi Chawla. 2020. Themis: Fair and Efficient GPU Cluster Scheduling. In 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201920)."},{"key":"e_1_3_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3387514.3405859"},{"key":"e_1_3_2_1_64_1","volume-title":"Pointer Sentinel Mixture Models. In International Conference on Learning Representations (ICLR \u201917)","author":"Merity Stephen","year":"2017","unstructured":"Stephen Merity , Caiming Xiong , James Bradbury , and Richard Socher . 2017 . Pointer Sentinel Mixture Models. In International Conference on Learning Representations (ICLR \u201917) . Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer Sentinel Mixture Models. In International Conference on Learning Representations (ICLR \u201917)."},{"key":"e_1_3_2_1_65_1","volume-title":"Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201922)","author":"Mohan Jayashree","year":"2022","unstructured":"Jayashree Mohan , Amar Phanishayee , Janardhan Kulkarni , and Vijay Chidambaram . 2022 . Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201922) . Jayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, and Vijay Chidambaram. 2022. Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201922)."},{"key":"e_1_3_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3477132.3483588"},{"key":"e_1_3_2_1_67_1","volume-title":"Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201920)","author":"Narayanan Deepak","year":"2020","unstructured":"Deepak Narayanan , Keshav Santhanam , Fiodar Kazhamiaka , Amar Phanishayee , and Matei Zaharia . 2020 . Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201920) . Deepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee, and Matei Zaharia. 2020. Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201920)."},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/375360.375365"},{"key":"e_1_3_2_1_69_1","volume-title":"Proceedings of the 38th International Conference on Machine Learning (ICML \u201921)","author":"Nori Harsha","year":"2021","unstructured":"Harsha Nori , Rich Caruana , Zhiqi Bu , Judy Hanwen Shen , and Janardhan Kulkarni . 2021 . Accuracy, Interpretability, and Differential Privacy via Explainable Boosting . In Proceedings of the 38th International Conference on Machine Learning (ICML \u201921) . Harsha Nori, Rich Caruana, Zhiqi Bu, Judy Hanwen Shen, and Janardhan Kulkarni. 2021. Accuracy, Interpretability, and Differential Privacy via Explainable Boosting. In Proceedings of the 38th International Conference on Machine Learning (ICML \u201921)."},{"key":"e_1_3_2_1_70_1","volume-title":"Proceedings of the Thirteenth EuroSys Conference (EuroSys \u201918)","author":"Park Jun Woo","unstructured":"Jun Woo Park , Alexey Tumanov , Angela Jiang , Michael A. Kozuch , and Gregory R. Ganger . 2018. 3Sigma: Distribution-Based Cluster Scheduling for Runtime Uncertainty . In Proceedings of the Thirteenth EuroSys Conference (EuroSys \u201918) . Jun Woo Park, Alexey Tumanov, Angela Jiang, Michael A. Kozuch, and Gregory R. Ganger. 2018. 3Sigma: Distribution-Based Cluster Scheduling for Runtime Uncertainty. In Proceedings of the Thirteenth EuroSys Conference (EuroSys \u201918)."},{"key":"e_1_3_2_1_71_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas Kopf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In Advances in Neural Information Processing Systems (NeurIPS \u201919). Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems (NeurIPS \u201919)."},{"key":"e_1_3_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378505"},{"key":"e_1_3_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3190508.3190517"},{"key":"e_1_3_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3052895"},{"key":"e_1_3_2_1_75_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)","author":"Qi Charles R.","unstructured":"Charles R. Qi , Hao Su , Kaichun Mo , and Leonidas J. Guibas . 2017. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917) . Charles R. Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017. PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201917)."},{"key":"e_1_3_2_1_76_1","volume-title":"Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201921)","author":"Qiao Aurick","unstructured":"Aurick Qiao , Sang Keun Choe , Suhas Jayaram Subramanya , Willie Neiswanger , Qirong Ho , Hao Zhang , Gregory R. Ganger , and Eric P. Xing . 2021 . Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201921) . Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing. 2021. Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201921)."},{"key":"e_1_3_2_1_77_1","volume-title":"Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks. In International Conference on Learning Representations (ICLR \u201916)","author":"Radford Alec","year":"2016","unstructured":"Alec Radford , Luke Metz , and Soumith Chintala . 2016 . Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks. In International Conference on Learning Representations (ICLR \u201916) . Alec Radford, Luke Metz, and Soumith Chintala. 2016. Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks. In International Conference on Learning Representations (ICLR \u201916)."},{"key":"e_1_3_2_1_78_1","volume-title":"100,000+ Questions for Machine Comprehension of Text. CoRR, abs\/1606.05250","author":"Rajpurkar Pranav","year":"2016","unstructured":"Pranav Rajpurkar , Jian Zhang , Konstantin Lopyrev , and Percy Liang . 2016. SQuAD : 100,000+ Questions for Machine Comprehension of Text. CoRR, abs\/1606.05250 ( 2016 ). Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. SQuAD: 100,000+ Questions for Machine Comprehension of Text. CoRR, abs\/1606.05250 (2016)."},{"key":"e_1_3_2_1_79_1","volume-title":"Proceedings of the ACM Symposium on Cloud Computing (SoCC \u201912)","author":"Reiss Charles","unstructured":"Charles Reiss , Alexey Tumanov , Gregory R. Ganger , Randy H. Katz , and Michael A. Kozuch . 2012. Heterogeneity and Dynamicity of Clouds at Scale: Google Trace Analysis . In Proceedings of the ACM Symposium on Cloud Computing (SoCC \u201912) . Charles Reiss, Alexey Tumanov, Gregory R. Ganger, Randy H. Katz, and Michael A. Kozuch. 2012. Heterogeneity and Dynamicity of Clouds at Scale: Google Trace Analysis. In Proceedings of the ACM Symposium on Cloud Computing (SoCC \u201912)."},{"key":"e_1_3_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939778"},{"key":"e_1_3_2_1_81_1","doi-asserted-by":"crossref","unstructured":"Cynthia Rudin. 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence 206\u2013215. \t\t\t\t  Cynthia Rudin. 2019. Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence 206\u2013215.","DOI":"10.1038\/s42256-019-0048-x"},{"key":"e_1_3_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_2_1_83_1","volume-title":"Proximal Policy Optimization Algorithms. CoRR, abs\/1707.06347","author":"Schulman John","year":"2017","unstructured":"John Schulman , Filip Wolski , Prafulla Dhariwal , Alec Radford , and Oleg Klimov . 2017. Proximal Policy Optimization Algorithms. CoRR, abs\/1707.06347 ( 2017 ). John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017. Proximal Policy Optimization Algorithms. CoRR, abs\/1707.06347 (2017)."},{"key":"e_1_3_2_1_84_1","volume-title":"Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads. CoRR, abs\/2202.07848","author":"Shukla Dharma","year":"2022","unstructured":"Dharma Shukla , Muthian Sivathanu , Srinidhi Viswanatha , Bhargav Gulavani , Rimma Nehme , Amey Agrawal , Chen Chen , Nipun Kwatra , Ramachandran Ramjee , Pankaj Sharma , Atul Katiyar , Vipul Modi , Vaibhav Sharma , Abhishek Singh , Shreshth Singhal , Kaustubh Welankar , Lu Xun , Ravi Anupindi , Karthik Elangovan , Hasibur Rahman , Zhou Lin , Rahul Seetharaman , Cheng Xu , Eddie Ailijiang , Suresh Krishnappa , and Mark Russinovich . 2022 . Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads. CoRR, abs\/2202.07848 (2022). Dharma Shukla, Muthian Sivathanu, Srinidhi Viswanatha, Bhargav Gulavani, Rimma Nehme, Amey Agrawal, Chen Chen, Nipun Kwatra, Ramachandran Ramjee, Pankaj Sharma, Atul Katiyar, Vipul Modi, Vaibhav Sharma, Abhishek Singh, Shreshth Singhal, Kaustubh Welankar, Lu Xun, Ravi Anupindi, Karthik Elangovan, Hasibur Rahman, Zhou Lin, Rahul Seetharaman, Cheng Xu, Eddie Ailijiang, Suresh Krishnappa, and Mark Russinovich. 2022. Singularity: Planet-Scale, Preemptive and Elastic Scheduling of AI Workloads. CoRR, abs\/2202.07848 (2022)."},{"key":"e_1_3_2_1_85_1","volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations (ICLR \u201915)","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman . 2015 . Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations (ICLR \u201915) . Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations (ICLR \u201915)."},{"key":"e_1_3_2_1_86_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning (ICML \u201919)","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le . 2019 . EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks . In Proceedings of the 36th International Conference on Machine Learning (ICML \u201919) . Mingxing Tan and Quoc Le. 2019. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning (ICML \u201919)."},{"key":"e_1_3_2_1_87_1","volume-title":"Michael A. Kozuch, and Gregory R. Ganger.","author":"Tumanov Alexey","year":"2016","unstructured":"Alexey Tumanov , Angela Jiang , Jun Woo Park , Michael A. Kozuch, and Gregory R. Ganger. 2016 . JamaisVu: Robust Scheduling with Auto-Estimated Job Runtimes. Carnegie Mellon University . Alexey Tumanov, Angela Jiang, Jun Woo Park, Michael A. Kozuch, and Gregory R. Ganger. 2016. JamaisVu: Robust Scheduling with Auto-Estimated Job Runtimes. Carnegie Mellon University."},{"key":"e_1_3_2_1_88_1","volume-title":"\u0141 ukasz Kaiser, and Illia Polosukhin","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , \u0141 ukasz Kaiser, and Illia Polosukhin . 2017 . Attention is All you Need. In Advances in Neural Information Processing Systems (NeurIPS \u201917). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141 ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All you Need. In Advances in Neural Information Processing Systems (NeurIPS \u201917)."},{"key":"e_1_3_2_1_89_1","doi-asserted-by":"publisher","DOI":"10.1145\/2523616.2523633"},{"key":"e_1_3_2_1_90_1","volume-title":"Ernest: Efficient Performance Prediction for Large-Scale Advanced Analytics. In 13th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201916)","author":"Venkataraman Shivaram","year":"2016","unstructured":"Shivaram Venkataraman , Zongheng Yang , Michael Franklin , Benjamin Recht , and Ion Stoica . 2016 . Ernest: Efficient Performance Prediction for Large-Scale Advanced Analytics. In 13th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201916) . Shivaram Venkataraman, Zongheng Yang, Michael Franklin, Benjamin Recht, and Ion Stoica. 2016. Ernest: Efficient Performance Prediction for Large-Scale Advanced Analytics. In 13th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201916)."},{"key":"e_1_3_2_1_91_1","volume-title":"Proceedings of the 8th ACM International Conference on Autonomic Computing (ICAC \u201911)","author":"Verma Abhishek","unstructured":"Abhishek Verma , Ludmila Cherkasova , and Roy H. Campbell . 2011. ARIA: Automatic Resource Inference and Allocation for Mapreduce Environments . In Proceedings of the 8th ACM International Conference on Autonomic Computing (ICAC \u201911) . Abhishek Verma, Ludmila Cherkasova, and Roy H. Campbell. 2011. ARIA: Automatic Resource Inference and Allocation for Mapreduce Environments. In Proceedings of the 8th ACM International Conference on Autonomic Computing (ICAC \u201911)."},{"key":"e_1_3_2_1_92_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386367.3432588"},{"key":"e_1_3_2_1_93_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC41405.2020.00072"},{"key":"e_1_3_2_1_94_1","doi-asserted-by":"publisher","DOI":"10.5555\/3433701.3433820"},{"key":"e_1_3_2_1_95_1","volume-title":"MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201922)","author":"Weng Qizhen","year":"2022","unstructured":"Qizhen Weng , Wencong Xiao , Yinghao Yu , Wei Wang , Cheng Wang , Jian He , Yong Li , Liping Zhang , Wei Lin , and Yu Ding . 2022 . MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201922) . Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022. MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201922)."},{"key":"e_1_3_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446763"},{"key":"e_1_3_2_1_97_1","volume-title":"Gandiva: Introspective Cluster Scheduling for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201918)","author":"Xiao Wencong","year":"2018","unstructured":"Wencong Xiao , Romil Bhardwaj , Ramachandran Ramjee , Muthian Sivathanu , Nipun Kwatra , Zhenhua Han , Pratyush Patel , Xuan Peng , Hanyu Zhao , Quanlu Zhang , Fan Yang , and Lidong Zhou . 2018 . Gandiva: Introspective Cluster Scheduling for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201918) . Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, Fan Yang, and Lidong Zhou. 2018. Gandiva: Introspective Cluster Scheduling for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201918)."},{"key":"e_1_3_2_1_98_1","volume-title":"AntMan: Dynamic Scaling on GPU Clusters for Deep Learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201920)","author":"Xiao Wencong","year":"2020","unstructured":"Wencong Xiao , Shiru Ren , Yong Li , Yang Zhang , Pengyang Hou , Zhi Li , Yihui Feng , Wei Lin , and Yangqing Jia . 2020 . AntMan: Dynamic Scaling on GPU Clusters for Deep Learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201920) . Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020. AntMan: Dynamic Scaling on GPU Clusters for Deep Learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI \u201920)."},{"key":"e_1_3_2_1_99_1","article-title":"ASTRAEA: A Fair Deep Learning Scheduler for Multi-tenant GPU Clusters","author":"Ye Zhisheng","year":"2021","unstructured":"Zhisheng Ye , Peng Sun , Wei Gao , Tianwei Zhang , Xiaolin Wang , Shengen Yan , and Yingwei Luo . 2021 . ASTRAEA: A Fair Deep Learning Scheduler for Multi-tenant GPU Clusters . IEEE Transactions on Parallel and Distributed Systems. Zhisheng Ye, Peng Sun, Wei Gao, Tianwei Zhang, Xiaolin Wang, Shengen Yan, and Yingwei Luo. 2021. ASTRAEA: A Fair Deep Learning Scheduler for Multi-tenant GPU Clusters. IEEE Transactions on Parallel and Distributed Systems.","journal-title":"IEEE Transactions on Parallel and Distributed Systems."},{"key":"e_1_3_2_1_100_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3079202"},{"key":"e_1_3_2_1_101_1","volume-title":"SLURM: Simple Linux Utility for Resource Management. In Job Scheduling Strategies for Parallel Processing.","author":"Yoo Andy B.","year":"2003","unstructured":"Andy B. Yoo , Morris A. Jette , and Mark Grondona . 2003 . SLURM: Simple Linux Utility for Resource Management. In Job Scheduling Strategies for Parallel Processing. Andy B. Yoo, Morris A. Jette, and Mark Grondona. 2003. SLURM: Simple Linux Utility for Resource Management. In Job Scheduling Strategies for Parallel Processing."},{"key":"e_1_3_2_1_102_1","volume-title":"LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop. CoRR, abs\/1506.03365","author":"Yu Fisher","year":"2016","unstructured":"Fisher Yu , Ari Seff , Yinda Zhang , Shuran Song , Thomas Funkhouser , and Jianxiong Xiao . 2016 . LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop. CoRR, abs\/1506.03365 (2016). Fisher Yu, Ari Seff, Yinda Zhang, Shuran Song, Thomas Funkhouser, and Jianxiong Xiao. 2016. LSUN: Construction of a Large-scale Image Dataset using Deep Learning with Humans in the Loop. CoRR, abs\/1506.03365 (2016)."},{"key":"e_1_3_2_1_103_1","volume-title":"Proceedings of Machine Learning and Systems (MLSys \u201920)","author":"Yu Peifeng","year":"2020","unstructured":"Peifeng Yu and Mosharaf Chowdhury . 2020 . Fine-Grained GPU Sharing Primitives for Deep Learning Applications . In Proceedings of Machine Learning and Systems (MLSys \u201920) . Peifeng Yu and Mosharaf Chowdhury. 2020. Fine-Grained GPU Sharing Primitives for Deep Learning Applications. In Proceedings of Machine Learning and Systems (MLSys \u201920)."},{"key":"e_1_3_2_1_104_1","volume-title":"Proceedings of Machine Learning and Systems (MLSys \u201921)","author":"Yu Peifeng","year":"2021","unstructured":"Peifeng Yu , Jiachen Liu , and Mosharaf Chowdhury . 2021 . Fluid: Resource-aware Hyperparameter Tuning Engine . In Proceedings of Machine Learning and Systems (MLSys \u201921) . Peifeng Yu, Jiachen Liu, and Mosharaf Chowdhury. 2021. Fluid: Resource-aware Hyperparameter Tuning Engine. In Proceedings of Machine Learning and Systems (MLSys \u201921)."},{"key":"e_1_3_2_1_105_1","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446693"},{"key":"e_1_3_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDCS47774.2020.00069"},{"key":"e_1_3_2_1_107_1","doi-asserted-by":"publisher","DOI":"10.1145\/3544216.3544224"},{"key":"e_1_3_2_1_108_1","volume-title":"Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201923)","author":"Zheng Pengfei","year":"2023","unstructured":"Pengfei Zheng , Rui Pan , Tarannum Khan , Shivaram Venkataraman , and Aditya Akella . 2023 . Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201923) . Pengfei Zheng, Rui Pan, Tarannum Khan, Shivaram Venkataraman, and Aditya Akella. 2023. Shockwave: Fair and Efficient Cluster Scheduling for Dynamic Adaptation in Machine Learning. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI \u201923)."}],"event":{"name":"ASPLOS '23: 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2","location":"Vancouver BC Canada","acronym":"ASPLOS '23","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture","SIGOPS ACM Special Interest Group on Operating Systems","SIGPLAN ACM Special Interest Group on Programming Languages"]},"container-title":["Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575705","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3575693.3575705","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T18:43:52Z","timestamp":1750272232000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3575693.3575705"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,27]]},"references-count":108,"alternative-id":["10.1145\/3575693.3575705","10.1145\/3575693"],"URL":"https:\/\/doi.org\/10.1145\/3575693.3575705","relation":{},"subject":[],"published":{"date-parts":[[2023,1,27]]},"assertion":[{"value":"2023-01-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}