{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T15:59:51Z","timestamp":1778255991114,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":37,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,11]],"date-time":"2020-10-11T00:00:00Z","timestamp":1602374400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"HKU 17204619","award":["Hong Kong RGC"],"award-info":[{"award-number":["Hong Kong RGC"]}]},{"name":"C5026-18G (CRF)","award":["Hong Kong RGC"],"award-info":[{"award-number":["Hong Kong RGC"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["2042019kf0016"],"award-info":[{"award-number":["2042019kf0016"]}]},{"name":"HKU 17225516","award":["Hong Kong RGC"],"award-info":[{"award-number":["Hong Kong RGC"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,11]]},"DOI":"10.1145\/3397166.3409128","type":"proceedings-article","created":{"date-parts":[[2020,10,8]],"date-time":"2020-10-08T02:02:42Z","timestamp":1602122562000},"page":"111-120","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["Online scheduling of heterogeneous distributed machine learning jobs"],"prefix":"10.1145","author":[{"given":"Qin","family":"Zhang","sequence":"first","affiliation":[{"name":"Wuhan University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ruiting","family":"Zhou","sequence":"additional","affiliation":[{"name":"Wuhan University, China and The Chinese University of Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chuan","family":"Wu","sequence":"additional","affiliation":[{"name":"The University of Hong Kong, Hong Kong"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Lei","family":"Jiao","sequence":"additional","affiliation":[{"name":"University of Oregon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zongpeng","family":"Li","sequence":"additional","affiliation":[{"name":"Wuhan University, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,10,11]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"[n.d.]. Distributed Training in TensorFlow. https:\/\/www.tensorflow.org\/guide\/distribute_strategy.  [n.d.]. Distributed Training in TensorFlow. https:\/\/www.tensorflow.org\/guide\/distribute_strategy."},{"key":"e_1_3_2_1_2_1","unstructured":"[n.d.]. Microsoft Cognitive Toolkit. https:\/\/www.microsoft.com\/en-us\/cognitive-toolkit\/.  [n.d.]. Microsoft Cognitive Toolkit. https:\/\/www.microsoft.com\/en-us\/cognitive-toolkit\/."},{"key":"e_1_3_2_1_3_1","unstructured":"[n.d.]. Technical report. https:\/\/1drv.ms\/b\/s!AvAD6Lae6eSxa3HIfAJ4xM3dBaU?e=fmPk0B.  [n.d.]. Technical report. https:\/\/1drv.ms\/b\/s!AvAD6Lae6eSxa3HIfAJ4xM3dBaU?e=fmPk0B."},{"key":"e_1_3_2_1_4_1","volume-title":"Proc. of USENIX OSDI.","author":"Abadi Martin","year":"2016","unstructured":"Martin Abadi , Paul Barham , Jianmin Chen , Zhifeng Chen , Andy Davis , Jeffrey Dean , Matthieu Devin , Sanjay Ghemawat , Geoffrey Irving , Michael Isard , Manjunath Kudlur , Josh Levenberg , Rajat Monga , Sherry Moore , Derek G. Murray , Benoit Steiner , Paul Tucker , Vijay Vasudevan , Pete Warden , Martin Wicke , Yuan Yu , and Xiaoqiang Zheng . 2016 . TensorFlow: A System for Large-Scale Machine Learning . In Proc. of USENIX OSDI. Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A System for Large-Scale Machine Learning. In Proc. of USENIX OSDI."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8682911"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/2764468.2764535"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2019.8737460"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2018.8486422"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2018.8486026"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2019.8737587"},{"key":"e_1_3_2_1_11_1","volume-title":"MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. In NIPS Workshop on Machine Learning Systems (LearningSys).","author":"Chen Tianqi","year":"2016","unstructured":"Tianqi Chen , Mu Li , Yutian Li , Min Lin , Naiyan Wang , Minjie Wang , Tianjun Xiao , Bing Xu , Chiyuan Zhang , and Zheng Zhang . 2016 . MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. In NIPS Workshop on Machine Learning Systems (LearningSys). Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2016. MXNet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. In NIPS Workshop on Machine Learning Systems (LearningSys)."},{"key":"e_1_3_2_1_12_1","volume-title":"Proc. of USENIX OSDI.","author":"Chilimbi Trishul","year":"2014","unstructured":"Trishul Chilimbi , Yutaka Suzue , Johnson Apacible , and Karthik Kalyanaraman . 2014 . Project Adam: Building an Efficient and Scalable Deep Learning Training System . In Proc. of USENIX OSDI. Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, and Karthik Kalyanaraman. 2014. Project Adam: Building an Efficient and Scalable Deep Learning Training System. In Proc. of USENIX OSDI."},{"key":"e_1_3_2_1_13_1","volume-title":"Proc. of USENIX NSDI.","author":"Chowdhury Mosharaf","year":"2016","unstructured":"Mosharaf Chowdhury , Zhenhua Liu , Ali Ghodsi , and Ion Stoica . 2016 . HUG: Multi-Resource Fairness for Correlated and Elastic Demands . In Proc. of USENIX NSDI. Mosharaf Chowdhury, Zhenhua Liu, Ali Ghodsi, and Ion Stoica. 2016. HUG: Multi-Resource Fairness for Correlated and Elastic Demands. In Proc. of USENIX NSDI."},{"key":"e_1_3_2_1_14_1","volume-title":"Proc. of ACM ICML.","author":"Gao Yuanxiang","year":"2018","unstructured":"Yuanxiang Gao , Li Chen , and Baochun Li . 2018 . Spotlight: Optimizing Device Placement for Training Deep Neural Networks . In Proc. of ACM ICML. Yuanxiang Gao, Li Chen, and Baochun Li. 2018. Spotlight: Optimizing Device Placement for Training Deep Neural Networks. In Proc. of ACM ICML."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1008861.1008867"},{"key":"e_1_3_2_1_16_1","volume-title":"Proc. of USENIX NSDI.","author":"Ghodsi Ali","year":"2011","unstructured":"Ali Ghodsi , Matei Zaharia , Benjamin Hindman , Andy Konwinski , Scott Shenker , and Ion Stoica . 2011 . Dominant Resource Fairness: Fair Allocation of Multiple Resource Types . In Proc. of USENIX NSDI. Ali Ghodsi, Matei Zaharia, Benjamin Hindman, Andy Konwinski, Scott Shenker, and Ion Stoica. 2011. Dominant Resource Fairness: Fair Allocation of Multiple Resource Types. In Proc. of USENIX NSDI."},{"key":"e_1_3_2_1_17_1","volume-title":"Scheduling to minimize average completion time: Off-line and on-line approximation algorithms. Mathematics of operations research 22, 3","author":"Hall Leslie A.","year":"1997","unstructured":"Leslie A. Hall , Andreas S. Schulz , David B. Shmoys , and Joel Wein . 1997. Scheduling to minimize average completion time: Off-line and on-line approximation algorithms. Mathematics of operations research 22, 3 ( 1997 ), 513--544. Leslie A. Hall, Andreas S. Schulz, David B. Shmoys, and Joel Wein. 1997. Scheduling to minimize average completion time: Off-line and on-line approximation algorithms. Mathematics of operations research 22, 3 (1997), 513--544."},{"key":"e_1_3_2_1_18_1","volume-title":"Proc. of USENIX NSDI.","author":"Hindman Benjamin","year":"2011","unstructured":"Benjamin Hindman , Andy Konwinski , Matei Zaharia , Ali Ghodsi , Anthony D. Joseph , Randy Katz , Scott Shenker , and Ion Stoica . 2011 . Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center . In Proc. of USENIX NSDI. Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, Anthony D. Joseph, Randy Katz, Scott Shenker, and Ion Stoica. 2011. Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center. In Proc. of USENIX NSDI."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.284"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/2742343"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNET.2017.2707142"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/IWQoS.2018.8624119"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2640087.2644155"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/FOCS.2017.34"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3190508.3190517"},{"key":"e_1_3_2_1_26_1","volume-title":"Lau","author":"Shi Weijie","year":"2014","unstructured":"Weijie Shi , Linquan Zhang , Chuan Wu , Zongpeng Li , and Francis C.M . Lau . 2014 . An online auction framework for dynamic resource provisioning in cloud computing. In Proc. of ACM SIGMETRICS. Weijie Shi, Linquan Zhang, Chuan Wu, Zongpeng Li, and Francis C.M. Lau. 2014. An online auction framework for dynamic resource provisioning in cloud computing. In Proc. of ACM SIGMETRICS."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2523616.2523633"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2741948.2741964"},{"key":"e_1_3_2_1_29_1","volume-title":"Proc. of IEEE\/ACM IWQoS.","author":"Wang Tiantian","year":"2020","unstructured":"Tiantian Wang , Zhuzhong Qian , Lei Jiao , Xin Li , and Sanglu Lu . 2020 . Geo-Clone: online task replication and scheduling for geo-distributed analytics under uncertainties . In Proc. of IEEE\/ACM IWQoS. Tiantian Wang, Zhuzhong Qian, Lei Jiao, Xin Li, and Sanglu Lu. 2020. Geo-Clone: online task replication and scheduling for geo-distributed analytics under uncertainties. In Proc. of IEEE\/ACM IWQoS."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/2783258.2783270"},{"key":"e_1_3_2_1_31_1","volume-title":"Proc. of USENIX HotCloud.","author":"Zaharia Matei","year":"2010","unstructured":"Matei Zaharia , Mosharaf Chowdhury , Michael J. Franklin , Scott Shenker , and Ion Stoica . 2010 . Spark: Cluster computing with working sets . In Proc. of USENIX HotCloud. Matei Zaharia, Mosharaf Chowdhury, Michael J. Franklin, Scott Shenker, and Ion Stoica. 2010. Spark: Cluster computing with working sets. In Proc. of USENIX HotCloud."},{"key":"e_1_3_2_1_32_1","volume-title":"Proc. of USENIX ATC.","author":"Zhang Hao","year":"2017","unstructured":"Hao Zhang , Zeyu Zheng , Shizhen Xu , Wei Dai , Qirong Ho , Xiaodan Liang , Zhiting Hu , Jinliang Wei , Pengtao Xie , and Eric P Xing . 2017 . Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters . In Proc. of USENIX ATC. Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P Xing. 2017. Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters. In Proc. of USENIX ATC."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2745844.2745855"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3078505.3078529"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2896377.2901452"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNET.2016.2609844"},{"key":"e_1_3_2_1_37_1","volume-title":"Chang","author":"Zou ShangXuan","year":"2017","unstructured":"ShangXuan Zou , ChunYen Chen , JuiLin Wu , ChunNan Chou , ChiaChin Tsao , KuanChieh Tung , TingWei Lin , ChengLung Sung , and Edward Y . Chang . 2017 . Distributed training large-scale deep architectures. In Proc. of Springer ADMA. ShangXuan Zou, ChunYen Chen, JuiLin Wu, ChunNan Chou, ChiaChin Tsao, KuanChieh Tung, TingWei Lin, ChengLung Sung, and Edward Y. Chang. 2017. Distributed training large-scale deep architectures. In Proc. of Springer ADMA."}],"event":{"name":"Mobihoc '20: The Twenty-first ACM International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing","location":"Virtual Event USA","acronym":"Mobihoc '20","sponsor":["SIGMOBILE ACM Special Interest Group on Mobility of Systems, Users, Data and Computing"]},"container-title":["Proceedings of the Twenty-First International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3397166.3409128","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3397166.3409128","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:41:25Z","timestamp":1750200085000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3397166.3409128"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,11]]},"references-count":37,"alternative-id":["10.1145\/3397166.3409128","10.1145\/3397166"],"URL":"https:\/\/doi.org\/10.1145\/3397166.3409128","relation":{},"subject":[],"published":{"date-parts":[[2020,10,11]]},"assertion":[{"value":"2020-10-11","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}