{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:22:21Z","timestamp":1750220541647,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":66,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,4,19]],"date-time":"2021-04-19T00:00:00Z","timestamp":1618790400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,4,19]]},"DOI":"10.1145\/3447545.3451185","type":"proceedings-article","created":{"date-parts":[[2021,4,11]],"date-time":"2021-04-11T16:51:45Z","timestamp":1618159905000},"page":"57-63","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["GradeML: Towards Holistic Performance Analysis for Machine Learning Workflows"],"prefix":"10.1145","author":[{"given":"Tim","family":"Hegeman","sequence":"first","affiliation":[{"name":"Vrije Universiteit Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matthijs","family":"Jansen","sequence":"additional","affiliation":[{"name":"Vrije Universiteit Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Alexandru","family":"Iosup","sequence":"additional","affiliation":[{"name":"Vrije Universiteit Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Animesh","family":"Trivedi","sequence":"additional","affiliation":[{"name":"Vrije Universiteit Amsterdam, Amsterdam, Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,4,19]]},"reference":[{"doi-asserted-by":"crossref","unstructured":"Agelastos et al. 2014. The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications. In SC.  Agelastos et al. 2014. The Lightweight Distributed Metric Service: A Scalable Infrastructure for Continuous Monitoring of Large Scale Computing Systems and Applications. In SC.","key":"e_1_3_2_1_1_1","DOI":"10.1109\/SC.2014.18"},{"key":"e_1_3_2_1_2_1","volume-title":"Software Engineering for Machine Learning: A Case Study. In 2019 IEEE\/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP).","author":"Amershi","year":"2019","unstructured":"Amershi et al. 2019 . Software Engineering for Machine Learning: A Case Study. In 2019 IEEE\/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). Amershi et al. 2019. Software Engineering for Machine Learning: A Case Study. In 2019 IEEE\/ACM 41st International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP)."},{"doi-asserted-by":"crossref","unstructured":"Bal et al. 2016. A Medium-Scale Distributed System for Computer Science Research: Infrastructure for the Long Term. IEEE Computer (2016).  Bal et al. 2016. A Medium-Scale Distributed System for Computer Science Research: Infrastructure for the Long Term. IEEE Computer (2016).","key":"e_1_3_2_1_3_1","DOI":"10.1109\/MC.2016.127"},{"key":"e_1_3_2_1_4_1","volume-title":"Challenges and Experiences with MLOps for Performance Diagnostics in Hybrid-Cloud Enterprise Software Deployments. In 2020 USENIX Conference on Operational Machine Learning (OpML 20)","author":"Banerjee","year":"2020","unstructured":"Banerjee et al. 2020 . Challenges and Experiences with MLOps for Performance Diagnostics in Hybrid-Cloud Enterprise Software Deployments. In 2020 USENIX Conference on Operational Machine Learning (OpML 20) . Banerjee et al. 2020. Challenges and Experiences with MLOps for Performance Diagnostics in Hybrid-Cloud Enterprise Software Deployments. In 2020 USENIX Conference on Operational Machine Learning (OpML 20)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_5_1","DOI":"10.1145\/3097983.3098021"},{"key":"e_1_3_2_1_6_1","volume-title":"Continuous Training for Production ML in the TensorFlow Extended (TFX) Platform. In 2019 USENIX Conference on Operational Machine Learning (OpML 19)","author":"Baylor","year":"2019","unstructured":"Baylor et al. 2019 . Continuous Training for Production ML in the TensorFlow Extended (TFX) Platform. In 2019 USENIX Conference on Operational Machine Learning (OpML 19) . Baylor et al. 2019. Continuous Training for Production ML in the TensorFlow Extended (TFX) Platform. In 2019 USENIX Conference on Operational Machine Learning (OpML 19)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_7_1","DOI":"10.1109\/IPDPS.2019.00018"},{"doi-asserted-by":"crossref","unstructured":"Chu et al. 2016. Data cleaning: Overview and emerging challenges. In SIGMOD.  Chu et al. 2016. Data cleaning: Overview and emerging challenges. In SIGMOD.","key":"e_1_3_2_1_8_1","DOI":"10.1145\/2882903.2912574"},{"key":"e_1_3_2_1_9_1","volume-title":"Training","volume":"100","author":"Coleman","year":"2017","unstructured":"Coleman et al. 2017 . Dawnbench: An end-to-end deep learning benchmark and competition . Training , Vol. 100 (2017). Coleman et al. 2017. Dawnbench: An end-to-end deep learning benchmark and competition. Training, Vol. 100 (2017)."},{"key":"e_1_3_2_1_10_1","volume-title":"Proceedings of the 14th USENIX Conference on Networked Systems Design and Implementation.","author":"Crankshaw","year":"2017","unstructured":"Crankshaw et al. 2017 . Clipper: A Low-Latency Online Prediction Serving System . In Proceedings of the 14th USENIX Conference on Networked Systems Design and Implementation. Crankshaw et al. 2017. Clipper: A Low-Latency Online Prediction Serving System. In Proceedings of the 14th USENIX Conference on Networked Systems Design and Implementation."},{"unstructured":"Dai et al. 2011. HiTune: Dataflow-Based Performance Analysis for Big Data Cloud. In ATC.  Dai et al. 2011. HiTune: Dataflow-Based Performance Analysis for Big Data Cloud. In ATC.","key":"e_1_3_2_1_11_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_12_1","DOI":"10.1145\/3394486.3403375"},{"key":"e_1_3_2_1_13_1","volume-title":"LISA 2012","author":"Gardu","year":"2012","unstructured":"Gardu n o et al. 2012 . Theia: Visual Signatures for Problem Diagnosis in Large Hadoop Clusters. In Strategies, Tools, and Techniques: Proceedings of the 26th Large Installation System Administration Conference , LISA 2012 , San Diego, CA, USA, December 9--14 , 2012. Gardu n o et al. 2012. Theia: Visual Signatures for Problem Diagnosis in Large Hadoop Clusters. In Strategies, Tools, and Techniques: Proceedings of the 26th Large Installation System Administration Conference, LISA 2012, San Diego, CA, USA, December 9--14, 2012."},{"key":"e_1_3_2_1_14_1","volume-title":"Data shapley: Equitable valuation of data for machine learning. arXiv preprint arXiv:1904.02868","author":"Zou Ghorbani","year":"2019","unstructured":"Ghorbani and Zou . 2019. Data shapley: Equitable valuation of data for machine learning. arXiv preprint arXiv:1904.02868 ( 2019 ). Ghorbani and Zou. 2019. Data shapley: Equitable valuation of data for machine learning. arXiv preprint arXiv:1904.02868 (2019)."},{"key":"e_1_3_2_1_15_1","volume-title":"A 20-Year Community Roadmap for Artificial Intelligence Research in the US. arxiv","author":"Gil Yolanda","year":"1908","unstructured":"Yolanda Gil and Bart Selman . 2019. A 20-Year Community Roadmap for Artificial Intelligence Research in the US. arxiv : 1908 .02624 [cs.CY] Yolanda Gil and Bart Selman. 2019. A 20-Year Community Roadmap for Artificial Intelligence Research in the US. arxiv: 1908.02624 [cs.CY]"},{"unstructured":"Guo et al. 2011. G2: A Graph Processing System for Diagnosing Distributed Systems. In ATC.  Guo et al. 2011. G2: A Graph Processing System for Diagnosing Distributed Systems. In ATC.","key":"e_1_3_2_1_16_1"},{"key":"e_1_3_2_1_17_1","volume-title":"Pipedream: Fast and efficient pipeline parallel dnn training. arXiv preprint arXiv:1806.03377","author":"Harlap","year":"2018","unstructured":"Harlap et al. 2018 . Pipedream: Fast and efficient pipeline parallel dnn training. arXiv preprint arXiv:1806.03377 (2018). Harlap et al. 2018. Pipedream: Fast and efficient pipeline parallel dnn training. arXiv preprint arXiv:1806.03377 (2018)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_18_1","DOI":"10.1109\/CVPR.2016.90"},{"doi-asserted-by":"crossref","unstructured":"Hegeman et al. 2020. Grade10: A Framework for Performance Characterization of Distributed Graph Processing. In CLUSTER.  Hegeman et al. 2020. Grade10: A Framework for Performance Characterization of Distributed Graph Processing. In CLUSTER.","key":"e_1_3_2_1_19_1","DOI":"10.1109\/CLUSTER49012.2020.00016"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_20_1","DOI":"10.1145\/3293883.3295710"},{"unstructured":"Hopsworks. 2021. Hopsworks. https:\/\/www.hopsworks.ai\/.  Hopsworks. 2021. Hopsworks. https:\/\/www.hopsworks.ai\/.","key":"e_1_3_2_1_21_1"},{"key":"e_1_3_2_1_22_1","volume-title":"Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Systems.","author":"Huang","year":"2019","unstructured":"Huang et al. 2019 . Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Systems. Huang et al. 2019. Gpipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Systems."},{"doi-asserted-by":"crossref","unstructured":"Jansen et al. 2020. DDLBench: Towards a Scalable Benchmarking Infrastructure for Distributed Deep Learning. In ISC.  Jansen et al. 2020. DDLBench: Towards a Scalable Benchmarking Infrastructure for Distributed Deep Learning. In ISC.","key":"e_1_3_2_1_23_1","DOI":"10.1109\/DLS51937.2020.00009"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_24_1","DOI":"10.1145\/3361525.3361538"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_25_1","DOI":"10.5555\/3358807.3358888"},{"unstructured":"Jiang et al. 2020. Hpc ai500: The methodology tools roofline performance models and metrics for benchmarking hpc ai systems. arXiv preprint arXiv:2007.00279 (2020).  Jiang et al. 2020. Hpc ai500: The methodology tools roofline performance models and metrics for benchmarking hpc ai systems. arXiv preprint arXiv:2007.00279 (2020).","key":"e_1_3_2_1_26_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_27_1","DOI":"10.1109\/BigData.2018.8622396"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_28_1","DOI":"10.1145\/3394486.3403290"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","DOI":"10.1145\/3409963","volume-title":"Proceedings of the 11th ACM SIGOPS Asia-Pacific Workshop on Systems.","author":"Lee Kim","year":"2020","unstructured":"Kim and Lee . 2020 . Reducing Tail Latency of DNN-Based Recommender Systems Using in-Storage Processing . In Proceedings of the 11th ACM SIGOPS Asia-Pacific Workshop on Systems. Kim and Lee. 2020. Reducing Tail Latency of DNN-Based Recommender Systems Using in-Storage Processing. In Proceedings of the 11th ACM SIGOPS Asia-Pacific Workshop on Systems."},{"key":"e_1_3_2_1_30_1","volume-title":"The Vampir Performance Analysis Tool-Set. In Tools for High Performance Computing - Proceedings of the 2nd International Workshop on Parallel Tools for High Performance Computing","author":"Kn\u00fc","year":"2008","unstructured":"Kn\u00fc pfer et al. 2008 . The Vampir Performance Analysis Tool-Set. In Tools for High Performance Computing - Proceedings of the 2nd International Workshop on Parallel Tools for High Performance Computing , July 2008 , HLRS, Stuttgart. Kn\u00fc pfer et al. 2008. The Vampir Performance Analysis Tool-Set. In Tools for High Performance Computing - Proceedings of the 2nd International Workshop on Parallel Tools for High Performance Computing, July 2008, HLRS, Stuttgart."},{"key":"e_1_3_2_1_31_1","article-title":"Auto-WEKA 2.0: Automatic model selection and hyperparameter optimization in WEKA","volume":"18","author":"Kotthoff","year":"2017","unstructured":"Kotthoff et al. 2017 . Auto-WEKA 2.0: Automatic model selection and hyperparameter optimization in WEKA . The Journal of Machine Learning Research , Vol. 18 (2017). Kotthoff et al. 2017. Auto-WEKA 2.0: Automatic model selection and hyperparameter optimization in WEKA. The Journal of Machine Learning Research, Vol. 18 (2017).","journal-title":"The Journal of Machine Learning Research"},{"key":"e_1_3_2_1_32_1","volume-title":"Proceedings of the VLDB Endowment","volume":"11","year":"2018","unstructured":"Kraska. 2018 . Northstar: An interactive data science system . Proceedings of the VLDB Endowment , Vol. 11 (2018). Kraska. 2018. Northstar: An interactive data science system. Proceedings of the VLDB Endowment, Vol. 11 (2018)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_33_1","DOI":"10.1145\/3187009.3177737"},{"unstructured":"Li et al. 2019. Across-Stack Profiling and Characterization of Machine Learning Models on GPUs. CoRR Vol. abs\/1908.06869 (2019).  Li et al. 2019. Across-Stack Profiling and Characterization of Machine Learning Models on GPUs. CoRR Vol. abs\/1908.06869 (2019).","key":"e_1_3_2_1_34_1"},{"key":"e_1_3_2_1_35_1","volume-title":"MLOp Lifecycle Scheme for Vision-based Inspection Process in Manufacturing. In 2019 USENIX Conference on Operational Machine Learning (OpML 19)","author":"Lim","year":"2019","unstructured":"Lim et al. 2019 . MLOp Lifecycle Scheme for Vision-based Inspection Process in Manufacturing. In 2019 USENIX Conference on Operational Machine Learning (OpML 19) . Lim et al. 2019. MLOp Lifecycle Scheme for Vision-based Inspection Process in Manufacturing. In 2019 USENIX Conference on Operational Machine Learning (OpML 19)."},{"key":"e_1_3_2_1_36_1","volume-title":"Medical Image Analysis","volume":"42","author":"Litjens","year":"2017","unstructured":"Litjens et al. 2017 . A survey on deep learning in medical image analysis . Medical Image Analysis , Vol. 42 (2017). Litjens et al. 2017. A survey on deep learning in medical image analysis. Medical Image Analysis, Vol. 42 (2017)."},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_37_1","DOI":"10.1145\/2815400.2815415"},{"key":"e_1_3_2_1_38_1","volume-title":"Proceedings of the 12th USENIX Conference on Networked Systems Design and Implementation.","author":"Mace","year":"2015","unstructured":"Mace et al. 2015 b. Retro: Targeted Resource Management in Multi-tenant Distributed Systems . In Proceedings of the 12th USENIX Conference on Networked Systems Design and Implementation. Mace et al. 2015b. Retro: Targeted Resource Management in Multi-tenant Distributed Systems. In Proceedings of the 12th USENIX Conference on Networked Systems Design and Implementation."},{"unstructured":"Mai et al. 2020. KungFu: Making Training in Distributed Machine Learning Adaptive. In OSDI.  Mai et al. 2020. KungFu: Making Training in Distributed Machine Learning Adaptive. In OSDI.","key":"e_1_3_2_1_39_1"},{"unstructured":"Mattson et al. 2019. Mlperf training benchmark. arXiv preprint arXiv:1910.01500 (2019).  Mattson et al. 2019. Mlperf training benchmark. arXiv preprint arXiv:1910.01500 (2019).","key":"e_1_3_2_1_40_1"},{"key":"e_1_3_2_1_41_1","volume-title":"ACM Comput. Surv.","volume":"53","author":"Jacobsen Mayer","year":"2020","unstructured":"Mayer and Jacobsen . 2020 . Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques, and Tools . ACM Comput. Surv. , Vol. 53 (2020). Mayer and Jacobsen. 2020. Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques, and Tools. ACM Comput. Surv., Vol. 53 (2020)."},{"doi-asserted-by":"crossref","unstructured":"Mirgorodskiy et al. 2008. Diagnosing distributed systems with self-propelled instrumentation. In Middleware.  Mirgorodskiy et al. 2008. Diagnosing distributed systems with self-propelled instrumentation. In Middleware.","key":"e_1_3_2_1_42_1","DOI":"10.1007\/978-3-540-89856-6_5"},{"unstructured":"NVIDIA. 2021 a. NVIDIA Nsight. https:\/\/developer.nvidia.com\/tools-overview.  NVIDIA. 2021 a. NVIDIA Nsight. https:\/\/developer.nvidia.com\/tools-overview.","key":"e_1_3_2_1_43_1"},{"unstructured":"NVIDIA. 2021 b. NVIDIA Profiler. https:\/\/docs.nvidia.com\/cuda\/profiler-users-guide\/index.html.  NVIDIA. 2021 b. NVIDIA Profiler. https:\/\/docs.nvidia.com\/cuda\/profiler-users-guide\/index.html.","key":"e_1_3_2_1_44_1"},{"unstructured":"Ousterhout et al. 2015. Making Sense of Performance in Data Analytics Frameworks. In NSDI.  Ousterhout et al. 2015. Making Sense of Performance in Data Analytics Frameworks. In NSDI.","key":"e_1_3_2_1_45_1"},{"doi-asserted-by":"crossref","unstructured":"Pi et al. 2018. Profiling distributed systems in lightweight virtualized environments with logs and resource metrics. In HPDC.  Pi et al. 2018. Profiling distributed systems in lightweight virtualized environments with logs and resource metrics. In HPDC.","key":"e_1_3_2_1_46_1","DOI":"10.1145\/3220192.3220197"},{"key":"e_1_3_2_1_47_1","volume-title":"Proceedings of Machine Learning and Systems","volume":"1","author":"Polyzotis","year":"2019","unstructured":"Polyzotis et al. 2019 . Data validation for machine learning . Proceedings of Machine Learning and Systems , Vol. 1 (2019). Polyzotis et al. 2019. Data validation for machine learning. Proceedings of Machine Learning and Systems, Vol. 1 (2019)."},{"volume-title":"VTune performance analyzer essentials","year":"2005","unstructured":"Reinders. 2005. VTune performance analyzer essentials . Intel Press ( 2005 ). Reinders. 2005. VTune performance analyzer essentials. Intel Press (2005).","key":"e_1_3_2_1_48_1"},{"key":"e_1_3_2_1_49_1","volume-title":"Social Media: Facebook and Twitter Perspectives. Advances in Science, Technology and Engineering Systems Journal","author":"Salloum","year":"2017","unstructured":"Salloum et al. 2017 . A Survey of Text Mining in Social Media: Facebook and Twitter Perspectives. Advances in Science, Technology and Engineering Systems Journal , Vol. 2 (2017). Salloum et al. 2017. A Survey of Text Mining in Social Media: Facebook and Twitter Perspectives. Advances in Science, Technology and Engineering Systems Journal, Vol. 2 (2017)."},{"key":"e_1_3_2_1_50_1","volume-title":"Horovod: fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799","author":"Del Sergeev","year":"2018","unstructured":"Sergeev and Del . 2018. Horovod: fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799 ( 2018 ). Sergeev and Del. 2018. Horovod: fast and easy distributed deep learning in TensorFlow. arXiv preprint arXiv:1802.05799 (2018)."},{"key":"e_1_3_2_1_51_1","volume-title":"IJHPCA","volume":"20","author":"Malony Shende","year":"2006","unstructured":"Shende and Malony . 2006 . The TAU Parallel Performance System . IJHPCA , Vol. 20 (2006). Shende and Malony. 2006. The TAU Parallel Performance System. IJHPCA, Vol. 20 (2006)."},{"unstructured":"Sigelman et al. 2010. Dapper a large-scale distributed systems tracing infrastructure. (2010).  Sigelman et al. 2010. Dapper a large-scale distributed systems tracing infrastructure. (2010).","key":"e_1_3_2_1_52_1"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_53_1","DOI":"10.1145\/3180155.3180220"},{"doi-asserted-by":"publisher","key":"e_1_3_2_1_54_1","DOI":"10.1145\/3343737.3343743"},{"unstructured":"Wang et al. 2012. VScope: Middleware for Troubleshooting Time-Sensitive Data Center Applications. In Middleware 2012 - ACM\/IFIP\/USENIX 13th International Middleware Conference Montreal QC Canada December 3--7 2012. Proceedings Vol. 7662.  Wang et al. 2012. VScope: Middleware for Troubleshooting Time-Sensitive Data Center Applications. In Middleware 2012 - ACM\/IFIP\/USENIX 13th International Middleware Conference Montreal QC Canada December 3--7 2012. Proceedings Vol. 7662.","key":"e_1_3_2_1_55_1"},{"unstructured":"Wang et al. 2019. Benchmarking TPU GPU and CPU platforms for deep learning. arXiv preprint arXiv:1907.10701 (2019).  Wang et al. 2019. Benchmarking TPU GPU and CPU platforms for deep learning. arXiv preprint arXiv:1907.10701 (2019).","key":"e_1_3_2_1_56_1"},{"doi-asserted-by":"crossref","unstructured":"Wang et al. 2020. Metis: learning to schedule long-running applications in shared container clusters at scale. In SC.  Wang et al. 2020. Metis: learning to schedule long-running applications in shared container clusters at scale. In SC.","key":"e_1_3_2_1_57_1","DOI":"10.1109\/SC41405.2020.00072"},{"unstructured":"Yang et al. [n.d.]. End-to-end I\/O Monitoring on a Leading Supercomputer. In NSDI.  Yang et al. [n.d.]. End-to-end I\/O Monitoring on a Leading Supercomputer. In NSDI.","key":"e_1_3_2_1_58_1"},{"key":"e_1_3_2_1_59_1","volume-title":"Nanolog: A nanosecond scale logging system. In ATC.","author":"Yang","year":"2018","unstructured":"Yang et al. 2018 . Nanolog: A nanosecond scale logging system. In ATC. Yang et al. 2018. Nanolog: A nanosecond scale logging system. In ATC."},{"doi-asserted-by":"crossref","unstructured":"You et al. 2018. Imagenet training in minutes. In ICPP.  You et al. 2018. Imagenet training in minutes. In ICPP.","key":"e_1_3_2_1_60_1","DOI":"10.1145\/3225058.3225069"},{"key":"e_1_3_2_1_61_1","volume-title":"IEEE Data Eng. Bull.","volume":"41","author":"Zaharia","year":"2018","unstructured":"Zaharia et al. 2018 . Accelerating the Machine Learning Lifecycle with MLflow . IEEE Data Eng. Bull. , Vol. 41 (2018). Zaharia et al. 2018. Accelerating the Machine Learning Lifecycle with MLflow. IEEE Data Eng. Bull., Vol. 41 (2018)."},{"key":"e_1_3_2_1_62_1","volume-title":"MIMP: Deadline and Interference Aware Scheduling of Hadoop Virtual Machines. In CCGrid.","author":"Zhang","year":"2014","unstructured":"Zhang et al. 2014 . MIMP: Deadline and Interference Aware Scheduling of Hadoop Virtual Machines. In CCGrid. Zhang et al. 2014. MIMP: Deadline and Interference Aware Scheduling of Hadoop Virtual Machines. In CCGrid."},{"key":"e_1_3_2_1_63_1","volume-title":"Model-Switching: Dealing with Fluctuating Workloads in Machine-Learning-as-a-Service Systems. In 12th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 20)","author":"Zhang","year":"2020","unstructured":"Zhang et al. 2020 . Model-Switching: Dealing with Fluctuating Workloads in Machine-Learning-as-a-Service Systems. In 12th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 20) . Zhang et al. 2020. Model-Switching: Dealing with Fluctuating Workloads in Machine-Learning-as-a-Service Systems. In 12th USENIX Workshop on Hot Topics in Cloud Computing (HotCloud 20)."},{"unstructured":"Zhao et al. 2016. Non-Intrusive Performance Profiling for Entire Software Stacks Based on the Flow Reconstruction Principle. In OSDI.  Zhao et al. 2016. Non-Intrusive Performance Profiling for Entire Software Stacks Based on the Flow Reconstruction Principle. In OSDI.","key":"e_1_3_2_1_64_1"},{"doi-asserted-by":"crossref","unstructured":"Zhao et al. 2017. Log20: Fully automated optimal placement of log printing statements under specified overhead threshold. In SOSP.  Zhao et al. 2017. Log20: Fully automated optimal placement of log printing statements under specified overhead threshold. In SOSP.","key":"e_1_3_2_1_65_1","DOI":"10.1145\/3132747.3132778"},{"key":"e_1_3_2_1_66_1","volume-title":"Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training. arXiv preprint arXiv:2006.03318","author":"Zhu","year":"2020","unstructured":"Zhu et al. 2020 . Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training. arXiv preprint arXiv:2006.03318 (2020). Zhu et al. 2020. Daydream: Accurately Estimating the Efficacy of Optimizations for DNN Training. arXiv preprint arXiv:2006.03318 (2020)."}],"event":{"sponsor":["SIGMETRICS ACM Special Interest Group on Measurement and Evaluation","SIGSOFT ACM Special Interest Group on Software Engineering"],"acronym":"ICPE '21","name":"ICPE '21: ACM\/SPEC International Conference on Performance Engineering","location":"Virtual Event France"},"container-title":["Companion of the ACM\/SPEC International Conference on Performance Engineering"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447545.3451185","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3447545.3451185","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:25:10Z","timestamp":1750195510000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3447545.3451185"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,19]]},"references-count":66,"alternative-id":["10.1145\/3447545.3451185","10.1145\/3447545"],"URL":"https:\/\/doi.org\/10.1145\/3447545.3451185","relation":{},"subject":[],"published":{"date-parts":[[2021,4,19]]},"assertion":[{"value":"2021-04-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}