{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,15]],"date-time":"2026-07-15T03:24:41Z","timestamp":1784085881305,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":33,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,4,26]],"date-time":"2021-04-26T00:00:00Z","timestamp":1619395200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,4,26]]},"DOI":"10.1145\/3437984.3458837","type":"proceedings-article","created":{"date-parts":[[2021,4,25]],"date-time":"2021-04-25T09:56:04Z","timestamp":1619344564000},"page":"80-88","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":31,"title":["Interference-Aware Scheduling for Inference Serving"],"prefix":"10.1145","author":[{"given":"Daniel","family":"Mendoza","sequence":"first","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Francisco","family":"Romero","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Qian","family":"Li","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Neeraja J.","family":"Yadwadkar","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christos","family":"Kozyrakis","sequence":"additional","affiliation":[{"name":"Stanford University"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,4,26]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2018. NVIDIA TensorRT: Programmable Inference Accelerator. https:\/\/developer.nvidia.com\/tensorrt.  2018. NVIDIA TensorRT: Programmable Inference Accelerator. https:\/\/developer.nvidia.com\/tensorrt."},{"key":"e_1_3_2_1_2_1","volume-title":"CherryPick: Adaptively Unearthing the Best Cloud Configurations for Big Data Analytics. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)","author":"Alipourfard Omid","year":"2017","unstructured":"Omid Alipourfard , Hongqiang Harry Liu , Jianshu Chen , Shivaram Venkataraman , Minlan Yu , and Ming Zhang . 2017 . CherryPick: Adaptively Unearthing the Best Cloud Configurations for Big Data Analytics. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . USENIX Association, Boston, MA, 469--482. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/alipourfard Omid Alipourfard, Hongqiang Harry Liu, Jianshu Chen, Shivaram Venkataraman, Minlan Yu, and Ming Zhang. 2017. CherryPick: Adaptively Unearthing the Best Cloud Configurations for Big Data Analytics. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 469--482. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/alipourfard"},{"key":"e_1_3_2_1_3_1","unstructured":"AWS [n.d.]. AWS Neuron. https:\/\/github.com\/aws\/aws-neuron-sdk.  AWS [n.d.]. AWS Neuron. https:\/\/github.com\/aws\/aws-neuron-sdk."},{"key":"e_1_3_2_1_4_1","unstructured":"AWS 2018. AWS Inferentia. https:\/\/aws.amazon.com\/machine-learning\/inferentia\/.  AWS 2018. AWS Inferentia. https:\/\/aws.amazon.com\/machine-learning\/inferentia\/."},{"key":"e_1_3_2_1_5_1","unstructured":"AWS 2019. Deliver high performance ML inference with AWS Inferentia. https:\/\/d1.awsstatic.com\/events\/reinvent\/2019\/REPEAT_1_Deliver_high_performance_ML_inference_with_AWS_Inferentia_CMP324-R1.pdf.  AWS 2019. Deliver high performance ML inference with AWS Inferentia. https:\/\/d1.awsstatic.com\/events\/reinvent\/2019\/REPEAT_1_Deliver_high_performance_ML_inference_with_AWS_Inferentia_CMP324-R1.pdf."},{"key":"e_1_3_2_1_6_1","volume-title":"Active Learning - Modern Learning Theory","author":"Balcan Maria-Florina","unstructured":"Maria-Florina Balcan and Ruth Urner . 2016. Active Learning - Modern Learning Theory . Springer New York , New York, NY , 8--13. https:\/\/doi.org\/10.1007\/978-1-4939-2864-4_769 Maria-Florina Balcan and Ruth Urner. 2016. Active Learning - Modern Learning Theory. Springer New York, New York, NY, 8--13. https:\/\/doi.org\/10.1007\/978-1-4939-2864-4_769"},{"key":"e_1_3_2_1_7_1","volume-title":"Noise reduction in speech processing","author":"Benesty Jacob","unstructured":"Jacob Benesty , Jingdong Chen , Yiteng Huang , and Israel Cohen . 2009. Pearson correlation coefficient . In Noise reduction in speech processing . Springer , 37--40. Jacob Benesty, Jingdong Chen, Yiteng Huang, and Israel Cohen. 2009. Pearson correlation coefficient. In Noise reduction in speech processing. Springer, 37--40."},{"key":"e_1_3_2_1_8_1","volume-title":"TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen , Thierry Moreau , Ziheng Jiang , Lianmin Zheng , Eddie Yan , Haichen Shen , Meghan Cowan , Leyuan Wang , Yuwei Hu , Luis Ceze , Carlos Guestrin , and Arvind Krishnamurthy . 2018 . TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . USENIX Association, Carlsbad, CA, 578--594. https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/chen Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). USENIX Association, Carlsbad, CA, 578--594. https:\/\/www.usenix.org\/conference\/osdi18\/presentation\/chen"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3419111.3421285"},{"key":"e_1_3_2_1_10_1","volume-title":"Clipper: A Low-Latency Online Prediction Serving System. In 14th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2017","author":"Crankshaw Daniel","year":"2017","unstructured":"Daniel Crankshaw , Xin Wang , Giulio Zhou , Michael J. Franklin , Joseph E. Gonzalez , and Ion Stoica . 2017 . Clipper: A Low-Latency Online Prediction Serving System. In 14th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2017 , Boston, MA, USA , March 27-29, 2017. 613--627. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/crankshaw Daniel Crankshaw, Xin Wang, Giulio Zhou, Michael J. Franklin, Joseph E. Gonzalez, and Ion Stoica. 2017. Clipper: A Low-Latency Online Prediction Serving System. In 14th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2017, Boston, MA, USA, March 27-29, 2017. 613--627. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/crankshaw"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2556583"},{"key":"e_1_3_2_1_12_1","first-page":"4","article-title":"Quasar","volume":"49","author":"Delimitrou Christina","year":"2014","unstructured":"Christina Delimitrou and Christos Kozyrakis . 2014 . Quasar : Resource-Efficient and QoS-Aware Cluster Management. SIGPLAN Not. 49 , 4 (Feb. 2014), 127--144. https:\/\/doi.org\/10.1145\/2644865.2541941 Christina Delimitrou and Christos Kozyrakis. 2014. Quasar: Resource-Efficient and QoS-Aware Cluster Management. SIGPLAN Not. 49, 4 (Feb. 2014), 127--144. https:\/\/doi.org\/10.1145\/2644865.2541941","journal-title":"Resource-Efficient and QoS-Aware Cluster Management. SIGPLAN Not."},{"key":"e_1_3_2_1_13_1","volume-title":"Load Balancing in Cloud Computing Environment Using Improved Weighted Round Robin Algorithm for Nonpreemptive Dependent Tasks. The Scientific World Journal 2016 (03","author":"Chitra Devi D.","year":"2016","unstructured":"D. Chitra Devi and V. Rhymend Uthariaraj . 2016. Load Balancing in Cloud Computing Environment Using Improved Weighted Round Robin Algorithm for Nonpreemptive Dependent Tasks. The Scientific World Journal 2016 (03 Feb 2016 ), 3896065. https:\/\/doi.org\/10.1155\/2016\/3896065 D. Chitra Devi and V. Rhymend Uthariaraj. 2016. Load Balancing in Cloud Computing Environment Using Improved Weighted Round Robin Algorithm for Nonpreemptive Dependent Tasks. The Scientific World Journal 2016 (03 Feb 2016), 3896065. https:\/\/doi.org\/10.1155\/2016\/3896065"},{"key":"e_1_3_2_1_14_1","unstructured":"Google [n.d.]. Google Cloud Platform. https:\/\/cloud.google.com\/compute.  Google [n.d.]. Google Cloud Platform. https:\/\/cloud.google.com\/compute."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Udit Gupta Samuel Hsia Vikram Saraph Xiaodong Wang Brandon Reagen Gu-Yeon Wei Hsien-Hsin S. Lee David Brooks and Carole-Jean Wu. 2020. DeepRecSys: A System for Optimizing End-To-End At-scale Neural Recommendation Inference. arXiv:2001.02772 [cs.DC]  Udit Gupta Samuel Hsia Vikram Saraph Xiaodong Wang Brandon Reagen Gu-Yeon Wei Hsien-Hsin S. Lee David Brooks and Carole-Jean Wu. 2020. DeepRecSys: A System for Optimizing End-To-End At-scale Neural Recommendation Inference. arXiv:2001.02772 [cs.DC]","DOI":"10.1109\/ISCA45697.2020.00084"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1136\/svn-2017-000101"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3230543.3230574"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_1_19_1","volume-title":"Yanqi Zhou, and Mike Burrows.","author":"Kaufman Samuel J.","year":"2020","unstructured":"Samuel J. Kaufman , Phitchaya Mangpo Phothilimthana , Yanqi Zhou, and Mike Burrows. 2020 . A Learned Performance Model for the Tensor Processing Unit . arXiv:2008.01040 [cs.PF] Samuel J. Kaufman, Phitchaya Mangpo Phothilimthana, Yanqi Zhou, and Mike Burrows. 2020. A Learned Performance Model for the Tensor Processing Unit. arXiv:2008.01040 [cs.PF]"},{"key":"e_1_3_2_1_20_1","unstructured":"Keras [n.d.]. Keras. https:\/\/github.com\/fchollet\/keras.  Keras [n.d.]. Keras. https:\/\/github.com\/fchollet\/keras."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341302.3342080"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378522"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/71.963420"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"crossref","unstructured":"T. Patel and D. Tiwari. 2020. CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale Computers. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). 193--206. https:\/\/doi.org\/10.1109\/HPCA47549.2020.00025  T. Patel and D. Tiwari. 2020. CLITE: Efficient and QoS-Aware Co-Location of Multiple Latency-Critical Jobs for Warehouse Scale Computers. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). 193--206. https:\/\/doi.org\/10.1109\/HPCA47549.2020.00025","DOI":"10.1109\/HPCA47549.2020.00025"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3243176.3243183"},{"key":"e_1_3_2_1_26_1","volume-title":"Managed & Model-less Inference Serving. CoRR abs\/1905.13348","author":"Romero Francisco","year":"2019","unstructured":"Francisco Romero , Qian Li , Neeraja J. Yadwadkar , and Christos Kozyrakis . 2019. INFaaS : Managed & Model-less Inference Serving. CoRR abs\/1905.13348 ( 2019 ). arXiv:1905.13348 http:\/\/arxiv.org\/abs\/1905.13348 Francisco Romero, Qian Li, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2019. INFaaS: Managed & Model-less Inference Serving. CoRR abs\/1905.13348 (2019). arXiv:1905.13348 http:\/\/arxiv.org\/abs\/1905.13348"},{"key":"e_1_3_2_1_27_1","volume-title":"Llama: A Heterogeneous & Serverless Framework for Auto-Tuning Video Analytics Pipelines. arXiv:2102.01887 [cs.DC]","author":"Romero Francisco","year":"2021","unstructured":"Francisco Romero , Mark Zhao , Neeraja J. Yadwadkar , and Christos Kozyrakis . 2021 . Llama: A Heterogeneous & Serverless Framework for Auto-Tuning Video Analytics Pipelines. arXiv:2102.01887 [cs.DC] Francisco Romero, Mark Zhao, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2021. Llama: A Heterogeneous & Serverless Framework for Auto-Tuning Video Analytics Pipelines. arXiv:2102.01887 [cs.DC]"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/645530.655646"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Subhadra Shaw and A. Singh. 2014. A survey on scheduling and load balancing techniques in cloud computing environment. 87--95. https:\/\/doi.org\/10.1109\/ICCCT.2014.7001474  Subhadra Shaw and A. Singh. 2014. A survey on scheduling and load balancing techniques in cloud computing environment. 87--95. https:\/\/doi.org\/10.1109\/ICCCT.2014.7001474","DOI":"10.1109\/ICCCT.2014.7001474"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/11564096_42"},{"key":"e_1_3_2_1_31_1","volume-title":"Characterization and Prediction of Performance Interference on Mediated Passthrough GPUs for Interference-aware Scheduler. In 11th USENIX Workshop on Hot Topics in Cloud Computing (Hot-Cloud 19)","author":"Xu Xin","year":"2019","unstructured":"Xin Xu , Na Zhang , Michael Cui , Michael He , and Ridhi Surana . 2019 . Characterization and Prediction of Performance Interference on Mediated Passthrough GPUs for Interference-aware Scheduler. In 11th USENIX Workshop on Hot Topics in Cloud Computing (Hot-Cloud 19) . USENIX Association, Renton, WA. https:\/\/www.usenix.org\/conference\/hotcloud19\/presentation\/xu-xin Xin Xu, Na Zhang, Michael Cui, Michael He, and Ridhi Surana. 2019. Characterization and Prediction of Performance Interference on Mediated Passthrough GPUs for Interference-aware Scheduler. In 11th USENIX Workshop on Hot Topics in Cloud Computing (Hot-Cloud 19). USENIX Association, Renton, WA. https:\/\/www.usenix.org\/conference\/hotcloud19\/presentation\/xu-xin"},{"key":"e_1_3_2_1_32_1","volume-title":"SLO-Aware Machine Learning Inference Serving. In 2019 USENIX Annual Technical Conference (USENIX ATC 19)","author":"Zhang Chengliang","year":"2019","unstructured":"Chengliang Zhang , Minchen Yu , Wei Wang , and Feng Yan . 2019 . MArk: Exploiting Cloud Services for Cost-Effective , SLO-Aware Machine Learning Inference Serving. In 2019 USENIX Annual Technical Conference (USENIX ATC 19) . USENIX Association, Renton, WA, 1049--1062. https:\/\/www.usenix.org\/conference\/atc19\/presentation\/zhang-chengliang Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019. MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving. In 2019 USENIX Annual Technical Conference (USENIX ATC 19). USENIX Association, Renton, WA, 1049--1062. https:\/\/www.usenix.org\/conference\/atc19\/presentation\/zhang-chengliang"},{"key":"e_1_3_2_1_33_1","volume-title":"14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)","author":"Zhang Haoyu","unstructured":"Haoyu Zhang , Ganesh Ananthanarayanan , Peter Bodik , Matthai Philipose , Paramvir Bahl , and Michael J. Freedman . 2017. Live Video Analytics at Scale with Approximation and Delay-Tolerance . In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . USENIX Association, Boston, MA, 377--392. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/zhang Haoyu Zhang, Ganesh Ananthanarayanan, Peter Bodik, Matthai Philipose, Paramvir Bahl, and Michael J. Freedman. 2017. Live Video Analytics at Scale with Approximation and Delay-Tolerance. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17). USENIX Association, Boston, MA, 377--392. https:\/\/www.usenix.org\/conference\/nsdi17\/technical-sessions\/presentation\/zhang"}],"event":{"name":"EuroSys '21: Sixteenth European Conference on Computer Systems","location":"Online United Kingdom","acronym":"EuroSys '21","sponsor":["SIGOPS ACM Special Interest Group on Operating Systems"]},"container-title":["Proceedings of the 1st Workshop on Machine Learning and Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3437984.3458837","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3437984.3458837","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:18Z","timestamp":1750193238000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3437984.3458837"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,26]]},"references-count":33,"alternative-id":["10.1145\/3437984.3458837","10.1145\/3437984"],"URL":"https:\/\/doi.org\/10.1145\/3437984.3458837","relation":{},"subject":[],"published":{"date-parts":[[2021,4,26]]},"assertion":[{"value":"2021-04-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}