{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T13:09:18Z","timestamp":1780664958043,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":72,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,4,26]],"date-time":"2026-04-26T00:00:00Z","timestamp":1777161600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,4,27]]},"DOI":"10.1145\/3767295.3769353","type":"proceedings-article","created":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T20:20:04Z","timestamp":1777062004000},"page":"1349-1364","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0008-0895-4585","authenticated-orcid":false,"given":"Tian","family":"Xia","sequence":"first","affiliation":[{"name":"UC Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7738-2498","authenticated-orcid":false,"given":"Ziming","family":"Mao","sequence":"additional","affiliation":[{"name":"UC Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-8647-2680","authenticated-orcid":false,"given":"Jamison","family":"Kerney","sequence":"additional","affiliation":[{"name":"UC Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-7770-0571","authenticated-orcid":false,"given":"Ethan J.","family":"Jackson","sequence":"additional","affiliation":[{"name":"UC Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-1488-4871","authenticated-orcid":false,"given":"Zhifei","family":"Li","sequence":"additional","affiliation":[{"name":"Renmin University of China, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-6163-0569","authenticated-orcid":false,"given":"Jiarong","family":"Xing","sequence":"additional","affiliation":[{"name":"Rice University AND UC Berkeley, Houston, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1357-7533","authenticated-orcid":false,"given":"Scott","family":"Shenker","sequence":"additional","affiliation":[{"name":"ICSI AND UC Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5373-0088","authenticated-orcid":false,"given":"Ion","family":"Stoica","sequence":"additional","affiliation":[{"name":"UC Berkeley, Berkeley, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,4,26]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2024. SGLang v0.4: Zero-Overhead Batch Scheduler Cache-Aware Load Balancer Faster Structured Outputs. https:\/\/lmsys.org\/blog\/2024-12-04-sglang-v0-4\/. Accessed: 2025-05-12."},{"key":"e_1_3_2_1_2_1","unstructured":"2025. Google Kubernetes Engine (GKE). https:\/\/cloud.google.com\/kubernetes-engine?hl=en. Accessed: 2025-05-12."},{"key":"e_1_3_2_1_3_1","unstructured":"AIME. 2025. CLOUD VS. ON-PREMISE - Total Cost of Ownership Analysis. https:\/\/www.aime.info\/blog\/en\/cloud-vs-on-premise-total-cost-of-ownership-analysis\/ Accessed: 2025-05-13."},{"key":"e_1_3_2_1_4_1","unstructured":"Amazon Web Service. 2025. AWS Inferentia. https:\/\/aws.amazon.com\/ai\/machine-learning\/inferentia\/ Accessed: 2025-05-15."},{"key":"e_1_3_2_1_5_1","unstructured":"Amazon Web Services. 2025. Amazon Bedrock. https:\/\/aws.amazon.com\/bedrock\/ Accessed: 2025-05-09."},{"key":"e_1_3_2_1_6_1","unstructured":"Amazon Web Services. 2025. Amazon Route 53. https:\/\/aws.amazon.com\/route53\/ Accessed: 2025-05-12."},{"key":"e_1_3_2_1_7_1","unstructured":"Amazon Web Services. 2025. Cloud Computing with AWS \u2014 Build Deploy and Manage Websites Apps or Processes on AWS Secure Reliable Network. https:\/\/aws.amazon.com\/ Accessed: 2025-05-13."},{"key":"e_1_3_2_1_8_1","unstructured":"Apple Inc. 2025. Use ChatGPT with Apple Intelligence on iPhone. https:\/\/support.apple.com\/guide\/iphone\/use-chatgpt-with-apple-intelligence-iph00fd3c8c2\/ios Accessed: 2025-05-12."},{"key":"e_1_3_2_1_9_1","unstructured":"Mohammad Beigi Sijia Wang Ying Shen Zihao Lin Adithya Kulkarni Jianfeng He Feng Chen Ming Jin Jin-Hee Cho Dawei Zhou Chang-Tien Lu and Lifu Huang. 2024. Rethinking the Uncertainty: A Critical Review and Analysis in the Era of Large Language Models. arXiv:2410.20199 [cs.AI] https:\/\/arxiv.org\/abs\/2410.20199"},{"key":"e_1_3_2_1_10_1","volume-title":"Proceedings of the 26th ACM Symposium on Operating Systems Principles.","author":"Bugnion Edouard","year":"2017","unstructured":"Edouard Bugnion. 2017. ZygOS: Achieving Low Tail Latency for Microsecond-scale Networked Tasks. In Proceedings of the 26th ACM Symposium on Operating Systems Principles."},{"key":"e_1_3_2_1_11_1","unstructured":"Brian Buntz. 2024. Musk unveils world's largest AI cluster OpenAI eyes premium subscriptions. https:\/\/www.rdworldonline.com\/this-week-in-ai-musk-unveils-worlds-largest-ai-cluster-openai-eyes-2000-month-subscriptions\/ Accessed: 2025-05-13."},{"key":"e_1_3_2_1_12_1","unstructured":"Shiyi Cao Yichuan Wang Ziming Mao Pin-Lun Hsu Liangsheng Yin Tian Xia Dacheng Li Shu Liu Yineng Zhang Yang Zhou et al. 2025. Locality-aware Fair Scheduling in LLM Serving. arXiv preprint arXiv:2501.14312 (2025)."},{"key":"e_1_3_2_1_13_1","volume-title":"Multi-model machine learning inference serving with gpu spatial partitioning. arXiv preprint arXiv:2109.01611","author":"Choi Seungbeom","year":"2021","unstructured":"Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh. 2021. Multi-model machine learning inference serving with gpu spatial partitioning. arXiv preprint arXiv:2109.01611 (2021)."},{"key":"e_1_3_2_1_14_1","volume-title":"Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:2110.14168","author":"Cobbe Karl","year":"2021","unstructured":"Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training Verifiers to Solve Math Word Problems. arXiv preprint arXiv:2110.14168 (2021)."},{"key":"e_1_3_2_1_15_1","volume-title":"Consensus: AI-Powered Academic Search Engine. https:\/\/consensus.app\/ Accessed: 2025-05-12.","year":"2025","unstructured":"Consensus. 2025. Consensus: AI-Powered Academic Search Engine. https:\/\/consensus.app\/ Accessed: 2025-05-12."},{"key":"e_1_3_2_1_16_1","volume-title":"Cursor: The AI Code Editor","year":"2025","unstructured":"Cursor. 2025. Cursor: The AI Code Editor. https:\/\/www.cursor.com\/en Accessed: 2025-05-15."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1137\/20m1323746"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3419111.3421284"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1654059.1654113"},{"key":"e_1_3_2_1_20_1","unstructured":"Doil Kim. 2025. Benchmarking LLM Serving Performance: A Comprehensive Guide. https:\/\/medium.com\/@kimdoil1211\/benchmarking-llm-serving-performance-a-comprehensive-guide-db94b1bfe8cf Accessed: 2025-09-11."},{"key":"e_1_3_2_1_21_1","unstructured":"Envoy Project. 2025. Load Balancers \u2014 Envoy Proxy Documentation. https:\/\/www.envoyproxy.io\/docs\/envoy\/latest\/intro\/arch_overview\/upstream\/load_balancing\/load_balancers Accessed: 2025-05-13."},{"key":"e_1_3_2_1_22_1","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Fried Joshua","year":"2020","unstructured":"Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, and Adam Belay. 2020. Caladan: Mitigating interference at microsecond timescales. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 281\u2013297."},{"key":"e_1_3_2_1_23_1","first-page":"325","article-title":"Prompt cache: Modular attention reuse for low-latency inference","volume":"6","author":"Gim In","year":"2024","unstructured":"In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong. 2024. Prompt cache: Modular attention reuse for low-latency inference. Proceedings of Machine Learning and Systems 6 (2024), 325\u2013338.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_2_1_24_1","unstructured":"GitHub. 2025. GitHub Copilot: Your AI pair programmer. https:\/\/github.com\/features\/copilot Accessed: 2025-05-12."},{"key":"e_1_3_2_1_25_1","unstructured":"Google Cloud. 2025. About the Gateway API | GKE networking. https:\/\/cloud.google.com\/kubernetes-engine\/docs\/concepts\/gateway-api Accessed: 2025-05-12."},{"key":"e_1_3_2_1_26_1","unstructured":"Google Cloud. 2025. Compute Engine Reservations Overview. https:\/\/cloud.google.com\/compute\/docs\/instances\/reservations-overview Accessed: 2025-05-12."},{"key":"e_1_3_2_1_27_1","unstructured":"Google Cloud Platform. 2025. Google Cloud Platform \u2014 Future-proof infrastructure. Powerful data and analytics. No ops just code. https:\/\/cloud.google.com\/ Accessed: 2025-05-13."},{"key":"e_1_3_2_1_28_1","volume-title":"Introducing Gemini: Your New Personal AI Assistant. https:\/\/gemini.google\/assistant\/ Accessed: 2025-05-12.","author":"Google","year":"2025","unstructured":"Google LLC. 2025. Introducing Gemini: Your New Personal AI Assistant. https:\/\/gemini.google\/assistant\/ Accessed: 2025-05-12."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"crossref","unstructured":"Norman P. Jouppi George Kurian Sheng Li Peter Ma Rahul Nagarajan Lifeng Nai Nishant Patil Suvinay Subramanian Andy Swing Brian Towles Cliff Young Xiang Zhou Zongwei Zhou and David Patterson. 2023. TPU v4: An Optically Reconfigurable Supercomputer for Machine Learning with Hardware Support for Embeddings. arXiv:2304.01433 [cs.AR] https:\/\/arxiv.org\/abs\/2304.01433","DOI":"10.1145\/3579371.3589350"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/258533.258660"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3379483"},{"key":"e_1_3_2_1_32_1","unstructured":"Kubernetes Authors. 2025. Services Load Balancing and Networking. https:\/\/kubernetes.io\/docs\/concepts\/services-networking\/ Accessed: 2025-05-13."},{"key":"e_1_3_2_1_33_1","volume-title":"Joseph E. Gonzalez, Hao Zhang, and Ion Stoica.","author":"Kwon Woosuk","year":"2023","unstructured":"Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient Memory Management for Large Language Model Serving with PagedAttention. arXiv:2309.06180 [cs.LG] https:\/\/arxiv.org\/abs\/2309.06180"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2851141.2851151"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3689031.3717459"},{"key":"e_1_3_2_1_36_1","volume-title":"19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22)","author":"McClure Sarah","year":"2022","unstructured":"Sarah McClure, Amy Ousterhout, Scott Shenker, and Sylvia Ratnasamy. 2022. Efficient scheduling policies for Microsecond-Scale tasks. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22). 1\u201318."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3620665.3640411"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/3588195.3595940"},{"key":"e_1_3_2_1_39_1","unstructured":"Evan Morikawa. 2023. Behind the Scenes Scaling ChatGPT. https:\/\/youtu.be\/PeKMEXUrlq4?t=833. Accessed: 2025-05-25."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/1858378.1858441"},{"key":"e_1_3_2_1_41_1","volume-title":"Vesper: Measuring Time-to-Interactivity for Web Pages. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18)","author":"Netravali Ravi","year":"2018","unstructured":"Ravi Netravali, Vikram Nathan, James Mickens, and Hari Balakrishnan. 2018. Vesper: Measuring Time-to-Interactivity for Web Pages. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18). USENIX Association, Renton, WA, 217\u2013231. https:\/\/www.usenix.org\/conference\/nsdi18\/presentation\/netravali-vesper"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2024.05.019"},{"key":"e_1_3_2_1_43_1","volume-title":"Flexllm: A system for co-serving large language model inference and parameter-efficient finetuning. arXiv preprint arXiv:2402.18789","author":"Oliaro Gabriele","year":"2024","unstructured":"Gabriele Oliaro, Xupeng Miao, Xinhao Cheng, Vineeth Kada, Ruohan Gao, Yingyi Huang, Remi Delacourt, April Yang, Yingcheng Wang, Mengdi Wu, et al. 2024. Flexllm: A system for co-serving large language model inference and parameter-efficient finetuning. arXiv preprint arXiv:2402.18789 (2024)."},{"key":"e_1_3_2_1_44_1","volume-title":"16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19)","author":"Ousterhout Amy","year":"2019","unstructured":"Amy Ousterhout, Joshua Fried, Jonathan Behrens, Adam Belay, and Hari Balakrishnan. 2019. Shenango: Achieving high CPU efficiency for latency-sensitive datacenter workloads. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19). 361\u2013378."},{"key":"e_1_3_2_1_45_1","volume-title":"Hierarchical Autoscaling for Large Language Model Serving with Chiron. arXiv preprint arXiv:2501.08090","author":"Patke Archit","year":"2025","unstructured":"Archit Patke, Dhemath Reddy, Saurabh Jha, Chandra Narayanaswami, Zbigniew Kalbarczyk, and Ravishankar Iyer. 2025. Hierarchical Autoscaling for Large Language Model Serving with Chiron. arXiv preprint arXiv:2501.08090 (2025)."},{"key":"e_1_3_2_1_46_1","volume-title":"Perplexity: AI-powered answer engine that provides accurate, trusted, and real-time answers to any question. https:\/\/www.perplexity.ai Accessed: 2025-05-15.","author":"Perplexity","year":"2025","unstructured":"Perplexity AI. 2025. Perplexity: AI-powered answer engine that provides accurate, trusted, and real-time answers to any question. https:\/\/www.perplexity.ai Accessed: 2025-05-15."},{"key":"e_1_3_2_1_47_1","volume-title":"Sam Altman: OpenAI Has Reached Roughly 800 Million Users. https:\/\/www.pymnts.com\/artificial-intelligence-2\/2025\/sam-altman-openai-has-reached-roughly-800-million-users\/ Accessed: 2025-05-13.","author":"PYMNTS.","year":"2025","unstructured":"PYMNTS. 2025. Sam Altman: OpenAI Has Reached Roughly 800 Million Users. https:\/\/www.pymnts.com\/artificial-intelligence-2\/2025\/sam-altman-openai-has-reached-roughly-800-million-users\/ Accessed: 2025-05-13."},{"key":"e_1_3_2_1_48_1","volume-title":"Ray Serve: Scalable and Programmable Serving for ML Models. https:\/\/docs.ray.io\/en\/latest\/serve\/index.html Accessed: 2025-05-13.","author":"Project Ray","year":"2025","unstructured":"Ray Project. 2025. Ray Serve: Scalable and Programmable Serving for ML Models. https:\/\/docs.ray.io\/en\/latest\/serve\/index.html Accessed: 2025-05-13."},{"key":"e_1_3_2_1_49_1","unstructured":"Replit Inc. 2025. Intro to Ghostwriter. https:\/\/replit.com\/learn\/intro-to-ghostwriter Accessed: 2025-05-12."},{"key":"e_1_3_2_1_50_1","unstructured":"Sherwood News. 2024. Just four companies are hoarding tens of billions of dollars worth of Nvidia GPU chips. https:\/\/sherwood.news\/tech\/companies-hoarding-nvidia-gpu-chips-meta-tesla\/ Accessed: 2025-05-14."},{"key":"e_1_3_2_1_51_1","unstructured":"Anton Shilov. 2024. TikTok owner ByteDance taps TSMC to make its own AI GPUs to stop relying on Nvidia. https:\/\/www.tomshardware.com\/tech-industry\/artificial-intelligence\/tiktok-owner-bytedance-taps-tsmc-to-make-its-own-ai-gpus-to-stop-relying-on-nvidia-the-company-has-reportedly-spent-over-dollar2-billion-on-nvidia-ai-gpus Accessed: 2025-05-12."},{"key":"e_1_3_2_1_52_1","volume-title":"Gim, Seung seob Lee, and Anurag Khandelwal.","author":"Soleimani Mahdi","year":"2025","unstructured":"Mahdi Soleimani, Grace Jia, In Gim, Seung seob Lee, and Anurag Khandelwal. 2025. Wiretapping LLMs: Network Side-Channel Attacks on Interactive LLM Services. Cryptology ePrint Archive, Paper 2025\/167. https:\/\/eprint.iacr.org\/2025\/167"},{"key":"e_1_3_2_1_53_1","volume-title":"Preble: Efficient Distributed Prompt Scheduling for LLM Serving.","author":"Srivatsa Vikranth","year":"2024","unstructured":"Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, and Yiying Zhang. 2024. Preble: Efficient Distributed Prompt Scheduling for LLM Serving. (2024). arXiv:2407.00023 [cs.DC] https:\/\/arxiv.org\/abs\/2407.00023"},{"key":"e_1_3_2_1_54_1","unstructured":"Steve Roberts. 2023. Amazon Code Whisperer Free for Individual Use is Now Generally Available. https:\/\/aws.amazon.com\/blogs\/aws\/amazon-codewhisperer-free-for-individual-use-is-now-generally-available\/ Accessed: 2025-05-12."},{"key":"e_1_3_2_1_55_1","volume-title":"Chord: a scalable peer-to-peer lookup protocol for internet applications","author":"Stoica Ion","year":"2003","unstructured":"Ion Stoica, Robert Morris, David Liben-Nowell, David R Karger, M Frans Kaashoek, Frank Dabek, and Hari Balakrishnan. 2003. Chord: a scalable peer-to-peer lookup protocol for internet applications. IEEE\/ACM Transactions on networking 11, 1 (2003), 17\u201332."},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"crossref","unstructured":"Jovan Stojkovic Chaojie Zhang \u00cd\u00f1igo Goiri Josep Torrellas and Esha Choukse. 2024. DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency. arXiv:2408.00741 [cs.AI] https:\/\/arxiv.org\/abs\/2408.00741","DOI":"10.1109\/HPCA61900.2025.00102"},{"key":"e_1_3_2_1_57_1","volume-title":"Tabnine: AI Code Assistant. https:\/\/www.tabnine.com\/ Accessed: 2025-05-12.","author":"Tabnine Inc.","year":"2025","unstructured":"Tabnine Inc. 2025. Tabnine: AI Code Assistant. https:\/\/www.tabnine.com\/ Accessed: 2025-05-12."},{"key":"e_1_3_2_1_58_1","volume-title":"Windsurf: The Most Powerful AI Code Editor","year":"2025","unstructured":"Windsurf. 2025. Windsurf: The Most Powerful AI Code Editor. https:\/\/windsurf.com\/ Accessed: 2025-05-12."},{"key":"e_1_3_2_1_59_1","unstructured":"Guanlong Wu Zheng Zhang Yao Zhang Weili Wang Jianyu Niu Ye Wu and Yinqian Zhang. 2025. I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM Serving. In NDSS. https:\/\/www.ndss-symposium.org\/ndss-paper\/i-know-what-you-asked-prompt-leakage-via-kv-cache-sharing-in-multi-tenant-llm-serving\/"},{"key":"e_1_3_2_1_60_1","unstructured":"Yuxing Xiang Xue Li Kun Qian Wenyuan Yu Ennan Zhai and Xin Jin. 2025. ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production. arXiv:2505.09999 [cs.DC] https:\/\/arxiv.org\/abs\/2505.09999"},{"key":"e_1_3_2_1_61_1","volume-title":"13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18)","author":"Xiao Wencong","year":"2018","unstructured":"Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, et al. 2018. Gandiva: Introspective cluster scheduling for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18). 595\u2013610."},{"key":"e_1_3_2_1_62_1","volume-title":"14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20)","author":"Xiao Wencong","year":"2020","unstructured":"Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020. AntMan: Dynamic scaling on GPU clusters for deep learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20). 533\u2013548."},{"key":"e_1_3_2_1_63_1","volume-title":"20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23)","author":"Yang Zongheng","year":"2023","unstructured":"Zongheng Yang, Zhanghao Wu, Michael Luo, Wei-Lin Chiang, Romil Bhardwaj, Woosuk Kwon, Siyuan Zhuang, Frank Sifei Luan, Gautam Mittal, Scott Shenker, et al. 2023. SkyPilot: An intercloud broker for sky computing. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23). 437\u2013455."},{"key":"e_1_3_2_1_64_1","unstructured":"Shunyu Yao Dian Yu Jeffrey Zhao Izhak Shafran Thomas L. Griffiths Yuan Cao and Karthik Narasimhan. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. arXiv:2305.10601 [cs.CL] https:\/\/arxiv.org\/abs\/2305.10601"},{"key":"e_1_3_2_1_65_1","volume-title":"Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)","author":"Yu Gyeong-In","year":"2022","unstructured":"Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022. Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22). USENIX Association, Carlsbad, CA, 521\u2013538. https:\/\/www.usenix.org\/conference\/osdi22\/presentation\/yu"},{"key":"e_1_3_2_1_66_1","volume-title":"Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving. arXiv preprint arXiv:2505.04021","author":"Yu Shan","year":"2025","unstructured":"Shan Yu, Jiarong Xing, Yifan Qiao, Mingyuan Ma, Yangmin Li, Yang Wang, Shuo Yang, Zhiqiang Xie, Shiyi Cao, Ke Bao, et al. 2025. Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving. arXiv preprint arXiv:2505.04021 (2025)."},{"key":"e_1_3_2_1_67_1","volume-title":"2019 USENIX Annual Technical Conference (USENIX ATC 19)","author":"Zhang Chengliang","year":"2019","unstructured":"Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019. MArk: Exploiting cloud services for Cost-Effective, SLO-Aware machine learning inference serving. In 2019 USENIX Annual Technical Conference (USENIX ATC 19). 1049\u20131062."},{"key":"e_1_3_2_1_68_1","volume-title":"The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Bl8u7ZRlbM","author":"Zhao Wenting","year":"2024","unstructured":"Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024. WildChat: 1M ChatGPT Interaction Logs in the Wild. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=Bl8u7ZRlbM"},{"key":"e_1_3_2_1_69_1","volume-title":"Muxflow: Efficient and safe gpu sharing in large-scale production deep learning clusters. arXiv preprint arXiv:2303.13803","author":"Zhao Yihao","year":"2023","unstructured":"Yihao Zhao, Xin Liu, Shufan Liu, Xiang Li, Yibo Zhu, Gang Huang, Xuanzhe Liu, and Xin Jin. 2023. Muxflow: Efficient and safe gpu sharing in large-scale production deep learning clusters. arXiv preprint arXiv:2303.13803 (2023)."},{"key":"e_1_3_2_1_70_1","unstructured":"Lianmin Zheng Wei-Lin Chiang Ying Sheng Siyuan Zhuang Zhanghao Wu Yonghao Zhuang Zi Lin Zhuohan Li Dacheng Li Eric. P Xing Hao Zhang Joseph E. Gonzalez and Ion Stoica. 2023. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. arXiv:2306.05685 [cs.CL]"},{"key":"e_1_3_2_1_71_1","first-page":"62557","article-title":"Sglang: Efficient execution of structured language model programs","volume":"37","author":"Zheng Lianmin","year":"2024","unstructured":"Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Livia Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al. 2024. Sglang: Efficient execution of structured language model programs. Advances in Neural Information Processing Systems 37 (2024), 62557\u201362583.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_72_1","first-page":"65517","article-title":"Response length perception and sequence scheduling: An llm-empowered llm inference pipeline","volume":"36","author":"Zheng Zangwei","year":"2023","unstructured":"Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo, Xin Jiang, and Yang You. 2023. Response length perception and sequence scheduling: An llm-empowered llm inference pipeline. Advances in Neural Information Processing Systems 36 (2023), 65517\u201365530.","journal-title":"Advances in Neural Information Processing Systems"}],"event":{"name":"EUROSYS '26: 21st European Conference on Computer Systems","location":"McEwan Hall\/The University of Edinburgh Edinburgh Scotland UK","acronym":"EUROSYS '26","sponsor":["SIGOPS ACM Special Interest Group on Operating Systems"]},"container-title":["Proceedings of the 21st European Conference on Computer Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3767295.3769353","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T12:24:04Z","timestamp":1780662244000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3767295.3769353"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,26]]},"references-count":72,"alternative-id":["10.1145\/3767295.3769353","10.1145\/3767295"],"URL":"https:\/\/doi.org\/10.1145\/3767295.3769353","relation":{},"subject":[],"published":{"date-parts":[[2026,4,26]]},"assertion":[{"value":"2026-04-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}