{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T16:46:25Z","timestamp":1782405985321,"version":"3.54.5"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T00:00:00Z","timestamp":1782345600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62202123, and 62402141"],"award-info":[{"award-number":["62202123, and 62402141"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"Open Research Fund of Pengcheng Laboratory","doi-asserted-by":"publisher","award":["2025KF1A0020"],"award-info":[{"award-number":["2025KF1A0020"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"Joint Funds of the National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U22A2036"],"award-info":[{"award-number":["U22A2036"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>\n                    The proliferation of deep learning inference services in power-constrained environments necessitates GPU management strategies that maximize throughput within strict power envelopes. Existing approaches often treat frequency scaling and resource partitioning as orthogonal problems or rely on static hardware assumptions, leading to suboptimal energy efficiency. This article presents\n                    <jats:sc>PctoDL<\/jats:sc>\n                    , a power-aware scheduling system that maximizes aggregate inference throughput by jointly optimizing spatial resource partitioning, batch size, and SM\/memory frequency settings. To address the throughput\u2013power tradeoff in power-constrained multi-tenant inference,\n                    <jats:sc>PctoDL<\/jats:sc>\n                    couples resource partitioning with coordinated frequency control under a fixed power cap. It combines a physics-informed iterative greedy partitioning algorithm, a thermodynamic model-predictive controller for runtime frequency regulation, and an online joint optimization mechanism for adaptive refinement. On the NVIDIA RTX 3080 Ti platform,\n                    <jats:sc>PctoDL<\/jats:sc>\n                    improves average throughput over BatchDVFS by 108.41%, with a peak gain of 262.74%. On the NVIDIA A100 platform, it delivers an average gain of 19.74% and a maximum gain of 57.03%. Compared with Morak\u2019s coarse-grained partitioning approach,\n                    <jats:sc>PctoDL<\/jats:sc>\n                    achieves average\/peak gains of 79.05%\/137.93% on the RTX 3080 Ti and 26.33%\/70.21% on the A100.\n                  <\/jats:p>","DOI":"10.1145\/3805802","type":"journal-article","created":{"date-parts":[[2026,4,13]],"date-time":"2026-04-13T11:06:09Z","timestamp":1776078369000},"page":"1-25","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["PctoDL: Adaptive GPU Throughput Optimization for Deep Learning Inference with Power Constraints"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0043-4370","authenticated-orcid":false,"given":"Meng","family":"Hao","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-8077-0047","authenticated-orcid":false,"given":"Zikun","family":"Wu","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-0415-327X","authenticated-orcid":false,"given":"Xueyang","family":"Tian","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-0071-6218","authenticated-orcid":false,"given":"Siyu","family":"Yang","sequence":"additional","affiliation":[{"name":"Computer Science and Technology, Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-0954-4001","authenticated-orcid":false,"given":"Guotong","family":"Guo","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8386-0131","authenticated-orcid":false,"given":"Hongwei","family":"Yang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6494-775X","authenticated-orcid":false,"given":"Hui","family":"He","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3298-5134","authenticated-orcid":false,"given":"Yiming","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Cyberspace Science, Harbin Industry University","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2478-2153","authenticated-orcid":false,"given":"Farui","family":"Wang","sequence":"additional","affiliation":[{"name":"Computer Science and Technology, Harbin Institute of Technology Shenzhen","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7502-7094","authenticated-orcid":false,"given":"Desheng","family":"Wang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Shenzhen, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4783-876X","authenticated-orcid":false,"given":"Weizhe","family":"Zhang","sequence":"additional","affiliation":[{"name":"Harbin Institute of Technology","place":["Harbin, China"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,25]]},"reference":[{"key":"e_1_3_1_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/IGCC.2018.8752132"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.312"},{"key":"e_1_3_1_4_2","first-page":"199","volume-title":"Proceedings of the 2022 USENIX Annual Technical Conference","author":"Choi Seungbeom","year":"2022","unstructured":"Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh. 2022. Serving heterogeneous machine learning models on \\(\\lbrace\\) Multi-GPU \\(\\rbrace\\) servers with \\(\\lbrace\\) Spatio-Temporal \\(\\rbrace\\) sharing. In Proceedings of the 2022 USENIX Annual Technical Conference. 199\u2013216."},{"key":"e_1_3_1_5_2","first-page":"613","volume-title":"Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation","author":"Crankshaw Daniel","year":"2017","unstructured":"Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J. Franklin, Joseph E. Gonzalez, and Ion Stoica. 2017. Clipper: A \\(\\lbrace\\) low-latency \\(\\rbrace\\) online prediction serving system. In Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation. 613\u2013627."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2020.3047638"},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/3419111.3421284"},{"key":"e_1_3_1_9_2","volume-title":"Caltech-256 Object Category Dataset","author":"Griffin Gregory","year":"2007","unstructured":"Gregory Griffin, Alex Holub, Pietro Perona, et\u00a0al. 2007. Caltech-256 Object Category Dataset. Technical Report. Technical Report 7694, California Institute of Technology Pasadena."},{"key":"e_1_3_1_10_2","first-page":"443","volume-title":"Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation","author":"Gujarati Arpan","year":"2020","unstructured":"Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020. Serving \\(\\lbrace\\) DNNs \\(\\rbrace\\) like clockwork: Performance predictability from the bottom up. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation. 443\u2013462."},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00059"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_1_14_2","unstructured":"Brody Huval Tao Wang Sameep Tandon Jeff Kiske Will Song Joel Pazhayampallil Mykhaylo Andriluka Pranav Rajpurkar Toki Migimatsu Royce Cheng-Yue Fernando Mujica Adam Coates and Andrew Y. Ng. 2015. An empirical evaluation of deep learning on highway driving. Retrieved from https:\/\/arxiv.org\/abs\/1504.01716"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.parco.2022.102954"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/MASS58611.2023.00074"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2021.3084813"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3772052.3772228"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2022.3144614"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.13"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/Cloud-Summit61220.2024.00018"},{"key":"e_1_3_1_22_2","unstructured":"NVIDIA Corporation.2019. Multi-Process Service . Retrieved April 5 2026 from https:\/\/docs.nvidia.com\/deploy\/pdf\/CUDA_Multi_Process_Service_Overview.pdf"},{"key":"e_1_3_1_23_2","unstructured":"NVIDIA Corporation.2022. NVIDIA Multi-Instance GPU User Guide. Retrieved April 5 2026 from https:\/\/docs.nvidia.com\/datacenter\/tesla\/mig-user-guide\/"},{"key":"e_1_3_1_24_2","unstructured":"NVIDIA Corporation.2024. NVIDIA Nsight compute . Retrieved April 5 2026 from https:\/\/docs.nvidia.com\/nsight-compute\/NsightCompute\/index.html"},{"key":"e_1_3_1_25_2","unstructured":"NVIDIA Corporation. 2024. NVML API Reference Manual. Retrieved April 5 2026 from https:\/\/docs.nvidia.com\/deploy\/nvml-api\/index.html. Accessed April 14 2024."},{"key":"e_1_3_1_26_2","article-title":"Pytorch: An imperative style, high-performance deep learning library","volume":"32","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et\u00a0al. 2019. Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_27_2","article-title":"Faster r-cnn: Towards real-time object detection with region proposal networks","volume":"28","author":"Ren Shaoqing","year":"2015","unstructured":"Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015. Faster r-cnn: Towards real-time object detection with region proposal networks. Advances in Neural Information Processing Systems 28 (2015).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00474"},{"key":"e_1_3_1_29_2","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very deep convolutional networks for large-scale image recognition. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.52"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/3627703.3629578"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00293"},{"key":"e_1_3_1_33_2","first-page":"10096","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Tan Mingxing","year":"2021","unstructured":"Mingxing Tan and Quoc Le. 2021. Efficientnetv2: Smaller models and faster training. In Proceedings of the International Conference on Machine Learning. PMLR, 10096\u201310106."},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCC\/SmartCity\/DSS.2019.00334"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1145\/3307772.3328315"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20053-3_27"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSUSC.2023.3314916"},{"issue":"11","key":"e_1_3_1_38_2","first-page":"2943","article-title":"Dynamic GPU energy optimization for machine learning training workloads","volume":"33","author":"Wang Farui","year":"2021","unstructured":"Farui Wang, Weizhe Zhang, Shichao Lai, Meng Hao, and Zheng Wang. 2021. Dynamic GPU energy optimization for machine learning training workloads. IEEE Transactions on Parallel and Distributed Systems 33, 11 (2021), 2943\u20132954.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","unstructured":"Yiming Wang Weizhe Zhang Meng Hao Weizhi Kong and Yuan Wen. 2025. Dynamic power management through multi-agent deep reinforcement learning for heterogeneous systems. ACM Transactions on Architecture and Code Optimization 22 2 (2025) 1\u201323.","DOI":"10.1145\/3716872"},{"key":"e_1_3_1_40_2","first-page":"69","volume-title":"Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation","author":"Wu Bingyang","year":"2023","unstructured":"Bingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu, and Xin Jin. 2023. Transparent \\(\\lbrace\\) GPU \\(\\rbrace\\) sharing in container clouds for deep learning workloads. In Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation. 69\u201385."},{"key":"e_1_3_1_41_2","unstructured":"Yonghui Wu Mike Schuster Zhifeng Chen Quoc V. Le Mohammad Norouzi Wolfgang Macherey Maxim Krikun Yuan Cao Qin Gao Klaus Macherey Jeff Klingner Apurva Shah Melvin Johnson Xiaobing Liu \u0141ukasz Kaiser Stephan Gouws Yoshikiyo Kato Taku Kudo Hideto Kazawa Keith Stevens George Kurian Nishant Patil Wei Wang Cliff Young Jason Smith Jason Riesa Alex Rudnick Oriol Vinyals Greg Corrado Macduff Hughes and Jeffrey Dean. 2016. Google\u2019s neural machine translation system: Bridging the gap between human and machine translation. Retrieved from https:\/\/arxiv.org\/abs\/1609.08144"},{"key":"e_1_3_1_42_2","first-page":"38087","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xiao Guangxuan","year":"2023","unstructured":"Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. Smoothquant: Accurate and efficient post-training quantization for large language models. In Proceedings of the International Conference on Machine Learning. PMLR, 38087\u201338099."},{"key":"e_1_3_1_43_2","first-page":"533","volume-title":"Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation","author":"Xiao Wencong","year":"2020","unstructured":"Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020. \\(\\lbrace\\) AntMan \\(\\rbrace\\) : Dynamic scaling on \\(\\lbrace\\) GPU \\(\\rbrace\\) clusters for deep learning. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation. 533\u2013548."},{"key":"e_1_3_1_44_2","first-page":"2397","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xiong Caiming","year":"2016","unstructured":"Caiming Xiong, Stephen Merity, and Richard Socher. 2016. Dynamic memory networks for visual and textual question answering. In Proceedings of the International Conference on Machine Learning. PMLR, 2397\u20132406."},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2022.3232715"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/2925426.2926265"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3332466.3374520"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2022.01.004"},{"key":"e_1_3_1_49_2","first-page":"119","volume-title":"Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation","author":"You Jie","year":"2023","unstructured":"Jie You, Jae-Won Chung, and Mosharaf Chowdhury. 2023. Zeus: Understanding and optimizing \\(\\lbrace\\) GPU \\(\\rbrace\\) energy consumption of \\(\\lbrace\\) DNN \\(\\rbrace\\) training. In Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation. 119\u2013139."},{"key":"e_1_3_1_50_2","doi-asserted-by":"crossref","unstructured":"Yongkang Zhang Haoxuan Yu Chenxia Han Cheng Wang Baotong Lu Yunzhe Li Zhifeng Jiang Yang Li Xiaowen Chu and Huaicheng Li. 2025. SGDRC: Software-defined dynamic resource control for concurrent DNN inference on nvidia gpus. In Proceedings of the 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming. 267\u2013281.","DOI":"10.1145\/3710848.3710863"},{"key":"e_1_3_1_51_2","doi-asserted-by":"crossref","unstructured":"Qinxin Zhou Yunfang Zhang Xinzi Xu Qichen Zhang Huaying Wu Yong Lian and Yang Zhao. 2024. A lightweight DRDPG-based RL DVFS for video rendering on CPU-GPU integrated SoC. IEEE Transactions on Circuits and Systems I: Regular Papers 71 5 (2024) 2119\u20132131.","DOI":"10.1109\/TCSI.2024.3351791"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2015.7056028"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3805802","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,25]],"date-time":"2026-06-25T15:54:45Z","timestamp":1782402885000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805802"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,25]]},"references-count":51,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3805802"],"URL":"https:\/\/doi.org\/10.1145\/3805802","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,25]]},"assertion":[{"value":"2025-07-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-03-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-25","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}