{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T13:47:19Z","timestamp":1782481639378,"version":"3.54.5"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T00:00:00Z","timestamp":1782432000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"Google PhD Fellowship 2022 program"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2026,6,30]]},"abstract":"<jats:p>Due to the significant cost associated with power consumption, hardware overprovisioning is widely used to cap processor power consumption and improve the average power utilization of servers in data centers and nodes in HPC clusters. However, uniformly capping power across sockets in multiprocessor servers can lead to performance degradation for co-running applications due to workload variability. Existing solutions primarily focus on cluster-level power management, making limited use of power scheduling within multi-socket servers or processor frequency scaling to regulate power consumption under power constraints.<\/jats:p>\n                  <jats:p>This article introduces Fulcrum, a novel power management library for co-running parallel applications on multi-socket, multi-core servers, independent of the underlying parallel programming model. Fulcrum dynamically redistributes power on a power-constrained multi-socket server to maximize throughput and fairness for co-running applications without requiring prior knowledge of application characteristics. Fulcrum periodically profiles hardware performance monitoring counters and independently adjusts core and uncore frequencies for each application based on its power sensitivity and degree of parallelism, thereby maximizing overall system throughput. After optimizing application-level power usage, Fulcrum enhances application-level fairness by dynamically redistributing power among applications, thereby maximizing aggregate throughput while adhering to the server\u2019s global power budget. We evaluated Fulcrum across various exascale proxy application mixes and power caps on a four-socket, 72-core Intel Cooper Lake processor. Our results show that Fulcrum improves system throughput (geometric mean) by 26.3% under low power caps and by up to 5.3% under higher power caps, while delivering power efficiency improvements of 27.7\u20138.4% at the respective power caps. Moreover, Fulcrum outperforms the two state-of-the-art approaches at both power caps, achieving geometric mean throughput improvements of 3.9\u201316.4% and power efficiency gains of 10.3\u201318.6%, while maintaining comparable fairness.<\/jats:p>\n                  <jats:p\/>","DOI":"10.1145\/3818685","type":"journal-article","created":{"date-parts":[[2026,5,26]],"date-time":"2026-05-26T11:41:57Z","timestamp":1779795717000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Power Scheduling for Maximizing Throughput and Fairness in Co-running Applications"],"prefix":"10.1145","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0891-7847","authenticated-orcid":false,"given":"Sunil","family":"Kumar","sequence":"first","affiliation":[{"name":"CSE, Indraprastha Institute of Information Technology Delhi","place":["New Delhi, India"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3042-4202","authenticated-orcid":false,"given":"Vivek","family":"Kumar","sequence":"additional","affiliation":[{"name":"CSE, Indraprastha Institute of Information Technology Delhi","place":["New Delhi, India"]}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1084-5683","authenticated-orcid":false,"given":"Sridutt","family":"Bhalachandra","sequence":"additional","affiliation":[{"name":"The University of North Carolina at Chapel Hill","place":["Chapel Hill, United States"]}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,26]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"[n.d.]. SchedMD. Retrieved December 26 2025 from https:\/\/slurm.schedmd.com\/SLUG15\/Power_mgmt.pdf"},{"key":"e_1_3_2_3_2","unstructured":"November 2025. TOP500. Retrieved December 26 2025 from https:\/\/top500.org\/lists\/top500\/2025\/11\/"},{"key":"e_1_3_2_4_2","volume-title":"ECP Proxy Applications","unstructured":"Released. ECP Proxy Applications. Retrieved December 26, 2025 from https:\/\/proxyapps.exascaleproject.org\/ecp-proxy-apps-suite\/"},{"key":"e_1_3_2_5_2","unstructured":"AMD. 2025. AMD HSMP. Retrieved April 10 2026 from https:\/\/github.com\/amd\/amd_hsmp"},{"key":"e_1_3_2_6_2","volume-title":"Arm\u00ae System Control and Management Interface (SCMI) Specification","author":"Limited Arm","unstructured":"Arm Limited. [n.d.]. Arm\u00ae System Control and Management Interface (SCMI) Specification. Arm Limited. Retrieved from https:\/\/developer.arm.com\/architectures\/system-architectures\/software-standards\/scmi"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.23919\/SpringSim.2019.8732878"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2017.114"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.suscom.2023.100865"},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/1840845.1840883"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581784.3607091"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-58667-0_21"},{"key":"e_1_3_2_13_2","volume-title":"PathFinder Graph Search Mini-App","unstructured":"ECP-copa. v1 release. PathFinder Graph Search Mini-App. Retrieved December 26, 2025 from https:\/\/github.com\/Mantevo\/PathFinder"},{"key":"e_1_3_2_14_2","volume-title":"A simple proxy for the force computations in typical molecular dynamics applications","unstructured":"ECP-copa. v2.0 release. A simple proxy for the force computations in typical molecular dynamics applications. Retrieved December 26, 2025 from https:\/\/github.com\/Mantevo\/miniMD"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2014.07.003"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807643"},{"key":"e_1_3_2_17_2","volume-title":"Co-design proxy application for the heterogeneous multiscale method","unstructured":"ExMatEx. commit id d8776bf. Co-design proxy application for the heterogeneous multiscale method. Retrieved December 26, 2025 from https:\/\/github.com\/exmatex\/CoHMM\/tree\/sad"},{"key":"e_1_3_2_18_2","volume-title":"Unstructured mesh hydrodynamics for advanced architectures","author":"Ferenbaugh Charles R.","unstructured":"Charles R. Ferenbaugh. v0.9 release. Unstructured mesh hydrodynamics for advanced architectures. Retrieved December 26, 2025 from https:\/\/github.com\/lanl\/PENNANT"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2005.57"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1145\/2967938.2967961"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356150"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1145\/3208040.3208047"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4302-6638-9"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW55747.2022.00164"},{"issue":"11","key":"e_1_3_2_25_2","first-page":"1","article-title":"Intel\u00ae 64 and ia-32 architectures software developer\u2019s manual","volume":"2","author":"Guide Part","year":"2011","unstructured":"Part Guide. 2011. Intel\u00ae 64 and ia-32 architectures software developer\u2019s manual. Volume 3B: System Programming Guide, Part 2, 11 (2011), 1\u201364.","journal-title":"Volume 3B: System Programming Guide, Part"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3302424.3303981"},{"key":"e_1_3_2_27_2","volume-title":"Performance Characterterics and Viability of the Method of Characteristics (MOC) for 3d Neutron Transport Calculations","author":"Gunow Geoffrey","unstructured":"Geoffrey Gunow and John Tramm. v4 release. Performance Characterterics and Viability of the Method of Characteristics (MOC) for 3d Neutron Transport Calculations. Retrieved December 26, 2025 from https:\/\/github.com\/ANL-CESAR\/SimpleMOC"},{"issue":"3","key":"e_1_3_2_28_2","article-title":"The uncore: A modular approach to feeding the high performance cores","volume":"14","author":"Hill David L.","year":"2010","unstructured":"David L. Hill, Derek Bachand, Selim Bilgin, Robert Greiner, Per Hammarlund, Thomas Huff, Steve Kulick, and Robert Safranek. 2010. The uncore: A modular approach to feeding the high performance cores. Intel Technology Journal 14, 3 (2010), 30\u201349.","journal-title":"Intel Technology Journal"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1145\/3173162.3173190"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid59990.2024.00039"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICAC.2019.00015"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807638"},{"key":"e_1_3_2_33_2","unstructured":"Intel. Accessed 2021. Intel xeon processor E5 v3 family uncore performance monitoring. Retrieved December 26 2025 from https:\/\/www.intel.com\/content\/dam\/www\/public\/us\/en\/zip\/xeon-e5-v3-uncore-performance-monitoring.zip"},{"key":"e_1_3_2_34_2","first-page":"1","volume-title":"Proceedings of the ATM Forum Contribution","author":"Jain Raj","year":"1999","unstructured":"Raj Jain, Arjan Durresi, and Gojko Babic. 1999. Throughput fairness index: An explanation. In Proceedings of the ATM Forum Contribution. 1\u201313."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2008.4536223"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2008.4658633"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3476163"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-99854-6_24"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10586-007-0045-4"},{"key":"e_1_3_2_40_2","volume-title":"Finite Element Mini-Application","unstructured":"Mantevo. v2.2.0 release. Finite Element Mini-Application. Retrieved December 26, 2025 from https:\/\/github.com\/Mantevo\/miniFE"},{"key":"e_1_3_2_41_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-20119-1_28"},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS47924.2020.00086"},{"key":"e_1_3_2_43_2","volume-title":"OpenMP Application Programming Interface Version 5.0","author":"ARB OpenMP","year":"2018","unstructured":"OpenMP ARB. November 2018. OpenMP Application Programming Interface Version 5.0. Retrieved December 26, 2025 from https:\/\/www.openmp.org\/wp-content\/uploads\/OpenMP-API-Specification-5.0.pdf"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3307681.3326607"},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1145\/2464996.2465009"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1145\/3721145.3734532"},{"key":"e_1_3_2_47_2","unstructured":"Xu Peter. 2025. Tuning UEFI Settings for Performance and Energy Efficiency on 5th Gen AMD EPYC Processor-Based ThinkSystem Servers. Retrieved December 26 2025 from https:\/\/lenovopress.lenovo.com\/lp2210.pdf"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/1542275.1542340"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3453483.3454109"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00031"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2013.6702684"},{"key":"e_1_3_2_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/2594291.2594292"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3545008.3545047"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1145\/1346281.1346317"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-24449-0_22"},{"key":"e_1_3_2_56_2","volume-title":"Proceedings of the High Performance Computing Symposium (HPC\u201918)","author":"Sundriyal Vaibhav","year":"2018","unstructured":"Vaibhav Sundriyal, Masha Sosonkina, Bryce M. Westheimer, and Mark Gordon. 2018. Comparisons of core and uncore frequency scaling modes in quantum chemistry application gamess. In Proceedings of the High Performance Computing Symposium (HPC\u201918). Society for Computer Simulation International, Article 13 (2018), 11 pages."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/CGC.2012.62"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-15976-8_3"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-57675-2_5"},{"key":"e_1_3_2_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/SAMOS.2016.7818333"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1007\/10968987_3"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1145\/2872362.2872375"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3225058.3225098"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3295500.3356174"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1145\/3712285.3759879"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2017.124"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3818685","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,26]],"date-time":"2026-06-26T12:56:33Z","timestamp":1782478593000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3818685"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,26]]},"references-count":65,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,30]]}},"alternative-id":["10.1145\/3818685"],"URL":"https:\/\/doi.org\/10.1145\/3818685","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,26]]},"assertion":[{"value":"2025-12-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-26","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}