{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T22:39:20Z","timestamp":1779143960149,"version":"3.51.4"},"reference-count":38,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2020,4,27]],"date-time":"2020-04-27T00:00:00Z","timestamp":1587945600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"DFG project CELERITY","award":["360291326"],"award-info":[{"award-number":["360291326"]}]},{"DOI":"10.13039\/501100004543","name":"China Scholarship Council","doi-asserted-by":"publisher","award":["CSC201709110138"],"award-info":[{"award-number":["CSC201709110138"]}],"id":[{"id":"10.13039\/501100004543","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Computation"],"abstract":"<jats:p>Energy optimization is an increasingly important aspect of today\u2019s high-performance computing applications. In particular, dynamic voltage and frequency scaling (DVFS) has become a widely adopted solution to balance performance and energy consumption, and hardware vendors provide management libraries that allow the programmer to change both memory and core frequencies manually to minimize energy consumption while maximizing performance. This article focuses on modeling the energy consumption and speedup of GPU applications while using different frequency configurations. The task is not straightforward, because of the large set of possible and uniformly distributed configurations and because of the multi-objective nature of the problem, which minimizes energy consumption and maximizes performance. This article proposes a machine learning-based method to predict the best core and memory frequency configurations on GPUs for an input OpenCL kernel. The method is based on two models for speedup and normalized energy predictions over the default frequency configuration. Those are later combined into a multi-objective approach that predicts a Pareto-set of frequency configurations. Results show that our approach is very accurate at predicting extema and the Pareto set, and finds frequency configurations that dominate the default configuration in either energy or performance.<\/jats:p>","DOI":"10.3390\/computation8020037","type":"journal-article","created":{"date-parts":[[2020,4,28]],"date-time":"2020-04-28T05:05:32Z","timestamp":1588050332000},"page":"37","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Accurate Energy and Performance Prediction for Frequency-Scaled GPU Kernels"],"prefix":"10.3390","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0118-0974","authenticated-orcid":false,"given":"Kaijie","family":"Fan","sequence":"first","affiliation":[{"name":"Faculty of Electrical Engineering and Computer Science, Technische Universit\u00e4t Berlin, Einsteinufer 17-6.0G, 10587 Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8869-6705","authenticated-orcid":false,"given":"Biagio","family":"Cosenza","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of Salerno, 84084 Fisciano (Salerno), Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ben","family":"Juurlink","sequence":"additional","affiliation":[{"name":"Faculty of Electrical Engineering and Computer Science, Technische Universit\u00e4t Berlin, Einsteinufer 17-6.0G, 10587 Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,4,27]]},"reference":[{"key":"ref_1","unstructured":"Intel (2020, April 24). RAPL (Running Average Power Limit) Power Meter. Available online: https:\/\/01.org\/rapl-power-meter."},{"key":"ref_2","unstructured":"NVIDIA (2020, April 24). NVIDIA Management Library (NVML). Available online: https:\/\/developer.nvidia.com\/nvidia-management-library-nvml."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Fan, K., Cosenza, B., and Juurlink, B.H.H. (2019, January 5\u20138). Predictable GPUs Frequency Scaling for Energy and Performance. Proceedings of the 48th International Conference on Parallel Processing, ICPP, Kyoto, Japan.","DOI":"10.1145\/3337821.3337833"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Mei, X., Yung, L.S., Zhao, K., and Chu, X. (2013, January 3\u20136). A Measurement Study of GPU DVFS on Energy Conservation. Proceedings of the Workshop on Power-Aware Computing and Systems, Berkleley, CA, USA.","DOI":"10.1145\/2525526.2525852"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"e4143","DOI":"10.1002\/cpe.4143","article-title":"Evaluation of DVFS techniques on modern HPC processors and accelerators for energy-aware applications","volume":"29","author":"Calore","year":"2017","journal-title":"Concurr. Comput. Pract. Exp."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ge, R., Vogt, R., Majumder, J., Alam, A., Burtscher, M., and Zong, Z. (2013, January 1\u20134). Effects of Dynamic Voltage and Frequency Scaling on a K20 GPU. Proceedings of the 42nd International Conference on Parallel Processing, ICPP, Lyon, France.","DOI":"10.1109\/ICPP.2013.98"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Tiwari, A., Laurenzano, M., Peraza, J., Carrington, L., and Snavely, A. (2012, January 1\u20133). Green Queue: Customized Large-Scale Clock Frequency Scaling. Proceedings of the International Conference on Cloud and Green Computing, CGC, Xiangtan, China.","DOI":"10.1109\/CGC.2012.62"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Vysocky, O., Beseda, M., R\u00edha, L., Zapletal, J., Lysaght, M., and Kannan, V. (2017, January 22\u201325). MERIC and RADAR Generator: Tools for Energy Evaluation and Runtime Tuning of HPC Applications. Proceedings of the High Performance Computing in Science and Engineering\u2014Third International Conference, HPCSE, Karolinka, Czech Republic. Revised Selected Papers.","DOI":"10.1007\/978-3-319-97136-0_11"},{"key":"ref_9","unstructured":"Hamano, T., Endo, T., and Matsuoka, S. (2009, January 23\u201329). Power-aware dynamic task scheduling for heterogeneous accelerated clusters. Proceedings of the 23rd IEEE International Symposium on Parallel and Distributed Processing, IPDPS, Rome, Italy."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Lopes, A., Pratas, F., Sousa, L., and Ilic, A. (2017, January 24\u201325). Exploring GPU performance, power and energy-efficiency bounds with Cache-aware Roofline Modeling. Proceedings of the 2017 IEEE International Symposium on Performance Analysis of Systems and Software, ISPASS, Santa Rosa, CA, USA.","DOI":"10.1109\/ISPASS.2017.7975297"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Ma, K., Li, X., Chen, W., Zhang, C., and Wang, X. (2012, January 10\u201313). GreenGPU: A Holistic Approach to Energy Efficiency in GPU-CPU Heterogeneous Architectures. Proceedings of the 41st International Conference on Parallel Processing, ICPP, Pittsburgh, PA, USA.","DOI":"10.1109\/ICPP.2012.31"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Song, S., Lee, M., Kim, J., Seo, W., Cho, Y., and Ryu, S. (2014, January 24\u201328). Energy-efficient scheduling for memory-intensive GPGPU workloads. Proceedings of the Design, Automation & Test in Europe Conference & Exhibition, Dresden, Germany.","DOI":"10.7873\/DATE.2014.032"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Choi, J., and Vuduc, R.W. (2016, January 23\u201327). Analyzing the Energy Efficiency of the Fast Multipole Method Using a DVFS-Aware Energy Model. Proceedings of the IEEE International Parallel and Distributed Processing Symposium Workshops, Chicago, IL, USA.","DOI":"10.1109\/IPDPSW.2016.206"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Lee, J.H., Nigania, N., Kim, H., Patel, K., and Kim, H. (2015). OpenCL Performance Evaluation on Modern Multicore CPUs. Sci. Program., 2015.","DOI":"10.1155\/2015\/859491"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"834","DOI":"10.1016\/j.parco.2013.08.009","article-title":"An application-centric evaluation of OpenCL on multi-core CPUs","volume":"39","author":"Shen","year":"2013","journal-title":"Parallel Comput."},{"key":"ref_16","unstructured":"Harris, G., Panangadan, A.V., and Prasanna, V.K. (2016, January 16\u201318). GPU-Accelerated Parameter Optimization for Classification Rule Learning. Proceedings of the International Florida Artificial Intelligence Research Society Conference, FLAIRS, Key Largo, FL, USA."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Pohl, A., Cosenza, B., and Juurlink, B.H.H. (2019, January 21\u201325). Portable Cost Modeling for Auto-Vectorizers. Proceedings of the 27th IEEE International Symposium on Modeling, Analysis, and Simulation of Computer and Telecommunication Systems, MASCOTS 2019, Rennes, France.","DOI":"10.1109\/MASCOTS.2019.00046"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Panneerselvam, S., and Swift, M.M. (2016, January 11\u201315). Rinnegan: Efficient Resource Use in Heterogeneous Architectures. Proceedings of the 2016 International Conference on Parallel Architectures and Compilation, PACT, Haifa, Israel.","DOI":"10.1145\/2967938.2967964"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Wang, Q., and Chu, X. (2018, January 11\u201313). GPGPU Performance Estimation with Core and Memory Frequency Scaling. Proceedings of the 24th IEEE International Conference on Parallel and Distributed Systems, ICPADS 2018, Singapore.","DOI":"10.1109\/PADSW.2018.8645000"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Kofler, K., Grasso, I., Cosenza, B., and Fahringer, T. (2013, January 10\u201314). An automatic input-sensitive approach for heterogeneous task partitioning. Proceedings of the International Conference on Supercomputing, ICS\u201913, Eugene, OR, USA.","DOI":"10.1145\/2464996.2465007"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Ge, R., Feng, X., and Cameron, K.W. (2009, January 25\u201329). Modeling and evaluating energy-performance efficiency of parallel processing on multicore based power aware systems. Proceedings of the 23rd IEEE International Symposium on Parallel and Distributed Processing, IPDPS, Rome, Italy.","DOI":"10.1109\/IPDPS.2009.5160979"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Bhattacharyya, A., Kwasniewski, G., and Hoefler, T. (2015, January 18\u201321). Using Compiler Techniques to Improve Automatic Performance Modeling. Proceedings of the International Conference on Parallel Architecture and Compilation, San Francisco, CA, USA.","DOI":"10.1109\/PACT.2015.39"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"De Mesmay, F., Voronenko, Y., and P\u00fcschel, M. (2010, January 19\u201323). Offline library adaptation using automatically generated heuristics. Proceedings of the 24th IEEE International Symposium on Parallel and Distributed Processing, IPDPS, Atlanta, GA, USA.","DOI":"10.1109\/IPDPS.2010.5470479"},{"key":"ref_24","first-page":"104:1","article-title":"e-PAL: An Active Learning Approach to the Multi-Objective Optimization Problem","volume":"17","author":"Zuluaga","year":"2016","journal-title":"J. Mach. Learn. Res."},{"key":"ref_25","unstructured":"Grewe, D., and O\u2019Boyle, M.F.P. (April, January 26). A Static Task Partitioning Approach for Heterogeneous Systems Using OpenCL. Proceedings of the 20th International Conference on Compiler Construction, CC, Saarbr\u00fccken, Germany."},{"key":"ref_26","first-page":"3537","article-title":"Analytical Processor Performance and Power Modeling Using Micro-Architecture Independent Characteristics","volume":"65","author":"Eyerman","year":"2016","journal-title":"IEEE Trans. Comput."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Abe, Y., Sasaki, H., Kato, S., Inoue, K., Edahiro, M., and Peres, M. (2014, January 19\u201323). Power and Performance Characterization and Modeling of GPU-Accelerated Systems. Proceedings of the IEEE 28th International Parallel and Distributed Processing Symposium, Phoenix, AZ, USA.","DOI":"10.1109\/IPDPS.2014.23"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"159150","DOI":"10.1109\/ACCESS.2019.2951218","article-title":"GPU Static Modeling using PTX and Deep Structured Learning","volume":"7","author":"Guerreiro","year":"2019","journal-title":"IEEE Access"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Guerreiro, J., Ilic, A., Roma, N., and Tomas, P. (2018, January 24\u201328). GPGPU Power Modelling for Multi-Domain Voltage-Frequency Scaling. Proceedings of the 24th IEEE International Symposium on High-Performance Computing Architecture, HPCA, Vienna, Austria.","DOI":"10.1109\/HPCA.2018.00072"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wu, G.Y., Greathouse, J.L., Lyashevsky, A., Jayasena, N., and Chiou, D. (2015, January 2). GPGPU performance and power estimation using machine learning. Proceedings of the 21st IEEE International Symposium on High Performance Computer Architecture, HPCA 2015, Burlingame, CA, USA.","DOI":"10.1109\/HPCA.2015.7056063"},{"key":"ref_31","unstructured":"Isci, C., and Martonosi, M. (2003, January 5). Runtime Power Monitoring in High-End Processors: Methodology and Empirical Data. Proceedings of the 36th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO 36), San Diego, CA, USA."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Cummins, C., Petoumenos, P., Wang, Z., and Leather, H. (2017, January 4\u20138). Synthesizing benchmarks for predictive modeling. Proceedings of the International Symposium on Code Generation and Optimization, CGO, Austin, TX, USA.","DOI":"10.1109\/CGO.2017.7863731"},{"key":"ref_33","unstructured":"Cosenza, B., Durillo, J.J., Ermon, S., and Juurlink, B.H.H. (June, January 29). Autotuning Stencil Computations with Structural Ordinal Regression Learning. Proceedings of the IEEE International Parallel and Distributed Processing Symposium, IPDPS, Orlando, FL, USA."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1023\/B:STCO.0000035301.49549.88","article-title":"A tutorial on support vector regression","volume":"14","author":"Smola","year":"2004","journal-title":"Stat. Comput."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"13:1","DOI":"10.1145\/2792984","article-title":"Many-Objective Evolutionary Algorithms: A Survey","volume":"48","author":"Li","year":"2015","journal-title":"ACM Comput. Surv."},{"key":"ref_36","first-page":"1","article-title":"Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning","volume":"18","author":"Nogueira","year":"2017","journal-title":"J. Mach. Learn. Res."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1109\/TEVC.2003.810758","article-title":"Performance Assessment of Multiobjective Optimizers: An Analysis and Review","volume":"7","author":"Zitzler","year":"2003","journal-title":"Trans. Evol. Comp."},{"key":"ref_38","unstructured":"Zitzler, E. (1999). Evolutionary Algorithms for Multiobjective Optimization: Methods and Applications. [Ph.D. Thesis, University of Zurich]."}],"container-title":["Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2079-3197\/8\/2\/37\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T14:09:19Z","timestamp":1760364559000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2079-3197\/8\/2\/37"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,4,27]]},"references-count":38,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2020,6]]}},"alternative-id":["computation8020037"],"URL":"https:\/\/doi.org\/10.3390\/computation8020037","relation":{},"ISSN":["2079-3197"],"issn-type":[{"value":"2079-3197","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,4,27]]}}}