{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T03:35:46Z","timestamp":1782963346838,"version":"3.54.5"},"reference-count":47,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2023,4,20]],"date-time":"2023-04-20T00:00:00Z","timestamp":1681948800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001602","name":"Science Foundation Ireland","doi-asserted-by":"publisher","award":["20\/FFP-P\/8683"],"award-info":[{"award-number":["20\/FFP-P\/8683"]}],"id":[{"id":"10.13039\/501100001602","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001602","name":"Science Foundation Ireland","doi-asserted-by":"publisher","award":["21\/RDD\/664"],"award-info":[{"award-number":["21\/RDD\/664"]}],"id":[{"id":"10.13039\/501100001602","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001603","name":"Sustainable Energy Authority of Ireland (SEAI)","doi-asserted-by":"publisher","award":["20\/FFP-P\/8683"],"award-info":[{"award-number":["20\/FFP-P\/8683"]}],"id":[{"id":"10.13039\/501100001603","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001603","name":"Sustainable Energy Authority of Ireland (SEAI)","doi-asserted-by":"publisher","award":["21\/RDD\/664"],"award-info":[{"award-number":["21\/RDD\/664"]}],"id":[{"id":"10.13039\/501100001603","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>The energy consumption of Information and Communications Technology (ICT) presents a new grand technological challenge. The two main approaches to tackle the challenge include the development of energy-efficient hardware and software. The development of energy-efficient software employing application-level energy optimization techniques has become an important category owing to the paradigm shift in the composition of digital platforms from single-core processors to heterogeneous platforms integrating multicore CPUs and graphics processing units (GPUs). In this work, we present an overview of application-level bi-objective optimization methods for energy and performance that address two fundamental challenges, non-linearity and heterogeneity, inherent in modern high-performance computing (HPC) platforms. Applying the methods requires energy profiles of the application\u2019s computational kernels executing on the different compute devices of the HPC platform. Therefore, we summarize the research innovations in the three mainstream component-level energy measurement methods and present their accuracy and performance tradeoffs. Finally, scaling the optimization methods for energy and performance is crucial to achieving energy efficiency objectives and meeting quality-of-service requirements in modern HPC platforms and cloud computing infrastructures. We introduce the building blocks needed to achieve this scaling and conclude with the challenges to scaling. Briefly, two significant challenges are described, namely fast optimization methods and accurate component-level energy runtime measurements, especially for components running on accelerators.<\/jats:p>","DOI":"10.3390\/info14040248","type":"journal-article","created":{"date-parts":[[2023,4,20]],"date-time":"2023-04-20T05:35:46Z","timestamp":1681968946000},"page":"248","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["Energy-Efficient Parallel Computing: Challenges to Scaling"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9460-3897","authenticated-orcid":false,"given":"Alexey","family":"Lastovetsky","sequence":"first","affiliation":[{"name":"School of Computer Science, University College Dublin, D04V1W8 Dublin, Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9181-3290","authenticated-orcid":false,"given":"Ravi Reddy","family":"Manumachu","sequence":"additional","affiliation":[{"name":"School of Computer Science, University College Dublin, D04V1W8 Dublin, Ireland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,4,20]]},"reference":[{"key":"ref_1","first-page":"19","article-title":"New perspectives on internet electricity use in 2030","volume":"3","author":"Andrae","year":"2020","journal-title":"Eng. Appl. Sci. Lett."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1119","DOI":"10.1109\/TPDS.2016.2608824","article-title":"New Model-Based Methods and Algorithms for Performance and Energy Optimization of Data Parallel Applications on Homogeneous Multicore Clusters","volume":"28","author":"Lastovetsky","year":"2017","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"160","DOI":"10.1109\/TC.2017.2742513","article-title":"Bi-Objective Optimization of Data-Parallel Applications on Homogeneous Multicore Clusters for Performance and Energy","volume":"67","author":"Reddy","year":"2018","journal-title":"IEEE Trans. Comput."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Miettinen, K. (1998). Nonlinear Multiobjective Optimization, Springer.","DOI":"10.1007\/978-1-4615-5563-6"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Talbi, E.G. (2009). Metaheuristics: From Design to Implementation, John Wiley & Sons.","DOI":"10.1002\/9780470496916"},{"key":"ref_6","unstructured":"Fahad, M., and Manumachu, R.R. (2023). HCLWattsUp: Energy API Using System-Level Physical Power Measurements Provided by Power Meters, Heterogeneous Computing Laboratory, University College Dublin."},{"key":"ref_7","unstructured":"OpenBLAS (2022, December 01). OpenBLAS: An Optimized BLAS Library. Available online: https:\/\/github.com\/xianyi\/OpenBLAS."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Fahad, M., Shahid, A., Reddy, R., and Lastovetsky, A. (2019). A Comparative Study of Methods for Measurement of Energy of Computing. Energies, 12.","DOI":"10.3390\/en12112204"},{"key":"ref_9","unstructured":"Top500 (2022, December 01). The Top500 Supercomputers List. Available online: https:\/\/www.top500.org."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1007\/s11265-015-1051-z","article-title":"OpenDwarfs: Characterization of Dwarf-Based Benchmarks on Fixed and Reconfigurable Architectures","volume":"85","author":"Krommydas","year":"2016","journal-title":"J. Signal Process. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1046","DOI":"10.1007\/s10766-016-0464-z","article-title":"GHOST: Building Blocks for High Performance Sparse Linear Algebra on Heterogeneous Systems","volume":"45","author":"Kreutzer","year":"2016","journal-title":"Int. J. Parallel Program."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1490","DOI":"10.1016\/j.cma.2011.01.013","article-title":"A new era in scientific computing: Domain decomposition methods in hybrid CPU\u2013GPU architectures","volume":"200","author":"Papadrakakis","year":"2011","journal-title":"Comput. Methods Appl. Mech. Eng."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"543","DOI":"10.1109\/TPDS.2020.3027338","article-title":"Bi-Objective Optimization of Data-Parallel Applications on Heterogeneous HPC Platforms for Performance and Energy Through Workload Distribution","volume":"32","author":"Khaleghzadeh","year":"2021","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Khaleghzadeh, H., Reddy, R., and Lastovetsky, A. (2022). Efficient exact algorithms for continuous bi-objective performance-energy optimization of applications with linear energy and monotonically increasing performance profiles on heterogeneous high performance computing platforms. Concurr. Comput. Pract. Exp., e7285.","DOI":"10.1002\/cpe.7285"},{"key":"ref_15","unstructured":"FFTW (2022, December 01). FFTW: A Fast, Free C FFT Library. Available online: https:\/\/www.fftw.org."},{"key":"ref_16","unstructured":"Lastovetsky, A.L., and Reddy, R. (2004, January 13\u201315). Data partitioning with a realistic performance model of networks of heterogeneous computers. Proceedings of the Parallel and Distributed Processing Symposium, Hong Kong, China."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"76","DOI":"10.1177\/1094342006074864","article-title":"Data partitioning with a functional performance model of heterogeneous processors","volume":"21","author":"Lastovetsky","year":"2007","journal-title":"Int. J. High Perform. Comput. Appl."},{"key":"ref_18","unstructured":"Lastovetsky, A., and Twamley, J. (2005). High Performance Computational Science and Engineering, Springer."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"2176","DOI":"10.1109\/TPDS.2018.2827055","article-title":"A novel data-partitioning algorithm for performance optimization of data-parallel applications on heterogeneous HPC platforms","volume":"29","author":"Khaleghzadeh","year":"2018","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"e5928","DOI":"10.1002\/cpe.5928","article-title":"A novel data partitioning algorithm for dynamic energy optimization on heterogeneous high-performance computing platforms","volume":"32","author":"Khaleghzadeh","year":"2020","journal-title":"Concurr. Comput. Pract. Exp."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"93793","DOI":"10.1109\/ACCESS.2020.2994953","article-title":"Accurate Energy Modelling of Hybrid Parallel Applications on Modern Heterogeneous Computing Platforms Using System-Level Measurements","volume":"8","author":"Fahad","year":"2020","journal-title":"IEEE Access"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"20","DOI":"10.1109\/MM.2012.12","article-title":"Power-Management Architecture of the Intel Microarchitecture Code-Named Sandy Bridge","volume":"32","author":"Rotem","year":"2012","journal-title":"IEEE Micro"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Treibig, J., Hager, G., and Wellein, G. (2010, January 13\u201316). Likwid: A lightweight performance-oriented tool suite for x86 multicore environments. Proceedings of the Parallel Processing Workshops (ICPPW), San Diego, CA, USA.","DOI":"10.1109\/ICPPW.2010.38"},{"key":"ref_24","unstructured":"Devices, A.M. (2022, December 01). AMD uProf User Guide. Available online: https:\/\/www.amd.com\/content\/dam\/amd\/en\/documents\/developer\/uprof-v4.0-gaGA-user-guide.pdf."},{"key":"ref_25","unstructured":"Devices, A.M. (2022, December 01). BIOS and Kernel Developer\u2019s Guide (BKDG) for AMD Family 15h Models 00h-0Fh Processors. Available online: https:\/\/www.amd.com\/system\/files\/TechDocs\/42301_15h_Mod_00h-0Fh_BKDG.pdf."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Hackenberg, D., Ilsche, T., Sch\u00f6ne, R., Molka, D., Schmidt, M., and Nagel, W.E. (2013, January 21\u201323). Power measurement techniques on standard compute nodes: A quantitative comparison. Proceedings of the 2013 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS), Austin, TX, USA.","DOI":"10.1109\/ISPASS.2013.6557170"},{"key":"ref_27","unstructured":"Intel Corporation (2022, December 01). Intel Xeon Phi Coprocessor System Software Developers Guide. Available online: https:\/\/www.intel.com\/content\/dam\/develop\/external\/us\/en\/documents\/intel-xeon-phi-coprocessor-quick-start-developers-guide-windows-v1-2.pdf."},{"key":"ref_28","unstructured":"Intel Corporation (2022, December 01). Intel Manycore Platform Software Stack (Intel MPSS). Available online: https:\/\/www.intel.com\/content\/www\/us\/en\/developer\/articles\/tool\/manycore-platform-software-stack-mpss.html."},{"key":"ref_29","unstructured":"Nvidia (2022, December 01). Nvidia Management Library: NVML API Reference Guide. Available online: https:\/\/docs.nvidia.com\/deploy\/nvml-api\/index.html."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2983575","article-title":"A Survey of Power and Energy Predictive Models in HPC Systems and Applications","volume":"50","author":"Pietri","year":"2017","journal-title":"ACM Comput. Surv."},{"key":"ref_31","first-page":"50","article-title":"Additivity: A Selection Criterion for Performance Events for Reliable Energy Predictive Modeling","volume":"4","author":"Shahid","year":"2017","journal-title":"Supercomput. Front. Innov. Int. J."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"63149","DOI":"10.1109\/ACCESS.2021.3075139","article-title":"Energy Predictive Models of Computing: Theory, Practical Implications and Experimental Analysis on Multicore Processors","volume":"9","author":"Shahid","year":"2021","journal-title":"IEEE Access"},{"key":"ref_33","unstructured":"Malyshkin, V. Improving the Accuracy of Energy Predictive Models for Multicore CPUs Using Additivity of Performance Monitoring Counters. Proceedings of the Parallel Computing Technologies."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1016\/j.jpdc.2021.01.007","article-title":"Improving the accuracy of energy predictive models for multicore CPUs by combining utilization and performance events model variables","volume":"151","author":"Shahid","year":"2021","journal-title":"J. Parallel Distrib. Comput."},{"key":"ref_35","unstructured":"NVIDIA Corporation (2022, December 01). CUDA Profiling Tools Interface (CUPTI)\u20141.0. Available online: https:\/\/developer.nvidia.com\/cuda-profiling-tools-interface."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Nagasaka, H., Maruyama, N., Nukada, A., Endo, T., and Matsuoka, S. (2010, January 15\u201318). Statistical power modeling of GPU kernels using performance counters. Proceedings of the International Green Computing Conference and Workshops (IGCC), Chicago, IL, USA.","DOI":"10.1109\/GREENCOMP.2010.5598315"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Hu, Y., Li, B., and Peng, L. (2011, January 28\u201330). Performance and Power Analysis of ATI GPU: A Statistical Approach. Proceedings of the IEEE Sixth International Conference on Networking, Architecture, and Storage, IEEE Computer Society, Washington, DC, USA.","DOI":"10.1109\/NAS.2011.51"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Song, S., Su, C., Rountree, B., and Cameron, K.W. (2013, January 20\u201324). A Simplified and Accurate Model of Power-Performance Efficiency on Emergent GPU Architectures. Proceedings of the 27th IEEE International Parallel and Distributed Processing Symposium (IPDPS), Cambridge, MA, USA.","DOI":"10.1109\/IPDPS.2013.73"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Al-Hashimi, M., Saleh, M., Abulnaja, O., and Aljabri, N. (2014, January 15\u201317). Evaluation of control loop statements power efficiency: An experimental study. Proceedings of the 2014 9th International Conference on Informatics and Systems, Cairo, Egypt.","DOI":"10.1109\/INFOS.2014.7036676"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Coplin, J., and Burtscher, M. (2015, January 7). Effects of Source-Code Optimizations on GPU Performance and Energy Consumption. Proceedings of the 8th Workshop on General Purpose Processing Using GPUs, San Francisco, CA, USA.","DOI":"10.1145\/2716282.2716292"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Lee, S., Kim, K., Koo, G., Jeon, H., Ro, W.W., and Annavaram, M. (2015, January 13\u201317). Warped-Compression: Enabling power efficient GPUs through register compression. Proceedings of the 2015 ACM\/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA), Portland, OR, USA.","DOI":"10.1145\/2749469.2750417"},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"14919","DOI":"10.1007\/s11227-022-04473-9","article-title":"Investigating the effect of varying block size on power and energy consumption of GPU kernels","volume":"78","author":"Ikram","year":"2022","journal-title":"J. Supercomput."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"2506","DOI":"10.1109\/TC.2014.2375202","article-title":"Data partitioning on multicore and multi-GPU platforms using functional performance models","volume":"64","author":"Zhong","year":"2015","journal-title":"Comput. IEEE Trans."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"David, H., Gorbatov, E., Hanebutte, U.R., Khanna, R., and Le, C. (2010, January 18\u201320). RAPL: Memory power estimation and capping. Proceedings of the 2010 ACM\/IEEE International Symposium on Low-Power Electronics and Design (ISLPED), Austin, TX, USA.","DOI":"10.1145\/1840845.1840883"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Gough, C., Steiner, I., and Saunders, W. (2015). Energy Efficient Servers Blueprints for Data Center Optimization, Apress.","DOI":"10.1007\/978-1-4302-6638-9"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Hackenberg, D., Sch\u00f6ne, R., Ilsche, T., Molka, D., Schuchart, J., and Geyer, R. (2015, January 25\u201329). An Energy Efficiency Feature Survey of the Intel Haswell Processor. Proceedings of the 2015 IEEE International Parallel and Distributed Processing Symposium Workshop, Washington, DC, USA.","DOI":"10.1109\/IPDPSW.2015.70"},{"key":"ref_47","unstructured":"Reddy, R., and Lastovetsky, A. (June, January 30). On Energy Nonproportionality of CPUs and GPUs. Proceedings of the 2022 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW), Lyon, France."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/4\/248\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:19:44Z","timestamp":1760123984000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/14\/4\/248"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,20]]},"references-count":47,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2023,4]]}},"alternative-id":["info14040248"],"URL":"https:\/\/doi.org\/10.3390\/info14040248","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,4,20]]}}}