{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,3,26]],"date-time":"2025-03-26T11:09:40Z","timestamp":1742987380398,"version":"3.40.3"},"publisher-location":"Cham","reference-count":38,"publisher":"Springer Nature Switzerland","isbn-type":[{"type":"print","value":"9783031226762"},{"type":"electronic","value":"9783031226779"}],"license":[{"start":{"date-parts":[[2023,1,1]],"date-time":"2023-01-01T00:00:00Z","timestamp":1672531200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,1,11]],"date-time":"2023-01-11T00:00:00Z","timestamp":1673395200000},"content-version":"vor","delay-in-days":10,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2023]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Although GPUs have been used to accelerate various convolutional neural network algorithms with good performance, the demand for performance improvement is still continuously increasing. CPU\/GPU overclocking technology brings opportunities for further performance improvement in CPU-GPU heterogeneous platforms. However, CPU\/GPU overclocking inevitably increases the power of the CPU\/GPU, which is not conducive to energy conservation, energy efficiency optimization, or even system stability. How to effectively constrain the total energy to remain roughly unchanged during the CPU\/GPU overclocking is a key issue in designing adaptive overclocking algorithms. There are two key factors during solving this key issue. Firstly, the dynamic power upper bound must be set to reflect the real-time behavior characteristics of the program so that algorithm can better meet the total energy unchanging constraints; secondly, instead of independently overclocking at both CPU and GPU sides, coordinately overclocking on CPU-GPU must be considered to adapt to real-time load balance for higher performance improvement and better energy constraints. This paper proposes an <jats:italic>Adaptive Overclocking Algorithm<\/jats:italic> (AOA) on CPU-GPU heterogeneous platforms to achieve the goal of performance improvement while the total energy remains roughly unchanged. AOA uses the function <jats:inline-formula><jats:alternatives><jats:tex-math>$$F_k$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                <mml:msub>\n                  <mml:mi>F<\/mml:mi>\n                  <mml:mi>k<\/mml:mi>\n                <\/mml:msub>\n              <\/mml:math><\/jats:alternatives><\/jats:inline-formula> to describe the variable power upper bound and introduces the load imbalance factor <jats:italic>W<\/jats:italic> to realize the CPU-GPU coordinated overclocking. Through the verification of several types convolutional neural network algorithms on two CPU-GPU heterogeneous platforms (Intel<jats:inline-formula><jats:alternatives><jats:tex-math>$$^\\circledR $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                <mml:msup>\n                  <mml:mrow\/>\n                  <mml:mo>\u00ae<\/mml:mo>\n                <\/mml:msup>\n              <\/mml:math><\/jats:alternatives><\/jats:inline-formula>Xeon E5-2660 &amp; NVIDIA<jats:inline-formula><jats:alternatives><jats:tex-math>$$^\\circledR $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                <mml:msup>\n                  <mml:mrow\/>\n                  <mml:mo>\u00ae<\/mml:mo>\n                <\/mml:msup>\n              <\/mml:math><\/jats:alternatives><\/jats:inline-formula>Tesla K80; Intel<jats:inline-formula><jats:alternatives><jats:tex-math>$$^\\circledR $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                <mml:msup>\n                  <mml:mrow\/>\n                  <mml:mo>\u00ae<\/mml:mo>\n                <\/mml:msup>\n              <\/mml:math><\/jats:alternatives><\/jats:inline-formula>Core\u2122i9-10920X &amp; NIVIDIA<jats:inline-formula><jats:alternatives><jats:tex-math>$$^\\circledR $$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                <mml:msup>\n                  <mml:mrow\/>\n                  <mml:mo>\u00ae<\/mml:mo>\n                <\/mml:msup>\n              <\/mml:math><\/jats:alternatives><\/jats:inline-formula>GeForce RTX 2080Ti), AOA achieves an average of 10.7% performance improvement and 4.4% energy savings. To verify the effectiveness of the AOA, we compare AOA with other methods including automatic boost, the highest overclocking and static optimal overclocking.<\/jats:p>","DOI":"10.1007\/978-3-031-22677-9_14","type":"book-chapter","created":{"date-parts":[[2023,1,10]],"date-time":"2023-01-10T09:04:32Z","timestamp":1673341472000},"page":"253-272","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["AOA: Adaptive Overclocking Algorithm on\u00a0CPU-GPU Heterogeneous Platforms"],"prefix":"10.1007","author":[{"given":"Zhixin","family":"Ou","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Juan","family":"Chen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuyang","family":"Sun","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tao","family":"Xu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guodong","family":"Jiang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhengyuan","family":"Tan","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinxin","family":"Qi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,1,11]]},"reference":[{"key":"14_CR1","unstructured":"OL. http:\/\/dag.wiee.rs\/home-made\/dstat\/. Accessed Dec 2021"},{"key":"14_CR2","unstructured":"Linux kernel profiling with perf. OL. https:\/\/perf.wiki.kernel.org\/index.php\/Tutorial. Accessed Dec 2021"},{"key":"14_CR3","unstructured":"Nvidia system management interface. OL. https:\/\/developer.nvidia.com\/nvidia-system-management-interface. Accessed Dec 2021"},{"key":"14_CR4","unstructured":"Abadi, M., et al.: Tensorflow: large-scale machine learning on heterogeneous distributed systems. arXiv abs\/1603.04467 (2016)"},{"key":"14_CR5","doi-asserted-by":"publisher","unstructured":"Acun, B., Miller, P., Kale, L.V.: Variation among processors under turbo boost in HPC systems. In: Proceedings of the 2016 International Conference on Supercomputing. ICS 2016. Association for Computing Machinery, New York (2016). https:\/\/doi.org\/10.1145\/2925426.2926289, https:\/\/doi-org-s.nudtproxy.yitlink.com\/10.1145\/2925426.2926289","DOI":"10.1145\/2925426.2926289"},{"key":"14_CR6","doi-asserted-by":"publisher","unstructured":"Chasapis, D., Moret\u00f3, M., Schulz, M., Rountree, B., Valero, M., Casas, M.: Power efficient job scheduling by predicting the impact of processor manufacturing variability. In: Proceedings of the ACM International Conference on Supercomputing. ICS 2019, pp. 296\u2013307. Association for Computing Machinery, New York (2019). https:\/\/doi.org\/10.1145\/3330345.3330372, https:\/\/doi-org-s.nudtproxy.yitlink.com\/10.1145\/3330345.3330372","DOI":"10.1145\/3330345.3330372"},{"issue":"6","key":"14_CR7","doi-asserted-by":"publisher","first-page":"1228","DOI":"10.1007\/s11704-018-7239-1","volume":"13","author":"J Chen","year":"2019","unstructured":"Chen, J., et al.: Analyzing time-dimension communication characterizations for representative scientific applications on supercomputer systems. Front. Comp. Sci. 13(6), 1228\u20131242 (2019)","journal-title":"Front. Comp. Sci."},{"key":"14_CR8","unstructured":"Chetlur, S., et al.: CUDNN: efficient primitives for deep learning. arXiv abs\/1410.0759 (2014)"},{"key":"14_CR9","doi-asserted-by":"publisher","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: ImageNet: a large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248\u2013255 (2009). https:\/\/doi.org\/10.1109\/CVPR.2009.5206848","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"14_CR10","unstructured":"Gad, E.A.: A work-stealing for dynamic workload balancing an CPU-GPU heterogeneous computing platforms. Thesis (2017). http:\/\/www.pqdtcn.com.nudtproxy.yitlink.com:80\/thesisDetails\/46952B07E4A7CC0D8C9AB6B408B99235"},{"key":"14_CR11","doi-asserted-by":"publisher","unstructured":"Gholkar, N., Mueller, F., Rountree, B.: Power tuning HPC jobs on power-constrained systems. In: 2016 International Conference on Parallel Architecture and Compilation Techniques (PACT), pp. 179\u2013190 (2016). https:\/\/doi.org\/10.1145\/2967938.2967961","DOI":"10.1145\/2967938.2967961"},{"key":"14_CR12","unstructured":"guassic: Text classification with CNN and RNN. https:\/\/github.com\/gaussic\/text-classification-cnn-rnn"},{"key":"14_CR13","doi-asserted-by":"publisher","unstructured":"He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 770\u2013778 (2016). https:\/\/doi.org\/10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"14_CR14","doi-asserted-by":"publisher","unstructured":"Inadomi, Y., et al.: Analyzing and mitigating the impact of manufacturing variability in power-constrained supercomputing. In: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. SC 2015. Association for Computing Machinery, New York (2015). https:\/\/doi.org\/10.1145\/2807591.2807638, https:\/\/doi-org-s.nudtproxy.yitlink.com\/10.1145\/2807591.2807638","DOI":"10.1145\/2807591.2807638"},{"key":"14_CR15","unstructured":"Intel\u00ae: Overclocking: Maximizing your performance. OL. https:\/\/www.intel.com\/content\/www\/us\/en\/gaming\/overclocking-intel-processors.html. Accessed Dec 2021"},{"key":"14_CR16","unstructured":"Intel\u00ae: Release notes (xtu-7.5.3.3-releasenotes.pdf). OLhttps:\/\/downloadmirror.intel.com\/29183\/XTU-7.5.3.3-ReleaseNotes.pdf. Accessed Dec 2021"},{"key":"14_CR17","doi-asserted-by":"crossref","unstructured":"Jia, Y., et al.: Caffe: convolutional architecture for fast feature embedding. arXiv abs\/1408.5093 (2014)","DOI":"10.1145\/2647868.2654889"},{"key":"14_CR18","doi-asserted-by":"publisher","unstructured":"Kodama, Y., Odajima, T., Arima, E., Sato, M.: Evaluation of power management control on the supercomputer Fugaku. In: 2020 IEEE International Conference on Cluster Computing (CLUSTER), pp. 484\u2013493 (2020). https:\/\/doi.org\/10.1109\/CLUSTER49012.2020.00069","DOI":"10.1109\/CLUSTER49012.2020.00069"},{"key":"14_CR19","doi-asserted-by":"publisher","unstructured":"Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. Commun. ACM 60(6), 84\u201390 (2017). https:\/\/doi.org\/10.1145\/3065386","DOI":"10.1145\/3065386"},{"key":"14_CR20","doi-asserted-by":"publisher","unstructured":"LeCun, Y., Kavukcuoglu, K., Farabet, C.: Convolutional networks and applications in vision. In: Proceedings of 2010 IEEE International Symposium on Circuits and Systems, pp. 253\u2013256 (2010). https:\/\/doi.org\/10.1109\/ISCAS.2010.5537907","DOI":"10.1109\/ISCAS.2010.5537907"},{"key":"14_CR21","doi-asserted-by":"publisher","unstructured":"Mittal, S., Vetter, J.S.: A survey of methods for analyzing and improving GPU energy efficiency. ACM Comput. Surv. 47(2) (2014). https:\/\/doi.org\/10.1145\/2636342, https:\/\/doi.org\/10.1145\/2636342","DOI":"10.1145\/2636342"},{"key":"14_CR22","unstructured":"NVIDIA\u00ae: GPU boost. OL. https:\/\/www.nvidia.com\/en-gb\/geforce\/technologies\/gpu-boost\/. Accessed Dec 2021"},{"key":"14_CR23","unstructured":"PyTorch: Imagenet training in PyTorch. OL. https:\/\/github.com\/pytorch\/examples\/tree\/master\/imagenet. Accessed Dec 2021"},{"key":"14_CR24","doi-asserted-by":"crossref","unstructured":"Ravichandran, D.S.M.R.M.E.C.S.: Processor Performance Enhancement Using Self-adaptive Clock Frequency, vol. 3, July 2010","DOI":"10.5120\/780-1104"},{"key":"14_CR25","doi-asserted-by":"crossref","unstructured":"Rodrigues, C.F., Riley, G., Luj\u00e1n, M.: Fine-grained energy profiling for deep convolutional neural networks on the Jetson tx1. In: 2017 IEEE International Symposium on Workload Characterization (IISWC), pp. 114\u2013115 (2017)","DOI":"10.1109\/IISWC.2017.8167764"},{"key":"14_CR26","doi-asserted-by":"crossref","unstructured":"Rouhani, B.D., Mirhoseini, A., Koushanfar, F.: Delight: adding energy dimension to deep neural networks. In: International Symposium on Low Power Electronics and Design (2016)","DOI":"10.1145\/2934583.2934599"},{"issue":"3","key":"14_CR27","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky, O., et al.: ImageNet large scale visual recognition challenge. Int. J. Comput. Vision 115(3), 211\u2013252 (2015). https:\/\/doi.org\/10.1007\/s11263-015-0816-y","journal-title":"Int. J. Comput. Vision"},{"key":"14_CR28","unstructured":"Simonyan, K., Zisserman, A.: Very deep convolutional networks for large-scale image recognition. Comput. Sci. (2014)"},{"key":"14_CR29","series-title":"Communications in Computer and Information Science","doi-asserted-by":"publisher","first-page":"196","DOI":"10.1007\/978-981-16-7443-3_12","volume-title":"Theoretical Computer Science","author":"Y Sun","year":"2021","unstructured":"Sun, Y., et al.: Evaluating performance, power and\u00a0energy of deep neural networks on\u00a0CPUs and GPUs. In: Cai, Z., Li, J., Zhang, J. (eds.) NCTCS 2021. CCIS, vol. 1494, pp. 196\u2013221. Springer, Singapore (2021). https:\/\/doi.org\/10.1007\/978-981-16-7443-3_12"},{"key":"14_CR30","doi-asserted-by":"publisher","unstructured":"Tang, Z., Wang, Y., Wang, Q., Chu, X.: The impact of GPU DVFs on the energy and performance of deep learning: an empirical study. In: Proceedings of the Tenth ACM International Conference on Future Energy Systems. e-Energy 2019, pp. 315\u2013325. Association for Computing Machinery, New York (2019). https:\/\/doi.org\/10.1145\/3307772.3328315","DOI":"10.1145\/3307772.3328315"},{"key":"14_CR31","doi-asserted-by":"publisher","unstructured":"Thomas, D., Shanmugasundaram, M.: A survey on different overclocking methods. In: 2018 Second International Conference on Electronics, Communication and Aerospace Technology (ICECA), pp. 1588\u20131592 (2018). https:\/\/doi.org\/10.1109\/ICECA.2018.8474921","DOI":"10.1109\/ICECA.2018.8474921"},{"key":"14_CR32","unstructured":"Wang, Y., et al.: E2-train: training state-of-the-art CNNs with over 80% energy savings. In: NeurIPS (2019)"},{"key":"14_CR33","doi-asserted-by":"crossref","unstructured":"Wu, F., Chen, J., Dong, Y., Zheng, W., Pan, X., Sun, Y.: Improve energy efficiency by processor overclocking and memory frequency scaling. In: 2018 IEEE 20th International Conference on High Performance Computing and Communications; IEEE 16th International Conference on Smart City; IEEE 4th International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS), pp. 960\u2013967 (2018)","DOI":"10.1109\/HPCC\/SmartCity\/DSS.2018.00159"},{"key":"14_CR34","doi-asserted-by":"publisher","unstructured":"Wu, F., et al.: A holistic energy-efficient approach for a processor-memory system. Tsinghua Sci. Technol. 24(4), 468\u2013483 (2019). https:\/\/doi.org\/10.26599\/TST.2018.9020104","DOI":"10.26599\/TST.2018.9020104"},{"key":"14_CR35","doi-asserted-by":"publisher","unstructured":"Yang, C., et al.: Adaptive optimization for petascale heterogeneous CPU\/GPU computing. In: 2010 IEEE International Conference on Cluster Computing, pp. 19\u201328 (2010). https:\/\/doi.org\/10.1109\/CLUSTER.2010.12","DOI":"10.1109\/CLUSTER.2010.12"},{"key":"14_CR36","unstructured":"Yang, F., Xu, Y., Meng, X., Gao, W., Mai, Q., Yang, C.: Nvidia tx2-based CPU, GPU coordinated frequency modulation energy-saving optimization method. Patent (2019). Patent Application Number: 201910360182.6. Publication Patent Number: CN 110308784 A"},{"key":"14_CR37","doi-asserted-by":"crossref","unstructured":"Yao, C., et al.: Evaluating and analyzing the energy efficiency of CNN inference on high-performance GPU. Concurrency and Computation: Practice and Experience (2020)","DOI":"10.1002\/cpe.6064"},{"key":"14_CR38","doi-asserted-by":"publisher","unstructured":"Zamani, H., Tripathy, D., Bhuyan, L., Chen, Z.: SAOU: safe adaptive overclocking and undervolting for energy-efficient GPU computing. In: Proceedings of the ACM\/IEEE International Symposium on Low Power Electronics and Design. ISLPED 2020, pp. 205\u2013210. Association for Computing Machinery, New York (2020). https:\/\/doi.org\/10.1145\/3370748.3406553, https:\/\/doi.org\/10.1145\/3370748.3406553","DOI":"10.1145\/3370748.3406553"}],"container-title":["Lecture Notes in Computer Science","Algorithms and Architectures for Parallel Processing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-3-031-22677-9_14","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,10]],"date-time":"2023-01-10T09:15:31Z","timestamp":1673342131000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-3-031-22677-9_14"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023]]},"ISBN":["9783031226762","9783031226779"],"references-count":38,"URL":"https:\/\/doi.org\/10.1007\/978-3-031-22677-9_14","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"type":"print","value":"0302-9743"},{"type":"electronic","value":"1611-3349"}],"subject":[],"published":{"date-parts":[[2023]]},"assertion":[{"value":"11 January 2023","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"ICA3PP","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Conference on Algorithms and Architectures for Parallel Processing","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Copenhagen","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Denmark","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2022","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"10 October 2022","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"12 October 2022","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"ica3pp2022","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Single-blind","order":1,"name":"type","label":"Type","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"EasyChair","order":2,"name":"conference_management_system","label":"Conference Management System","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"91","order":3,"name":"number_of_submissions_sent_for_review","label":"Number of Submissions Sent for Review","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"33","order":4,"name":"number_of_full_papers_accepted","label":"Number of Full Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"10","order":5,"name":"number_of_short_papers_accepted","label":"Number of Short Papers Accepted","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"36% - The value is computed by the equation \"Number of Full Papers Accepted \/ Number of Submissions Sent for Review * 100\" and then rounded to a whole number.","order":6,"name":"acceptance_rate_of_full_papers","label":"Acceptance Rate of Full Papers","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"3","order":7,"name":"average_number_of_reviews_per_paper","label":"Average Number of Reviews per Paper","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"5","order":8,"name":"average_number_of_papers_per_reviewer","label":"Average Number of Papers per Reviewer","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}},{"value":"Yes","order":9,"name":"external_reviewers_involved","label":"External Reviewers Involved","group":{"name":"ConfEventPeerReviewInformation","label":"Peer Review Information (provided by the conference organizers)"}}]}}