{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,29]],"date-time":"2025-09-29T08:04:53Z","timestamp":1759133093355,"version":"3.41.0"},"reference-count":40,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2013,12,1]],"date-time":"2013-12-01T00:00:00Z","timestamp":1385856000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004837","name":"Ministerio de Ciencia e Innovaci\u00f3n","doi-asserted-by":"publisher","award":["TIN2012-34557"],"award-info":[{"award-number":["TIN2012-34557"]}],"id":[{"id":"10.13039\/501100004837","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004963","name":"Seventh Framework Programme","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100004963","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100004543","name":"China Scholarship Council","doi-asserted-by":"publisher","award":["2010608015"],"award-info":[{"award-number":["2010608015"]}],"id":[{"id":"10.13039\/501100004543","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2013,12]]},"abstract":"<jats:p>Accurately determining the energy consumed by each task in a system will become of prominent importance in future multicore-based systems because it offers several benefits, including (i) better application energy\/performance optimizations, (ii) improved energy-aware task scheduling, and (iii) energy-aware billing in data centers. Unfortunately, existing methods for energy metering in multicores fail to provide accurate energy estimates for each task when several tasks run simultaneously.<\/jats:p><jats:p>This article makes a case for accurate Per-Task Energy Metering (PTEM) based on tracking the resource utilization and occupancy of each task. Different hardware implementations with different trade-offs between energy prediction accuracy and hardware-implementation complexity are proposed. Our evaluation shows that the energy consumed in a multicore by each task can be accurately measured. For a 32-core, 2-way, simultaneous multithreaded core setup, PTEM reduces the average accuracy error from more than 12% when our hardware support is not used to less than 4% when it is used. The maximum observed error for any task in the workload we used reduces from 58% down to 9% when our hardware support is used.<\/jats:p>","DOI":"10.1145\/2541228.2555291","type":"journal-article","created":{"date-parts":[[2014,1,14]],"date-time":"2014-01-14T13:39:57Z","timestamp":1389706797000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Hardware support for accurate per-task energy metering in multicore systems"],"prefix":"10.1145","volume":"10","author":[{"given":"Qixiao","family":"Liu","sequence":"first","affiliation":[{"name":"Barcelona Supercomputing Center (BSC-CNS) and Universitat Polit\u00e8cnica de Catalunya (UPC), Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Miquel","family":"Moreto","sequence":"additional","affiliation":[{"name":"UPC and BSC-CNS, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Victor","family":"Jimenez","sequence":"additional","affiliation":[{"name":"BSC-CNS, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jaume","family":"Abella","sequence":"additional","affiliation":[{"name":"BSC-CNS, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francisco J.","family":"Cazorla","sequence":"additional","affiliation":[{"name":"Spanish National Research Council (IIIA-CSIC) and BSC-CNS, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mateo","family":"Valero","sequence":"additional","affiliation":[{"name":"UPC and BSC-CNS, Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2013,12]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/1061267.1061271"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/1095408.1095420"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2007.443"},{"volume-title":"Proceedings of the Intersociety Conference on Thermal and Thermomechanical Phenomena in Electronics Systems, 439--444","author":"Belady C.","key":"e_1_2_1_4_1","unstructured":"Belady , C. and Malone , C . 2006. Data center power projections to 2014 . Proceedings of the Intersociety Conference on Thermal and Thermomechanical Phenomena in Electronics Systems, 439--444 . Belady, C. and Malone, C. 2006. Data center power projections to 2014. Proceedings of the Intersociety Conference on Thermal and Thermomechanical Phenomena in Electronics Systems, 439--444."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/566726.566736"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2011.03.007"},{"volume-title":"IEEE\/SEMI Advanced Semiconductor Manufacturing Conference. 387--392","author":"Bickford J.","key":"e_1_2_1_7_1","unstructured":"Bickford , J. , Rosner , R. , Hedberg , E. , Yoder , J. , and Barnett , T . 2008. Sram redundancy - silicon area versus number of repairs trade-off . In IEEE\/SEMI Advanced Semiconductor Manufacturing Conference. 387--392 . Bickford, J., Rosner, R., Hedberg, E., Yoder, J., and Barnett, T. 2008. Sram redundancy - silicon area versus number of repairs trade-off. In IEEE\/SEMI Advanced Semiconductor Manufacturing Conference. 387--392."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2011.47"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/339647.339657"},{"volume-title":"USENIX Annual Technical Conference. 21--21","author":"Carroll A.","key":"e_1_2_1_10_1","unstructured":"Carroll , A. and Heiser , G . 2010. An analysis of power consumption in a smartphone . In USENIX Annual Technical Conference. 21--21 . Carroll, A. and Heiser, G. 2010. An analysis of power consumption in a smartphone. In USENIX Annual Technical Conference. 21--21."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPADS.2011.28"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2006.39"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1133572.1133597"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2011.29"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2011.48"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555756"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1840845.1840883"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.29"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2011.35"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2010.38"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807128.1807136"},{"key":"e_1_2_1_22_1","volume-title":"Growth in data center electricity use 2005 to","author":"Koomey J.","year":"2010","unstructured":"Koomey , J. 2011. Growth in data center electricity use 2005 to 2010 . Analytics Press . Koomey, J. 2011. Growth in data center electricity use 2005 to 2010. Analytics Press."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.34"},{"key":"e_1_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Liu Q. Jimenez V. Moreto M. Abella J. Cazorla F. and Valero M. 2013. Per-task energy accounting in computing systems. IEEE Computer Architecture Letters (to appear). http:\/\/people.ac.upc.edu\/jabella\/camerareadyIEEECAL.pdf. Liu Q. Jimenez V. Moreto M. Abella J. Cazorla F. and Valero M. 2013. Per-task energy accounting in computing systems. IEEE Computer Architecture Letters (to appear). http:\/\/people.ac.upc.edu\/jabella\/camerareadyIEEECAL.pdf.","DOI":"10.1109\/L-CA.2013.24"},{"volume-title":"USENIX Annual Technical Conference. 12--12","author":"McCullough J.","key":"e_1_2_1_25_1","unstructured":"McCullough , J. , Agarwal , Y. , Chandrashekar , J. , Kuppuswamy , S. , Snoeren , A. , and Gupta , R . 2011. Evaluating the effectiveness of model-based power characterization . In USENIX Annual Technical Conference. 12--12 . McCullough, J., Agarwal, Y., Chandrashekar, J., Kuppuswamy, S., Snoeren, A., and Gupta, R. 2011. Evaluating the effectiveness of model-based power characterization. In USENIX Annual Technical Conference. 12--12."},{"volume-title":"11th Workshop on the Use of High Performance Computing in Meteorology, Reading.","author":"Michalakes J.","key":"e_1_2_1_26_1","unstructured":"Michalakes , J. , Dudhia , J. , Gill , D. , Henderson , T. , Klemp , J. , Skamarock , W. , and Wang , W . The weather research and forecast model: software architecture and performance . In 11th Workshop on the Use of High Performance Computing in Meteorology, Reading. Michalakes, J., Dudhia, J., Gill, D., Henderson, T., Klemp, J., Skamarock, W., and Wang, W. The weather research and forecast model: software architecture and performance. In 11th Workshop on the Use of High Performance Computing in Meteorology, Reading."},{"key":"e_1_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Moreto M. Cazorla F. Ramirez A. and Valero M. 2008. Mlp-aware dynamic cache partitioning. In HiPEAC. Moreto M. Cazorla F. Ramirez A. and Valero M. 2008. Mlp-aware dynamic cache partitioning. In HiPEAC.","DOI":"10.1109\/PACT.2007.4336246"},{"key":"e_1_2_1_28_1","unstructured":"Muralimanohar N. and Balasubramonian R. 2009. Cacti 6.0: A tool to understand large caches. HP Tech Report HPL-2009-85. Muralimanohar N. and Balasubramonian R. 2009. Cacti 6.0: A tool to understand large caches. HP Tech Report HPL-2009-85."},{"key":"e_1_2_1_29_1","article-title":"The implementation of a 2-core multi-threaded itanium family processor","author":"Naffziger S.","year":"2005","unstructured":"Naffziger , S. , Stackhouse , B. , Grutkowski , T. , Josephson , D. , Desai , J. , Alon , E. , and Horowitz , M. 2005 . The implementation of a 2-core multi-threaded itanium family processor . IEEE Journal of Solid-State Circuits, 182--183. Naffziger, S., Stackhouse, B., Grutkowski, T., Josephson, D., Desai, J., Alon, E., and Horowitz, M. 2005. The implementation of a 2-core multi-threaded itanium family processor. IEEE Journal of Solid-State Circuits, 182--183.","journal-title":"IEEE Journal of Solid-State Circuits, 182--183."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2007.910967"},{"key":"e_1_2_1_31_1","unstructured":"Nokia. 2012. Energy profiler. Nokia. 2012. Energy profiler."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1966445.1966460"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1346281.1346289"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.sysarc.2006.11.006"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451116.2451124"},{"key":"e_1_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Sherwood T. Perelman E. and Calder B. 2001. Basic block distribution analysis to find periodic behavior and simulation points in applications. 3--14. Sherwood T. Perelman E. and Calder B. 2001. Basic block distribution analysis to find periodic behavior and simulation points in applications. 3--14.","DOI":"10.1109\/PACT.2001.953283"},{"key":"e_1_2_1_37_1","unstructured":"Singhal R. 2008. Inside intel next generation nehalem microarchitecture. In Intel Developer Forum. Singhal R. 2008. Inside intel next generation nehalem microarchitecture. In Intel Developer Forum."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/285930.286011"},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Udipi A. Muralimanohar N. and Balasubramonian R. 2010. Towards scalable energy-efficient bus-based on-chip networks. In HPCA. Udipi A. Muralimanohar N. and Balasubramonian R. 2010. Towards scalable energy-efficient bus-based on-chip networks. In HPCA.","DOI":"10.1109\/HPCA.2010.5416639"},{"key":"e_1_2_1_40_1","unstructured":"Weste N. and Eshraghian K. 1988. Principles of CMOS VLSI Design. A Systems Perspective. Addison-Wesley. Weste N. and Eshraghian K. 1988. Principles of CMOS VLSI Design. A Systems Perspective. Addison-Wesley."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2541228.2555291","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2541228.2555291","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T07:35:00Z","timestamp":1750232100000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2541228.2555291"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,12]]},"references-count":40,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2013,12]]}},"alternative-id":["10.1145\/2541228.2555291"],"URL":"https:\/\/doi.org\/10.1145\/2541228.2555291","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"type":"print","value":"1544-3566"},{"type":"electronic","value":"1544-3973"}],"subject":[],"published":{"date-parts":[[2013,12]]},"assertion":[{"value":"2013-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-10-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-12-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}