{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:49:25Z","timestamp":1750308565719,"version":"3.41.0"},"reference-count":27,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2015,1,21]],"date-time":"2015-01-21T00:00:00Z","timestamp":1421798400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2015,1,21]]},"abstract":"<jats:p>Several studies and recent real-world designs have promoted sharing of underutilized resources between cores in a multicore processor to achieve better performance\/power. It has been argued that when utilization of such resources is low, sharing has a negligible impact on performance while offering considerable area and power benefits. In this article, we investigate the performance and performance\/watt implications of sharing large and underutilized resources between pairs of cores in a multicore. We first study sharing of the entire floating-point datapath (including reservation stations and execution units) by two cores, similar to AMD\u2019s Bulldozer. We find that while this architecture results in power savings for certain workload combinations, it also results in significant performance loss of up to 28%. Next, we study an alternative sharing architecture where only the floating-point execution units are shared, while the individual cores retain their reservation stations. This reduces the highest performance loss to 14%. We then extend the study to include sharing of other large execution units that are used infrequently, namely, the integer multiply and divide units. Subsequently, we analyze the impact of sharing hardware resources in Simultaneously Multithreaded (SMT) processors where multiple threads run concurrently on the same core. We observe that sharing improves performance\/watt at a negligible performance cost only if the shared units have high throughput. Sharing low-throughput units reduces both performance and performance\/watt. To increase the throughput of the shared units, we propose the use of Dynamic Voltage and Frequency Boosting (DVFB) of only the shared units that can be placed on a separate voltage island. Our results indicate that the use of DVFB improves both performance and performance\/watt by as much as 22% and 10%, respectively.<\/jats:p>","DOI":"10.1145\/2680543","type":"journal-article","created":{"date-parts":[[2015,1,28]],"date-time":"2015-01-28T14:05:51Z","timestamp":1422453951000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Does the Sharing of Execution Units Improve Performance\/Power of Multicores?"],"prefix":"10.1145","volume":"14","author":[{"given":"Rance","family":"Rodrigues","sequence":"first","affiliation":[{"name":"Department of Electrical and Computer Engineering, University of Massachusetts, Amherst MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Israel","family":"Koren","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, University of Massachusetts, Amherst MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sandip","family":"Kundu","sequence":"additional","affiliation":[{"name":"Department of Electrical and Computer Engineering, University of Massachusetts, Amherst MA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2015,1,21]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/SAMOS.2011.6045477"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/339647.339657"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2011.23"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1077603.1077657"},{"key":"e_1_2_1_5_1","volume-title":"CASH: Revisiting Hardware Sharing in Single-Chip Parallel Processor. Technical Report. J. Instruction-Level Parallelism.","author":"Dolbeau R.","year":"2002","unstructured":"R. Dolbeau and A. Seznec . 2002 . CASH: Revisiting Hardware Sharing in Single-Chip Parallel Processor. Technical Report. J. Instruction-Level Parallelism. R. Dolbeau and A. Seznec. 2002. CASH: Revisiting Hardware Sharing in Single-Chip Parallel Processor. Technical Report. J. Instruction-Level Parallelism."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1952998.1952999"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1629911.1630120"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2009.2022531"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2008.4771786"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2004.842831"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2012.6169037"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICGCS.2010.5543066"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the IEEE 14th International Symposium on High Performance Computer Architecture (HPCA\u201908)","author":"Kim W.","year":"2008","unstructured":"W. Kim , M. S. Gupta , G.-Y. Wei , and D. Brooks . 2008. System level analysis of fast, per-core DVFS using on-chip switching regulators . In Proceedings of the IEEE 14th International Symposium on High Performance Computer Architecture (HPCA\u201908) . 123--134. DOI: http:\/\/dx.doi.org\/10.1109\/HPCA. 2008 .4658633. 10.1109\/HPCA.2008.4658633 W. Kim, M. S. Gupta, G.-Y. Wei, and D. Brooks. 2008. System level analysis of fast, per-core DVFS using on-chip switching regulators. In Proceedings of the IEEE 14th International Symposium on High Performance Computer Architecture (HPCA\u201908). 123--134. DOI: http:\/\/dx.doi.org\/10.1109\/HPCA.2008.4658633."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2004.12"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/774572.774601"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/232973.232993"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the 2004 IEEE International on Solid-State Circuits Conference, 2004","volume":"1","author":"Lichtenau C.","year":"2004","unstructured":"C. Lichtenau , M. I. Ringler , T. Pfluger , S. Geissler , R. Hilgendorf , J. Heaslip , U. Weiss , P. Sandon , N. Rohrer , E. Cohen , and M. Canada . 2004. PowerTune: Advanced frequency and power scaling on 64b PowerPC microprocessor . In Proceedings of the 2004 IEEE International on Solid-State Circuits Conference, 2004 . Digest of Technical Papers. 356--357 , Vol. 1 . DOI: http:\/\/dx.doi.org\/10.1109\/ISSCC. 2004 .1332741. 10.1109\/ISSCC.2004.1332741 C. Lichtenau, M. I. Ringler, T. Pfluger, S. Geissler, R. Hilgendorf, J. Heaslip, U. Weiss, P. Sandon, N. Rohrer, E. Cohen, and M. Canada. 2004. PowerTune: Advanced frequency and power scaling on 64b PowerPC microprocessor. In Proceedings of the 2004 IEEE International on Solid-State Circuits Conference, 2004. Digest of Technical Papers. 356--357, Vol. 1. DOI: http:\/\/dx.doi.org\/10.1109\/ISSCC.2004.1332741."},{"key":"e_1_2_1_19_1","unstructured":"J. Renau. 2005. SESC: SuperESCalar Simulator. Retrieved from http:\/\/sourceforge.net\/projects\/sesc\/.  J. Renau. 2005. SESC: SuperESCalar Simulator. Retrieved from http:\/\/sourceforge.net\/projects\/sesc\/."},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.18"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2390191.2390196"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.5555\/874076.876477"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/946246.946573"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1577129.1577137"},{"volume-title":"The Standard Performance Evaluation Corporation (Spec CPI2000 suite).","year":"2000","key":"e_1_2_1_26_1","unstructured":"SPEC2000. 2000 . The Standard Performance Evaluation Corporation (Spec CPI2000 suite). Retrieved from https:\/\/www.spec.org\/cpu2000\/. SPEC2000. 2000. The Standard Performance Evaluation Corporation (Spec CPI2000 suite). Retrieved from https:\/\/www.spec.org\/cpu2000\/."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/225830.224449"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815965"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/225830.223990"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2680543","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2680543","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T19:04:16Z","timestamp":1750273456000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2680543"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,1,21]]},"references-count":27,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2015,1,21]]}},"alternative-id":["10.1145\/2680543"],"URL":"https:\/\/doi.org\/10.1145\/2680543","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"type":"print","value":"1539-9087"},{"type":"electronic","value":"1558-3465"}],"subject":[],"published":{"date-parts":[[2015,1,21]]},"assertion":[{"value":"2013-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-12-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2015-01-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}