{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,24]],"date-time":"2026-04-24T01:23:16Z","timestamp":1776993796447,"version":"3.51.4"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2013,1,1]],"date-time":"2013-01-01T00:00:00Z","timestamp":1356998400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000028","name":"Semiconductor Research Corporation","doi-asserted-by":"publisher","award":["1985.001"],"award-info":[{"award-number":["1985.001"]}],"id":[{"id":"10.13039\/100000028","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["903191"],"award-info":[{"award-number":["903191"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Des. Autom. Electron. Syst."],"published-print":{"date-parts":[[2013,1]]},"abstract":"<jats:p>Asymmetric multi-core processors (AMPs) have been shown to outperform symmetric ones in terms of performance and performance\/watt. Improved performance and power efficiency are achieved when the program threads are matched to their most suitable cores. Since the computational needs of a program may change during its execution, the best thread to core assignment will likely change with time. We have, therefore, developed an online program phase classification scheme that allows the swapping of threads when the current needs of the threads justify a change in the assignment.<\/jats:p>\n          <jats:p>The architectural differences among the cores in an AMP can never match the diversity that exists among different programs and even between different phases of the same program. Consider, for example, a program (or a program phase) that has a high instruction-level parallelism (ILP) and will exhibit high power efficiency if executed on a powerful core. We can not, however, include such powerful cores in the designed AMP, since they will remain underutilized most of the time, and they are not power efficient when the programs do not exhibit a high degree of ILP. Thus, we must expect to see program phases where the designed cores will be unable to support the ILP that the program can exhibit. We, therefore, propose in this article a dynamic morphing scheme. This scheme will allow a core to gain control of a functional unit that is ordinarily under the control of a neighboring core during periods of intense computation with high ILP. This way, we dynamically adjust the hardware resources to the current needs of the application.<\/jats:p>\n          <jats:p>Our results show that combining online phase classification and dynamic core morphing can significantly improve the performance\/watt of most multithreaded workloads.<\/jats:p>","DOI":"10.1145\/2390191.2390196","type":"journal-article","created":{"date-parts":[[2013,1,15]],"date-time":"2013-01-15T15:32:11Z","timestamp":1358263931000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":13,"title":["Improving performance per watt of asymmetric multi-core processors via online program phase classification and adaptive core morphing"],"prefix":"10.1145","volume":"18","author":[{"given":"Rance","family":"Rodrigues","sequence":"first","affiliation":[{"name":"University of Massachusetts at Amherst, Amherst, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arunachalam","family":"Annamalai","sequence":"additional","affiliation":[{"name":"University of Massachusetts at Amherst, Amherst, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Israel","family":"Koren","sequence":"additional","affiliation":[{"name":"University of Massachusetts at Amherst, Amherst, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sandip","family":"Kundu","sequence":"additional","affiliation":[{"name":"University of Massachusetts at Amherst, Amherst, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2013,1,16]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.36"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.51"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/1128022.1128029"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/339647.339657"},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1629911.1630149"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1077603.1077657"},{"key":"e_1_2_2_7_1","volume-title":"Proceedings of the IEEE International Conference on Computer Design (ICCD).","author":"Das A.","unstructured":"Das , A. , Rodrigues , R. , Koren , I. , and Kundu , S . 2010. A study on performance benefits of core morphing in an asymmetric multicore processor . In Proceedings of the IEEE International Conference on Computer Design (ICCD). Das, A., Rodrigues, R., Koren, I., and Kundu, S. 2010. A study on performance benefits of core morphing in an asymmetric multicore processor. In Proceedings of the IEEE International Conference on Computer Design (ICCD)."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815966"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2010.30"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/1128020.1128563"},{"key":"e_1_2_2_11_1","unstructured":"Held J. Bautista J. and Koehl S. 2006. White paper from a few cores to many: A tera-scale computing research review.  Held J. Bautista J. and Koehl S. 2006. White paper from a few cores to many: A tera-scale computing research review."},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1147\/rd.483.0425"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2008.209"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1273440.1250686"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/383082.383119"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1785481.1785573"},{"key":"e_1_2_2_17_1","volume-title":"Proceedings of the 6th International Conference on High-Performance and Embedded Architectures and Compliers (HiPEAC), 84--110","author":"Khan O.","unstructured":"Khan , O. and Kundu , S . 2011. Microvisor: A runtime architecture for thermal management in chip multiprocessors . In Proceedings of the 6th International Conference on High-Performance and Embedded Architectures and Compliers (HiPEAC), 84--110 . Khan, O. and Kundu, S. 2011. Microvisor: A runtime architecture for thermal management in chip multiprocessors. In Proceedings of the 6th International Conference on High-Performance and Embedded Architectures and Compliers (HiPEAC), 84--110."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.5555\/1331699.1331733"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1755913.1755928"},{"key":"e_1_2_2_20_1","volume-title":"Proceedings of the 36th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO-36)","author":"Kumar R.","unstructured":"Kumar , R. , Farkas , K. , Jouppi , N. , Ranganathan , P. , and Tullsen , D . 2003. Single-isa heterogeneous multi-core architectures: The potential for processor power reduction . In Proceedings of the 36th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO-36) . Kumar, R., Farkas, K., Jouppi, N., Ranganathan, P., and Tullsen, D. 2003. Single-isa heterogeneous multi-core architectures: The potential for processor power reduction. In Proceedings of the 36th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO-36)."},{"key":"e_1_2_2_21_1","volume-title":"Proceedings of the 31st Annual International Symposium on Computer Architecture.","author":"Kumar R.","unstructured":"Kumar , R. , Tullsen , D. , Ranganathan , P. , Jouppi , N. , and Farkas , K . 2004. Single-isa heterogeneous multi-core architectures for multithreaded workload performance . In Proceedings of the 31st Annual International Symposium on Computer Architecture. Kumar, R., Tullsen, D., Ranganathan, P., Jouppi, N., and Farkas, K. 2004. Single-isa heterogeneous multi-core architectures for multithreaded workload performance. In Proceedings of the 31st Annual International Symposium on Computer Architecture."},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1152154.1152162"},{"key":"e_1_2_2_23_1","volume-title":"Proceedings of the 30th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO-30)","author":"Lee C.","unstructured":"Lee , C. , Potkonjak , M. , and Mangione-Smith , W. H . 1997. Mediabench: A tool for evaluating and synthesizing multimedia and communicatons systems . In Proceedings of the 30th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO-30) . Lee, C., Potkonjak, M., and Mangione-Smith, W. H. 1997. Mediabench: A tool for evaluating and synthesizing multimedia and communicatons systems. In Proceedings of the 30th Annual ACM\/IEEE International Symposium on Microarchitecture (MICRO-30)."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/1362622.1362694"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854329"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2009.44"},{"key":"e_1_2_2_27_1","volume-title":"Proceedings of the IEEE 15th International Symposium on High Performance Computer Architecture (HPCA'09)","author":"Najaf","unstructured":"Najaf -abadi, H. and Rotenberg , E . 2009. Architectural contesting . In Proceedings of the IEEE 15th International Symposium on High Performance Computer Architecture (HPCA'09) . 189--200. Najaf-abadi, H. and Rotenberg, E. 2009. Architectural contesting. In Proceedings of the IEEE 15th International Symposium on High Performance Computer Architecture (HPCA'09). 189--200."},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/1299042.1299107"},{"key":"e_1_2_2_29_1","volume-title":"Sesc: Superescalar simulator","author":"Renau J.","year":"2005","unstructured":"Renau , J. 2005 . Sesc: Superescalar simulator . http:\/\/lacoma.cs.uiuc.edu\/~paulsack\/sescdoc\/. Renau, J. 2005. Sesc: Superescalar simulator. http:\/\/lacoma.cs.uiuc.edu\/~paulsack\/sescdoc\/."},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.18"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1755913.1755929"},{"key":"e_1_2_2_32_1","volume-title":"Proceedings of the IEEE 14th International Symposium on High Performance Computer Architecture (HPCA'08)","author":"Salverda P.","unstructured":"Salverda , P. and Zilles , C . 2008. Fundamental performance constraints in horizontal fusion of in-order cores . In Proceedings of the IEEE 14th International Symposium on High Performance Computer Architecture (HPCA'08) . 252--263. Salverda, P. and Zilles, C. 2008. Fundamental performance constraints in horizontal fusion of in-order cores. In Proceedings of the IEEE 14th International Symposium on High Performance Computer Architecture (HPCA'08). 252--263."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1531793.1531804"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/859618.859657"},{"key":"e_1_2_2_35_1","unstructured":"Shivakumar P. Jouppi N. P. and Shivakumar P. 2001. Cacti 3.0: An integrated cache timing power and area model. Tech. rep.  Shivakumar P. Jouppi N. P. and Shivakumar P. 2001. Cacti 3.0: An integrated cache timing power and area model. Tech. rep."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/1577129.1577137"},{"key":"e_1_2_2_37_1","unstructured":"SPEC2000. The standard performance evaluation corporation (spec cpi2000 suite).  SPEC2000. The standard performance evaluation corporation (spec cpi2000 suite)."},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/1945023.1945032"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.5555\/1874620.1874924"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815965"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/1854273.1854283"}],"container-title":["ACM Transactions on Design Automation of Electronic Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2390191.2390196","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2390191.2390196","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T08:35:45Z","timestamp":1750235745000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2390191.2390196"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,1]]},"references-count":41,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2013,1]]}},"alternative-id":["10.1145\/2390191.2390196"],"URL":"https:\/\/doi.org\/10.1145\/2390191.2390196","relation":{},"ISSN":["1084-4309","1557-7309"],"issn-type":[{"value":"1084-4309","type":"print"},{"value":"1557-7309","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,1]]},"assertion":[{"value":"2012-03-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2012-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2013-01-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}