{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:38:37Z","timestamp":1750307917056,"version":"3.41.0"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2007,3,1]],"date-time":"2007-03-01T00:00:00Z","timestamp":1172707200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGARCH Comput. Archit. News"],"published-print":{"date-parts":[[2007,3]]},"abstract":"<jats:p>The scaling of technology and the diminishing return of complicated uniprocessors have driven the industry towards multicore processors. While multithreaded applications can naturally leverage the enhanced throughput of multi-core processors, a large number of important applications are single-threaded, which cannot automatically harness the potential of multi-core processors. In this paper, we propose a compiler-driven heterogeneous multicore architecture, consisting of tightly-integrated VLIW (Very Long Instruction Word) and superscalar processors on a single chip, to automatically boost the performance of single-threaded applications without compromising the capability to support multithreaded programs. In the proposed multi-core architecture, while the high-performance VLIW core is used to run code segments with high instruction-level parallelism (ILP) extracted by the compiler; the superscalar core can be exploited to deal with the runtime events that are typically difficult for the VLIW core to handle, such as L2 cache misses. Our initial experimental results by running the preexecution thread on the superscalar core to mitigate the L2 cache misses of the main thread on the VLIW core indicate that the proposed VLIW\/superscalar multi-core processor can automatically improve the performance of single-threaded general-purpose applications by up to 40.8%.<\/jats:p>","DOI":"10.1145\/1241601.1241603","type":"journal-article","created":{"date-parts":[[2007,6,6]],"date-time":"2007-06-06T14:37:16Z","timestamp":1181140636000},"page":"141-148","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["Hybrid multi-core architecture for boosting single-threaded performance"],"prefix":"10.1145","volume":"35","author":[{"given":"Jun","family":"Yan","sequence":"first","affiliation":[{"name":"Southern Illinois University Carbondale, Carbondale, IL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wei","family":"Zhang","sequence":"additional","affiliation":[{"name":"Southern Illinois University Carbondale, Carbondale, IL"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2007,3]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/998680.1006707"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2005.13"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/237090.237140"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2000.899310"},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4757-5676-0","volume-title":"Loop parallelization","author":"Banerjee U.","year":"1994","unstructured":"U. Banerjee . Loop parallelization . Kluwer Academic Publishers , Boston . MA, 1994 . U. Banerjee. Loop parallelization. Kluwer Academic Publishers, Boston. MA, 1994."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/781498.781500"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/520793.825732"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/996841.996851"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/822079.822712"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.795219"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1105734.1105741"},{"key":"e_1_2_1_13_1","volume-title":"Advanced compiler design and implementation","author":"Muchnick S. S.","year":"1997","unstructured":"S. S. Muchnick . Advanced compiler design and implementation . Morgan Kaufmann Publishers , 1997 . S. S. Muchnick. Advanced compiler design and implementation. Morgan Kaufmann Publishers, 1997."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/155090.155119"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.5555\/255235.255275"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/225160.225189"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/512529.512544"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/379240.379250"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/379240.379248"},{"key":"e_1_2_1_20_1","volume-title":"Proc of the 6th Workshop on Multithreaded Execution. Architecture and Compilation","author":"Brown J.","year":"2002","unstructured":"J. Brown , H. Wang , G. Chrysos , P. Wang and J. Shen . Speculative precomputation on chip multiprocessors . In Proc of the 6th Workshop on Multithreaded Execution. Architecture and Compilation , Nov 2002 . J. Brown, H. Wang, G. Chrysos, P. Wang and J. Shen. Speculative precomputation on chip multiprocessors. In Proc of the 6th Workshop on Multithreaded Execution. Architecture and Compilation, Nov 2002."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2005.17"},{"key":"e_1_2_1_23_1","volume-title":"Microprocessor Report","author":"Gwennap L.","year":"1996","unstructured":"L. Gwennap . Digital 21264 sets new standard . Microprocessor Report , Oct. 28, 1996 . L. Gwennap. Digital 21264 sets new standard. Microprocessor Report, Oct. 28, 1996."},{"volume-title":"model TM5800","year":"2001","key":"e_1_2_1_24_1","unstructured":"Crusoe processor , model TM5800 . Transmeta Corp. White Paper , 2001 . Crusoe processor, model TM5800. Transmeta Corp. White Paper, 2001."},{"key":"e_1_2_1_25_1","volume-title":"Embedded computing: a VLIW approach to architecture, compilers, and tools","author":"Fisher J. A.","year":"2005","unstructured":"J. A. Fisher , P. Faraboschi and C. Young . Embedded computing: a VLIW approach to architecture, compilers, and tools . Morgan Kaufmann Publishers , 2005 . J. A. Fisher, P. Faraboschi and C. Young. Embedded computing: a VLIW approach to architecture, compilers, and tools. Morgan Kaufmann Publishers, 2005."},{"key":"e_1_2_1_26_1","unstructured":"Trimaran homepage http:\/\/www.trimaran.org  Trimaran homepage http:\/\/www.trimaran.org"},{"key":"e_1_2_1_28_1","unstructured":"SPEC homepage http:\/\/www.spec.org.  SPEC homepage http:\/\/www.spec.org."},{"key":"e_1_2_1_29_1","volume-title":"IBM Research Report","author":"Altman E.","year":"1999","unstructured":"E. Altman , M. Gschwind , S. Sathaye , S. Kosonocky , A. Bright , J. Fritts , P. Ledak , D. Appenzeller , C. Agricola and Z. Filan . BOA: the architecture of a binary translation processor . IBM Research Report , 1999 . E. Altman, M. Gschwind, S. Sathaye, S. Kosonocky, A. Bright, J. Fritts, P. Ledak, D. Appenzeller, C. Agricola and Z. Filan. BOA: the architecture of a binary translation processor. IBM Research Report, 1999."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/358923.358939"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/378993.379237"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/1055626.1055644"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/115952.115979"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.5555\/646348.690381"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/223982.224451"},{"key":"e_1_2_1_36_1","first-page":"30","article-title":"Trace processors","author":"Rotenberg E.","year":"1997","unstructured":"E. Rotenberg , Q. Jacobson , Y. Sazeides , and J. Smith . Trace processors . In Proc. of Micro 30 , 1997 . E. Rotenberg, Q. Jacobson, Y. Sazeides, and J. Smith. Trace processors. In Proc. of Micro 30, 1997.","journal-title":"Proc. of Micro"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.5555\/563998.564005"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2002.997877"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/192724.192731"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/800046.801649"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF01205185"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.5555\/144953.144998"},{"key":"e_1_2_1_43_1","volume-title":"Aug","author":"Dual-Core Processor 0","year":"2004","unstructured":"OMAP591 0 Dual-Core Processor Datasheet. Texas Instruments White Paper , Aug 2004 . OMAP5910 Dual-Core Processor Datasheet. Texas Instruments White Paper, Aug 2004."},{"volume-title":"Philips White Paper","year":"2006","key":"e_1_2_1_44_1","unstructured":"Nexperia PNX4103 mobile multimedia processor . Philips White Paper , Jan 2006 . Nexperia PNX4103 mobile multimedia processor. Philips White Paper, Jan 2006."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.5555\/255235.255260"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/12.868027"},{"key":"e_1_2_1_47_1","volume-title":"Dec.","author":"Ju R.","year":"2001","unstructured":"R. Ju , S. Chan , C. Wu , R. Lian and T. Tuo . Open research compiler (ORC) for Itanium processor family. MICRO-34 Tutorial , Dec. 2001 . R. Ju, S. Chan, C. Wu, R. Lian and T. Tuo. Open research compiler (ORC) for Itanium processor family. MICRO-34 Tutorial, Dec. 2001."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2005.51"},{"issue":"4","key":"e_1_2_1_49_1","article-title":"Introduction to the Cell multiprocessor","volume":"49","author":"Kahle J. A.","year":"2005","unstructured":"J. A. Kahle , M. N. Day , H. P. Hofstee , C. R. Johns , T. R. Maeurer and D. Shippy . Introduction to the Cell multiprocessor . In IBM Journal of Research and Development , Vol. 49 , No. 4\/5 , July\/September , 2005 . J. A. Kahle, M. N. Day, H. P. Hofstee, C. R. Johns, T. R. Maeurer and D. Shippy. Introduction to the Cell multiprocessor. In IBM Journal of Research and Development, Vol. 49, No. 4\/5, July\/September, 2005.","journal-title":"IBM Journal of Research and Development"}],"container-title":["ACM SIGARCH Computer Architecture News"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1241601.1241603","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1241601.1241603","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T14:51:26Z","timestamp":1750258286000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1241601.1241603"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,3]]},"references-count":46,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2007,3]]}},"alternative-id":["10.1145\/1241601.1241603"],"URL":"https:\/\/doi.org\/10.1145\/1241601.1241603","relation":{},"ISSN":["0163-5964"],"issn-type":[{"type":"print","value":"0163-5964"}],"subject":[],"published":{"date-parts":[[2007,3]]},"assertion":[{"value":"2007-03-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}