{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,21]],"date-time":"2025-11-21T11:24:09Z","timestamp":1763724249154,"version":"3.45.0"},"reference-count":42,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2009,6,1]],"date-time":"2009-06-01T00:00:00Z","timestamp":1243814400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000143","name":"Division of Computing and Communication Foundations","doi-asserted-by":"publisher","award":["CCF-0811702CNS-0509356"],"award-info":[{"award-number":["CCF-0811702CNS-0509356"]}],"id":[{"id":"10.13039\/100000143","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000144","name":"Division of Computer and Network Systems","doi-asserted-by":"publisher","award":["CCF-0811702CNS-0509356"],"award-info":[{"award-number":["CCF-0811702CNS-0509356"]}],"id":[{"id":"10.13039\/100000144","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2009,6]]},"abstract":"<jats:p>\n                    Functional full-system simulators are powerful and versatile research tools for accelerating architectural exploration and advanced software development. Their main shortcoming is limited throughput when simulating large multiprocessor systems with hundreds or thousands of processors or when instrumentation is introduced. We propose the\n                    <jats:sc>ProtoFlex<\/jats:sc>\n                    simulation architecture, which uses FPGAs to accelerate full-system multiprocessor simulation and to facilitate high-performance instrumentation. Prior FPGA approaches that prototype a complete system in hardware are either too complex when scaling to large-scale configurations or require significant effort to provide full-system support. In contrast,\n                    <jats:sc>ProtoFlex<\/jats:sc>\n                    virtualizes the execution of many logical processors onto a consolidated number of multiple-context execution engines on the FPGA. Through virtualization, the number of engines can be judiciously scaled, as needed, to deliver on necessary simulation performance at a large savings in complexity. Further, to achieve low-complexity full-system support, a hybrid simulation technique called transplanting allows implementing in the FPGA only the frequently encountered behaviors, while a software simulator preserves the abstraction of a complete system.\n                  <\/jats:p>\n                  <jats:p>\n                    We have created a first instance of the\n                    <jats:sc>ProtoFlex<\/jats:sc>\n                    simulation architecture, which is an FPGA-based, full-system functional simulator for a 16-way UltraSPARC III symmetric multiprocessor server, hosted on a single Xilinx Virtex-II XCV2P70 FPGA. On average, the simulator achieves a 38x speedup (and as high as 49\u00d7) over comparable software simulation across a suite of applications, including OLTP on a commercial database server. We also demonstrate the advantages of minimal-overhead FPGA-accelerated instrumentation through a CMP cache simulation technique that runs orders-of-magnitude faster than software.\n                  <\/jats:p>","DOI":"10.1145\/1534916.1534925","type":"journal-article","created":{"date-parts":[[2009,6,16]],"date-time":"2009-06-16T08:58:25Z","timestamp":1245142705000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":76,"title":["ProtoFlex"],"prefix":"10.1145","volume":"2","author":[{"given":"Eric S.","family":"Chung","sequence":"first","affiliation":[{"name":"Computer Architecture Laboratory at Carnegie Mellon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael K.","family":"Papamichael","sequence":"additional","affiliation":[{"name":"Computer Architecture Laboratory at Carnegie Mellon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eriko","family":"Nurvitadhi","sequence":"additional","affiliation":[{"name":"Computer Architecture Laboratory at Carnegie Mellon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James C.","family":"Hoe","sequence":"additional","affiliation":[{"name":"Computer Architecture Laboratory at Carnegie Mellon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ken","family":"Mai","sequence":"additional","affiliation":[{"name":"Computer Architecture Laboratory at Carnegie Mellon"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Babak","family":"Falsafi","sequence":"additional","affiliation":[{"name":"Parallel Systems Architecture Laboratory \u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2009,6]]},"reference":[{"issue":"4","key":"e_1_2_1_1_1","first-page":"3","article-title":"Advanced Micro Devices","volume":"4","author":"AMD.","year":"2008","unstructured":"AMD. 2008. Advanced Micro Devices, SimNow Simulator 4.4.3. User\u2019s manual.","journal-title":"SimNow Simulator"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/342001.339696"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/1247360.1247401"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2006.82"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1054907.1054910"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/MDT.2005.30"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2008.20"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/643114.643116"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/1331699.1331723"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1344671.1344684"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1273440.1250722"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.982918"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/956417.956541"},{"volume-title":"Proceedings of the Conference on Field Programmable Logic and Applications.","author":"Krasnov A.","key":"e_1_2_1_14_1","unstructured":"Krasnov, A., Schultz, A., Wawrzynek, J., Gibeling, G., and Droz, P. 2007. RAMP Blue: A message-passing manycore system in FPGAs. In Proceedings of the Conference on Field Programmable Logic and Applications."},{"key":"e_1_2_1_15_1","volume-title":"Proceedings of the 4th Annual Workshop on Modeling, Benchmarking and Simulation.","author":"Lantz R.","year":"2008","unstructured":"Lantz, R. 2008. Fast functional simulation with parallel Embra. In Proceedings of the 4th Annual Workshop on Modeling, Benchmarking and Simulation."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/238793.238822"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1216919.1216927"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.982916"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1105734.1105747"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/4434.895100"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1250734.1250746"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","unstructured":"Nussbaum F. Fedorova A. and Small C. 2004. An overview of the Sam CMT simulator kit. Tech. rep. TR-2004-133 Sun Microsystems Research Labs.","DOI":"10.5555\/1698188"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/201310.201321"},{"key":"e_1_2_1_24_1","doi-asserted-by":"crossref","unstructured":"Over A. Clarke B. and Strazdins P. 2007. A comparison of two approaches to parallel simulation of multiprocessors. ispass 0 12--22.","DOI":"10.1109\/ISPASS.2007.363732"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2004.28"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2008.4510733"},{"volume-title":"Proceedings of the 12th International Symposium on High-Performance Computer Architecture, 29--40","author":"Penry D.","key":"e_1_2_1_27_1","unstructured":"Penry, D., Fay, D., Hodgdon, D., Wells, R., Schelle, G., August, D., and Connors, D. 2006. Exploiting parallelism and structure to accelerate the simulation of chip multi-processors. In Proceedings of the 12th International Symposium on High-Performance Computer Architecture, 29--40."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/166962.166979"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/88.473612"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","unstructured":"Smith B. 1985. In The Architecture of HEP on Parallel MIMD Computation: HEP Supercomputer and its Applications. Massachusetts Institute of Technology Cambridge MA 41--55.","DOI":"10.5555\/4832.4834"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/178243.178260"},{"volume-title":"Proceedings of the 3rd Workshop on Architectural Research Prototyping.","author":"Tan Z.","key":"e_1_2_1_32_1","unstructured":"Tan, Z., Asanovi\u0107, K., and Patterson, D. 2008. An FPGA host-multithreaded functional model for SPARC v8. In Proceedings of the 3rd Workshop on Architectural Research Prototyping."},{"key":"e_1_2_1_33_1","unstructured":"Thornton J. E. 1995. Parallel operation in the control data 6600. 5--12."},{"key":"e_1_2_1_34_1","unstructured":"Vahia D. and Hartke P. 2007. OpenSPARC T1 on Xilinx FPGAs--Updates. June 2007 RAMP Retreat."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2007.346205"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/1341312.1341325"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2007.39"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/1216919.1216936"},{"volume-title":"Proceedings of the Tutorial in the International Symposium on Microarchitecture (MICRO-38)","author":"Wenisch T.","key":"e_1_2_1_39_1","unstructured":"Wenisch, T. and Wunderlich, R. 2005. SimFlex: Fast, accurate and flexible simulation of computer systems. In Proceedings of the Tutorial in the International Symposium on Microarchitecture (MICRO-38)."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2006.79"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/233008.233025"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2007.363733"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1534916.1534925","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1534916.1534925","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1534916.1534925","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,11,18]],"date-time":"2025-11-18T09:27:00Z","timestamp":1763458020000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1534916.1534925"}},"subtitle":["Towards Scalable, Full-System Multiprocessor Simulations Using FPGAs"],"short-title":[],"issued":{"date-parts":[[2009,6]]},"references-count":42,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2009,6]]}},"alternative-id":["10.1145\/1534916.1534925"],"URL":"https:\/\/doi.org\/10.1145\/1534916.1534925","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"type":"print","value":"1936-7406"},{"type":"electronic","value":"1936-7414"}],"subject":[],"published":{"date-parts":[[2009,6]]},"assertion":[{"value":"2008-06-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2008-11-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2009-06-01","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}