{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T21:11:17Z","timestamp":1780089077148,"version":"3.54.0"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"2","content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Meas. Anal. Comput. Syst."],"published-print":{"date-parts":[[2026,5,29]]},"abstract":"<jats:p>Architectural simulation has become the critical bottleneck limiting design space exploration for high-performance computing systems. The growing complexity of modern GPUs and AI accelerators\u2014with hundreds to thousands of tightly-coupled components\u2014demands simulation frameworks that deliver efficient parallelism and scalable single-node execution. Existing frameworks such as SST, gem5, and GPGPU-Sim fail to meet these requirements: SST focuses on multi-node MPI scalability but struggles with intra-node scaling; GPGPU-Sim remains largely single-threaded. Critically, these frameworks provide fixed threading models with no mechanism for users to optimize simulation performance for their specific workload patterns.<\/jats:p>\n                  <jats:p>\n                    We introduce ACALSim, a scalable parallel simulation framework designed to accelerate design space exploration for high-performance systems. As a\n                    <jats:italic toggle=\"yes\">framework<\/jats:italic>\n                    contribution, ACALSim provides infrastructure and APIs for building high-performance simulators\u2014timing model accuracy is the responsibility of simulator developers, not the framework. ACALSim's key innovation is a\n                    <jats:bold>pluggable thread management architecture<\/jats:bold>\n                    that enables developers to implement custom scheduling strategies optimized for specific simulation patterns\u2014a capability absent in existing frameworks. This is complemented by (1) event-driven execution with fast-forward to eliminate idle cycle overhead, (2) a shared-memory data model enabling zero-copy communication, and (3) a two-phase parallel execution model for deterministic thread scaling. We demonstrate ACALSim's effectiveness through HPCSim, a GPU simulator targeting A100-class architectures. Direct comparison with an SST implementation\u2014using identical shared timing cores to isolate framework overhead\u2014shows ACALSim achieves\n                    <jats:bold>over 14\u00d7 speedup<\/jats:bold>\n                    with 41% lower memory footprint, while hardware validation confirms 0.72--1.22\u00d7 cycle count correlation with A100 measurements. While SST fails to complete 256+ thread block workloads within practical time limits, ACALSim simulates full LLaMA transformer layers (single block) in 17.7 minutes for LLaMA-7B and 30.4 minutes for LLaMA-13B\u2014enabling practical design space exploration that SST cannot achieve.\n                  <\/jats:p>","DOI":"10.1145\/3805626","type":"journal-article","created":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T20:34:18Z","timestamp":1780086858000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["ACALSim: A Scalable Parallel Simulation Framework for High-Performance System Design Space Exploration"],"prefix":"10.1145","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-1399-6037","authenticated-orcid":false,"given":"Wei-Fen","family":"Lin","sequence":"first","affiliation":[{"name":"Mijotech Inc., Tainan, Taiwan and Taiwan High-Performance Computing Education Association, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1814-8797","authenticated-orcid":false,"given":"Jen-Chien","family":"Chang","sequence":"additional","affiliation":[{"name":"National Cheng Kung University, Tainan, Taiwan and Taiwan High-Performance Computing Education Association, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5932-454X","authenticated-orcid":false,"given":"Yen-Po","family":"Chen","sequence":"additional","affiliation":[{"name":"National Cheng Kung University, Tainan, Taiwan and Taiwan High-Performance Computing Education Association, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9399-8605","authenticated-orcid":false,"given":"Zi-Yi","family":"Tai","sequence":"additional","affiliation":[{"name":"Taiwan High-Performance Computing Education Association, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-2327-6711","authenticated-orcid":false,"given":"Yu-Chen","family":"Chang","sequence":"additional","affiliation":[{"name":"National Yang Ming Chiao Tung University, Hsin-Chu, Taiwan and Taiwan High-Performance Computing Education Association, Taipei, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-0295-4878","authenticated-orcid":false,"given":"Chia-Pao","family":"Chiang","sequence":"additional","affiliation":[{"name":"National Cheng Kung University, Tainan, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-1071-8524","authenticated-orcid":false,"given":"Yu-Yang","family":"Lee","sequence":"additional","affiliation":[{"name":"National Cheng Kung University, Tainan, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-2649-6408","authenticated-orcid":false,"given":"Yu-Jie","family":"Wang","sequence":"additional","affiliation":[{"name":"National Cheng Kung University, Tainan, Taiwan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,5,29]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2009.4919648"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO61859.2024.00021"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICSAMOS.2010.5642102"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0743-7315(02)00004-7"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASAP.2015.7245728"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/SCW63240.2024.00129"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC63097.2024.00012"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2023.3256796"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3579371.3589348"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2024.3476390"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2024.3484648"},{"key":"e_1_2_1_12_1","unstructured":"A. Ishii et al. 2018. NVSwitch and DGX-2: NVLink-Switching chip and scale-up compute server. In Hot Chips 30."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00014"},{"key":"e_1_2_1_14_1","volume-title":"Proc. ACM\/IEEE 47th Annu. Int. Symp. Comput. Archit. 473-486","author":"Khairy M.","unstructured":"M. Khairy, Z. Shen, T. M. Aamodt, and T. G. Rogers. 2020. Accel-sim: An extensible simulation framework for validated GPU modeling. In Proc. ACM\/IEEE 47th Annu. Int. Symp. Comput. Archit. 473-486."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2020.2985963"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3669940.3707265"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS57955.2024.00064"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/2.982916"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS48437.2020.00029"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3059962"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/SAMOS.2017.8344612"},{"key":"e_1_2_1_22_1","unstructured":"NVIDIA. 2024. NVIDIA DGX H100\/H200 User Guide. https:\/\/docs.nvidia.com\/dgx\/dgxh100-user-guide\/dgxh100-user-guide.pdf."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/500001.500018"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2019.00042"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1964218.1964225"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS48437.2020.00016"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485963"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00077"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2014.6844466"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS57527.2023.00035"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICICDT.2011.5783207"}],"container-title":["Proceedings of the ACM on Measurement and Analysis of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3805626","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,29]],"date-time":"2026-05-29T20:38:03Z","timestamp":1780087083000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3805626"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,29]]},"references-count":31,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,5,29]]}},"alternative-id":["10.1145\/3805626"],"URL":"https:\/\/doi.org\/10.1145\/3805626","relation":{},"ISSN":["2476-1249"],"issn-type":[{"value":"2476-1249","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,5,29]]},"assertion":[{"value":"2026-05-29","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}