{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T12:48:12Z","timestamp":1782996492150,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":39,"publisher":"ACM","license":[{"start":{"date-parts":[[2026,7,5]],"date-time":"2026-07-05T00:00:00Z","timestamp":1783209600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"National Natural Science Foundation of China","award":["U23A6007"],"award-info":[{"award-number":["U23A6007"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2026,7,6]]},"DOI":"10.1145\/3797905.3800517","type":"proceedings-article","created":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T11:50:37Z","timestamp":1782993037000},"page":"215-226","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["GRASP: Fine-grained and Adaptive Sampled Simulation for GPU Performance Modeling"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1802-5188","authenticated-orcid":false,"given":"Ruini","family":"Xue","sequence":"first","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-5588-1596","authenticated-orcid":false,"given":"Lingwei","family":"Chao","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-7613-6156","authenticated-orcid":false,"given":"Zhenxing","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, 0009-0003-4135-6329, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-5725-0513","authenticated-orcid":false,"given":"Peilin","family":"Cai","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, 0009-0003-4135-6329, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-4667-8741","authenticated-orcid":false,"given":"Jiangying","family":"Xue","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, 0009-0003-4135-6329, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-4135-6329","authenticated-orcid":false,"given":"Tianyu","family":"Xiong","sequence":"additional","affiliation":[{"name":"University of Electronic Science and Technology of China, Chengdu, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,7,5]]},"reference":[{"key":"e_1_3_3_1_2_2","first-page":"265","volume-title":"Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI)","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek\u00a0Gordon Murray, Benoit Steiner, Paul\u00a0A. Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016. TensorFlow: A System for Large-Scale Machine Learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI). USENIX Association, Savannah, GA, USA, 265\u2013283."},{"key":"e_1_3_3_1_3_2","doi-asserted-by":"publisher","unstructured":"M. Abraham T. Murtola R. Schulz S. P\u00e1ll J. Smith B. Hess and E. Lindahl. 2015. GROMACS: High performance molecular simulations through multi-level parallelism from laptops to supercomputers. SoftwareX 1-2 (July 2015) 19\u201325. 10.1016\/j.softx.2015.06.001","DOI":"10.1016\/j.softx.2015.06.001"},{"key":"e_1_3_3_1_4_2","volume-title":"AMD APP SDK 3.0 Getting Started","year":"2017","unstructured":"AMD. 2017. AMD APP SDK 3.0 Getting Started. Advanced Micro Devices, Inc."},{"key":"e_1_3_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/2830772.2830780"},{"key":"e_1_3_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480100"},{"key":"e_1_3_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/SBAC-PAD.2014.30"},{"key":"e_1_3_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_3_1_9_2","first-page":"578","volume-title":"Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI)","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie\u00a0Q. Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. TVM: An Automated End-to-End Optimizing Compiler for Deep Learning. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI). USENIX Association, Carlsbad, CA, USA, 578\u2013594."},{"key":"e_1_3_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/1735688.1735702"},{"key":"e_1_3_3_1_11_2","doi-asserted-by":"publisher","unstructured":"Sambit Das Phani Motamarri Vishal Subramanian David\u00a0M. Rogers and Vikram Gavini. 2022. DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization. Computer Physics Communications 280 (2022) 108473. 10.1016\/j.cpc.2022.108473","DOI":"10.1016\/j.cpc.2022.108473"},{"key":"e_1_3_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/3038228.3038239"},{"key":"e_1_3_3_1_13_2","doi-asserted-by":"publisher","unstructured":"Lieven Eeckhout Nilanjan Goswami Tao Li Hai Jin Cheng-Zhong Xu and Junmin Wu. 2015. GPGPU-MiniBench: accelerating GPGPU micro-architecture simulation. IEEE Trans. Comput. 64 (Nov. 2015) 1\u20131. 10.1109\/TC.2015.2395427","DOI":"10.1109\/TC.2015.2395427"},{"key":"e_1_3_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC43674.2020.9286184"},{"key":"e_1_3_3_1_15_2","unstructured":"Zhangyin Feng Daya Guo Duyu Tang Nan Duan Xiaocheng Feng Ming Gong Linjun Shou Bing Qin Ting Liu Daxin Jiang and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. arxiv:https:\/\/arXiv.org\/abs\/2002.08155\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2002.08155"},{"key":"e_1_3_3_1_16_2","unstructured":"Daya Guo Shuo Ren Shuai Lu Zhangyin Feng Duyu Tang Shujie Liu Long Zhou Nan Duan Alexey Svyatkovskiy Shengyu Fu Michele Tufano Shao\u00a0Kun Deng Colin Clement Dawn Drain Neel Sundaresan Jian Yin Daxin Jiang and Ming Zhou. 2021. GraphCodeBERT: Pre-training Code Representations with Data Flow. arxiv:https:\/\/arXiv.org\/abs\/2009.08366\u00a0[cs.SE] https:\/\/arxiv.org\/abs\/2009.08366"},{"key":"e_1_3_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555775"},{"key":"e_1_3_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.59"},{"key":"e_1_3_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2014.53"},{"key":"e_1_3_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2015.14"},{"key":"e_1_3_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00047"},{"key":"e_1_3_3_1_23_2","series-title":"Advances in Parallel Computing","first-page":"93","volume-title":"Parallel Computing: On the Road to Exascale","author":"Krasnopolsky Boris","year":"2016","unstructured":"Boris Krasnopolsky and Alexey Medvedev. 2016. Acceleration of Large Scale OpenFOAM Simulations on Distributed Systems with Multicore CPUs and GPUs. In Parallel Computing: On the Road to Exascale. Advances in Parallel Computing, Vol.\u00a027. IOS Press, Amsterdam, The Netherlands, 93\u2013102."},{"key":"e_1_3_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3470496.3527384"},{"key":"e_1_3_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3614277"},{"key":"e_1_3_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1145\/3613424.3623773"},{"key":"e_1_3_3_1_27_2","doi-asserted-by":"crossref","unstructured":"Wenjie Liu Wim Heirman Stijn Eyerman Shoaib Akram and Lieven Eeckhout. 2021. Scale-model simulation. IEEE Computer Architecture Letters 20 2 (2021) 175\u2013178.","DOI":"10.1109\/LCA.2021.3133112"},{"key":"e_1_3_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS55109.2022.00006"},{"key":"e_1_3_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS57527.2023.00030"},{"key":"e_1_3_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS48437.2020.00017"},{"key":"e_1_3_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA57654.2024.00088"},{"key":"e_1_3_3_1_32_2","doi-asserted-by":"crossref","unstructured":"Timothy Sherwood Erez Perelman Greg Hamerly and Brad Calder. 2002. Automatically characterizing large scale program behavior. ACM SIGPLAN Notices 37 10 (2002) 45\u201357.","DOI":"10.1145\/605432.605403"},{"key":"e_1_3_3_1_33_2","volume-title":"3rd International Conference on Learning Representations (ICLR)","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In 3rd International Conference on Learning Representations (ICLR). ICLR, San Diego, CA, USA, 14\u00a0pages. arXiv:https:\/\/arXiv.org\/abs\/1409.1556."},{"key":"e_1_3_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322230"},{"key":"e_1_3_3_1_35_2","first-page":"1","volume-title":"Proceedings of the IEEE International Symposium on Workload Characterization (IISWC)","author":"Sun Yifan","year":"2016","unstructured":"Yifan Sun, Xiang Gong, Amir\u00a0Kavyan Ziabari, Leiming Yu, Xiangyu Li, Saoni Mukherjee, Carter McCardwell, Alejandro Villegas, and David Kaeli. 2016. HeteroMark, a benchmark suite for CPU-GPU collaborative computing. In Proceedings of the IEEE International Symposium on Workload Characterization (IISWC). IEEE, Los Alamitos, CA, USA, 1\u201310."},{"key":"e_1_3_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322230"},{"key":"e_1_3_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370865"},{"key":"e_1_3_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00077"},{"key":"e_1_3_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO50266.2020.00085"},{"key":"e_1_3_3_1_40_2","doi-asserted-by":"crossref","unstructured":"Yue Wang Weishi Wang Shafiq Joty and Steven C.\u00a0H. Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation. arxiv:https:\/\/arXiv.org\/abs\/2109.00859\u00a0[cs.CL] https:\/\/arxiv.org\/abs\/2109.00859","DOI":"10.18653\/v1\/2021.emnlp-main.685"}],"event":{"name":"ICS '26: 2026 International Conference on Supercomputing","location":"Belfast United Kingdom","acronym":"ICS '26","sponsor":["SIGHPC ACM Special Interest Group on High Performance Computing, Special Interest Group on High Performance Computing","SIGARCH ACM Special Interest Group on Computer Architecture"]},"container-title":["Proceedings of the 40th ACM International Conference on Supercomputing"],"original-title":[],"deposited":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T12:30:47Z","timestamp":1782995447000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3797905.3800517"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,5]]},"references-count":39,"alternative-id":["10.1145\/3797905.3800517","10.1145\/3797905"],"URL":"https:\/\/doi.org\/10.1145\/3797905.3800517","relation":{},"subject":[],"published":{"date-parts":[[2026,7,5]]},"assertion":[{"value":"2026-07-05","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}