{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T18:09:26Z","timestamp":1785953366864,"version":"3.56.0"},"publisher-location":"New York, NY, USA","reference-count":82,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,10,28]],"date-time":"2023-10-28T00:00:00Z","timestamp":1698451200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","award":["2154973, 1725657, 2011146, 1910413, 2312157"],"award-info":[{"award-number":["2154973, 1725657, 2011146, 1910413, 2312157"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,10,28]]},"DOI":"10.1145\/3613424.3614269","type":"proceedings-article","created":{"date-parts":[[2023,12,8]],"date-time":"2023-12-08T17:22:15Z","timestamp":1702056135000},"page":"1163-1177","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":16,"title":["IDYLL: Enhancing Page Translation in Multi-GPUs via Light Weight PTE Invalidations"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6281-6799","authenticated-orcid":false,"given":"Bingyao","family":"Li","sequence":"first","affiliation":[{"name":"University of Pittsburgh, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0034-0358","authenticated-orcid":false,"given":"Yanan","family":"Guo","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-4540-2146","authenticated-orcid":false,"given":"Yueqi","family":"Wang","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5709-2992","authenticated-orcid":false,"given":"Aamer","family":"Jaleel","sequence":"additional","affiliation":[{"name":"NVIDIA, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8372-6541","authenticated-orcid":false,"given":"Jun","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3385-2053","authenticated-orcid":false,"given":"Xulong","family":"Tang","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,12,8]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378468"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS49936.2021.00023"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458817.3480855"},{"key":"e_1_3_2_1_4_1","unstructured":"AMD. 2015. AMD APP SDK OpenCL Optimization Guide.  AMD. 2015. AMD APP SDK OpenCL Optimization Guide."},{"key":"e_1_3_2_1_5_1","volume-title":"Optimizing the TLB Shootdown Algorithm with Page Access Tracking. In 2017 USENIX Annual Technical Conference (USENIX ATC 17)","author":"Amit Nadav","year":"2017","unstructured":"Nadav Amit . 2017 . Optimizing the TLB Shootdown Algorithm with Page Access Tracking. In 2017 USENIX Annual Technical Conference (USENIX ATC 17) . USENIX Association, Santa Clara, CA, 27\u201339. https:\/\/www.usenix.org\/conference\/atc17\/technical-sessions\/presentation\/amit Nadav Amit. 2017. Optimizing the TLB Shootdown Algorithm with Page Access Tracking. In 2017 USENIX Annual Technical Conference (USENIX ATC 17). USENIX Association, Santa Clara, CA, 27\u201339. https:\/\/www.usenix.org\/conference\/atc17\/technical-sessions\/presentation\/amit"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3140659.3080231"},{"key":"e_1_3_2_1_7_1","volume-title":"Mosaic: A GPU Memory Manager with Application-Transparent Support for Multiple Page Sizes. In 2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 136\u2013150","author":"Ausavarungnirun Rachata","year":"2017","unstructured":"Rachata Ausavarungnirun , Joshua Landgraf , Vance Miller , Saugata Ghose , Jayneel Gandhi , Christopher\u00a0 J Rossbach , and Onur Mutlu . 2017 . Mosaic: A GPU Memory Manager with Application-Transparent Support for Multiple Page Sizes. In 2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 136\u2013150 . Rachata Ausavarungnirun, Joshua Landgraf, Vance Miller, Saugata Ghose, Jayneel Gandhi, Christopher\u00a0J Rossbach, and Onur Mutlu. 2017. Mosaic: A GPU Memory Manager with Application-Transparent Support for Multiple Page Sizes. In 2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 136\u2013150."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815970"},{"key":"e_1_3_2_1_9_1","volume-title":"2011 38th Annual International Symposium on Computer Architecture (ISCA). 307\u2013317","author":"Barr W","year":"2011","unstructured":"Thomas\u00a0 W Barr , Alan\u00a0 L Cox , and Scott Rixner . 2011 . SpecTLB: A mechanism for speculative address translation . In 2011 38th Annual International Symposium on Computer Architecture (ISCA). 307\u2013317 . Thomas\u00a0W Barr, Alan\u00a0L Cox, and Scott Rixner. 2011. SpecTLB: A mechanism for speculative address translation. In 2011 38th Annual International Symposium on Computer Architecture (ISCA). 307\u2013317."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00055"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3410463.3414639"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485943"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1122445.1122456"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/2540708.2540741"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.2200\/S00795ED1V01Y201708CAC042"},{"key":"e_1_3_2_1_16_1","volume-title":"2019 IEEE\/ACM Workshop on Memory Centric High Performance Computing (MCHPC). IEEE, 50\u201357","author":"Chien Steven","year":"2019","unstructured":"Steven Chien , Ivy Peng , and Stefano Markidis . 2019 . Performance evaluation of advanced features in CUDA unified memory . In 2019 IEEE\/ACM Workshop on Memory Centric High Performance Computing (MCHPC). IEEE, 50\u201357 . Steven Chien, Ivy Peng, and Stefano Markidis. 2019. Performance evaluation of advanced features in CUDA unified memory. In 2019 IEEE\/ACM Workshop on Memory Centric High Performance Computing (MCHPC). IEEE, 50\u201357."},{"key":"e_1_3_2_1_17_1","volume-title":"Cache hierarchy and memory subsystem of the AMD Opteron processor","author":"Conway Pat","year":"2010","unstructured":"Pat Conway , Nathan Kalyanasundharam , Gregg Donley , Kevin Lepak , and Bill Hughes . 2010. Cache hierarchy and memory subsystem of the AMD Opteron processor . IEEE micro 30, 2 ( 2010 ), 16\u201329. Pat Conway, Nathan Kalyanasundharam, Gregg Donley, Kevin Lepak, and Bill Hughes. 2010. Cache hierarchy and memory subsystem of the AMD Opteron processor. IEEE micro 30, 2 (2010), 16\u201329."},{"key":"e_1_3_2_1_18_1","unstructured":"NVIDIA Corp. 2020. NVIDIA A100 Tensor Core GPU Architecture. https:\/\/images.nvidia.cn\/aem-dam\/en-zz\/Solutions\/data-center\/nvidia-ampere-architecture-whitepaper.pdf  NVIDIA Corp. 2020. NVIDIA A100 Tensor Core GPU Architecture. https:\/\/images.nvidia.cn\/aem-dam\/en-zz\/Solutions\/data-center\/nvidia-ampere-architecture-whitepaper.pdf"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037704"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10070956"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1735688.1735702"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2499368.2451157"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3038228.3038239"},{"key":"e_1_3_2_1_24_1","volume-title":"Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU Systems. arXiv preprint arXiv:2203.15981","author":"Dutta Sankha\u00a0Baran","year":"2022","unstructured":"Sankha\u00a0Baran Dutta , Hoda Naghibijouybari , Arjun Gupta , Nael Abu-Ghazaleh , Andres Marquez , and Kevin Barker . 2022. Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU Systems. arXiv preprint arXiv:2203.15981 ( 2022 ). Sankha\u00a0Baran Dutta, Hoda Naghibijouybari, Arjun Gupta, Nael Abu-Ghazaleh, Andres Marquez, and Kevin Barker. 2022. Spy in the GPU-box: Covert and Side Channel Attacks on Multi-GPU Systems. arXiv preprint arXiv:2203.15981 (2022)."},{"key":"e_1_3_2_1_25_1","volume-title":"Medical image processing on the GPU\u2013Past, present and future. Medical image analysis 17, 8","author":"Eklund Anders","year":"2013","unstructured":"Anders Eklund , Paul Dufort , Daniel Forsberg , and Stephen\u00a0 M LaConte . 2013. Medical image processing on the GPU\u2013Past, present and future. Medical image analysis 17, 8 ( 2013 ), 1073\u20131094. Anders Eklund, Paul Dufort, Daniel Forsberg, and Stephen\u00a0M LaConte. 2013. Medical image processing on the GPU\u2013Past, present and future. Medical image analysis 17, 8 (2013), 1073\u20131094."},{"key":"e_1_3_2_1_26_1","volume-title":"2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). IEEE, 202\u2013211","author":"Feng Siying","year":"2019","unstructured":"Siying Feng , Subhankar Pal , Yichen Yang , and Ronald\u00a0 G Dreslinski . 2019 . Parallelism analysis of prominent desktop applications: An 18-year perspective . In 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). IEEE, 202\u2013211 . Siying Feng, Subhankar Pal, Yichen Yang, and Ronald\u00a0G Dreslinski. 2019. Parallelism analysis of prominent desktop applications: An 18-year perspective. In 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). IEEE, 202\u2013211."},{"key":"e_1_3_2_1_27_1","volume-title":"2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 451\u2013461","author":"Ganguly Debashis","year":"2020","unstructured":"Debashis Ganguly , Ziyu Zhang , Jun Yang , and Rami Melhem . 2020 . Adaptive page migration for irregular data-intensive applications under GPU memory oversubscription . In 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 451\u2013461 . Debashis Ganguly, Ziyu Zhang, Jun Yang, and Rami Melhem. 2020. Adaptive page migration for irregular data-intensive applications under GPU memory oversubscription. In 2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS). IEEE, 451\u2013461."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2591635.2667189"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3373376.3378494"},{"key":"e_1_3_2_1_30_1","unstructured":"Intel. 2018. The Future of Core Intel GPUs 10nm and Hybrid x86. https:\/\/www.anandtech.com\/show\/13699\/intel-architecture-day-2018-core-future-hybrid-x86\/5  Intel. 2018. The Future of Core Intel GPUs 10nm and Hybrid x86. https:\/\/www.anandtech.com\/show\/13699\/intel-architecture-day-2018-core-future-hybrid-x86\/5"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2749471"},{"key":"e_1_3_2_1_32_1","volume-title":"2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1022\u20131036","author":"Khairy Mahmoud","year":"2020","unstructured":"Mahmoud Khairy , Vadim Nikiforov , David Nellans , and Timothy\u00a0 G Rogers . 2020 . Locality-centric data and threadblock management for massive GPUs . In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1022\u20131036 . Mahmoud Khairy, Vadim Nikiforov, David Nellans, and Timothy\u00a0G Rogers. 2020. Locality-centric data and threadblock management for massive GPUs. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1022\u20131036."},{"key":"e_1_3_2_1_33_1","unstructured":"Ian King. 2017. Chipmakers Nvidia AMD Ride Cryptocurrency Wave\u2014for Now. www.bloomberg.com\/news\/articles\/2017-07-17\/chipmakers-nvidia-amd-ride-cryptocurrency-wave-for-now.  Ian King. 2017. Chipmakers Nvidia AMD Ride Cryptocurrency Wave\u2014for Now. www.bloomberg.com\/news\/articles\/2017-07-17\/chipmakers-nvidia-amd-ride-cryptocurrency-wave-for-now."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173162.3173198"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.5555\/3026877.3026931"},{"key":"e_1_3_2_1_36_1","unstructured":"Ya Le and Xuan Yang. 2015. Tiny imagenet visual recognition challenge. http:\/\/cs231n.stanford.edu\/  Ya Le and Xuan Yang. 2015. Tiny imagenet visual recognition challenge. http:\/\/cs231n.stanford.edu\/"},{"key":"e_1_3_2_1_37_1","volume-title":"SnakeByte: A TLB Design with Adaptive and Recursive Page Merging in GPUs. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 1195\u20131207","author":"Lee Jiwon","year":"2023","unstructured":"Jiwon Lee , Ju\u00a0Min Lee , Yunho Oh , William\u00a0 J Song , and Won\u00a0Woo Ro . 2023 . SnakeByte: A TLB Design with Adaptive and Recursive Page Merging in GPUs. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 1195\u20131207 . Jiwon Lee, Ju\u00a0Min Lee, Yunho Oh, William\u00a0J Song, and Won\u00a0Woo Ro. 2023. SnakeByte: A TLB Design with Adaptive and Recursive Page Merging in GPUs. In 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 1195\u20131207."},{"key":"e_1_3_2_1_38_1","volume-title":"Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUs. In The 60th ACM\/IEEE Design Automation Conference (DAC).","author":"Li Bingyao","year":"2023","unstructured":"Bingyao Li , Yueqi Wang , and Xulong Tang . 2023 . Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUs. In The 60th ACM\/IEEE Design Automation Conference (DAC). Bingyao Li, Yueqi Wang, and Xulong Tang. 2023. Orchestrated Scheduling and Partitioning for Improved Address Translation in GPUs. In The 60th ACM\/IEEE Design Automation Conference (DAC)."},{"key":"e_1_3_2_1_39_1","volume-title":"Optimizing Data Layout for Training Deep Neural Networks. In Companion Proceedings of the Web Conference","author":"Li Bingyao","year":"2022","unstructured":"Bingyao Li , Qi Xue , Geng Yuan , Sheng Li , Xiaolong Ma , Yanzhi Wang , and Xulong Tang . 2022 . Optimizing Data Layout for Training Deep Neural Networks. In Companion Proceedings of the Web Conference 2022. 548\u2013554. Bingyao Li, Qi Xue, Geng Yuan, Sheng Li, Xiaolong Ma, Yanzhi Wang, and Xulong Tang. 2022. Optimizing Data Layout for Training Deep Neural Networks. In Companion Proceedings of the Web Conference 2022. 548\u2013554."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA56546.2023.10071054"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480083"},{"key":"e_1_3_2_1_42_1","volume-title":"The Eleventh International Conference on Learning Representations.","author":"Li Sheng","year":"2022","unstructured":"Sheng Li , Geng Yuan , Yue Dai , Youtao Zhang , Yanzhi Wang , and Xulong Tang . 2022 . SmartFRZ: An Efficient Training Framework using Attention-Based Layer Freezing . In The Eleventh International Conference on Learning Representations. Sheng Li, Geng Yuan, Yue Dai, Youtao Zhang, Yanzhi Wang, and Xulong Tang. 2022. SmartFRZ: An Efficient Training Framework using Attention-Based Layer Freezing. In The Eleventh International Conference on Learning Representations."},{"key":"e_1_3_2_1_43_1","volume-title":"Unified-TP: A Unified TLB and Page Table Cache Structure for Efficient Address Translation. In 2020 IEEE 38th International Conference on Computer Design (ICCD). IEEE, 255\u2013262","author":"Ma Zhulin","year":"2020","unstructured":"Zhulin Ma , Yujuan Tan , Hong Jiang , Zhichao Yan , Duo Liu , Xianzhang Chen , Qingfeng Zhuge , Edwin Hsing-Mean Sha , and Chengliang Wang . 2020 . Unified-TP: A Unified TLB and Page Table Cache Structure for Efficient Address Translation. In 2020 IEEE 38th International Conference on Computer Design (ICCD). IEEE, 255\u2013262 . Zhulin Ma, Yujuan Tan, Hong Jiang, Zhichao Yan, Duo Liu, Xianzhang Chen, Qingfeng Zhuge, Edwin Hsing-Mean Sha, and Chengliang Wang. 2020. Unified-TP: A Unified TLB and Page Table Cache Structure for Efficient Address Translation. In 2020 IEEE 38th International Conference on Computer Design (ICCD). IEEE, 255\u2013262."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358294"},{"key":"e_1_3_2_1_45_1","volume-title":"2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 507\u2013519","author":"Mazumdar Chandrashis","year":"2021","unstructured":"Chandrashis Mazumdar , Prachatos Mitra , and Arkaprava Basu . 2021 . Dead page and dead block predictors: Cleaning tlbs and caches together . In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 507\u2013519 . Chandrashis Mazumdar, Prachatos Mitra, and Arkaprava Basu. 2021. Dead page and dead block predictors: Cleaning tlbs and caches together. In 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). IEEE, 507\u2013519."},{"key":"e_1_3_2_1_46_1","volume-title":"Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture. 123\u2013135","author":"Milic Ugljesa","year":"2017","unstructured":"Ugljesa Milic , Oreste Villa , Evgeny Bolotin , Akhil Arunkumar , Eiman Ebrahimi , Aamer Jaleel , Alex Ramirez , and David Nellans . 2017 . Beyond the socket: NUMA-aware GPUs . In Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture. 123\u2013135 . Ugljesa Milic, Oreste Villa, Evgeny Bolotin, Akhil Arunkumar, Eiman Ebrahimi, Aamer Jaleel, Alex Ramirez, and David Nellans. 2017. Beyond the socket: NUMA-aware GPUs. In Proceedings of the 50th Annual IEEE\/ACM International Symposium on Microarchitecture. 123\u2013135."},{"key":"e_1_3_2_1_47_1","volume-title":"GPS: A Global Publish-Subscribe Model for Multi-GPU Memory Management. In MICRO-54: 54th Annual IEEE\/ACM International Symposium on Microarchitecture. 46\u201358","author":"Muthukrishnan Harini","year":"2021","unstructured":"Harini Muthukrishnan , Daniel Lustig , David Nellans , and Thomas Wenisch . 2021 . GPS: A Global Publish-Subscribe Model for Multi-GPU Memory Management. In MICRO-54: 54th Annual IEEE\/ACM International Symposium on Microarchitecture. 46\u201358 . Harini Muthukrishnan, Daniel Lustig, David Nellans, and Thomas Wenisch. 2021. GPS: A Global Publish-Subscribe Model for Multi-GPU Memory Management. In MICRO-54: 54th Annual IEEE\/ACM International Symposium on Microarchitecture. 46\u201358."},{"key":"e_1_3_2_1_48_1","unstructured":"NVIDIA. 2018. DB2 Launch Datasheet Deep Learning Letter WEB. https:\/\/www.scribd.com\/document\/336084072\/61681-DB2-Launch-Datasheet-Deep-Learning-Letter-WEB-NVidia-Deep-Learning-Box#  NVIDIA. 2018. DB2 Launch Datasheet Deep Learning Letter WEB. https:\/\/www.scribd.com\/document\/336084072\/61681-DB2-Launch-Datasheet-Deep-Learning-Letter-WEB-NVidia-Deep-Learning-Box#"},{"key":"e_1_3_2_1_49_1","unstructured":"NVIDIA. 2018. NVIDIA DGX-2. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/dgx-2\/dgx-2-print-datasheet-738070-nvidia-a4-web-uk.pdf  NVIDIA. 2018. NVIDIA DGX-2. https:\/\/www.nvidia.com\/content\/dam\/en-zz\/Solutions\/Data-Center\/dgx-2\/dgx-2-print-datasheet-738070-nvidia-a4-web-uk.pdf"},{"key":"e_1_3_2_1_50_1","unstructured":"NVIDIA. 2022. NVIDIA Linux Open GPU Kernel Module Source. https:\/\/github.com\/NVIDIA\/open-gpu-kernel-modules  NVIDIA. 2022. NVIDIA Linux Open GPU Kernel Module Source. https:\/\/github.com\/NVIDIA\/open-gpu-kernel-modules"},{"key":"e_1_3_2_1_51_1","unstructured":"NVIDIA Corp. 2018. EVERYTHING YOU NEED TO KNOW ABOUT UNIFIED MEMORY. https:\/\/on-demand.gputechconf.com\/gtc\/2018\/presentation\/s8430-everything-you-need-to-know-about-unified-memory.pdf  NVIDIA Corp. 2018. EVERYTHING YOU NEED TO KNOW ABOUT UNIFIED MEMORY. https:\/\/on-demand.gputechconf.com\/gtc\/2018\/presentation\/s8430-everything-you-need-to-know-about-unified-memory.pdf"},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3445814.3446709"},{"key":"e_1_3_2_1_53_1","volume-title":"2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 193\u2013206","author":"Parasar Mayank","year":"2018","unstructured":"Mayank Parasar , Abhishek Bhattacharjee , and Tushar Krishna . 2018 . SEESAW: Using Superpages to Improve VIPT Caches . In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 193\u2013206 . Mayank Parasar, Abhishek Bhattacharjee, and Tushar Krishna. 2018. SEESAW: Using Superpages to Improve VIPT Caches. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 193\u2013206."},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080217"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503222.3507718"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2014.6835964"},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.32"},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2830772.2830773"},{"key":"e_1_3_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541942"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2014.6835965"},{"key":"e_1_3_2_1_61_1","volume-title":"Improving GPU Multi-tenancy with Page Walk Stealing. In 2021 IEEE 27th International Symposium on High Performance Computer Architecture (HPCA).","author":"Pratheek B","year":"2021","unstructured":"B Pratheek , Neha Jawalkar , and Arkaprava Basu . 2021 . Improving GPU Multi-tenancy with Page Walk Stealing. In 2021 IEEE 27th International Symposium on High Performance Computer Architecture (HPCA). B Pratheek, Neha Jawalkar, and Arkaprava Basu. 2021. Improving GPU Multi-tenancy with Page Walk Stealing. In 2021 IEEE 27th International Symposium on High Performance Computer Architecture (HPCA)."},{"key":"e_1_3_2_1_62_1","volume-title":"Designing Virtual Memory System of MCM GPUs. In 2022 55th IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 404\u2013422","author":"Pratheek B","year":"2022","unstructured":"B Pratheek , Neha Jawalkar , and Arkaprava Basu . 2022 . Designing Virtual Memory System of MCM GPUs. In 2022 55th IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 404\u2013422 . B Pratheek, Neha Jawalkar, and Arkaprava Basu. 2022. Designing Virtual Memory System of MCM GPUs. In 2022 55th IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 404\u2013422."},{"key":"e_1_3_2_1_63_1","unstructured":"Nikolay Sakharnykh. 2017. Unified Memory on Pascal and Volta. http:\/\/on-demand.gputechconf.com\/gtc\/2017\/presentation\/s7285-nikolay-sakharnykh-unified-memory-on-pascal-and-volta.pdf  Nikolay Sakharnykh. 2017. Unified Memory on Pascal and Volta. http:\/\/on-demand.gputechconf.com\/gtc\/2017\/presentation\/s7285-nikolay-sakharnykh-unified-memory-on-pascal-and-volta.pdf"},{"key":"e_1_3_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.58"},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2018.00025"},{"key":"e_1_3_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2018.00036"},{"key":"e_1_3_2_1_67_1","volume-title":"Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems. 1093\u20131108","author":"Skarlatos Dimitrios","year":"2020","unstructured":"Dimitrios Skarlatos , Apostolos Kokolis , Tianyin Xu , and Josep Torrellas . 2020 . Elastic cuckoo page tables: Rethinking virtual memory translation for parallelism . In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems. 1093\u20131108 . Dimitrios Skarlatos, Apostolos Kokolis, Tianyin Xu, and Josep Torrellas. 2020. Elastic cuckoo page tables: Rethinking virtual memory translation for parallelism. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems. 1093\u20131108."},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322230"},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2016.7581262"},{"key":"e_1_3_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1145\/3410463.3414633"},{"key":"e_1_3_2_1_71_1","unstructured":"Tech Power Up. 2017. ETH Mining: Lower VRAM GPUs to be Rendered Unprofitable in Time. www.techpowerup.com\/234482\/eth-mining-lower-vram-gpus-to-be-rendered-unprofitable-in-time  Tech Power Up. 2017. ETH Mining: Lower VRAM GPUs to be Rendered Unprofitable in Time. www.techpowerup.com\/234482\/eth-mining-lower-vram-gpus-to-be-rendered-unprofitable-in-time"},{"key":"e_1_3_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2008.16"},{"key":"e_1_3_2_1_73_1","unstructured":"Tyler Allen. 2023. UVM performance evaluation. https:\/\/github.com\/tallendev\/uvm-eval  Tyler Allen. 2023. UVM performance evaluation. https:\/\/github.com\/tallendev\/uvm-eval"},{"key":"e_1_3_2_1_74_1","volume-title":"2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 85\u201398","author":"Vavouliotis Georgios","year":"2021","unstructured":"Georgios Vavouliotis , Lluc Alvarez , Vasileios Karakostas , Konstantinos Nikas , Nectarios Koziris , Daniel\u00a0 A Jim\u00e9nez , and Marc Casas . 2021 . Exploiting page table locality for agile TLB prefetching . In 2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 85\u201398 . Georgios Vavouliotis, Lluc Alvarez, Vasileios Karakostas, Konstantinos Nikas, Nectarios Koziris, Daniel\u00a0A Jim\u00e9nez, and Marc Casas. 2021. Exploiting page table locality for agile TLB prefetching. In 2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 85\u201398."},{"key":"e_1_3_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2016.7482091"},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322247"},{"key":"e_1_3_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322223"},{"key":"e_1_3_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080211"},{"key":"e_1_3_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372224.3419192"},{"key":"e_1_3_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173162.3173195"},{"key":"e_1_3_2_1_81_1","volume-title":"2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 339\u2013351","author":"Young Vinson","year":"2018","unstructured":"Vinson Young , Aamer Jaleel , Evgeny Bolotin , Eiman Ebrahimi , David Nellans , and Oreste Villa . 2018 . Combining HW\/SW mechanisms to improve NUMA performance of multi-GPU systems . In 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 339\u2013351 . Vinson Young, Aamer Jaleel, Evgeny Bolotin, Eiman Ebrahimi, David Nellans, and Oreste Villa. 2018. Combining HW\/SW mechanisms to improve NUMA performance of multi-GPU systems. In 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 339\u2013351."},{"key":"e_1_3_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1145\/3466752.3480056"}],"event":{"name":"MICRO '23: 56th Annual IEEE\/ACM International Symposium on Microarchitecture","location":"Toronto ON Canada","acronym":"MICRO '23","sponsor":["SIGMICRO ACM Special Interest Group on Microarchitectural Research and Processing"]},"container-title":["56th Annual IEEE\/ACM International Symposium on Microarchitecture"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3613424.3614269","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3613424.3614269","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3613424.3614269","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:21Z","timestamp":1750178781000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3613424.3614269"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,28]]},"references-count":82,"alternative-id":["10.1145\/3613424.3614269","10.1145\/3613424"],"URL":"https:\/\/doi.org\/10.1145\/3613424.3614269","relation":{},"subject":[],"published":{"date-parts":[[2023,10,28]]},"assertion":[{"value":"2023-12-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}