{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,5]],"date-time":"2026-03-05T15:46:45Z","timestamp":1772725605705,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":73,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,9,30]],"date-time":"2020-09-30T00:00:00Z","timestamp":1601424000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Science Foundation","award":["1763681 1629129 1931531 1629915"],"award-info":[{"award-number":["1763681 1629129 1931531 1629915"]}]},{"name":"University of Pittsburgh","award":["Startup fund"],"award-info":[{"award-number":["Startup fund"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,9,30]]},"DOI":"10.1145\/3410463.3414633","type":"proceedings-article","created":{"date-parts":[[2020,9,30]],"date-time":"2020-09-30T10:43:04Z","timestamp":1601462584000},"page":"191-204","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["Enhancing Address Translations in Throughput Processors via Compression"],"prefix":"10.1145","author":[{"given":"Xulong","family":"Tang","sequence":"first","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ziyu","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weizheng","family":"Xu","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mahmut Taylan","family":"Kandemir","sequence":"additional","affiliation":[{"name":"The Pennsylvania State University, University Park, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rami","family":"Melhem","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Pittsburgh, Pittsburgh, PA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,9,30]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2015.7056046"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080209"},{"key":"e_1_3_2_1_3_1","unstructured":"AMD Corp. 2016. I\/O Virtualization Technology(IOMMU) Specification. https:\/\/www.amd.com\/system\/files\/TechDocs\/48882_IOMMU.pdf  AMD Corp. 2016. I\/O Virtualization Technology(IOMMU) Specification. https:\/\/www.amd.com\/system\/files\/TechDocs\/48882_IOMMU.pdf"},{"key":"e_1_3_2_1_4_1","unstructured":"AMD Corp. 2017. Radeons Next-generation Vega Architecture. https:\/\/radeon.com\/_downloads\/vega-whitepaper-11.6.17.pdf  AMD Corp. 2017. Radeons Next-generation Vega Architecture. https:\/\/radeon.com\/_downloads\/vega-whitepaper-11.6.17.pdf"},{"key":"e_1_3_2_1_5_1","volume-title":"Mosaic: A GPU Memory Manager with Application-Transparent Support for Multiple Page Sizes. In 2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 136--150","author":"Ausavarungnirun R.","unstructured":"R. Ausavarungnirun , J. Landgraf , V. Miller , S. Ghose , J. Gandhi , C. J. Rossbach , and O. Mutlu . 2017 . Mosaic: A GPU Memory Manager with Application-Transparent Support for Multiple Page Sizes. In 2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 136--150 . R. Ausavarungnirun, J. Landgraf, V. Miller, S. Ghose, J. Gandhi, C. J. Rossbach, and O. Mutlu. 2017. Mosaic: A GPU Memory Manager with Application-Transparent Support for Multiple Page Sizes. In 2017 50th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 136--150."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173162.3173169"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815970"},{"key":"e_1_3_2_1_8_1","volume-title":"2011 38th Annual International Symposium on Computer Architecture (ISCA). 307--317","author":"Barr T. W.","unstructured":"T. W. Barr , A. L. Cox , and S. Rixner . 2011. SpecTLB: A mechanism for speculative address translation . In 2011 38th Annual International Symposium on Computer Architecture (ISCA). 307--317 . T. W. Barr, A. L. Cox, and S. Rixner. 2011. SpecTLB: A mechanism for speculative address translation. In 2011 38th Annual International Symposium on Computer Architecture (ISCA). 307--317."},{"key":"e_1_3_2_1_9_1","volume-title":"Proceedings of the 40th Annual International Symposium on Computer Architecture (Tel-Aviv, Israel) (ISCA '13)","author":"Basu Arkaprava","unstructured":"Arkaprava Basu , Jayneel Gandhi , Jichuan Chang , Mark D. Hill , and Michael M. Swift . 2013. Efficient Virtual Memory for Big Memory Servers . In Proceedings of the 40th Annual International Symposium on Computer Architecture (Tel-Aviv, Israel) (ISCA '13) . ACM, New York, NY, USA, 237--248. https:\/\/doi.org\/10.1145\/2485922.2485943 10.1145\/2485922.2485943 Arkaprava Basu, Jayneel Gandhi, Jichuan Chang, Mark D. Hill, and Michael M. Swift. 2013. Efficient Virtual Memory for Big Memory Servers. In Proceedings of the 40th Annual International Symposium on Computer Architecture (Tel-Aviv, Israel) (ISCA '13). ACM, New York, NY, USA, 237--248. https:\/\/doi.org\/10.1145\/2485922.2485943"},{"key":"e_1_3_2_1_10_1","volume-title":"Scalable Distributed Last-Level TLBs Using Low-Latency Interconnects. In 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 271--284","author":"Bharadwaj S.","year":"2018","unstructured":"S. Bharadwaj , G. Cox , T. Krishna , and A. Bhattacharjee . 2018 . Scalable Distributed Last-Level TLBs Using Low-Latency Interconnects. In 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 271--284 . https:\/\/doi.org\/10.1109\/MICRO. 2018 .00030 10.1109\/MICRO.2018.00030 S. Bharadwaj, G. Cox, T. Krishna, and A. Bhattacharjee. 2018. Scalable Distributed Last-Level TLBs Using Low-Latency Interconnects. In 2018 51st Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 271--284. https:\/\/doi.org\/10.1109\/MICRO.2018.00030"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2540708.2540741"},{"key":"e_1_3_2_1_12_1","volume-title":"Translation-Triggered Prefetching. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems (Xi'an, China) (ASPLOS '17)","author":"Bhattacharjee Abhishek","year":"2017","unstructured":"Abhishek Bhattacharjee . 2017 . Translation-Triggered Prefetching. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems (Xi'an, China) (ASPLOS '17) . ACM, New York, NY, USA, 63--76. https:\/\/doi.org\/10.1145\/3037697.3037705 10.1145\/3037697.3037705 Abhishek Bhattacharjee. 2017. Translation-Triggered Prefetching. In Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems (Xi'an, China) (ASPLOS '17). ACM, New York, NY, USA, 63--76. https:\/\/doi.org\/10.1145\/3037697.3037705"},{"key":"e_1_3_2_1_13_1","volume-title":"Furenlid","author":"Caucci Luca","year":"2015","unstructured":"Luca Caucci and Lars R . Furenlid . 2015 . GPU programming for biomedical imaging. In Medical Applications of Radiation Detectors V,, H. Bradford Barber, Lars R. Furenlid, and Hans N. Roehrig (Eds.), Vol. 9594 . International Society for Optics and Photonics, SPIE , 79--93. https:\/\/doi.org\/10.1117\/12.2195217 10.1117\/12.2195217 Luca Caucci and Lars R. Furenlid. 2015. GPU programming for biomedical imaging. In Medical Applications of Radiation Detectors V,, H. Bradford Barber, Lars R. Furenlid, and Hans N. Roehrig (Eds.), Vol. 9594. International Society for Optics and Photonics, SPIE, 79--93. https:\/\/doi.org\/10.1117\/12.2195217"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2013.6704684"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037704"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2737924.2737989"},{"key":"e_1_3_2_1_18_1","volume-title":"2015 ACM\/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA). 92--104","author":"Du Z.","unstructured":"Z. Du , R. Fasthuber , T. Chen , P. Ienne , L. Li , T. Luo , X. Feng , Y. Chen , and O. Temam . 2015. ShiDianNao: Shifting vision processing closer to the sensor . In 2015 ACM\/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA). 92--104 . https:\/\/doi.org\/10.1145\/2749460779.2750389 10.1145\/2749460779.2750389 Z. Du, R. Fasthuber, T. Chen, P. Ienne, L. Li, T. Luo, X. Feng, Y. Chen, and O. Temam. 2015. ShiDianNao: Shifting vision processing closer to the sensor. In 2015 ACM\/IEEE 42nd Annual International Symposium on Computer Architecture (ISCA). 92--104. https:\/\/doi.org\/10.1145\/2749460779.2750389"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322224"},{"key":"#cr-split#-e_1_3_2_1_20_1.1","doi-asserted-by":"crossref","unstructured":"S. Grauer-Gray L. Xu R. Searles S. Ayalasomayajula and J. Cavazos. 2012. Auto-tuning a high-level language targeted to GPU codes. In 2012 Innovative Parallel Computing (InPar). 1--10. https:\/\/doi.org\/10.1109\/InPar.2012.6339595 10.1109\/InPar.2012.6339595","DOI":"10.1109\/InPar.2012.6339595"},{"key":"#cr-split#-e_1_3_2_1_20_1.2","doi-asserted-by":"crossref","unstructured":"S. Grauer-Gray L. Xu R. Searles S. Ayalasomayajula and J. Cavazos. 2012. Auto-tuning a high-level language targeted to GPU codes. In 2012 Innovative Parallel Computing (InPar). 1--10. https:\/\/doi.org\/10.1109\/InPar.2012.6339595","DOI":"10.1109\/InPar.2012.6339595"},{"key":"e_1_3_2_1_21_1","volume-title":"Supporting Address Translation for Accelerator-Centric Architectures. In 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA). 37--48","author":"Hao Y.","year":"2017","unstructured":"Y. Hao , Z. Fang , G. Reinman , and J. Cong . 2017 . Supporting Address Translation for Accelerator-Centric Architectures. In 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA). 37--48 . https:\/\/doi.org\/10.1109\/HPCA. 2017 .19 10.1109\/HPCA.2017.19 Y. Hao, Z. Fang, G. Reinman, and J. Cong. 2017. Supporting Address Translation for Accelerator-Centric Architectures. In 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA). 37--48. https:\/\/doi.org\/10.1109\/HPCA.2017.19"},{"key":"e_1_3_2_1_22_1","volume-title":"Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems","author":"Haria Swapnil","unstructured":"Swapnil Haria , Mark D. Hill , and Michael M. Swift . 2018. Devirtualizing Memory in Heterogeneous Systems . In Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems ( Williamsburg, VA, USA) (ASPLOS '18). ACM, New York, NY, USA, 637--650. https:\/\/doi.org\/10.1145\/3173162.3173194 10.1145\/3173162.3173194 Swapnil Haria, Mark D. Hill, and Michael M. Swift. 2018. Devirtualizing Memory in Heterogeneous Systems. In Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems (Williamsburg, VA, USA) (ASPLOS '18). ACM, New York, NY, USA, 637--650. https:\/\/doi.org\/10.1145\/3173162.3173194"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2591635.2667189"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2745844.2745867"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2749471"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2731186.2731192"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3192366.3192386"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2009.4919639"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173162.3173198"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/3026877.3026931"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2611758"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358294"},{"key":"e_1_3_2_1_33_1","volume-title":"ACM Comput. Surv.","volume":"47","author":"Mittal Sparsh","year":"2015","unstructured":"Sparsh Mittal and Jeffrey S. Vetter . 2015. A Survey of CPU-GPU Heterogeneous Computing Techniques . ACM Comput. Surv. , Vol. 47 , 4, Article 69 ( July 2015 ), 35 pages. https:\/\/doi.org\/10.1145\/2788396 10.1145\/2788396 Sparsh Mittal and Jeffrey S. Vetter. 2015. A Survey of CPU-GPU Heterogeneous Computing Techniques. ACM Comput. Surv., Vol. 47, 4, Article 69 (July 2015), 35 pages. https:\/\/doi.org\/10.1145\/2788396"},{"key":"e_1_3_2_1_34_1","unstructured":"NVIDIA Corp. 2016. NVIDIA Tesla P100. https:\/\/images.nvidia.com\/content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf  NVIDIA Corp. 2016. NVIDIA Tesla P100. https:\/\/images.nvidia.com\/content\/pdf\/tesla\/whitepaper\/pascal-architecture-whitepaper.pdf"},{"key":"e_1_3_2_1_35_1","unstructured":"NVIDIA Corp. 2018. NVIDIA Pascal Architecture. https:\/\/www.nvidia.com\/en-us\/data-center\/pascal-gpu-architecture\/  NVIDIA Corp. 2018. NVIDIA Pascal Architecture. https:\/\/www.nvidia.com\/en-us\/data-center\/pascal-gpu-architecture\/"},{"key":"e_1_3_2_1_36_1","volume-title":"2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 193--206","author":"Parasar M.","unstructured":"M. Parasar , A. Bhattacharjee , and T. Krishna . 2018. SEESAW: Using Superpages to Improve VIPT Caches . In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 193--206 . M. Parasar, A. Bhattacharjee, and T. Krishna. 2018. SEESAW: Using Superpages to Improve VIPT Caches. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 193--206."},{"key":"e_1_3_2_1_37_1","volume-title":"2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA). 444--456","author":"Park C. H.","unstructured":"C. H. Park , T. Heo , J. Jeong , and J. Huh . 2017. Hybrid TLB coalescing: Improving TLB translation coverage under diverse fragmented memory allocations . In 2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA). 444--456 . https:\/\/doi.org\/10.1145\/3079856.3080217 10.1145\/3079856.3080217 C. H. Park, T. Heo, J. Jeong, and J. Huh. 2017. Hybrid TLB coalescing: Improving TLB translation coverage under diverse fragmented memory allocations. In 2017 ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA). 444--456. https:\/\/doi.org\/10.1145\/3079856.3080217"},{"key":"e_1_3_2_1_38_1","volume-title":"Automation Test in Europe Conference Exhibition (DATE). 1341--1346","author":"Park E.","unstructured":"E. Park , J. Ahn , S. Hong , S. Yoo , and S. Lee . 2015. Memory fast-forward: A low cost special function unit to enhance energy efficiency in GPU for big data processing. In 2015 Design , Automation Test in Europe Conference Exhibition (DATE). 1341--1346 . E. Park, J. Ahn, S. Hong, S. Yoo, and S. Lee. 2015. Memory fast-forward: A low cost special function unit to enhance energy efficiency in GPU for big data processing. In 2015 Design, Automation Test in Europe Conference Exhibition (DATE). 1341--1346."},{"key":"e_1_3_2_1_39_1","volume-title":"Proceedings of the 2016 International Conference on Parallel Architectures and Compilation (PACT).","author":"Pattnaik Ashutosh","unstructured":"Ashutosh Pattnaik , Xulong Tang , Adwait Jog , Onur Kayiran , Asit K. Mishra , Mahmut T. Kandemir , Onur Mutlu , and Chita R. Das . 2016. Scheduling Techniques for GPU Architectures with Processing-In-Memory Capabilities . In Proceedings of the 2016 International Conference on Parallel Architectures and Compilation (PACT). Ashutosh Pattnaik, Xulong Tang, Adwait Jog, Onur Kayiran, Asit K. Mishra, Mahmut T. Kandemir, Onur Mutlu, and Chita R. Das. 2016. Scheduling Techniques for GPU Architectures with Processing-In-Memory Capabilities. In Proceedings of the 2016 International Conference on Parallel Architectures and Compilation (PACT)."},{"key":"e_1_3_2_1_40_1","volume-title":"Proceedings of the 46th International Symposium on Computer Architecture.","author":"Pattnaik Ashutosh","unstructured":"Ashutosh Pattnaik , Xulong Tang , Onur Kayiran , Adwait Jog , Asit Mishra , Mahmut T. Kandemir , Anand Sivasubramaniam , and Chita R. Das . 2019. Opportunistic Computing in GPU Architectures . In Proceedings of the 46th International Symposium on Computer Architecture. Ashutosh Pattnaik, Xulong Tang, Onur Kayiran, Adwait Jog, Asit Mishra, Mahmut T. Kandemir, Anand Sivasubramaniam, and Chita R. Das. 2019. Opportunistic Computing in GPU Architectures. In Proceedings of the 46th International Symposium on Computer Architecture."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370816.2370870"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2014.6835964"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2017.2712140"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.32"},{"key":"e_1_3_2_1_45_1","volume-title":"2015 48th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 1--12","author":"Pham B.","unstructured":"B. Pham , J. Vesel\u00fd , G. H. Loh , and A. Bhattacharjee . 2015. Large pages and lightweight memory management in virtualized environments: Can you have it both ways? . In 2015 48th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 1--12 . https:\/\/doi.org\/10.1145\/2830772.2830773 10.1145\/2830772.2830773 B. Pham, J. Vesel\u00fd, G. H. Loh, and A. Bhattacharjee. 2015. Large pages and lightweight memory management in virtualized environments: Can you have it both ways?. In 2015 48th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). 1--12. https:\/\/doi.org\/10.1145\/2830772.2830773"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541942"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2015.44"},{"key":"e_1_3_2_1_48_1","volume-title":"Near-Memory Address Translation. In 2017 26th International Conference on Parallel Architectures and Compilation Techniques (PACT). 303--317","author":"Picorel J.","year":"2017","unstructured":"J. Picorel , D. Jevdjic , and B. Falsafi . 2017 . Near-Memory Address Translation. In 2017 26th International Conference on Parallel Architectures and Compilation Techniques (PACT). 303--317 . https:\/\/doi.org\/10.1109\/PACT. 2017 .56 10.1109\/PACT.2017.56 J. Picorel, D. Jevdjic, and B. Falsafi. 2017. Near-Memory Address Translation. In 2017 26th International Conference on Parallel Architectures and Compilation Techniques (PACT). 303--317. https:\/\/doi.org\/10.1109\/PACT.2017.56"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2014.2299539"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2014.6835965"},{"key":"e_1_3_2_1_51_1","volume-title":"Aamodt","author":"Rogers Timothy G.","year":"2012","unstructured":"Timothy G. Rogers , Mike O'Connor , and Tor M . Aamodt . 2012 . Cache-Conscious Wavefront Scheduling. In MICRO. Timothy G. Rogers, Mike O'Connor, and Tor M. Aamodt. 2012. Cache-Conscious Wavefront Scheduling. In MICRO."},{"key":"e_1_3_2_1_52_1","volume-title":"Architecture-Centric Bottleneck Analysis for Deep Neural Network Applications. In 2019 IEEE 26th International Conference on High Performance Computing, Data, and Analytics (HiPC). IEEE, 205--214","author":"Ryoo Jihyun","year":"2019","unstructured":"Jihyun Ryoo , Mengran Fan , Xulong Tang , Huaipan Jiang , Meena Arunachalam , Sharada Naveen , and Mahmut T Kandemir . 2019 . Architecture-Centric Bottleneck Analysis for Deep Neural Network Applications. In 2019 IEEE 26th International Conference on High Performance Computing, Data, and Analytics (HiPC). IEEE, 205--214 . Jihyun Ryoo, Mengran Fan, Xulong Tang, Huaipan Jiang, Meena Arunachalam, Sharada Naveen, and Mahmut T Kandemir. 2019. Architecture-Centric Bottleneck Analysis for Deep Neural Network Applications. In 2019 IEEE 26th International Conference on High Performance Computing, Data, and Analytics (HiPC). IEEE, 205--214."},{"key":"e_1_3_2_1_53_1","volume-title":"Dynamic Aggregation of Virtual Addresses in TLB Using TCAM Cells. In 21st International Conference on VLSI Design (VLSID 2008","author":"Samanta R.","year":"2008","unstructured":"R. Samanta , J. Surprise , and R. Mahapatr . 2008 . Dynamic Aggregation of Virtual Addresses in TLB Using TCAM Cells. In 21st International Conference on VLSI Design (VLSID 2008 ). 243--248. https:\/\/doi.org\/10.1109\/VLSI. 2008 .57 10.1109\/VLSI.2008.57 R. Samanta, J. Surprise, and R. Mahapatr. 2008. Dynamic Aggregation of Virtual Addresses in TLB Using TCAM Cells. In 21st International Conference on VLSI Design (VLSID 2008). 243--248. https:\/\/doi.org\/10.1109\/VLSI.2008.57"},{"key":"e_1_3_2_1_54_1","volume-title":"10th Dimacs Implementation Challenge-Graph Partitioning and Graph Clustering. (2012)","author":"Sanders Peter","unstructured":"Peter Sanders and Christian Schulz . 2012. 10th Dimacs Implementation Challenge-Graph Partitioning and Graph Clustering. (2012) . Peter Sanders and Christian Schulz. 2012. 10th Dimacs Implementation Challenge-Graph Partitioning and Graph Clustering. (2012)."},{"key":"e_1_3_2_1_55_1","volume-title":"ActivePointers: A Case for Software Address Translation on GPUs. In 2016 ACM\/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA). 596--608","author":"Shahar S.","year":"2016","unstructured":"S. Shahar , S. Bergman , and M. Silberstein . 2016 . ActivePointers: A Case for Software Address Translation on GPUs. In 2016 ACM\/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA). 596--608 . https:\/\/doi.org\/10.1109\/ISCA. 2016 .58 10.1109\/ISCA.2016.58 S. Shahar, S. Bergman, and M. Silberstein. 2016. ActivePointers: A Case for Software Address Translation on GPUs. In 2016 ACM\/IEEE 43rd Annual International Symposium on Computer Architecture (ISCA). 596--608. https:\/\/doi.org\/10.1109\/ISCA.2016.58"},{"key":"e_1_3_2_1_56_1","volume-title":"Scheduling Page Table Walks for Irregular GPU Applications. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 180--192","author":"Shin S.","year":"2018","unstructured":"S. Shin , G. Cox , M. Oskin , G. H. Loh , Y. Solihin , A. Bhattacharjee , and A. Basu . 2018 . Scheduling Page Table Walks for Irregular GPU Applications. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 180--192 . https:\/\/doi.org\/10.1109\/ISCA. 2018 .00025 10.1109\/ISCA.2018.00025 S. Shin, G. Cox, M. Oskin, G. H. Loh, Y. Solihin, A. Bhattacharjee, and A. Basu. 2018. Scheduling Page Table Walks for Irregular GPU Applications. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 180--192. https:\/\/doi.org\/10.1109\/ISCA.2018.00025"},{"key":"e_1_3_2_1_57_1","volume-title":"Geng Daniel Liu, and Wen-mei W Hwu","author":"Stratton John A","year":"2012","unstructured":"John A Stratton , Christopher Rodrigues , I- Jui Sung , Nady Obeid , Li-Wen Chang , Nasser Anssari , Geng Daniel Liu, and Wen-mei W Hwu . 2012 . Parboil : A revised benchmark suite for scientific and commercial throughput computing. Center for Reliable and High-Performance Computing , Vol. 127 (2012). John A Stratton, Christopher Rodrigues, I-Jui Sung, Nady Obeid, Li-Wen Chang, Nasser Anssari, Geng Daniel Liu, and Wen-mei W Hwu. 2012. Parboil: A revised benchmark suite for scientific and commercial throughput computing. Center for Reliable and High-Performance Computing, Vol. 127 (2012)."},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195708"},{"key":"e_1_3_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3309697.3331487"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123954"},{"key":"e_1_3_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2017.14"},{"key":"e_1_3_2_1_62_1","volume-title":"Proceedings of the 2019 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems (SIGMETRICS).","author":"Tang Xulong","unstructured":"Xulong Tang , Ashutosh Pattnaik , Onur Kayiran , Adwait Jog , Mahmut Taylan Kandemir , and Chita R. Das . 2019 b. Quantifying Data Locality in Dynamic Parallelism in GPUs . In Proceedings of the 2019 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems (SIGMETRICS). Xulong Tang, Ashutosh Pattnaik, Onur Kayiran, Adwait Jog, Mahmut Taylan Kandemir, and Chita R. Das. 2019 b. Quantifying Data Locality in Dynamic Parallelism in GPUs. In Proceedings of the 2019 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems (SIGMETRICS)."},{"key":"e_1_3_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/3314221.3314599"},{"key":"e_1_3_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2008.16"},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPASS.2016.7482091"},{"key":"e_1_3_2_1_66_1","volume-title":"Generic System Calls for GPUs. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 843--856","author":"Vesel\u00fd J.","year":"2018","unstructured":"J. Vesel\u00fd , A. Basu , A. Bhattacharjee , G. H. Loh , M. Oskin , and S. K. Reinhardt . 2018 . Generic System Calls for GPUs. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 843--856 . https:\/\/doi.org\/10.1109\/ISCA. 2018 .00075 10.1109\/ISCA.2018.00075 J. Vesel\u00fd, A. Basu, A. Bhattacharjee, G. H. Loh, M. Oskin, and S. K. Reinhardt. 2018. Generic System Calls for GPUs. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA). 843--856. https:\/\/doi.org\/10.1109\/ISCA.2018.00075"},{"key":"e_1_3_2_1_67_1","volume-title":"2015 International Conference on Hardware\/Software Codesign and System Synthesis (CODESISSS). 45--54","author":"Vogel P.","year":"2015","unstructured":"P. Vogel , A. Marongiu , and L. Benini . 2015. Lightweight virtual memory support for many-core accelerators in heterogeneous embedded SoCs . In 2015 International Conference on Hardware\/Software Codesign and System Synthesis (CODESISSS). 45--54 . https:\/\/doi.org\/10.1109\/CODESISSS. 2015 .7331367 10.1109\/CODESISSS.2015.7331367 P. Vogel, A. Marongiu, and L. Benini. 2015. Lightweight virtual memory support for many-core accelerators in heterogeneous embedded SoCs. In 2015 International Conference on Hardware\/Software Codesign and System Synthesis (CODESISSS). 45--54. https:\/\/doi.org\/10.1109\/CODESISSS.2015.7331367"},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178491"},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322223"},{"key":"#cr-split#-e_1_3_2_1_70_1.1","doi-asserted-by":"crossref","unstructured":"S. Zhang Y. Yang L. Shen and Z. Wang. 2018. Efficient Data Communication between CPU and GPU through Transparent Partial-Page Migration. In 2018 IEEE 20th International Conference on High Performance Computing and Communications; IEEE 16th International Conference on Smart City; IEEE 4th International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS). 618--625. https:\/\/doi.org\/10.1109\/HPCC\/SmartCity\/DSS.2018.00112 10.1109\/HPCC","DOI":"10.1109\/HPCC\/SmartCity\/DSS.2018.00112"},{"key":"#cr-split#-e_1_3_2_1_70_1.2","doi-asserted-by":"crossref","unstructured":"S. Zhang Y. Yang L. Shen and Z. Wang. 2018. Efficient Data Communication between CPU and GPU through Transparent Partial-Page Migration. In 2018 IEEE 20th International Conference on High Performance Computing and Communications; IEEE 16th International Conference on Smart City; IEEE 4th International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS). 618--625. https:\/\/doi.org\/10.1109\/HPCC\/SmartCity\/DSS.2018.00112","DOI":"10.1109\/HPCC\/SmartCity\/DSS.2018.00112"},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2016.7446077"}],"event":{"name":"PACT '20: International Conference on Parallel Architectures and Compilation Techniques","location":"Virtual Event GA USA","acronym":"PACT '20","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture"]},"container-title":["Proceedings of the ACM International Conference on Parallel Architectures and Compilation Techniques"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3410463.3414633","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3410463.3414633","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3410463.3414633","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:31:51Z","timestamp":1750195911000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3410463.3414633"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,30]]},"references-count":73,"alternative-id":["10.1145\/3410463.3414633","10.1145\/3410463"],"URL":"https:\/\/doi.org\/10.1145\/3410463.3414633","relation":{},"subject":[],"published":{"date-parts":[[2020,9,30]]},"assertion":[{"value":"2020-09-30","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}