{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:31:01Z","timestamp":1750221061965,"version":"3.41.0"},"reference-count":143,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2018,8,28]],"date-time":"2018-08-28T00:00:00Z","timestamp":1535414400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGOPS Oper. Syst. Rev."],"published-print":{"date-parts":[[2018,8,28]]},"abstract":"<jats:p>Contemporary discrete GPUs support rich memory management features such as virtual memory and demand paging. These features simplify GPU programming by providing a virtual address space abstraction similar to CPUs and eliminating manual memory management, but they introduce high performance overheads during (1) address translation and (2) page faults. A GPU relies on high degrees of thread-level parallelism (TLP) to hide memory latency. Address translation can undermine TLP, as a single miss in the translation lookaside buffer (TLB) invokes an expensive serialized page table walk that often stalls multiple threads. Demand paging can also undermine TLP, as multiple threads often stall while they wait for an expensive data transfer over the system I\/O (e.g., PCIe) bus when the GPU demands a page.<\/jats:p>\n          <jats:p>In modern GPUs, we face a trade-off on how the page size used for memory management affects address translation and demand paging. The address translation overhead is lower when we employ a larger page size (e.g., 2MB large pages, compared with conventional 4KB base pages), which increases TLB coverage and thus reduces TLB misses. Conversely, the demand paging overhead is lower when we employ a smaller page size, which decreases the system I\/O bus transfer latency. Support for multiple page sizes can help relax the page size trade-off so that address translation and demand paging optimizations work together synergistically. However, existing page coalescing (i.e., merging base pages into a large page) and splintering (i.e., splitting a large page into base pages) policies require costly base page migrations that undermine the benefits multiple page sizes provide. In this paper, we observe that GPGPU applications present an opportunity to support multiple page sizes without costly data migration, as the applications perform most of their memory allocation en masse (i.e., they allocate a large number of base pages at once).We show that this en masse allocation allows us to create intelligent memory allocation policies which ensure that base pages that are contiguous in virtual memory are allocated to contiguous physical memory pages. As a result, coalescing and splintering operations no longer need to migrate base pages.<\/jats:p>","DOI":"10.1145\/3273982.3273986","type":"journal-article","created":{"date-parts":[[2018,8,30]],"date-time":"2018-08-30T13:45:11Z","timestamp":1535636711000},"page":"27-44","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Mosaic"],"prefix":"10.1145","volume":"52","author":[{"given":"Rachata","family":"Ausavarungnirun","sequence":"first","affiliation":[{"name":"Carnegie Mellon University &amp; King Mongkut University of Technology North Bangkok"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joshua","family":"Landgraf","sequence":"additional","affiliation":[{"name":"University of Texas at Austin"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Vance","family":"Miller","sequence":"additional","affiliation":[{"name":"University of Texas at Austin"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Saugata","family":"Ghose","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jayneel","family":"Gandhi","sequence":"additional","affiliation":[{"name":"VMware Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christopher J.","family":"Rossbach","sequence":"additional","affiliation":[{"name":"University of Texas at Austin &amp; VMware Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Onur","family":"Mutlu","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University &amp; ETH Z\u00fcrich"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,8,28]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"\"NVIDIA GRID \" http:\/\/www.nvidia.com\/object\/grid-boards.html.  \"NVIDIA GRID \" http:\/\/www.nvidia.com\/object\/grid-boards.html."},{"key":"e_1_2_1_2_1","unstructured":"A. Abrevaya \"Linux Transparent Huge Pages JEMalloc and NuoDB \" http:\/\/www.nuodb.com\/techblog\/ linux-transparent-huge-pages-jemalloc-and-nuodb 2014  A. Abrevaya \"Linux Transparent Huge Pages JEMalloc and NuoDB \" http:\/\/www.nuodb.com\/techblog\/ linux-transparent-huge-pages-jemalloc-and-nuodb 2014"},{"key":"e_1_2_1_3_1","unstructured":"Advanced Micro Devices \"AMD Accelerated Processing Units.\" {Online}. Available: http:\/\/www.amd.com\/us\/products\/technologies\/apu\/Pages\/apu.aspx  Advanced Micro Devices \"AMD Accelerated Processing Units.\" {Online}. Available: http:\/\/www.amd.com\/us\/products\/technologies\/apu\/Pages\/apu.aspx"},{"key":"e_1_2_1_4_1","unstructured":"Advanced Micro Devices Inc. \"OpenCL: The Future of Accelerated Application Performance Is Now \" https:\/\/www.amd.com\/Documents\/FirePro_OpenCL_ Whitepaper.pdf.  Advanced Micro Devices Inc. \"OpenCL: The Future of Accelerated Application Performance Is Now \" https:\/\/www.amd.com\/Documents\/FirePro_OpenCL_ Whitepaper.pdf."},{"volume-title":"Unlocking Bandwidth for GPUs in CC-NUMA Systems,\" in HPCA","year":"2015","author":"Agarwal N.","key":"e_1_2_1_5_1"},{"volume-title":"Revisiting Hardware-Assisted Page Walks for Virtualized Systems,\" in ISCA","year":"2012","author":"Ahn J.","key":"e_1_2_1_6_1"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2015.2401022"},{"key":"e_1_2_1_8_1","unstructured":"Apple Inc. \"Huge Page Support in Mac OS X \" http:\/\/blog.couchbase.com\/ often-overlooked-linux-os-tweaks 2014.  Apple Inc. \"Huge Page Support in Mac OS X \" http:\/\/blog.couchbase.com\/ often-overlooked-linux-os-tweaks 2014."},{"key":"e_1_2_1_9_1","unstructured":"ARM Holdings \"ARM Cortex-A Series \" http:\/\/infocenter.arm.com\/help\/topic\/ com.arm.doc.den0024a\/DEN0024A_v8_architecture_PG.pdf 2015.  ARM Holdings \"ARM Cortex-A Series \" http:\/\/infocenter.arm.com\/help\/topic\/ com.arm.doc.den0024a\/DEN0024A_v8_architecture_PG.pdf 2015."},{"volume-title":"Carnegie Mellon Univ.","year":"2017","author":"Ausavarungnirun R.","key":"e_1_2_1_10_1"},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","DOI":"10.1145\/2366231.2337207","volume-title":"Staged Memory Scheduling: Achieving High Performance and Scalability in Heterogeneous Systems,\" in ISCA","author":"Ausavarungnirun R.","year":"2012"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2015.38"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123939.3123975"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173162.3173169"},{"volume-title":"Analyzing CUDA Workloads Using a Detailed GPU Simulator,\" in ISPASS","year":"2009","author":"Bakhoda A.","key":"e_1_2_1_15_1"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815970"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2000064.2000101"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485943"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2540708.2540741"},{"volume-title":"Shared Last-level TLBs for Chip Multiprocessors,\" in HPCA","year":"2011","author":"Bhattacharjee A.","key":"e_1_2_1_20_1"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2009.26"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/1736020.1736060"},{"volume-title":"Applying AMD's \"Kaveri\" APU for Heterogeneous Computing,\" in HOTCHIP","year":"2014","author":"Bouvier D.","key":"e_1_2_1_23_1"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2011.2"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/1941553.1941562"},{"volume-title":"Low-Cost Inter-Linked Subarrays (LISA): Enabling Fast Inter-Subarray Data Movement in DRAM,\" in HPCA","year":"2016","author":"Chang K. K.","key":"e_1_2_1_26_1"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/IISWC.2009.5306797"},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","unstructured":"M. Clark \"A New X86 Core Architecture for the Next Generation of Computing \" in HotChips 2016  M. Clark \"A New X86 Core Architecture for the Next Generation of Computing \" in HotChips 2016","DOI":"10.1109\/HOTCHIPS.2016.7936224"},{"volume-title":"Supporting Address Translation for Accelerator-Centric Architectures,\" in HPCA","year":"2017","author":"Cong J.","key":"e_1_2_1_29_1"},{"key":"e_1_2_1_30_1","unstructured":"J. Corbet \"Transparent Hugepages \" https:\/\/lwn.net\/Articles\/359158\/ 2009.  J. Corbet \"Transparent Hugepages \" https:\/\/lwn.net\/Articles\/359158\/ 2009."},{"key":"e_1_2_1_31_1","unstructured":"Couchbase Inc. \"Often Overlooked Linux OS Tweaks \" http:\/\/blog.couchbase. com\/often-overlooked-linux-os-tweaks 2014.  Couchbase Inc. \"Often Overlooked Linux OS Tweaks \" http:\/\/blog.couchbase. com\/often-overlooked-linux-os-tweaks 2014."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037704"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1735688.1735702"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669150"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1815976"},{"volume-title":"Supporting Superpages in Non-Contiguous Physical Memory,\" in HPCA","year":"2015","author":"Du Y.","key":"e_1_2_1_36_1"},{"volume-title":"rCUDA: Reducing the Number of GPU-based Accelerators in High Performance Clusters,\" in HPCS","year":"2010","author":"Duato J.","key":"e_1_2_1_37_1"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2008.44"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/L-CA.2013.9"},{"volume-title":"of the IEEE","author":"Flynn M.","key":"e_1_2_1_40_1"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.12"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.37"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.67"},{"volume-title":"Large Pages May Be Harmful on NUMA Systems,\" in USENIX ATC","year":"2014","author":"Gaud F.","key":"e_1_2_1_44_1"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/277650.277748"},{"key":"e_1_2_1_46_1","unstructured":"M. Gorman \"Huge Pages Part 2 (Interfaces) \" https:\/\/lwn.net\/Articles\/375096\/ 2010.  M. Gorman \"Huge Pages Part 2 (Interfaces) \" https:\/\/lwn.net\/Articles\/375096\/ 2010."},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/1375634.1375641"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-24322-6_24"},{"volume-title":"Graphics Accelerated VDI with the Visual Performance of a Workstation","year":"2014","author":"Herrera A.","key":"e_1_2_1_49_1"},{"volume-title":"intel.com\/content\/dam\/www\/public\/us\/en\/documents\/white-papers\/ ia-introduction-basics-paper.pdf","year":"2014","author":"Intel Corp.","key":"e_1_2_1_50_1"},{"key":"e_1_2_1_51_1","unstructured":"Intel Corp. \"Intel\u00ae 64 and IA-32 Architectures Optimization Reference Manual \" https:\/\/www.intel.com\/content\/dam\/www\/public\/us\/en\/documents\/ manuals\/64-ia-32-architectures-optimization-manual.pdf 2016.  Intel Corp. \"Intel\u00ae 64 and IA-32 Architectures Optimization Reference Manual \" https:\/\/www.intel.com\/content\/dam\/www\/public\/us\/en\/documents\/ manuals\/64-ia-32-architectures-optimization-manual.pdf 2016."},{"key":"e_1_2_1_52_1","unstructured":"Intel Corp. \"6th Generation Intel\u00ae CoreTM Processor Family Datasheet Vol. 1 \" http:\/\/www.intel.com\/content\/dam\/www\/public\/us\/en\/documents\/ datasheets\/desktop-6th-gen-core-family-datasheet-vol-1.pdf 2017.  Intel Corp. \"6th Generation Intel\u00ae CoreTM Processor Family Datasheet Vol. 1 \" http:\/\/www.intel.com\/content\/dam\/www\/public\/us\/en\/documents\/ datasheets\/desktop-6th-gen-core-family-datasheet-vol-1.pdf 2017."},{"key":"e_1_2_1_53_1","unstructured":"Intel Corporation \"Sandy Bridge Intel Processor Graphics Performance Developer's Guide.\" {Online}. Available: http:\/\/software.intel.com\/file\/34436  Intel Corporation \"Sandy Bridge Intel Processor Graphics Performance Developer's Guide.\" {Online}. Available: http:\/\/software.intel.com\/file\/34436"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/2228360.2228513"},{"volume-title":"Pennsylvania State Univ.","year":"2015","author":"Jog A.","key":"e_1_2_1_55_1"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/2818950.2818979"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485951"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451116.2451158"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/2896377.2901468"},{"key":"e_1_2_1_60_1","doi-asserted-by":"crossref","DOI":"10.1145\/545214.545237","volume-title":"Going the Distance for TLB Prefetching: An Application-Driven Study,\" in ISCA","author":"Kandiraju G. B.","year":"2002"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2749471"},{"volume-title":"Energy-Efficient Address Translation,\" in HPCA","year":"2016","author":"Karakostas V.","key":"e_1_2_1_62_1"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2013.115"},{"key":"e_1_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.62"},{"volume-title":"khronos.org\/registry\/cl\/specs\/opencl-1.0.29.pdf","year":"2008","author":"Khronos OpenCL Working Group","key":"e_1_2_1_66_1"},{"volume-title":"ATLAS: A Scalable and High- Performance Scheduling Algorithm for Multiple Memory Controllers,\" in HPCA","year":"2010","author":"Kim Y.","key":"e_1_2_1_67_1"},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2010.51"},{"key":"e_1_2_1_69_1","unstructured":"D. Kroft \"Lockup-Free Instruction Fetch\/Prefetch Cache Organization \" in ISCA 1981.   D. Kroft \"Lockup-Free Instruction Fetch\/Prefetch Cache Organization \" in ISCA 1981."},{"volume-title":"Coordinated and Efficient Huge Page Management with Ingens,\" in OSDI","year":"2016","author":"Kwon Y.","key":"e_1_2_1_70_1"},{"volume-title":"Advanced Micro Devices","year":"2012","author":"Kyriazis G.","key":"e_1_2_1_71_1"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2015.51"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/2628071.2628075"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2008.31"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/2445572.2445574"},{"key":"e_1_2_1_76_1","unstructured":"Mark Mumy \"SAP IQ and Linux Hugepages\/Transparent Hugepages \" http:\/\/scn.sap.com\/people\/markmumy\/blog\/2014\/05\/22\/ sap-iq-and-linux-hugepagestransparent-hugepages SAP SE 2014.  Mark Mumy \"SAP IQ and Linux Hugepages\/Transparent Hugepages \" http:\/\/scn.sap.com\/people\/markmumy\/blog\/2014\/05\/22\/ sap-iq-and-linux-hugepagestransparent-hugepages SAP SE 2014."},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2549523"},{"key":"e_1_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1145\/2892242.2892258"},{"key":"e_1_2_1_79_1","unstructured":"Microsoft Corp. Large-Page Support in Windows https:\/\/msdn.microsoft.com\/ en-us\/library\/windows\/desktop\/aa366720(v=vs.85).aspx.  Microsoft Corp. Large-Page Support in Windows https:\/\/msdn.microsoft.com\/ en-us\/library\/windows\/desktop\/aa366720(v=vs.85).aspx."},{"key":"e_1_2_1_80_1","unstructured":"R. Mijat \"Take GPU Processing Power Beyond Graphics with Mali GPU Computing \" 2012.  R. Mijat \"Take GPU Processing Power Beyond Graphics with Mali GPU Computing \" 2012."},{"key":"e_1_2_1_81_1","unstructured":"MongoDB Inc. \"Disable Transparent Huge Pages (THP) \" https:\/\/docs.mongodb. org\/manual\/tutorial\/transparent-huge-pages\/ 2017.  MongoDB Inc. \"Disable Transparent Huge Pages (THP) \" https:\/\/docs.mongodb. org\/manual\/tutorial\/transparent-huge-pages\/ 2017."},{"key":"e_1_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155664"},{"key":"e_1_2_1_83_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.40"},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2008.7"},{"key":"e_1_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1145\/2155620.2155656"},{"volume-title":"Transparent Operating System Support for Superpages,\" in OSDI","year":"2002","author":"Navarro J.","key":"e_1_2_1_86_1"},{"volume-title":"Yak: A High-Performance Big-Data-Friendly Garbage Collector,\" in OSDI","year":"2016","author":"Nguyen K.","key":"e_1_2_1_87_1"},{"key":"e_1_2_1_88_1","unstructured":"NVIDIA Corp. \"CUDA C\/C++ SDK Code Samples \" http:\/\/developer.nvidia.com\/ cuda-cc-sdk-code-samples 2011.  NVIDIA Corp. \"CUDA C\/C++ SDK Code Samples \" http:\/\/developer.nvidia.com\/ cuda-cc-sdk-code-samples 2011."},{"volume-title":"Fermi,\" http:\/\/www.nvidia.com\/content\/pdf\/fermi_white_papers\/nvidia_fermi_ compute_architecture_whitepaper.pdf","year":"2011","author":"NVIDIA Corp.","key":"e_1_2_1_89_1"},{"volume-title":"Kepler GK110,\" http:\/\/www.nvidia.com\/content\/PDF\/kepler\/ NVIDIA-Kepler-GK110-Architecture-Whitepaper.pdf","year":"2012","author":"NVIDIA Corp.","key":"e_1_2_1_90_1"},{"key":"e_1_2_1_91_1","unstructured":"NVIDIA Corp. \"NVIDIA GeForce GTX 750 Ti \" http:\/\/ international.download.nvidia.com\/geforce-com\/international\/pdfs\/ GeForce-GTX-750-Ti-Whitepaper.pdf 2014.  NVIDIA Corp. \"NVIDIA GeForce GTX 750 Ti \" http:\/\/ international.download.nvidia.com\/geforce-com\/international\/pdfs\/ GeForce-GTX-750-Ti-Whitepaper.pdf 2014."},{"key":"e_1_2_1_92_1","unstructured":"NVIDIA Corp. \"CUDA C Programming Guide \" http:\/\/docs.nvidia.com\/cuda\/ cuda-c-programming-guide\/index.html 2015.  NVIDIA Corp. \"CUDA C Programming Guide \" http:\/\/docs.nvidia.com\/cuda\/ cuda-c-programming-guide\/index.html 2015."},{"key":"e_1_2_1_93_1","unstructured":"NVIDIA Corp. \"NVIDIA RISC-V Story \" https:\/\/riscv.org\/wp-content\/uploads\/ 2016\/07\/Tue1100_Nvidia_RISCV_Story_V2.pdf 2016.  NVIDIA Corp. \"NVIDIA RISC-V Story \" https:\/\/riscv.org\/wp-content\/uploads\/ 2016\/07\/Tue1100_Nvidia_RISCV_Story_V2.pdf 2016."},{"key":"e_1_2_1_94_1","unstructured":"NVIDIA Corp. \"NVIDIA Tesla P100 \" https:\/\/images.nvidia.com\/content\/pdf\/ tesla\/whitepaper\/pascal-architecture-whitepaper.pdf 2016.  NVIDIA Corp. \"NVIDIA Tesla P100 \" https:\/\/images.nvidia.com\/content\/pdf\/ tesla\/whitepaper\/pascal-architecture-whitepaper.pdf 2016."},{"volume-title":"download.nvidia.com\/geforce-com\/international\/pdfs\/GeForce_GTX_ 1080_Whitepaper_FINAL.pdf","year":"2017","author":"NVIDIA Corp.","key":"e_1_2_1_95_1"},{"key":"e_1_2_1_96_1","unstructured":"NVIDIA Corporation \"NVIDIA Tegra K1 \" http:\/\/www.nvidia.com\/content\/pdf\/ tegra_white_papers\/tegra-k1-whitepaper-v1.0.pdf.  NVIDIA Corporation \"NVIDIA Tegra K1 \" http:\/\/www.nvidia.com\/content\/pdf\/ tegra_white_papers\/tegra-k1-whitepaper-v1.0.pdf."},{"key":"e_1_2_1_97_1","unstructured":"NVIDIA Corporation \"NVIDIA\u01cdo Tegra\u01cdo X1 \" https:\/\/international.download. nvidia.com\/pdf\/tegra\/Tegra-X1-whitepaper-v1.0.pdf.  NVIDIA Corporation \"NVIDIA\u01cdo Tegra\u01cdo X1 \" https:\/\/international.download. nvidia.com\/pdf\/tegra\/Tegra-X1-whitepaper-v1.0.pdf."},{"key":"e_1_2_1_98_1","unstructured":"NVIDIA Corporation \"Multi-Process Service \" https:\/\/docs.nvidia.com\/deploy\/ pdf\/CUDA_Multi_Process_Service_Overview.pdf 2015.  NVIDIA Corporation \"Multi-Process Service \" https:\/\/docs.nvidia.com\/deploy\/ pdf\/CUDA_Multi_Process_Service_Overview.pdf 2015."},{"key":"e_1_2_1_99_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451116.2451160"},{"volume-title":"Prediction-Based Superpage-Friendly TLB Designs,\" in HPCA","year":"2015","author":"Papadopoulou M.-M.","key":"e_1_2_1_100_1"},{"key":"e_1_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1145\/2694344.2694346"},{"key":"e_1_2_1_102_1","unstructured":"PCI-SIG \"PCI Express Base Specification Revision 3.1a \" 2015.  PCI-SIG \"PCI Express Base Specification Revision 3.1a \" 2015."},{"volume-title":"A Case for Toggle-aware Compression for GPU Systems,\" in HPCA","year":"2016","author":"Pekhimenko G.","key":"e_1_2_1_103_1"},{"volume-title":"Percona LLC","year":"2014","author":"Zaitsev Peter","key":"e_1_2_1_104_1"},{"volume-title":"Increasing TLB Reach by Exploiting Clustering in Page Translations,\" in HPCA","year":"2014","author":"Pham B.","key":"e_1_2_1_105_1"},{"key":"e_1_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.32"},{"key":"e_1_2_1_107_1","doi-asserted-by":"publisher","DOI":"10.1145\/2830772.2830773"},{"key":"e_1_2_1_108_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541942"},{"volume-title":"Supporting x86-64 Address Translation for 100s of GPU Lanes,\" in HPCA","year":"2014","author":"Power J.","key":"e_1_2_1_109_1"},{"key":"e_1_2_1_110_1","unstructured":"PowerVR \"PowerVR Hardware Architecture Overview for Developers \" 2016 http:\/\/cdn.imgtec.com\/sdk-documentation\/PowerVR+Hardware. Architecture+Overview+for+Developers.pdf.  PowerVR \"PowerVR Hardware Architecture Overview for Developers \" 2016 http:\/\/cdn.imgtec.com\/sdk-documentation\/PowerVR+Hardware. Architecture+Overview+for+Developers.pdf."},{"key":"e_1_2_1_111_1","unstructured":"Redis Labs \"Redis Latency Problems Troubleshooting \" http:\/\/redis.io\/topics\/ latency.  Redis Labs \"Redis Latency Problems Troubleshooting \" http:\/\/redis.io\/topics\/ latency."},{"key":"e_1_2_1_112_1","doi-asserted-by":"publisher","DOI":"10.1145\/339647.339668"},{"volume-title":"dissertation","year":"2015","author":"Rogers T. G.","key":"e_1_2_1_113_1"},{"key":"e_1_2_1_114_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.16"},{"key":"e_1_2_1_115_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522715"},{"key":"e_1_2_1_116_1","unstructured":"SAFARI Research Group \"Mosaic - GitHub Repository \" https:\/\/github.com\/ CMU-SAFARI\/Mosaic\/.  SAFARI Research Group \"Mosaic - GitHub Repository \" https:\/\/github.com\/ CMU-SAFARI\/Mosaic\/."},{"key":"e_1_2_1_117_1","unstructured":"SAFARI Research Group \"SAFARI Software Tools - GitHub Repository \" https: \/\/github.com\/CMU-SAFARI\/.  SAFARI Research Group \"SAFARI Software Tools - GitHub Repository \" https: \/\/github.com\/CMU-SAFARI\/."},{"key":"e_1_2_1_118_1","doi-asserted-by":"publisher","DOI":"10.1145\/339647.339666"},{"key":"e_1_2_1_119_1","doi-asserted-by":"crossref","DOI":"10.1145\/2540708.2540725","volume-title":"RowClone: Fast and Energy-Efficient In-DRAM Bulk Data Copy and Initialization,\" in ISCA","author":"Seshadri V.","year":"2013"},{"volume-title":"Simple Operations in Memory to Reduce Data Movement,\" in Advances in Computers","year":"2017","author":"Seshadri V.","key":"e_1_2_1_120_1"},{"key":"e_1_2_1_121_1","volume-title":"Pentium Pro Processor System Architecture","author":"Shanley T.","year":"1996","edition":"1"},{"volume-title":"ALPHA Architecture Reference Manual","year":"1998","author":"Sites R. L.","key":"e_1_2_1_122_1"},{"key":"e_1_2_1_123_1","doi-asserted-by":"crossref","unstructured":"B. Smith \"Architecture and Applications of the HEP Multiprocessor Computer System \" SPIE 1981  B. Smith \"Architecture and Applications of the HEP Multiprocessor Computer System \" SPIE 1981","DOI":"10.1117\/12.932535"},{"volume-title":"Shared Resource MIMD Computer,\" in ICPP","year":"1978","author":"Smith B. J.","key":"e_1_2_1_124_1"},{"key":"e_1_2_1_125_1","unstructured":"Splunk Inc. \"Transparent Huge Memory Pages and Splunk Performance \" http:\/\/docs.splunk.com\/Documentation\/Splunk\/6.1.3\/ReleaseNotes\/ SplunkandTHP 2013.  Splunk Inc. \"Transparent Huge Memory Pages and Splunk Performance \" http:\/\/docs.splunk.com\/Documentation\/Splunk\/6.1.3\/ReleaseNotes\/ SplunkandTHP 2013."},{"key":"e_1_2_1_126_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2010.26"},{"key":"e_1_2_1_128_1","doi-asserted-by":"publisher","DOI":"10.1145\/2584665"},{"key":"e_1_2_1_129_1","doi-asserted-by":"publisher","DOI":"10.1145\/195473.195531"},{"key":"e_1_2_1_130_1","doi-asserted-by":"publisher","DOI":"10.1145\/1464039.1464045"},{"volume-title":"Scott Foresman & Co","year":"1970","author":"Thornton J. E.","key":"e_1_2_1_131_1"},{"key":"e_1_2_1_132_1","doi-asserted-by":"publisher","DOI":"10.1145\/2847255"},{"volume-title":"Observations and Opportunities in Architecting Shared Virtual Memory for Heterogeneous Systems,\" in ISPASS","year":"2016","author":"Vesely J.","key":"e_1_2_1_133_1"},{"volume-title":"Zorua: A Holistic Approach to Resource Virtualization in GPUs,\" in MICRO","year":"2016","author":"Vijaykumar N.","key":"e_1_2_1_134_1"},{"key":"e_1_2_1_135_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750399"},{"volume-title":"Documentation: Configure Memory Management,\" https: \/\/docs.voltdb.com\/AdminGuide\/adminmemmgt.php.","key":"e_1_2_1_136_1"},{"volume-title":"GPU Virtualization for High Performance General Purpose Computing on the ESX Hypervisor,\" in HPC","year":"2014","author":"Vu L.","key":"e_1_2_1_137_1"},{"key":"e_1_2_1_138_1","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2015.2477405"},{"key":"e_1_2_1_139_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080203"},{"key":"e_1_2_1_140_1","unstructured":"S. Wasson. (2011 Oct.) AMD's A8-3800 Fusion APU. {Online}. Available: http:\/\/techreport.com\/articles.x\/21730  S. Wasson. (2011 Oct.) AMD's A8-3800 Fusion APU. {Online}. Available: http:\/\/techreport.com\/articles.x\/21730"},{"key":"e_1_2_1_141_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.29"},{"key":"e_1_2_1_142_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173162.3173195"},{"key":"e_1_2_1_143_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669119"},{"volume-title":"Towards High Performance Paged Memory for GPUs,\" in HPCA","year":"2016","author":"Zheng T.","key":"e_1_2_1_144_1"},{"volume-title":"Controller for a Synchronous DRAM That Maximizes Throughput by Allowing Memory Requests and Commands to Be Issued Out of Order,\" US Patent No. 5,630,096","year":"1997","author":"Zuravleff W. K.","key":"e_1_2_1_145_1"}],"container-title":["ACM SIGOPS Operating Systems Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3273982.3273986","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3273982.3273986","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:44:44Z","timestamp":1750207484000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3273982.3273986"}},"subtitle":["Enabling Application-Transparent Support for Multiple Page Sizes in Throughput Processors"],"short-title":[],"issued":{"date-parts":[[2018,8,28]]},"references-count":143,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,8,28]]}},"alternative-id":["10.1145\/3273982.3273986"],"URL":"https:\/\/doi.org\/10.1145\/3273982.3273986","relation":{},"ISSN":["0163-5980"],"issn-type":[{"type":"print","value":"0163-5980"}],"subject":[],"published":{"date-parts":[[2018,8,28]]},"assertion":[{"value":"2018-08-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}