{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,31]],"date-time":"2026-01-31T10:32:08Z","timestamp":1769855528683,"version":"3.49.0"},"publisher-location":"New York, NY, USA","reference-count":92,"publisher":"ACM","license":[{"start":{"date-parts":[[2018,11,1]],"date-time":"2018-11-01T00:00:00Z","timestamp":1541030400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2018,11]]},"DOI":"10.1145\/3243176.3243188","type":"proceedings-article","created":{"date-parts":[[2018,10,10]],"date-time":"2018-10-10T13:32:32Z","timestamp":1539178352000},"page":"1-14","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":25,"title":["In-DRAM near-data approximate acceleration for GPUs"],"prefix":"10.1145","author":[{"given":"Amir","family":"Yazdanbakhsh","sequence":"first","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Choungki","family":"Song","sequence":"additional","affiliation":[{"name":"University of Wisconsin-Madison"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jacob","family":"Sacks","sequence":"additional","affiliation":[{"name":"Georgia Institute of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pejman","family":"Lotfi-Kamran","sequence":"additional","affiliation":[{"name":"Institute for Research in Fundamental Sciences (IPM)"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hadi","family":"Esmaeilzadeh","sequence":"additional","affiliation":[{"name":"University of California"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nam Sung","family":"Kim","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana-Champaign"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,11]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2015. NanGate FreePDK45 Open Cell Library. http:\/\/www.nangate.com. (2015). http:\/\/www.nangate.com\/?page_id=2325  2015. NanGate FreePDK45 Open Cell Library. http:\/\/www.nangate.com. (2015). http:\/\/www.nangate.com\/?page_id=2325"},{"key":"e_1_3_2_1_2_1","unstructured":"2015. NVIDIA Corporation. CUDA Programming Guide. http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide. (2015). http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide  2015. NVIDIA Corporation. CUDA Programming Guide. http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide. (2015). http:\/\/docs.nvidia.com\/cuda\/cuda-c-programming-guide"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750386"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750397"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Ren\u00e9e St Amant Amir Yazdanbakhsh Jongse Park Bradley Thwaites Hadi Esmaeilzadeh Arjang Hassibi Luis Ceze and Doug Burger. 2014. General-Purpose Code Acceleration with Limited-Precision Analog Computation. In ISCA.   Ren\u00e9e St Amant Amir Yazdanbakhsh Jongse Park Bradley Thwaites Hadi Esmaeilzadeh Arjang Hassibi Luis Ceze and Doug Burger. 2014. General-Purpose Code Acceleration with Limited-Precision Analog Computation. In ISCA .","DOI":"10.1109\/ISCA.2014.6853213"},{"key":"e_1_3_2_1_6_1","volume-title":"Jung Ho Ahn, and Nam Sung Kim.","author":"Asghari-Moghaddam Hadi","year":"2016","unstructured":"Hadi Asghari-Moghaddam , Young Hoon Son , Jung Ho Ahn, and Nam Sung Kim. 2016 . Chameleon : Versatile and Practical Near-DRAM Acceleration Architecture for Large Memory Systems. In MICRO. Hadi Asghari-Moghaddam, Young Hoon Son, Jung Ho Ahn, and Nam Sung Kim. 2016. Chameleon: Versatile and Practical Near-DRAM Acceleration Architecture for Large Memory Systems. In MICRO."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","unstructured":"A Bakhoda G.L. Yuan W.W.L. Fung H. Wong and T.M. Aamodt. 2009. Analyzing CUDA Workloads using a Detailed GPU Simulator. In ISPASS.  A Bakhoda G.L. Yuan W.W.L. Fung H. Wong and T.M. Aamodt. 2009. Analyzing CUDA Workloads using a Detailed GPU Simulator. In ISPASS .","DOI":"10.1109\/ISPASS.2009.4919648"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508148.2485923"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.12"},{"key":"e_1_3_2_1_10_1","volume-title":"Memory System on Fusion APUs. AMD Fusion developer summit","author":"Boudier Pierre","year":"2011","unstructured":"Pierre Boudier and Graham Sellers . 2011. Memory System on Fusion APUs. AMD Fusion developer summit ( 2011 ). Pierre Boudier and Graham Sellers. 2011. Memory System on Fusion APUs. AMD Fusion developer summit (2011)."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/2509136.2509546"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"crossref","unstructured":"K. K. Chang P. J. Nair D. Lee S. Ghose M. K. Qureshi and O. Mutlu. 2016. Low-Cost Inter-Linked Subarrays (LISA): Enabling fast inter-subarray data movement in DRAM. In HPCA.  K. K. Chang P. J. Nair D. Lee S. Ghose M. K. Qureshi and O. Mutlu. 2016. Low-Cost Inter-Linked Subarrays (LISA): Enabling fast inter-subarray data movement in DRAM. In HPCA .","DOI":"10.1109\/HPCA.2016.7446095"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2014.16"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2014.11"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.40"},{"key":"e_1_3_2_1_16_1","unstructured":"Radoslav Danilak. 2009. System and Method for Hardware-based GPU Pagingto System Memory. US7623134 B1.  Radoslav Danilak. 2009. System and Method for Hardware-based GPU Pagingto System Memory. US7623134 B1."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/192161.192194"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/514191.514197"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Zidong Du Avinash Lingamneni Yunji Chen Krishna Palem Olivier Temam and Chengyong Wu. 2014. Leveraging the Error Resilience of Machine-Learning Applications for Designing Highly Energy Efficient Accelerators. In ASP-DAC.  Zidong Du Avinash Lingamneni Yunji Chen Krishna Palem Olivier Temam and Chengyong Wu. 2014. Leveraging the Error Resilience of Machine-Learning Applications for Designing Highly Energy Efficient Accelerators. In ASP-DAC .","DOI":"10.1109\/ASPDAC.2014.6742890"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2015.21"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CICC.1992.591879"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.48"},{"key":"e_1_3_2_1_23_1","volume-title":"K. Morrow, and Nam Sung Kim.","author":"Farmahini-Farahani A.","year":"2015","unstructured":"A. Farmahini-Farahani , Jung Ho Ahn , K. Morrow, and Nam Sung Kim. 2015 . DRAMA : An Architecture for Accelerated Processing Near Memory. CAL 14, 1 (2015). A. Farmahini-Farahani, Jung Ho Ahn, K. Morrow, and Nam Sung Kim. 2015. DRAMA: An Architecture for Accelerated Processing Near Memory. CAL 14, 1 (2015)."},{"key":"e_1_3_2_1_24_1","volume-title":"K. Morrow, and Nam Sung Kim.","author":"Farmahini-Farahani A.","year":"2015","unstructured":"A. Farmahini-Farahani , Jung Ho Ahn , K. Morrow, and Nam Sung Kim. 2015 . NDA : Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules. In HPCA. A. Farmahini-Farahani, Jung Ho Ahn, K. Morrow, and Nam Sung Kim. 2015. NDA: Near-DRAM Acceleration Architecture Leveraging Commodity DRAM Devices and Standard Memory Modules. In HPCA."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Yusuke Fujii Takuya Azumi Nobuhiko Nishio Shinpei Kato and Masato Edahiro. 2013. Data Transfer Matters for GPU Computing. In ICPADS.   Yusuke Fujii Takuya Azumi Nobuhiko Nishio Shinpei Kato and Masato Edahiro. 2013. Data Transfer Matters for GPU Computing. In ICPADS .","DOI":"10.1109\/ICPADS.2013.47"},{"key":"e_1_3_2_1_26_1","volume-title":"HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing. In HPCA.","author":"Gao M.","year":"2016","unstructured":"M. Gao and Ch. Kozyrakis . 2016 . HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing. In HPCA. M. Gao and Ch. Kozyrakis. 2016. HRL: Efficient and Flexible Reconfigurable Logic for Near-Data Processing. In HPCA."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037702"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2012.51"},{"key":"e_1_3_2_1_29_1","volume-title":"BRAINIAC: Bringing Reliable Accuracy Into Neurally-Implemented Approximate Computing. In HPCA.","author":"Grigorian Beayna","year":"2015","unstructured":"Beayna Grigorian , Nazanin Farahpour , and Glenn Reinman . 2015 . BRAINIAC: Bringing Reliable Accuracy Into Neurally-Implemented Approximate Computing. In HPCA. Beayna Grigorian, Nazanin Farahpour, and Glenn Reinman. 2015. BRAINIAC: Bringing Reliable Accuracy Into Neurally-Implemented Approximate Computing. In HPCA."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Beayna Grigorian and Glenn Reinman. 2014. Accelerating Divergent Applications on SIMD Architectures using Neural Networks. In ICCD.  Beayna Grigorian and Glenn Reinman. 2014. Accelerating Divergent Applications on SIMD Architectures using Neural Networks. In ICCD .","DOI":"10.1109\/ICCD.2014.6974700"},{"key":"e_1_3_2_1_31_1","unstructured":"Q. Guo N. Alachiotis B. Akin F. Sadi G. Xu T-M. Low L. Pileggi J. Hoe and F. Franchetti. 2014. 3D-Stacked Memory-Side Acceleration: Accelerator and System Design. In WoNDP.  Q. Guo N. Alachiotis B. Akin F. Sadi G. Xu T-M. Low L. Pileggi J. Hoe and F. Franchetti. 2014. 3D-Stacked Memory-Side Acceleration: Accelerator and System Design. In WoNDP ."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485939"},{"key":"e_1_3_2_1_33_1","volume-title":"Dally","author":"Han Song","year":"2016","unstructured":"Song Han , Huizi Mao , and William J . Dally . 2016 . Deep Compression : Compressing Deep Neural Networks with Pruning, Trained Quantization, and Huffman Coding. In ICLR. Song Han, Huizi Mao, and William J. Dally. 2016. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization, and Huffman Coding. In ICLR."},{"key":"e_1_3_2_1_34_1","volume-title":"Inside Pascal: Nvidia's Newest Computing Platform. https:\/\/devblogs.nvidia.com\/parallelforall\/inside-pascal\/.","author":"Harris Mark","year":"2016","unstructured":"Mark Harris . 2016 . Inside Pascal: Nvidia's Newest Computing Platform. https:\/\/devblogs.nvidia.com\/parallelforall\/inside-pascal\/. (2016). https:\/\/devblogs.nvidia.com\/parallelforall\/inside-pascal\/ Mark Harris. 2016. Inside Pascal: Nvidia's Newest Computing Platform. https:\/\/devblogs.nvidia.com\/parallelforall\/inside-pascal\/. (2016). https:\/\/devblogs.nvidia.com\/parallelforall\/inside-pascal\/"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2818950.2818952"},{"key":"e_1_3_2_1_36_1","unstructured":"Mark Horowitz. {n. d.}. Energy Table for 45nm Process. ({n. d.}).  Mark Horowitz. {n. d.}. Energy Table for 45nm Process. ({n. d.})."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"crossref","unstructured":"Rui Hou Lixin Zhang Michael C Huang Kun Wang Hubertus Franke Yi Ge and Xiaotao Chang. 2011. Efficient Data Streaming with On-chip Accelerators: Opportunities and Challenges. In HPCA.   Rui Hou Lixin Zhang Michael C Huang Kun Wang Hubertus Franke Yi Ge and Xiaotao Chang. 2011. Efficient Data Streaming with On-chip Accelerators: Opportunities and Challenges. In HPCA .","DOI":"10.1109\/HPCA.2011.5749739"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.27"},{"key":"e_1_3_2_1_39_1","volume-title":"Hynix GDDR5 SGRAM Part H5GQ1H24AFR Revision 1.0","year":"2015","unstructured":"Hynix. Hynix GDDR5 SGRAM Part H5GQ1H24AFR Revision 1.0 . 2015 . (2015). Hynix. Hynix GDDR5 SGRAM Part H5GQ1H24AFR Revision 1.0. 2015. (2015)."},{"key":"e_1_3_2_1_40_1","volume-title":"CMOS Processors and Memories","author":"Iniewski Krzysztof","unstructured":"Krzysztof Iniewski . 2010. CMOS Processors and Memories . Springer Science & Business Media . Krzysztof Iniewski. 2010. CMOS Processors and Memories. Springer Science & Business Media."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1535\/itj.1002.04"},{"key":"e_1_3_2_1_42_1","unstructured":"JEDEC. October 2013. High Bandwidth Memory DRAM. http:\/\/www.jedec.org\/standards-documents\/docs\/jesd235. (October 2013).  JEDEC. October 2013. High Bandwidth Memory DRAM. http:\/\/www.jedec.org\/standards-documents\/docs\/jesd235. (October 2013)."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2451116.2451118"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"crossref","unstructured":"Yi Kang Wei Huang Seung-Moon Yoo Diana Keen Zhenzhou Ge Vinh Lam Pratap Pattnaik and Josep Torrellas. 2012. FlexRAM: Toward an Advanced Intelligent Memory System. In ICCD.  Yi Kang Wei Huang Seung-Moon Yoo Diana Keen Zhenzhou Ge Vinh Lam Pratap Pattnaik and Josep Torrellas. 2012. FlexRAM: Toward an Advanced Intelligent Memory System. In ICCD .","DOI":"10.1109\/ICCD.2012.6378608"},{"key":"e_1_3_2_1_45_1","volume-title":"Gdev: First-Class GPU Resource Management in the Operating System. In USENIX.","author":"Kato Shinpei","year":"2012","unstructured":"Shinpei Kato , Michael McThrow , Carlos Maltzahn , and Scott Brandt . 2012 . Gdev: First-Class GPU Resource Management in the Operating System. In USENIX. Shinpei Kato, Michael McThrow, Carlos Maltzahn, and Scott Brandt. 2012. Gdev: First-Class GPU Resource Management in the Operating System. In USENIX."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2011.89"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.41"},{"key":"e_1_3_2_1_48_1","unstructured":"K.W. Kim. 2004. Apparatus for Pipe Latch Control Circuit in Synchronous Memory Device. (2004). https:\/\/www.google.com\/patents\/US6724684US6724684B2.  K.W. Kim. 2004. Apparatus for Pipe Latch Control Circuit in Synchronous Memory Device. (2004). https:\/\/www.google.com\/patents\/US6724684US6724684B2."},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"crossref","unstructured":"K. Koo S. Ok Y. Kang S. Kim C. Song H. Lee H. Kim Y. Kim J. Lee S. Oak Y. Lee J. Lee J. Lee H. Lee J. Jang J. Jung B. Choi Y. Kim Y. Hur Y. Kim B. Chung and Y. Kim. 2012. A 1.2V 38nm 2.4Gb\/s\/pin 2Gb DDR4 SDRAM with Bank Group and x4 Half-Page Architecture. In ISSCC. 40--41.  K. Koo S. Ok Y. Kang S. Kim C. Song H. Lee H. Kim Y. Kim J. Lee S. Oak Y. Lee J. Lee J. Lee H. Lee J. Jang J. Jung B. Choi Y. Kim Y. Hur Y. Kim B. Chung and Y. Kim. 2012. A 1.2V 38nm 2.4Gb\/s\/pin 2Gb DDR4 SDRAM with Bank Group and x4 Half-Page Architecture. In ISSCC . 40--41.","DOI":"10.1109\/ISSCC.2012.6176869"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"crossref","unstructured":"D. U. Lee K. W. Kim K. W. Kim H. Kim J. Y. Kim Y. J. Park J. H. Kim D. S. Kim H. B. Park J. W. Shin J. H. Cho K. H. Kwon M. J. Kim J. Lee K. W. Park B. Chung and S. Hong. 2014. 25.2 A 1.2V 8Gb 8-Channel 128GB\/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I\/O Test Methods using 29nm Process and TSV. In ISSCC.  D. U. Lee K. W. Kim K. W. Kim H. Kim J. Y. Kim Y. J. Park J. H. Kim D. S. Kim H. B. Park J. W. Shin J. H. Cho K. H. Kwon M. J. Kim J. Lee K. W. Park B. Chung and S. Hong. 2014. 25.2 A 1.2V 8Gb 8-Channel 128GB\/s High-Bandwidth Memory (HBM) Stacked DRAM with Effective Microbump I\/O Test Methods using 29nm Process and TSV. In ISSCC .","DOI":"10.1109\/ISSCC.2014.6757501"},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485964"},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/1669112.1669172"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/1250662.1250701"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/1384529.1375496"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2012.118"},{"key":"e_1_3_2_1_56_1","volume-title":"Gokhan Memik, and Nikos Hardavellas.","author":"Liu Song","year":"2011","unstructured":"Song Liu , Brian Leung , Alexander Neckar , Seda Ogrenci Memik , Gokhan Memik, and Nikos Hardavellas. 2011 . Hardware\/Software Techniques for DRAM Thermal Management. In HPCA. Song Liu, Brian Leung, Alexander Neckar, Seda Ogrenci Memik, Gokhan Memik, and Nikos Hardavellas. 2011. Hardware\/Software Techniques for DRAM Thermal Management. In HPCA."},{"key":"e_1_3_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/339647.339673"},{"key":"e_1_3_2_1_58_1","unstructured":"K Man. {n. d.}. Bensley FB-DIMM Performance\/Thermal Management. In Intel Developer Forum.  K Man. {n. d.}. Bensley FB-DIMM Performance\/Thermal Management. In Intel Developer Forum ."},{"key":"e_1_3_2_1_59_1","volume-title":"EMEURO: A Framework for Generating Multi-Purpose Accelerators via Deep Learning. In CGO.","author":"McAfee Lawrence","year":"2015","unstructured":"Lawrence McAfee and Kunle Olukotun . 2015 . EMEURO: A Framework for Generating Multi-Purpose Accelerators via Deep Learning. In CGO. Lawrence McAfee and Kunle Olukotun. 2015. EMEURO: A Framework for Generating Multi-Purpose Accelerators via Deep Learning. In CGO."},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2016.2549523"},{"key":"e_1_3_2_1_61_1","volume-title":"SNNAP: Approximate Computing on Programmable SoCs via Neural Acceleration. In HPCA.","author":"Moreau Thierry","year":"2015","unstructured":"Thierry Moreau , Mark Wyse , Jacob Nelson , Adrian Sampson , Hadi Esmaeilzadeh , Luis Ceze , and Mark Oskin . 2015 . SNNAP: Approximate Computing on Programmable SoCs via Neural Acceleration. In HPCA. Thierry Moreau, Mark Wyse, Jacob Nelson, Adrian Sampson, Hadi Esmaeilzadeh, Luis Ceze, and Mark Oskin. 2015. SNNAP: Approximate Computing on Programmable SoCs via Neural Acceleration. In HPCA."},{"key":"e_1_3_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/2485922.2485927"},{"key":"e_1_3_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2007.30"},{"key":"e_1_3_2_1_64_1","unstructured":"Lifeng Nai Ramyad Hadidi Jaewoong Sim Hyojong Kim Pranith Kumar and Hyesoon Kim. 2017. GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks. In HPCA.  Lifeng Nai Ramyad Hadidi Jaewoong Sim Hyojong Kim Pranith Kumar and Hyesoon Kim. 2017. GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks. In HPCA ."},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1147\/JRD.2015.2409732"},{"key":"e_1_3_2_1_66_1","first-page":"107","article-title":"A 7 Gb\/s\/pin 1 Gbit GDDR5 SDRAM With 2.5 ns Bank to Bank Active Time and No Bank Group Restriction","volume":"46","author":"Oh T. Y.","year":"2011","unstructured":"T. Y. Oh , Y. S. Sohn , S. J. Bae , M. S. Park , J. H. Lim , Y. K. Cho , D. H. Kim , D. M. Kim , H. R. Kim , H. J. Kim , J. H. Kim , J. K. Kim , Y. S. Kim , B. C. Kim , S. H. Kwak , J. H. Lee , J. Y. Lee , C. H. Shin , Y. Yang , B. S. Cho , S. Y. Bang , H. J. Yang , Y. R. Choi , G. S. Moon , C. G. Park , S. W. Hwang , J. D. Lim , K. I. Park , J. S. Choi , and Y. H. Jun . 2011 . A 7 Gb\/s\/pin 1 Gbit GDDR5 SDRAM With 2.5 ns Bank to Bank Active Time and No Bank Group Restriction . JSSC 46 , 1 (2011), 107 -- 118 . T. Y. Oh, Y. S. Sohn, S. J. Bae, M. S. Park, J. H. Lim, Y. K. Cho, D. H. Kim, D. M. Kim, H. R. Kim, H. J. Kim, J. H. Kim, J. K. Kim, Y. S. Kim, B. C. Kim, S. H. Kwak, J. H. Lee, J. Y. Lee, C. H. Shin, Y. Yang, B. S. Cho, S. Y. Bang, H. J. Yang, Y. R. Choi, G. S. Moon, C. G. Park, S. W. Hwang, J. D. Lim, K. I. Park, J. S. Choi, and Y. H. Jun. 2011. A 7 Gb\/s\/pin 1 Gbit GDDR5 SDRAM With 2.5 ns Bank to Bank Active Time and No Bank Group Restriction. JSSC 46, 1 (2011), 107--118.","journal-title":"JSSC"},{"key":"e_1_3_2_1_67_1","volume-title":"ISSCC'10","author":"Oh T. Y.","unstructured":"T. Y. Oh , Y. S. Sohn , S. J. Bae , M. S. Park , J. H. Lim , Y. K. Cho , D. H. Kim , D. M. Kim , H. R. Kim , H. J. Kim , J. H. Kim , J. K. Kim , Y. S. Kim , B. C. Kim , S. H. Kwak , J. H. Lee , J. Y. Lee , C. H. Shin , Y. S. Yang , B. S. Cho , S. Y. Bang , H. J. Yang , Y. R. Choi , G. S. Moon , C. G. Park , S. W. Hwang , J. D. Lim , K. I. Park , J. S. Choi , and Y. H. Jun . {n. d.}. A 7Gb\/s\/pin GDDR5 SDRAM with 2.5ns Bank-to-Bank Active Time and no Bank-group Restriction . In ISSCC'10 . T. Y. Oh, Y. S. Sohn, S. J. Bae, M. S. Park, J. H. Lim, Y. K. Cho, D. H. Kim, D. M. Kim, H. R. Kim, H. J. Kim, J. H. Kim, J. K. Kim, Y. S. Kim, B. C. Kim, S. H. Kwak, J. H. Lee, J. Y. Lee, C. H. Shin, Y. S. Yang, B. S. Cho, S. Y. Bang, H. J. Yang, Y. R. Choi, G. S. Moon, C. G. Park, S. W. Hwang, J. D. Lim, K. I. Park, J. S. Choi, and Y. H. Jun. {n. d.}. A 7Gb\/s\/pin GDDR5 SDRAM with 2.5ns Bank-to-Bank Active Time and no Bank-group Restriction. In ISSCC'10."},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/279358.279387"},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/2786805.2786807"},{"key":"e_1_3_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1109\/40.592312"},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/2654822.2541942"},{"key":"e_1_3_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/2541940.2541942"},{"key":"e_1_3_2_1_73_1","doi-asserted-by":"crossref","unstructured":"Jason Power Mark D Hill and David A Wood. 2014. Supporting x86-64 Address Translation for 100s of GPU Lanes. In HPCA.  Jason Power Mark D Hill and David A Wood. 2014. Supporting x86-64 Address Translation for 100s of GPU Lanes. In HPCA .","DOI":"10.1109\/HPCA.2014.6835965"},{"key":"e_1_3_2_1_74_1","volume-title":"Comparing Implementations of Near-Data Computing with In-Memory MapReduce Workloads. Micro","author":"Pugsley S.H.","year":"2014","unstructured":"S.H. Pugsley , J. Jestes , R. Balasubramonian , V. Srinivasan , A. Buyuktosunoglu , A. Davis , and Feifei Li. 2014. Comparing Implementations of Near-Data Computing with In-Memory MapReduce Workloads. Micro , IEEE 34, 4 ( 2014 ). S.H. Pugsley, J. Jestes, R. Balasubramonian, V. Srinivasan, A. Buyuktosunoglu, A. Davis, and Feifei Li. 2014. Comparing Implementations of Near-Data Computing with In-Memory MapReduce Workloads. Micro, IEEE 34, 4 (2014)."},{"key":"e_1_3_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555754.1555801"},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.16"},{"key":"e_1_3_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1145\/1993498.1993518"},{"key":"e_1_3_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2013.6522329"},{"key":"e_1_3_2_1_79_1","unstructured":"C. Shelor K. Kavi and Adavally S. 2015. Dataflow based Near Data Processing using Coarse Grain Reconfigurable Logic. In WoNDP.  C. Shelor K. Kavi and Adavally S. 2015. Dataflow based Near Data Processing using Coarse Grain Reconfigurable Logic. In WoNDP ."},{"key":"e_1_3_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2013.6522351"},{"key":"e_1_3_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508148.2485955"},{"key":"e_1_3_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1145\/2716282.2716283"},{"key":"e_1_3_2_1_83_1","volume-title":"US20080028181 A1.","author":"Tong Peter C","year":"2008","unstructured":"Peter C Tong , Sonny S Yeoh , Kevin J Kranzusch , Gary D Lorensen , Kaymann L Woo , Ashish Kishen Kaul , Colyn S Case , Stefan A Gottschalk , and Dennis K Ma . 2008 . Dedicated Mechanism for Page Mapping in a GPU . US20080028181 A1. Peter C Tong, Sonny S Yeoh, Kevin J Kranzusch, Gary D Lorensen, Kaymann L Woo, Ashish Kishen Kaul, Colyn S Case, Stefan A Gottschalk, and Dennis K Ma. 2008. Dedicated Mechanism for Page Mapping in a GPU. US20080028181 A1."},{"key":"e_1_3_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750399"},{"key":"e_1_3_2_1_85_1","unstructured":"Nicholas Wilt. 2013. The CUDA Handbook: A Comprehensive Guide to GPU Programming. Pearson Education.  Nicholas Wilt. 2013. The CUDA Handbook: A Comprehensive Guide to GPU Programming . Pearson Education."},{"key":"e_1_3_2_1_86_1","doi-asserted-by":"crossref","unstructured":"Henry Wong Misel-Myrto Papadopoulou Maryam Sadooghi-Alvandi and Andreas Moshovos. 2010. Demystifying GPU Microarchitecture through Microbenchmarking. In ISPASS.  Henry Wong Misel-Myrto Papadopoulou Maryam Sadooghi-Alvandi and Andreas Moshovos. 2010. Demystifying GPU Microarchitecture through Microbenchmarking. In ISPASS .","DOI":"10.1109\/ISPASS.2010.5452013"},{"key":"e_1_3_2_1_87_1","volume-title":"AxBench: A Multi-Platform Benchmark Suite for Approximate Computing: Acceleration for GPU Throughput Processors","author":"Yazdanbakhsh Amir","year":"2016","unstructured":"Amir Yazdanbakhsh , Divya Mahajan , Pejman Lotfi-Kamran , and Hadi Esmaeilzadeh . 2016. AxBench: A Multi-Platform Benchmark Suite for Approximate Computing: Acceleration for GPU Throughput Processors . IEEE Design and Test ( 2016 ). Amir Yazdanbakhsh, Divya Mahajan, Pejman Lotfi-Kamran, and Hadi Esmaeilzadeh. 2016. AxBench: A Multi-Platform Benchmark Suite for Approximate Computing: Acceleration for GPU Throughput Processors. IEEE Design and Test (2016)."},{"key":"e_1_3_2_1_88_1","volume-title":"Axilog: Language Support for Approximate Hardware Design. In DATE.","author":"Yazdanbakhsh Amir","year":"2015","unstructured":"Amir Yazdanbakhsh , Divya Mahajan , Bradley Thwaites , Jongse Park , Anandhavel Nagendrakumar , Sindhuja Sethuraman , Kartik Ramkrishnan , Nishanthi Ravindran , Rudra Jariwala , Abbas Rahimi , Hadi Esmaeilzadeh , and Kia Bazargan . 2015 . Axilog: Language Support for Approximate Hardware Design. In DATE. Amir Yazdanbakhsh, Divya Mahajan, Bradley Thwaites, Jongse Park, Anandhavel Nagendrakumar, Sindhuja Sethuraman, Kartik Ramkrishnan, Nishanthi Ravindran, Rudra Jariwala, Abbas Rahimi, Hadi Esmaeilzadeh, and Kia Bazargan. 2015. Axilog: Language Support for Approximate Hardware Design. In DATE."},{"key":"e_1_3_2_1_89_1","doi-asserted-by":"publisher","DOI":"10.1145\/2830772.2830810"},{"key":"e_1_3_2_1_90_1","doi-asserted-by":"publisher","DOI":"10.1145\/2836168"},{"key":"e_1_3_2_1_91_1","doi-asserted-by":"publisher","DOI":"10.1145\/2600212.2600213"},{"key":"e_1_3_2_1_92_1","unstructured":"Qiuling Zhu T. Graf H.E. Sumbul L. Pileggi and F. Franchetti. 2013. Accelerating Sparse Matrix-Matrix Multiplication with 3D-Stacked Logic-in-Memory Hardware. In HPEC.  Qiuling Zhu T. Graf H.E. Sumbul L. Pileggi and F. Franchetti. 2013. Accelerating Sparse Matrix-Matrix Multiplication with 3D-Stacked Logic-in-Memory Hardware. In HPEC ."}],"event":{"name":"PACT '18: International conference on Parallel Architectures and Compilation Techniques","location":"Limassol Cyprus","acronym":"PACT '18","sponsor":["SIGARCH ACM Special Interest Group on Computer Architecture","IFIP WG 10.3 IFIP WG 10.3","IEEE CS"]},"container-title":["Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3243176.3243188","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3243176.3243188","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:57:39Z","timestamp":1750208259000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3243176.3243188"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,11]]},"references-count":92,"alternative-id":["10.1145\/3243176.3243188","10.1145\/3243176"],"URL":"https:\/\/doi.org\/10.1145\/3243176.3243188","relation":{},"subject":[],"published":{"date-parts":[[2018,11]]},"assertion":[{"value":"2018-11-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}