{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,10]],"date-time":"2026-02-10T18:38:21Z","timestamp":1770748701812,"version":"3.50.0"},"reference-count":179,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2022,6,6]],"date-time":"2022-06-06T00:00:00Z","timestamp":1654473600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"H2020 research and innovation programme","award":["732631"],"award-info":[{"award-number":["732631"]}]},{"name":"European Commission under Marie Sklodowska-Curie Innovative Training Networks European Industrial Doctorate","award":["676240"],"award-info":[{"award-number":["676240"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2022,12,31]]},"abstract":"<jats:p>\n            Ongoing climate change calls for fast and accurate weather and climate modeling. However, when solving large-scale weather prediction simulations, state-of-the-art CPU and GPU implementations suffer from limited performance and high energy consumption. These implementations are dominated by complex irregular memory access patterns and low arithmetic intensity that pose fundamental challenges to acceleration. To overcome these challenges, we propose and evaluate the use of near-memory acceleration using a reconfigurable fabric with high-bandwidth memory (HBM). We focus on compound stencils that are fundamental kernels in weather prediction models. By using high-level synthesis techniques, we develop NERO, an field-programmable gate array+HBM-based accelerator connected through Open Coherent Accelerator Processor Interface to an IBM POWER9 host system. Our experimental results show that NERO outperforms a 16-core POWER9 system by\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\( 5.3\\times \\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            and\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\( 12.7\\times \\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            when running two different compound stencil kernels. NERO reduces the energy consumption by\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\( 12\\times \\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            and\n            <jats:inline-formula content-type=\"math\/tex\">\n              <jats:tex-math notation=\"LaTeX\" version=\"MathJax\">\\( 35\\times \\)<\/jats:tex-math>\n            <\/jats:inline-formula>\n            for the same two kernels over the POWER9 system with an energy efficiency of 1.61 GFLOPS\/W and 21.01 GFLOPS\/W. We conclude that employing near-memory acceleration solutions for weather prediction modeling \u00a0is\u00a0promising as a means to achieve both high performance and high energy efficiency.\n          <\/jats:p>","DOI":"10.1145\/3501804","type":"journal-article","created":{"date-parts":[[2022,2,9]],"date-time":"2022-02-09T19:55:30Z","timestamp":1644436530000},"page":"1-27","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["Accelerating Weather Prediction Using Near-Memory Reconfigurable Fabric"],"prefix":"10.1145","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3502-7401","authenticated-orcid":false,"given":"Gagandeep","family":"Singh","sequence":"first","affiliation":[{"name":"ETH Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2979-5946","authenticated-orcid":false,"given":"Dionysios","family":"Diamantopoulos","sequence":"additional","affiliation":[{"name":"IBM Research Europe, Z\u00fcrich Lab, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6514-1571","authenticated-orcid":false,"given":"Juan","family":"G\u00f3mez-Luna","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6815-7835","authenticated-orcid":false,"given":"Christoph","family":"Hagleitner","sequence":"additional","affiliation":[{"name":"IBM Research Europe, Z\u00fcrich Lab, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2518-6847","authenticated-orcid":false,"given":"Sander","family":"Stuijk","sequence":"additional","affiliation":[{"name":"Eindhoven Univesity of Technology, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4506-5732","authenticated-orcid":false,"given":"Henk","family":"Corporaal","sequence":"additional","affiliation":[{"name":"Eindhoven Univesity of Technology, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0075-2312","authenticated-orcid":false,"given":"Onur","family":"Mutlu","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,6,6]]},"reference":[{"key":"e_1_3_2_2_2","volume-title":"ADM-PCIE-9H7-High-Speed Communications Hub. Retrieved from","unstructured":"ADM-PCIE-9H7-High-Speed Communications Hub. Retrieved fromhttps:\/\/www.alpha-data.com\/dcp\/products.php?product=adm-pcie-9h7."},{"key":"e_1_3_2_3_2","volume-title":"ADM-PCIE-9V3-High-Performance Network Accelerator. Retrieved from","unstructured":"ADM-PCIE-9V3-High-Performance Network Accelerator. Retrieved fromhttps:\/\/www.alpha-data.com\/dcp\/products.php?product=adm-pcie-9v3."},{"key":"e_1_3_2_4_2","volume-title":"AXI High Bandwidth Memory Controller v1.0. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/ip_documentation\/hbm\/v1_0\/pg276-axi-hbm.pdf","unstructured":"AXI High Bandwidth Memory Controller v1.0. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/ip_documentation\/hbm\/v1_0\/pg276-axi-hbm.pdf."},{"key":"e_1_3_2_5_2","volume-title":"AXI Reference Guide. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/ip_documentation\/ug761_axi_reference_guide.pdf","unstructured":"AXI Reference Guide. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/ip_documentation\/ug761_axi_reference_guide.pdf."},{"key":"e_1_3_2_6_2","volume-title":"CentOS-7 (2009) Release Notes. Retrieved from https:\/\/wiki.centos.org\/Manuals\/ReleaseNotes\/CentOS7.2009","unstructured":"CentOS-7 (2009) Release Notes. Retrieved from https:\/\/wiki.centos.org\/Manuals\/ReleaseNotes\/CentOS7.2009."},{"key":"e_1_3_2_7_2","volume-title":"GCC, the GNU Compiler Collection. Retrieved from https:\/\/gcc.gnu.org\/","unstructured":"GCC, the GNU Compiler Collection. Retrieved from https:\/\/gcc.gnu.org\/."},{"key":"e_1_3_2_8_2","volume-title":"High Bandwidth Memory (HBM) DRAM (JESD235). Retrieved from https:\/\/www.jedec.org\/document_search?search_api_views_fulltext=jesd235","unstructured":"High Bandwidth Memory (HBM) DRAM (JESD235). Retrieved from https:\/\/www.jedec.org\/document_search?search_api_views_fulltext=jesd235."},{"key":"e_1_3_2_9_2","volume-title":"High Bandwidth Memory (HBM) DRAM. Retrieved from https:\/\/www.jedec.org\/sites\/default\/files\/JESD235B-HBM_Ballout.zip","unstructured":"High Bandwidth Memory (HBM) DRAM. Retrieved from https:\/\/www.jedec.org\/sites\/default\/files\/JESD235B-HBM_Ballout.zip."},{"key":"e_1_3_2_10_2","volume-title":"IBM XL C\/C++ for Linux. Retrieved from https:\/\/www.ibm.com\/products\/xl-cpp-linux-compiler-power","unstructured":"IBM XL C\/C++ for Linux. Retrieved from https:\/\/www.ibm.com\/products\/xl-cpp-linux-compiler-power."},{"key":"e_1_3_2_11_2","volume-title":"Intel Stratix 10 MX FPGAs. Retrieved from https:\/\/www.intel.com\/content\/www\/us\/en\/products\/programmable\/sip\/stratix-10-mx.html","unstructured":"Intel Stratix 10 MX FPGAs. Retrieved from https:\/\/www.intel.com\/content\/www\/us\/en\/products\/programmable\/sip\/stratix-10-mx.html."},{"key":"e_1_3_2_12_2","volume-title":"Intel\u00ae Xeon Phi\u2122 Processor 7230 (16GB, 1.30 GHz, 64 core). Retrieved from https:\/\/www.intel.com\/content\/www\/us\/en\/products\/sku\/94034\/intel-xeon-phi-processor-7230-16gb-1-30-ghz-64-core\/specifications.html","unstructured":"Intel\u00ae Xeon Phi\u2122 Processor 7230 (16GB, 1.30 GHz, 64 core). Retrieved from https:\/\/www.intel.com\/content\/www\/us\/en\/products\/sku\/94034\/intel-xeon-phi-processor-7230-16gb-1-30-ghz-64-core\/specifications.html."},{"key":"e_1_3_2_13_2","volume-title":"NVIDIA\u00ae TESLA\u00ae P100 GPU Accelerator. Retrieved from https:\/\/images.nvidia.com\/content\/tesla\/pdf\/nvidia-tesla-p100-PCIe-datasheet.pdf","unstructured":"NVIDIA\u00ae TESLA\u00ae P100 GPU Accelerator. Retrieved from https:\/\/images.nvidia.com\/content\/tesla\/pdf\/nvidia-tesla-p100-PCIe-datasheet.pdf."},{"key":"e_1_3_2_14_2","volume-title":"OC-Accel. Retrieved from https:\/\/opencapi.github.io\/oc-accel-doc\/","unstructured":"OC-Accel. Retrieved from https:\/\/opencapi.github.io\/oc-accel-doc\/."},{"key":"e_1_3_2_15_2","volume-title":"OpenPOWER Work Groups. Retrieved from https:\/\/openpowerfoundation.org\/technical\/working-groups","unstructured":"OpenPOWER Work Groups. Retrieved from https:\/\/openpowerfoundation.org\/technical\/working-groups."},{"key":"e_1_3_2_16_2","volume-title":"RDIMM. Retrieved from https:\/\/www.micron.com\/products\/dram-modules\/rdimm","unstructured":"RDIMM. Retrieved from https:\/\/www.micron.com\/products\/dram-modules\/rdimm."},{"key":"e_1_3_2_17_2","volume-title":"Ubuntu 20.04.3 LTS (Focal Fossa). Retrieved from https:\/\/releases.ubuntu.com\/20.04\/","unstructured":"Ubuntu 20.04.3 LTS (Focal Fossa). Retrieved from https:\/\/releases.ubuntu.com\/20.04\/."},{"key":"e_1_3_2_18_2","volume-title":"UltraScale Architecture Memory Resources. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/user_guides\/ug573-ultrascale-memory-resources.pdf","unstructured":"UltraScale Architecture Memory Resources. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/user_guides\/ug573-ultrascale-memory-resources.pdf."},{"key":"e_1_3_2_19_2","volume-title":"Virtex UltraScale+ HBM FPGA: A Revolutionary Increase in Memory Performance. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/white_papers\/wp485-hbm.pdf","unstructured":"Virtex UltraScale+ HBM FPGA: A Revolutionary Increase in Memory Performance. Retrieved from https:\/\/www.xilinx.com\/support\/documentation\/white_papers\/wp485-hbm.pdf."},{"key":"e_1_3_2_20_2","volume-title":"Virtex UltraScale+. Retrieved from https:\/\/www.xilinx.com\/products\/silicon-devices\/fpga\/virtex-ultrascale-plus.html","unstructured":"Virtex UltraScale+. Retrieved from https:\/\/www.xilinx.com\/products\/silicon-devices\/fpga\/virtex-ultrascale-plus.html."},{"key":"e_1_3_2_21_2","volume-title":"Vivado High-Level Synthesis. Retrieved from https:\/\/www.xilinx.com\/products\/design-tools\/vivado\/integration\/esl-design.html","unstructured":"Vivado High-Level Synthesis. Retrieved from https:\/\/www.xilinx.com\/products\/design-tools\/vivado\/integration\/esl-design.html."},{"key":"e_1_3_2_22_2","volume-title":"Xilinx VCU1525. Retrieved from https:\/\/www.xilinx.com\/products\/boards-and-kits\/ vcu1525-a.html","unstructured":"Xilinx VCU1525. Retrieved from https:\/\/www.xilinx.com\/products\/boards-and-kits\/ vcu1525-a.html."},{"key":"e_1_3_2_23_2","volume-title":"Xilinx Virtex UltraScale+. Retrieved from https:\/\/www.xilinx.com\/products\/silicon-devices\/fpga\/virtex-ultrascale-plus.html","unstructured":"Xilinx Virtex UltraScale+. Retrieved from https:\/\/www.xilinx.com\/products\/silicon-devices\/fpga\/virtex-ultrascale-plus.html."},{"key":"e_1_3_2_24_2","volume-title":"Xilinx Vivado. Retrieved from https:\/\/www.xilinx.com\/support\/download.html","unstructured":"Xilinx Vivado. Retrieved from https:\/\/www.xilinx.com\/support\/download.html."},{"key":"e_1_3_2_25_2","volume-title":"ISCA","author":"Ahn Junwhan","unstructured":"Junwhan Ahn, Sungpack Hong, Sungjoo Yoo, Onur Mutlu, and Kiyoung Choi. 2015. A scalable processing-in-memory accelerator for parallel graph processing. In ISCA."},{"key":"e_1_3_2_26_2","volume-title":"ISCA","author":"Ahn Junwhan","unstructured":"Junwhan Ahn, Sungjoo Yoo, Onur Mutlu, and Kiyoung Choi. 2015. PIM-Enabled instructions: A low-overhead, locality-aware processing-in-memory architecture. In ISCA."},{"key":"e_1_3_2_27_2","volume-title":"ISCA","author":"Akin Berkin","unstructured":"Berkin Akin, Franz Franchetti, and James C. Hoe. Data reorganization in memory using 3D-stacked DRAM. 2015. In ISCA."},{"key":"e_1_3_2_28_2","volume-title":"MICRO","author":"Alian M.","unstructured":"M. Alian, S. W. Min, H. Asgharimoghaddam, A. Dhar, D. K. Wang, T. Roewer, A. McPadden, O. O\u2019Halloran, D. Chen, J. Xiong, D. Kim, W. Hwu, and N. S. Kim. 2018. Application-Transparent near-memory processing architecture with memory channel network. In MICRO."},{"key":"e_1_3_2_29_2","volume-title":"IEEE Micro","author":"Alser Mohammed","unstructured":"Mohammed Alser, Z\u00fclal Bing\u00f6l, Damla Senol Cali, Jeremie Kim, Saugata Ghose, Can Alkan, and Onur Mutlu. 2020. Accelerating genome analysis: A primer on an ongoing journey. In IEEE Micro."},{"key":"e_1_3_2_30_2","volume-title":"Bioinformatics","author":"Alser Mohammed","unstructured":"Mohammed Alser, Hasan Hassan, Akash Kumar, Onur Mutlu, and Can Alkan. 2019. Shouji: A fast and efficient pre-alignment filter for sequence alignment. Bioinformatics 35, 21 (2019), 4255\u20134263."},{"key":"e_1_3_2_31_2","volume-title":"Bioinformatics","author":"Alser Mohammed","unstructured":"Mohammed Alser, Hasan Hassan, Hongyi Xin, O\u01e7uz Ergin, Onur Mutlu, and Can Alkan. 2017. GateKeeper: A new hardware architecture for accelerating pre-alignment in DNA short read mapping. Bioinformatics 33, 21 (2017), 3355\u20133363."},{"key":"e_1_3_2_32_2","volume-title":"Genome Biology","author":"Alser Mohammed","unstructured":"Mohammed Alser, Jeremy Rotman, Kodi Taraszka, Huwenbo Shi, Pelin Icer Baykal, Harry Taegyun Yang, Victor Xue, Sergey Knyazev, Benjamin D. Singer, Brunilda Balliu, et al. 2020. Technology dictates algorithms: Recent developments in read alignment. In Genome Biology, Vol. 22. 1\u201334."},{"key":"e_1_3_2_33_2","volume-title":"Bioinformatics","author":"Alser Mohammed","unstructured":"Mohammed Alser, Taha Shahroodi, Juan Gomez-Luna, Can Alkan, and Onur Mutlu. 2020. SneakySnake: A fast and accurate universal genome pre-alignment filter for CPUs, GPUs, and FPGAs. Bioinformatics 36, 22\u201323 (2020), 5282\u20135290."},{"key":"e_1_3_2_34_2","volume-title":"DAC","author":"Angizi Shaahin","unstructured":"Shaahin Angizi, Jiao Sun, Wei Zhang, and Deliang Fan. 2019. AlignS: A processing-in-memory accelerator for DNA short read alignment leveraging SOT-MRAM. In DAC."},{"key":"e_1_3_2_35_2","volume-title":"PACT","author":"Ansel Jason","unstructured":"Jason Ansel, Shoaib Kamil, Kalyan Veeramachaneni, Jonathan Ragan-Kelley, Jeffrey Bosboom, Una-May O\u2019Reilly, and Saman Amarasinghe. 2014. OpenTuner: An extensible framework for program autotuning. In PACT."},{"key":"e_1_3_2_36_2","volume-title":"PACT","author":"Armejach Adri\u00e0","unstructured":"Adri\u00e0 Armejach, Helena Caminal, Juan M. Cebrian, Rekai Gonz\u00e1lez-Alberquilla, Chris Adeniyi-Jones, Mateo Valero, Marc Casas, and Miquel Moret\u00f3. 2018. Stencil codes on a vector length agnostic architecture. In PACT."},{"key":"e_1_3_2_37_2","volume-title":"MICRO","author":"Asghari-Moghaddam Hadi","unstructured":"Hadi Asghari-Moghaddam, Young Hoon Son, Jung Ho Ahn, and Nam Sung Kim. 2016. Chameleon: Versatile and practical Near-DRAM acceleration architecture for large memory systems. In MICRO."},{"key":"e_1_3_2_38_2","volume-title":"SIGMOD","author":"Babarinsa Oreoluwatomiwa O.","unstructured":"Oreoluwatomiwa O. Babarinsa and Stratos Idreos. 2015. JAFAR: Near-Data processing for databases. In SIGMOD."},{"key":"e_1_3_2_39_2","volume-title":"TC","author":"Barnes George H.","unstructured":"George H. Barnes, Richard M. Brown, Maso Kato, David J. Kuck, Daniel L. Slotnick, and Richard A. Stokes. 1968. The ILLIAC IV computer. In TC."},{"key":"e_1_3_2_40_2","volume-title":"OFA","author":"Benton Brad","unstructured":"Brad Benton. 2017. CCIX, Gen-Z, OpenCAPI: Overview and comparison. In OFA."},{"key":"e_1_3_2_41_2","volume-title":"MICRO","author":"Besta Maciej","unstructured":"Maciej Besta, Raghavendra Kanakagiri, Grzegorz Kwasniewski, Rachata Ausavarungnirun, Jakub Ber\u00e1nek, Konstantinos Kanellopoulos, Kacper Janda, Zur Vonarburg-Shmaria, Lukas Gianinazzi, Ioana Stefan, et al. 2021. SISA: Set-Centric instruction set architecture for graph mining on processing-in-memory systems. In MICRO."},{"key":"e_1_3_2_42_2","volume-title":"ISC","author":"Bianco M.","unstructured":"M. Bianco, T. Diamanti, O. Fuhrer, T. Gysi, X. Lapillonne, C. Osuna, and T. Schulthess. 2013. A GPU capable version of the COSMO weather model. In ISC."},{"key":"e_1_3_2_43_2","volume-title":"JCP","author":"Bonaventura Luca","unstructured":"Luca Bonaventura. 2000. A semi-implicit semi-Lagrangian scheme using the height coordinate for a nonhydrostatic and fully elastic model of atmospheric flows. In JCP."},{"key":"e_1_3_2_44_2","volume-title":"PACT","author":"Boroumand Amirali","unstructured":"Amirali Boroumand, Saugata Ghose, Berkin Akin, Ravi Narayanaswami, Geraldo F. Oliveira, Xiaoyu Ma, Eric Shiu, and Onur Mutlu. 2021. Google neural network models for edge devices: Analyzing and mitigating machine learning inference bottlenecks. In PACT."},{"key":"e_1_3_2_45_2","volume-title":"ASPLOS","author":"Boroumand Amirali","unstructured":"Amirali Boroumand, Saugata Ghose, Youngsok Kim, Rachata Ausavarungnirun, Eric Shiu, Rahul Thakur, Daehyun Kim, Aki Kuusela, Allan Knies, Parthasarathy Ranganathan, and Onur Mutlu. 2018. Google workloads for consumer devices: Mitigating data movement bottlenecks. In ASPLOS."},{"key":"e_1_3_2_46_2","volume-title":"ISCA","author":"Boroumand Amirali","unstructured":"Amirali Boroumand, Saugata Ghose, Minesh Patel, Hasan Hassan, Brandon Lucia, Rachata Ausavarungnirun, Kevin Hsieh, Nastaran Hajinazar, Krishna T Malladi, Hongzhong Zheng, et al. 2019. CoNDA: Efficient cache coherence support for near-data accelerators. In ISCA."},{"key":"e_1_3_2_47_2","volume-title":"CAL","author":"Boroumand Amirali","unstructured":"Amirali Boroumand, Saugata Ghose, Minesh Patel, Hasan Hassan, Brandon Lucia, Kevin Hsieh, Krishna T. Malladi, Hongzhong Zheng, and Onur Mutlu. 2016. LazyPIM: An efficient cache coherence mechanism for processing-in-memory. In CAL."},{"key":"e_1_3_2_48_2","volume-title":"MICRO","author":"Cali Damla Senol","unstructured":"Damla Senol Cali, Gurpreet S. Kalsi, Z\u00fclal Bing\u00f6l, Can Firtina, Lavanya Subramanian, Jeremie S. Kim, Rachata Ausavarungnirun, Mohammed Alser, Juan G\u00f3mez Luna, Amirali Boroumand, Anant Nori, Allison Scibisz, Sreenivas Subramoney, Can Alkan, Saugata Ghose, and Onur Mutlu. 2020. GenASM: A high-performance, low-power approximate string matching acceleration framework for genome sequence analysis. In MICRO."},{"key":"e_1_3_2_49_2","volume-title":"MICRO","author":"Caulfield A. M.","unstructured":"A. M. Caulfield, E. S. Chung, A. Putnam, H. Angepat, J. Fowers, M. Haselman, S. Heil, M. Humphrey, P. Kaur, J. Kim, D. Lo, T. Massengill, K. Ovtcharov, M. Papamichael, L. Woods, S. Lanka, D. Chiou, and D. Burger. 2016. A cloud-scale acceleration architecture. In MICRO."},{"key":"e_1_3_2_50_2","volume-title":"ICPE","author":"Chang Li-Wen","unstructured":"Li-Wen Chang, Juan G\u00f3mez-Luna, Izzat El Hajj, Sitao Huang, Deming Chen, and Wen-mei Hwu. 2017. Collaborative computing for heterogeneous integrated systems. In ICPE."},{"key":"e_1_3_2_51_2","volume-title":"ISCA","author":"Chi Ping","unstructured":"Ping Chi, Shuangchen Li, Cong Xu, Tao Zhang, Jishen Zhao, Yongpan Liu, Yu Wang, and Yuan Xie. 2016. PRIME: A novel processing-in-memory architecture for neural network computation in ReRAM-based main memory. In ISCA."},{"key":"e_1_3_2_52_2","volume-title":"ICCAD","author":"Chi Yuze","unstructured":"Yuze Chi, Jason Cong, Peng Wei, and Peipei Zhou. 2018. SODA: Stencil with optimized dataflow architecture. In ICCAD."},{"key":"e_1_3_2_53_2","volume-title":"DAC","author":"Choi Young-kyu","unstructured":"Young-kyu Choi, Jason Cong, Zhenman Fang, Yuchen Hao, Glenn Reinman, and Peng Wei. 2016. A quantitative analysis on microarchitectures of modern CPU-FPGA platforms. In DAC."},{"key":"e_1_3_2_54_2","volume-title":"IPDPS","author":"Christen Matthias","unstructured":"Matthias Christen, Olaf Schenk, and Helmar Burkhart. 2011. PATUS: A code generation and autotuning framework for parallel iterative stencil computations on modern microarchitectures. In IPDPS."},{"key":"e_1_3_2_55_2","volume-title":"SIAM Review","author":"Datta Kaushik","unstructured":"Kaushik Datta, Shoaib Kamil, Samuel Williams, Leonid Oliker, John Shalf, and Katherine Yelick. 2009. Optimization and performance modeling of stencil computations on modern microprocessors. In SIAM Review."},{"key":"e_1_3_2_56_2","volume-title":"PPoPP","author":"Fine Licht Johannes de","unstructured":"Johannes de Fine Licht, Michaela Blott, and Torsten Hoefler. 2018. Designing scalable FPGA architectures using high-level synthesis. In PPoPP."},{"key":"e_1_3_2_57_2","volume-title":"CGO","author":"Fine Licht Johannes de","unstructured":"Johannes de Fine Licht, Andreas Kuster, Tiziano De Matteis, Tal Ben-Nun, Dominic Hofer, and Torsten Hoefler. 2021. StencilFlow: Mapping large stencil programs to distributed spatial computing systems. In CGO."},{"key":"e_1_3_2_58_2","volume-title":"COOL CHIPS","author":"Diamantopoulos Dionysios","unstructured":"Dionysios Diamantopoulos, Heiner Giefers, and Christoph Hagleitner. 2018. ecTALK: Energy efficient coherent transprecision accelerators\u2014The bidirectional long short-term memory neural network case. In COOL CHIPS."},{"key":"e_1_3_2_59_2","volume-title":"FPT","author":"Diamantopoulos Dionysios","unstructured":"Dionysios Diamantopoulos and Christoph Hagleitner. 2018. A system-level transprecision FPGA accelerator for BLSTM using on-chip memory reshaping. In FPT."},{"key":"e_1_3_2_60_2","volume-title":"DWD, GB Forschung und Entwicklung","author":"Doms G.","unstructured":"G. Doms and U. Sch\u00e4ttler. 1999. The nonhydrostatic limited-area model LM (Lokal-model) of the DWD. Part I: Scientific documentation. In DWD, GB Forschung und Entwicklung."},{"key":"e_1_3_2_61_2","volume-title":"ISCA","author":"Drumond Mario","unstructured":"Mario Drumond, Alexandros Daglis, Nooshin Mirzadeh, Dmitrii Ustiugov, Javier Picorel, Babak Falsafi, Boris Grot, and Dionisios Pnevmatikatos. 2017. The Mondrian data engine. In ISCA."},{"key":"e_1_3_2_62_2","volume-title":"JINST","author":"Duarte Javier","unstructured":"Javier Duarte, Song Han, Philip Harris, Sergo Jindariani, Edward Kreinar, Benjamin Kreis, Jennifer Ngadiuba, Maurizio Pierini, Ryan Rivera, Nhan Tran, and Z. Wu. 2018. Fast inference of deep neural networks in FPGAs for pinproceedings physics. In JINST."},{"key":"e_1_3_2_63_2","volume-title":"VLDB","author":"Fang Jian","unstructured":"Jian Fang, Yvo T. B. Mulder, Jan Hidders, Jinho Lee, and H. Peter Hofstee. 2020. In-memory database acceleration on FPGAs: A survey. In VLDB."},{"key":"e_1_3_2_64_2","volume-title":"HPCA","author":"Farmahini-Farahani A.","unstructured":"A. Farmahini-Farahani, J. H. Ahn, K. Morrow, and N. S. Kim. 2015. NDA: Near-DRAM acceleration architecture leveraging commodity DRAM devices and standard memory modules. In HPCA."},{"key":"e_1_3_2_65_2","volume-title":"ICCD","author":"Fernandez Ivan","unstructured":"Ivan Fernandez, Ricardo Quislant, Eladio Guti\u00e9rrez, Oscar Plata, Christina Giannoula, Mohammed Alser, Juan G\u00f3mez-Luna, and Onur Mutlu. 2020. NATSA: A near-data processing accelerator for time series analysis. In ICCD."},{"key":"e_1_3_2_66_2","volume-title":"Proceedings of the IEEE","author":"Flynn Michael J.","unstructured":"Michael J. Flynn. 1966. Very high-speed computing systems. Proceedings of the IEEE 54, 12 (1966), 1901\u20131909."},{"key":"e_1_3_2_67_2","volume-title":"FPGA","author":"Fu Haohuan","unstructured":"Haohuan Fu and Robert G. Clapp. 2011. Eliminating the memory bottleneck: An FPGA-based solution for 3D reverse time migration. In FPGA."},{"key":"e_1_3_2_68_2","volume-title":"FPGA","author":"Gaide Brian","unstructured":"Brian Gaide, Dinesh Gaitonde, Chirag Ravishankar, and Trevor Bauer. 2019. Xilinx adaptive compute acceleration platform: Versal\u2122 architecture. In FPGA."},{"key":"e_1_3_2_69_2","volume-title":"MICRO","author":"Gao Fei","unstructured":"Fei Gao, Georgios Tziantzioulis, and David Wentzlaff. 2019. ComputeDRAM: In-Memory compute using off-the-shelf DRAMs. In MICRO."},{"key":"e_1_3_2_70_2","volume-title":"PACT","author":"Gao Mingyu","unstructured":"Mingyu Gao, Grant Ayers, and Christos Kozyrakis. 2015. Practical near-data processing for in-memory analytics frameworks. In PACT."},{"key":"e_1_3_2_71_2","volume-title":"HPCA","author":"Gao M.","unstructured":"M. Gao and C. Kozyrakis. 2016. HRL: Efficient and flexible reconfigurable logic for near-data processing. In HPCA."},{"key":"e_1_3_2_72_2","volume-title":"IBM JRD","author":"Ghose Saugata","unstructured":"Saugata Ghose, Amirali Boroumand, Jeremie S. Kim, Juan G\u00f3mez-Luna, and Onur Mutlu. 2019. Processing-in-memory: A workload-driven perspective. In IBM JRD."},{"key":"e_1_3_2_73_2","volume-title":"POMACS","author":"Ghose Saugata","unstructured":"Saugata Ghose, Tianshi Li, Nastaran Hajinazar, Damla Senol Cali, and Onur Mutlu. 2019. Demystifying complex workload-DRAM interactions: An Experimental Study. In POMACS."},{"key":"e_1_3_2_74_2","volume-title":"HPCA","author":"Giannoula Christina","unstructured":"Christina Giannoula, Nandita Vijaykumar, Nikela Papadopoulou, Vasileios Karakostas, Ivan Fernandez, Juan G\u00f3mez-Luna, Lois Orosa, Nectarios Koziris, Georgios Goumas, and Onur Mutlu. 2021. SynCron: Efficient synchronization support for near-data-processing architectures. In HPCA."},{"key":"e_1_3_2_75_2","volume-title":"DATE","author":"Giefers Heiner","unstructured":"Heiner Giefers, Raphael Polig, and Christoph Hagleitner. 2015. Accelerating arithmetic kernels with coherent attached FPGA coprocessors. In DATE."},{"key":"e_1_3_2_76_2","doi-asserted-by":"crossref","unstructured":"Juan G\u00f3mez-Luna Izzat El Hajj Ivan Fernandez Christina Giannoula Geraldo F. Oliveira and Onur Mutlu. 2021. Benchmarking a new paradigm: An experimental analysis of a real processing-in-memory architecture arxiv.","DOI":"10.1109\/ACCESS.2022.3174101"},{"key":"e_1_3_2_77_2","volume-title":"CUT","author":"G\u00f3mez-Luna Juan","unstructured":"Juan G\u00f3mez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2021. Benchmarking memory-centric computing systems: Analysis of real processing-in-memory hardware. In CUT."},{"key":"e_1_3_2_78_2","volume-title":"ICS","author":"Gonz\u00e1lez Jos\u00e9","unstructured":"Jos\u00e9 Gonz\u00e1lez and Antonio Gonz\u00e1lez. 1997. Speculative execution via address prediction and data prefetching. In ICS."},{"key":"e_1_3_2_79_2","volume-title":"ISCA","author":"Gu Boncheol","unstructured":"Boncheol Gu, Andre S. Yoon, Duck-Ho Bae, Insoon Jo, Jinyoung Lee, Jonghyun Yoon, Jeong-Uk Kang, Moonsang Kwon, Chanho Yoon, Sangyeun Cho, Jaeheon Jeong, and Duckhyun Chang. 2016. Biscuit: A framework for near-data processing of big data workloads. In ISCA."},{"key":"e_1_3_2_80_2","volume-title":"SC","author":"Gysi Tobias","unstructured":"Tobias Gysi, Tobias Grosser, and Torsten Hoefler. 2015. MODESTO: Data-centric analytic optimization of complex stencil programs on heterogeneous architectures. In SC."},{"key":"e_1_3_2_81_2","volume-title":"ASPLOS","author":"Hajinazar Nastaran","unstructured":"Nastaran Hajinazar, Geraldo F. Oliveira, Sven Gregorio, Jo\u00e3o Ferreira, Nika Mansouri Ghiasi, Minesh Patel, Mohammed Alser, Saugata Ghose, Juan G\u00f3mez Luna, and Onur Mutlu. 2021. SIMDRAM: An end-to-end framework for bit-serial SIMD computing in DRAM. In ASPLOS."},{"key":"e_1_3_2_82_2","volume-title":"ISCA","author":"Hashemi Milad","unstructured":"Milad Hashemi, Eiman Ebrahimi, Onur Mutlu, Yale N. Patt, et al. 2016. Accelerating dependent cache misses with an enhanced memory controller. In ISCA."},{"key":"e_1_3_2_83_2","volume-title":"MICRO","author":"Hashemi Milad","unstructured":"Milad Hashemi, Onur Mutlu, and Yale N. Patt. 2016. Continuous runahead: Transparent hardware acceleration for memory intensive workloads. In MICRO."},{"key":"e_1_3_2_84_2","volume-title":"CC","author":"Henretty Tom","unstructured":"Tom Henretty, Kevin Stock, Louis-No\u00ebl Pouchet, Franz Franchetti, J. Ramanujam, and P. Sadayappan. 2011. Data layout transformation for stencil computations on short-vector SIMD architectures. In CC."},{"key":"e_1_3_2_85_2","volume-title":"IMAVIS","author":"Hermosilla Txomin","unstructured":"Txomin Hermosilla, E. Bermejo, A. Balaguer, and Luis A. Ruiz. 2008. Non-linear fourth-order image interpolation for subpixel edge detection and localization. In IMAVIS."},{"key":"e_1_3_2_86_2","volume-title":"ISCA","author":"Hsieh Kevin","unstructured":"Kevin Hsieh, Eiman Ebrahimi, Gwangsun Kim, Niladrish Chatterjee, Mike O\u2019Connor, Nandita Vijaykumar, Onur Mutlu, and Stephen W. Keckler. 2016. Transparent offloading and mapping (TOM): Enabling programmer-transparent near-data processing in GPU systems. In ISCA."},{"key":"e_1_3_2_87_2","volume-title":"ICCD","author":"Hsieh Kevin","unstructured":"Kevin Hsieh, Samira Khan, Nandita Vijaykumar, Kevin K. Chang, Amirali Boroumand, Saugata Ghose, and Onur Mutlu. 2016. Accelerating pointer chasing in 3D-Stacked memory: Challenges, mechanisms, evaluation. In ICCD."},{"key":"e_1_3_2_88_2","volume-title":"ICPE","author":"Huang Sitao","unstructured":"Sitao Huang, Li-Wen Chang, Izzat El Hajj, Simon Garcia de Gonzalo, Juan G\u00f3mez-Luna, Sai Rahul Chalamalasetti, Mohamed El-Hadedy, Dejan Milojicic, Onur Mutlu, Deming Chen, and Wen-mei Hwu. 2019. Analysis and modeling of collaborative execution strategies for heterogeneous CPU-FPGA architectures. In ICPE."},{"key":"e_1_3_2_89_2","volume-title":"Computers & Fluids","author":"Huynh H. T.","unstructured":"H. T. Huynh, Zhi J. Wang, and Peter E. Vincent. 2014. High-order methods for computational fluid dynamics: A brief review of compact differential formulations on unstructured grids. In Computers & Fluids."},{"key":"e_1_3_2_90_2","volume-title":"VLDB","author":"Istv\u00e1n Zsolt","unstructured":"Zsolt Istv\u00e1n, David Sidler, and Gustavo Alonso. 2017. Caribou: Intelligent distributed storage. In VLDB."},{"key":"e_1_3_2_91_2","volume-title":"FPGA","author":"Jiang Jiantong","unstructured":"Jiantong Jiang, Zeke Wang, Xue Liu, Juan G\u00f3mez-Luna, Nan Guan, Qingxu Deng, Wei Zhang, and Onur Mutlu. 2020. Boyi: A systematic framework for automatically deciding the right execution model of OpenCL applications on FPGAs. In FPGA."},{"key":"e_1_3_2_92_2","volume-title":"Computer","author":"Jongerius R.","unstructured":"R. Jongerius, S. Wijnholds, R. Nijboer, and H. Corporaal. 2014. An end-to-end computing model for the square kilometre array. In Computer."},{"key":"e_1_3_2_93_2","volume-title":"ISCA","author":"Jun Sang-Woo","unstructured":"Sang-Woo Jun, Ming Liu, Sungjin Lee, Jamey Hicks, John Ankcorn, Myron King, Shuotao Xu, et al. 2015. BlueDBM: An appliance for big data analytics. In ISCA."},{"key":"e_1_3_2_94_2","volume-title":"MSST","author":"Kang Yangwook","unstructured":"Yangwook Kang, Yang-suk Kee, Ethan L. Miller, and Chanik Park. 2013. Enabling cost-effective data processing with smart SSD. In MSST."},{"key":"e_1_3_2_95_2","volume-title":"FCCM","author":"Kara Kaan","unstructured":"Kaan Kara, Dan Alistarh, Gustavo Alonso, Onur Mutlu, and Ce Zhang. 2017. FPGA-accelerated dense linear machine learning: A precision-convergence trade-off. In FCCM."},{"key":"e_1_3_2_96_2","volume-title":"FPL","author":"Kara Kaan","unstructured":"Kaan Kara, Christoph Hagleitner, Dionysios Diamantopoulos, Dimitris Syrivelis, and Gustavo Alonso. 2020. High bandwidth memory on FPGAs: A data analytics perspective. In FPL."},{"key":"e_1_3_2_97_2","volume-title":"ISCA","author":"Ke L.","unstructured":"L. Ke, U. Gupta, B. Y. Cho, D. Brooks, V. Chandra, U. Diril, A. Firoozshahian, K. Hazelwood, B. Jia, H. S. Lee, M. Li, B. Maher, D. Mudigere, M. Naumov, M. Schatz, M. Smelyanskiy, X. Wang, B. Reagen, C. Wu, M. Hempstead, and X. Zhang. 2020. RecNMP: Accelerating personalized recommendation with near-memory processing. In ISCA."},{"key":"e_1_3_2_98_2","volume-title":"Atmosphere-Ocean","author":"Kehler Scott","unstructured":"Scott Kehler, John Hanesiak, Michelle Curry, David Sills, and Neil Taylor. 2016. High resolution deterministic prediction system (HRDPS) simulations of Manitoba lake breezes. In Atmosphere-Ocean."},{"key":"e_1_3_2_99_2","volume-title":"ISCA","author":"Kim Duckhwan","unstructured":"Duckhwan Kim, Jaeha Kung, Sek Chai, Sudhakar Yalamanchili, and Saibal Mukhopadhyay. 2016. Neurocube: A programmable digital neuromorphic architecture with high-density 3D memory. In ISCA."},{"key":"e_1_3_2_100_2","volume-title":"JSSC","author":"Kim J.","unstructured":"J. Kim, C. S. Oh, H. Lee, D. Lee, H. R. Hwang, S. Hwang, B. Na, J. Moon, J. Kim, H. Park, J. Ryu, K. Park, S. K. Kang, S. Kim, H. Kim, J. Bang, H. Cho, M. Jang, C. Han, J. LeeLee, J. S. Choi, and Y. Jun. 2012. A 1.2 V 12.8 GB\/s 2 Gb mobile wide-I\/O DRAM with 4 \\( \\times \\) 128 I\/Os using TSV based stacking. In JSSC."},{"key":"e_1_3_2_101_2","volume-title":"BMC Genomics","author":"Kim Jeremie S.","unstructured":"Jeremie S. Kim, Damla Senol Cali, Hongyi Xin, Donghyuk Lee, Saugata Ghose, Mohammed Alser, Hasan Hassan, Oguz Ergin, Can Alkan, and Onur Mutlu. 2018. GRIM-Filter: Fast seed location filtering in DNA read mapping using processing-in-memory technologies. BMC Genomics 19, 2 (2018), 23\u201340."},{"key":"e_1_3_2_102_2","volume-title":"MICRO","author":"Koo Gunjae","unstructured":"Gunjae Koo, Kiran Kumar Matam, Te I, H. V. Krishna Giri Narra, Jing Li, Hung-Wei Tseng, Steven Swanson, and Murali Annavaram. 2017. Summarizer: Trading communication with computing near storage. In MICRO."},{"key":"e_1_3_2_103_2","volume-title":"ISSCC","author":"Kwon Young-Cheon","unstructured":"Young-Cheon Kwon, Suk Han Lee, Jaehoon Lee, Sang-Hyuk Kwon, Je Min Ryu, Jong-Pil Son, Seongil O, Hak-Soo Yu, Haesuk Lee, Soo Young Kim, Youngmin Cho, Jin Guk Kim, Jongyoon Choi, Hyun-Sung Shin, Jin Kim, BengSeng Phuah, HyoungMin Kim, Myeong Jun Song, Ahn Choi, Daeho Kim, SooYoung Kim, Eun-Bong Kim, David Wang, Shinhaeng Kang, Yuhwan Ro, Seungwoo Seo, JoonHo Song, Jaeyoun Youn, Kyomin Sohn, and Nam Sung Kim. 2021. A 20nm 6GB function-in-memory DRAM, based on HBM2 with a 1.2TFLOPS programmable computing unit using bank-level parallelism, for machine learning applications. In ISSCC."},{"key":"e_1_3_2_104_2","volume-title":"FPGA","author":"Lai Yi-Hsiang","unstructured":"Yi-Hsiang Lai, Yuze Chi, Yuwei Hu, Jie Wang, Cody Hao Yu, Yuan Zhou, Jason Cong, and Zhiru Zhang. 2019. HeteroCL: A multi-paradigm programming infrastructure for software-defined reconfigurable computing. In FPGA."},{"key":"e_1_3_2_105_2","volume-title":"ACM TACO","author":"Lee Donghyuk","unstructured":"Donghyuk Lee, Saugata Ghose, Gennady Pekhimenko, Samira Khan, and Onur Mutlu. 2016. Simultaneous multi-layer access: Improving 3D-Stacked memory bandwidth at low cost. ACM TACO 12, 4 (2016), 1\u201329."},{"key":"e_1_3_2_106_2","volume-title":"ISSCC","author":"Lee D. U.","unstructured":"D. U. Lee, K. W. Kim, K. W. Kim, H. Kim, J. Y. Kim, Y. J. Park, J. H. Kim, D. S. Kim, H. B. Park, J. W. Shin, J. H. Cho, K. H. Kwon, M. J. Kim, J. Lee, K. W. Park, B. Chung, and S. Hong. 2014. 25.2 A 1.2V 8Gb 8-channel 128GB\/s high-bandwidth memory (HBM) stacked DRAM with effective microbump I\/O test methods using 29nm process and TSV. In ISSCC."},{"key":"e_1_3_2_107_2","volume-title":"VLDB","author":"Lee Jinho","unstructured":"Jinho Lee, Heesu Kim, Sungjoo Yoo, Kiyoung Choi, H. Peter Hofstee, Gi-Joon Nam, Mark R. Nutter, and Damir Jamsek. 2017. ExtraV: Boosting graph processing near storage with a coherent accelerator. In VLDB."},{"key":"e_1_3_2_108_2","volume-title":"PACT","author":"Lee Joo Hwan","unstructured":"Joo Hwan Lee, Jaewoong Sim, and Hyesoon Kim. 2015. BSSync: Processing near memory for machine learning workloads with bounded staleness consistency models. In PACT."},{"key":"e_1_3_2_109_2","volume-title":"ISCA","author":"Lee Sukhan","unstructured":"Sukhan Lee, Shin-haeng Kang, Jaehoon Lee, Hyeonsu Kim, Eojin Lee, Seungwoo Seo, Hosang Yoon, Seungwon Lee, Kyounghwan Lim, Hyunsung Shin, Jinhyun Kim, Seongil O, Anand Iyer, David Wang, Kyomin Sohn, and Nam Sung Kim. 2021. Hardware architecture and software stack for FIM based on commercial DRAM technology. In ISCA."},{"key":"e_1_3_2_110_2","volume-title":"IPDPS","author":"Lee Vincent T.","unstructured":"Vincent T. Lee, Amrita Mazumdar, Carlo C. del Mundo, Armin Alaghi, Luis Ceze, and Mark Oskin. 2018. Application codesign of near-data processing for similarity search. In IPDPS."},{"key":"e_1_3_2_111_2","volume-title":"FPGA","author":"Li Jiajie","unstructured":"Jiajie Li, Yuze Chi, and Jason Cong. 2020. HeteroHalide: From image processing DSL to efficient FPGA acceleration. In FPGA."},{"key":"e_1_3_2_112_2","volume-title":"DAC","author":"Li Shuangchen","unstructured":"Shuangchen Li, Cong Xu, Qiaosha Zou, Jishen Zhao, Yu Lu, and Yuan Xie. 2016. Pinatubo: A processing-in-memory architecture for bulk bitwise operations in emerging non-volatile memories. In DAC."},{"key":"e_1_3_2_113_2","volume-title":"MICRO","author":"Liu Jiawen","unstructured":"Jiawen Liu, Hengyu Zhao, Matheus A. Ogleari, Dong Li, and Jishen Zhao. 2018. Processing-in-Memory for energy-efficient neural network training: A heterogeneous approach. In MICRO."},{"key":"e_1_3_2_114_2","volume-title":"SPAA","author":"Liu Zhiyu","unstructured":"Zhiyu Liu, Irina Calciu, Maurice Herlihy, and Onur Mutlu. 2017. Concurrent data structures for near-memory computing. In SPAA."},{"key":"e_1_3_2_115_2","volume-title":"HOTI","author":"Mayhew David","unstructured":"David Mayhew and Venkata Krishnan. 2003. PCI express and advanced switching: Evolutionary path to building next generation interconnects. In HOTI."},{"key":"e_1_3_2_116_2","volume-title":"IJPP","author":"Meng Jiayuan","unstructured":"Jiayuan Meng and Kevin Skadron. 2011. A performance study for iterative stencil loops on GPUs with ghost zone optimizations. In IJPP."},{"key":"e_1_3_2_117_2","volume-title":"Deploy ML models to field-programmable gate arrays (FPGAs) with Azure Machine Learning. Retrieved from https:\/\/docs.microsoft.com\/en-us\/azure\/machine-learning\/how-to-deploy-fpga-web-service","unstructured":"Microsoft. Deploy ML models to field-programmable gate arrays (FPGAs) with Azure Machine Learning. Retrieved from https:\/\/docs.microsoft.com\/en-us\/azure\/machine-learning\/how-to-deploy-fpga-web-service."},{"key":"e_1_3_2_118_2","volume-title":"ACM TACO","author":"Morad Amir","unstructured":"Amir Morad, Leonid Yavits, and Ran Ginosar. 2015. GP-SIMD processing-in-memory. ACM TACO 11, 4 (2015), 1\u201326."},{"key":"e_1_3_2_119_2","volume-title":"DATE","author":"Mutlu Onur","unstructured":"Onur Mutlu. 2021. Intelligent architectures for intelligent computing systems. In DATE."},{"key":"e_1_3_2_120_2","volume-title":"DAC","author":"Mutlu Onur","unstructured":"Onur Mutlu, Saugata Ghose, Juan G\u00f3mez-Luna, and Rachata Ausavarungnirun. 2019. Enabling practical processing in and near memory for data-intensive computing. In DAC."},{"key":"e_1_3_2_121_2","volume-title":"MicPro","author":"Mutlu Onur","unstructured":"Onur Mutlu, Saugata Ghose, Juan G\u00f3mez-Luna, and Rachata Ausavarungnirun. 2019. Processing data where it makes sense: Enabling in-memory computation. In MicPro, Vol. 67. 28\u201341."},{"key":"e_1_3_2_122_2","volume-title":"Emerging Computing: From Devices to Systems-Looking Beyond Moore and Von Neumann.","author":"Mutlu Onur","unstructured":"Onur Mutlu, Saugata Ghose, Juan G\u00f3mez-Luna, and Rachata Ausavarungnirun. 2021. A modern primer on processing in memor. In Emerging Computing: From Devices to Systems-Looking Beyond Moore and Von Neumann. Springer."},{"key":"e_1_3_2_123_2","volume-title":"HPCA","author":"Nai L.","unstructured":"L. Nai, R. Hadidi, J. Sim, H. Kim, P. Kumar, and H. Kim. 2017. GraphPIM: Enabling instruction-level PIM offloading in graph computing frameworks. In HPCA."},{"key":"e_1_3_2_124_2","volume-title":"IBM JRD","author":"Nair R.","unstructured":"R. Nair, S. F. Antao, C. Bertolli, P. Bose, J. R. Brunheroto, T. Chen, C. Cher, C. H. A. Costa, J. Doi, C. Evangelinos, B. M. Fleischer, T. W. Fox, D. S. Gallo, L. Grinberg, J. A. Gunnels, A. C. Jacob, P. Jacob, H. M. Jacobson, T. Karkhanis, C. Kim, J. H. Moreno, J. K. O\u2019Brien, M. Ohmacht, Y. Park, D. A. Prener, B. S. Rosenburg, K. D. Ryu, O. Sallenave, M. J. Serrano, P. D. M. Siegl, K. Sugavanam, and Z. Sura. 2015. Active memory cube: A processing-in-memory architecture for exascale systems. IBM JRD 59, 2\/3 (2015), 17\u20131."},{"key":"e_1_3_2_125_2","doi-asserted-by":"crossref","unstructured":"F\u00e1bio C. P. Navarro Hussein Mohsen Chengfei Yan Shantao Li Mengting Gu William Meyerson and Mark Gerstein. 2019. Genomics and data science: An application within an umbrella. Genome Biology 20 1 (2019) 1\u201311.","DOI":"10.1186\/s13059-019-1724-1"},{"key":"e_1_3_2_126_2","volume-title":"NCAR Tech. Note","author":"Neale Richard B.","unstructured":"Richard B. Neale, Chih-Chieh Chen, Andrew Gettelman, Peter H. Lauritzen, Sungsu Park, David L. Williamson, Andrew J. Conley, Rolando Garcia, Doug Kinnison, Jean-Francois Lamarque, et al. 2010. Description of the NCAR Community atmosphere model (CAM 5.0). In NCAR Tech. Note."},{"key":"e_1_3_2_127_2","volume-title":"IEEE Access","author":"Oliveira Geraldo Francisco","unstructured":"Geraldo Francisco Oliveira, Juan G\u00f3mez-Luna, Lois Orosa, Saugata Ghose, Nandita Vijaykumar, Ivan Fernandez, Mohammad Sadrosadati, and Onur Mutlu. 2021. DAMOV: A new methodology and benchmark suite for evaluating data movement bottlenecks. In IEEE Access, Vol. 9. 134457\u2013134502."},{"key":"e_1_3_2_128_2","doi-asserted-by":"publisher","DOI":"10.1201\/9781315168784"},{"key":"e_1_3_2_129_2","volume-title":"MICRO","author":"Park Jaehyun","unstructured":"Jaehyun Park, Byeongho Kim, Sungmin Yun, Eojin Lee, Minsoo Rhu, and Jung Ho Ahn. 2021. TRiM: Enhancing processor-memory interfaces with scalable tensor reduction in memory. In MICRO."},{"key":"e_1_3_2_130_2","volume-title":"PACT","author":"Pattnaik Ashutosh","unstructured":"Ashutosh Pattnaik, Xulong Tang, Adwait Jog, Onur Kayiran, Asit K. Mishra, Mahmut T. Kandemir, Onur Mutlu, and Chita R. Das. 2016. Scheduling techniques for gpu architectures with processing-in-memory capabilities. In PACT."},{"key":"e_1_3_2_131_2","volume-title":"HCS","author":"Pawlowski J. T.","unstructured":"J. T. Pawlowski. 2011. Hybrid memory cube (HMC). In HCS."},{"key":"e_1_3_2_132_2","volume-title":"VLDB","author":"Pohl Constantin","unstructured":"Constantin Pohl, Kai-Uwe Sattler, and Goetz Graefe. 2019. Joins on high-bandwidth memory: A new level in the memory hierarchy. In VLDB."},{"key":"e_1_3_2_133_2","volume-title":"ISPASS","author":"Pugsley Seth H.","unstructured":"Seth H. Pugsley, Jeffrey Jestes, Huihui Zhang, Rajeev Balasubramonian, Vijayalakshmi Srinivasan, Alper Buyuktosunoglu, Al Davis, and Feifei Li. 2014. NDC: Analyzing the impact of 3D-Stacked memory+logic devices on mapreduce workloads. In ISPASS."},{"key":"e_1_3_2_134_2","volume-title":"H2RC","author":"Rojek Krzysztof","unstructured":"Krzysztof Rojek et al. 2019. CFD Acceleration with FPGA. In H2RC."},{"key":"e_1_3_2_135_2","volume-title":"IEEE Micro","author":"Sadasivam Satish Kumar","unstructured":"Satish Kumar Sadasivam, Brian W. Thompto, Ron Kalla, and William J. Starke. 2017. IBM POWER9 processor architecture. In IEEE Micro."},{"key":"e_1_3_2_136_2","volume-title":"TPDS","author":"Sano Kentaro","unstructured":"Kentaro Sano, Yoshiaki Hatsuda, and Satoru Yamamoto. 2014. Multi-FPGA accelerator for scalable stencil computation with constant memory bandwidth. In TPDS."},{"key":"e_1_3_2_137_2","volume-title":"DATE","author":"Santos Paulo C.","unstructured":"Paulo C. Santos, Geraldo F. Oliveira, Diego G. Tom\u00e9, Marco A. Z. Alves, Eduardo C. Almeida, and Luigi Carro. 2017. Operand size reconfiguration for big data processing in memory. In DATE."},{"key":"e_1_3_2_138_2","volume-title":"BAMS","author":"Sch\u00e4r Christoph","unstructured":"Christoph Sch\u00e4r, Oliver Fuhrer, Andrea Arteaga, Nikolina Ban, Christophe Charpilloz, Salvatore Di Girolamo, Laureline Hentgen, Torsten Hoefler, Xavier Lapillonne, David Leutwyler, Katherine Osterried, Davide Panosetti, Stefan Rudishli, Linda Schlemmer, Thomas C. Schulthess, Michael Sprenger, Stefano Ubbiali, and Heini Wernli. 2020. Kilometer-scale climate models: Prospects and challenges. In BAMS."},{"key":"e_1_3_2_139_2","volume-title":"CAL","author":"Seshadri Vivek","unstructured":"Vivek Seshadri, Kevin Hsieh, Amirali Boroum, Donghyuk Lee, Michael A. Kozuch, Onur Mutlu, Phillip B. Gibbons, and Todd C. Mowry. 2015. Fast bulk bitwise AND and OR in DRAM. In CAL."},{"key":"e_1_3_2_140_2","volume-title":"MICRO","author":"Seshadri Vivek","unstructured":"Vivek Seshadri, Yoongu Kim, Chris Fallin, Donghyuk Lee, Rachata Ausavarungnirun, Gennady Pekhimenko, Yixin Luo, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, et al. 2013. RowClone: Fast and energy-efficient In-DRAM bulk data copy and initialization. In MICRO."},{"key":"e_1_3_2_141_2","volume-title":"MICRO","author":"Seshadri Vivek","unstructured":"Vivek Seshadri, Donghyuk Lee, Thomas Mullins, Hasan Hassan, Amirali Boroumand, Jeremie Kim, Michael A. Kozuch, Onur Mutlu, Phillip B. Gibbons, and Todd C. Mowry. 2017. Ambit: In-Memory accelerator for bulk bitwise operations using commodity DRAM technology. In MICRO."},{"key":"e_1_3_2_142_2","unstructured":"Vivek Seshadri Donghyuk Lee Thomas Mullins Hasan Hassan Amirali Boroumand Jeremie Kim Michael A. Kozuch Onur Mutlu Phillip B. Gibbons and Todd C. Mowry. 2016. Buddy-RAM: Improving the performance and efficiency of bulk bitwise operations using DRAM (unpublished)."},{"key":"e_1_3_2_143_2","volume-title":"MICRO","author":"Seshadri Vivek","unstructured":"Vivek Seshadri, Thomas Mullins, Amirali Boroumand, Onur Mutlu, Phillip B. Gibbons, Michael A. Kozuch, and Todd C. Mowry. 2015. Gather-Scatter DRAM: In-DRAM address translation to improve the spatial locality of non-unit strided accesses. In MICRO."},{"key":"e_1_3_2_144_2","doi-asserted-by":"publisher","DOI":"10.1016\/bs.adcom.2017.04.004"},{"key":"e_1_3_2_145_2","unstructured":"Vivek Seshadri and Onur Mutlu. 2019. In-DRAM bulk bitwise execution engine. arxiv."},{"key":"e_1_3_2_146_2","volume-title":"CXL Consortium White Paper","author":"Sharma D. D.","year":"2019","unstructured":"D. D. Sharma. Compute express link. In CXL Consortium White Paper2019."},{"key":"e_1_3_2_147_2","volume-title":"TC","author":"Simon William Andrew","unstructured":"William Andrew Simon, Yasir Mahmood Qureshi, Marco Rios, Alexandre Levisse, Marina Zapater, and David Atienza. 2020. BLADE: An in-cache computing architecture for edge devices. In TC."},{"key":"e_1_3_2_148_2","volume-title":"DAC","author":"Singh Gagandeep","unstructured":"Gagandeep Singh et al. 2019. NAPEL: Near-memory computing application performance prediction via ensemble learning. In DAC."},{"key":"e_1_3_2_149_2","volume-title":"IEEE Micro","author":"Singh Gagandeep","unstructured":"Gagandeep Singh, Mohammed Alser, Damla Senol Cali, Dionysios Diamantopoulos, Juan G\u00f3mez-Luna, Henk Corporaal, and Onur Mutlu. 2021. FPGA-based near-memory acceleration of modern data-intensive applications. In IEEE Micro."},{"key":"e_1_3_2_150_2","volume-title":"MicPro","author":"Singh Gagandeep","unstructured":"Gagandeep Singh, Lorenzo Chelini, Stefano Corda, Ahsan Javed Awan, Sander Stuijk, Roel Jordans, Henk Corporaal, and Albert-Jan Boonstra. 2019. Near-Memory computing: Past, present, and future. In MicPro, Vol. 71. 102868."},{"key":"e_1_3_2_151_2","volume-title":"DSD","author":"Singh Gagandeep","unstructured":"Gagandeep Singh, Lorenzo Chelini, Stefano Corda, Ahsan Javed Awan, Sander Stuijk, Roel Jordans, Henk Corporaal, and Albert-Jan Boonstra. 2018. A review of near-memory computing architectures: Opportunities and challenges. In DSD."},{"key":"e_1_3_2_152_2","volume-title":"FPGA","author":"Singh Gagandeep","unstructured":"Gagandeep Singh, Dionysios Diamantopolous, Juan G\u00f3mez-Luna, Sander Stuijk, Onur Mutlu, and Henk Corporaal. 2021. Modeling FPGA-based systems via few-shot learning. In FPGA."},{"key":"e_1_3_2_153_2","volume-title":"FPL","author":"Singh Gagandeep","unstructured":"Gagandeep Singh, Dionysios Diamantopoulos, Christoph Hagleitner, Juan G\u00f3mez-Luna, Sander Stuijk, Onur Mutlu, and Henk Corporaal. 2020. NERO: A near high-bandwidth memory stencil accelerator for weather prediction modeling. In FPL."},{"key":"e_1_3_2_154_2","volume-title":"FPL","author":"Singh Gagandeep","unstructured":"Gagandeep Singh, Dionysios Diamantopoulos, Christoph Hagleitner, Sander Stuijk, and Henk Corporaal. 2019. NARMADA: Near-memory horizontal diffusion accelerator for scalable stencil computations. In FPL."},{"key":"e_1_3_2_155_2","volume-title":"Springer LNCS","author":"Singh Gagandeep","unstructured":"Gagandeep Singh, Dionysios Diamantopoulos, Sander Stuijk, Christoph Hagleitner, and Henk Corporaal. 2019. Low precision processing for high order stencil computations. In Springer LNCS."},{"key":"e_1_3_2_156_2","volume-title":"ICS","author":"Strzodka Robert","unstructured":"Robert Strzodka, Mohammed Shaheen, Dawid Pajak, and Hans-Peter Seidel. 2010. Cache oblivious parallelograms in iterative stencil computations. In ICS."},{"key":"e_1_3_2_157_2","volume-title":"IBM JRD","author":"Stuecheli Jeffrey","unstructured":"Jeffrey Stuecheli et al. 2018. IBM POWER9 opens up a new era of acceleration enablement: OpenCAPI. IBM JRD 62, 4\/5 (2018), 8\u20131."},{"key":"e_1_3_2_158_2","volume-title":"IBM JRD","author":"Stuecheli Jeffrey","unstructured":"Jeffrey Stuecheli, Bart Blaner, C. R. Johns, and M. S. Siegel. 2015. CAPI: A coherent accelerator processor interface. IBM JRD 59, 1 (2015), 7\u20131."},{"key":"e_1_3_2_159_2","volume-title":"MICRO","author":"Sukhwani B.","unstructured":"B. Sukhwani, T. Roewer, C. L. Haymes, K. Kim, A. J. McPadden, D. M. Dreps, D. Sanner, J. V. Lunteren, and S. Asaad. 2017. ConTutto\u2014A novel FPGA-based prototyping platform enabling innovation in the memory subsystem of a server class processor. In MICRO."},{"key":"e_1_3_2_160_2","volume-title":"PPAM","author":"Szustak Lukasz","unstructured":"Lukasz Szustak, Krzysztof Rojek, and Pawel Gepner. 2013. Using Intel Xeon Phi coprocessor to accelerate computations in MPDATA algorithm. In PPAM."},{"key":"e_1_3_2_161_2","volume-title":"SPAA","author":"Tang Yuan","unstructured":"Yuan Tang, Rezaul Alam Chowdhury, Bradley C. Kuszmaul, Chi-Keung Luk, and Charles E. Leiserson. 2011. The Pochoir stencil compiler. In SPAA."},{"key":"e_1_3_2_162_2","volume-title":"PASC","author":"Thaler Felix","unstructured":"Felix Thaler, Stefan Moosbrugger, Carlos Osuna, Mauro Bianco, Hannes Vogt, Anton Afanasyev, Lukas Mosimann, Oliver Fuhrer, Thomas C. Schulthess, and Torsten Hoefler. 2019. Porting the COSMO weather model to Manycore CPUs. In PASC."},{"key":"e_1_3_2_163_2","volume-title":"Watson Sci. Comput. Lab. Report, Columbia University","author":"Thomas Llewellyn","unstructured":"Llewellyn Thomas. 1949. Elliptic problems in linear differential equations over a network. In Watson Sci. Comput. Lab. Report, Columbia University."},{"key":"e_1_3_2_164_2","volume-title":"ISCA","author":"Tsai Po-An","unstructured":"Po-An Tsai et al. 2017. Jenga: Software-defined cache hierarchies. In ISCA."},{"key":"e_1_3_2_165_2","volume-title":"ISCA","author":"Tullsen Dean M.","unstructured":"Dean M. Tullsen, Susan J. Eggers, and Henry M. Levy. 1995. Simultaneous multithreading: Maximizing on-chip parallelism. In ISCA."},{"key":"e_1_3_2_166_2","volume-title":"DATE","author":"Lunteren Jan van","unstructured":"Jan van Lunteren, Ronald Luijten, Dionysios Diamantopoulos, Florian Auernhammer, Christoph Hagleitner, Lorenzo Chelini, Stefano Corda, and Gagandeep Singh. 2019. Coherently attached programmable near-memory acceleration platform and its application to stencil processing. In DATE."},{"key":"e_1_3_2_167_2","volume-title":"SC","author":"Volkov Vasily","unstructured":"Vasily Volkov and James W. Demmel. 2008. Benchmarking GPUs to tune dense linear algebra. In SC."},{"key":"e_1_3_2_168_2","volume-title":"SC","author":"Wahib Mohamed","unstructured":"Mohamed Wahib and Naoya Maruyama. 2014. Scalable kernel fusion for memory-bound GPU applications. In SC."},{"key":"e_1_3_2_169_2","volume-title":"IEEE Access","author":"Waidyasooriya Hasitha Muthumala","unstructured":"Hasitha Muthumala Waidyasooriya and Masanori Hariyama. 2019. Multi-FPGA accelerator architecture for stencil computation exploiting spacial and temporal scalability. IEEE Access 7 (2019), 53188\u201353201."},{"key":"e_1_3_2_170_2","volume-title":"TPDS","author":"Waidyasooriya H. M.","unstructured":"H. M. Waidyasooriya, Y. Takei, S. Tatsumi, and M. Hariyama. 2017. OpenCL-based FPGA-platform for stencil computation and its optimization methodology. In TPDS."},{"key":"e_1_3_2_171_2","volume-title":"DAC","author":"Wang Shuo","unstructured":"Shuo Wang and Yun Liang. 2017. A comprehensive framework for synthesizing stencil algorithms on FPGAs using OpenCL model. In DAC."},{"key":"e_1_3_2_172_2","volume-title":"FCCM","author":"Wang Zeke","unstructured":"Zeke Wang, Hongjing Huang, Jie Zhang, and Gustavo Alonso. 2020. Shuhai: Benchmarking high bandwidth memory on FPGAs. In FCCM."},{"key":"e_1_3_2_173_2","volume-title":"Euro-Par","author":"Wenzel Lukas","unstructured":"Lukas Wenzel, Robert Schmid, Balthasar Martin, Max Plauth, Felix Eberhardt, and Andreas Polze. 2018. Getting started with CAPI SNAP: Hardware development for software engineers. In Euro-Par."},{"key":"e_1_3_2_174_2","volume-title":"CACM","author":"Williams Samuel","unstructured":"Samuel Williams, Andrew Waterman, and David Patterson. 2009. Roofline: An insightful visual performance model for multicore architectures. In CACM."},{"key":"e_1_3_2_175_2","volume-title":"ISCA","author":"Wu Lingxi","unstructured":"Lingxi Wu, Rasool Sharifi, Marzieh Lenjani, Kevin Skadron, and Ashish Venkat. 2021. Sieve: Scalable In-situ DRAM-based accelerator designs for massively parallel k-mer matching. In ISCA."},{"key":"e_1_3_2_176_2","volume-title":"ACM TACO","author":"Xu Jingheng","unstructured":"Jingheng Xu, Haohuan Fu, Wen Shi, Lin Gan, Yuxuan Li, Wayne Luk, and Guangwen Yang. 2018. Performance tuning and analysis for stencil-based applications on POWER8 processor. ACM TACO 15, 4 (2018), 1\u201325."},{"key":"e_1_3_2_177_2","volume-title":"HPDC","author":"Zhang Dongping","unstructured":"Dongping Zhang, Nuwan Jayasena, Alexander Lyashevsky, Joseph L. Greathouse, Lifan Xu, and Michael Ignatowski. 2014. TOP-PIM: Throughput-oriented programmable processing in memory. In HPDC."},{"key":"e_1_3_2_178_2","volume-title":"Weather and Forecasting","author":"Zhang Jun A.","unstructured":"Jun A. Zhang, Frank D. Marks, Jason A. Sippel, Robert F. Rogers, Xuejin Zhang, Sundararaman G. Gopalakrishnan, Zhan Zhang, and Vijay Tallapragada. 2018. Evaluating the impact of improvement in the horizontal diffusion parameterization on hurricane prediction in the operational hurricane weather research and forecast (HWRF) model. In Weather and Forecasting."},{"key":"e_1_3_2_179_2","volume-title":"VLSI","author":"Zhu Maohua","unstructured":"Maohua Zhu, Youwei Zhuo, Chao Wang, Wenguang Chen, and Yuan Xie. 2018. Performance evaluation and optimization of HBM-enabled GPU for data-intensive applications. In VLSI."},{"key":"e_1_3_2_180_2","volume-title":"FPGA","author":"Zohouri Hamid Reza","unstructured":"Hamid Reza Zohouri, Artur Podobas, and Satoshi Matsuoka. 2018. Combined spatial and temporal blocking for high-performance stencil computation on FPGAs using OpenCL. In FPGA."}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3501804","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3501804","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:09:48Z","timestamp":1750183788000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3501804"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,6,6]]},"references-count":179,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2022,12,31]]}},"alternative-id":["10.1145\/3501804"],"URL":"https:\/\/doi.org\/10.1145\/3501804","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,6,6]]},"assertion":[{"value":"2021-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-11-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-06-06","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}