{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T22:45:09Z","timestamp":1781131509964,"version":"3.54.1"},"reference-count":150,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2024,12,10]],"date-time":"2024-12-10T00:00:00Z","timestamp":1733788800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Meas. Anal. Comput. Syst."],"published-print":{"date-parts":[[2024,12,10]]},"abstract":"<jats:p>Graph Neural Networks (GNNs) are emerging models to analyze graph-structure data. GNN execution involves both compute-intensive and memory-intensive kernels. The latter kernels dominate execution time, because they are significantly bottlenecked by data movement between memory and processors. Processing-In-Memory (PIM) systems can alleviate this data movement bottleneck by placing simple processors near or inside to memory arrays. This work investigates the potential of PIM systems to alleviate the data movement bottleneck in GNNs, and introduces PyGim, an efficient and easy-to-use GNN library for real PIM systems. We propose intelligent parallelization techniques for memory-intensive kernels of GNNs tailored for real PIM systems, and develop an easy-to-use Python API for them. PyGim employs a cooperative GNN execution, in which the compute- and memory-intensive kernels are executed in processor-centric and memory-centric computing systems, respectively, to fully exploit the hardware capabilities. PyGim integrates a lightweight autotuner to tune the parallelization strategy of the memory-intensive kernel of GNNs and enable high programming ease. We extensively evaluate PyGim on a real-world PIM system that has 16 PIM DIMMs with 1992 PIM cores connected to a Host CPU. In GNN inference, we demonstrate that it outperforms prior state-of-the-art PIM works by on average 4.38\u00d7 (up to 7.20\u00d7), and state-of-the-art PyTorch running on Host by on average 3.04\u00d7 (up to 3.44\u00d7). PyGim improves energy efficiency by 2.86\u00d7 (up to 3.68\u00d7) and 1.55\u00d7 (up to 1.75\u00d7) over prior PIM and PyTorch Host schemes, respectively. In memory-intensive kernel of GNNs, PyGim provides 11.6\u00d7 higher resource utilization in PIM system than that of PyTorch library (optimized CUDA implementation) in GPU systems. Our work provides useful recommendations for software, system and hardware designers. PyGim is publicly and freely available at https:\/\/github.com\/CMU-SAFARI\/PyGim facilitate the widespread use of PIM systems in GNNs.<\/jats:p>","DOI":"10.1145\/3700434","type":"journal-article","created":{"date-parts":[[2024,12,13]],"date-time":"2024-12-13T12:12:12Z","timestamp":1734091932000},"page":"1-36","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["PyGim : An Efficient Graph Neural Network Library for Real Processing-In-Memory Architectures"],"prefix":"10.1145","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0162-4547","authenticated-orcid":false,"given":"Christina","family":"Giannoula","sequence":"first","affiliation":[{"name":"University of Toronto, ETH Z\u00fcrich, Vector Institute, &amp; CentML, Toronto, ON, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0899-1592","authenticated-orcid":false,"given":"Peiming","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Toronto, Toronto, ON, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6133-5670","authenticated-orcid":false,"given":"Ivan","family":"Fernandez","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center &amp; Universitat Polit\u00e8cnica de Catalunya, Barcelona, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-9581-9088","authenticated-orcid":false,"given":"Jiacheng","family":"Yang","sequence":"additional","affiliation":[{"name":"University of Toronto &amp; Vector Institute, Toronto, ON, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4899-4994","authenticated-orcid":false,"given":"Sankeerth","family":"Durvasula","sequence":"additional","affiliation":[{"name":"University of Toronto &amp; Vector Institute, Toronto, ON, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5754-0501","authenticated-orcid":false,"given":"Yu Xin","family":"Li","sequence":"additional","affiliation":[{"name":"University of Toronto, Toronto, ON, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4029-0175","authenticated-orcid":false,"given":"Mohammad","family":"Sadrosadati","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Z\u00fcrich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6514-1571","authenticated-orcid":false,"given":"Juan Gomez","family":"Luna","sequence":"additional","affiliation":[{"name":"NVIDIA, Z\u00fcrich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0075-2312","authenticated-orcid":false,"given":"Onur","family":"Mutlu","sequence":"additional","affiliation":[{"name":"ETH Zurich, Z\u00fcrich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3839-0919","authenticated-orcid":false,"given":"Gennady","family":"Pekhimenko","sequence":"additional","affiliation":[{"name":"University of Toronto, Vector Institute, &amp; CentML, Toronto, ON, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,12,13]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Tensorflow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. arXiv","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al. 2016. Tensorflow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. arXiv (2016)."},{"key":"e_1_2_1_2_1","unstructured":"Junwhan Ahn Sungpack Hong Sungjoo Yoo Onur Mutlu and Kiyoung Choi. 2015. A Scalable Processing-In-Memory Accelerator for Parallel Graph Processing. In ISCA."},{"key":"e_1_2_1_3_1","volume-title":"Exploiting Locality in Sparse Matrix-Matrix Multiplication on Many-Core Architectures. TPDS","author":"Akbudak Kadir","year":"2017","unstructured":"Kadir Akbudak and Cevdet Aykanat. 2017. Exploiting Locality in Sparse Matrix-Matrix Multiplication on Many-Core Architectures. TPDS (2017)."},{"key":"e_1_2_1_4_1","volume-title":"Jung Ho Ahn, and Nam Sung Kim.","author":"Asghari-Moghaddam Hadi","year":"2016","unstructured":"Hadi Asghari-Moghaddam, Young Hoon Son, Jung Ho Ahn, and Nam Sung Kim. 2016. Chameleon: Versatile and Practical Near-DRAM Acceleration Architecture for Large Memory Systems. In MICRO."},{"key":"e_1_2_1_5_1","doi-asserted-by":"crossref","unstructured":"Adam Auten Matthew Tomei and Rakesh Kumar. 2020. Hardware Acceleration of Graph Neural Networks. In DAC.","DOI":"10.1109\/DAC18072.2020.9218751"},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Daehyeon Baek Soojin Hwang Taekyung Heo Daehoon Kim and Jaehyuk Huh. 2021. InnerSP: A Memory Efficient Sparse Matrix Multiplication Accelerator With Locality-Aware Inner Product Processing. In PACT.","DOI":"10.1109\/PACT52795.2021.00016"},{"key":"e_1_2_1_7_1","unstructured":"Riyadh Baghdadi Massinissa Merouani Mohamed-Hicham Leghettas Kamel Abdous Taha Arbaoui Karima Benatchba et al. 2021. A Deep Learning Based Cost Model for Automatic Code Optimization. MLSys (2021)."},{"key":"e_1_2_1_8_1","volume":"202","author":"Bharadwaj V.","unstructured":"V. Bharadwaj, A. Buluc, and J. Demmel. 2022. Distributed-Memory Sparse Kernels for Machine Learning. In IPDPS.","journal-title":"J. Demmel."},{"key":"e_1_2_1_9_1","volume-title":"Neural Networks for Pattern Recognition","author":"Bishop Christopher M","unstructured":"Christopher M Bishop. 1995. Neural Networks for Pattern Recognition. Oxford University Press."},{"key":"e_1_2_1_10_1","doi-asserted-by":"crossref","unstructured":"\u00c5ke Bj\u00f6rck. 1996. Numerical Methods for Least Squares Problems. In SIAM.","DOI":"10.1137\/1.9781611971484"},{"key":"e_1_2_1_11_1","doi-asserted-by":"crossref","unstructured":"Charles Block Gerasimos Gerogiannis Charith Mendis Ariful Azad and Josep Torrellas. 2024. Two-Face: Combining Collective and One-Sided Communication for Efficient Distributed SpMM. In ASPLOS.","DOI":"10.1145\/3620665.3640427"},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Amirali Boroumand Saugata Ghose Youngsok Kim Rachata Ausavarungnirun Eric Shiu Rahul Thakur Daehyun Kim Aki Kuusela Allan Knies Parthasarathy Ranganathan and Onur Mutlu. 2018. Google Workloads for Consumer Devices: Mitigating Data Movement Bottlenecks. In ASPLOS.","DOI":"10.1145\/3173162.3173177"},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Dan Chen Haiheng He Hai Jin Long Zheng Yu Huang Xinyang Shen and Xiaofei Liao. 2023. MetaNMP: Leveraging Cartesian-Like Product to Accelerate HGNNs with Near-Memory Processing. In ISCA.","DOI":"10.1145\/3579371.3589091"},{"key":"e_1_2_1_14_1","volume-title":"Yuxin Guo, and Onur Mutlu.","author":"Chen Jinfan","year":"2023","unstructured":"Jinfan Chen, Juan G\u00f3mez-Luna, Izzat El Hajj, Yuxin Guo, and Onur Mutlu. 2023. SimplePIM: A Software Framework for Productive and Efficient Processing-in-Memory. In PACT."},{"key":"e_1_2_1_15_1","volume-title":"Mxnet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. arXiv","author":"Chen Tianqi","year":"2015","unstructured":"Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang. 2015. Mxnet: A Flexible and Efficient Machine Learning Library for Heterogeneous Distributed Systems. arXiv (2015)."},{"key":"e_1_2_1_16_1","volume-title":"Learning to Optimize Tensor Programs. NIPS","author":"Chen Tianqi","year":"2018","unstructured":"Tianqi Chen, Lianmin Zheng, Eddie Yan, Ziheng Jiang, Thierry Moreau, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy. 2018. Learning to Optimize Tensor Programs. NIPS (2018)."},{"key":"e_1_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Wei-Lin Chiang Xuanqing Liu Si Si Yang Li Samy Bengio and Cho-Jui Hsieh. 2019. Cluster-GCN: An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks. In SIGKDD.","DOI":"10.1145\/3292500.3330925"},{"key":"e_1_2_1_18_1","unstructured":"Benjamin Y. Cho Yongkee Kwon Sangkug Lym and Mattan Erez. 2020. Near Data Acceleration with Concurrent Host Access. In ISCA."},{"key":"e_1_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Jiwon Choe Amy Huang Tali Moreshet Maurice Herlihy and R. Iris Bahar. 2019. Concurrent Data Structures with Near-Data-Processing: An Architecture-Aware Implementation. In SPAA.","DOI":"10.1145\/3323165.3323191"},{"key":"e_1_2_1_20_1","unstructured":"Ctranslate2. 2023. Ctranslate2. https:\/\/github.com\/OpenNMT\/CTranslate2"},{"key":"e_1_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Leonardo Dagum and Ramesh Menon. 1998. OpenMP: An Industry-Standard API for Shared-Memory Programming. In IEEE Comput. Sci. Eng.","DOI":"10.1109\/99.660313"},{"key":"e_1_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Steven Dalton Luke Olson and Nathan Bell. 2015. Optimizing Sparse Matrix-Matrix Multiplication for the GPU. ACM Trans. Math. Softw. (2015).","DOI":"10.1145\/2699470"},{"key":"e_1_2_1_23_1","volume-title":"Mark Indovina, Sai Manoj Pudukotai Dinakarrao, and Amlan Ganguly.","author":"Das Prangon","year":"2022","unstructured":"Prangon Das, Purab Ranjan Sutradhar, Mark Indovina, Sai Manoj Pudukotai Dinakarrao, and Amlan Ganguly. 2022. Implementation and Evaluation of Deep Neural Networks in Commercially Available Processing in Memory Hardware. In SOCC."},{"key":"e_1_2_1_24_1","volume-title":"Davis and Yifan Hu","author":"Timothy","year":"2011","unstructured":"Timothy A. Davis and Yifan Hu. 2011. The University of Florida Sparse Matrix Collection. In TOMS."},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","unstructured":"F. Devaux. 2019. The True Processing In Memory Accelerator. In Hot Chips.","DOI":"10.1109\/HOTCHIPS.2019.8875680"},{"key":"e_1_2_1_26_1","volume-title":"Onur Mutlu, and Izzat El Hajj.","author":"Diab Safaa","year":"2023","unstructured":"Safaa Diab, Amir Nassereldine, Mohammed Alser, Juan G\u00f3mez Luna, Onur Mutlu, and Izzat El Hajj. 2023. A Framework for High-Throughput Sequence Alignment Using Real Processing-in-Memory Systems. Bioinformatics (2023)."},{"key":"e_1_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Mario Drumond Alexandros Daglis Nooshin Mirzadeh Dmitrii Ustiugov Javier Picorel Babak Falsafi Boris Grot and Dionisios Pnevmatikatos. 2017. The Mondrian Data Engine. In ISCA.","DOI":"10.1145\/3079856.3080233"},{"key":"e_1_2_1_28_1","volume-title":"Energy Efficiency Impact of Processing in Memory: A Comprehensive Review of Workloads on the UPMEM Architecture","author":"Falevoz Yann","unstructured":"Yann Falevoz and Julien Legriel. 2023. Energy Efficiency Impact of Processing in Memory: A Comprehensive Review of Workloads on the UPMEM Architecture. In Euro-PAR. Springer."},{"key":"e_1_2_1_29_1","volume-title":"Graph Neural Networks for Social Recommendation. In The World Wide Web Conference.","author":"Fan Wenqi","year":"2019","unstructured":"Wenqi Fan, Yao Ma, Qing Li, Yuan He, Eric Zhao, Jiliang Tang, and Dawei Yin. 2019. Graph Neural Networks for Social Recommendation. In The World Wide Web Conference."},{"key":"e_1_2_1_30_1","volume-title":"Juan G\u00f3mez-Luna, Eladio Gutierrez, Oscar Plata, and Onur Mutlu.","author":"Fernandez Ivan","year":"2024","unstructured":"Ivan Fernandez, Christina Giannoula, Aditya Manglik, Ricardo Quislant, Nika Mansouri Ghiasi, Juan G\u00f3mez-Luna, Eladio Gutierrez, Oscar Plata, and Onur Mutlu. 2024. MATSA: An MRAM-Based Energy-Efficient Accelerator for Time Series Analysis. IEEE Access (2024)."},{"key":"e_1_2_1_31_1","volume-title":"NATSA: A Near-Data Processing Accelerator for Time Series Analysis. In ICCD.","author":"Fernandez Ivan","year":"2020","unstructured":"Ivan Fernandez, Ricardo Quislant, Christina Giannoula, Mohammed Alser, Juan G\u00f3mez-Luna, Eladio Guti\u00e9rrez, Oscar Plata, and Onur Mutlu. 2020. NATSA: A Near-Data Processing Accelerator for Time Series Analysis. In ICCD."},{"key":"e_1_2_1_32_1","volume-title":"Lenssen","author":"Fey Matthias","year":"2019","unstructured":"Matthias Fey and Jan E. Lenssen. 2019. Fast Graph Representation Learning with PyTorch Geometric. In ICLR."},{"key":"e_1_2_1_33_1","unstructured":"Trevor Gale Matei Zaharia Cliff Young and Erich Elsen. [n. d.]. Sparse GPU Kernels for Deep Learning. In SC."},{"key":"e_1_2_1_34_1","unstructured":"Mingyu Gao Grant Ayers and Christos Kozyrakis. 2015. Practical Near-Data Processing for In-Memory Analytics Frameworks. In PACT."},{"key":"e_1_2_1_35_1","volume-title":"Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network Training. In ATC.","author":"Geoffrey X Yu","year":"2021","unstructured":"X Yu Geoffrey, Yubo Gao, Pavel Golikov, and Gennady Pekhimenko. 2021. Habitat: A Runtime-Based Computational Performance Predictor for Deep Neural Network Training. In ATC."},{"key":"e_1_2_1_36_1","volume-title":"SPADE: A Flexible and Scalable Accelerator for SpMM and SDDMM. In ISCA.","author":"Gerogiannis Gerasimos","year":"2023","unstructured":"Gerasimos Gerogiannis, Serif Yesil, Damitha Lenadora, Dingyuan Cao, Charith Mendis, and Josep Torrellas. 2023. SPADE: A Flexible and Scalable Accelerator for SpMM and SDDMM. In ISCA."},{"key":"e_1_2_1_37_1","doi-asserted-by":"crossref","unstructured":"Saugata Ghose Amirali Boroumand Jeremie Kim Juan G\u00f3mez-Luna and Onur Mutlu. 2019. Processing-in-Memory: A Workload-Driven Perspective. In IBM JRD.","DOI":"10.1147\/JRD.2019.2934048"},{"key":"e_1_2_1_38_1","doi-asserted-by":"crossref","unstructured":"Christina Giannoula Ivan Fernandez Juan G\u00f3mez-Luna Nectarios Koziris Georgios Goumas and Onur Mutlu. 2022. Towards Efficient Sparse Matrix Vector Multiplication on Real Processing-In-Memory Architectures. In SIGMETRICS.","DOI":"10.1145\/3489048.3522661"},{"key":"e_1_2_1_39_1","volume-title":"Nectarios Koziris, Georgios Goumas, and Onur Mutlu.","author":"Giannoula Christina","year":"2022","unstructured":"Christina Giannoula, Ivan Fernandez, Juan G\u00f3mez Luna, Nectarios Koziris, Georgios Goumas, and Onur Mutlu. 2022. SparseP: Towards Efficient Sparse Matrix Vector Multiplication on Real Processing-in-Memory Architectures. POMACS (2022)."},{"key":"e_1_2_1_40_1","doi-asserted-by":"crossref","unstructured":"Christina Giannoula Nandita Vijaykumar Nikela Papadopoulou Vasileios Karakostas Ivan Fernandez Juan G\u00f3mez-Luna Lois Orosa Nectarios Koziris Georgios Goumas and Onur Mutlu. 2021. SynCron: Efficient Synchronization Support for Near-Data-Processing Architectures. In HPCA.","DOI":"10.1109\/HPCA51647.2021.00031"},{"key":"e_1_2_1_41_1","volume-title":"Jiacheng Yang, Sankeerth Durvasula, Yu Xin Li, Mohammad Sadrosadati, Juan Gomez Luna, Onur Mutlu, and Gennady Pekhimenko.","author":"Giannoula Christina","year":"2024","unstructured":"Christina Giannoula, Peiming Yang, Ivan Fernandez Vega, Jiacheng Yang, Sankeerth Durvasula, Yu Xin Li, Mohammad Sadrosadati, Juan Gomez Luna, Onur Mutlu, and Gennady Pekhimenko. 2024. Accelerating Graph Neural Networks on Real Processing-In-Memory Systems. https:\/\/arxiv.org\/abs\/2402.16731"},{"key":"e_1_2_1_42_1","doi-asserted-by":"crossref","unstructured":"Maya Gokhale Scott Lloyd and Chris Hajas. 2015. Near Memory Data Structure Rearrangement. In MEMSYS.","DOI":"10.1145\/2818950.2818986"},{"key":"e_1_2_1_43_1","volume-title":"Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu.","author":"G\u00f3mez-Luna Juan","year":"2021","unstructured":"Juan G\u00f3mez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2021. Benchmarking a New Paradigm: An Experimental Analysis of a Real Processing-in-Memory Architecture. In CoRR. https:\/\/arxiv.org\/abs\/2105.03814"},{"key":"e_1_2_1_44_1","volume-title":"Graphite: Optimizing Graph Neural Networks on CPUs through Cooperative Software-Hardware Techniques. In ISCA.","author":"Gong Zhangxiaowen","year":"2022","unstructured":"Zhangxiaowen Gong, Houxiang Ji, Yao Yao, Christopher W. Fletcher, Christopher J. Hughes, and Josep Torrellas. 2022. Graphite: Optimizing Graph Neural Networks on CPUs through Cooperative Software-Hardware Techniques. In ISCA."},{"key":"e_1_2_1_45_1","unstructured":"SAFARI Research Group. 2022. PyGim Software Package. https:\/\/github.com\/Carnegie Mellon University-SAFARI\/PyGim"},{"key":"e_1_2_1_46_1","unstructured":"Zhixiang Gu Jose Moreira David Edelsohn and Ariful Azad. 2020. Bandwidth Optimized Parallel Algorithms for Sparse Matrix-Matrix Multiplication Using Propagation Blocking. In SPAA."},{"key":"e_1_2_1_47_1","volume-title":"Deep Learning with Keras","author":"Gulli Antonio","unstructured":"Antonio Gulli and Sujit Pal. 2017. Deep Learning with Keras. Packt Publishing Ltd."},{"key":"e_1_2_1_48_1","doi-asserted-by":"crossref","unstructured":"Juan G\u00f3mez-Luna Yuxin Guo Sylvan Brocard Julien Legriel Remy Cimadomo Geraldo F. Oliveira Gagandeep Singh and Onur Mutlu. 2023. Evaluating Machine Learning Workloads on Memory-Centric Computing Systems. In ISPASS.","DOI":"10.1109\/ISPASS57527.2023.00013"},{"key":"e_1_2_1_49_1","volume-title":"Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu.","author":"G\u00f3mez-Luna Juan","year":"2022","unstructured":"Juan G\u00f3mez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2022. Benchmarking a New Paradigm: Experimental Analysis and Characterization of a Real Processing-in-Memory System. IEEE Access (2022)."},{"key":"e_1_2_1_50_1","volume-title":"NIPS","volume":"30","author":"Hamilton Will","year":"2017","unstructured":"Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive Representation Learning on Large Graphs. NIPS, Vol. 30 (2017)."},{"key":"e_1_2_1_51_1","volume-title":"Newton: A DRAM-Maker's Accelerator-in-Memory (AiM) Architecture for Machine Learning. In MICRO.","author":"He Mingxuan","year":"2020","unstructured":"Mingxuan He, Choungki Song, Ilkon Kim, Chunseok Jeong, Seho Kim, Il Park, Mithuna Thottethodi, and T. N. Vijaykumar. 2020. Newton: A DRAM-Maker's Accelerator-in-Memory (AiM) Architecture for Machine Learning. In MICRO."},{"key":"e_1_2_1_52_1","unstructured":"Guseul Heo Sangyeop Lee Jaehong Cho Hyunmin Choi Sanghyeon Lee Hyungkyu Ham Gwangsun Kim Divya Mahajan and Jongse Park. 2024. NeuPIMs: NPU-PIM Heterogeneous Acceleration for Batched LLM Inferencing. In ASPLOS."},{"key":"e_1_2_1_53_1","doi-asserted-by":"crossref","unstructured":"Changwan Hong Aravind Sukumaran-Rajam Israt Nisa Kunal Singh and P. Sadayappan. 2019. Adaptive Sparse Tiling for Sparse Matrix Multiplication. In PpopP.","DOI":"10.1145\/3293883.3295712"},{"key":"e_1_2_1_54_1","doi-asserted-by":"crossref","unstructured":"Kevin Hsieh Samira Khan Nandita Vijaykumar Kevin Chang Amirali Boroumand Saugata Ghose and Onur Mutlu. 2016. Accelerating Pointer Chasing in 3D-stacked Memory: Challenges Mechanisms Evaluation. In ICCD.","DOI":"10.1109\/ICCD.2016.7753257"},{"key":"e_1_2_1_55_1","doi-asserted-by":"crossref","unstructured":"Kezhao Huang Jidong Zhai Zhen Zheng Youngmin Yi and Xipeng Shen. 2021. Understanding and Bridging the Gaps in Current GNN Performance Optimizations. In PpopP.","DOI":"10.1145\/3437801.3441585"},{"key":"e_1_2_1_56_1","doi-asserted-by":"crossref","unstructured":"Bongjoon Hyun Taehun Kim Dongjae Lee and Minsoo Rhu. 2024. Pathfinding Future PIM Architectures by Demystifying a Commercial PIM Technology. In HPCA.","DOI":"10.1109\/HPCA57654.2024.00029"},{"key":"e_1_2_1_57_1","volume-title":"Yelick","author":"Im Eun-Jin","year":"1999","unstructured":"Eun-Jin Im and Katherine A. Yelick. 1999. Optimizing Sparse Matrix Vector Multiplication on SMP. In PPSC."},{"key":"e_1_2_1_58_1","doi-asserted-by":"crossref","unstructured":"Maurus Item Geraldo F. Oliveira Juan G\u00f3mez-Luna Mohammad Sadrosadati Yuxin Guo and Onur Mutlu. 2023. TransPimLib: Efficient Transcendental Functions for Processing-in-Memory Systems. In ISPASS.","DOI":"10.1109\/ISPASS57527.2023.00031"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_2_1_60_1","unstructured":"Zhihao Jia Sina Lin Mingyu Gao Matei A. Zaharia and Alexander Aiken. 2020. Improving the Accuracy Scalability and Performance of Graph Neural Networks with Roc. In MLSys."},{"key":"e_1_2_1_61_1","unstructured":"Muhammad Attahir Jibril Hani Al-Sayeh and Kai-Uwe Sattler. 2024. Accelerating Aggregation Using a Real Processing-in-Memory System. In ICDE."},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/3652963.3655079"},{"key":"e_1_2_1_63_1","volume-title":"Taha Shahroodi, Juan Gomez Luna, and Onur Mutlu.","author":"Kanellopoulos Konstantinos","year":"2019","unstructured":"Konstantinos Kanellopoulos, Nandita Vijaykumar, Christina Giannoula, Roknoddin Azizi, Skanda Koppula, Nika Mansouri Ghiasi, Taha Shahroodi, Juan Gomez Luna, and Onur Mutlu. 2019. Smash: Co-Designing Software Compression and Hardware-Accelerated Indexing for Efficient Sparse Matrix Operations. In MICRO."},{"key":"e_1_2_1_64_1","volume-title":"A Learned Performance Model for Tensor Processing Units. MLSys","author":"Kaufman Sam","year":"2021","unstructured":"Sam Kaufman, Phitchaya Phothilimthana, Yanqi Zhou, Charith Mendis, Sudip Roy, Amit Sabne, and Mike Burrows. 2021. A Learned Performance Model for Tensor Processing Units. MLSys (2021)."},{"key":"e_1_2_1_65_1","volume-title":"Mark Hempstead, Brandon Reagen, Xuan Zhang, David Brooks, Vikas Chandra, Utku Diril, et al.","author":"Ke Liu","year":"2020","unstructured":"Liu Ke, Udit Gupta, Carole-Jean Wu, Benjamin Youngjae Cho, Mark Hempstead, Brandon Reagen, Xuan Zhang, David Brooks, Vikas Chandra, Utku Diril, et al. 2020. RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing. In ISCA."},{"key":"e_1_2_1_66_1","unstructured":"Kashif Nizam Khan Mikael Hirki Tapio Niemi Jukka K Nurminen and Zhonghong Ou. 2018. Rapl in Action: Experiences in Using RAPL for Power Measurements. In TOMPECS."},{"key":"e_1_2_1_67_1","volume-title":"GRIP: A Graph Neural Network Accelerator Architecture","author":"Kiningham Kevin","year":"2022","unstructured":"Kevin Kiningham, Philip Levis, and Christopher R\u00e9. 2022. GRIP: A Graph Neural Network Accelerator Architecture. IEEE Trans. Comput. (2022)."},{"key":"e_1_2_1_68_1","volume-title":"Semi-Supervised Classification with Graph Convolutional Networks. arXiv","author":"Kipf Thomas N","year":"2016","unstructured":"Thomas N Kipf and Max Welling. 2016. Semi-Supervised Classification with Graph Convolutional Networks. arXiv (2016)."},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3133901"},{"key":"e_1_2_1_70_1","doi-asserted-by":"crossref","unstructured":"Penporn Koanantakool Ariful Azad Aydin Bulu\u00e7 Dmitriy Morozov Sang-Yun Oh Leonid Oliker and Katherine Yelick. 2016. Communication-Avoiding Parallel Sparse-Dense Matrix-Matrix Multiplication. In IPDPS.","DOI":"10.1109\/IPDPS.2016.117"},{"key":"e_1_2_1_71_1","doi-asserted-by":"crossref","unstructured":"Youngeun Kwon Yunjae Lee and Minsoo Rhu. 2019. TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning. In MICRO.","DOI":"10.1145\/3352460.3358284"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358284"},{"key":"e_1_2_1_73_1","unstructured":"Young-Cheon Kwon Suk Han Lee Jaehoon Lee Sang-Hyuk Kwon Je Min Ryu Jong-Pil Son O Seongil Hak-Soo Yu Haesuk Lee Soo Young Kim Youngmin Cho Jin Guk Kim Jongyoon Choi Hyun-Sung Shin Jin Kim BengSeng Phuah HyoungMin Kim Myeong Jun Song Ahn Choi Daeho Kim SooYoung Kim Eun-Bong Kim David Wang Shinhaeng Kang Yuhwan Ro Seungwoo Seo JoonHo Song Jaeyoun Youn Kyomin Sohn and Nam Sung Kim. 2021. 25.4 A 20nm 6GB Function-In-Memory DRAM Based on HBM2 with a 1.2TFLOPS Programmable Computing Unit Using Bank-Level Parallelism for Machine Learning Applications. In ISSCC."},{"key":"e_1_2_1_74_1","doi-asserted-by":"crossref","unstructured":"Daniel Langr and Pavel Tvrd\u00edk. 2016. Evaluation Criteria for Sparse Matrix Storage Formats. In TPDS.","DOI":"10.1109\/TPDS.2015.2401575"},{"key":"e_1_2_1_75_1","unstructured":"Sukhan Lee Shin-haeng Kang Jaehoon Lee Hyeonsu Kim Eojin Lee Seungwoo Seo Hosang Yoon Seungwon Lee Kyounghwan Lim Hyunsung Shin Jinhyun Kim O Seongil Anand Iyer David Wang Kyomin Sohn and Nam Sung Kim. 2021. Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product. In ISCA."},{"key":"e_1_2_1_76_1","unstructured":"Seongju Lee Kyuyoung Kim Sanghoon Oh Joonhong Park Gimoon Hong Dongyoon Ka Kyudong Hwang Jeongje Park Kyeongpil Kang Jungyeon Kim et al. 2022. A 1ynm 1.25 V 8GB 16GB\/s\/Pin GDDR6-Based Accelerator-in-Memory Supporting 1Tflops MAC Operation and Various Activation Functions for Deep-Learning Applications. In ISSCC."},{"key":"e_1_2_1_77_1","unstructured":"Yunjae Lee Jinha Chung and Minsoo Rhu. 2022. SmartSAGE: Training Large-Scale Graph Neural Networks Using In-Storage Processing Architectures. In ISCA."},{"key":"e_1_2_1_78_1","unstructured":"Damitha Lenadora Vimarsh Sathia Gerasimos Gerogiannis Serif Yesil Josep Torrellas and Charith Mendis. 2024. SENSEi: Input-Sensitive Compilation for Accelerating GNNs. In arXiv."},{"key":"e_1_2_1_79_1","volume-title":"GLIST: Towards In-Storage Graph Learning. In ATC.","author":"Li Cangyuan","year":"2021","unstructured":"Cangyuan Li, Ying Wang, Cheng Liu, Shengwen Liang, Huawei Li, and Xiaowei Li. 2021. GLIST: Towards In-Storage Graph Learning. In ATC."},{"key":"e_1_2_1_80_1","unstructured":"Cong Li Zhe Zhou Xingchen Li Guangyu Sun and Dimin Niu. 2023. NMExplorer: An Efficient Exploration Framework for DIMM-Based Near-Memory Tensor Reduction. In DAC."},{"key":"e_1_2_1_81_1","unstructured":"Cong Li Zhe Zhou Yang Wang Fan Yang Ting Cao Mao Yang Yun Liang and Guangyu Sun. 2024. PIM-DL: Expanding the Applicability of Commodity DRAM-PIMs for Deep Learning via Algorithm-System Co-Optimization. In ASPLOS."},{"key":"e_1_2_1_82_1","unstructured":"Cong Li Zhe Zhou Size Zheng Jiaxi Zhang Yun Liang and Guangyu Sun. 2024. SpecPIM: Accelerating Speculative Inference on PIM-Enabled System via Architecture-Dataflow Co-Exploration. In ASPLOS."},{"key":"e_1_2_1_83_1","volume-title":"GCNAX: A Flexible and Energy-Efficient Accelerator for Graph Convolutional Neural Networks. In HPCA.","author":"Li Jiajun","year":"2021","unstructured":"Jiajun Li, Ahmed Louri, Avinash Karanth, and Razvan Bunescu. 2021. GCNAX: A Flexible and Energy-Efficient Accelerator for Graph Convolutional Neural Networks. In HPCA."},{"key":"e_1_2_1_84_1","volume-title":"Engn: A High-Throughput and Energy-Efficient Accelerator for Large Graph Neural Networks","author":"Liang Shengwen","year":"2020","unstructured":"Shengwen Liang, Ying Wang, Cheng Liu, Lei He, LI Huawei, Dawen Xu, and Xiaowei Li. 2020. Engn: A High-Throughput and Energy-Efficient Accelerator for Large Graph Neural Networks. IEEE Trans. Comput. (2020)."},{"key":"e_1_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1145\/3589258"},{"key":"e_1_2_1_86_1","doi-asserted-by":"crossref","unstructured":"Y. Lin and V. Prasanna. 2023. HyScale-GNN: A Scalable Hybrid GNN Training System on Single-Node Heterogeneous Architecture. In IPDPS.","DOI":"10.1109\/IPDPS54959.2023.00062"},{"key":"e_1_2_1_87_1","volume-title":"Dong Li, and Jishen Zhao.","author":"Liu Jiawen","year":"2018","unstructured":"Jiawen Liu, Hengyu Zhao, Matheus Almeida Ogleari, Dong Li, and Jishen Zhao. 2018. Processing-in-Memory for Energy-Efficient Neural Network Training: A Heterogeneous Approach. In MICRO."},{"key":"e_1_2_1_88_1","volume-title":"BGL: GPU-Efficient GNN Training by Optimizing Graph Data I\/O and Preprocessing. In NSDI.","author":"Liu Tianfeng","year":"2023","unstructured":"Tianfeng Liu, Yangrui Chen, Dan Li, Chuan Wu, Yibo Zhu, Jun He, Yanghua Peng, Hongzheng Chen, Hongzhi Chen, and Chuanxiong Guo. 2023. BGL: GPU-Efficient GNN Training by Optimizing Graph Data I\/O and Preprocessing. In NSDI."},{"key":"e_1_2_1_89_1","unstructured":"Weifeng Liu and Brian Vinter. 2014. An Efficient GPU General Sparse Matrix-Matrix Multiplication for Irregular Data. In IPDPS."},{"key":"e_1_2_1_90_1","doi-asserted-by":"crossref","unstructured":"Zhiyu Liu Irina Calciu Maurice Herlihy and Onur Mutlu. 2017. Concurrent Data Structures for Near-Memory Computing. In SPAA.","DOI":"10.1145\/3087556.3087582"},{"key":"e_1_2_1_91_1","unstructured":"Dominik Marek Loroch Norbert Wehn Franz-Josef Pfreundt and Janis Keuper. 2017. TensorQuant - A Simulation Toolbox for Deep Neural Network Quantization. In arXiv."},{"key":"e_1_2_1_92_1","volume-title":"Neugraph: Parallel Deep Neural Network Computation on Large Graphs. In ATC.","author":"Ma Lingxiao","year":"2019","unstructured":"Lingxiao Ma, Zhi Yang, Youshan Miao, Jilong Xue, Ming Wu, Lidong Zhou, and Yafei Dai. 2019. Neugraph: Parallel Deep Neural Network Computation on Large Graphs. In ATC."},{"key":"e_1_2_1_93_1","unstructured":"Vasimuddin Md Sanchit Misra Guixiang Ma Ramanarayan Mohanty Evangelos Georganas Alexander Heinecke Dhiraj Kalamkar Nesreen K. Ahmed and Sasikanth Avancha. 2021. DistGNN: Scalable Distributed Training for Large-Scale Graph Neural Networks. In SC."},{"key":"e_1_2_1_94_1","volume-title":"Discovering Protein Drug Targets Using Knowledge Graph Embeddings. Bioinformatics","author":"Mohamed Sameh K","year":"2020","unstructured":"Sameh K Mohamed, V\u00edt Nov\u00e1vcek, and Aayah Nounu. 2020. Discovering Protein Drug Targets Using Knowledge Graph Embeddings. Bioinformatics (2020)."},{"key":"e_1_2_1_95_1","doi-asserted-by":"crossref","unstructured":"P. Mpakos D. Galanopoulos P. Anastasiadis N. Papadopoulou N. Koziris and G. Goumas. 2023. Feature-Based SpMV Performance Analysis on Contemporary Devices. In IPDPS.","DOI":"10.1109\/IPDPS54959.2023.00072"},{"key":"e_1_2_1_96_1","doi-asserted-by":"crossref","unstructured":"Onur Mutlu Saugata Ghose Juan G\u00f3mez-Luna and Rachata Ausavarungnirun. 2019. Processing Data Where It Makes Sense: Enabling In-Memory Computation. In MICPRO.","DOI":"10.1016\/j.micpro.2019.01.009"},{"key":"e_1_2_1_97_1","doi-asserted-by":"crossref","unstructured":"Onur Mutlu Saugata Ghose Juan G\u00f3mez-Luna and Rachata Ausavarungnirun. 2021. A Modern Primer on Processing in Memory. In Emerging Computing: From Devices to Systems - Looking Beyond Moore and Von Neumann. https:\/\/arxiv.org\/pdf\/2012.03112.pdf","DOI":"10.1007\/978-981-16-7487-7_7"},{"key":"e_1_2_1_98_1","volume-title":"Emerging Computing: From Devices to Systems - Looking Beyond Moore and Von Neumann","author":"Mutlu Onur","year":"2021","unstructured":"Onur Mutlu, Saugata Ghose, Juan G\u00f3mez-Luna, and R. Ausavarungnirun. 2021. A Modern Primer on Processing in Memory. Emerging Computing: From Devices to Systems - Looking Beyond Moore and Von Neumann (2021)."},{"key":"e_1_2_1_99_1","unstructured":"R. Nair S. F. Antao C. Bertolli P. Bose J. R. Brunheroto T. Chen C.-Y. Cher C. H. A. Costa J. Doi C. Evangelinos and et al. 2015. Active Memory Cube: A Processing-in-Memory Architecture for Exascale Systems. In IBM JRD."},{"key":"e_1_2_1_100_1","volume-title":"Cusparse Library. In GPU Technology Conference.","author":"Naumov Maxim","year":"2010","unstructured":"Maxim Naumov, L Chien, Philippe Vandermersch, and Ujval Kapasi. 2010. Cusparse Library. In GPU Technology Conference."},{"key":"e_1_2_1_101_1","doi-asserted-by":"crossref","unstructured":"Yuyao Niu Zhengyang Lu Haonan Ji Shuhui Song Zhou Jin and Weifeng Liu. 2022. TileSpGEMM: A Tiled Algorithm for Parallel Sparse General Matrix-Matrix Multiplication on GPUs. In PpopP.","DOI":"10.1145\/3503221.3508431"},{"key":"e_1_2_1_102_1","volume":"202","author":"Noh S.","unstructured":"S. Noh, J. Hong, C. Lim, S. Park, J. Kim, H. Kim, Y. Kim, and J. Lee. 2024. PID-Comm: A Fast and Flexible Collective Communication Framework for Commodity Processing-in-DIMM Devices. In ISCA.","journal-title":"J. Lee."},{"key":"e_1_2_1_103_1","volume-title":"Outerspace: An Outer Product Based Sparse Matrix Multiplication Accelerator. In HPCA.","author":"Pal Subhankar","year":"2018","unstructured":"Subhankar Pal, Jonathan Beaumont, Dong-Hyeon Park, Aporva Amarnath, Siying Feng, Chaitali Chakrabarti, Hun-Seok Kim, David Blaauw, Trevor Mudge, and Ronald Dreslinski. 2018. Outerspace: An Outer Product Based Sparse Matrix Multiplication Accelerator. In HPCA."},{"key":"e_1_2_1_104_1","volume-title":"Yongsuk Kwon, Nam Sung Kim, and Jung Ho Ahn.","author":"Park Jaehyun","year":"2024","unstructured":"Jaehyun Park, Jaewan Choi, Kwanhee Kyung, Michael Jaemin Kim, Yongsuk Kwon, Nam Sung Kim, and Jung Ho Ahn. 2024. AttAcc! Unleashing the Power of PIM for Batched Transformer-Based Generative Model Inference. In ASPLOS."},{"key":"e_1_2_1_105_1","volume-title":"Proc. VLDB Endow.","author":"Park Yeonhong","year":"2022","unstructured":"Yeonhong Park, Sunhong Min, and Jae W. Lee. 2022. Ginex: SSD-Enabled Billion-Scale Graph Neural Network Training on a Single Machine via Provably Optimal In-Memory Caching. Proc. VLDB Endow. (2022)."},{"key":"e_1_2_1_106_1","volume-title":"Pytorch: An Imperative Style","author":"Paszke Adam","year":"2019","unstructured":"Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. 2019. Pytorch: An Imperative Style, High-Performance Deep Learning Library. NIPS (2019)."},{"key":"e_1_2_1_107_1","unstructured":"PeakPerf. 2021. PeakPerf. https:\/\/github.com\/Dr-Noob\/peakperf.git"},{"key":"e_1_2_1_108_1","volume-title":"Heath","author":"Pinar Ali","year":"1999","unstructured":"Ali Pinar and Michael T. Heath. 1999. Improving Performance of Sparse Matrix-Vector Multiplication. In SC."},{"key":"e_1_2_1_109_1","volume-title":"Pooch and Al Nieder","author":"Udo","year":"1973","unstructured":"Udo W. Pooch and Al Nieder. 1973. A Survey of Indexing Techniques for Sparse Matrices. In ACM Comput. Surv."},{"key":"e_1_2_1_110_1","unstructured":"PyG. 2024. PyG Website. https:\/\/pyg.org\/"},{"key":"e_1_2_1_111_1","doi-asserted-by":"crossref","unstructured":"Guocheng Qian Abdulellah Abualshour Guohao Li Ali Thabet and Bernard Ghanem. 2021. Pu-GCN: Point Cloud Upsampling Using Graph Convolutional Networks. In CVPR.","DOI":"10.1109\/CVPR46437.2021.01151"},{"key":"e_1_2_1_112_1","unstructured":"Zheng Qu Dimin Niu Shuangchen Li Hongzhong Zheng and Yuan Xie. 2023. TT-GNN: Efficient On-Chip Graph Neural Network Training via Embedding Reformation and Hardware Optimization. In MICRO."},{"key":"e_1_2_1_113_1","doi-asserted-by":"crossref","unstructured":"Steve Rhyner Haocong Luo Juan G\u00f3mez-Luna Mohammad Sadrosadati Jiawei Jiang Ataberk Olgun Harshita Gupta Ce Zhang and Onur Mutlu. 2024. Analysis of Distributed Optimization Algorithms on a Real Processing-In-Memory System. In PACT.","DOI":"10.1145\/3656019.3676947"},{"key":"e_1_2_1_114_1","doi-asserted-by":"crossref","unstructured":"Oguz Selvitopi Benjamin Brock Israt Nisa Alok Tripathy Katherine Yelick and Aydin Bulucc. 2021. Distributed-Memory Parallel Algorithms for Sparse Times Tall-Skinny-Dense Matrix Multiplication. In ICS.","DOI":"10.1145\/3447818.3461472"},{"key":"e_1_2_1_115_1","volume-title":"Owens","author":"Sengupta Shubhabrata","year":"2007","unstructured":"Shubhabrata Sengupta, Mark Harris, Yao Zhang, and John D. Owens. 2007. Scan Primitives for GPU Computing. In GH."},{"key":"e_1_2_1_116_1","unstructured":"Chao Shang Jie Chen and Jinbo Bi. 2021. Discrete Graph Structure Learning for Forecasting Multiple Time Series. In ICLR."},{"key":"e_1_2_1_117_1","doi-asserted-by":"crossref","unstructured":"Yongwon Shin Juseong Park Sungjun Cho and Hyojin Sung. 2023. PIMFlow: Compiler and Runtime Support for CNN Models on Processing-in-Memory DRAM. In CGO.","DOI":"10.1145\/3579990.3580009"},{"key":"e_1_2_1_118_1","volume-title":"Sextans: A Streaming Accelerator for General-Purpose Sparse-Matrix Dense-Matrix Multiplication. In SIGDA.","author":"Song Linghao","year":"2022","unstructured":"Linghao Song, Yuze Chi, Atefeh Sohrabizadeh, Young-kyu Choi, Jason Lau, and Jason Cong. 2022. Sextans: A Streaming Accelerator for General-Purpose Sparse-Matrix Dense-Matrix Multiplication. In SIGDA."},{"key":"e_1_2_1_119_1","volume-title":"Cambricon-G: A Polyvalent Energy-Efficient Accelerator for Dynamic Graph Neural Networks. TCAD","author":"Song Xinkai","year":"2021","unstructured":"Xinkai Song, Tian Zhi, Zhe Fan, Zhenxing Zhang, Xi Zeng, Wei Li, Xing Hu, Zidong Du, Qi Guo, and Yunji Chen. 2021. Cambricon-G: A Polyvalent Energy-Efficient Accelerator for Dynamic Graph Neural Networks. TCAD (2021)."},{"key":"e_1_2_1_120_1","volume-title":"Matraptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise Product. In MICRO.","author":"Srivastava Nitish","year":"2020","unstructured":"Nitish Srivastava, Hanchen Jin, Jie Liu, David Albonesi, and Zhiru Zhang. 2020. Matraptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise Product. In MICRO."},{"key":"e_1_2_1_121_1","doi-asserted-by":"crossref","unstructured":"Jacob R Stevens Dipankar Das Sasikanth Avancha Bharat Kaul and Anand Raghunathan. 2021. GNNerator: A Hardware\/Software Framework for Accelerating Graph Neural Networks. In DAC.","DOI":"10.1109\/DAC18074.2021.9586122"},{"key":"e_1_2_1_122_1","doi-asserted-by":"crossref","unstructured":"Jonathan M Stokes Kevin Yang Kyle Swanson Wengong Jin Andres Cubillos-Ruiz Nina M Donghia Craig R MacNair Shawn French Lindsey A Carfrae Zohar Bloom-Ackermann et al. 2020. A Deep Learning Approach to Antibiotic Discovery. Cell (2020).","DOI":"10.1016\/j.cell.2020.04.001"},{"key":"e_1_2_1_123_1","volume-title":"An Adaptive Concurrent Priority Queue for NUMA Architectures. In International Conference on Computing Frontiers.","author":"Strati Foteini","year":"2019","unstructured":"Foteini Strati, Christina Giannoula, Dimitrios Siakavaras, Georgios Goumas, and Nectarios Koziris. 2019. An Adaptive Concurrent Priority Queue for NUMA Architectures. In International Conference on Computing Frontiers."},{"key":"e_1_2_1_124_1","unstructured":"stream. 2021. STREAM. https:\/\/github.com\/jeffhammond\/STREAM.git"},{"key":"e_1_2_1_125_1","doi-asserted-by":"crossref","unstructured":"Damian Szklarczyk Annika L Gable David Lyon Alexander Junge Stefan Wyder Jaime Huerta-Cepas Milan Simonovic Nadezhda T Doncheva John H Morris Peer Bork et al. 2019. STRING v11: Protein--Protein Association Networks with Increased Coverage Supporting Functional Discovery in Genome-Wide Experimental Datasets. Nucleic Acids Research (2019).","DOI":"10.1093\/nar\/gky1131"},{"key":"e_1_2_1_126_1","volume-title":"Xibai Li, and Rick Siow Mong Goh.","author":"Tang Wai Teng","year":"2015","unstructured":"Wai Teng Tang, Ruizhe Zhao, Mian Lu, Yun Liang, Huynh Phung Huyng, Xibai Li, and Rick Siow Mong Goh. 2015. Optimizing and Auto-Tuning Scale-Free Sparse Matrix-Vector Multiplication on Intel Xeon Phi. In CGO."},{"key":"e_1_2_1_127_1","volume-title":"Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads. In OSDI.","author":"Thorpe John","year":"2021","unstructured":"John Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng, Guanzhou Hu, Zhihao Jia, Jinliang Wei, Keval Vora, Ravi Netravali, Miryung Kim, and Guoqing Harry Xu. 2021. Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads. In OSDI."},{"key":"e_1_2_1_128_1","doi-asserted-by":"crossref","unstructured":"Boyu Tian Yiwei Li Li Jiang Shuangyu Cai and Mingyu Gao. 2024. NDPBridge: Enabling Cross-Bank Coordination in Near-DRAM-Bank Processing Architectures. In ISCA.","DOI":"10.1109\/ISCA59077.2024.00052"},{"key":"e_1_2_1_129_1","volume-title":"G-NMP: Accelerating Graph Neural Networks with DIMM-Based Near-Memory Processing. Journal of Systems Architecture","author":"Tian Teng","year":"2022","unstructured":"Teng Tian, Xiaotian Wang, Letian Zhao, Wei Wu, Xuecang Zhang, Fangmin Lu, Tianqi Wang, and Xi Jin. 2022. G-NMP: Accelerating Graph Neural Networks with DIMM-Based Near-Memory Processing. Journal of Systems Architecture (2022)."},{"key":"e_1_2_1_130_1","doi-asserted-by":"crossref","unstructured":"Alok Tripathy Katherine Yelick and Aydin Bulucc. 2020. Reducing Communication in Graph Neural Network Training. In SC.","DOI":"10.1109\/SC41405.2020.00074"},{"key":"e_1_2_1_131_1","unstructured":"UPMEM. 2020. UPMEM Website. https:\/\/www.upmem.com"},{"key":"e_1_2_1_132_1","volume-title":"Robert A Van De Geijn, Francisco D Igual, Mikhail Smelyanskiy, Xianyi Zhang, Michael Kistler, Vernon Austel, et al.","author":"Van Zee Field G","year":"2016","unstructured":"Field G Van Zee, Tyler M Smith, Bryan Marker, Tze Meng Low, Robert A Van De Geijn, Francisco D Igual, Mikhail Smelyanskiy, Xianyi Zhang, Michael Kistler, Vernon Austel, et al. 2016. The BLIS Framework: Experiments in Portability. TOMS (2016)."},{"key":"e_1_2_1_133_1","unstructured":"Petar Velickovic Guillem Cucurull Arantxa Casanova Adriana Romero Pietro Lio Yoshua Bengio et al. 2017. Graph Attention Networks. Stat (2017)."},{"key":"e_1_2_1_134_1","doi-asserted-by":"crossref","unstructured":"Endong Wang Qing Zhang Bo Shen Guangyong Zhang Xiaowei Lu Qing Wu and Yajuan Wang. 2014. Intel Math Kernel Library.","DOI":"10.1007\/978-3-319-06486-4_7"},{"key":"e_1_2_1_135_1","doi-asserted-by":"crossref","unstructured":"Shu Wu Yuyuan Tang Yanqiao Zhu Liang Wang Xing Xie and Tieniu Tan. 2019. Session-Based Recommendation with Graph Neural Networks. In AAAI.","DOI":"10.1609\/aaai.v33i01.3301346"},{"key":"e_1_2_1_136_1","volume-title":"How Powerful Are Graph Neural Networks? arXiv","author":"Xu Keyulu","year":"2018","unstructured":"Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2018. How Powerful Are Graph Neural Networks? arXiv (2018)."},{"key":"e_1_2_1_137_1","volume-title":"Owens","author":"Yang Carl","year":"2018","unstructured":"Carl Yang, Aydin Bulucc, and John D. Owens. 2018. Design Principles for Sparse Matrix Multiplication on the GPU. In Euro-PAR."},{"key":"e_1_2_1_138_1","unstructured":"Zhilin Yang William Cohen and Ruslan Salakhudinov. 2016. Revisiting Semi-Supervised Learning with Graph Embeddings. In ICML."},{"key":"e_1_2_1_139_1","unstructured":"Zihao Ye Ruihang Lai Junru Shao Tianqi Chen and Luis Ceze. 2023. SparseTIR: Composable Abstractions for Sparse Compilation in Deep Learning. In ASPLOS."},{"key":"e_1_2_1_140_1","doi-asserted-by":"crossref","unstructured":"Serif Yesil Jos\u00e9 E Moreira and Josep Torrellas. 2022. Dense Dynamic Blocks: Optimizing SpMM for Processors with Vector and Matrix Units Using Machine Learning Techniques. In ICS.","DOI":"10.1145\/3524059.3532369"},{"key":"e_1_2_1_141_1","doi-asserted-by":"crossref","unstructured":"Rex Ying Ruining He Kaifeng Chen Pong Eksombatchai William L Hamilton and Jure Leskovec. 2018. Graph Convolutional Neural Networks for Web-Scale Recommender Systems. In SIGKDD.","DOI":"10.1145\/3219819.3219890"},{"key":"e_1_2_1_142_1","volume-title":"Jung Ho Ahn, and Eojin Lee","author":"Yun Sungmin","year":"2023","unstructured":"Sungmin Yun, Hwayong Nam, Jaehyun Park, Byeongho Kim, Jung Ho Ahn, and Eojin Lee. 2023. GraNDe: Efficient Near-Data Processing Architecture for Graph Neural Networks. IEEE Trans. Comput. (2023)."},{"key":"e_1_2_1_143_1","volume-title":"GraphSaint: Graph Sampling Based Inductive Learning Method. arXiv","author":"Zeng Hanqing","year":"2019","unstructured":"Hanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan, and Viktor Prasanna. 2019. GraphSaint: Graph Sampling Based Inductive Learning Method. arXiv (2019)."},{"key":"e_1_2_1_144_1","volume-title":"Serving Graph Neural Networks With Distributed Fog Servers for Smart IoT Services","author":"Zeng Liekang","year":"2023","unstructured":"Liekang Zeng, Xu Chen, Peng Huang, Ke Luo, Xiaoxi Zhang, and Zhi Zhou. 2023. Serving Graph Neural Networks With Distributed Fog Servers for Smart IoT Services. IEEE\/ACM Transactions on Networking (2023)."},{"key":"e_1_2_1_145_1","volume-title":"TLP: A Deep Learning-Based Cost Model for Tensor Program Tuning. In ASPLOS.","author":"Zhai Yi","year":"2023","unstructured":"Yi Zhai, Yu Zhang, Shuo Liu, Xiaomeng Chu, Jie Peng, Jianmin Ji, and Yanyong Zhang. 2023. TLP: A Deep Learning-Based Cost Model for Tensor Program Tuning. In ASPLOS."},{"key":"e_1_2_1_146_1","doi-asserted-by":"crossref","unstructured":"Mingxing Zhang Youwei Zhuo Chao Wang Mingyu Gao Yongwei Wu Kang Chen Christos Kozyrakis and Xuehai Qian. 2018. GraphP: Reducing Communication for PIM-Based Graph Processing with Efficient Data Partition. In HPCA.","DOI":"10.1109\/HPCA.2018.00053"},{"key":"e_1_2_1_147_1","doi-asserted-by":"crossref","unstructured":"Tianyi Zhang Zhiqiu Lin Guandao Yang and Christopher De Sa. 2019. QPyTorch: A Low-Precision Arithmetic Simulation Framework. In arXiv.","DOI":"10.1109\/EMC2-NIPS53020.2019.00010"},{"key":"e_1_2_1_148_1","doi-asserted-by":"crossref","unstructured":"D. Zheng C. Ma M. Wang J. Zhou Q. Su X. Song Q. Gan Z. Zhang and G. Karypis. 2020. DistDGL: Distributed Graph Neural Network Training for Billion-Scale Graphs. In IA3.","DOI":"10.1109\/IA351965.2020.00011"},{"key":"e_1_2_1_149_1","doi-asserted-by":"publisher","DOI":"10.1145\/3559009.3569670"},{"key":"e_1_2_1_150_1","doi-asserted-by":"crossref","unstructured":"Youwei Zhuo Chao Wang Mingxing Zhang Rui Wang Dimin Niu Yanzhi Wang and Xuehai Qian. 2019. GraphQ: Scalable PIM-based Graph Processing. In MICRO.","DOI":"10.1145\/3352460.3358256"}],"container-title":["Proceedings of the ACM on Measurement and Analysis of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3700434","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3700434","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T00:14:19Z","timestamp":1755908059000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3700434"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,10]]},"references-count":150,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,12,10]]}},"alternative-id":["10.1145\/3700434"],"URL":"https:\/\/doi.org\/10.1145\/3700434","relation":{},"ISSN":["2476-1249"],"issn-type":[{"value":"2476-1249","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,10]]},"assertion":[{"value":"2024-12-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}