{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T23:50:35Z","timestamp":1783036235950,"version":"3.54.6"},"reference-count":122,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2022,2,24]],"date-time":"2022-02-24T00:00:00Z","timestamp":1645660800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Meas. Anal. Comput. Syst."],"published-print":{"date-parts":[[2022,2,24]]},"abstract":"<jats:p>Several manufacturers have already started to commercialize near-bank Processing-In-Memory (PIM) architectures, after decades of research efforts. Near-bank PIM architectures place simple cores close to DRAM banks. Recent research demonstrates that they can yield significant performance and energy improvements in parallel applications by alleviating data access costs. Real PIM systems can provide high levels of parallelism, large aggregate memory bandwidth and low memory access latency, thereby being a good fit to accelerate the Sparse Matrix Vector Multiplication (SpMV) kernel. SpMV has been characterized as one of the most significant and thoroughly studied scientific computation kernels. It is primarily a memory-bound kernel with intensive memory accesses due its algorithmic nature, the compressed matrix format used, and the sparsity patterns of the input matrices given. This paper provides the first comprehensive analysis of SpMV on a real-world PIM architecture, and presents SparseP, the first SpMV library for real PIM architectures. We make three key contributions. First, we implement a wide variety of software strategies on SpMV for a multithreaded PIM core, including (1) various compressed matrix formats, (2) load balancing schemes across parallel threads and (3) synchronization approaches, and characterize the computational limits of a single multithreaded PIM core. Second, we design various load balancing schemes across multiple PIM cores, and two types of data partitioning techniques to execute SpMV on thousands of PIM cores: (1) 1D-partitioned kernels to perform the complete SpMV computation only using PIM cores, and (2) 2D-partitioned kernels to strive a balance between computation and data transfer costs to PIM-enabled memory. Third, we compare SpMV execution on a real-world PIM system with 2528 PIM cores to an Intel Xeon CPU and an NVIDIA Tesla V100 GPU to study the performance and energy efficiency of various devices, i.e., both memory-centric PIM systems and conventional processor-centric CPU\/GPU systems, for the SpMV kernel. SparseP software package provides 25 SpMV kernels for real PIM systems supporting the four most widely used compressed matrix formats, i.e., CSR, COO, BCSR and BCOO, and a wide range of data types. SparseP is publicly and freely available at https:\/\/github.com\/CMU-SAFARI\/SparseP. Our extensive evaluation using 26 matrices with various sparsity patterns provides new insights and recommendations for software designers and hardware architects to efficiently accelerate the SpMV kernel on real PIM systems.<\/jats:p>","DOI":"10.1145\/3508041","type":"journal-article","created":{"date-parts":[[2022,2,28]],"date-time":"2022-02-28T23:44:29Z","timestamp":1646091869000},"page":"1-49","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":77,"title":["SparseP"],"prefix":"10.1145","volume":"6","author":[{"given":"Christina","family":"Giannoula","sequence":"first","affiliation":[{"name":"ETH Z\u00fcrich &amp; National Technical University of Athens, Athens, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ivan","family":"Fernandez","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich &amp; University of Malaga, Malaga, Spain"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Juan G\u00f3mez","family":"Luna","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Z\u00fcrich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nectarios","family":"Koziris","sequence":"additional","affiliation":[{"name":"National Technical University of Athens, Athens, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Georgios","family":"Goumas","sequence":"additional","affiliation":[{"name":"National Technical University of Athens, Athens, Greece"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Onur","family":"Mutlu","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Z\u00fcrich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,2,28]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Junwhan Ahn Sungpack Hong Sungjoo Yoo Onur Mutlu and Kiyoung Choi. 2015. A Scalable Processing-In-Memory Accelerator for Parallel Graph Processing. In ISCA .  Junwhan Ahn Sungpack Hong Sungjoo Yoo Onur Mutlu and Kiyoung Choi. 2015. A Scalable Processing-In-Memory Accelerator for Parallel Graph Processing. In ISCA ."},{"key":"e_1_2_1_2_1","doi-asserted-by":"crossref","unstructured":"Bahar Asgari Ramyad Hadidi Joshua Dierberger Charlotte Steinichen and Hyesoon Kim. 2020 a. Copernicus: Characterizing the Performance Implications of Compression Formats Used in Sparse Workloads. In CoRR . https:\/\/arxiv.org\/abs\/2011.10932  Bahar Asgari Ramyad Hadidi Joshua Dierberger Charlotte Steinichen and Hyesoon Kim. 2020 a. Copernicus: Characterizing the Performance Implications of Compression Formats Used in Sparse Workloads. In CoRR . https:\/\/arxiv.org\/abs\/2011.10932","DOI":"10.1109\/IISWC53511.2021.00012"},{"key":"e_1_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Bahar Asgari Ramyad Hadidi Tushar Krishna Hyesoon Kim and Sudhakar Yalamanchili. 2020 b. ALRESCHA: A Lightweight Reconfigurable Sparse-Computation Accelerator. In HPCA .  Bahar Asgari Ramyad Hadidi Tushar Krishna Hyesoon Kim and Sudhakar Yalamanchili. 2020 b. ALRESCHA: A Lightweight Reconfigurable Sparse-Computation Accelerator. In HPCA .","DOI":"10.1109\/HPCA47549.2020.00029"},{"key":"e_1_2_1_4_1","volume-title":"Jung Ho Ahn, and Nam Sung Kim.","author":"Asghari-Moghaddam Hadi","year":"2016","unstructured":"Hadi Asghari-Moghaddam , Young Hoon Son , Jung Ho Ahn, and Nam Sung Kim. 2016 . Chameleon : Versatile and PracticalNear-DRAM Acceleration Architecture for Large Memory Systems. In MICRO . Hadi Asghari-Moghaddam, Young Hoon Son, Jung Ho Ahn, and Nam Sung Kim. 2016. Chameleon: Versatile and PracticalNear-DRAM Acceleration Architecture for Large Memory Systems. In MICRO ."},{"key":"e_1_2_1_5_1","volume-title":"Ribbens","author":"Belgin Mehmet","year":"2009","unstructured":"Mehmet Belgin , Godmar Back , and Calvin J . Ribbens . 2009 . Pattern-Based Sparse Matrix Representation for Memory-Efficient SMVM Kernels. In ICS . Mehmet Belgin, Godmar Back, and Calvin J. Ribbens. 2009. Pattern-Based Sparse Matrix Representation for Memory-Efficient SMVM Kernels. In ICS ."},{"key":"e_1_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Akrem Benatia Weixing Ji and Yizhuo Wang. 2019. Sparse Matrix Partitioning for Optimizing SpMV on CPU-GPU Heterogeneous Platforms. In IJHPCA .  Akrem Benatia Weixing Ji and Yizhuo Wang. 2019. Sparse Matrix Partitioning for Optimizing SpMV on CPU-GPU Heterogeneous Platforms. In IJHPCA .","DOI":"10.1177\/1094342019886628"},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Akrem Benatia Weixing Ji Yizhuo Wang and Feng Shi. 2016. Sparse Matrix Format Selection with Multiclass SVM for SpMV on GPU. In ICPP .  Akrem Benatia Weixing Ji Yizhuo Wang and Feng Shi. 2016. Sparse Matrix Format Selection with Multiclass SVM for SpMV on GPU. In ICPP .","DOI":"10.1109\/ICPP.2016.64"},{"key":"e_1_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Akrem Benatia Weixing Ji Yizhuo Wang and Feng Shi. 2018. BestSF: A Sparse Meta-Format for Optimizing SpMV on GPU. In TACO .  Akrem Benatia Weixing Ji Yizhuo Wang and Feng Shi. 2018. BestSF: A Sparse Meta-Format for Optimizing SpMV on GPU. In TACO .","DOI":"10.1145\/3226228"},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Maciej Besta Florian Marending Edgar Solomonik and Torsten Hoefler. 2017. SlimSell: A Vectorizable Graph Representation for Breadth-First Search. In IPDPS .  Maciej Besta Florian Marending Edgar Solomonik and Torsten Hoefler. 2017. SlimSell: A Vectorizable Graph Representation for Breadth-First Search. In IPDPS .","DOI":"10.1109\/IPDPS.2017.93"},{"key":"e_1_2_1_10_1","volume-title":"Bisseling and Wouter Meesen","author":"Rob","year":"2005","unstructured":"Rob H. Bisseling and Wouter Meesen . 2005 . Communication Balancing in Parallel Sparse Matrix-Vector Multiplication. In ETNA. Electronic Transactions on Numerical Analysis . Rob H. Bisseling and Wouter Meesen. 2005. Communication Balancing in Parallel Sparse Matrix-Vector Multiplication. In ETNA. Electronic Transactions on Numerical Analysis ."},{"key":"e_1_2_1_11_1","series-title":"SIAM .","volume-title":"Numerical Methods for Least Squares Problems","author":"Bj\u00f6rck \u00c5ke","unstructured":"\u00c5ke Bj\u00f6rck . 1996. Numerical Methods for Least Squares Problems . In SIAM . \u00c5ke Bj\u00f6rck. 1996. Numerical Methods for Least Squares Problems. In SIAM ."},{"key":"e_1_2_1_12_1","doi-asserted-by":"crossref","unstructured":"Jeff Bolz Ian Farmer Eitan Grinspun and Peter Schr\u00f6der. 2003 a. Sparse Matrix Solvers on the GPU: Conjugate Gradients and Multigrid. In SIGGRAPH .  Jeff Bolz Ian Farmer Eitan Grinspun and Peter Schr\u00f6der. 2003 a. Sparse Matrix Solvers on the GPU: Conjugate Gradients and Multigrid. In SIGGRAPH .","DOI":"10.1145\/1201775.882364"},{"key":"e_1_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Jeff Bolz Ian Farmer Eitan Grinspun and Peter Schr\u00f6der. 2003 b. Sparse Matrix Solvers on the GPU: Conjugate Gradients and Multigrid. In ACM Transactions on Graphics .  Jeff Bolz Ian Farmer Eitan Grinspun and Peter Schr\u00f6der. 2003 b. Sparse Matrix Solvers on the GPU: Conjugate Gradients and Multigrid. In ACM Transactions on Graphics .","DOI":"10.1145\/1201775.882364"},{"key":"e_1_2_1_14_1","doi-asserted-by":"crossref","unstructured":"Sergey Brin and Lawrence Page. 1998. The Anatomy of a Large-scale Hypertextual Web Search Engine. In WWW .  Sergey Brin and Lawrence Page. 1998. The Anatomy of a Large-scale Hypertextual Web Search Engine. In WWW .","DOI":"10.1016\/S0169-7552(98)00110-X"},{"key":"e_1_2_1_15_1","doi-asserted-by":"crossref","unstructured":"Aydin Bulu\u00e7 Samuel Williams Leonid Oliker and James Demmel. 2011. Reduced-Bandwidth Multithreaded Algorithms for Sparse Matrix-Vector Multiplication. In IPDPS .  Aydin Bulu\u00e7 Samuel Williams Leonid Oliker and James Demmel. 2011. Reduced-Bandwidth Multithreaded Algorithms for Sparse Matrix-Vector Multiplication. In IPDPS .","DOI":"10.1109\/IPDPS.2011.73"},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Beata Bylina Jaroslaw Bylina Przemyslaw Stpiczy'ski and Dominik Szakowski. 2014. Performance Analysis of Multicore and Multinodal Implementation of SpMV Operation. In FedCSIS .  Beata Bylina Jaroslaw Bylina Przemyslaw Stpiczy'ski and Dominik Szakowski. 2014. Performance Analysis of Multicore and Multinodal Implementation of SpMV Operation. In FedCSIS .","DOI":"10.15439\/2014F313"},{"key":"e_1_2_1_17_1","unstructured":"Benjamin Y. Cho Yongkee Kwon Sangkug Lym and Mattan Erez. 2020. Near Data Acceleration with Concurrent Host Access. In ISCA .  Benjamin Y. Cho Yongkee Kwon Sangkug Lym and Mattan Erez. 2020. Near Data Acceleration with Concurrent Host Access. In ISCA ."},{"key":"e_1_2_1_18_1","volume-title":"Vuduc","author":"Choi Jee W.","year":"2010","unstructured":"Jee W. Choi , Amik Singh , and Richard W . Vuduc . 2010 . Model-Driven Autotuning of Sparse Matrix-Vector Multiply on GPUs. In PpopP . Jee W. Choi, Amik Singh, and Richard W. Vuduc. 2010. Model-Driven Autotuning of Sparse Matrix-Vector Multiply on GPUs. In PpopP ."},{"key":"e_1_2_1_19_1","unstructured":"CSR5. 2015. CSR5 Cuda . https:\/\/github.com\/weifengliu-ssslab\/Benchmark_SpMV_using_CSR5  CSR5. 2015. CSR5 Cuda . https:\/\/github.com\/weifengliu-ssslab\/Benchmark_SpMV_using_CSR5"},{"key":"e_1_2_1_20_1","unstructured":"cuSparse. 2021. cuSparse . https:\/\/docs.nvidia.com\/cuda\/cusparse\/index.html  cuSparse. 2021. cuSparse . https:\/\/docs.nvidia.com\/cuda\/cusparse\/index.html"},{"key":"e_1_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Leonardo Dagum and Ramesh Menon. 1998. OpenMP: An Industry-Standard API for Shared-Memory Programming. In IEEE Comput. Sci. Eng.  Leonardo Dagum and Ramesh Menon. 1998. OpenMP: An Industry-Standard API for Shared-Memory Programming. In IEEE Comput. Sci. Eng.","DOI":"10.1109\/99.660313"},{"key":"e_1_2_1_22_1","volume-title":"Davis and Yifan Hu","author":"Timothy","year":"2011","unstructured":"Timothy A. Davis and Yifan Hu . 2011 . The University of Florida Sparse Matrix Collection . In TOMS . Timothy A. Davis and Yifan Hu. 2011. The University of Florida Sparse Matrix Collection. In TOMS ."},{"key":"e_1_2_1_23_1","doi-asserted-by":"crossref","unstructured":"F. Devaux. 2019. The True Processing In Memory Accelerator. In Hot Chips .  F. Devaux. 2019. The True Processing In Memory Accelerator. In Hot Chips .","DOI":"10.1109\/HOTCHIPS.2019.8875680"},{"key":"e_1_2_1_24_1","unstructured":"Jack Dongarra Andrew Lumsdaine Xinhui Niu Roldan Pozoz and Karin Remington. 1994. Sparse Matrix Libraries in C  Jack Dongarra Andrew Lumsdaine Xinhui Niu Roldan Pozoz and Karin Remington. 1994. Sparse Matrix Libraries in C"},{"key":"e_1_2_1_25_1","unstructured":"for High Performance Architectures. In Mathematics .  for High Performance Architectures. In Mathematics ."},{"key":"e_1_2_1_26_1","doi-asserted-by":"crossref","unstructured":"Athena Elafrou G. Goumas and N. Koziris. 2017. Performance Analysis and Optimization of Sparse Matrix-Vector Multiplication on Modern Multi- and Many-Core Processors. In ICPP .  Athena Elafrou G. Goumas and N. Koziris. 2017. Performance Analysis and Optimization of Sparse Matrix-Vector Multiplication on Modern Multi- and Many-Core Processors. In ICPP .","DOI":"10.1109\/IPDPSW.2017.134"},{"key":"e_1_2_1_27_1","doi-asserted-by":"crossref","unstructured":"Athena Elafrou Georgios Goumas and Nectarios Koziris. 2019. Conflict-Free Symmetric Sparse Matrix-Vector Multiplication on Multicore Architectures. In SC .  Athena Elafrou Georgios Goumas and Nectarios Koziris. 2019. Conflict-Free Symmetric Sparse Matrix-Vector Multiplication on Multicore Architectures. In SC .","DOI":"10.1145\/3295500.3356148"},{"key":"e_1_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Athena Elafrou Vasileios Karakasis Theodoros Gkountouvas Kornilios Kourtis Georgios Goumas and Nectarios Koziris. 2018. SparseX: A Library for High-Performance Sparse Matrix-Vector Multiplication on Multicore Platforms. In ACM TOMS .  Athena Elafrou Vasileios Karakasis Theodoros Gkountouvas Kornilios Kourtis Georgios Goumas and Nectarios Koziris. 2018. SparseX: A Library for High-Performance Sparse Matrix-Vector Multiplication on Multicore Platforms. In ACM TOMS .","DOI":"10.1145\/3134442"},{"key":"e_1_2_1_29_1","doi-asserted-by":"crossref","unstructured":"R. D. Falgout. 2006. An Introduction to Algebraic Multigrid. In Computing in Science Engineering .  R. D. Falgout. 2006. An Introduction to Algebraic Multigrid. In Computing in Science Engineering .","DOI":"10.1109\/MCSE.2006.105"},{"key":"e_1_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Robert D Falgout and Ulrike Meier Yang. 2002. hypre: A Library of High Performance Preconditioners. In ICCS .  Robert D Falgout and Ulrike Meier Yang. 2002. hypre: A Library of High Performance Preconditioners. In ICCS .","DOI":"10.1007\/3-540-47789-6_66"},{"key":"e_1_2_1_31_1","volume-title":"NATSA: A Near-Data Processing Accelerator for Time Series Analysis. In ICCD .","author":"Fernandez Ivan","year":"2020","unstructured":"Ivan Fernandez , Ricardo Quislant , Christina Giannoula , Mohammed Alser , Juan G\u00f3mez-Luna , Eladio Guti\u00e9rrez , Oscar Plata , and Onur Mutlu . 2020 . NATSA: A Near-Data Processing Accelerator for Time Series Analysis. In ICCD . Ivan Fernandez, Ricardo Quislant, Christina Giannoula, Mohammed Alser, Juan G\u00f3mez-Luna, Eladio Guti\u00e9rrez, Oscar Plata, and Onur Mutlu. 2020. NATSA: A Near-Data Processing Accelerator for Time Series Analysis. In ICCD ."},{"key":"e_1_2_1_32_1","doi-asserted-by":"crossref","unstructured":"Jeremy Fowers Kalin Ovtcharov Karin Strauss Eric S. Chung and Greg Stitt. 2014. A High Memory Bandwidth FPGA Accelerator for Sparse Matrix-Vector Multiplication. In FCCM .  Jeremy Fowers Kalin Ovtcharov Karin Strauss Eric S. Chung and Greg Stitt. 2014. A High Memory Bandwidth FPGA Accelerator for Sparse Matrix-Vector Multiplication. In FCCM .","DOI":"10.1109\/FCCM.2014.23"},{"key":"e_1_2_1_33_1","doi-asserted-by":"crossref","unstructured":"Daichi Fujiki Niladrish Chatterjee Donghyuk Lee and Mike O'Connor. 2019. Near-Memory Data Transformation for Efficient Sparse Matrix Multi-Vector Multiplication. In SC .  Daichi Fujiki Niladrish Chatterjee Donghyuk Lee and Mike O'Connor. 2019. Near-Memory Data Transformation for Efficient Sparse Matrix Multi-Vector Multiplication. In SC .","DOI":"10.1145\/3295500.3356154"},{"key":"e_1_2_1_34_1","unstructured":"Mingyu Gao Grant Ayers and Christos Kozyrakis. 2015. Practical Near-Data Processing for In-Memory Analytics Frameworks. In PACT .  Mingyu Gao Grant Ayers and Christos Kozyrakis. 2015. Practical Near-Data Processing for In-Memory Analytics Frameworks. In PACT ."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3037697.3037702"},{"key":"e_1_2_1_36_1","volume-title":"Nectarios Koziris, Georgios Goumas, and Onur Mutlu.","author":"Giannoula Christina","year":"2022","unstructured":"Christina Giannoula , Ivan Fernandez , Juan G\u00f3 mez-Luna , Nectarios Koziris, Georgios Goumas, and Onur Mutlu. 2022 . SparseP: Towards Efficient Sparse Matrix Vector Multiplication on Real Processing-In-Memory Systems. In CoRR . https:\/\/arxiv.org\/abs\/2201.05072 Christina Giannoula, Ivan Fernandez, Juan G\u00f3 mez-Luna, Nectarios Koziris, Georgios Goumas, and Onur Mutlu. 2022. SparseP: Towards Efficient Sparse Matrix Vector Multiplication on Real Processing-In-Memory Systems. In CoRR . https:\/\/arxiv.org\/abs\/2201.05072"},{"key":"e_1_2_1_37_1","volume-title":"Lois Orosa, Nectarios Koziris, Georgios I. Goumas, and Onur Mutlu.","author":"Giannoula Christina","year":"2021","unstructured":"Christina Giannoula , Nandita Vijaykumar , Nikela Papadopoulou , Vasileios Karakostas , Ivan Fernandez , Juan G\u00f3 mez-Luna , Lois Orosa, Nectarios Koziris, Georgios I. Goumas, and Onur Mutlu. 2021 . SynCron: Efficient Synchronization Support for Near-Data-Processing Architectures. In HPCA . Christina Giannoula, Nandita Vijaykumar, Nikela Papadopoulou, Vasileios Karakostas, Ivan Fernandez, Juan G\u00f3 mez-Luna, Lois Orosa, Nectarios Koziris, Georgios I. Goumas, and Onur Mutlu. 2021. SynCron: Efficient Synchronization Support for Near-Data-Processing Architectures. In HPCA ."},{"key":"e_1_2_1_38_1","volume-title":"Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu.","author":"Luna Juan G\u00f3","year":"2021","unstructured":"Juan G\u00f3 mez- Luna , Izzat El Hajj , Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2021 . Benchmarking a New Paradigm : An Experimental Analysis of a Real Processing-in-Memory Architecture. In CoRR . https:\/\/arxiv.org\/abs\/2105.03814 Juan G\u00f3 mez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2021. Benchmarking a New Paradigm: An Experimental Analysis of a Real Processing-in-Memory Architecture. In CoRR . https:\/\/arxiv.org\/abs\/2105.03814"},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Georgios Goumas Kornilios Kourtis Nikos Anastopoulos Vasileios Karakasis and Nectarios Koziris. 2009. Performance Evaluation of the Sparse Matrix-Vector Multiplication on Modern Architectures. In J. Supercomput.  Georgios Goumas Kornilios Kourtis Nikos Anastopoulos Vasileios Karakasis and Nectarios Koziris. 2009. Performance Evaluation of the Sparse Matrix-Vector Multiplication on Modern Architectures. In J. Supercomput.","DOI":"10.1109\/PDP.2008.41"},{"key":"e_1_2_1_40_1","doi-asserted-by":"crossref","unstructured":"Paul Grigoras Pavel Burovskiy Eddie Hung and Wayne Luk. 2015. Accelerating SpMV on FPGAs by Compressing Nonzero Values. In FCCM .  Paul Grigoras Pavel Burovskiy Eddie Hung and Wayne Luk. 2015. Accelerating SpMV on FPGAs by Compressing Nonzero Values. In FCCM .","DOI":"10.1109\/FCCM.2015.30"},{"key":"e_1_2_1_41_1","unstructured":"SAFARI Research Group. 2022. SparseP Software Package . https:\/\/github.com\/Carnegie Mellon University-SAFARI\/SparseP  SAFARI Research Group. 2022. SparseP Software Package . https:\/\/github.com\/Carnegie Mellon University-SAFARI\/SparseP"},{"key":"e_1_2_1_42_1","volume-title":"A Performance Modeling and Optimization Analysis Tool for Sparse Matrix-Vector Multiplication on GPUs","author":"Guo Ping","unstructured":"Ping Guo , Liqiang Wang , and Po Chen . 2014. A Performance Modeling and Optimization Analysis Tool for Sparse Matrix-Vector Multiplication on GPUs . In IEEE TPDS . Ping Guo, Liqiang Wang, and Po Chen. 2014. A Performance Modeling and Optimization Analysis Tool for Sparse Matrix-Vector Multiplication on GPUs. In IEEE TPDS ."},{"key":"e_1_2_1_43_1","doi-asserted-by":"crossref","unstructured":"Udit Gupta Xiaodong Wang Maxim Naumov Carole-Jean Wu Brandon Reagen David Brooks Bradford Cottel Kim M. Hazelwood Bill Jia Hsien-Hsin S. Lee Andrey Malevich Dheevatsa Mudigere Mikhail Smelyanskiy Liang Xiong and Xuan Zhang. 2019. The Architectural Implications of Facebook's DNN-based Personalized Recommendation. In CoRR .  Udit Gupta Xiaodong Wang Maxim Naumov Carole-Jean Wu Brandon Reagen David Brooks Bradford Cottel Kim M. Hazelwood Bill Jia Hsien-Hsin S. Lee Andrey Malevich Dheevatsa Mudigere Mikhail Smelyanskiy Liang Xiong and Xuan Zhang. 2019. The Architectural Implications of Facebook's DNN-based Personalized Recommendation. In CoRR .","DOI":"10.1109\/HPCA47549.2020.00047"},{"key":"e_1_2_1_44_1","volume-title":"Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu.","author":"G\u00f3mez-Luna Juan","year":"2021","unstructured":"Juan G\u00f3mez-Luna , Izzat El Hajj , Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2021 . Benchmarking Memory-Centric Computing Systems: Analysis of Real Processing-In-Memory Hardware. In IGSC . Juan G\u00f3mez-Luna, Izzat El Hajj, Ivan Fernandez, Christina Giannoula, Geraldo F. Oliveira, and Onur Mutlu. 2021. Benchmarking Memory-Centric Computing Systems: Analysis of Real Processing-In-Memory Hardware. In IGSC ."},{"key":"e_1_2_1_45_1","volume-title":"Fletcher","author":"Hegde Kartik","year":"2019","unstructured":"Kartik Hegde , Hadi Asghari-Moghaddam , Michael Pellauer , Neal Crago , Aamer Jaleel , Edgar Solomonik , Joel Emer , and Christopher W . Fletcher . 2019 . ExTensor: An Accelerator for Sparse Tensor Algebra. In MICRO . Kartik Hegde, Hadi Asghari-Moghaddam, Michael Pellauer, Neal Crago, Aamer Jaleel, Edgar Solomonik, Joel Emer, and Christopher W. Fletcher. 2019. ExTensor: An Accelerator for Sparse Tensor Algebra. In MICRO ."},{"key":"e_1_2_1_46_1","volume-title":"PASTIX: A High-Performance Parallel Direct Solver for Sparse Symmetric Positive Definite Systems. In PMAA .","author":"H\u00e9non Pascal","year":"2002","unstructured":"Pascal H\u00e9non , Pierre Ramet , and Jean Roman . 2002 . PASTIX: A High-Performance Parallel Direct Solver for Sparse Symmetric Positive Definite Systems. In PMAA . Pascal H\u00e9non, Pierre Ramet, and Jean Roman. 2002. PASTIX: A High-Performance Parallel Direct Solver for Sparse Symmetric Positive Definite Systems. In PMAA ."},{"key":"e_1_2_1_47_1","volume-title":"Israt Nisa, Shivani Sabhlok, \u00dcmit V. cCataly\u00fcrek, Srinivasan Parthasarathy, and P. Sadayappan.","author":"Hong Changwan","year":"2018","unstructured":"Changwan Hong , Aravind Sukumaran-Rajam , Bortik Bandyopadhyay , Jinsung Kim , S\u00fcreyya Emre Kurt , Israt Nisa, Shivani Sabhlok, \u00dcmit V. cCataly\u00fcrek, Srinivasan Parthasarathy, and P. Sadayappan. 2018 a. Efficient Sparse-Matrix Multi-Vector Product on GPUs. In HPDC . Changwan Hong, Aravind Sukumaran-Rajam, Bortik Bandyopadhyay, Jinsung Kim, S\u00fcreyya Emre Kurt, Israt Nisa, Shivani Sabhlok, \u00dcmit V. cCataly\u00fcrek, Srinivasan Parthasarathy, and P. Sadayappan. 2018a. Efficient Sparse-Matrix Multi-Vector Product on GPUs. In HPDC ."},{"key":"e_1_2_1_48_1","volume-title":"Israt Nisa, Shivani Sabhlok, \u00dcmit V. cCataly\u00fcrek, Srinivasan Parthasarathy, and P. Sadayappan.","author":"Hong Changwan","year":"2018","unstructured":"Changwan Hong , Aravind Sukumaran-Rajam , Bortik Bandyopadhyay , Jinsung Kim , S\u00fcreyya Emre Kurt , Israt Nisa, Shivani Sabhlok, \u00dcmit V. cCataly\u00fcrek, Srinivasan Parthasarathy, and P. Sadayappan. 2018 b. Efficient Sparse-Matrix Multi-Vector Product on GPUs. In HPDC . Changwan Hong, Aravind Sukumaran-Rajam, Bortik Bandyopadhyay, Jinsung Kim, S\u00fcreyya Emre Kurt, Israt Nisa, Shivani Sabhlok, \u00dcmit V. cCataly\u00fcrek, Srinivasan Parthasarathy, and P. Sadayappan. 2018b. Efficient Sparse-Matrix Multi-Vector Product on GPUs. In HPDC ."},{"key":"e_1_2_1_49_1","volume-title":"Yelick","author":"Im Eun-Jin","year":"1999","unstructured":"Eun-Jin Im and Katherine A . Yelick . 1999 . Optimizing Sparse Matrix Vector Multiplication on SMP. In PPSC. Eun-Jin Im and Katherine A. Yelick. 1999. Optimizing Sparse Matrix Vector Multiplication on SMP. In PPSC."},{"key":"e_1_2_1_50_1","volume-title":"Sparsity: Optimization Framework for Sparse Matrix Kernels. In The International Journal of High Performance Computing Applications .","author":"Im Eun-Jin","year":"2004","unstructured":"Eun-Jin Im , Katherine Yelick , and Richard Vuduc . 2004 . Sparsity: Optimization Framework for Sparse Matrix Kernels. In The International Journal of High Performance Computing Applications . Eun-Jin Im, Katherine Yelick, and Richard Vuduc. 2004. Sparsity: Optimization Framework for Sparse Matrix Kernels. In The International Journal of High Performance Computing Applications ."},{"key":"e_1_2_1_51_1","unstructured":"Sivaramakrishna Bharadwaj Indarapu Manoj Maramreddy and Kishore Kothapalli. 2014. Architecture- and Workload- Aware Heterogeneous Algorithms for Sparse Matrix Vector Multiplication. In COMPUTE .  Sivaramakrishna Bharadwaj Indarapu Manoj Maramreddy and Kishore Kothapalli. 2014. Architecture- and Workload- Aware Heterogeneous Algorithms for Sparse Matrix Vector Multiplication. In COMPUTE ."},{"key":"e_1_2_1_52_1","volume-title":"Taha Shahroodi, Juan Gomez Luna, and Onur Mutlu.","author":"Kanellopoulos Konstantinos","year":"2019","unstructured":"Konstantinos Kanellopoulos , Nandita Vijaykumar , Christina Giannoula , Roknoddin Azizi , Skanda Koppula , Nika Mansouri Ghiasi , Taha Shahroodi, Juan Gomez Luna, and Onur Mutlu. 2019 . SMASH : Co-Designing Software Compression and Hardware-Accelerated Indexing for Efficient Sparse Matrix Operations. In MICRO . Konstantinos Kanellopoulos, Nandita Vijaykumar, Christina Giannoula, Roknoddin Azizi, Skanda Koppula, Nika Mansouri Ghiasi, Taha Shahroodi, Juan Gomez Luna, and Onur Mutlu. 2019. SMASH: Co-Designing Software Compression and Hardware-Accelerated Indexing for Efficient Sparse Matrix Operations. In MICRO ."},{"key":"e_1_2_1_53_1","doi-asserted-by":"crossref","unstructured":"Vasileios Karakasis Georgios Goumas and Nectarios Koziris. 2009. Perfomance Models for Blocked Sparse Matrix-Vector Multiplication Kernels. In ICPP .  Vasileios Karakasis Georgios Goumas and Nectarios Koziris. 2009. Perfomance Models for Blocked Sparse Matrix-Vector Multiplication Kernels. In ICPP .","DOI":"10.1109\/ICPP.2009.21"},{"key":"e_1_2_1_54_1","volume-title":"Semi-Two-Dimensional Partitioning for Parallel Sparse Matrix-Vector Multiplication. In IPDPS Workshop .","author":"Kayaaslan Enver","year":"2015","unstructured":"Enver Kayaaslan , Bora U\u00e7ar , and Cevdet Aykanat . 2015 . Semi-Two-Dimensional Partitioning for Parallel Sparse Matrix-Vector Multiplication. In IPDPS Workshop . Enver Kayaaslan, Bora U\u00e7ar, and Cevdet Aykanat. 2015. Semi-Two-Dimensional Partitioning for Parallel Sparse Matrix-Vector Multiplication. In IPDPS Workshop ."},{"key":"e_1_2_1_55_1","volume-title":"Mark Hempstead, Brandon Reagen, Xuan Zhang, David Brooks, Vikas Chandra, Utku Diril, et almbox.","author":"Ke Liu","year":"2020","unstructured":"Liu Ke , Udit Gupta , Carole-Jean Wu , Benjamin Youngjae Cho , Mark Hempstead, Brandon Reagen, Xuan Zhang, David Brooks, Vikas Chandra, Utku Diril, et almbox. 2020 . RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing. In ISCA . Liu Ke, Udit Gupta, Carole-Jean Wu, Benjamin Youngjae Cho, Mark Hempstead, Brandon Reagen, Xuan Zhang, David Brooks, Vikas Chandra, Utku Diril, et almbox. 2020. RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing. In ISCA ."},{"key":"e_1_2_1_56_1","unstructured":"Kashif Nizam Khan Mikael Hirki Tapio Niemi Jukka K Nurminen and Zhonghong Ou. 2018. Rapl in Action: Experiences in Using RAPL for Power Measurements. In TOMPECS .  Kashif Nizam Khan Mikael Hirki Tapio Niemi Jukka K Nurminen and Zhonghong Ou. 2018. Rapl in Action: Experiences in Using RAPL for Power Measurements. In TOMPECS ."},{"key":"e_1_2_1_57_1","doi-asserted-by":"crossref","unstructured":"Yoongu Kim Vivek Seshadri Donghyuk Lee Jamie Liu and Onur Mutlu. 2012. A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM. In ISCA .  Yoongu Kim Vivek Seshadri Donghyuk Lee Jamie Liu and Onur Mutlu. 2012. A Case for Exploiting Subarray-Level Parallelism (SALP) in DRAM. In ISCA .","DOI":"10.1109\/ISCA.2012.6237032"},{"key":"e_1_2_1_58_1","doi-asserted-by":"crossref","unstructured":"David R Kincaid Thomas C Oppe and David M Young. 1989. Itpackv 2D User's Guide .  David R Kincaid Thomas C Oppe and David M Young. 1989. Itpackv 2D User's Guide .","DOI":"10.2172\/7093021"},{"key":"e_1_2_1_59_1","volume-title":"TACO: A Tool to Generate Tensor Algebra Kernels . In ASE .","author":"Kjolstad Fredrik","year":"2017","unstructured":"Fredrik Kjolstad , Stephen Chou , David Lugato , Shoaib Kamil , and Saman Amarasinghe . 2017 . TACO: A Tool to Generate Tensor Algebra Kernels . In ASE . Fredrik Kjolstad, Stephen Chou, David Lugato, Shoaib Kamil, and Saman Amarasinghe. 2017. TACO: A Tool to Generate Tensor Algebra Kernels . In ASE ."},{"key":"e_1_2_1_60_1","doi-asserted-by":"crossref","unstructured":"Kornilios Kourtis Georgios Goumas and Nectarios Koziris. 2008. Optimizing Sparse Matrix-Vector Multiplication Using Index and Value Compression. In CF .  Kornilios Kourtis Georgios Goumas and Nectarios Koziris. 2008. Optimizing Sparse Matrix-Vector Multiplication Using Index and Value Compression. In CF .","DOI":"10.1109\/ICPP.2008.62"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/1941553.1941587"},{"key":"e_1_2_1_62_1","doi-asserted-by":"crossref","unstructured":"Youngeun Kwon Yunjae Lee and Minsoo Rhu. 2019. TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning. In MICRO .  Youngeun Kwon Yunjae Lee and Minsoo Rhu. 2019. TensorDIMM: A Practical Near-Memory Processing Architecture for Embeddings and Tensor Operations in Deep Learning. In MICRO .","DOI":"10.1145\/3352460.3358284"},{"key":"e_1_2_1_63_1","unstructured":"Young-Cheon Kwon Suk Han Lee Jaehoon Lee Sang-Hyuk Kwon Je Min Ryu Jong-Pil Son O Seongil Hak-Soo Yu Haesuk Lee Soo Young Kim Youngmin Cho Jin Guk Kim Jongyoon Choi Hyun-Sung Shin Jin Kim BengSeng Phuah HyoungMin Kim Myeong Jun Song Ahn Choi Daeho Kim SooYoung Kim Eun-Bong Kim David Wang Shinhaeng Kang Yuhwan Ro Seungwoo Seo JoonHo Song Jaeyoun Youn Kyomin Sohn and Nam Sung Kim. 2021. 25.4 A 20nm 6GB Function-In-Memory DRAM Based on HBM2 with a 1.2TFLOPS Programmable Computing Unit Using Bank-Level Parallelism for Machine Learning Applications. In ISSCC .  Young-Cheon Kwon Suk Han Lee Jaehoon Lee Sang-Hyuk Kwon Je Min Ryu Jong-Pil Son O Seongil Hak-Soo Yu Haesuk Lee Soo Young Kim Youngmin Cho Jin Guk Kim Jongyoon Choi Hyun-Sung Shin Jin Kim BengSeng Phuah HyoungMin Kim Myeong Jun Song Ahn Choi Daeho Kim SooYoung Kim Eun-Bong Kim David Wang Shinhaeng Kang Yuhwan Ro Seungwoo Seo JoonHo Song Jaeyoun Youn Kyomin Sohn and Nam Sung Kim. 2021. 25.4 A 20nm 6GB Function-In-Memory DRAM Based on HBM2 with a 1.2TFLOPS Programmable Computing Unit Using Bank-Level Parallelism for Machine Learning Applications. In ISSCC ."},{"key":"e_1_2_1_64_1","doi-asserted-by":"crossref","unstructured":"Daniel Langr and Pavel Tvrd\u00edk. 2016. Evaluation Criteria for Sparse Matrix Storage Formats. In TPDS .  Daniel Langr and Pavel Tvrd\u00edk. 2016. Evaluation Criteria for Sparse Matrix Storage Formats. In TPDS .","DOI":"10.1109\/TPDS.2015.2401575"},{"key":"e_1_2_1_65_1","doi-asserted-by":"crossref","unstructured":"Dominique Lavenier Remy Cimadomo and Romaric Jodin. 2020. Variant Calling Parallelization on Processor-in-Memory Architecture. In BIBM.  Dominique Lavenier Remy Cimadomo and Romaric Jodin. 2020. Variant Calling Parallelization on Processor-in-Memory Architecture. In BIBM.","DOI":"10.1109\/BIBM49941.2020.9313351"},{"key":"e_1_2_1_66_1","unstructured":"Seyong Lee and Rudolf Eigenmann. 2008. Adaptive Runtime Tuning of Parallel Sparse Matrix-Vector Multiplication on Distributed Memory Systems. In ICS .  Seyong Lee and Rudolf Eigenmann. 2008. Adaptive Runtime Tuning of Parallel Sparse Matrix-Vector Multiplication on Distributed Memory Systems. In ICS ."},{"key":"e_1_2_1_67_1","volume-title":"H. Yoon, Seungwon Lee, K. Lim, Hyunsung Shin, Jinhyun Kim, O. Seongil, Anand Iyer, David Wang, K. Sohn, and N. Kim.","author":"Lee Sukhan","year":"2021","unstructured":"Sukhan Lee , Shin-Haeng Kang , Jaehoon Lee , H. Kim , Eojin Lee , Seung young Seo , H. Yoon, Seungwon Lee, K. Lim, Hyunsung Shin, Jinhyun Kim, O. Seongil, Anand Iyer, David Wang, K. Sohn, and N. Kim. 2021 . Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product. In ISCA . Sukhan Lee, Shin-Haeng Kang, Jaehoon Lee, H. Kim, Eojin Lee, Seung young Seo, H. Yoon, Seungwon Lee, K. Lim, Hyunsung Shin, Jinhyun Kim, O. Seongil, Anand Iyer, David Wang, K. Sohn, and N. Kim. 2021. Hardware Architecture and Software Stack for PIM Based on Commercial DRAM Technology: Industrial Product. In ISCA ."},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/2898361"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/2491956.2462181"},{"key":"e_1_2_1_70_1","volume-title":"Performance Analysis and Optimization for SpMV on GPU Using Probabilistic Modeling","author":"Li Kenli","unstructured":"Kenli Li , Wangdong Yang , and Keqin Li. 2015. Performance Analysis and Optimization for SpMV on GPU Using Probabilistic Modeling . In IEEE TPDS . Kenli Li, Wangdong Yang, and Keqin Li. 2015. Performance Analysis and Optimization for SpMV on GPU Using Probabilistic Modeling. In IEEE TPDS ."},{"key":"e_1_2_1_71_1","doi-asserted-by":"crossref","unstructured":"Colin Yu Lin Zheng Zhang Ngai Wong and Hayden Kwok-Hay So. 2010. Design Space Exploration for Sparse Matrix-Matrix Multiplication on FPGAs. In FPT .  Colin Yu Lin Zheng Zhang Ngai Wong and Hayden Kwok-Hay So. 2010. Design Space Exploration for Sparse Matrix-Matrix Multiplication on FPGAs. In FPT .","DOI":"10.1109\/FPT.2010.5681425"},{"key":"e_1_2_1_72_1","doi-asserted-by":"crossref","unstructured":"Greg Linden Brent Smith and Jeremy York. 2003. Amazon.com Recommendations: Item-to-Item Collaborative Filtering. In IC .  Greg Linden Brent Smith and Jeremy York. 2003. Amazon.com Recommendations: Item-to-Item Collaborative Filtering. In IC .","DOI":"10.1109\/MIC.2003.1167344"},{"key":"e_1_2_1_73_1","doi-asserted-by":"crossref","unstructured":"Baoyuan Liu Min Wang Hassan Foroosh Marshall Tappen and Marianna Pensky. 2015. Sparse Convolutional Neural Networks. In CVPR .  Baoyuan Liu Min Wang Hassan Foroosh Marshall Tappen and Marianna Pensky. 2015. Sparse Convolutional Neural Networks. In CVPR .","DOI":"10.1109\/CVPR.2015.7298681"},{"key":"e_1_2_1_74_1","unstructured":"Changxi Liu Biwei Xie Xin Liu Wei Xue Hailong Yang and Xu Liu. 2018. Towards Efficient SpMV on Sunway Manycore Architectures. In ICS .  Changxi Liu Biwei Xie Xin Liu Wei Xue Hailong Yang and Xu Liu. 2018. Towards Efficient SpMV on Sunway Manycore Architectures. In ICS ."},{"key":"e_1_2_1_75_1","unstructured":"Weifeng Liu and Brian Vinter. 2014. An Efficient GPU General Sparse Matrix-Matrix Multiplication for Irregular Data. In IPDPS .  Weifeng Liu and Brian Vinter. 2014. An Efficient GPU General Sparse Matrix-Matrix Multiplication for Irregular Data. In IPDPS ."},{"key":"e_1_2_1_76_1","unstructured":"Weifeng Liu and Brian Vinter. 2015a. CSR5: An Efficient Storage Format for Cross-Platform Sparse Matrix-Vector Multiplication. In ICS .  Weifeng Liu and Brian Vinter. 2015a. CSR5: An Efficient Storage Format for Cross-Platform Sparse Matrix-Vector Multiplication. In ICS ."},{"key":"e_1_2_1_77_1","unstructured":"Weifeng Liu and Brian Vinter. 2015b. CSR5: An Efficient Storage Format for Cross-Platform Sparse Matrix-Vector Multiplication. In ICS .  Weifeng Liu and Brian Vinter. 2015b. CSR5: An Efficient Storage Format for Cross-Platform Sparse Matrix-Vector Multiplication. In ICS ."},{"key":"e_1_2_1_78_1","doi-asserted-by":"crossref","unstructured":"Marco Maggioni and Tanya Berger-Wolf. 2013. AdELL: An Adaptive Warp-Balancing ELL Format for Efficient Sparse Matrix-Vector Multiplication on GPUs. In ICPP .  Marco Maggioni and Tanya Berger-Wolf. 2013. AdELL: An Adaptive Warp-Balancing ELL Format for Efficient Sparse Matrix-Vector Multiplication on GPUs. In ICPP .","DOI":"10.1109\/ICPP.2013.10"},{"key":"e_1_2_1_79_1","doi-asserted-by":"crossref","unstructured":"Duane Merrill and Michael Garland. 2016. Merge-Based Parallel Sparse Matrix-Vector Multiplication. In SC .  Duane Merrill and Michael Garland. 2016. Merge-Based Parallel Sparse Matrix-Vector Multiplication. In SC .","DOI":"10.1109\/SC.2016.57"},{"key":"e_1_2_1_80_1","doi-asserted-by":"crossref","unstructured":"Anurag Mukkara Nathan Beckmann Maleen Abeydeera Xiaosong Ma and Daniel Sanchez. 2018. Exploiting Locality in Graph Analytics through Hardware-Accelerated Traversal Scheduling. In MICRO .  Anurag Mukkara Nathan Beckmann Maleen Abeydeera Xiaosong Ma and Daniel Sanchez. 2018. Exploiting Locality in Graph Analytics through Hardware-Accelerated Traversal Scheduling. In MICRO .","DOI":"10.1109\/MICRO.2018.00010"},{"key":"e_1_2_1_81_1","doi-asserted-by":"crossref","unstructured":"Onur Mutlu Saugata Ghose Juan G\u00f3mez-Luna and Rachata Ausavarungnirun. 2021. A Modern Primer on Processing in Memory. In Emerging Computing: From Devices to Systems - Looking Beyond Moore and Von Neumann. https:\/\/arxiv.org\/pdf\/2012.03112.pdf  Onur Mutlu Saugata Ghose Juan G\u00f3mez-Luna and Rachata Ausavarungnirun. 2021. A Modern Primer on Processing in Memory. In Emerging Computing: From Devices to Systems - Looking Beyond Moore and Von Neumann. https:\/\/arxiv.org\/pdf\/2012.03112.pdf","DOI":"10.1007\/978-981-16-7487-7_7"},{"key":"e_1_2_1_82_1","doi-asserted-by":"crossref","unstructured":"Naveen Namashivayam Sanyam Mehta and Pen-Chung Yew. 2021. Variable-Sized Blocks for Locality-Aware SpMV . In CGO .  Naveen Namashivayam Sanyam Mehta and Pen-Chung Yew. 2021. Variable-Sized Blocks for Locality-Aware SpMV . In CGO .","DOI":"10.1109\/CGO51591.2021.9370327"},{"key":"e_1_2_1_83_1","unstructured":"Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia Liang Xiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. In CoRR .  Maxim Naumov Dheevatsa Mudigere Hao-Jun Michael Shi Jianyu Huang Narayanan Sundaraman Jongsoo Park Xiaodong Wang Udit Gupta Carole-Jean Wu Alisson G. Azzolini Dmytro Dzhulgakov Andrey Mallevich Ilia Cherniavskii Yinghai Lu Raghuraman Krishnamoorthi Ansha Yu Volodymyr Kondratenko Stephanie Pereira Xianjie Chen Wenlin Chen Vijay Rao Bill Jia Liang Xiong and Misha Smelyanskiy. 2019. Deep Learning Recommendation Model for Personalization and Recommendation Systems. In CoRR ."},{"key":"e_1_2_1_84_1","doi-asserted-by":"crossref","unstructured":"Yuyao Niu Zhengyang Lu Meichen Dong Zhou Jin Weifeng Liu and Guangming Tan. 2021. TileSpMV: A Tiled Algorithm for Sparse Matrix-Vector Multiplication on GPUs. In IPDPS .  Yuyao Niu Zhengyang Lu Meichen Dong Zhou Jin Weifeng Liu and Guangming Tan. 2021. TileSpMV: A Tiled Algorithm for Sparse Matrix-Vector Multiplication on GPUs. In IPDPS .","DOI":"10.1109\/IPDPS49936.2021.00016"},{"key":"e_1_2_1_85_1","doi-asserted-by":"crossref","unstructured":"Eriko Nurvitadhi Asit Mishra Yu Wang Ganesh Venkatesh and Debbie Marr. 2016. Hardware Accelerator for Analytics of Sparse Data. In DAC .  Eriko Nurvitadhi Asit Mishra Yu Wang Ganesh Venkatesh and Debbie Marr. 2016. Hardware Accelerator for Analytics of Sparse Data. In DAC .","DOI":"10.3850\/9783981537079_0766"},{"key":"e_1_2_1_86_1","unstructured":"NVIDIA. 2016. NVIDIA System Management Interface Program . http:\/\/developer.download.nvidia.com\/compute\/DCGM\/docs\/nvidia-smi-367.38.pdf .  NVIDIA. 2016. NVIDIA System Management Interface Program . http:\/\/developer.download.nvidia.com\/compute\/DCGM\/docs\/nvidia-smi-367.38.pdf ."},{"key":"e_1_2_1_87_1","volume-title":"Kogge","author":"Page Brian A.","year":"2018","unstructured":"Brian A. Page and Peter M . Kogge . 2018 . Scalability of Hybrid Sparse Matrix Dense Vector (SpMV) Multiplication. In HPCS . Brian A. Page and Peter M. Kogge. 2018. Scalability of Hybrid Sparse Matrix Dense Vector (SpMV) Multiplication. In HPCS ."},{"key":"e_1_2_1_88_1","unstructured":"Subhankar Pal Jonathan Beaumont Dong-Hyeon Park Aporva Amarnath Siying Feng Chaitali Chakrabarti Hun-Seok Kim David Blaauw Trevor Mudge and Ronald Dreslinski. 2018. OuterSPACE: An Outer Product Based Sparse Matrix Multiplication Accelerator. In HPCA .  Subhankar Pal Jonathan Beaumont Dong-Hyeon Park Aporva Amarnath Siying Feng Chaitali Chakrabarti Hun-Seok Kim David Blaauw Trevor Mudge and Ronald Dreslinski. 2018. OuterSPACE: An Outer Product Based Sparse Matrix Multiplication Accelerator. In HPCA ."},{"key":"e_1_2_1_89_1","unstructured":"peakperf. 2021. peakperf. https:\/\/github.com\/Dr-Noob\/peakperf.git  peakperf. 2021. peakperf. https:\/\/github.com\/Dr-Noob\/peakperf.git"},{"key":"e_1_2_1_90_1","volume-title":"Heath","author":"Pinar Ali","year":"1999","unstructured":"Ali Pinar and Michael T . Heath . 1999 . Improving Performance of Sparse Matrix-Vector Multiplication. In SC . Ali Pinar and Michael T. Heath. 1999. Improving Performance of Sparse Matrix-Vector Multiplication. In SC ."},{"key":"e_1_2_1_91_1","volume-title":"Pooch and Al Nieder","author":"Udo","year":"1973","unstructured":"Udo W. Pooch and Al Nieder . 1973 . A Survey of Indexing Techniques for Sparse Matrices. In ACM Comput. Surv . Udo W. Pooch and Al Nieder. 1973. A Survey of Indexing Techniques for Sparse Matrices. In ACM Comput. Surv."},{"key":"e_1_2_1_92_1","volume-title":"SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training. In HPCA .","author":"Qin Eric","year":"2020","unstructured":"Eric Qin , Ananda Samajdar , Hyoukjun Kwon , Vineet Nadella , Sudarshan Srinivasan , Dipankar Das , Bharat Kaul , and Tushar Krishna . 2020 . SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training. In HPCA . Eric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella, Sudarshan Srinivasan, Dipankar Das, Bharat Kaul, and Tushar Krishna. 2020. SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training. In HPCA ."},{"key":"e_1_2_1_93_1","volume-title":"James C. Hoe, Larry Pileggi, and Franz Franchetti.","author":"Sadi Fazle","year":"2019","unstructured":"Fazle Sadi , Joe Sweeney , Tze Meng Low , James C. Hoe, Larry Pileggi, and Franz Franchetti. 2019 . Efficient SpMV Operation for Large and Highly Sparse Matrices Using Scalable Multi-Way Merge Parallelization. In MICRO . Fazle Sadi, Joe Sweeney, Tze Meng Low, James C. Hoe, Larry Pileggi, and Franz Franchetti. 2019. Efficient SpMV Operation for Large and Highly Sparse Matrices Using Scalable Multi-Way Merge Parallelization. In MICRO ."},{"key":"e_1_2_1_94_1","unstructured":"SciPy. 2021. List-of-list Sparse Matrix .  SciPy. 2021. List-of-list Sparse Matrix ."},{"key":"e_1_2_1_95_1","doi-asserted-by":"crossref","unstructured":"Naser Sedaghati Te Mu Louis-Noel Pouchet Srinivasan Parthasarathy and P. Sadayappan. 2015. Automatic Selection of Sparse Matrix Representation on GPUs. In ICS .  Naser Sedaghati Te Mu Louis-Noel Pouchet Srinivasan Parthasarathy and P. Sadayappan. 2015. Automatic Selection of Sparse Matrix Representation on GPUs. In ICS .","DOI":"10.1145\/2751205.2751244"},{"key":"e_1_2_1_96_1","volume-title":"Owens","author":"Sengupta Shubhabrata","year":"2007","unstructured":"Shubhabrata Sengupta , Mark Harris , Yao Zhang , and John D . Owens . 2007 . Scan Primitives for GPU Computing. In GH . Shubhabrata Sengupta, Mark Harris, Yao Zhang, and John D. Owens. 2007. Scan Primitives for GPU Computing. In GH ."},{"key":"e_1_2_1_97_1","unstructured":"A. Smith. 2019. 6 New Facts About Facebook . http:\/\/mediashift.org  A. Smith. 2019. 6 New Facts About Facebook . http:\/\/mediashift.org"},{"key":"e_1_2_1_98_1","doi-asserted-by":"crossref","unstructured":"Markus Steinberger Rhaleb Zayer and Hans-Peter Seidel. 2017. Globally Homogeneous Locally Adaptive Sparse Matrix-Vector Multiplication on the GPU. In ICS .  Markus Steinberger Rhaleb Zayer and Hans-Peter Seidel. 2017. Globally Homogeneous Locally Adaptive Sparse Matrix-Vector Multiplication on the GPU. In ICS .","DOI":"10.1145\/3079079.3079086"},{"key":"e_1_2_1_99_1","unstructured":"stream. 2021. stream. https:\/\/github.com\/jeffhammond\/STREAM.git  stream. 2021. stream. https:\/\/github.com\/jeffhammond\/STREAM.git"},{"key":"e_1_2_1_100_1","unstructured":"Bor-Yiing Su and Kurt Keutzer. 2012. ClSpMV: A Cross-Platform OpenCL SpMV Framework on GPUs. In ICS .  Bor-Yiing Su and Kurt Keutzer. 2012. ClSpMV: A Cross-Platform OpenCL SpMV Framework on GPUs. In ICS ."},{"key":"e_1_2_1_101_1","doi-asserted-by":"publisher","DOI":"10.1145\/3218823"},{"key":"e_1_2_1_102_1","volume-title":"Xibai Li, and Rick Siow Mong Goh.","author":"Tang Wai Teng","year":"2015","unstructured":"Wai Teng Tang , Ruizhe Zhao , Mian Lu , Yun Liang , Huynh Phung Huyng , Xibai Li, and Rick Siow Mong Goh. 2015 . Optimizing and Auto-Tuning Scale-Free Sparse Matrix-Vector Multiplication on Intel Xeon Phi. In CGO . Wai Teng Tang, Ruizhe Zhao, Mian Lu, Yun Liang, Huynh Phung Huyng, Xibai Li, and Rick Siow Mong Goh. 2015. Optimizing and Auto-Tuning Scale-Free Sparse Matrix-Vector Multiplication on Intel Xeon Phi. In CGO ."},{"key":"e_1_2_1_103_1","doi-asserted-by":"crossref","unstructured":"Yaman Umuroglu and Magnus Jahre. 2014. An Energy Efficient Column-Major Backend for FPGA SpMV Accelerators. In ICCD .  Yaman Umuroglu and Magnus Jahre. 2014. An Energy Efficient Column-Major Backend for FPGA SpMV Accelerators. In ICCD .","DOI":"10.1109\/ICCD.2014.6974716"},{"key":"e_1_2_1_104_1","unstructured":"UPMEM. 2018. Introduction to UPMEM PIM. Processing-in-memory (PIM) on DRAM Accelerator (White Paper) .  UPMEM. 2018. Introduction to UPMEM PIM. Processing-in-memory (PIM) on DRAM Accelerator (White Paper) ."},{"key":"e_1_2_1_105_1","unstructured":"UPMEM. 2020. UPMEM Website . https:\/\/www.upmem.com  UPMEM. 2020. UPMEM Website . https:\/\/www.upmem.com"},{"key":"e_1_2_1_106_1","volume-title":"UPMEM User Manual. Version","author":"UPMEM.","year":"2021","unstructured":"UPMEM. 2021. UPMEM User Manual. Version 2021 .3 . UPMEM. 2021. UPMEM User Manual. Version 2021.3 ."},{"key":"e_1_2_1_107_1","doi-asserted-by":"crossref","unstructured":"R. Vuduc J.W. Demmel K.A. Yelick S. Kamil R. Nishtala and B. Lee. 2002. Performance Optimizations and Bounds for Sparse Matrix-Vector Multiply. In SC .  R. Vuduc J.W. Demmel K.A. Yelick S. Kamil R. Nishtala and B. Lee. 2002. Performance Optimizations and Bounds for Sparse Matrix-Vector Multiply. In SC .","DOI":"10.1109\/SC.2002.10025"},{"key":"e_1_2_1_108_1","volume-title":"Demmel","author":"Vuduc Richard Wilson","year":"2003","unstructured":"Richard Wilson Vuduc and James W . Demmel . 2003 . Automatic Performance Tuning of Sparse Matrix Kernels. In PhD Thesis . Richard Wilson Vuduc and James W. Demmel. 2003. Automatic Performance Tuning of Sparse Matrix Kernels. In PhD Thesis ."},{"key":"e_1_2_1_109_1","volume-title":"Vuduc and Hyun-Jin Moon","author":"Richard","year":"2005","unstructured":"Richard W. Vuduc and Hyun-Jin Moon . 2005 . Fast Sparse Matrix-Vector Multiplication by Exploiting Variable Block Structure. In HPCC . Richard W. Vuduc and Hyun-Jin Moon. 2005. Fast Sparse Matrix-Vector Multiplication by Exploiting Variable Block Structure. In HPCC ."},{"key":"e_1_2_1_110_1","doi-asserted-by":"crossref","unstructured":"Jeremiah Willcock and Andrew Lumsdaine. 2006. Accelerating Sparse Matrix Computations via Data Compression. In ICS .  Jeremiah Willcock and Andrew Lumsdaine. 2006. Accelerating Sparse Matrix Computations via Data Compression. In ICS .","DOI":"10.1145\/1183401.1183444"},{"key":"e_1_2_1_111_1","doi-asserted-by":"crossref","unstructured":"Samuel Williams Leonid Oliker Richard Vuduc John Shalf Katherine Yelick and James Demmel. 2007. Optimization of Sparse Matrix-Vector Multiplication on Emerging Multicore Platforms. In SC .  Samuel Williams Leonid Oliker Richard Vuduc John Shalf Katherine Yelick and James Demmel. 2007. Optimization of Sparse Matrix-Vector Multiplication on Emerging Multicore Platforms. In SC .","DOI":"10.1145\/1362622.1362674"},{"key":"e_1_2_1_112_1","unstructured":"Tianji Wu Bo Wang Yi Shan Feng Yan Yu Wang and Ningyi Xu. 2010. Efficient PageRank and SpMV Computation on AMD GPUs. In ICPP .  Tianji Wu Bo Wang Yi Shan Feng Yan Yu Wang and Ningyi Xu. 2010. Efficient PageRank and SpMV Computation on AMD GPUs. In ICPP ."},{"key":"e_1_2_1_113_1","unstructured":"Xinfeng Xie Zheng Liang Peng Gu Abanti Basak Lei Deng Ling Liang Xing Hu and Yuan Xie. 2021. SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator. In HPCA.  Xinfeng Xie Zheng Liang Peng Gu Abanti Basak Lei Deng Ling Liang Xing Hu and Yuan Xie. 2021. SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator. In HPCA."},{"key":"e_1_2_1_114_1","unstructured":"Shengen Yan Chao Li Yunquan Zhang and Huiyang Zhou. 2014a. YaSpMV: Yet Another SpMV Framework on GPUs. In PPoPP .  Shengen Yan Chao Li Yunquan Zhang and Huiyang Zhou. 2014a. YaSpMV: Yet Another SpMV Framework on GPUs. In PPoPP ."},{"key":"e_1_2_1_115_1","unstructured":"Shengen Yan Chao Li Yunquan Zhang and Huiyang Zhou. 2014b. YaSpMV: Yet Another SpMV Framework on GPUs. In PPoPP .  Shengen Yan Chao Li Yunquan Zhang and Huiyang Zhou. 2014b. YaSpMV: Yet Another SpMV Framework on GPUs. In PPoPP ."},{"key":"e_1_2_1_116_1","doi-asserted-by":"crossref","unstructured":"Wangdong Yang Kenli Li and Keqin Li. 2017. A Hybrid Computing Method of SpMV on CPU--GPU Heterogeneous Computing Systems. In JPDC .  Wangdong Yang Kenli Li and Keqin Li. 2017. A Hybrid Computing Method of SpMV on CPU--GPU Heterogeneous Computing Systems. In JPDC .","DOI":"10.1016\/j.jpdc.2016.12.023"},{"key":"e_1_2_1_117_1","unstructured":"Wangdong Yang Kenli Li Yan Liu Lin Shi and Lanjun Wan. 2014. Optimization of Quasi-Diagonal Matrix-Vector Multiplication on GPU. In Int. J. High Perform. Comput. Appl.  Wangdong Yang Kenli Li Yan Liu Lin Shi and Lanjun Wan. 2014. Optimization of Quasi-Diagonal Matrix-Vector Multiplication on GPU. In Int. J. High Perform. Comput. Appl."},{"key":"e_1_2_1_118_1","volume-title":"Performance Optimization Using Partitioned SpMV on GPUs and Multicore CPUs","author":"Yang Wangdong","unstructured":"Wangdong Yang , Kenli Li , Zeyao Mo , and Keqin Li. 2015. Performance Optimization Using Partitioned SpMV on GPUs and Multicore CPUs . In IEEE Transactions on Computers . Wangdong Yang, Kenli Li, Zeyao Mo, and Keqin Li. 2015. Performance Optimization Using Partitioned SpMV on GPUs and Multicore CPUs. In IEEE Transactions on Computers ."},{"key":"e_1_2_1_119_1","doi-asserted-by":"crossref","unstructured":"Shijin Zhang Zidong Du Lei Zhang Huiying Lan Shaoli Liu Ling Li Qi Guo Tianshi Chen and Yunji Chen. 2016. Cambricon-X: An Accelerator for Sparse Neural Networks. In MICRO .  Shijin Zhang Zidong Du Lei Zhang Huiying Lan Shaoli Liu Ling Li Qi Guo Tianshi Chen and Yunji Chen. 2016. Cambricon-X: An Accelerator for Sparse Neural Networks. In MICRO .","DOI":"10.1109\/MICRO.2016.7783723"},{"key":"e_1_2_1_120_1","doi-asserted-by":"crossref","unstructured":"Yue Zhao Jiajia Li Chunhua Liao and Xipeng Shen. 2018a. Bridging the Gap between Deep Learning and Sparse Matrix Format Selection. In PPoPP .  Yue Zhao Jiajia Li Chunhua Liao and Xipeng Shen. 2018a. Bridging the Gap between Deep Learning and Sparse Matrix Format Selection. In PPoPP .","DOI":"10.2172\/1426119"},{"key":"e_1_2_1_121_1","doi-asserted-by":"crossref","unstructured":"Yue Zhao Jiajia Li Chunhua Liao and Xipeng Shen. 2018b. Bridging the Gap between Deep Learning and Sparse Matrix Format Selection. In PPoPP .  Yue Zhao Jiajia Li Chunhua Liao and Xipeng Shen. 2018b. Bridging the Gap between Deep Learning and Sparse Matrix Format Selection. In PPoPP .","DOI":"10.2172\/1426119"},{"key":"e_1_2_1_122_1","doi-asserted-by":"crossref","unstructured":"Yue Zhao Weijie Zhou Xipeng Shen and Graham Yiu. 2018c. Overhead-Conscious Format Selection for SpMV-Based Applications. In IPDPS .  Yue Zhao Weijie Zhou Xipeng Shen and Graham Yiu. 2018c. Overhead-Conscious Format Selection for SpMV-Based Applications. In IPDPS .","DOI":"10.1109\/IPDPS.2018.00104"}],"container-title":["Proceedings of the ACM on Measurement and Analysis of Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3508041","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3508041","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:12:29Z","timestamp":1750191149000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3508041"}},"subtitle":["Towards Efficient Sparse Matrix Vector Multiplication on Real Processing-In-Memory Architectures"],"short-title":[],"issued":{"date-parts":[[2022,2,24]]},"references-count":122,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,2,24]]}},"alternative-id":["10.1145\/3508041"],"URL":"https:\/\/doi.org\/10.1145\/3508041","relation":{},"ISSN":["2476-1249"],"issn-type":[{"value":"2476-1249","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,24]]},"assertion":[{"value":"2022-02-28","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}