{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,5]],"date-time":"2026-08-05T18:09:24Z","timestamp":1785953364085,"version":"3.56.0"},"publisher-location":"New York, NY, USA","reference-count":79,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","award":["1937301, 2028602, CCF-1563078, 1563113"],"award-info":[{"award-number":["1937301, 2028602, CCF-1563078, 1563113"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]},{"name":"DARPA\/AFRL","award":["FA8650-18-2-7865"],"award-info":[{"award-number":["FA8650-18-2-7865"]}]},{"name":"DARPA","award":["FA-8750-17-2-0095, FA-8750-14-2-0240"],"award-info":[{"award-number":["FA-8750-17-2-0095, FA-8750-14-2-0240"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,18]]},"DOI":"10.1145\/3466752.3480047","type":"proceedings-article","created":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T19:16:55Z","timestamp":1634498215000},"page":"1022-1035","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":37,"title":["Capstan: A Vector RDA for Sparsity"],"prefix":"10.1145","author":[{"given":"Alexander","family":"Rucker","sequence":"first","affiliation":[{"name":"Stanford University, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Matthew","family":"Vilim","sequence":"additional","affiliation":[{"name":"Stanford University, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tian","family":"Zhao","sequence":"additional","affiliation":[{"name":"Stanford University, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yaqi","family":"Zhang","sequence":"additional","affiliation":[{"name":"Stanford University, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Raghu","family":"Prabhakar","sequence":"additional","affiliation":[{"name":"SambaNova Systems, Inc., United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kunle","family":"Olukotun","sequence":"additional","affiliation":[{"name":"Stanford University, United States of America"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"PIUMA: Programmable Integrated Unified Memory Architecture. arXiv preprint arXiv:2010.06277(2020).","author":"Aananthakrishnan Sriram","year":"2020","unstructured":"Sriram Aananthakrishnan , Nesreen\u00a0 K Ahmed , Vincent Cave , Marcelo Cintra , Yigit Demir , Kristof\u00a0Du Bois , Stijn Eyerman , Joshua\u00a0 B Fryman , Ivan Ganev , Wim Heirman , 2020 . PIUMA: Programmable Integrated Unified Memory Architecture. arXiv preprint arXiv:2010.06277(2020). Sriram Aananthakrishnan, Nesreen\u00a0K Ahmed, Vincent Cave, Marcelo Cintra, Yigit Demir, Kristof\u00a0Du Bois, Stijn Eyerman, Joshua\u00a0B Fryman, Ivan Ganev, Wim Heirman, 2020. PIUMA: Programmable Integrated Unified Memory Architecture. arXiv preprint arXiv:2010.06277(2020)."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2872887.2750386"},{"key":"e_1_3_2_1_3_1","volume-title":"Proceedings of ISCA-43","author":"Albericio J","year":"2016","unstructured":"J Albericio , P Judd , T Hetherington , T Aamodt , N Jerger , and A Moshovos . 2016 . Cnvlutin: Zero-Neuron-Free Deep Convolutional Neural Network Computing . In Proceedings of ISCA-43 . J Albericio, P Judd, T Hetherington, T Aamodt, N Jerger, and A Moshovos. 2016. Cnvlutin: Zero-Neuron-Free Deep Convolutional Neural Network Computing. In Proceedings of ISCA-43."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/161541.161736"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3317550.3321441"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/1654059.1654112"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2004.65"},{"key":"e_1_3_2_1_8_1","unstructured":"Cerebras. 2020. The Cerebras CS-1 Product Overview. https:\/\/secureservercdn.net\/192.169.220.245\/a7b.fcb.myftpupload.com\/wp-content\/uploads\/2020\/01\/The-Cerebras-CS-1-Product-Overview-rev20200112.pdf  Cerebras. 2020. The Cerebras CS-1 Product Overview. https:\/\/secureservercdn.net\/192.169.220.245\/a7b.fcb.myftpupload.com\/wp-content\/uploads\/2020\/01\/The-Cerebras-CS-1-Product-Overview-rev20200112.pdf"},{"key":"e_1_3_2_1_9_1","volume-title":"Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks. In IEEE International Solid-State Circuits Conference, ISSCC","author":"Krishna Yu-Hsin","year":"2016","unstructured":"Chen, Yu-Hsin and Krishna , Tushar and Emer , Joel and Sze , Vivienne. 2016 . Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks. In IEEE International Solid-State Circuits Conference, ISSCC 2016, Digest of Technical Papers. 262\u2013263. Chen, Yu-Hsin and Krishna, Tushar and Emer, Joel and Sze, Vivienne. 2016. Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks. In IEEE International Solid-State Circuits Conference, ISSCC 2016, Digest of Technical Papers. 262\u2013263."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/HCS49909.2020.9220622"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3276493"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/1950413.1950435"},{"key":"e_1_3_2_1_13_1","volume-title":"PolyGraph: Exposing the Value of Flexibility for Graph Processing Accelerators. In 2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 595\u2013608","author":"Dadu Vidushi","year":"2021","unstructured":"Vidushi Dadu , Sihao Liu , and Tony Nowatzki . 2021 . PolyGraph: Exposing the Value of Flexibility for Graph Processing Accelerators. In 2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 595\u2013608 . Vidushi Dadu, Sihao Liu, and Tony Nowatzki. 2021. PolyGraph: Exposing the Value of Flexibility for Graph Processing Accelerators. In 2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). IEEE, 595\u2013608."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358276"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847339"},{"key":"e_1_3_2_1_16_1","volume-title":"Principles and Practices of Interconnection Networks","author":"Dally William\u00a0James","unstructured":"William\u00a0James Dally and Brian\u00a0Patrick Towles . 2004. Principles and Practices of Interconnection Networks . Elsevier . William\u00a0James Dally and Brian\u00a0Patrick Towles. 2004. Principles and Practices of Interconnection Networks. Elsevier."},{"key":"e_1_3_2_1_17_1","unstructured":"Shail Dave Riyadh Baghdadi Tony Nowatzki Sasikanth Avancha Aviral Shrivastava and Baoxin Li. 2020. Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights. arXiv preprint arXiv:2007.00864(2020).  Shail Dave Riyadh Baghdadi Tony Nowatzki Sasikanth Avancha Aviral Shrivastava and Baoxin Li. 2020. Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights. arXiv preprint arXiv:2007.00864(2020)."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2049662.2049670"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.1999.765937"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2012.51"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/355791.355796"},{"key":"e_1_3_2_1_22_1","volume-title":"Graphicionado: A High-Performance and Energy-Efficient Accelerator for Graph Analytics. In 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1\u201313","author":"Ham Tae\u00a0Jun","year":"2016","unstructured":"Tae\u00a0Jun Ham , Lisa Wu , Narayanan Sundaram , Nadathur Satish , and Margaret Martonosi . 2016 . Graphicionado: A High-Performance and Energy-Efficient Accelerator for Graph Analytics. In 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1\u201313 . Tae\u00a0Jun Ham, Lisa Wu, Narayanan Sundaram, Nadathur Satish, and Margaret Martonosi. 2016. Graphicionado: A High-Performance and Energy-Efficient Accelerator for Graph Analytics. In 2016 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 1\u201313."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3007787.3001163"},{"key":"e_1_3_2_1_24_1","unstructured":"Song Han Huizi Mao and William\u00a0J Dally. 2015. Deep compression: Compressing Deep Neural Networks with Pruning Trained Quantization and Huffman Coding. arXiv preprint arXiv:1510.00149(2015).  Song Han Huizi Mao and William\u00a0J Dally. 2015. Deep compression: Compressing Deep Neural Networks with Pruning Trained Quantization and Huffman Coding. arXiv preprint arXiv:1510.00149(2015)."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_1_26_1","volume-title":"Optimizing Energy Efficient Low-Swing Interconnect for Sub-Threshold FPGAs. In 2015 25th International Conference on Field Programmable Logic and Applications (FPL). 1\u20134. https:\/\/doi.org\/10","author":"Qi He","year":"2015","unstructured":"He Qi , O. Ayorinde , Yu Huang , and B. Calhoun . 2015 . Optimizing Energy Efficient Low-Swing Interconnect for Sub-Threshold FPGAs. In 2015 25th International Conference on Field Programmable Logic and Applications (FPL). 1\u20134. https:\/\/doi.org\/10 .1109\/FPL. 2015 .7293979 10.1109\/FPL.2015.7293979 He Qi, O. Ayorinde, Yu Huang, and B. Calhoun. 2015. Optimizing Energy Efficient Low-Swing Interconnect for Sub-Threshold FPGAs. In 2015 25th International Conference on Field Programmable Logic and Applications (FPL). 1\u20134. https:\/\/doi.org\/10.1109\/FPL.2015.7293979"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/965145.801294"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358275"},{"key":"e_1_3_2_1_29_1","unstructured":"Forrest\u00a0N Iandola Song Han Matthew\u00a0W Moskewicz Khalid Ashraf William\u00a0J Dally and Kurt Keutzer. 2016. SqueezeNet: AlexNet-Level Accuracy with 50x Fewer Parameters and< 0.5 MB Model Size. arXiv preprint arXiv:1602.07360(2016).  Forrest\u00a0N Iandola Song Han Matthew\u00a0W Moskewicz Khalid Ashraf William\u00a0J Dally and Kurt Keutzer. 2016. SqueezeNet: AlexNet-Level Accuracy with 50x Fewer Parameters and< 0.5 MB Model Size. arXiv preprint arXiv:1602.07360(2016)."},{"key":"e_1_3_2_1_30_1","unstructured":"Intel. [n.d.]. Intel Math Kernel Library. https:\/\/software.intel.com\/en-us\/mkl  Intel. [n.d.]. Intel Math Kernel Library. https:\/\/software.intel.com\/en-us\/mkl"},{"key":"e_1_3_2_1_31_1","unstructured":"Intel. [n.d.]. Intel Xeon Processor E7-8890 v3. https:\/\/ark.intel.com\/content\/www\/us\/en\/ark\/products\/84685\/intel-xeon-processor-e7-8890-v3-45m-cache-2-50-ghz.html  Intel. [n.d.]. Intel Xeon Processor E7-8890 v3. https:\/\/ark.intel.com\/content\/www\/us\/en\/ark\/products\/84685\/intel-xeon-processor-e7-8890-v3-45m-cache-2-50-ghz.html"},{"key":"e_1_3_2_1_32_1","unstructured":"Nan Jiang Daniel Becker George Michelogiannakis and William\u00a0J. Dally. 2011. Performance Implications of Age-Based Allocation in On-Chip-Networks. (2011).  Nan Jiang Daniel Becker George Michelogiannakis and William\u00a0J. Dally. 2011. Performance Implications of Age-Based Allocation in On-Chip-Networks. (2011)."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/3352460.3358286"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.5555\/305219.305248"},{"key":"e_1_3_2_1_36_1","volume-title":"Mathematical Foundations of the GraphBLAS. In 2016 IEEE High Performance Extreme Computing Conference (HPEC). IEEE, 1\u20139.","author":"Kepner Jeremy","year":"2016","unstructured":"Jeremy Kepner , Peter Aaltonen , David Bader , Aydin Bulu\u00e7 , Franz Franchetti , John Gilbert , Dylan Hutchison , Manoj Kumar , Andrew Lumsdaine , Henning Meyerhenke , 2016 . Mathematical Foundations of the GraphBLAS. In 2016 IEEE High Performance Extreme Computing Conference (HPEC). IEEE, 1\u20139. Jeremy Kepner, Peter Aaltonen, David Bader, Aydin Bulu\u00e7, Franz Franchetti, John Gilbert, Dylan Hutchison, Manoj Kumar, Andrew Lumsdaine, Henning Meyerhenke, 2016. Mathematical Foundations of the GraphBLAS. In 2016 IEEE High Performance Extreme Computing Conference (HPEC). IEEE, 1\u20139."},{"key":"e_1_3_2_1_37_1","unstructured":"Nitish\u00a0Shirish Keskar Dheevatsa Mudigere Jorge Nocedal Mikhail Smelyanskiy and Ping Tak\u00a0Peter Tang. 2017. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. arxiv:1609.04836\u00a0[cs.LG]  Nitish\u00a0Shirish Keskar Dheevatsa Mudigere Jorge Nocedal Mikhail Smelyanskiy and Ping Tak\u00a0Peter Tang. 2017. On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima. arxiv:1609.04836\u00a0[cs.LG]"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2015.2414456"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASE.2017.8115709"},{"key":"e_1_3_2_1_41_1","volume-title":"Spatial: A Language and Compiler for Application Accelerators. In ACM SIGPLAN Notices, Vol.\u00a053. ACM, 296\u2013311.","author":"Koeplinger David","year":"2018","unstructured":"David Koeplinger , Matthew Feldman , Raghu Prabhakar , Yaqi Zhang , Stefan Hadjis , Ruben Fiszel , Tian Zhao , Luigi Nardi , Ardavan Pedram , Christos Kozyrakis , 2018 . Spatial: A Language and Compiler for Application Accelerators. In ACM SIGPLAN Notices, Vol.\u00a053. ACM, 296\u2013311. David Koeplinger, Matthew Feldman, Raghu Prabhakar, Yaqi Zhang, Stefan Hadjis, Ruben Fiszel, Tian Zhao, Luigi Nardi, Ardavan Pedram, Christos Kozyrakis, 2018. Spatial: A Language and Compiler for Application Accelerators. In ACM SIGPLAN Notices, Vol.\u00a053. ACM, 296\u2013311."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2749469.2750374"},{"key":"e_1_3_2_1_43_1","article-title":"GST","volume":"35","author":"Krajcevski Pavel","year":"2016","unstructured":"Pavel Krajcevski , Srihari Pratapa , and Dinesh Manocha . 2016 . GST : GPU-Decodable Supercompressed Textures. ACM Trans. Graph. 35 , 6, Article 230 (Nov. 2016), 10\u00a0pages. https:\/\/doi.org\/10.1145\/2980179.2982439 10.1145\/2980179.2982439 Pavel Krajcevski, Srihari Pratapa, and Dinesh Manocha. 2016. GST: GPU-Decodable Supercompressed Textures. ACM Trans. Graph. 35, 6, Article 230 (Nov. 2016), 10\u00a0pages. https:\/\/doi.org\/10.1145\/2980179.2982439","journal-title":"GPU-Decodable Supercompressed Textures. ACM Trans. Graph."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1080\/15427951.2009.10129177"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/2717764.2717783"},{"key":"e_1_3_2_1_46_1","volume-title":"An FPGA Architecture for the PageRank Eigenvector Problem. In 2008 International Conference on Field Programmable Logic and Applications. IEEE, 523\u2013526","author":"McGettrick Seamas","year":"2008","unstructured":"Seamas McGettrick , Dermot Geraghty , and Ciaran McElroy . 2008 . An FPGA Architecture for the PageRank Eigenvector Problem. In 2008 International Conference on Field Programmable Logic and Applications. IEEE, 523\u2013526 . Seamas McGettrick, Dermot Geraghty, and Ciaran McElroy. 2008. An FPGA Architecture for the PageRank Eigenvector Problem. In 2008 International Conference on Field Programmable Logic and Applications. IEEE, 523\u2013526."},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/90.769767"},{"key":"e_1_3_2_1_48_1","volume-title":"GraphGen: An FPGA Framework for Vertex-Centric Graph Computation. In 2014 IEEE 22nd Annual International Symposium on Field-Programmable Custom Computing Machines. IEEE, 25\u201328","author":"Nurvitadhi Eriko","year":"2014","unstructured":"Eriko Nurvitadhi , Gabriel Weisz , Yu Wang , Skand Hurkat , Marie Nguyen , James\u00a0 C Hoe , Jos\u00e9\u00a0 F Mart\u00ednez , and Carlos Guestrin . 2014 . GraphGen: An FPGA Framework for Vertex-Centric Graph Computation. In 2014 IEEE 22nd Annual International Symposium on Field-Programmable Custom Computing Machines. IEEE, 25\u201328 . Eriko Nurvitadhi, Gabriel Weisz, Yu Wang, Skand Hurkat, Marie Nguyen, James\u00a0C Hoe, Jos\u00e9\u00a0F Mart\u00ednez, and Carlos Guestrin. 2014. GraphGen: An FPGA Framework for Vertex-Centric Graph Computation. In 2014 IEEE 22nd Annual International Symposium on Field-Programmable Custom Computing Machines. IEEE, 25\u201328."},{"key":"e_1_3_2_1_49_1","unstructured":"Nvidia. [n.d.]. Nvidia Tesla V100 GPU Architecture. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf  Nvidia. [n.d.]. Nvidia Tesla V100 GPU Architecture. https:\/\/images.nvidia.com\/content\/volta-architecture\/pdf\/volta-architecture-whitepaper.pdf"},{"key":"e_1_3_2_1_50_1","unstructured":"Nvidia. 2019. The API reference guide for cuSPARSE.  Nvidia. 2019. The API reference guide for cuSPARSE."},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847337"},{"key":"e_1_3_2_1_52_1","volume-title":"OuterSPACE: An Outer Product Based Sparse Matrix Multiplication Accelerator. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 724\u2013736","author":"Pal Subhankar","year":"2018","unstructured":"Subhankar Pal , Jonathan Beaumont , Dong-Hyeon Park , Aporva Amarnath , Siying Feng , Chaitali Chakrabarti , Hun-Seok Kim , David Blaauw , Trevor Mudge , and Ronald Dreslinski . 2018 . OuterSPACE: An Outer Product Based Sparse Matrix Multiplication Accelerator. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 724\u2013736 . Subhankar Pal, Jonathan Beaumont, Dong-Hyeon Park, Aporva Amarnath, Siying Feng, Chaitali Chakrabarti, Hun-Seok Kim, David Blaauw, Trevor Mudge, and Ronald Dreslinski. 2018. OuterSPACE: An Outer Product Based Sparse Matrix Multiplication Accelerator. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 724\u2013736."},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080254"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3297858.3304025"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080256"},{"key":"e_1_3_2_1_56_1","volume-title":"Mapping the Gnutella Network: Macroscopic Properties of Large-Scale Peer-to-Peer Systems. In international workshop on peer-to-peer systems. Springer, 85\u201393","author":"Ripeanu Matei","year":"2002","unstructured":"Matei Ripeanu and Ian Foster . 2002 . Mapping the Gnutella Network: Macroscopic Properties of Large-Scale Peer-to-Peer Systems. In international workshop on peer-to-peer systems. Springer, 85\u201393 . Matei Ripeanu and Ian Foster. 2002. Mapping the Gnutella Network: Macroscopic Properties of Large-Scale Peer-to-Peer Systems. In international workshop on peer-to-peer systems. Springer, 85\u201393."},{"key":"e_1_3_2_1_57_1","unstructured":"SambaNova. 2021. Accelerated Computing with a Reconfigurable Dataflow Architecture. https:\/\/sambanova.ai\/wp-content\/uploads\/2021\/04\/SambaNova_RDA_Whitepaper.pdf  SambaNova. 2021. Accelerated Computing with a Reconfigurable Dataflow Architecture. https:\/\/sambanova.ai\/wp-content\/uploads\/2021\/04\/SambaNova_RDA_Whitepaper.pdf"},{"key":"e_1_3_2_1_58_1","volume-title":"Skewed-Associative Caches. In International Conference on Parallel Architectures and Languages Europe. Springer, 305\u2013316","author":"Seznec Andr\u00e9","year":"1993","unstructured":"Andr\u00e9 Seznec and Francois Bodin . 1993 . Skewed-Associative Caches. In International Conference on Parallel Architectures and Languages Europe. Springer, 305\u2013316 . Andr\u00e9 Seznec and Francois Bodin. 1993. Skewed-Associative Caches. In International Conference on Parallel Architectures and Languages Europe. Springer, 305\u2013316."},{"key":"e_1_3_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.5555\/1015090.1015309"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA.2018.00052"},{"key":"e_1_3_2_1_61_1","volume-title":"MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise Product. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 766\u2013780","author":"Srivastava Nitish","year":"2020","unstructured":"Nitish Srivastava , Hanchen Jin , Jie Liu , David Albonesi , and Zhiru Zhang . 2020 . MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise Product. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 766\u2013780 . Nitish Srivastava, Hanchen Jin, Jie Liu, David Albonesi, and Zhiru Zhang. 2020. MatRaptor: A Sparse-Sparse Matrix Multiplication Accelerator Based on Row-Wise Product. In 2020 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). IEEE, 766\u2013780."},{"key":"e_1_3_2_1_62_1","volume-title":"WaveScalar. In Proceedings of the 36th annual IEEE\/ACM International Symposium on Microarchitecture. IEEE Computer Society, 291","author":"Swanson Steven","year":"2003","unstructured":"Steven Swanson , Ken Michelson , Andrew Schwerin , and Mark Oskin . 2003 . WaveScalar. In Proceedings of the 36th annual IEEE\/ACM International Symposium on Microarchitecture. IEEE Computer Society, 291 . Steven Swanson, Ken Michelson, Andrew Schwerin, and Mark Oskin. 2003. WaveScalar. In Proceedings of the 36th annual IEEE\/ACM International Symposium on Microarchitecture. IEEE Computer Society, 291."},{"key":"e_1_3_2_1_63_1","volume-title":"Proceedings of HotChips, Vol.\u00a013","author":"Taylor Michael","year":"2001","unstructured":"Michael Taylor , Jason Kim , Jason Miller , Fae Ghodrat , Ben Greenwald , Paul Johnson , Walter Lee , Albert Ma , Nathan Shnidman , David Wentzlaff , 2001 . The Raw Processor: A Composeable 32-bit Fabric for Embedded and General Purpose Computing . In Proceedings of HotChips, Vol.\u00a013 . Michael Taylor, Jason Kim, Jason Miller, Fae Ghodrat, Ben Greenwald, Paul Johnson, Walter Lee, Albert Ma, Nathan Shnidman, David Wentzlaff, 2001. The Raw Processor: A Composeable 32-bit Fabric for Embedded and General Purpose Computing. In Proceedings of HotChips, Vol.\u00a013."},{"key":"e_1_3_2_1_64_1","series-title":"SIAM Journal on scientific and Statistical Computing 13, 2","volume-title":"Bi-CGSTAB: A Fast and Smoothly Converging Variant of Bi-CG for the Solution of Nonsymmetric Linear Systems","author":"Vorst A Van\u00a0der","year":"1992","unstructured":"Henk\u00a0 A Van\u00a0der Vorst . 1992. Bi-CGSTAB: A Fast and Smoothly Converging Variant of Bi-CG for the Solution of Nonsymmetric Linear Systems . SIAM Journal on scientific and Statistical Computing 13, 2 ( 1992 ), 631\u2013644. Henk\u00a0A Van\u00a0der Vorst. 1992. Bi-CGSTAB: A Fast and Smoothly Converging Variant of Bi-CG for the Solution of Nonsymmetric Linear Systems. SIAM Journal on scientific and Statistical Computing 13, 2 (1992), 631\u2013644."},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00035"},{"key":"e_1_3_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3289602.3294007"},{"key":"e_1_3_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/2851141.2851145"},{"key":"e_1_3_2_1_68_1","unstructured":"Bob Wheeler. 2020. Growing AI Diversity and Complexity Demands Flexible Data-Center Accelerators. https:\/\/www.simplemachines.ai\/sites\/default\/files\/SMI%20white%20paper-revised.pdf  Bob Wheeler. 2020. Growing AI Diversity and Complexity Demands Flexible Data-Center Accelerators. https:\/\/www.simplemachines.ai\/sites\/default\/files\/SMI%20white%20paper-revised.pdf"},{"key":"e_1_3_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/1519065.1519089"},{"key":"e_1_3_2_1_70_1","doi-asserted-by":"crossref","unstructured":"Guowei Zhang Nithya Attaluri Joel Emer and Daniel Sanchez. 2021. Exploiting Gustavson\u2019s Algorithm to Accelerate Sparse Matrix Multiplication. https:\/\/asplos-conference.org\/abstracts\/asplos21-paper95-extended_abstract.pdf  Guowei Zhang Nithya Attaluri Joel Emer and Daniel Sanchez. 2021. Exploiting Gustavson\u2019s Algorithm to Accelerate Sparse Matrix Multiplication. https:\/\/asplos-conference.org\/abstracts\/asplos21-paper95-extended_abstract.pdf","DOI":"10.1145\/3445814.3446702"},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021737"},{"key":"e_1_3_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.5555\/3327345.3327423"},{"key":"e_1_3_2_1_73_1","volume-title":"GraphP: Reducing Communication for PIM-Based Graph Processing with Efficient Data Partition. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 544\u2013557","author":"Zhang Mingxing","year":"2018","unstructured":"Mingxing Zhang , Youwei Zhuo , Chao Wang , Mingyu Gao , Yongwei Wu , Kang Chen , Christos Kozyrakis , and Xuehai Qian . 2018 . GraphP: Reducing Communication for PIM-Based Graph Processing with Efficient Data Partition. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 544\u2013557 . Mingxing Zhang, Youwei Zhuo, Chao Wang, Mingyu Gao, Yongwei Wu, Kang Chen, Christos Kozyrakis, and Xuehai Qian. 2018. GraphP: Reducing Communication for PIM-Based Graph Processing with Efficient Data Partition. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 544\u2013557."},{"key":"e_1_3_2_1_74_1","volume-title":"Cambricon-X: An Accelerator for Sparse Neural Networks. In The 49th Annual IEEE\/ACM International Symposium on Microarchitecture. IEEE Press, 20","author":"Zhang Shijin","year":"2016","unstructured":"Shijin Zhang , Zidong Du , Lei Zhang , Huiying Lan , Shaoli Liu , Ling Li , Qi Guo , Tianshi Chen , and Yunji Chen . 2016 . Cambricon-X: An Accelerator for Sparse Neural Networks. In The 49th Annual IEEE\/ACM International Symposium on Microarchitecture. IEEE Press, 20 . Shijin Zhang, Zidong Du, Lei Zhang, Huiying Lan, Shaoli Liu, Ling Li, Qi Guo, Tianshi Chen, and Yunji Chen. 2016. Cambricon-X: An Accelerator for Sparse Neural Networks. In The 49th Annual IEEE\/ACM International Symposium on Microarchitecture. IEEE Press, 20."},{"key":"e_1_3_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322249"},{"key":"e_1_3_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1145\/3276491"},{"key":"e_1_3_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00085"},{"key":"e_1_3_2_1_78_1","volume-title":"SpArch: Efficient Architecture for Sparse Matrix Multiplication. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 261\u2013274","author":"Zhang Zhekai","year":"2020","unstructured":"Zhekai Zhang , Hanrui Wang , Song Han , and William\u00a0 J Dally . 2020 . SpArch: Efficient Architecture for Sparse Matrix Multiplication. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 261\u2013274 . Zhekai Zhang, Hanrui Wang, Song Han, and William\u00a0J Dally. 2020. SpArch: Efficient Architecture for Sparse Matrix Multiplication. In 2020 IEEE International Symposium on High Performance Computer Architecture (HPCA). IEEE, 261\u2013274."},{"key":"e_1_3_2_1_79_1","unstructured":"Tian Zhao Yaqi Zhang and Kunle Olukotun. 2019. Serving Recurrent Neural Networks Efficiently with a Spatial Accelerator. arXiv preprint arXiv:1909.13654(2019).  Tian Zhao Yaqi Zhang and Kunle Olukotun. 2019. Serving Recurrent Neural Networks Efficiently with a Spatial Accelerator. arXiv preprint arXiv:1909.13654(2019)."},{"key":"e_1_3_2_1_80_1","volume-title":"Accelerating Large-Scale Single-Source Shortest Path on FPGA. In 2015 IEEE International Parallel and Distributed Processing Symposium Workshop. IEEE, 129\u2013136","author":"Zhou Shijie","year":"2015","unstructured":"Shijie Zhou , Charalampos Chelmis , and Viktor\u00a0 K Prasanna . 2015 . Accelerating Large-Scale Single-Source Shortest Path on FPGA. In 2015 IEEE International Parallel and Distributed Processing Symposium Workshop. IEEE, 129\u2013136 . Shijie Zhou, Charalampos Chelmis, and Viktor\u00a0K Prasanna. 2015. Accelerating Large-Scale Single-Source Shortest Path on FPGA. In 2015 IEEE International Parallel and Distributed Processing Symposium Workshop. IEEE, 129\u2013136."}],"event":{"name":"MICRO '21: 54th Annual IEEE\/ACM International Symposium on Microarchitecture","location":"Virtual Event Greece","acronym":"MICRO '21","sponsor":["SIGMICRO ACM Special Interest Group on Microarchitectural Research and Processing"]},"container-title":["MICRO-54: 54th Annual IEEE\/ACM International Symposium on Microarchitecture"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3466752.3480047","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/abs\/10.1145\/3466752.3480047","content-type":"text\/html","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3466752.3480047","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3466752.3480047","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:24:52Z","timestamp":1750195492000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3466752.3480047"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":79,"alternative-id":["10.1145\/3466752.3480047","10.1145\/3466752"],"URL":"https:\/\/doi.org\/10.1145\/3466752.3480047","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}