{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,8]],"date-time":"2026-05-08T15:49:37Z","timestamp":1778255377564,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":70,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,10,28]],"date-time":"2023-10-28T00:00:00Z","timestamp":1698451200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,10,28]]},"DOI":"10.1145\/3613424.3614263","type":"proceedings-article","created":{"date-parts":[[2023,12,8]],"date-time":"2023-12-08T17:22:15Z","timestamp":1702056135000},"page":"353-366","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":12,"title":["Sparse-DySta: Sparsity-Aware Dynamic and Static Scheduling for Sparse Multi-DNN Workloads"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2387-5611","authenticated-orcid":false,"given":"Hongxiang","family":"Fan","sequence":"first","affiliation":[{"name":"Samsung AI Cambridge and University of Cambridge, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5181-6251","authenticated-orcid":false,"given":"Stylianos I.","family":"Venieris","sequence":"additional","affiliation":[{"name":"Samsung AI, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2900-430X","authenticated-orcid":false,"given":"Alexandros","family":"Kouris","sequence":"additional","affiliation":[{"name":"Samsung AI and Imperial College London, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2728-8273","authenticated-orcid":false,"given":"Nicholas","family":"Lane","sequence":"additional","affiliation":[{"name":"University of Cambridge and Samsung AI, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,12,8]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3487552.3487863"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA45697.2020.00081"},{"key":"e_1_3_2_1_3_1","volume-title":"Conference on Machine Learning and Systems (MLSys).","author":"Blalock Davis","year":"2020","unstructured":"Davis Blalock , Jose\u00a0Javier Gonzalez\u00a0Ortiz , Jonathan Frankle , and John Guttag . 2020 . What is the State of Neural Network Pruning? . In Conference on Machine Learning and Systems (MLSys). Davis Blalock, Jose\u00a0Javier Gonzalez\u00a0Ortiz, Jonathan Frankle, and John Guttag. 2020. What is the State of Neural Network Pruning?. In Conference on Machine Learning and Systems (MLSys)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2016.2616357"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2019.2910232"},{"key":"e_1_3_2_1_6_1","volume-title":"Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing. In USENIX Annual Technical Conference (ATC).","author":"Choi Seungbeom","year":"2022","unstructured":"Seungbeom Choi , Sunho Lee , Yeonjae Kim , Jongse Park , Youngjin Kwon , and Jaehyuk Huh . 2022 . Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing. In USENIX Annual Technical Conference (ATC). Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh. 2022. Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing. In USENIX Annual Technical Conference (ATC)."},{"key":"e_1_3_2_1_7_1","volume-title":"PREMA: A Predictive Multi-task Scheduling Algorithm for Preemptible Neural Processing Units. In IEEE International Symposium on High Performance Computer Architecture (HPCA).","author":"Choi Yujeong","year":"2020","unstructured":"Yujeong Choi and Minsoo Rhu . 2020 . PREMA: A Predictive Multi-task Scheduling Algorithm for Preemptible Neural Processing Units. In IEEE International Symposium on High Performance Computer Architecture (HPCA). Yujeong Choi and Minsoo Rhu. 2020. PREMA: A Predictive Multi-task Scheduling Algorithm for Preemptible Neural Processing Units. In IEEE International Symposium on High Performance Computer Architecture (HPCA)."},{"key":"e_1_3_2_1_8_1","volume-title":"Masa: Responsive Multi-DNN Inference on the Edge. In IEEE International Conference on Pervasive Computing (PerCom).","author":"Cox Bart","year":"2021","unstructured":"Bart Cox , Jeroen Galjaard , Amirmasoud Ghiassi , Robert Birke , and Lydia\u00a0 Y Chen . 2021 . Masa: Responsive Multi-DNN Inference on the Edge. In IEEE International Conference on Pervasive Computing (PerCom). Bart Cox, Jeroen Galjaard, Amirmasoud Ghiassi, Robert Birke, and Lydia\u00a0Y Chen. 2021. Masa: Responsive Multi-DNN Inference on the Edge. In IEEE International Conference on Pervasive Computing (PerCom)."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2021.3098483"},{"key":"e_1_3_2_1_10_1","volume-title":"GoSPA: An Energy-efficient High-performance Globally Optimized SParse Convolutional Neural Network Accelerator. In Annual International Symposium on Computer Architecture (ISCA).","author":"Deng Chunhua","year":"2021","unstructured":"Chunhua Deng , Yang Sui , Siyu Liao , Xuehai Qian , and Bo Yuan . 2021 . GoSPA: An Energy-efficient High-performance Globally Optimized SParse Convolutional Neural Network Accelerator. In Annual International Symposium on Computer Architecture (ISCA). Chunhua Deng, Yang Sui, Siyu Liao, Xuehai Qian, and Bo Yuan. 2021. GoSPA: An Energy-efficient High-performance Globally Optimized SParse Convolutional Neural Network Accelerator. In Annual International Symposium on Computer Architecture (ISCA)."},{"key":"e_1_3_2_1_11_1","volume-title":"ImageNet: A Large-Scale Hierarchical Image Database. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Deng Jia","year":"2009","unstructured":"Jia Deng 2009 . ImageNet: A Large-Scale Hierarchical Image Database. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Jia Deng 2009. ImageNet: A Large-Scale Hierarchical Image Database. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_1_12_1","volume-title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In ACL.","author":"Devlin Jacob","year":"2019","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2019 . BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In ACL. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In ACL."},{"key":"e_1_3_2_1_13_1","unstructured":"Lukasz Dudziak Thomas Chau Mohamed Abdelfattah Royson Lee Hyeji Kim and Nicholas Lane. 2020. BRP-NAS: Prediction-based NAS using GCNs. In Advances in Neural Information Processing Systems (NeurIPS).  Lukasz Dudziak Thomas Chau Mohamed Abdelfattah Royson Lee Hyeji Kim and Nicholas Lane. 2020. BRP-NAS: Prediction-based NAS using GCNs. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2008.44"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO56248.2022.00050"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.3390\/electronics10040520"},{"key":"e_1_3_2_1_17_1","volume-title":"Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural Networks. In IEEE\/ACM International Symposium on Microarchitecture (MICRO).","author":"Ghodrati Soroush","year":"2020","unstructured":"Soroush Ghodrati , Byung\u00a0Hoon Ahn , Joon\u00a0Kyung Kim , Sean Kinzer , Brahmendra\u00a0Reddy Yatham , Navateja Alla , Hardik Sharma , Mohammad Alian , Eiman Ebrahimi , Nam\u00a0Sung Kim , 2020 . Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural Networks. In IEEE\/ACM International Symposium on Microarchitecture (MICRO). Soroush Ghodrati, Byung\u00a0Hoon Ahn, Joon\u00a0Kyung Kim, Sean Kinzer, Brahmendra\u00a0Reddy Yatham, Navateja Alla, Hardik Sharma, Mohammad Alian, Eiman Ebrahimi, Nam\u00a0Sung Kim, 2020. Planaria: Dynamic Architecture Fission for Spatial Multi-Tenant Acceleration of Deep Neural Networks. In IEEE\/ACM International Symposium on Microarchitecture (MICRO)."},{"key":"e_1_3_2_1_18_1","volume-title":"INFER: INterFerence-aware Estimation of Runtime for Concurrent CNN Execution on DPUs. In IEEE International Conference on Field-Programmable Technology (ICFPT).","author":"Goel Shikha","year":"2020","unstructured":"Shikha Goel , Rajesh Kedia , M Balakrishnan , and Rijurekha Sen . 2020 . INFER: INterFerence-aware Estimation of Runtime for Concurrent CNN Execution on DPUs. In IEEE International Conference on Field-Programmable Technology (ICFPT). Shikha Goel, Rajesh Kedia, M Balakrishnan, and Rijurekha Sen. 2020. INFER: INterFerence-aware Estimation of Runtime for Concurrent CNN Execution on DPUs. In IEEE International Conference on Field-Programmable Technology (ICFPT)."},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA47549.2020.00035"},{"key":"e_1_3_2_1_20_1","volume-title":"Lightweight Self-Attention Mechanism in Neural Networks. In ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA).","author":"Ham Tae\u00a0Jun","year":"2021","unstructured":"Tae\u00a0Jun Ham , Yejin Lee , Seong\u00a0Hoon Seo , Soosung Kim , Hyunji Choi , Sung\u00a0Jun Jung , and Jae\u00a0 W Lee . 2021 . ELSA: Hardware-Software Co-Design for Efficient , Lightweight Self-Attention Mechanism in Neural Networks. In ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). Tae\u00a0Jun Ham, Yejin Lee, Seong\u00a0Hoon Seo, Soosung Kim, Hyunji Choi, Sung\u00a0Jun Jung, and Jae\u00a0W Lee. 2021. ELSA: Hardware-Software Co-Design for Efficient, Lightweight Self-Attention Mechanism in Neural Networks. In ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)."},{"key":"e_1_3_2_1_21_1","volume-title":"Trained Quantization and Huffman Coding. In International Conference on Representation Learning (ICLR).","author":"Han Song","year":"2015","unstructured":"Song Han , Huizi Mao , and William\u00a0 J Dally . 2015 . Deep Compression: Compressing Deep Neural Networks with Pruning , Trained Quantization and Huffman Coding. In International Conference on Representation Learning (ICLR). Song Han, Huizi Mao, and William\u00a0J Dally. 2015. Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. In International Conference on Representation Learning (ICLR)."},{"key":"e_1_3_2_1_22_1","volume-title":"Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"He Kaiming","year":"2016","unstructured":"Kaiming He , Xiangyu Zhang , Shaoqing Ren , and Jian Sun . 2016 . Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_3_2_1_23_1","volume-title":"Channel Pruning for Accelerating Very Deep Neural Networks. In IEEE International Conference on Computer Vision (ICCV). 1389\u20131397","author":"He Yihui","year":"2017","unstructured":"Yihui He , Xiangyu Zhang , and Jian Sun . 2017 . Channel Pruning for Accelerating Very Deep Neural Networks. In IEEE International Conference on Computer Vision (ICCV). 1389\u20131397 . Yihui He, Xiangyu Zhang, and Jian Sun. 2017. Channel Pruning for Accelerating Very Deep Neural Networks. In IEEE International Conference on Computer Vision (ICCV). 1389\u20131397."},{"key":"e_1_3_2_1_24_1","volume-title":"MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv","author":"Howard G","year":"2017","unstructured":"Andrew\u00a0 G Howard , Menglong Zhu , Bo Chen , Dmitry Kalenichenko , Weijun Wang , Tobias Weyand , Marco Andreetto , and Hartwig Adam . 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv ( 2017 ). Andrew\u00a0G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv (2017)."},{"key":"e_1_3_2_1_25_1","volume-title":"IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW).","author":"Ignatov Andrey","unstructured":"Andrey Ignatov , Radu Timofte , Andrei Kulik , Seungsoo Yang , Ke Wang , Felix Baum , Max Wu , Lirong Xu , and Luc Van\u00a0Gool . [n. d.]. AI Benchmark : All About Deep Learning on Smartphones in 2019 . In IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW). Andrey Ignatov, Radu Timofte, Andrei Kulik, Seungsoo Yang, Ke Wang, Felix Baum, Max Wu, Lirong Xu, and Luc Van\u00a0Gool. [n. d.]. AI Benchmark: All About Deep Learning on Smartphones in 2019. In IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW)."},{"key":"e_1_3_2_1_26_1","volume-title":"MLPerf Mobile Inference Benchmark: An Industry-Standard Open-Source Machine Learning Benchmark for On-Device AI. In Conference on Machine Learning and Systems (MLSys).","author":"Janapa\u00a0Reddi Vijay","year":"2022","unstructured":"Vijay Janapa\u00a0Reddi , David Kanter , Peter Mattson , Jared Duke , Thai Nguyen , Ramesh Chukka , Ken Shiring , Koan-Sin Tan , Mark Charlebois , William Chou , 2022 . MLPerf Mobile Inference Benchmark: An Industry-Standard Open-Source Machine Learning Benchmark for On-Device AI. In Conference on Machine Learning and Systems (MLSys). Vijay Janapa\u00a0Reddi, David Kanter, Peter Mattson, Jared Duke, Thai Nguyen, Ramesh Chukka, Ken Shiring, Koan-Sin Tan, Mark Charlebois, William Chou, 2022. MLPerf Mobile Inference Benchmark: An Industry-Standard Open-Source Machine Learning Benchmark for On-Device AI. In Conference on Machine Learning and Systems (MLSys)."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3307650.3322214"},{"key":"e_1_3_2_1_28_1","volume-title":"Sparsity-Aware and Re-configurable NPU Architecture for Samsung Flagship Mobile SoC. In ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA).","author":"Jang Jun-Woo","year":"2021","unstructured":"Jun-Woo Jang , Sehwan Lee , Dongyoung Kim , Hyunsun Park , Ali\u00a0Shafiee Ardestani , Yeongjae Choi , Channoh Kim , Yoojin Kim , Hyeongseok Yu , Hamzah Abdel-Aziz , 2021 . Sparsity-Aware and Re-configurable NPU Architecture for Samsung Flagship Mobile SoC. In ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). Jun-Woo Jang, Sehwan Lee, Dongyoung Kim, Hyunsun Park, Ali\u00a0Shafiee Ardestani, Yeongjae Choi, Channoh Kim, Yoojin Kim, Hyeongseok Yu, Hamzah Abdel-Aziz, 2021. Sparsity-Aware and Re-configurable NPU Architecture for Samsung Flagship Mobile SoC. In ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3498361.3538948"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3579371.3589350"},{"key":"e_1_3_2_1_31_1","volume-title":"MAGMA: An Optimization Framework for Mapping Multiple DNNs on Multiple Accelerator Cores. In IEEE International Symposium on High-Performance Computer Architecture (HPCA).","author":"Kao Sheng-Chun","year":"2022","unstructured":"Sheng-Chun Kao and Tushar Krishna . 2022 . MAGMA: An Optimization Framework for Mapping Multiple DNNs on Multiple Accelerator Cores. In IEEE International Symposium on High-Performance Computer Architecture (HPCA). Sheng-Chun Kao and Tushar Krishna. 2022. MAGMA: An Optimization Framework for Mapping Multiple DNNs on Multiple Accelerator Cores. In IEEE International Symposium on High-Performance Computer Architecture (HPCA)."},{"key":"e_1_3_2_1_32_1","volume-title":"Design Space Exploration of FPGA-Based System with Multiple DNN Accelerators","author":"Kedia Rajesh","year":"2020","unstructured":"Rajesh Kedia , Shikha Goel , M. Balakrishnan , Kolin Paul , and Rijurekha Sen . 2020. Design Space Exploration of FPGA-Based System with Multiple DNN Accelerators . IEEE Embedded Systems Letters (ESL) ( 2020 ). Rajesh Kedia, Shikha Goel, M. Balakrishnan, Kolin Paul, and Rijurekha Sen. 2020. Design Space Exploration of FPGA-Based System with Multiple DNN Accelerators. IEEE Embedded Systems Letters (ESL) (2020)."},{"key":"e_1_3_2_1_33_1","volume-title":"ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS).","author":"Kim Seah","year":"2024","unstructured":"Seah Kim , Hyoukjun Kwon , Jinook Song , Jihyuck Jo , Yu-Hsin Chen , Liangzhen Lai , and Vikas Chandra . 2024 . SDRM3: A Dynamic Scheduler for Dynamic Real-time Multi-model ML Workloads . In ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS). Seah Kim, Hyoukjun Kwon, Jinook Song, Jihyuck Jo, Yu-Hsin Chen, Liangzhen Lai, and Vikas Chandra. 2024. SDRM3: A Dynamic Scheduler for Dynamic Real-time Multi-model ML Workloads. In ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)."},{"key":"e_1_3_2_1_34_1","volume-title":"Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks. In International Conference on Machine Learning (ICML).","author":"Kurtz Mark","year":"2020","unstructured":"Mark Kurtz , Justin Kopinsky , Rati Gelashvili , Alexander Matveev , John Carr , Michael Goin , William Leiserson , Sage Moore , Nir Shavit , and Dan Alistarh . 2020 . Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks. In International Conference on Machine Learning (ICML). Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Nir Shavit, and Dan Alistarh. 2020. Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks. In International Conference on Machine Learning (ICML)."},{"key":"e_1_3_2_1_35_1","volume-title":"Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks. In International Conference on Machine Learning (ICML). 5533\u20135543","author":"Kurtz Mark","year":"2020","unstructured":"Mark Kurtz , Justin Kopinsky , Rati Gelashvili , Alexander Matveev , John Carr , Michael Goin , William Leiserson , Sage Moore , Nir Shavit , and Dan Alistarh . 2020 . Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks. In International Conference on Machine Learning (ICML). 5533\u20135543 . Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Nir Shavit, and Dan Alistarh. 2020. Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural Networks. In International Conference on Machine Learning (ICML). 5533\u20135543."},{"key":"e_1_3_2_1_36_1","volume-title":"Heterogeneous Dataflow Accelerators for Multi-DNN Workloads. In IEEE International Symposium on High-Performance Computer Architecture (HPCA).","author":"Kwon Hyoukjun","year":"2021","unstructured":"Hyoukjun Kwon , Liangzhen Lai , Michael Pellauer , Tushar Krishna , Yu-Hsin Chen , and Vikas Chandra . 2021 . Heterogeneous Dataflow Accelerators for Multi-DNN Workloads. In IEEE International Symposium on High-Performance Computer Architecture (HPCA). Hyoukjun Kwon, Liangzhen Lai, Michael Pellauer, Tushar Krishna, Yu-Hsin Chen, and Vikas Chandra. 2021. Heterogeneous Dataflow Accelerators for Multi-DNN Workloads. In IEEE International Symposium on High-Performance Computer Architecture (HPCA)."},{"key":"e_1_3_2_1_37_1","volume-title":"XRBench: An Extended Reality (XR) Machine Learning Benchmark Suite for the Metaverse. In Conference on Machine Learning and Systems (MLSys).","author":"Kwon Hyoukjun","year":"2023","unstructured":"Hyoukjun Kwon , Krishnakumar Nair , Jamin Seo , Jason Yik , Debabrata Mohapatra , Dongyuan Zhan , Jinook Song , Peter Capak , Peizhao Zhang , Peter Vajda , 2023 . XRBench: An Extended Reality (XR) Machine Learning Benchmark Suite for the Metaverse. In Conference on Machine Learning and Systems (MLSys). Hyoukjun Kwon, Krishnakumar Nair, Jamin Seo, Jason Yik, Debabrata Mohapatra, Dongyuan Zhan, Jinook Song, Peter Capak, Peizhao Zhang, Peter Vajda, 2023. XRBench: An Extended Reality (XR) Machine Learning Benchmark Suite for the Metaverse. In Conference on Machine Learning and Systems (MLSys)."},{"key":"e_1_3_2_1_38_1","volume-title":"Dataflow Mirroring: Architectural Support for Highly Efficient Fine-Grained Spatial Multitasking on Systolic-Array NPUs. In Design Automation Conference (DAC).","author":"Lee Jounghoo","year":"2021","unstructured":"Jounghoo Lee , Jinwoo Choi , Jaeyeon Kim , Jinho Lee , and Youngsok Kim . 2021 . Dataflow Mirroring: Architectural Support for Highly Efficient Fine-Grained Spatial Multitasking on Systolic-Array NPUs. In Design Automation Conference (DAC). Jounghoo Lee, Jinwoo Choi, Jaeyeon Kim, Jinho Lee, and Youngsok Kim. 2021. Dataflow Mirroring: Architectural Support for Highly Efficient Fine-Grained Spatial Multitasking on Systolic-Array NPUs. In Design Automation Conference (DAC)."},{"key":"e_1_3_2_1_39_1","volume-title":"BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL.","author":"Lewis Mike","year":"2020","unstructured":"Mike Lewis , Yinhan Liu , Naman Goyal , Marjan Ghazvininejad , Abdelrahman Mohamed , Omer Levy , Ves Stoyanov , and Luke Zettlemoyer . 2020 . BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL. Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. In ACL."},{"key":"e_1_3_2_1_40_1","volume-title":"Microsoft COCO: Common Objects in Context. In European Conference on Computer Vision (ECCV). Springer, 740\u2013755","author":"Lin Tsung-Yi","year":"2014","unstructured":"Tsung-Yi Lin , Michael Maire , Serge Belongie , James Hays , Pietro Perona , Deva Ramanan , Piotr Doll\u00e1r , and C\u00a0Lawrence Zitnick . 2014 . Microsoft COCO: Common Objects in Context. In European Conference on Computer Vision (ECCV). Springer, 740\u2013755 . Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C\u00a0Lawrence Zitnick. 2014. Microsoft COCO: Common Objects in Context. In European Conference on Computer Vision (ECCV). Springer, 740\u2013755."},{"key":"e_1_3_2_1_41_1","volume-title":"SSD: Single Shot Multibox Detector. In European Conference on Computer Vision (ECCV). Springer, 21\u201337","author":"Liu Wei","year":"2016","unstructured":"Wei Liu , Dragomir Anguelov , Dumitru Erhan , Christian Szegedy , Scott Reed , Cheng-Yang Fu , and Alexander\u00a0 C Berg . 2016 . SSD: Single Shot Multibox Detector. In European Conference on Computer Vision (ECCV). Springer, 21\u201337 . Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander\u00a0C Berg. 2016. SSD: Single Shot Multibox Detector. In European Conference on Computer Vision (ECCV). Springer, 21\u201337."},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2018.10.010"},{"key":"e_1_3_2_1_43_1","volume-title":"Sanger: A Co-Design Framework for Enabling Sparse Attention Using Reconfigurable Architecture. In 54th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO).","author":"Lu Liqiang","year":"2021","unstructured":"Liqiang Lu , Yicheng Jin , Hangrui Bi , Zizhang Luo , Peng Li , Tao Wang , and Yun Liang . 2021 . Sanger: A Co-Design Framework for Enabling Sparse Attention Using Reconfigurable Architecture. In 54th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO). Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. 2021. Sanger: A Co-Design Framework for Enabling Sparse Attention Using Reconfigurable Architecture. In 54th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO)."},{"key":"e_1_3_2_1_44_1","unstructured":"Gaurav Menghani. [n. d.]. Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller Faster and Better. ACM Computing Surveys (CSUR) ([n. d.]).  Gaurav Menghani. [n. d.]. Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller Faster and Better. ACM Computing Surveys (CSUR) ([n. d.])."},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"crossref","unstructured":"Francisco Mu\u00f1oz Mart\u00ednez Raveesh Garg Michael Pellauer Jos\u00e9\u00a0L. Abell\u00e1n Manuel\u00a0E. Acacio and Tushar Krishna. 2023. Flexagon: A Multi-Dataflow Sparse-Sparse Matrix Multiplication Accelerator for Efficient DNN Processing. In ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS).  Francisco Mu\u00f1oz Mart\u00ednez Raveesh Garg Michael Pellauer Jos\u00e9\u00a0L. Abell\u00e1n Manuel\u00a0E. Acacio and Tushar Krishna. 2023. Flexagon: A Multi-Dataflow Sparse-Sparse Matrix Multiplication Accelerator for Efficient DNN Processing. In ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS).","DOI":"10.1145\/3582016.3582069"},{"key":"e_1_3_2_1_46_1","volume-title":"STONNE: Enabling Cycle-Level Microarchitectural Simulation for DNN Inference Accelerators. In IEEE International Symposium on Workload Characterization (IISWC). IEEE, 201\u2013213","author":"Mu\u00f1oz-Mart\u00ednez Francisco","year":"2021","unstructured":"Francisco Mu\u00f1oz-Mart\u00ednez , Jos\u00e9\u00a0 L Abell\u00e1n , Manuel\u00a0 E Acacio , and Tushar Krishna . 2021 . STONNE: Enabling Cycle-Level Microarchitectural Simulation for DNN Inference Accelerators. In IEEE International Symposium on Workload Characterization (IISWC). IEEE, 201\u2013213 . Francisco Mu\u00f1oz-Mart\u00ednez, Jos\u00e9\u00a0L Abell\u00e1n, Manuel\u00a0E Acacio, and Tushar Krishna. 2021. STONNE: Enabling Cycle-Level Microarchitectural Simulation for DNN Inference Accelerators. In IEEE International Symposium on Workload Characterization (IISWC). IEEE, 201\u2013213."},{"key":"e_1_3_2_1_47_1","volume-title":"Rectified Linear Units Improve Restricted Boltzmann Machines. In International Conference on Machine Learning (ICML).","author":"Nair Vinod","year":"2010","unstructured":"Vinod Nair and Geoffrey\u00a0 E Hinton . 2010 . Rectified Linear Units Improve Restricted Boltzmann Machines. In International Conference on Machine Learning (ICML). Vinod Nair and Geoffrey\u00a0E Hinton. 2010. Rectified Linear Units Improve Restricted Boltzmann Machines. In International Conference on Machine Learning (ICML)."},{"key":"e_1_3_2_1_48_1","unstructured":"NVIDIA. 2021. Accelerating Inference with Sparsity using Ampere and TensorRT. https:\/\/developer.nvidia.com\/blog\/accelerating-inference-with-sparsity-using-ampere-and-tensorrt\/. Accessed: 2023\/10\/05 06:02:02.  NVIDIA. 2021. Accelerating Inference with Sparsity using Ampere and TensorRT. https:\/\/developer.nvidia.com\/blog\/accelerating-inference-with-sparsity-using-ampere-and-tensorrt\/. Accessed: 2023\/10\/05 06:02:02."},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCA51647.2021.00056"},{"key":"e_1_3_2_1_50_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In Advances in Neural Information Processing Systems (NeurIPS). Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_2_1_51_1","volume-title":"SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training. In IEEE International Symposium on High Performance Computer Architecture (HPCA).","author":"Qin Eric","year":"2020","unstructured":"Eric Qin , Ananda Samajdar , Hyoukjun Kwon , Vineet Nadella , Sudarshan Srinivasan , Dipankar Das , Bharat Kaul , and Tushar Krishna . 2020 . SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training. In IEEE International Symposium on High Performance Computer Architecture (HPCA). Eric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella, Sudarshan Srinivasan, Dipankar Das, Bharat Kaul, and Tushar Krishna. 2020. SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN Training. In IEEE International Symposium on High Performance Computer Architecture (HPCA)."},{"key":"e_1_3_2_1_52_1","volume-title":"DOTA: Detect and Omit Weak Attentions for Scalable Transformer Acceleration. In 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS).","author":"Qu Zheng","year":"2022","unstructured":"Zheng Qu , Liu Liu , Fengbin Tu , Zhaodong Chen , Yufei Ding , and Yuan Xie . 2022 . DOTA: Detect and Omit Weak Attentions for Scalable Transformer Acceleration. In 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS). Zheng Qu, Liu Liu, Fengbin Tu, Zhaodong Chen, Yufei Ding, and Yuan Xie. 2022. DOTA: Detect and Omit Weak Attentions for Scalable Transformer Acceleration. In 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)."},{"key":"e_1_3_2_1_53_1","volume-title":"Language models are unsupervised multitask learners. OpenAI blog 1, 8","author":"Radford Alec","year":"2019","unstructured":"Alec Radford , Jeffrey Wu , Rewon Child , David Luan , Dario Amodei , and Ilya Sutskever . 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 ( 2019 ), 9. Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. OpenAI blog 1, 8 (2019), 9."},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_3_2_1_55_1","volume-title":"MLPerf Inference Benchmark. In International Symposium on Computer Architecture (ISCA).","author":"Reddi Vijay\u00a0Janapa","year":"2020","unstructured":"Vijay\u00a0Janapa Reddi , Christine Cheng , David Kanter , Peter Mattson , Guenther Schmuelling , Carole-Jean Wu , Brian Anderson , Maximilien Breughe , Mark Charlebois , William Chou , Ramesh Chukka , Cody Coleman , Sam Davis , Pan Deng , Greg Diamos , Jared Duke , Dave Fick , J.\u00a0 Scott Gardner , Itay Hubara , Sachin Idgunji , Thomas\u00a0 B. Jablin , Jeff Jiao , Tom\u00a0 St. John , Pankaj Kanwar , David Lee , Jeffery Liao , Anton Lokhmotov , Francisco Massa , Peng Meng , Paulius Micikevicius , Colin Osborne , Gennady Pekhimenko , Arun Tejusve\u00a0Raghunath Rajan , Dilip Sequeira , Ashish Sirasao , Fei Sun , Hanlin Tang , Michael Thomson , Frank Wei , Ephrem Wu , Lingjie Xu , Koichi Yamada , Bing Yu , George Yuan , Aaron Zhong , Peizhao Zhang , and Yuchen Zhou . 2020 . MLPerf Inference Benchmark. In International Symposium on Computer Architecture (ISCA). Vijay\u00a0Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson, Guenther Schmuelling, Carole-Jean Wu, Brian Anderson, Maximilien Breughe, Mark Charlebois, William Chou, Ramesh Chukka, Cody Coleman, Sam Davis, Pan Deng, Greg Diamos, Jared Duke, Dave Fick, J.\u00a0Scott Gardner, Itay Hubara, Sachin Idgunji, Thomas\u00a0B. Jablin, Jeff Jiao, Tom\u00a0St. John, Pankaj Kanwar, David Lee, Jeffery Liao, Anton Lokhmotov, Francisco Massa, Peng Meng, Paulius Micikevicius, Colin Osborne, Gennady Pekhimenko, Arun Tejusve\u00a0Raghunath Rajan, Dilip Sequeira, Ashish Sirasao, Fei Sun, Hanlin Tang, Michael Thomson, Frank Wei, Ephrem Wu, Lingjie Xu, Koichi Yamada, Bing Yu, George Yuan, Aaron Zhong, Peizhao Zhang, and Yuchen Zhou. 2020. MLPerf Inference Benchmark. In International Symposium on Computer Architecture (ISCA)."},{"key":"e_1_3_2_1_56_1","volume-title":"INFaaS: Automated Model-less Inference Serving. In USENIX Annual Technical Conference (ATC).","author":"Romero Francisco","year":"2021","unstructured":"Francisco Romero , Qian Li , Neeraja\u00a0 J. Yadwadkar , and Christos Kozyrakis . 2021 . INFaaS: Automated Model-less Inference Serving. In USENIX Annual Technical Conference (ATC). Francisco Romero, Qian Li, Neeraja\u00a0J. Yadwadkar, and Christos Kozyrakis. 2021. INFaaS: Automated Model-less Inference Serving. In USENIX Annual Technical Conference (ATC)."},{"key":"e_1_3_2_1_57_1","volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations (ICLR).","author":"Simonyan Karen","year":"2015","unstructured":"Karen Simonyan and Andrew Zisserman . 2015 . Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations (ICLR). Karen Simonyan and Andrew Zisserman. 2015. Very Deep Convolutional Networks for Large-Scale Image Recognition. In International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2018.00072"},{"key":"e_1_3_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2022.3176845"},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/W18-5446"},{"key":"e_1_3_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/3214306"},{"key":"e_1_3_2_1_62_1","volume-title":"SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning. IEEE International Symposium on High Performance Computer Architecture (HPCA)","author":"Wang Hanrui","year":"2021","unstructured":"Hanrui Wang , Zhekai Zhang , and Song Han . 2021 . SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning. IEEE International Symposium on High Performance Computer Architecture (HPCA) (2021). Hanrui Wang, Zhekai Zhang, and Song Han. 2021. SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head Pruning. IEEE International Symposium on High Performance Computer Architecture (HPCA) (2021)."},{"key":"e_1_3_2_1_63_1","volume-title":"Deep Retinex Decomposition for Low-Light Enhancement. In British Machine Vision Conference (BMVC).","author":"Wei Chen","year":"2018","unstructured":"Chen Wei , Wenjing Wang , Wenhan Yang , and Jiaying Liu . 2018 . Deep Retinex Decomposition for Low-Light Enhancement. In British Machine Vision Conference (BMVC). Chen Wei, Wenjing Wang, Wenhan Yang, and Jiaying Liu. 2018. Deep Retinex Decomposition for Low-Light Enhancement. In British Machine Vision Conference (BMVC)."},{"key":"e_1_3_2_1_64_1","volume-title":"Sparseloop: An Analytical Approach To Sparse Tensor Accelerator Modeling. In IEEE\/ACM International Symposium on Microarchitecture (MICRO).","author":"Wu Yannan\u00a0Nellie","year":"2022","unstructured":"Yannan\u00a0Nellie Wu , Po-An Tsai , Angshuman Parashar , Vivienne Sze , and Joel\u00a0 S Emer . 2022 . Sparseloop: An Analytical Approach To Sparse Tensor Accelerator Modeling. In IEEE\/ACM International Symposium on Microarchitecture (MICRO). Yannan\u00a0Nellie Wu, Po-An Tsai, Angshuman Parashar, Vivienne Sze, and Joel\u00a0S Emer. 2022. Sparseloop: An Analytical Approach To Sparse Tensor Accelerator Modeling. In IEEE\/ACM International Symposium on Microarchitecture (MICRO)."},{"key":"e_1_3_2_1_65_1","volume-title":"Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple Tasks. In Design Automation Conference (DAC).","author":"Yang Lei","year":"2020","unstructured":"Lei Yang 2020 . Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple Tasks. In Design Automation Conference (DAC). Lei Yang 2020. Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple Tasks. In Design Automation Conference (DAC)."},{"key":"e_1_3_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372224.3419192"},{"key":"e_1_3_2_1_67_1","volume-title":"Serving Multi-DNN Workloads on FPGAs: a Coordinated Architecture, Scheduling, and Mapping Perspective","author":"Zeng Shulin","year":"2022","unstructured":"Shulin Zeng , Guohao Dai , Niansong Zhang , Xinhao Yang , Haoyu Zhang , Zhenhua Zhu , Huazhong Yang , and Yu Wang . 2022. Serving Multi-DNN Workloads on FPGAs: a Coordinated Architecture, Scheduling, and Mapping Perspective . IEEE Transactions on Computers (TC) ( 2022 ). Shulin Zeng, Guohao Dai, Niansong Zhang, Xinhao Yang, Haoyu Zhang, Zhenhua Zhu, Huazhong Yang, and Yu Wang. 2022. Serving Multi-DNN Workloads on FPGAs: a Coordinated Architecture, Scheduling, and Mapping Perspective. IEEE Transactions on Computers (TC) (2022)."},{"key":"e_1_3_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/3458864.3467882"},{"key":"e_1_3_2_1_69_1","volume-title":"International Conference on Representation Learning (ICLR).","author":"Zhou Aojun","year":"2021","unstructured":"Aojun Zhou , Yukun Ma , Junnan Zhu , Jianbo Liu , Zhijie Zhang , Kun Yuan , Wenxiu Sun , and Hongsheng Li . 2021 . Learning N: M Fine-grained Structured Sparse Neural Networks from Scratch . In International Conference on Representation Learning (ICLR). Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li. 2021. Learning N: M Fine-grained Structured Sparse Neural Networks from Scratch. In International Conference on Representation Learning (ICLR)."},{"key":"e_1_3_2_1_70_1","volume-title":"Energon: Towards Efficient Acceleration of Transformers Using Dynamic Sparse Attention","author":"Zhou Zhe","year":"2021","unstructured":"Zhe Zhou , Junlin Liu , Zhenyu Gu , and Guangyu Sun . 2021 . Energon: Towards Efficient Acceleration of Transformers Using Dynamic Sparse Attention . IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) ( 2021). Zhe Zhou, Junlin Liu, Zhenyu Gu, and Guangyu Sun. 2021. Energon: Towards Efficient Acceleration of Transformers Using Dynamic Sparse Attention. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems (TCAD) (2021)."}],"event":{"name":"MICRO '23: 56th Annual IEEE\/ACM International Symposium on Microarchitecture","location":"Toronto ON Canada","acronym":"MICRO '23","sponsor":["SIGMICRO ACM Special Interest Group on Microarchitectural Research and Processing"]},"container-title":["56th Annual IEEE\/ACM International Symposium on Microarchitecture"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3613424.3614263","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3613424.3614263","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:21Z","timestamp":1750178781000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3613424.3614263"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,10,28]]},"references-count":70,"alternative-id":["10.1145\/3613424.3614263","10.1145\/3613424"],"URL":"https:\/\/doi.org\/10.1145\/3613424.3614263","relation":{},"subject":[],"published":{"date-parts":[[2023,10,28]]},"assertion":[{"value":"2023-12-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}