{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T23:08:30Z","timestamp":1784934510828,"version":"3.55.0"},"reference-count":62,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,2,15]],"date-time":"2024-02-15T00:00:00Z","timestamp":1707955200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&D Program of China","doi-asserted-by":"crossref","award":["2022YFB4501404"],"award-info":[{"award-number":["2022YFB4501404"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100005090","name":"Beijing Nova Program","doi-asserted-by":"crossref","award":["20230484420, and 20220484054"],"award-info":[{"award-number":["20230484420, and 20220484054"]}],"id":[{"id":"10.13039\/501100005090","id-type":"DOI","asserted-by":"crossref"}]},{"name":"CAS Project for Young Scientists in Basic Research","award":["YSBR- 029"],"award-info":[{"award-number":["YSBR- 029"]}]},{"name":"CAS Project for Youth Innovation Promotion Association and Open Research Projects of Zhejiang Lab","award":["2022PB0AB01"],"award-info":[{"award-number":["2022PB0AB01"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2024,3,31]]},"abstract":"<jats:p>Dataflow architectures can achieve much better performance and higher efficiency than general-purpose core, approaching the performance of a specialized design while retaining programmability. However, advanced application scenarios place higher demands on the hardware in terms of cross-domain and multi-batch processing. In this article, we propose a unified scale-vector architecture that can work in multiple modes and adapt to diverse algorithms and requirements efficiently. First, a novel reconfigurable interconnection structure is proposed, which can organize execution units into different cluster typologies as a way to accommodate different data-level parallelism. Second, we decouple threads within each DFG node into consecutive pipeline stages and provide architectural support. By time-multiplexing during these stages, dataflow hardware can achieve much higher utilization and performance. In addition, the task-based program model can also exploit multi-level parallelism and deploy applications efficiently. Evaluated in a wide range of benchmarks, including digital signal processing algorithms, CNNs, and scientific computing algorithms, our design attains up to 11.95\u00d7 energy efficiency (performance-per-watt) improvement over GPU (V100), and 2.01\u00d7 energy efficiency improvement over state-of-the-art dataflow architectures.<\/jats:p>","DOI":"10.1145\/3637906","type":"journal-article","created":{"date-parts":[[2023,12,18]],"date-time":"2023-12-18T11:52:30Z","timestamp":1702900350000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Improving Utilization of Dataflow Unit for Multi-Batch Processing"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5950-7370","authenticated-orcid":false,"given":"Zhihua","family":"Fan","sequence":"first","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4069-2251","authenticated-orcid":false,"given":"Wenming","family":"Li","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-8853-9915","authenticated-orcid":false,"given":"Zhen","family":"Wang","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-6744-1225","authenticated-orcid":false,"given":"Yu","family":"Yang","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4598-1685","authenticated-orcid":false,"given":"Xiaochun","family":"Ye","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5219-0908","authenticated-orcid":false,"given":"Dongrui","family":"Fan","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1953-1392","authenticated-orcid":false,"given":"Ninghui","family":"Sun","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0494-6332","authenticated-orcid":false,"given":"Xuejun","family":"An","sequence":"additional","affiliation":[{"name":"State Key Lab of Processors, Institute of Computing Technology, Chinese Academy of Sciences and University of Chinese Academy of Sciences, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,2,15]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"crossref","first-page":"1160","DOI":"10.1109\/MICRO56248.2022.00083","volume-title":"55th IEEE\/ACM International Symposium on Microarchitecture (MICRO 2022) (Chicago, IL, October 1-5, 2022","author":"Baskaran Saambhavi","year":"2022","unstructured":"Saambhavi Baskaran, Mahmut Taylan Kandemir, and Jack Sampson. 2022. An architecture interface and offload model for low-overhead, near-data, distributed accelerators. In 55th IEEE\/ACM International Symposium on Microarchitecture (MICRO 2022) (Chicago, IL, October 1-5, 2022). IEEE, 1160\u20131177."},{"key":"e_1_3_2_3_2","first-page":"269","volume-title":"Architectural Support for Programming Languages and Operating Systems (ASPLOS 2014) (Salt Lake City, UT, March 1-5, 2014)","author":"Chen Tianshi","year":"2014","unstructured":"Tianshi Chen, Zidong Du, Ninghui Sun, Jia Wang, Chengyong Wu, Yunji Chen, and Olivier Temam. 2014. DianNao: A small-footprint high-throughput accelerator for ubiquitous machine-learning. In Architectural Support for Programming Languages and Operating Systems (ASPLOS 2014) (Salt Lake City, UT, March 1-5, 2014), Rajeev Balasubramonian, Al Davis, and Sarita V. Adve (Eds.). ACM, 269\u2013284."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/2996864"},{"key":"e_1_3_2_5_2","first-page":"609","volume-title":"47th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO 2014) (Cambridge, United Kingdom, December 13-17, 2014)","author":"Chen Yunji","year":"2014","unstructured":"Yunji Chen, Tao Luo, Shaoli Liu, Shijin Zhang, Liqiang He, Jia Wang, Ling Li, Tianshi Chen, Zhiwei Xu, Ninghui Sun, and Olivier Temam. 2014. DaDianNao: A machine-learning supercomputer. In 47th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO 2014) (Cambridge, United Kingdom, December 13-17, 2014). IEEE Computer Society, 609\u2013622."},{"key":"e_1_3_2_6_2","first-page":"1","volume-title":"ASPLOS","author":"Dadu Vidushi","year":"2022","unstructured":"Vidushi Dadu and Tony Nowatzki. 2022. TaskStream: Accelerating task-parallel workloads by recovering program structure. In ASPLOS. 1\u201313."},{"key":"e_1_3_2_7_2","unstructured":"Groq Dale Southard Ecosystem Solutions Distinguished Architect. 2019. Tensor streaming architecture delivers unmatched Performance for compute-intensive workloads. https:\/\/groq.com\/wp-content\/uploads\/2019\/10\/Groq_Whitepaper_2019Oct.pdf"},{"key":"e_1_3_2_8_2","first-page":"1110","volume-title":"48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021)","author":"Deng Chunhua","year":"2021","unstructured":"Chunhua Deng, Yang Sui, Siyu Liao, Xuehai Qian, and Bo Yuan. 2021. GoSPA: An energy-efficient high-performance globally optimized sparse convolutional neural network accelerator. In 48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021). IEEE, 1110\u20131123."},{"issue":"4","key":"e_1_3_2_9_2","doi-asserted-by":"crossref","first-page":"668","DOI":"10.1109\/JPROC.1999.752522","article-title":"Design of ion-implanted MOSFET\u2019s with very small physical dimensions","volume":"87","author":"Dennard Robert H.","year":"1974","unstructured":"Robert H. Dennard, Fritz H. Gaensslen, Hwa-Nien Yu, V. Leo Rideout, Ernest Bassous, and Andre R. Leblanc. 1974. Design of ion-implanted MOSFET\u2019s with very small physical dimensions. Proc. IEEE 87, 4 (1974), 668\u2013678.","journal-title":"Proc. IEEE"},{"key":"e_1_3_2_10_2","series-title":"Programming Symposium, Proceedings Colloque sur la Programmation (Paris, France, April 9-11, 1974)","first-page":"362","volume":"19","author":"Dennis Jack B.","year":"1974","unstructured":"Jack B. Dennis. 1974. First version of a data flow procedure language. In Programming Symposium, Proceedings Colloque sur la Programmation (Paris, France, April 9-11, 1974)(Lecture Notes in Computer Science, Vol. 19), Bernard J. Robinet (Ed.). Springer, 362\u2013376."},{"key":"e_1_3_2_11_2","first-page":"596","volume-title":"IEEE International Symposium on High Performance Computer Architecture (HPCA 2018) (Vienna, Austria, February 24-28, 2018)","author":"Fan Dongrui","year":"2018","unstructured":"Dongrui Fan, Wenming Li, Xiaochun Ye, Da Wang, Hao Zhang, Zhimin Tang, and Ninghui Sun. 2018. SmarCo: An efficient many-core processor for high-throughput applications in datacenters. In IEEE International Symposium on High Performance Computer Architecture (HPCA 2018) (Vienna, Austria, February 24-28, 2018). IEEE Computer Society, 596\u2013607."},{"key":"e_1_3_2_12_2","first-page":"1","volume-title":"EuroPar","author":"Fan Zhihua","year":"2023","unstructured":"Zhihua Fan and Wenming Li. 2023. Improving utilization of dataflow architectures through software and hardware co-design. In EuroPar. 1\u201314."},{"key":"e_1_3_2_13_2","first-page":"1","volume-title":"25th IEEE International Symposium on High Performance Computer Architecture (HPCA 2019) (Washington, DC, February 16-20, 2019)","author":"Fuchs Adi","year":"2019","unstructured":"Adi Fuchs and David Wentzlaff. 2019. The accelerator wall: Limits of chip specialization. In 25th IEEE International Symposium on High Performance Computer Architecture (HPCA 2019) (Washington, DC, February 16-20, 2019). IEEE, 1\u201314."},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1109\/SBAC-PADW.2014.30","volume-title":"2014 International Symposium on Computer Architecture and High Performance Computing Workshop","author":"Giorgi Roberto","year":"2014","unstructured":"Roberto Giorgi and Paolo Faraboschi. 2014. An introduction to DF-threads and their execution model. In 2014 International Symposium on Computer Architecture and High Performance Computing Workshop. 60\u201365."},{"key":"e_1_3_2_15_2","first-page":"876","volume-title":"IEEE International Symposium on High-Performance Computer Architecture (HPCA 2022) (Seoul, South Korea, April 2-6, 2022)","author":"Gudaparthi Sumanth","year":"2022","unstructured":"Sumanth Gudaparthi, Sarabjeet Singh, Surya Narayanan, Rajeev Balasubramonian, and Visvesh Sathe. 2022. CANDLES: Channel-aware novel dataflow-microarchitecture co-design for low energy sparse neural network acceleration. In IEEE International Symposium on High-Performance Computer Architecture (HPCA 2022) (Seoul, South Korea, April 2-6, 2022). IEEE, 876\u2013891."},{"key":"e_1_3_2_16_2","first-page":"835","volume-title":"55th IEEE\/ACM International Symposium on Microarchitecture (MICRO 2022) (Chicago, IL, October 1-5, 2022)","author":"Haj-Yahya Jawad","year":"2022","unstructured":"Jawad Haj-Yahya, Haris Volos, Davide B. Bartolini, Georgia Antoniou, Jeremie S. Kim, Zhe Wang, Kleovoulos Kalaitzidis, Tom Rollet, Zhirui Chen, Ye Geng, Onur Mutlu, and Yiannakis Sazeides. 2022. AgileWatts: An energy-efficient CPU core idle-state architecture for latency-sensitive server applications. In 55th IEEE\/ACM International Symposium on Microarchitecture (MICRO 2022) (Chicago, IL, October 1-5, 2022). IEEE, 835\u2013850."},{"key":"e_1_3_2_17_2","first-page":"191","volume-title":"MICRO","author":"Ham Tae Jun","year":"2015","unstructured":"Tae Jun Ham, Juan L. Arag\u00f3n, and Margaret Martonosi. 2015. DeSC: Decoupled supply-compute communication management for heterogeneous architectures. In MICRO. 191\u2013203."},{"key":"e_1_3_2_18_2","first-page":"37","volume-title":"37th International Symposium on Computer Architecture (ISCA 2010), (Saint-Malo, France","author":"Hameed Rehan","year":"2010","unstructured":"Rehan Hameed, Wajahat Qadeer, Megan Wachs, Omid Azizi, Alex Solomatnikov, Benjamin C. Lee, Stephen Richardson, Christos Kozyrakis, and Mark Horowitz. 2010. Understanding sources of inefficiency in general-purpose chips. In 37th International Symposium on Computer Architecture (ISCA 2010), (Saint-Malo, France, June 19-23, 2010), Andr\u00e9 Seznec, Uri C. Weiser, and Ronny Ronen (Eds.). ACM, 37\u201347."},{"key":"e_1_3_2_19_2","first-page":"57","volume-title":"55th IEEE\/ACM International Symposium on Microarchitecture (MICRO 2022) (Chicago, IL, October 1-5, 2022)","author":"Hao Yifan","year":"2022","unstructured":"Yifan Hao, Yongwei Zhao, Chenxiao Liu, Zidong Du, Shuyao Cheng, Xiaqing Li, Xing Hu, Qi Guo, Zhiwei Xu, and Tianshi Chen. 2022. Cambricon-P: A bitflow architecture for arbitrary precision computing. In 55th IEEE\/ACM International Symposium on Microarchitecture (MICRO 2022) (Chicago, IL, October 1-5, 2022). IEEE, 57\u201372."},{"key":"e_1_3_2_20_2","first-page":"54","volume-title":"IEEE International Symposium on High-Performance Computer Architecture (HPCA 2021) (Seoul, South Korea, February 27 - March 3, 2021)","author":"Kinzer Sean","year":"2021","unstructured":"Sean Kinzer, Joon Kyung Kim, Soroush Ghodrati, Brahmendra Reddy Yatham, Alric Althoff, Divya Mahajan, Sorin Lerner, and Hadi Esmaeilzadeh. 2021. A computational stack for cross-domain acceleration. In IEEE International Symposium on High-Performance Computer Architecture (HPCA 2021) (Seoul, South Korea, February 27 - March 3, 2021). IEEE, 54\u201370."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3065386"},{"key":"e_1_3_2_22_2","first-page":"1359","volume-title":"55th IEEE\/ACM International Symposium on Microarchitecture (MICRO 2022) (Chicago, IL, October 1-5, 2022)","author":"Lee Hunjun","year":"2022","unstructured":"Hunjun Lee, Minseop Kim, Dongmoon Min, Joonsung Kim, Jongwon Back, Honam Yoo, Jong-Ho Lee, and Jangwoo Kim. 2022. 3D-FPIM: An extreme energy-efficient DNN acceleration system using 3D NAND flash-based in-situ PIM unit. In 55th IEEE\/ACM International Symposium on Microarchitecture (MICRO 2022) (Chicago, IL, October 1-5, 2022). IEEE, 1359\u20131376."},{"key":"e_1_3_2_23_2","first-page":"169","volume-title":"IEEE International Symposium on High-Performance Computer Architecture (HPCA 2022) (Seoul, South Korea, April 2-6, 2022)","author":"Lee Yejin","year":"2022","unstructured":"Yejin Lee, Hyunji Choi, Sunhong Min, Hyunseung Lee, Sangwon Beak, Dawoon Jeong, Jae W. Lee, and Tae Jun Ham. 2022. ANNA: Specialized architecture for approximate nearest neighbor search. In IEEE International Symposium on High-Performance Computer Architecture (HPCA 2022) (Seoul, South Korea, April 2-6, 2022). IEEE, 169\u2013183."},{"key":"e_1_3_2_24_2","first-page":"992","volume-title":"54th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201921) (Virtual Event, Greece, October 18-22, 2021)","author":"Li Shiyu","year":"2021","unstructured":"Shiyu Li, Edward Hanson, Xuehai Qian, Hai (Helen) Li, and Yiran Chen. 2021. ESCALATE: Boosting the efficiency of sparse CNN accelerator with kernel decomposition. In 54th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201921) (Virtual Event, Greece, October 18-22, 2021). ACM, 992\u20131004."},{"key":"e_1_3_2_25_2","first-page":"352","volume-title":"25th IET Irish Signals and Systems Conference 2014 and 2014 China-Ireland International Conference on Information and Communications Technologies","author":"Loughlin Declan","year":"2014","unstructured":"Declan Loughlin, Aedan Coffey, Frank Callaly, Darren Lyons, and Fearghal Morgan. 2014. Xilinx vivado high level synthesis: Case studies. In 25th IET Irish Signals and Systems Conference 2014 and 2014 China-Ireland International Conference on Information and Communications Technologies. 352\u2013356."},{"key":"e_1_3_2_26_2","first-page":"977","volume-title":"54th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201921) (Virtual Event, Greece, October 18-22, 2021)","author":"Lu Liqiang","year":"2021","unstructured":"Liqiang Lu, Yicheng Jin, Hangrui Bi, Zizhang Luo, Peng Li, Tao Wang, and Yun Liang. 2021. Sanger: A co-design framework for enabling sparse attention using reconfigurable architecture. In 54th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201921) (Virtual Event, Greece, October 18-22, 2021). ACM, 977\u2013991."},{"key":"e_1_3_2_27_2","first-page":"553","volume-title":"2017 IEEE International Symposium on High Performance Computer Architecture (HPCA 2017) (Austin, TX, February 4-8, 2017)","author":"Lu Wenyan","year":"2017","unstructured":"Wenyan Lu, Guihai Yan, Jiajun Li, Shijun Gong, Yinhe Han, and Xiaowei Li. 2017. FlexFlow: A flexible dataflow accelerator architecture for convolutional neural networks. In 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA 2017) (Austin, TX, February 4-8, 2017). IEEE Computer Society, 553\u2013564."},{"key":"e_1_3_2_28_2","first-page":"790","volume-title":"48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021)","author":"Ma Xiaohan","year":"2021","unstructured":"Xiaohan Ma, Chang Si, Ying Wang, Cheng Liu, and Lei Zhang. 2021. NASA: Accelerating neural network design with a NAS processor. In 48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021). IEEE, 790\u2013803."},{"key":"e_1_3_2_29_2","first-page":"195","volume-title":"56th Annual Design Automation Conference 2019 (DAC 2019) (Las Vegas, NV, June 02-06, 2019)","author":"Man Xingchen","year":"2019","unstructured":"Xingchen Man, Leibo Liu, Jianfeng Zhu, and Shaojun Wei. 2019. A general pattern-based dynamic compilation framework for coarse-grained reconfigurable architectures. In 56th Annual Design Automation Conference 2019 (DAC 2019) (Las Vegas, NV, June 02-06, 2019). ACM, 195."},{"key":"e_1_3_2_30_2","first-page":"259","volume-title":"49th Annual International Symposium on Computer Architecture (ISCA\u201922) (New York,, June 18-22, 2022)","author":"Man Xingchen","year":"2022","unstructured":"Xingchen Man, Jianfeng Zhu, Guihuan Song, Shouyi Yin, Shaojun Wei, and Leibo Liu. 2022. CaSMap: Agile mapper for reconfigurable spatial architectures by automatically clustering intermediate representations and scattering mapping process. In 49th Annual International Symposium on Computer Architecture (ISCA\u201922) (New York,, June 18-22, 2022), Valentina Salapura, Mohamed Zahran, Fred Chong, and Lingjia Tang (Eds.). ACM, 259\u2013273."},{"issue":"4","key":"e_1_3_2_31_2","first-page":"52","article-title":"Architectural support for data-driven execution","volume":"11","author":"Matheou George","year":"2015","unstructured":"George Matheou and Paraskevas Evripidou. 2015. Architectural support for data-driven execution. ACM Trans. Archit. Code Optim. 11, 4, Article 52 (Jan 2015), 25 pages.","journal-title":"ACM Trans. Archit. Code Optim."},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.1998.658762"},{"key":"e_1_3_2_33_2","first-page":"45:1\u201345:12","volume-title":"49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO 2016) (Taipei, Taiwan, October 15-19, 2016)","author":"Murray Sean","year":"2016","unstructured":"Sean Murray, William Floyd-Jones, Ying Qi, George Dimitri Konidaris, and Daniel J. Sorin. 2016. The microarchitecture of a real-time robot motion planning accelerator. In 49th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO 2016) (Taipei, Taiwan, October 15-19, 2016). IEEE Computer Society, 45:1\u201345:12."},{"key":"e_1_3_2_34_2","first-page":"340","volume-title":"38th IEEE International Conference on Computer Design (ICCD 2020) (Hartford, CT, October 18-21, 2020)","author":"Muthappa Ponnanna Kelettira","year":"2020","unstructured":"Ponnanna Kelettira Muthappa, Florian Neugebauer, Ilia Polian, and John P. Hayes. 2020. Hardware-based fast real-time image classification with stochastic computing. In 38th IEEE International Conference on Computer Design (ICCD 2020) (Hartford, CT, October 18-21, 2020). IEEE, 340\u2013347."},{"key":"e_1_3_2_35_2","first-page":"150","volume-title":"2017 IEEE International Parallel and Distributed Processing Symposium Workshops","author":"Nestorov Anna Maria","year":"2017","unstructured":"Anna Maria Nestorov, Enrico Reggiani, Hristina Palikareva, Pavel Burovskiy, Tobias Becker, and Marco D. Santambrogio. 2017. A scalable dataflow implementation of Curran\u2019s approximation algorithm. In 2017 IEEE International Parallel and Distributed Processing Symposium Workshops. 150\u2013157."},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","first-page":"596","DOI":"10.1109\/MICRO50266.2020.00056","volume-title":"53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO 2020), (Athens, Greece, October 17-21, 2020)","author":"Nguyen Quan M.","year":"2020","unstructured":"Quan M. Nguyen and Daniel S\u00e1nchez. 2020. Pipette: Improving core utilization on irregular applications through intra-core pipeline parallelism. In 53rd Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO 2020), (Athens, Greece, October 17-21, 2020). IEEE, 596\u2013608."},{"key":"e_1_3_2_37_2","doi-asserted-by":"crossref","first-page":"1064","DOI":"10.1145\/3466752.3480048","volume-title":"54th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201921) (Virtual Event, Greece, October 18-22, 2021)","author":"Nguyen Quan M.","year":"2021","unstructured":"Quan M. Nguyen and Daniel Sanchez. 2021. Fifer: Practical acceleration of irregular applications on reconfigurable architectures. In 54th Annual IEEE\/ACM International Symposium on Microarchitecture (MICRO\u201921) (Virtual Event, Greece, October 18-22, 2021). ACM, 1064\u20131077."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/325096.325117"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2019.2952839"},{"key":"e_1_3_2_40_2","first-page":"389","volume-title":"ISCA","year":"2017","unstructured":"Raghu Prabhakar, Yaqi Zhang, David Koeplinger, Matt Feldman, Tian Zhao, S. tefan Hadjis, Ardavan Pedram, Christos Kozyrakis, and Kunle Olukotun. 2017. Plasticine: A reconfigurable architecture for parallel patterns. In ISCA. ACM, 389\u2013402."},{"key":"e_1_3_2_41_2","volume-title":"International Symposium on Code Generation and Optimization (CGO\u201906)","author":"Smith A.","year":"2006","unstructured":"A. Smith, J. Burrill, J. Gibson, B. Maher, N. Nethercote, B. Yoder, D. Burger, and K. S. McKinley. 2006. Compiling for EDGE architectures. In International Symposium on Code Generation and Optimization (CGO\u201906). 185\u2013195."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1145\/1067649.801719"},{"key":"e_1_3_2_43_2","first-page":"1","volume-title":"Euro-Par 2013 Parallel Processing","year":"2013","unstructured":"Joshua Suettlerlein, St\u00e9phane Zuckerman, and Guang R. Gao. 2013. An implementation of the codelet model. In Euro-Par 2013 Parallel Processing. 1\u201314."},{"key":"e_1_3_2_44_2","first-page":"304","volume-title":"IEEE International Symposium on High-Performance Computer Architecture(HPCA 2022) (Seoul, South Korea, April 2-6, 2022)","author":"Tan Cheng","year":"2022","unstructured":"Cheng Tan, Nicolas Bohm Agostini, Tong Geng, Chenhao Xie, Jiajia Li, Ang Li, Kevin J. Barker, and Antonino Tumeo. 2022. DRIPS: Dynamic rebalancing of pipelined streaming applications on CGRAs. In IEEE International Symposium on High-Performance Computer Architecture(HPCA 2022) (Seoul, South Korea, April 2-6, 2022). IEEE, 304\u2013316."},{"key":"e_1_3_2_45_2","first-page":"1013","volume-title":"48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021)","author":"Tan Zhanhong","year":"2021","unstructured":"Zhanhong Tan, Hongyu Cai, Runpei Dong, and Kaisheng Ma. 2021. NN-baton: DNN workload orchestration and chiplet granularity exploration for multichip accelerators. In 48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021). IEEE, 1013\u20131026."},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2002.997877"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511807213"},{"key":"e_1_3_2_48_2","first-page":"402","volume-title":"48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021","author":"Vilim Matthew","year":"2021","unstructured":"Matthew Vilim, Alexander Rucker, and Kunle Olukotun. 2021. Aurochs: An architecture for dataflow threads. In 48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021). IEEE, 402\u2013415."},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","first-page":"435","DOI":"10.1109\/ICCD.2017.77","volume-title":"2017 IEEE International Conference on Computer Design (ICCD)","author":"Voss Nils","year":"2017","unstructured":"Nils Voss, Marco Bacis, Oskar Mencer, Georgi Gaydadjiev, and Wayne Luk. 2017. Convolutional neural networks on dataflow engines. In 2017 IEEE International Conference on Computer Design (ICCD). 435\u2013438."},{"key":"e_1_3_2_50_2","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1109\/FCCM.2019.00021","volume-title":"2019 IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)","author":"Voss Nils","year":"2019","unstructured":"Nils Voss, Pablo Quintana, Oskar Mencer, Wayne Luk, and Georgi Gaydadjiev. 2019. Memory mapping for multi-die FPGAs. In 2019 IEEE 27th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM). 78\u201386."},{"key":"e_1_3_2_51_2","first-page":"776","volume-title":"48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021","author":"Wang Xingbin","year":"2021","unstructured":"Xingbin Wang, Boyan Zhao, Rui Hou, Amro Awad, Zhihong Tian, and Dan Meng. 2021. NASGuard: A novel accelerator architecture for robust neural architecture search (NAS) networks. In 48th ACM\/IEEE Annual International Symposium on Computer Architecture (ISCA 2021) (Valencia, Spain, June 14-18, 2021). IEEE, 776\u2013789."},{"key":"e_1_3_2_52_2","first-page":"703","volume-title":"HPCA","year":"2020","unstructured":"Jian Weng, Sihao Liu, Zhengrong Wang, Vidushi Dadu, and Tony N. Owatzki. 2020. A hybrid systolic-dataflow architecture for inductive matrix algorithms. In HPCA. 703\u2013716."},{"key":"e_1_3_2_53_2","first-page":"831","volume-title":"2022 Design, Automation & Test in Europe Conference & Exhibition (DATE 2022) (Antwerp, Belgium, March 14-23, 2022)","author":"Wu Xinxin","year":"2022","unstructured":"Xinxin Wu, Zhihua Fan, Tianyu Liu, Wenming Li, Xiaochun Ye, and Dongrui Fan. 2022. LRP: Predictive output activation based on SVD approach for CNN s acceleration. In 2022 Design, Automation & Test in Europe Conference & Exhibition (DATE 2022) (Antwerp, Belgium, March 14-23, 2022), Cristiana Bolchini, Ingrid Verbauwhede, and Ioana Vatajelu (Eds.). IEEE, 831\u2013836."},{"key":"e_1_3_2_54_2","first-page":"1003","volume-title":"IEEE International Symposium on High-Performance Computer Architecture (HPCA 2023) (Montreal, QC, Canada, February 25 - March 1, 2023","author":"Yao Jianguo","year":"2023","unstructured":"Jianguo Yao, Hao Zhou, Yalin Zhang, Ying Li, Chuang Feng, Shi Chen, Jiaoyan Chen, Yongdong Wang, and Qiaojuan Hu. 2023. High performance and power efficient accelerator for cloud inference. In IEEE International Symposium on High-Performance Computer Architecture (HPCA 2023) (Montreal, QC, Canada, February 25 - March 1, 2023). IEEE, 1003\u20131016."},{"key":"e_1_3_2_55_2","first-page":"650","volume-title":"2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture","author":"Yazdanbakhsh Amir","year":"2018","unstructured":"Amir Yazdanbakhsh, Kambiz Samadi, Nam Sung Kim, Hadi Esmaeilzadeh, Hajar Falahati, and Philip J. Wolfe. 2018. GANAX: A unified MIMD-SIMD acceleration for generative adversarial networks. In 2018 ACM\/IEEE 45th Annual International Symposium on Computer Architecture. 650\u2013661."},{"key":"e_1_3_2_56_2","first-page":"273","volume-title":"International Symposium on Low Power Electronics and Design (ISLPED) (Beijing, China, September 4-6, 2013)","author":"Ye Xiaochun","year":"2013","unstructured":"Xiaochun Ye, Dongrui Fan, Ninghui Sun, Shibin Tang, Mingzhe Zhang, and Hao Zhang. 2013. SimICT: A fast and flexible framework for performance and power evaluation of large-scale architecture. In International Symposium on Low Power Electronics and Design (ISLPED) (Beijing, China, September 4-6, 2013), Pai H. Chou, Ru Huang, Yuan Xie, and Tanay Karnik (Eds.). IEEE, 273\u2013278."},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.future.2020.03.023"},{"key":"e_1_3_2_58_2","first-page":"1394","volume-title":"DATE","author":"Yin Chen","year":"2021","unstructured":"Chen Yin and Qin Wang. 2021. Subgraph decoupling and rescheduling for increased utilization in CGRA architecture. In DATE. 1394\u20131399."},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2018.2821561"},{"key":"e_1_3_2_60_2","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1145\/2684746.2689060","volume-title":"2015 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201915)","author":"Zhang Chen","year":"2015","unstructured":"Chen Zhang, Peng Li, Guangyu Sun, Yijin Guan, Bingjun Xiao, and Jason Cong. 2015. Optimizing FPGA-based accelerator design for deep convolutional neural networks. In 2015 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201915) (Monterey, California) . ACM, New York, 161\u2013170."},{"key":"e_1_3_2_61_2","first-page":"552","volume-title":": The 49th Annual International Symposium on Computer Architecture (ISCA\u201922) (New York, June 18-22, 2022","author":"Zhang Yunan","year":"2022","unstructured":"Yunan Zhang, Po-An Tsai, and Hung-Wei Tseng. 2022. SIMD \\({}^{\\mbox{2}}\\) : A generalized matrix instruction set for accelerating tensor computation beyond GEMM. In : The 49th Annual International Symposium on Computer Architecture (ISCA\u201922) (New York, June 18-22, 2022), Valentina Salapura, Mohamed Zahran, Fred Chong, and Lingjia Tang (Eds.). ACM, 552\u2013566."},{"key":"e_1_3_2_62_2","doi-asserted-by":"crossref","first-page":"1041","DOI":"10.1109\/ISCA52012.2021.00085","volume-title":"2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA)","author":"Zhang Yaqi","year":"2021","unstructured":"Yaqi Zhang, Nathan Zhang, Tian Zhao, Matt Vilim, Muhammad Shahbaz, and Kunle Olukotun. 2021. SARA: Scaling a reconfigurable dataflow accelerator. In 2021 ACM\/IEEE 48th Annual International Symposium on Computer Architecture (ISCA). 1041\u20131054."},{"key":"e_1_3_2_63_2","first-page":"475","volume-title":"IEEE International Symposium on High-Performance Computer Architecture (HPCA 2022) (Seoul, South Korea, April 2-6, 2022","author":"Zheng Shixuan","year":"2022","unstructured":"Shixuan Zheng, Xianjue Zhang, Leibo Liu, Shaojun Wei, and Shouyi Yin. 2022. Atomic dataflow based graph-level workload orchestration for scalable DNN accelerators. In IEEE International Symposium on High-Performance Computer Architecture (HPCA 2022) (Seoul, South Korea, April 2-6, 2022). IEEE, 475\u2013489."}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3637906","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3637906","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:49:03Z","timestamp":1750286943000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3637906"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,15]]},"references-count":62,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,3,31]]}},"alternative-id":["10.1145\/3637906"],"URL":"https:\/\/doi.org\/10.1145\/3637906","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,15]]},"assertion":[{"value":"2023-05-04","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-11","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}