{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T16:14:51Z","timestamp":1781885691636,"version":"3.54.5"},"reference-count":194,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2023,1,24]],"date-time":"2023-01-24T00:00:00Z","timestamp":1674518400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100006602","name":"Air Force Research Laboratory","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100006602","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100000185","name":"Defense Advanced Research Projects Agency","doi-asserted-by":"crossref","award":["FA8650-18-2-7860"],"award-info":[{"award-number":["FA8650-18-2-7860"]}],"id":[{"id":"10.13039\/100000185","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2023,3,31]]},"abstract":"<jats:p>Process technology-driven performance and energy efficiency improvements have slowed down as we approach physical design limits. General-purpose manycore architectures attempt to circumvent this challenge, but they have a significant performance and energy-efficient gap compared to special-purpose solutions. Domain-specific architectures (DSAs), an instance of heterogeneous architectures, efficiently combine general-purpose cores and specialized hardware accelerators to boost energy efficiency and provide programming flexibility. Indeed, the hardware, software, and systems aspects in DSAs are highly tailored to maximize the energy efficiency of applications in a target domain. As DSAs and their conceptualization advance rapidly, there is a strong need to understand the research problems that need immediate attention. This article discusses the primary research directions in the design and runtime management of DSAs. Then, it surveys some promising approaches and highlights the outstanding research needs.<\/jats:p>","DOI":"10.1145\/3563946","type":"journal-article","created":{"date-parts":[[2022,9,21]],"date-time":"2022-09-21T11:56:26Z","timestamp":1663761386000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":37,"title":["Domain-Specific Architectures: Research Problems and Promising Approaches"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2419-1860","authenticated-orcid":false,"given":"Anish","family":"Krishnakumar","sequence":"first","affiliation":[{"name":"University of Wisconsin\u2013Madison"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5045-5535","authenticated-orcid":false,"given":"Umit","family":"Ogras","sequence":"additional","affiliation":[{"name":"University of Wisconsin\u2013Madison"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1826-7646","authenticated-orcid":false,"given":"Radu","family":"Marculescu","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5593-9694","authenticated-orcid":false,"given":"Mike","family":"Kishinevsky","sequence":"additional","affiliation":[{"name":"Intel Corporation"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7845-2187","authenticated-orcid":false,"given":"Trevor","family":"Mudge","sequence":"additional","affiliation":[{"name":"University of Michigan"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,1,24]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Apple. [n.d.]. Apple Secure Enclave. Retrieved May 15 2022 from https:\/\/support.apple.com\/guide\/security\/secure-enclave-sec59b0b31ff\/web."},{"key":"e_1_3_1_3_2","unstructured":"Cadence. [n.d.]. ARM CoreLink Interconnects Whitepaper. Retrieved May 15 2022 from https:\/\/ip.cadence.com\/uploads\/251\/white-paper-interconnect-solutions-debugging-issues-advanced-ARM-CoreLink-pdf."},{"key":"e_1_3_1_4_2","unstructured":"ARM. [n.d.]. ARM TrustZone. Retrived May 15 2022 from https:\/\/developer.arm.com\/documentation\/PRD29-GENC-009492\/c\/TrustZone-Hardware-Architecture."},{"key":"e_1_3_1_5_2","unstructured":"Google. [n.d.]. Google\u2019s Thrust Towards Open-Source Hardware. Retrieved May 15 2022 from https:\/\/opensource.googleblog.com\/2019\/05\/google-fosters-open-source-hardware.html."},{"key":"e_1_3_1_6_2","unstructured":"Aakash Jani. 2022. Year in Review: PC Processors Adopt Hybrid CPUs. Retrieved May 15 2022 from https:\/\/www.techinsights.com\/blog\/year-review-pc-processors-adopt-hybrid-cpus."},{"key":"e_1_3_1_7_2","unstructured":"Retrieved May 15 2022 from https:\/\/futurenetworks.ieee.org\/images\/files\/pdf\/FirstResponder\/Tom-Rondeau-DARPA.pdf."},{"key":"e_1_3_1_8_2","unstructured":"Siemens. [n.d.]. Veloce2 Emulator. Retrieved May 15 2022 from https:\/\/www.mentor.com\/products\/fv\/emulation-systems\/veloce."},{"key":"e_1_3_1_9_2","unstructured":"Synopsys. [n.d.]. ZeBu Server 4. Retrieved May 15 2022 from https:\/\/www.synopsys.com\/verification\/emulation\/zebu-server.html."},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3326334"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2021.3085505"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2013.57"},{"issue":"8","key":"e_1_3_1_13_2","first-page":"1248","article-title":"DS3: A system-level domain-specific system-on-chip simulation framework","volume":"69","author":"Arda Samet","year":"2020","unstructured":"Samet Arda, Anish Krishnakumar, Ahmet Alper Goksoy, Joshua Mack, Nirmal Kumbhare, Anderson Luiz Sartor, Ali Akoglu, Radu Marculescu, and Umit Y. Ogras. 2020. DS3: A system-level domain-specific system-on-chip simulation framework. IEEE Transactions on Computers 69, 8 (2020), 1248\u20131262.","journal-title":"IEEE Transactions on Computers"},{"key":"e_1_3_1_14_2","unstructured":"Krste Asanovic Ras Bodik Bryan Christopher Catanzaro Joseph James Gebis Parry Husbands Kurt Keutzer David A. Patterson et\u00a0al. 2006. The Landscape of Parallel Computing Research: A View from Berkeley . Technical Report No. UCB\/EECS-2006-183. EECS Department University of California Berkeley."},{"key":"e_1_3_1_15_2","unstructured":"Krste Asanovi\u0107 and David A. Patterson. 2014. Instruction Sets Should Be Free: The Case for RISC-V . Technical Report No. UCB\/EECS-2014-146. EECS Department University of California Berkeley."},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/1255456.1255463"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTCSA.2007.21"},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.1631"},{"key":"e_1_3_1_19_2","first-page":"1","volume-title":"Proceedings of the 57th ACM\/IEEE Design Automation Conference (DAC\u201920)","author":"al Rick Bahr, Clark Barrett, Nikhil Bhagdikar, Alex Carsello, Ross Daly, Caleb Donovick, David Durst, et","year":"2020","unstructured":"Rick Bahr, Clark Barrett, Nikhil Bhagdikar, Alex Carsello, Ross Daly, Caleb Donovick, David Durst, et al. 2020. Creating an agile hardware design flow. In Proceedings of the 57th ACM\/IEEE Design Automation Conference (DAC\u201920). 1\u20136."},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.5555\/1134160.1134163"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.aei.2007.08.005"},{"key":"e_1_3_1_22_2","first-page":"10","volume-title":"Proceedings of the USENIX Annual Technical Conference: FREENIX Track","volume":"41","author":"Bellard Fabrice","year":"2005","unstructured":"Fabrice Bellard. 2005. QEMU, A fast and portable dynamic translator. In Proceedings of the USENIX Annual Technical Conference: FREENIX Track, Vol. 41. 10\u20135555."},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2009.2030268"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2019.8747521"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/92.845896"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2017.2770163"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/PDP.2010.56"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.1145\/3544016"},{"key":"e_1_3_1_29_2","unstructured":"Pradip Bose Augusto Vega Sarita Adve Vikram Adve Sasa Misailovic Luca Carloni Ken Shepard David Brooks Vijay Janapa Reddi and Gu-Yeon Wei. 2021. Secure and resilient SoCs for autonomous vehicles. In International Workshop on Domain Specific System Architecture (DOSSA) in conjunction with IEEE International Symposium on High-Performance Computer Architecture (HPCA) . https:\/\/scholar.google.com\/scholar?hl=en&as_sdt=0%2C50&q=Secure+and+resilient+SoCs+for+autonomous+vehicles&btnG=."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1006\/jpdc.2000.1714"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/PACT.2011.15"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/HCS52781.2021.9567455"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/HCS52781.2021.9567066"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/2514740"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS.2009.5118013"},{"key":"e_1_3_1_36_2","first-page":"1","volume-title":"Proceedings of the ACM\/EDAC\/IEEE Design Automation Conference (DAC\u201916)","author":"Carloni Luca P.","year":"2016","unstructured":"Luca P. Carloni. 2016. The case for embedded scalable platforms. In Proceedings of the ACM\/EDAC\/IEEE Design Automation Conference (DAC\u201916). 1\u20136."},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2011.2173941"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE48585.2020.9116469"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-020-01555-w"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD51958.2021.9643465"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1145\/3300189.3300192"},{"key":"e_1_3_1_42_2","doi-asserted-by":"crossref","unstructured":"Kuan-Yu Chen Chi-Sheng Yang Yu-Hsiu Sun Chien-Wei Tseng Morteza Fayazi Xin He Siying Feng et\u00a0al. 2022. A 507 GMACs\/J 256-core domain adaptive systolic-array-processor for wireless communication and linear-algebra kernels in 12nm FINFET. In Proceedings of the 2022 IEEE Symposium on VLSI Technology and Circuits (VLSI Technology and Circuits\u201922) .","DOI":"10.1109\/VLSITechnologyandCir46769.2022.9830330"},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/MDAT.2017.2735383"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3195970.3195986"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2008.2003301"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2014.12"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/MDT.2010.141"},{"key":"e_1_3_1_48_2","first-page":"800","article-title":"TensorFlow Lite Micro: Embedded machine learning for TinyML systems","volume":"3","author":"David Robert","year":"2021","unstructured":"Robert David, Jared Duke, Advait Jain, Vijay Janapa Reddi, Nat Jeffries, Jian Li, Nick Kreeger, et\u00a0al. 2021. TensorFlow Lite Micro: Embedded machine learning for TinyML systems. Proceedings of Machine Learning and Systems 3 (2021), 800\u2013811.","journal-title":"Proceedings of Machine Learning and Systems"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/1978802.1978814"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.3390\/fi13060146"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1049\/iet-cdt.2019.0037"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1145\/2968456.2968459"},{"key":"e_1_3_1_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/1046192.1046204"},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/RTSS.2018.00019"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.7873\/DATE.2015.0450"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/2000064.2000108"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1145\/2567895"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/24039.24041"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.microrel.2012.01.005"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.bdr.2017.01.005"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1998.681704"},{"key":"e_1_3_1_62_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4419-0504-8"},{"key":"e_1_3_1_63_2","doi-asserted-by":"publisher","DOI":"10.1109\/92.486082"},{"key":"e_1_3_1_64_2","volume-title":"SpecC: Specification Language and Methodology","author":"Gajski Daniel D.","year":"2012","unstructured":"Daniel D. Gajski, Jianwen Zhu, Rainer D\u00f6mer, Andreas Gerstlauer, and Shuqing Zhao. 2012. SpecC: Specification Language and Methodology. Springer Science & Business Media."},{"key":"e_1_3_1_65_2","volume-title":"GNU Scientific Library","author":"Galassi Mark","year":"2002","unstructured":"Mark Galassi, Jim Davies, James Theiler, Brian Gough, Gerard Jungman, Patrick Alken, Michael Booth, Fabrice Rossi, and Rhys Ulerich. 2002. GNU Scientific Library. Network Theory Limited."},{"key":"e_1_3_1_66_2","volume-title":"Comparison of Several FFT Libraries in C\/C++","year":"2020","unstructured":"P. Gambron and S. Thorne. 2020. Comparison of Several FFT Libraries in C\/C++. Technical Report. STFC."},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2005.160"},{"key":"e_1_3_1_68_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCAS.2007.910029"},{"key":"e_1_3_1_69_2","doi-asserted-by":"publisher","DOI":"10.1109\/PARELEC.2006.54"},{"key":"e_1_3_1_70_2","doi-asserted-by":"publisher","DOI":"10.1109\/RSP.2010.5656355"},{"key":"e_1_3_1_71_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2009.2026356"},{"key":"e_1_3_1_72_2","volume-title":"Transaction-Level Modeling with SystemC","year":"2005","unstructured":"Frank Ghenassia (Ed.). 2005. Transaction-Level Modeling with SystemC. Vol. 2. Springer."},{"key":"e_1_3_1_73_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3073893"},{"key":"e_1_3_1_74_2","doi-asserted-by":"publisher","DOI":"10.1109\/NOCS.2018.8512153"},{"key":"e_1_3_1_75_2","unstructured":"Daniel S. Green. 2018. Heterogeneous Integration at DARPA: Pathfinding and Progress in Assembly Approaches . DARPA."},{"key":"e_1_3_1_76_2","doi-asserted-by":"publisher","DOI":"10.1145\/2304576.2304621"},{"key":"e_1_3_1_77_2","doi-asserted-by":"publisher","DOI":"10.1145\/3386377"},{"key":"e_1_3_1_78_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.vlsi.2004.06.001"},{"key":"e_1_3_1_79_2","volume-title":"System Design with SystemCTM","author":"Gr\u00f6tker Thorsten","year":"2007","unstructured":"Thorsten Gr\u00f6tker, Stan Liao, Grant Martin, and Stuart Swan. 2007. System Design with SystemCTM. Springer Science & Business Media."},{"key":"e_1_3_1_80_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4020-6488-3_3"},{"key":"e_1_3_1_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2986214"},{"key":"e_1_3_1_82_2","doi-asserted-by":"publisher","DOI":"10.1145\/2661430"},{"key":"e_1_3_1_83_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2012.213"},{"key":"e_1_3_1_84_2","unstructured":"ODROID Wiki. [n.d.]. Hardkernel. ODROID-XU3. Retrieved May 15 2022 from https:\/\/wiki.odroid.com\/old_product\/odroid-xu3\/odroid-xu3."},{"key":"e_1_3_1_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_86_2","volume-title":"Proceedings of the ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA\u201918)","author":"Hennessy John","year":"2018","unstructured":"John Hennessy and David Patterson. 2018. A new golden age for computer architecture: Domain-specific hardware\/software co-design, enhanced. In Proceedings of the ACM\/IEEE 45th Annual International Symposium on Computer Architecture (ISCA\u201918)."},{"key":"e_1_3_1_87_2","doi-asserted-by":"publisher","DOI":"10.1145\/3282307"},{"key":"e_1_3_1_88_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2005.844106"},{"key":"e_1_3_1_89_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-80475-6"},{"key":"e_1_3_1_90_2","doi-asserted-by":"publisher","DOI":"10.1145\/2039370.2039409"},{"key":"e_1_3_1_91_2","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1109\/DATE.2009.5090632","volume-title":"Proceedings of the Design, Automation, and Test in Europe Conference and Exhibition (DATE\u201909)","author":"Huang Lin","year":"2009","unstructured":"Lin Huang, Feng Yuan, and Qiang Xu. 2009. Lifetime reliability-aware task allocation and scheduling for MPSoC platforms. In Proceedings of the Design, Automation, and Test in Europe Conference and Exhibition (DATE\u201909). 51\u201356."},{"key":"e_1_3_1_92_2","first-page":"219","volume-title":"Proceedings of the International Conference on Smart Card Research and Advanced Applications","author":"Hutter Michael","year":"2013","unstructured":"Michael Hutter and J\u00f6rn-Marc Schmidt. 2013. The temperature side channel and heating fault attacks. In Proceedings of the International Conference on Smart Card Research and Advanced Applications. 219\u2013235."},{"key":"e_1_3_1_93_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2010.31"},{"key":"e_1_3_1_94_2","volume-title":"Intel Xeon Phi Processor High Performance Programming: Knights Landing Edition","year":"2016","unstructured":"James Jeffers, James Reinders, and Avinash Sodani. 2016. Intel Xeon Phi Processor High Performance Programming: Knights Landing Edition. Morgan Kaufmann."},{"key":"e_1_3_1_95_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-007-0139-z"},{"key":"e_1_3_1_96_2","doi-asserted-by":"publisher","DOI":"10.1145\/68182.68207"},{"key":"e_1_3_1_97_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA52012.2021.00010"},{"key":"e_1_3_1_98_2","doi-asserted-by":"publisher","DOI":"10.1109\/NOCARC.2018.8541158"},{"key":"e_1_3_1_99_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSSC.2022.3140241"},{"key":"e_1_3_1_100_2","doi-asserted-by":"publisher","DOI":"10.1145\/2187671.2187675"},{"key":"e_1_3_1_101_2","doi-asserted-by":"publisher","DOI":"10.1145\/3178487.3178493"},{"key":"e_1_3_1_102_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2020.3012861"},{"key":"e_1_3_1_103_2","article-title":"ImageNet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems 25 (2012).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_104_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVT.2020.2999533"},{"key":"e_1_3_1_105_2","article-title":"HERO: Heterogeneous embedded research platform for exploring RISC-V manycore accelerators on FPGA","author":"Kurth Andreas","year":"2017","unstructured":"Andreas Kurth, Pirmin Vogel, Alessandro Capotondi, Andrea Marongiu, and Luca Benini. 2017. HERO: Heterogeneous embedded research platform for exploring RISC-V manycore accelerators on FPGA. arXiv preprint arXiv:1712.06497 (2017).","journal-title":"arXiv preprint arXiv:1712.06497"},{"key":"e_1_3_1_106_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2011.5979711"},{"key":"e_1_3_1_107_2","doi-asserted-by":"publisher","DOI":"10.5555\/977395.977673"},{"key":"e_1_3_1_108_2","doi-asserted-by":"publisher","DOI":"10.1145\/1594835.1504194"},{"key":"e_1_3_1_109_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2015.08.004"},{"key":"e_1_3_1_110_2","doi-asserted-by":"publisher","DOI":"10.1145\/3173162.3173191"},{"key":"e_1_3_1_111_2","doi-asserted-by":"publisher","DOI":"10.1109\/TETC.2014.2348182"},{"key":"e_1_3_1_112_2","doi-asserted-by":"publisher","DOI":"10.1145\/3357375"},{"key":"e_1_3_1_113_2","article-title":"Performant, multi-objective scheduling of highly interleaved task graphs on heterogeneous system on chip devices","author":"Mack Joshua","year":"2021","unstructured":"Joshua Mack, Samet Arda, Umit Y. Ogras, and Ali Akoglu. 2021. Performant, multi-objective scheduling of highly interleaved task graphs on heterogeneous system on chip devices. IEEE Transactions on Parallel and Distributed Systems 33 (2021), 2148\u20132162.","journal-title":"IEEE Transactions on Parallel and Distributed Systems"},{"key":"e_1_3_1_114_2","doi-asserted-by":"publisher","DOI":"10.1145\/3529257"},{"key":"e_1_3_1_115_2","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPSW50202.2020.00016"},{"key":"e_1_3_1_116_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477006"},{"key":"e_1_3_1_117_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2019.8714921"},{"key":"e_1_3_1_118_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2019.2926106"},{"key":"e_1_3_1_119_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-69131-8_3"},{"key":"e_1_3_1_120_2","doi-asserted-by":"publisher","DOI":"10.1145\/3400302.3415753"},{"key":"e_1_3_1_121_2","doi-asserted-by":"publisher","DOI":"10.1145\/3005745.3005750"},{"key":"e_1_3_1_122_2","doi-asserted-by":"publisher","DOI":"10.1145\/3341302.3342080"},{"key":"e_1_3_1_123_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2008.2010691"},{"key":"e_1_3_1_124_2","doi-asserted-by":"publisher","DOI":"10.1089\/106652700750050826"},{"key":"e_1_3_1_125_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3112876"},{"key":"e_1_3_1_126_2","doi-asserted-by":"publisher","DOI":"10.1145\/1118890.1118892"},{"key":"e_1_3_1_127_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00521-018-3761-1"},{"key":"e_1_3_1_128_2","doi-asserted-by":"publisher","DOI":"10.1145\/2788396"},{"key":"e_1_3_1_129_2","doi-asserted-by":"publisher","DOI":"10.1145\/3358203"},{"key":"e_1_3_1_130_2","article-title":"FluidFFT: Common API (C++ and Python) for fast Fourier transform HPC libraries","author":"Mohanan Ashwin Vishnu","year":"2018","unstructured":"Ashwin Vishnu Mohanan, Cyrille Bonamy, and Pierre Augier. 2018. FluidFFT: Common API (C++ and Python) for fast Fourier transform HPC libraries. arXiv preprint arXiv:1807.01775 (2018).","journal-title":"arXiv preprint arXiv:1807.01775"},{"key":"e_1_3_1_131_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2009.2032372"},{"key":"e_1_3_1_132_2","doi-asserted-by":"publisher","DOI":"10.1109\/SP40000.2020.00057"},{"key":"e_1_3_1_133_2","doi-asserted-by":"publisher","DOI":"10.1145\/3243734.3243831"},{"key":"e_1_3_1_134_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3058217"},{"key":"e_1_3_1_135_2","first-page":"128","volume-title":"Proceedings of the Science and Information Conference","author":"O\u2019Mahony Niall","year":"2019","unstructured":"Niall O\u2019Mahony, Sean Campbell, Anderson Carvalho, Suman Harapanahalli, Gustavo Velasco Hernandez, Lenka Krpalkova, Daniel Riordan, and Joseph Walsh. 2019. Deep learning vs. traditional computer vision. In Proceedings of the Science and Information Conference. 128\u2013144."},{"key":"e_1_3_1_136_2","doi-asserted-by":"publisher","DOI":"10.1049\/iet-cdt.2014.0074"},{"key":"e_1_3_1_137_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394885.3431595"},{"key":"e_1_3_1_138_2","doi-asserted-by":"publisher","DOI":"10.1109\/MDAT.2020.2976669"},{"key":"e_1_3_1_139_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISSCC.2018.8310168"},{"key":"e_1_3_1_140_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2983308"},{"key":"e_1_3_1_141_2","volume-title":"XXV Congreso Argentino de Ciencias de la Computaci\u00f3n (CACIC\u201919).","author":"Puig Mart\u00edn Pi","year":"2019","unstructured":"Mart\u00edn Pi Puig, Laura Cristina De Giusti, Marcelo Naiouf, and Armando Eduardo De Giusti. 2019. A study of hardware performance counters selection for cross architectural GPU power modeling. In XXV Congreso Argentino de Ciencias de la Computaci\u00f3n (CACIC\u201919)."},{"key":"e_1_3_1_142_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2006.16"},{"key":"e_1_3_1_143_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2017.12.020"},{"key":"e_1_3_1_144_2","doi-asserted-by":"publisher","DOI":"10.1145\/3107953"},{"key":"e_1_3_1_145_2","first-page":"1","volume-title":"Proceedings of the Embedded Systems Conference","author":"Punkka Timo","year":"2012","unstructured":"Timo Punkka. 2012. Agile hardware and co-design. In Proceedings of the Embedded Systems Conference. 1\u20138."},{"key":"e_1_3_1_146_2","doi-asserted-by":"publisher","DOI":"10.1145\/3150211"},{"key":"e_1_3_1_147_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMSCS.2017.2755619"},{"key":"e_1_3_1_148_2","doi-asserted-by":"publisher","DOI":"10.1109\/41.744370"},{"key":"e_1_3_1_149_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2007.895245"},{"issue":"3","key":"e_1_3_1_150_2","first-page":"103","article-title":"Efficient FPGA implementation of FFT\/IFFT processor","volume":"3","author":"Saeed Ahmed","year":"2009","unstructured":"Ahmed Saeed, M. Elbably, G. Abdelfadeel, and M. I. Eladawy. 2009. Efficient FPGA implementation of FFT\/IFFT processor. International Journal of Circuits, Systems and Signal Processing 3, 3 (2009), 103\u2013110.","journal-title":"International Journal of Circuits, Systems and Signal Processing"},{"key":"e_1_3_1_151_2","doi-asserted-by":"publisher","DOI":"10.1145\/2993452.2994309"},{"key":"e_1_3_1_152_2","doi-asserted-by":"publisher","DOI":"10.1109\/RSP.2014.6966902"},{"key":"e_1_3_1_153_2","doi-asserted-by":"publisher","DOI":"10.1109\/LCA.2020.2992182"},{"key":"e_1_3_1_154_2","doi-asserted-by":"publisher","DOI":"10.5555\/3195638.3195697"},{"key":"e_1_3_1_155_2","article-title":"Very deep convolutional networks for large-scale image recognition","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014).","journal-title":"arXiv preprint arXiv:1409.1556"},{"key":"e_1_3_1_156_2","doi-asserted-by":"publisher","DOI":"10.1145\/2463209.2488734"},{"key":"e_1_3_1_157_2","doi-asserted-by":"crossref","unstructured":"David B. Skillicorn and Domenico Talia. 1998. Models and languages for parallel computation. ACM Computing Surveys 2 (1998) 123\u2013169.","DOI":"10.1145\/280277.280278"},{"key":"e_1_3_1_158_2","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2012.20"},{"key":"e_1_3_1_159_2","unstructured":"Ashley Stevens. 2014. Quality of Service (QoS) in ARM\u00ae Systems: An Overview . White Paper. ARM Cambridge UK."},{"key":"e_1_3_1_160_2","doi-asserted-by":"publisher","DOI":"10.1145\/2847263.2847276"},{"key":"e_1_3_1_161_2","doi-asserted-by":"publisher","DOI":"10.1145\/2584665"},{"key":"e_1_3_1_162_2","doi-asserted-by":"publisher","DOI":"10.1109\/ReCoSoC.2018.8449389"},{"key":"e_1_3_1_163_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2021.3114882"},{"key":"e_1_3_1_164_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_1_165_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10723-015-9334-y"},{"key":"e_1_3_1_166_2","doi-asserted-by":"publisher","DOI":"10.24251\/HICSS.2018.715"},{"key":"e_1_3_1_167_2","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2017.29"},{"key":"e_1_3_1_168_2","doi-asserted-by":"publisher","DOI":"10.1109\/71.993206"},{"key":"e_1_3_1_169_2","article-title":"RedMulE: A compact FP16 matrix-multiplication accelerator for adaptive deep learning on RISC-V-based ultra-low-power SoCs","author":"Tortorella Yvan","year":"2022","unstructured":"Yvan Tortorella, Luca Bertaccini, Davide Rossi, Luca Benini, and Francesco Conti. 2022. RedMulE: A compact FP16 matrix-multiplication accelerator for adaptive deep learning on RISC-V-based ultra-low-power SoCs. arXiv preprint arXiv:2204.11192 (2022).","journal-title":"arXiv preprint arXiv:2204.11192"},{"key":"e_1_3_1_170_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2017.2688340"},{"key":"e_1_3_1_171_2","first-page":"180","volume-title":"Open Architecture\/Open Business Model Net-Centric Systems and Defense Transformation 2019","author":"Uhrie Richard","year":"2019","unstructured":"Richard Uhrie, Daniel W. Bliss, Chaitali Chakrabarti, Umit Y. Ogras, and John Brunhaver. 2019. Machine understanding of domain computation for domain-specific system-on-chips (DSSoC). In Open Architecture\/Open Business Model Net-Centric Systems and Defense Transformation 2019, Vol. 11015. International Society for Optics and Photonics, SPIE, 180\u2013187."},{"key":"e_1_3_1_172_2","article-title":"Automated parallel kernel extraction from dynamic application traces","author":"Uhrie Richard","year":"2020","unstructured":"Richard Uhrie, Chaitali Chakrabarti, and John Brunhaver. 2020. Automated parallel kernel extraction from dynamic application traces. arXiv preprint arXiv:2001.09995 (2020).","journal-title":"arXiv preprint arXiv:2001.09995"},{"key":"e_1_3_1_173_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0022-0000(75)80008-0"},{"key":"e_1_3_1_174_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD.2010.5647727"},{"key":"e_1_3_1_175_2","doi-asserted-by":"publisher","DOI":"10.1145\/2103799.2103813"},{"key":"e_1_3_1_176_2","unstructured":"Augusto Vega John-David Wellman Hubertus Franke Alper Buyuktosunoglu Pradip Bose Aporva Amarnath Hiwot Kassa Subhankar Pal and Ronald Dreslinski. 2021. STOMP: Agile evaluation of scheduling policies in heterogeneous multi-processors. In Proceedings of the 3rd International Workshop on Domain Specific System Architecture in Conjunction with the 27th IEEE International Symposium on High-Performance Computer Architecture (DOSSA-3 @ HPCA\u201921) ."},{"key":"e_1_3_1_177_2","doi-asserted-by":"publisher","DOI":"10.1109\/CIT.2010.322"},{"key":"e_1_3_1_178_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM.2019.00019"},{"key":"e_1_3_1_179_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.vlsi.2022.03.002"},{"key":"e_1_3_1_180_2","doi-asserted-by":"crossref","first-page":"167","DOI":"10.1007\/978-3-319-06486-4_7","volume-title":"High-Performance Computing on the Intel\u00ae Xeon Phi \\(^{TM}\\)","author":"Wang Endong","year":"2014","unstructured":"Endong Wang, Qing Zhang, Bo Shen, Guangyong Zhang, Xiaowei Lu, Qing Wu, and Yajuan Wang. 2014. Intel math kernel library. In High-Performance Computing on the Intel\u00ae Xeon Phi \\(^{TM}\\) . Springer, 167\u2013188."},{"key":"e_1_3_1_181_2","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2013.74"},{"key":"e_1_3_1_182_2","article-title":"Benchmarking TPU, GPU, and CPU platforms for deep learning","author":"Wang Yu Emma","year":"2019","unstructured":"Yu Emma Wang, Gu-Yeon Wei, and David Brooks. 2019. Benchmarking TPU, GPU, and CPU platforms for deep learning. arXiv preprint arXiv:1907.10701 (2019).","journal-title":"arXiv preprint arXiv:1907.10701"},{"key":"e_1_3_1_183_2","doi-asserted-by":"publisher","DOI":"10.1145\/3061639.3062207"},{"key":"e_1_3_1_184_2","doi-asserted-by":"publisher","DOI":"10.1093\/cid\/cix731"},{"key":"e_1_3_1_185_2","article-title":"HiMap: Fast and scalable high-quality mapping on CGRA via hierarchical abstraction","author":"Wijerathne Dhananjaya","year":"2021","unstructured":"Dhananjaya Wijerathne, Zhaoying Li, Anuj Pathania, Tulika Mitra, and Lothar Thiele. 2021. HiMap: Fast and scalable high-quality mapping on CGRA via hierarchical abstraction. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems 41, 10 (2021), 3290\u20133303.","journal-title":"IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems"},{"key":"e_1_3_1_186_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCD.2011.6081395"},{"key":"e_1_3_1_187_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMSCS.2015.2487983"},{"key":"e_1_3_1_188_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVLSI.2019.2897650"},{"key":"e_1_3_1_189_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3071507"},{"key":"e_1_3_1_190_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS45731.2020.9180871"},{"key":"e_1_3_1_191_2","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"},{"key":"e_1_3_1_192_2","doi-asserted-by":"publisher","DOI":"10.1145\/3276491"},{"key":"e_1_3_1_193_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2020.2989149"},{"key":"e_1_3_1_194_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSC.2019.2963301"},{"key":"e_1_3_1_195_2","article-title":"DRHEFT: Deadline-constrained reliability-aware HEFT algorithm for real-time heterogeneous MPSoC systems","author":"Zhou Junlong","year":"2022","unstructured":"Junlong Zhou, Mingyue Zhang, Jin Sun, Tian Wang, Xiumin Zhou, and Shiyan Hu. 2022. DRHEFT: Deadline-constrained reliability-aware HEFT algorithm for real-time heterogeneous MPSoC systems. IEEE Transactions on Reliability 71, 1 (2022), 178\u2013189.","journal-title":"IEEE Transactions on Reliability"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3563946","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3563946","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:08:15Z","timestamp":1750183695000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3563946"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,24]]},"references-count":194,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,3,31]]}},"alternative-id":["10.1145\/3563946"],"URL":"https:\/\/doi.org\/10.1145\/3563946","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,24]]},"assertion":[{"value":"2022-07-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-08-10","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-01-24","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}