{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,21]],"date-time":"2026-04-21T21:32:57Z","timestamp":1776807177874,"version":"3.51.2"},"reference-count":34,"publisher":"Springer Science and Business Media LLC","issue":"9","license":[{"start":{"date-parts":[[2023,2,15]],"date-time":"2023-02-15T00:00:00Z","timestamp":1676419200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,2,15]],"date-time":"2023-02-15T00:00:00Z","timestamp":1676419200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"name":"Tampere University including Tampere University Hospital, Tampere University of Applied Sciences"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Sign Process Syst"],"published-print":{"date-parts":[[2023,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Modern High Level Synthesis (HLS) tools succeed well in their engineering productivity goal, but still require toolset and target technology specific modifications to the source code to guide the process towards an efficient implementation. Furthermore, their end result is a fixed function accelerator with limited field and runtime flexibility. In this paper we describe the status of AEx, a novel work-in-progress HLS tool developed in the FitOptiVis ECSEL JU project. AEx is based on automated exploration of architectures using a flexible and lightweight parallel co-processor template. We compare its current performance in CHStone C-language benchmarks to the state of the art FPGA HLS tool Vitis, provide ASIC implementation numbers, and identify the main remaining toolset features that are expected to dramatically further improve the performance. The potential is explored with a hand-optimized case study that shows only 1.64x performance slowdown with the programmable co-processor in comparison to the fixed function Vitis HLS result.<\/jats:p>","DOI":"10.1007\/s11265-023-01841-3","type":"journal-article","created":{"date-parts":[[2023,2,17]],"date-time":"2023-02-17T07:35:58Z","timestamp":1676619358000},"page":"1051-1065","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["AEx: Automated High-Level Synthesis of Compiler Programmable Co-Processors"],"prefix":"10.1007","volume":"95","author":[{"given":"Alex","family":"Hirvonen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8795-7435","authenticated-orcid":false,"given":"Topi","family":"Lepp\u00e4nen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kari","family":"Hepola","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joonas","family":"Multanen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joost","family":"Hoozemans","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pekka","family":"J\u00e4\u00e4skel\u00e4inen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,2,15]]},"reference":[{"key":"1841_CR1","doi-asserted-by":"crossref","unstructured":"Fahmy,\u00a0S. A., Vipin,\u00a0K., & Shreejith, S. (2015). Virtualized FPGA accelerators for efficient cloud computing. In 2015 IEEE 7th International Conference on Cloud Computing Technology and Science (CloudCom), pp. 430\u2013435.","DOI":"10.1109\/CloudCom.2015.60"},{"key":"1841_CR2","doi-asserted-by":"crossref","unstructured":"Ren, H. (2014) A brief introduction on contemporary high-level synthesis. In Proceedings of the 2014 IEEE International Conference on IC Design & Technology (ICICDT).","DOI":"10.1109\/ICICDT.2014.6838614"},{"key":"1841_CR3","volume-title":"Are we there yet?","author":"S Lahti","year":"2018","unstructured":"Lahti, S., Sj\u00f6vall, P., Vanne, J., & H\u00e4m\u00e4l\u00e4inen, T. D. (2018). Are we there yet? IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems: A study on the state of high-level synthesis."},{"key":"1841_CR4","doi-asserted-by":"crossref","unstructured":"Nane,\u00a0R., et al. (2016). A survey and evaluation of FPGA high-level synthesis tools. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 35(10).","DOI":"10.1109\/TCAD.2015.2513673"},{"key":"1841_CR5","doi-asserted-by":"crossref","unstructured":"Hirvonen,\u00a0A., Tervo,\u00a0K., Kultala,\u00a0H., & J\u00e4\u00e4skel\u00e4\u00efnen, P. (2019). AEx: Automated customization of exposed datapath soft-cores. In 2019 22nd Euromicro Conference on Digital System Design (DSD), pp. 35\u201342.","DOI":"10.1109\/DSD.2019.00016"},{"key":"1841_CR6","doi-asserted-by":"crossref","unstructured":"Al-Ars,\u00a0Z., et al. (2019). The FitOptiVis ECSEL project: Highly efficient distributed embedded image\/video processing in cyber-physical systems. In Proceedings of the 16th ACM International Conference on Computing Frontiers, CF \u201919, pp. 333\u2013338, New York, NY, USA. Association for Computing Machinery.","DOI":"10.1145\/3310273.3323437"},{"key":"1841_CR7","doi-asserted-by":"crossref","unstructured":"Hoogerbrugge,\u00a0J., & Corporaal, H. (1994). Register file port requirements of Transport Triggered Architectures. In IProceedings of the 27th Annual International Symposium on Microarchitecture.","DOI":"10.1145\/192724.192751"},{"key":"1841_CR8","doi-asserted-by":"crossref","unstructured":"Corporaal,\u00a0H., & Hoogerbrugge,\u00a0J. (2002). Code Generation for Transport Triggered Architectures, pp. 240\u2013259. Springer US, Boston, MA.","DOI":"10.1007\/978-1-4615-2323-9_14"},{"issue":"1","key":"1841_CR9","doi-asserted-by":"publisher","first-page":"49","DOI":"10.1007\/s11265-014-0924-x","volume":"80","author":"P J\u00e4\u00e4skel\u00e4inen","year":"2014","unstructured":"J\u00e4\u00e4skel\u00e4inen, P., Kultala, H., Viitanen, T., & Takala, J. (2014). Code density and energy efficiency of exposed datapath architectures. Journal of Signal Processing Systems, 80(1), 49\u201364.","journal-title":"Journal of Signal Processing Systems"},{"key":"1841_CR10","doi-asserted-by":"crossref","unstructured":"J\u00e4\u00e4skel\u00e4inen,\u00a0P., Tervo,\u00a0A., Vay\u00e1,\u00a0G. P., Viitanen,\u00a0T., Behmann,\u00a0N., Takala,\u00a0J., & Blume, H. (2018). Transport-triggered soft cores. In 2018 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW).","DOI":"10.1109\/IPDPSW.2018.00022"},{"key":"1841_CR11","doi-asserted-by":"crossref","unstructured":"Palesi,\u00a0M., & Givargis, T. (2002). Multi-objective design space exploration using genetic algorithms. In Proceedings of the Tenth International Symposium on Hardware\/Software Codesign, CODES \u201902, pp. 67\u201372, New York, NY, USA, 2002. Association for Computing Machinery.","DOI":"10.1145\/774789.774804"},{"key":"1841_CR12","doi-asserted-by":"crossref","unstructured":"Kumar,\u00a0A., & Chakarverty, S. (2011). Design optimization using genetic algorithm and cuckoo search. In 2011 IEEE International Conference on Electro\/Information Technology, pp. 1\u20135.","DOI":"10.1109\/EIT.2011.5978616"},{"key":"1841_CR13","doi-asserted-by":"crossref","unstructured":"Ferrandi,\u00a0F., Lanzi,\u00a0P. L., Loiacono,\u00a0D., Pilato,\u00a0C., & Sciuto, D. (2008). A multi-objective genetic algorithm for design space exploration in high-level synthesis. In 2008 IEEE Computer Society Annual Symposium on VLSI, pp. 417\u2013422.","DOI":"10.1109\/ISVLSI.2008.73"},{"key":"1841_CR14","doi-asserted-by":"crossref","unstructured":"Jordans,\u00a0R., J\u00f3\u017awiak,\u00a0L., & Corporaal, H. (2014). Instruction-set architecture exploration of VLIW ASIPs using a genetic algorithm. In 2014 3rd Mediterranean Conference on Embedded Computing (MECO), pp. 32\u201335.","DOI":"10.1109\/MECO.2014.6862720"},{"key":"1841_CR15","doi-asserted-by":"crossref","unstructured":"Viitanen,\u00a0T., Kultala,\u00a0H., J\u00e4\u00e4skel\u00e4inen,\u00a0P., & Takala, J. (2014). Heuristics for greedy transport triggered architecture interconnect exploration. In Proceedings of the International Conference on Compilers, Architecture and Synthesis for Embedded Systems (CASES). ACM.","DOI":"10.1145\/2656106.2656123"},{"key":"1841_CR16","doi-asserted-by":"crossref","unstructured":"J\u00e4\u00e4skel\u00e4inen,\u00a0P., Viitanen,\u00a0T., Takala,\u00a0J., & Berg, H. (2017). HW\/SW co-design toolset for customization of exposed datapath processors. Computing Platforms for Software-Defined Radio, 147\u2013164.","DOI":"10.1007\/978-3-319-49679-5_8"},{"key":"1841_CR17","doi-asserted-by":"crossref","unstructured":"Esko, O., J\u00e4\u00e4skel\u00e4inen, P., Huerta, P., dela Lama, C. S., Takala,\u00a0J., & Martinez, J. I. (2010). Customized exposed datapath soft-core design flow with compiler support. In Proceedings of the International Conference on Field Programmable Logic and Applications.","DOI":"10.1109\/FPL.2010.51"},{"key":"1841_CR18","unstructured":"Xilinx. Vitis High-Level Synthesis User Guide (UG1399). Retrieved February 14, 2023, from https:\/\/docs.xilinx.com\/r\/en-US\/ug1399-vitis-hls"},{"key":"1841_CR19","doi-asserted-by":"publisher","first-page":"242","DOI":"10.2197\/ipsjjip.17.242","volume":"17","author":"Y Hara","year":"2009","unstructured":"Hara, Y., Tomiyama, H., Honda, S., & Takada, H. (2009). Proposal and quantitative analysis of the CHStone benchmark program suite for practical C-based high-level synthesis. Journal of Information Processing, 17, 242\u2013254.","journal-title":"Journal of Information Processing"},{"key":"1841_CR20","unstructured":"ARM. Cortex-A9.\u00a0Retrieved February 14, 2023, from\u00a0https:\/\/developer.arm.com\/ip-products\/processors\/cortex-a\/cortex-a9"},{"key":"1841_CR21","unstructured":"Digilent. (2017). PYNQ-Z1 Board Reference Manual."},{"key":"1841_CR22","doi-asserted-by":"crossref","unstructured":"Jouppi,\u00a0N., & Wall, D. (1989). Available instruction-level parallelism for superscalar and superpipelined machines. In Proceedings of the third international conference on architectural support for programming languages and operating systems, ASPLOS III, pp. 272\u2013282. ACM.","DOI":"10.1145\/68182.68207"},{"key":"1841_CR23","doi-asserted-by":"crossref","unstructured":"Wall, D. W. (1991). Limits of instruction-level parallelism. SIGPLAN Notices, 26(4), 176\u2013188.","DOI":"10.1145\/106973.106991"},{"key":"1841_CR24","unstructured":"Waterman,\u00a0A., Lee,\u00a0Y., Patterson,\u00a0D. A., & Asanovic, K. (2011). The RISC-V instruction set manual, volume I: Base user-level ISA. EECS Department, UC Berkeley, Technical Report UCB\/EECS-2011-62."},{"key":"1841_CR25","doi-asserted-by":"crossref","unstructured":"Schiavone, P. D., Conti,\u00a0F., Rossi, D., Gautschi,\u00a0M., Pullini,\u00a0A., Flamand,\u00a0E., & Benini, L. (2017). Slow and steady wins the race? A comparison of ultra-low-power RISC-V cores for internet-of-things applications. In Proceedings of International Symposium on Power and Timing Modeling, Optimization and Simulation (PATMOS).","DOI":"10.1109\/PATMOS.2017.8106976"},{"key":"1841_CR26","unstructured":"Synopsys. Design Compiler Graphical.\u00a0Retrieved February 14, 2023, from\u00a0https:\/\/www.synopsys.com\/implementation-and-signoff\/rtl-synthesis-test\/design-compiler-graphical.html"},{"key":"1841_CR27","doi-asserted-by":"crossref","unstructured":"Lam, M. (1988). Software pipelining: an effective scheduling technique for VLIW machines. In Proceedings of the ACM SIGPLAN 1988 conference on programming language design and implementation, PLDI \u201988, pp. 318\u2013328. ACM.","DOI":"10.1145\/53990.54022"},{"key":"1841_CR28","doi-asserted-by":"crossref","unstructured":"Kultala,\u00a0H., Jaaskelainen,\u00a0P., & Takala, J. (2011). Operation set customization in retargetable compilers. In 2011 Conference Record of the Forty Fifth Asilomar Conference on Signals, Systems and Computers (ASILOMAR), pp. 761\u2013765. IEEE.","DOI":"10.1109\/ACSSC.2011.6190108"},{"key":"1841_CR29","doi-asserted-by":"crossref","unstructured":"Tervo,\u00a0K., Malik,\u00a0S., Leppanen,\u00a0T., & J\u00e4\u00e4skel\u00e4inen, P. (2020). TTA-SIMD soft core processors. In 2020 30th International Conference on Field-Programmable Logic and Applications (FPL), pp. 79\u201384.","DOI":"10.1109\/FPL50879.2020.00023"},{"key":"1841_CR30","doi-asserted-by":"crossref","unstructured":"Wolf,\u00a0D. L., Spang,\u00a0C., & Hochberger, C. (2020). Towards purposeful design space exploration of heterogeneous CGRAs: Clock frequency estimation. In 2020 57th ACM\/IEEE Design Automation Conference (DAC), pp. 1\u20136.","DOI":"10.1109\/DAC18072.2020.9218649"},{"key":"1841_CR31","doi-asserted-by":"crossref","unstructured":"Lattner, C., & Adve, V. (2004). LLVM: A compilation framework for lifelong program analysis & transformation.\u00a0In Proceedings of the International Symposium on Code Generation and Optimization, pp. 20\u201324. Palo Alto, CA.","DOI":"10.1109\/CGO.2004.1281665"},{"key":"1841_CR32","doi-asserted-by":"crossref","unstructured":"Nane,\u00a0R., Sima,\u00a0V., Olivier,\u00a0B., Meeuws,\u00a0R., Yankova,\u00a0Y., & Bertels, K. (2012). DWARV 2.0: A CoSy-based C-to-VHDL hardware compiler. In 22nd International Conference on Field Programmable Logic and Applications (FPL).","DOI":"10.1109\/FPL.2012.6339221"},{"key":"1841_CR33","doi-asserted-by":"crossref","unstructured":"Canis,\u00a0A., Choi,\u00a0J., Aldham,\u00a0M., Zhang,\u00a0V., Kammoona,\u00a0A., Anderson,\u00a0J. H., Brown,\u00a0S., & Czajkowski, T. (2011). LegUp: High-level synthesis for FPGA-based processor\/accelerator systems. In Proceedings of the 19th ACM\/SIGDA International Symposium on Field Programmable Gate Arrays, FPGA \u201911, New York, NY, USA. ACM.","DOI":"10.1145\/1950413.1950423"},{"key":"1841_CR34","doi-asserted-by":"crossref","unstructured":"Pilato,\u00a0C., & Ferrandi, F. (2013). Bambu: A modular framework for the high level synthesis of memory-intensive applications. In 23rd International Conference on Field Programmable Logic and Applications.","DOI":"10.1109\/FPL.2013.6645550"}],"container-title":["Journal of Signal Processing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11265-023-01841-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11265-023-01841-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11265-023-01841-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,9]],"date-time":"2023-10-09T08:17:27Z","timestamp":1696839447000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11265-023-01841-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,15]]},"references-count":34,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2023,9]]}},"alternative-id":["1841"],"URL":"https:\/\/doi.org\/10.1007\/s11265-023-01841-3","relation":{},"ISSN":["1939-8018","1939-8115"],"issn-type":[{"value":"1939-8018","type":"print"},{"value":"1939-8115","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,15]]},"assertion":[{"value":"1 April 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 April 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"16 January 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 February 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}