{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,11]],"date-time":"2026-03-11T01:55:16Z","timestamp":1773194116651,"version":"3.50.1"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,9,1]],"date-time":"2023-09-01T00:00:00Z","timestamp":1693526400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000266","name":"EPSRC","doi-asserted-by":"crossref","award":["EP\/P010040\/1, EP\/R006865\/1"],"award-info":[{"award-number":["EP\/P010040\/1, EP\/R006865\/1"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Reconfigurable Technol. Syst."],"published-print":{"date-parts":[[2023,12,31]]},"abstract":"<jats:p>\n            Recently, there is a trend to use high-level synthesis (HLS) tools to generate dynamically scheduled hardware. The generated hardware is made up of components connected using handshake signals. These handshake signals schedule the components at runtime when inputs become available. Such approaches promise superior performance on \u201cirregular\u201d source programs, such as those whose control flow depends on input data. This is at the cost of additional area. Current dynamic scheduling techniques are well able to exploit parallelism among instructions\n            <jats:italic>within<\/jats:italic>\n            each basic block (BB) of the source program, but parallelism\n            <jats:italic>between<\/jats:italic>\n            BBs is under-explored, due to the complexity in runtime control flows and memory dependencies. Existing tools allow some of the operations of different BBs to overlap, but to simplify the analysis required at compile time they require the BBs to\n            <jats:italic>start<\/jats:italic>\n            in strict program order, thus limiting the achievable parallelism and overall performance.\n          <\/jats:p>\n          <jats:p>We formulate a general dependency model suitable for comparing the ability of different dynamic scheduling approaches to extract maximal parallelism at runtime. Using this model, we explore a variety of mechanisms for runtime scheduling, incorporating and generalising existing approaches. In particular, we precisely identify the restrictions in existing scheduling implementation and define possible optimisation solutions. We identify two particularly promising examples where the compile-time overhead is small and the area overhead is minimal and yet we are able to significantly speed up execution time: (1)\u00a0parallelising consecutive independent loops; and (2)\u00a0parallelising independent inner-loop instances in a nested loop as individual threads. Using benchmark sets from related works, we compare our proposed toolflow against a state-of-the-art dynamic-scheduling HLS tool called Dynamatic. Our results show that, on average, our toolflow yields a 4\u00d7 speedup from (1) and a 2.9\u00d7 speedup from (2), with a negligible area overhead. This increases to a 14.3\u00d7 average speedup when combining (1) and (2).<\/jats:p>","DOI":"10.1145\/3599973","type":"journal-article","created":{"date-parts":[[2023,5,31]],"date-time":"2023-05-31T11:25:55Z","timestamp":1685532355000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["Parallelising Control Flow in Dynamic-scheduling High-level Synthesis"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2791-2555","authenticated-orcid":false,"given":"Jianyi","family":"Cheng","sequence":"first","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6659-8533","authenticated-orcid":false,"given":"Lana","family":"Josipovi\u0107","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6735-5533","authenticated-orcid":false,"given":"John","family":"Wickerson","sequence":"additional","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0201-310X","authenticated-orcid":false,"given":"George A.","family":"Constantinides","sequence":"additional","affiliation":[{"name":"Imperial College London, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,9]]},"reference":[{"key":"e_1_3_2_2_2","unstructured":"Amazon. 2022. Amazon EC2 F1 Instances. Retrieved from https:\/\/aws.amazon.com\/ec2\/instance-types\/f1\/."},{"key":"e_1_3_2_3_2","unstructured":"B. Barney. 2021. POSIX Threads Programming. Retrieved from https:\/\/computing.llnl.gov\/tutorials\/pthreads."},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/1375581.1375595"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/263953"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","first-page":"17","DOI":"10.1109\/ICSAMOS.2009.5289237","volume-title":"International Symposium on Systems, Architectures, Modeling, and Simulation","author":"Cabrera Daniel","year":"2009","unstructured":"Daniel Cabrera, Xavier Martorell, Georgi Gaydadjiev, Eduard Ayguade, and Daniel Jim\u00e9nez-Gonz\u00e1lez. 2009. OpenMP extensions for FPGA accelerators. In International Symposium on Systems, Architectures, Modeling, and Simulation. IEEE, 17\u201324."},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL.2014.6927490"},{"key":"e_1_3_2_8_2","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1145\/1950413.1950423","volume-title":"19th ACM\/SIGDA International Symposium on Field Programmable Gate Arrays (FPGA\u201911)","author":"Canis Andrew","year":"2011","unstructured":"Andrew Canis, Jongsok Choi, Mark Aldham, Victor Zhang, Ahmed Kammoona, Jason H. Anderson, Stephen Brown, and Tomasz Czajkowski. 2011. LegUp: High-level synthesis for FPGA-based processor\/accelerator systems. In 19th ACM\/SIGDA International Symposium on Field Programmable Gate Arrays (FPGA\u201911). ACM, 33\u201336."},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/43.945302"},{"key":"e_1_3_2_10_2","volume-title":"IEEE Hot Chips 26 Symposium (HCS\u201914)","author":"Castellana V. G.","year":"2014","unstructured":"V. G. Castellana, A. Tumeo, and F. Ferrandi. 2014. High-level synthesis of memory bound and irregular parallel applications with Bambu. In IEEE Hot Chips 26 Symposium (HCS\u201914). IEEE."},{"key":"e_1_3_2_11_2","unstructured":"Catapult High-Level Synthesis. 2021. Retrieved from https:\/\/www.mentor.com\/hls-lp\/catapult-high-level-synthesis."},{"key":"e_1_3_2_12_2","unstructured":"Celoxica. 2005. Handel-C. Retrieved from http:\/\/www.celoxica.com."},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.23919\/FPL.2017.8056841"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2021.3066466"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375297"},{"key":"e_1_3_2_16_2","volume-title":"32nd International Conference on Field-Programmable Logic and Applications (FPL\u201922)","author":"Cheng Jianyi","year":"2022","unstructured":"Jianyi Cheng, Lana Josipovi\u0107, George A. Constantinides, and John Wickerson. 2022. Dynamic inter-block scheduling for HLS. In 32nd International Conference on Field-Programmable Logic and Applications (FPL\u201922)."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPL53798.2021.00066"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM53951.2022.9786096"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPT.2013.6718365"},{"key":"e_1_3_2_20_2","volume-title":"Scalable Verification Techniques for Data-parallel Programs","author":"Chong N. Y. S.","year":"2014","unstructured":"N. Y. S. Chong. 2014. Scalable Verification Techniques for Data-parallel Programs. Doctoral Thesis. Imperial College London, London, UK."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/1146909.1147025"},{"key":"e_1_3_2_22_2","unstructured":"CIRCT contributors. 2023. CIRCT-based HLS Compilation Flows Debugging and Cosimulation Tools. Retrieved from https:\/\/github.com\/circt-hls\/circt-hls."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/MDT.2009.69"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/99.660313"},{"key":"e_1_3_2_25_2","doi-asserted-by":"crossref","first-page":"337","DOI":"10.1007\/978-3-540-78800-3_24","volume-title":"Tools and Algorithms for the Construction and Analysis of Systems","author":"Moura Leonardo de","year":"2008","unstructured":"Leonardo de Moura and Nikolaj Bj\u00f8rner. 2008. Z3: An efficient SMT solver. In Tools and Algorithms for the Construction and Analysis of Systems, C. R. Ramakrishnan and Jakob Rehof (Eds.). Springer Berlin, 337\u2013340."},{"key":"e_1_3_2_26_2","volume-title":"1st International Workshop on Polyhedral Compilation Techniques (IMPACT\u201911)","author":"Grosser Tobias","year":"2011","unstructured":"Tobias Grosser, Hongbin Zheng, Raghesh Aloor, Andreas Simb\u00fcrger, Armin Gr\u00f6\u00dflinger, and Louis-No\u00ebl Pouchet. 2011. Polly\u2014Polyhedral optimization in LLVM. In 1st International Workshop on Polyhedral Compilation Techniques (IMPACT\u201911)."},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICVD.2003.1183177"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS.2008.4541637"},{"key":"e_1_3_2_29_2","unstructured":"Ian Page and Wayne Luk. 1991. Compiling occam into field-programmable gate arrays. In FPGAs Oxford Workshop on Field Programmable Logic and Applications Vol. 15. Abingdon EE&CS Books Abingdon 271\u2013283."},{"key":"e_1_3_2_30_2","unstructured":"Intel Compiler. 2022. Retrieved from https:\/\/www.intel.com\/content\/www\/us\/en\/developer\/tools\/oneapi\/dpc-compiler.html#gs.sa60u7."},{"key":"e_1_3_2_31_2","unstructured":"Intel FPGA SDK for OpenCL Software Technology. 2021. Retrieved from https:\/\/www.intel.co.uk\/content\/www\/uk\/en\/software\/programmable\/sdk-for-opencl\/overview.html."},{"key":"e_1_3_2_32_2","unstructured":"Intel HLS Compiler. 2022. Retrieved from https:\/\/www.intel.co.uk\/content\/www\/uk\/en\/software\/programmable\/quartus-prime\/hls-compiler.html."},{"issue":"5","key":"e_1_3_2_33_2","article-title":"An out-of-order load-store queue for spatial computing","volume":"16","author":"Josipovi\u0107 Lana","year":"2017","unstructured":"Lana Josipovi\u0107, Philip Brisk, and Paolo Ienne. 2017. An out-of-order load-store queue for spatial computing. ACM Trans. Embed. Comput. Syst. 16, 5s (Sept.2017).","journal-title":"ACM Trans. Embed. Comput. Syst."},{"key":"e_1_3_2_34_2","doi-asserted-by":"crossref","first-page":"127","DOI":"10.1145\/3174243.3174264","volume-title":"ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201918)","author":"Josipovi\u0107 Lana","year":"2018","unstructured":"Lana Josipovi\u0107, Radhika Ghosal, and Paolo Ienne. 2018. Dynamically scheduled high-level synthesis. In ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays (FPGA\u201918). ACM, 127\u2013136."},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/FCCM53951.2022.9786084"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2021.3105574"},{"key":"e_1_3_2_37_2","unstructured":"K. Rustan M. Leino. 2008. This is Boogie 2. Retrieved from https:\/\/www.microsoft.com\/en-us\/research\/publication\/this-is-boogie-2-2\/."},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-95432-0_7"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/FPT.2006.270297"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/ASYNC48570.2021.00009"},{"key":"e_1_3_2_41_2","first-page":"72","volume-title":"IEEE 24th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201916)","author":"Liu J.","year":"2016","unstructured":"J. Liu, J. Wickerson, and G. A. Constantinides. 2016. Loop splitting for efficient pipelining in high-level synthesis. In IEEE 24th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM\u201916). 72\u201379. DOI:urlhttps:\/\/doi.org\/10.1109\/FCCM.2016.27"},{"key":"e_1_3_2_42_2","doi-asserted-by":"crossref","first-page":"251","DOI":"10.1109\/FCCM.2007.18","volume-title":"15th Annual IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM\u201907)","author":"Liu Q.","year":"2007","unstructured":"Q. Liu, G. A. Constantinides, K. Masselos, and P. Y. K. Cheung. 2007. Automatic on-chip memory minimization for data reuse. In 15th Annual IEEE Symposium on Field-Programmable Custom Computing Machines (FCCM\u201907). IEEE, 251\u2013260."},{"key":"e_1_3_2_43_2","unstructured":"Yury Markovskiy and Yatish Patel. 2002. Simple symmetric multithreading in xilinx FPGAs. (2002)."},{"key":"e_1_3_2_44_2","unstructured":"Microsoft. 2022. Project Catapult. Retrieved from https:\/\/www.microsoft.com\/en-us\/research\/project\/project-catapult\/."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/ITC-CSCC.2019.8793453"},{"key":"e_1_3_2_46_2","unstructured":"Louis-No\u00ebl Pouchet et\u00a0al. 2012. PolyBench: The polyhedral benchmark suite. Retrieved from http:\/\/www.cs.ucla.edu\/pouchet\/software\/polybench."},{"key":"e_1_3_2_47_2","unstructured":"Stratus High-Level Synthesis. 2021. Retrieved from https:\/\/www.cadence.com\/en_US\/home\/tools\/digital-design-and-signoff\/synthesis\/stratus-high-level-synthesis.html."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375317"},{"key":"e_1_3_2_49_2","volume-title":"IEEE 13th International Workshop on Logic Synthesis (IWLS\u201904)","author":"Venkataramani Girish","year":"2004","unstructured":"Girish Venkataramani, Mihai Budiu, Tiberiu Chelcea, and Seth Copen Goldstein. 2004. C to asynchronous dataflow circuits: An end-to-end toolflow. In IEEE 13th International Workshop on Logic Synthesis (IWLS\u201904)."},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3431920.3439292"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1145\/611817.611845"},{"key":"e_1_3_2_52_2","unstructured":"Xilinx Vivado HLS. 2022. Retrieved from https:\/\/www.xilinx.com\/support\/documentation-navigation\/design-hubs\/dh0012-vivado-high-level-synthesis-hub.html."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3373087.3375300"},{"key":"e_1_3_2_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCAD.2013.6691121"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1145\/3020078.3021734"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/2435264.2435271"}],"container-title":["ACM Transactions on Reconfigurable Technology and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3599973","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3599973","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:47:14Z","timestamp":1750178834000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3599973"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9]]},"references-count":55,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,12,31]]}},"alternative-id":["10.1145\/3599973"],"URL":"https:\/\/doi.org\/10.1145\/3599973","relation":{},"ISSN":["1936-7406","1936-7414"],"issn-type":[{"value":"1936-7406","type":"print"},{"value":"1936-7414","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9]]},"assertion":[{"value":"2022-11-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-10","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}