{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T19:50:36Z","timestamp":1781898636399,"version":"3.54.5"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"1-2","license":[{"start":{"date-parts":[[2026,3,2]],"date-time":"2026-03-02T00:00:00Z","timestamp":1772409600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,2]],"date-time":"2026-03-02T00:00:00Z","timestamp":1772409600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100000266","name":"EPSRC","doi-asserted-by":"crossref","award":["00855959"],"award-info":[{"award-number":["00855959"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Parallel Prog"],"published-print":{"date-parts":[[2026,6]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    FPGAs are promising accelerators for scientific computing tasks because of their potential for delivering high performance-per-Watt. However, programming for optimal performance remains a complex task. Our goal is to bring FPGAs within the reach of domain scientists by developing compilers targeting scientific Fortran code. In this paper, we present a novel approach to aggressively reduce memory utilisation of stencil-based finite-difference Fortran code through a compiler-based automatic program transformation that trades memory accesses for computation. The key contribution of this work is a set of type-driven rewrite rules that identify and eliminate the intermediate arrays in stencil computations and replace them with re-computation, thus reducing the number of memory accesses. The main novelty lies in the transformations to move stencil operations out of maps and folds and to fuse stencils. We demonstrate the effectiveness of our approach using a set of five 3-D and 2-D stencil benchmarks evaluated on an Intel Arria 10 FPGA board. Our transformation result on average in a\n                    <jats:inline-formula>\n                      <jats:tex-math>$$2.5\\times$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    reduction in DRAM usage, a\n                    <jats:inline-formula>\n                      <jats:tex-math>$$3.4\\times$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    increase in DSP usage, and a\n                    <jats:inline-formula>\n                      <jats:tex-math>$$25\\times$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    improvement in throughput. We show on a real-world exemplar, the Large Eddy Simulator for Urban Flows, that our algorithm successfully removes all intermediate arrays (18 in total), reducing the memory footprint by a factor of\n                    <jats:inline-formula>\n                      <jats:tex-math>$$4.5\\times$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    . The memory-reduced FPGA code, automatically transpiled from the original Fortran source, is\n                    <jats:inline-formula>\n                      <jats:tex-math>$$9.7\\times$$<\/jats:tex-math>\n                    <\/jats:inline-formula>\n                    faster than the original Fortran code, and performance competitive with a hand optimised FPGA version while supporting a four times larger domain size.\n                  <\/jats:p>","DOI":"10.1007\/s10766-025-00809-z","type":"journal-article","created":{"date-parts":[[2026,3,2]],"date-time":"2026-03-02T19:44:59Z","timestamp":1772480699000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Optimising Stencil Code on FPGAs by Trading Data Movement for Compute using Compiler Rewrite Rules"],"prefix":"10.1007","volume":"54","author":[{"given":"Robert","family":"Szafarczyk","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Syed Waqar","family":"Nabi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wim","family":"Vanderbauwhede","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,3,2]]},"reference":[{"issue":"7","key":"809_CR1","doi-asserted-by":"publisher","first-page":"2052","DOI":"10.1002\/cpe.3522","volume":"28","author":"W Vanderbauwhede","year":"2016","unstructured":"Vanderbauwhede, W., Takemi, T.: An analysis of the feasibility and benefits of GPU\/multicore acceleration of the weather research and forecasting model. Concurr. Comput. Pract. Exp. 28(7), 2052\u20132072 (2016)","journal-title":"Concurr. Comput. Pract. Exp."},{"issue":"1","key":"809_CR2","doi-asserted-by":"publisher","first-page":"114","DOI":"10.1007\/s10766-018-0572-z","volume":"47","author":"W Vanderbauwhede","year":"2019","unstructured":"Vanderbauwhede, W., Nabi, S.W., Urlea, C.: Type-driven automated program transformations and cost modelling for optimizing streaming programs on FPGAS. Int. J. Parallel Prog. 47(1), 114\u2013136 (2019)","journal-title":"Int. J. Parallel Prog."},{"key":"809_CR3","doi-asserted-by":"crossref","unstructured":"de la Cruz, R., Araya-Polo, M.: Modeling stencil computations on modern hpc architectures. In: International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems, pp. 149\u2013171 (2014)","DOI":"10.1007\/978-3-319-17248-4_8"},{"key":"809_CR4","doi-asserted-by":"publisher","unstructured":"Krueger, J., Donofrio, D., Shalf, J., Mohiyuddin, M., Williams, S., Oliker, L., Pfreundt, F.-J.: Hardware\/software co-design for energy-efficient seismic modeling. In: SC \u201911: Proceedings of 2011 International Conference for High Performance Computing, Networking, Storage and Analysis, pp. 1\u201312 (2011). https:\/\/doi.org\/10.1145\/2063384.2063482","DOI":"10.1145\/2063384.2063482"},{"key":"809_CR5","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2022.111287","volume":"466","author":"X Deng","year":"2022","unstructured":"Deng, X., Jiang, Z.-H., Vincent, P., Xiao, F., Yan, C.: A new paradigm of dissipation-adjustable, multi-scale resolving schemes for compressible flows. J. Comput. Phys. 466, 111287 (2022). https:\/\/doi.org\/10.1016\/j.jcp.2022.111287","journal-title":"J. Comput. Phys."},{"key":"809_CR6","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2021.110821","volume":"450","author":"L Lecointre","year":"2022","unstructured":"Lecointre, L., Vicquelin, R., Kudriakov, S., Studer, E., Tenaud, C.: High-order numerical scheme for compressible multi-component real gas flows using an extension of the Roe approximate Riemann solver and specific Monotonicity-Preserving constraints. J. Comput. Phys. 450, 110821 (2022). https:\/\/doi.org\/10.1016\/j.jcp.2021.110821","journal-title":"J. Comput. Phys."},{"key":"809_CR7","doi-asserted-by":"publisher","first-page":"302","DOI":"10.1016\/j.jcp.2012.10.048","volume":"235","author":"N Mai-Duy","year":"2013","unstructured":"Mai-Duy, N., Tran-Cong, T.: A compact five-point stencil based on integrated RBFs for 2D second-order differential problems. J. Comput. Phys. 235, 302\u2013321 (2013). https:\/\/doi.org\/10.1016\/j.jcp.2012.10.048","journal-title":"J. Comput. Phys."},{"issue":"12","key":"809_CR8","doi-asserted-by":"publisher","first-page":"4419","DOI":"10.1016\/j.jcp.2011.01.039","volume":"230","author":"Y Shen","year":"2011","unstructured":"Shen, Y., Zha, G.: Generalized finite compact difference scheme for shock\/complex flowfield interaction. J. Comput. Phys. 230(12), 4419\u20134436 (2011). https:\/\/doi.org\/10.1016\/j.jcp.2011.01.039","journal-title":"J. Comput. Phys."},{"key":"809_CR9","doi-asserted-by":"publisher","unstructured":"Rivera, G., Tseng, C.-W.: Tiling optimizations for 3d scientific computations. In: SC \u201900: Proceedings of the 2000 ACM\/IEEE Conference on Supercomputing, pp. 32\u201332 (2000). https:\/\/doi.org\/10.1109\/SC.2000.10015","DOI":"10.1109\/SC.2000.10015"},{"key":"809_CR10","doi-asserted-by":"publisher","unstructured":"Datta, K., Murphy, M., Volkov, V., Williams, S., Carter, J., Oliker, L., Patterson, D., Shalf, J., Yelick, K.: Stencil computation optimization and auto-tuning on state-of-the-art multicore architectures. In: SC \u201908: Proceedings of the 2008 ACM\/IEEE Conference on Supercomputing, pp. 1\u201312 (2008). https:\/\/doi.org\/10.1109\/SC.2008.5222004","DOI":"10.1109\/SC.2008.5222004"},{"key":"809_CR11","doi-asserted-by":"publisher","unstructured":"Holewinski, J., Pouchet, L.-N., Sadayappan, P.: High-performance code generation for stencil computations on gpu architectures. In: Proceedings of the 26th ACM International Conference on Supercomputing. ICS \u201912, pp. 311\u2013320. Association for Computing Machinery, New York, NY, USA (2012). https:\/\/doi.org\/10.1145\/2304576.2304619","DOI":"10.1145\/2304576.2304619"},{"key":"809_CR12","doi-asserted-by":"publisher","unstructured":"Krishnamoorthy, S., Baskaran, M., Bondhugula, U., Ramanujam, J., Rountev, A., Sadayappan, P.: Effective automatic parallelization of stencil computations. PLDI \u201907 (2007). https:\/\/doi.org\/10.1145\/1250734.1250761","DOI":"10.1145\/1250734.1250761"},{"key":"809_CR13","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.compfluid.2018.06.005","volume":"173","author":"W Vanderbauwhede","year":"2018","unstructured":"Vanderbauwhede, W., Davidson, G.: Domain-specific acceleration and auto-parallelization of legacy scientific code in FORTRAN 77 using source-to-source compilation. Comput. Fluids 173, 1\u20135 (2018). https:\/\/doi.org\/10.1016\/j.compfluid.2018.06.005","journal-title":"Comput. Fluids"},{"issue":"5","key":"809_CR14","doi-asserted-by":"publisher","first-page":"1014","DOI":"10.1109\/TPDS.2020.3039409","volume":"32","author":"J de Fine Licht","year":"2021","unstructured":"de Fine Licht, J., Besta, M., Meierhans, S., Hoefler, T.: Transformations of high-level synthesis codes for high-performance computing. IEEE Trans. Parallel Distrib. Syst. 32(5), 1014\u20131029 (2021). https:\/\/doi.org\/10.1109\/TPDS.2020.3039409","journal-title":"IEEE Trans. Parallel Distrib. Syst."},{"key":"809_CR15","doi-asserted-by":"publisher","unstructured":"Kenter, T., Mahale, G., Alhaddad, S., Grynko, Y., Schmitt, C., Afzal, A., Hannig, F., F\u00f6rstner, J., Plessl, C.: Opencl-based fpga design to accelerate the nodal discontinuous galerkin method for unstructured meshes. In: 2018 IEEE 26th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM), pp. 189\u2013196 (2018). https:\/\/doi.org\/10.1109\/FCCM.2018.00037","DOI":"10.1109\/FCCM.2018.00037"},{"key":"809_CR16","doi-asserted-by":"publisher","unstructured":"Weller, D., Oboril, F., Lukarski, D., Becker, J., Tahoori, M.: Energy efficient scientific computing on fpgas using opencl. In: Proceedings of the 2017 ACM\/SIGDA International Symposium on Field-Programmable Gate Arrays. FPGA \u201917, pp. 247\u2013256. Association for Computing Machinery, New York, NY, USA (2017). https:\/\/doi.org\/10.1145\/3020078.3021730","DOI":"10.1145\/3020078.3021730"},{"key":"809_CR17","doi-asserted-by":"publisher","unstructured":"Mullapudi, R.T., Vasista, V., Bondhugula, U.: Polymage: automatic optimization for image processing pipelines. ASPLOS \u201915 (2015). https:\/\/doi.org\/10.1145\/2694344.2694364","DOI":"10.1145\/2694344.2694364"},{"key":"809_CR18","doi-asserted-by":"publisher","unstructured":"Chugh, N., Vasista, V., Purini, S., Bondhugula, U.: A DSL compiler for accelerating image processing pipelines on fpgas. PACT \u201916 (2016). https:\/\/doi.org\/10.1145\/2967938.2967969","DOI":"10.1145\/2967938.2967969"},{"issue":"6","key":"809_CR19","doi-asserted-by":"publisher","first-page":"519","DOI":"10.1145\/2499370.2462176","volume":"48","author":"J Ragan-Kelley","year":"2013","unstructured":"Ragan-Kelley, J., Barnes, C., Adams, A., Paris, S., Durand, F., Amarasinghe, S.: Halide: a language and compiler for optimizing parallelism, locality, and recomputation in image processing pipelines. SIGPLAN Not. 48(6), 519\u2013530 (2013). https:\/\/doi.org\/10.1145\/2499370.2462176","journal-title":"SIGPLAN Not."},{"key":"809_CR20","doi-asserted-by":"publisher","unstructured":"Li, J., Chi, Y., Cong, J.: Heterohalide: from image processing dsl to efficient fpga acceleration. FPGA \u201920 (2020). https:\/\/doi.org\/10.1145\/3373087.3375320","DOI":"10.1145\/3373087.3375320"},{"key":"809_CR21","doi-asserted-by":"publisher","unstructured":"Chi, Y., Cong, J., Wei, P., Zhou, P.: Soda: stencil with optimized dataflow architecture. ICCAD \u201918 (2018). https:\/\/doi.org\/10.1145\/3240765.3240850","DOI":"10.1145\/3240765.3240850"},{"key":"809_CR22","doi-asserted-by":"publisher","first-page":"7348013","DOI":"10.1155\/2019\/7348013","volume":"2019","author":"SW Nabi","year":"2019","unstructured":"Nabi, S.W., Vanderbauwhede, W.: Automatic pipelining and vectorization of scientific code for FPGAs. Int. J. Reconfig. Comput. 2019, 7348013 (2019). https:\/\/doi.org\/10.1155\/2019\/7348013. (Publisher: Hindawi)","journal-title":"Int. J. Reconfig. Comput."},{"key":"809_CR23","doi-asserted-by":"publisher","unstructured":"Ben-Nun, T., de Fine Licht, J., Ziogas, A.N., Schneider, T., Hoefler, T.: Stateful dataflow multigraphs: a data-centric model for performance portability on heterogeneous architectures. SC \u201919 (2019). https:\/\/doi.org\/10.1145\/3295500.3356173","DOI":"10.1145\/3295500.3356173"},{"key":"809_CR24","doi-asserted-by":"publisher","unstructured":"de Fine\u00a0Licht, J., Kuster, A., De\u00a0Matteis, T., Ben-Nun, T., Hofer, D., Hoefler, T.: StencilFlow: mapping large stencil programs to distributed spatial computing systems, (2021). https:\/\/doi.org\/10.1109\/CGO51591.2021.9370315","DOI":"10.1109\/CGO51591.2021.9370315"},{"issue":"2","key":"809_CR25","doi-asserted-by":"publisher","first-page":"2988","DOI":"10.1007\/s11227-021-03839-9","volume":"78","author":"W Vanderbauwhede","year":"2022","unstructured":"Vanderbauwhede, W.: Making legacy Fortran code type safe through automated program transformation. J. Supercomput. 78(2), 2988\u20133028 (2022). https:\/\/doi.org\/10.1007\/s11227-021-03839-9","journal-title":"J. Supercomput."},{"issue":"1","key":"809_CR26","doi-asserted-by":"publisher","first-page":"114","DOI":"10.1007\/s10766-018-0572-z","volume":"47","author":"W Vanderbauwhede","year":"2019","unstructured":"Vanderbauwhede, W., Nabi, S.W., Urlea, C.: Type-Driven automated program transformations and cost modelling for optimizing streaming programs on FPGAs. Int. J. Parallel Prog. 47(1), 114\u2013136 (2019). https:\/\/doi.org\/10.1007\/s10766-018-0572-z","journal-title":"Int. J. Parallel Prog."},{"key":"809_CR27","unstructured":"Marlow, S., et al.: Haskell 2010 language report. (2010) Available on: https:\/\/www.haskell.org\/onlinereport\/haskell2010"},{"issue":"5","key":"809_CR28","doi-asserted-by":"publisher","first-page":"552","DOI":"10.1017\/S095679681300018X","volume":"23","author":"E Brady","year":"2013","unstructured":"Brady, E.: Idris, a general-purpose dependently typed programming language: design and implementation. J. Funct. Prog. 23(5), 552\u2013593 (2013)","journal-title":"J. Funct. Prog."},{"issue":"4","key":"809_CR29","doi-asserted-by":"publisher","first-page":"383","DOI":"10.1007\/s10766-006-0018-x","volume":"34","author":"C Grelck","year":"2006","unstructured":"Grelck, C., Scholz, S.-B.: SAC\u2014a functional array language for efficient multi-threaded execution. Int. J. Parallel Prog. 34(4), 383\u2013427 (2006)","journal-title":"Int. J. Parallel Prog."},{"key":"809_CR30","doi-asserted-by":"publisher","first-page":"2988","DOI":"10.1007\/s11227-021-03839-9","volume":"78","author":"W Vanderbauwhede","year":"2021","unstructured":"Vanderbauwhede, W.: Making legacy Fortran code type safe through automated program transformation. J. Supercomput. 78, 2988\u20133028 (2021)","journal-title":"J. Supercomput."},{"key":"809_CR31","doi-asserted-by":"crossref","unstructured":"Klop, J.W., Klop, J.: Term Rewriting Systems vol. Centrum voor Wiskunde en Informatica, (1990)","DOI":"10.1007\/3-540-54317-1_79"},{"key":"809_CR32","unstructured":"van Noort, T.R.: Dynamic Typing in Type-driven Programming vol. Radboud Universiteit Nijmegen, (2012). https:\/\/hdl.handle.net\/2066\/92736"},{"issue":"4","key":"809_CR33","doi-asserted-by":"publisher","first-page":"434","DOI":"10.1017\/S0956796814000185","volume":"24","author":"N Sculthorpe","year":"2014","unstructured":"Sculthorpe, N., Frisby, N., Gill, A.: The Kansas university rewrite engine: a Haskell-embedded strategic programming language with custom closed universes. J. Funct. Program. 24(4), 434\u2013473 (2014)","journal-title":"J. Funct. Program."},{"issue":"3\u20134","key":"809_CR34","first-page":"289","volume":"6","author":"A Sabry","year":"1993","unstructured":"Sabry, A., Felleisen, M.: Reasoning about programs in continuation-passing style. LISP Symb. Comput. 6(3\u20134), 289\u2013360 (1993)","journal-title":"LISP Symb. Comput."},{"issue":"1","key":"809_CR35","doi-asserted-by":"publisher","first-page":"127","DOI":"10.1007\/s10546-018-0344-8","volume":"168","author":"T Yoshida","year":"2018","unstructured":"Yoshida, T., Takemi, T., Horiguchi, M.: Large-eddy-simulation study of the effects of building-height variability on turbulent flows over an actual urban area. Bound.-Layer Meteorol. 168(1), 127\u2013153 (2018)","journal-title":"Bound.-Layer Meteorol."},{"issue":"3","key":"809_CR36","doi-asserted-by":"publisher","first-page":"99","DOI":"10.1175\/1520-0493(1963)091<0099:GCEWTP>2.3.CO;2","volume":"91","author":"J Smagorinsky","year":"1963","unstructured":"Smagorinsky, J.: General circulation experiments with the primitive equations: I. The basic experiment. Mon. Weather. Rev. 91(3), 99\u2013164 (1963). https:\/\/journals.ametsoc.org\/view\/journals\/mwre\/91\/3\/1520-0493_1963_091_0099_gcewtp_2_3_co_2.xml","journal-title":"Mon. Weather. Rev."},{"key":"809_CR37","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-00820-7","volume-title":"Ocean Modelling for Beginners: Using Open-source Software","author":"J K\u00e4mpf","year":"2009","unstructured":"K\u00e4mpf, J.: Ocean Modelling for Beginners: Using Open-source Software. Springer Science & Business Media, Germany (2009)"},{"key":"809_CR38","doi-asserted-by":"publisher","unstructured":"Che, S., Boyer, M., Meng, J., Tarjan, D., Sheaffer, J.W., Lee, S.-H., Skadron, K.: Rodinia: A benchmark suite for heterogeneous computing. In: 2009 IEEE International Symposium on Workload Characterization (IISWC), pp. 44\u201354 (2009). https:\/\/doi.org\/10.1109\/IISWC.2009.5306798","DOI":"10.1109\/IISWC.2009.5306798"},{"key":"809_CR39","doi-asserted-by":"publisher","unstructured":"Hutton, M., Schleicher, J., Lewis, D., Pedersen, B., Yuan, R., Kaptanoglu, S., Baeckler, G., Ratchev, B., Padalia, K., Bourgeault, M., Lee, A., Kim, H., Saini, R.: Improving FPGA performance and area using an adaptive logic module, vol. 3203, pp. 135\u2013144 (2004). https:\/\/doi.org\/10.1007\/978-3-540-30117-2_16","DOI":"10.1007\/978-3-540-30117-2_16"},{"issue":"6","key":"809_CR40","doi-asserted-by":"publisher","first-page":"15","DOI":"10.1145\/113445.113448","volume":"26","author":"G Goff","year":"1991","unstructured":"Goff, G., Kennedy, K., Tseng, C.-W.: Practical dependence testing. ACM SIGPLAN Not. 26(6), 15\u201329 (1991). https:\/\/doi.org\/10.1145\/113445.113448","journal-title":"ACM SIGPLAN Not."}],"container-title":["International Journal of Parallel Programming"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-025-00809-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10766-025-00809-z","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10766-025-00809-z.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T18:53:24Z","timestamp":1781895204000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10766-025-00809-z"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,2]]},"references-count":40,"journal-issue":{"issue":"1-2","published-print":{"date-parts":[[2026,6]]}},"alternative-id":["809"],"URL":"https:\/\/doi.org\/10.1007\/s10766-025-00809-z","relation":{},"ISSN":["0885-7458","1573-7640"],"issn-type":[{"value":"0885-7458","type":"print"},{"value":"1573-7640","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,2]]},"assertion":[{"value":"12 June 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 December 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 March 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no Conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}],"article-number":"1"}}