{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T07:39:14Z","timestamp":1740123554380,"version":"3.37.3"},"reference-count":34,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2022,11,29]],"date-time":"2022-11-29T00:00:00Z","timestamp":1669680000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,11,29]],"date-time":"2022-11-29T00:00:00Z","timestamp":1669680000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2023,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Heterogeneous computing is the major driving factor in designing new energy-efficient high-performance computing systems. Despite the broad adoption of GPUs and other specialized architectures, the interest in spatial architectures like field-programmable gate arrays (FPGAs) has grown. While combining high performance, low power consumption and high adaptability constitute an advantage, these devices still suffer from a weak software ecosystem, which forces application developers to use tools requiring deep knowledge of the underlying system, often leaving legacy code (e.g., Fortran applications) unsupported. By realizing this, we describe a methodology for porting Fortran (legacy) code on modern FPGA architectures, with the target of preserving performance\/power ratios. Aimed as an experience report, we considered an industrial computational fluid dynamics application to demonstrate that our methodology produces synthesizable OpenCL codes targeting Intel Arria10 and Stratix10 devices. Although performance gain is not far beyond that of the original CPU code (we obtained a relative speedup of <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\times$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u00d7<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>\u00a00.59 and <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\times$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u00d7<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>\u00a00.63, respectively, for a single optimized main kernel, while only on the Stratix10 we achieved <jats:inline-formula><jats:alternatives><jats:tex-math>$$\\times$$<\/jats:tex-math><mml:math xmlns:mml=\"http:\/\/www.w3.org\/1998\/Math\/MathML\">\n                  <mml:mo>\u00d7<\/mml:mo>\n                <\/mml:math><\/jats:alternatives><\/jats:inline-formula>\u00a02.56 by replicating the main optimized kernel 4 times), our results are quite encouraging to drawn the path for further investigations. This paper also reports some major criticalities in porting Fortran code on FPGA architectures.<\/jats:p>","DOI":"10.1007\/s11227-022-04925-2","type":"journal-article","created":{"date-parts":[[2022,11,30]],"date-time":"2022-11-30T06:15:27Z","timestamp":1669788927000},"page":"7461-7483","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Accelerating legacy applications with spatial computing devices"],"prefix":"10.1007","volume":"79","author":[{"given":"Paolo","family":"Savio","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8138-9403","authenticated-orcid":false,"given":"Alberto","family":"Scionti","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Giacomo","family":"Vitali","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Paolo","family":"Viviani","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chiara","family":"Vercellino","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Olivier","family":"Terzo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huy-Nam","family":"Nguyen","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Donato","family":"Magarielli","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ennio","family":"Spano","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michele","family":"Marconcini","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francesco","family":"Poli","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,11,29]]},"reference":[{"key":"4925_CR1","doi-asserted-by":"crossref","unstructured":"Horowitz M (2014) 1.1 Computing\u2019s energy problem (and what we can do about it). In: 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC). IEEE","DOI":"10.1109\/ISSCC.2014.6757323"},{"key":"4925_CR2","unstructured":"Al Kadi M (2018) FGPU: a flexible soft GPU architecture for general purpose computing on FPGAs"},{"issue":"2","key":"4925_CR3","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3457886","volume":"14","author":"R Ma","year":"2021","unstructured":"Ma R et al (2021) Specializing FGPU for persistent deep learning. ACM Trans Reconfig Technol Syst (TRETS) 14(2):1\u201323","journal-title":"ACM Trans Reconfig Technol Syst (TRETS)"},{"key":"4925_CR4","doi-asserted-by":"crossref","unstructured":"De Matteis T, de Fine Licht J, Hoefler T (2020) FBLAS: streaming linear algebra on FPGA. In: SC20: International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE","DOI":"10.1109\/SC41405.2020.00063"},{"key":"4925_CR5","doi-asserted-by":"crossref","unstructured":"Zohouri HR et al (2016) Evaluating and optimizing OpenCL kernels for high performance computing with FPGAs. In: SC\u201916: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE","DOI":"10.1109\/SC.2016.34"},{"key":"4925_CR6","unstructured":"https:\/\/spec.oneApi.com\/versions\/latest\/index.html"},{"key":"4925_CR7","doi-asserted-by":"crossref","unstructured":"Arnone A (1994) Viscous analysis of three-dimensional rotor flow using a multigrid method 435\u2013445","DOI":"10.1115\/1.2929430"},{"issue":"5","key":"4925_CR8","doi-asserted-by":"publisher","first-page":"705","DOI":"10.1007\/s00193-018-0883-4","volume":"29","author":"R Pacciani","year":"2019","unstructured":"Pacciani R, Marconcini M, Arnone A (2019) Comparison of the AUSM+-up and other advection schemes for turbomachinery applications. Shock Waves 29(5):705\u2013716","journal-title":"Shock Waves"},{"key":"4925_CR9","doi-asserted-by":"crossref","unstructured":"Poli F, Marconcini M, Pacciani R, Magarielli D, Spano E, Arnone A (2022) Exploiting GPU-based HPC architectures to accelerate an unsteady CFD solver for turbomachinery applications. In: IGTI ASME Turbo Expo, June 13\u201317, Rotterdam, The Netherlands, ASME paper GT2022-82569","DOI":"10.1115\/GT2022-82569"},{"key":"4925_CR10","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1016\/j.compfluid.2018.06.005","volume":"173","author":"W Vanderbauwhede","year":"2018","unstructured":"Vanderbauwhede W, Davidson G (2018) Domain-specific acceleration and auto-parallelization of legacy scientific code in FORTRAN 77 using source-to-source compilation. Comput Fluids 173:1\u20135","journal-title":"Comput Fluids"},{"key":"4925_CR11","unstructured":"Vanderbauwhede W, Nabi SW (2018) Towards automatic transformation of legacy scientific code into OpenCL for optimal performance on FPGAs. arXiv:1901.00416"},{"key":"4925_CR12","doi-asserted-by":"crossref","unstructured":"Thomas DB et al (2015) Transparent linking of compiled software and synthesized hardware. In: 2015 Design, Automation and Test in Europe Conference and Exhibition (DATE). IEEE","DOI":"10.7873\/DATE.2015.1125"},{"key":"4925_CR13","doi-asserted-by":"crossref","unstructured":"Nabi SW, Vander bauwhede W (2019) Automatic pipelining and vectorization of scientific code for FPGAs. Int J Reconfig Comput 2019","DOI":"10.1155\/2019\/7348013"},{"issue":"20","key":"4925_CR14","doi-asserted-by":"publisher","first-page":"e6570","DOI":"10.1002\/cpe.6570","volume":"34","author":"T Nguyen","year":"2022","unstructured":"Nguyen T et al (2022) FPGA-based HPC accelerators: an evaluation on performance and energy efficiency. Concurr Comput Pract Exp 34(20):e6570","journal-title":"Concurr Comput Pract Exp"},{"issue":"2","key":"4925_CR15","doi-asserted-by":"publisher","first-page":"600","DOI":"10.1109\/TPDS.2015.2407896","volume":"27","author":"Fernando A Escobar","year":"2015","unstructured":"Escobar Fernando A, Chang Xin, Valderrama Carlos (2015) Suitability analysis of FPGAs for heterogeneous platforms in HPC. IEEE Trans Parallel Distrib Syst 27(2):600\u2013612","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"4925_CR16","doi-asserted-by":"crossref","unstructured":"Weller D et al (2017) Energy efficient scientific computing on FPGAs using OpenCL. In: Proceedings of the 2017 ACM\/SIGDA international symposium on field-programmable gate arrays","DOI":"10.1145\/3020078.3021730"},{"key":"4925_CR17","doi-asserted-by":"crossref","unstructured":"Pell O et al (2013) Maximum performance computing with dataflow engines. In: High-performance computing using FPGAs. Springer, New York, pp 747\u2013774","DOI":"10.1007\/978-1-4614-1791-0_25"},{"key":"4925_CR18","doi-asserted-by":"crossref","unstructured":"Czajkowski TS et al (2012) From OpenCL to high-performance hardware on FPGAs. In: 22nd International Conference on Field Programmable Logic and Applications (FPL). IEEE","DOI":"10.1109\/FPL.2012.6339272"},{"key":"4925_CR19","unstructured":"Segal O et al (2015) Sparkcl: a unified programming framework for accelerators on heterogeneous clusters. arXiv:1505.01120"},{"key":"4925_CR20","doi-asserted-by":"crossref","unstructured":"Agron J (2009) Domain-specific language for HW\/SW co-design for FPGAs. In: IFIP Working Conference on Domain-Specific Languages. Springer, Berlin, Heidelberg","DOI":"10.1007\/978-3-642-03034-5_13"},{"key":"4925_CR21","doi-asserted-by":"crossref","unstructured":"Kulkarni C, Brebner G, Schelle G (2004) Mapping a domain specific language to a platform FPGA. In: Proceedings of the 41st Annual Design Automation Conference","DOI":"10.1145\/996566.996811"},{"key":"4925_CR22","doi-asserted-by":"crossref","unstructured":"Bachrach J et al (2012) Chisel: constructing hardware in a scala embedded language. In: DAC Design Automation Conference 2012. IEEE","DOI":"10.1145\/2228360.2228584"},{"issue":"3","key":"4925_CR23","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1016\/j.parco.2003.12.002","volume":"30","author":"M Cole","year":"2004","unstructured":"Cole M (2004) Bringing skeletons out of the closet: a pragmatic manifesto for skeletal parallel programming. Parallel Comput 30(3):389\u2013406","journal-title":"Parallel Comput"},{"key":"4925_CR24","doi-asserted-by":"crossref","unstructured":"Weinhardt M, Wayne L (1999) Memory access optimization and RAM inference for pipeline vectorization. In: International Workshop on Field Programmable Logic and Applications. Springer, Berlin, Heidelberg","DOI":"10.1007\/978-3-540-48302-1_7"},{"key":"4925_CR25","doi-asserted-by":"crossref","unstructured":"Liao C et al (2010) A ROSE-based OpenMP 3.0 research compiler supporting multiple runtime libraries. International workshop on OpenMP. Springer, Berlin, Heidelberg","DOI":"10.1007\/978-3-642-13217-9_2"},{"key":"4925_CR26","doi-asserted-by":"crossref","unstructured":"Orchard D, Rice A (2013) Upgrading fortran source code using automatic refactoring. In: Proceedings of the 2013 ACM workshop on Workshop on refactoring tools","DOI":"10.1145\/2541348.2541356"},{"key":"4925_CR27","doi-asserted-by":"crossref","unstructured":"Overbey J et al (2005) Refactorings for Fortran and high-performance computing. In: Proceedings of the second international workshop on Software engineering for high performance computing system applications","DOI":"10.1145\/1145319.1145331"},{"key":"4925_CR28","doi-asserted-by":"crossref","unstructured":"Mayer F et al (2022) The ORKA-HPC compiler-practical OpenMP for FPGAs. In: International workshop on languages and compilers for parallel computing. Springer, Cham","DOI":"10.1007\/978-3-030-99372-6_6"},{"key":"4925_CR29","doi-asserted-by":"crossref","unstructured":"Fatica M, Ruetsch G (2014) CUDA Fortran for scientists and engineers","DOI":"10.1016\/B978-0-12-415992-1.00017-1"},{"key":"4925_CR30","unstructured":"https:\/\/developer.nvidia.com\/cuda-fortran"},{"key":"4925_CR31","unstructured":"https:\/\/github.com\/wimvanderbauwhede\/OpenCLIntegration"},{"key":"4925_CR32","doi-asserted-by":"crossref","unstructured":"Kashino R et al (2022) Multi-hetero acceleration by GPU and FPGA for astrophysics simulation on oneAPI environment. In: International Conference on High Performance Computing in Asia-Pacific Region","DOI":"10.1145\/3492805.3492817"},{"key":"4925_CR33","doi-asserted-by":"crossref","unstructured":"Uguen Y, Petit E (2018) PyGA: a Python to FPGA compiler prototype. In: Proceedings of the 5th ACM SIGPLAN international workshop on artificial intelligence and empirical methods for software engineering and parallel computing systems","DOI":"10.1145\/3281070.3281072"},{"key":"4925_CR34","doi-asserted-by":"crossref","unstructured":"Brown N (2021) Porting incompressible flow matrix assembly to FPGAs for accelerating HPC engineering simulations. In: 2021 IEEE\/ACM international workshop on heterogeneous high-performance reconfigurable computing (H2RC). IEEE","DOI":"10.1109\/H2RC54759.2021.00007"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-04925-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-022-04925-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-04925-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,23]],"date-time":"2023-03-23T09:15:33Z","timestamp":1679562933000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-022-04925-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,11,29]]},"references-count":34,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2023,5]]}},"alternative-id":["4925"],"URL":"https:\/\/doi.org\/10.1007\/s11227-022-04925-2","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"type":"print","value":"0920-8542"},{"type":"electronic","value":"1573-0484"}],"subject":[],"published":{"date-parts":[[2022,11,29]]},"assertion":[{"value":"1 November 2022","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 November 2022","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}