{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T22:43:42Z","timestamp":1777675422015,"version":"3.51.4"},"reference-count":35,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2017,8,4]],"date-time":"2017-08-04T00:00:00Z","timestamp":1501804800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2018,1]]},"abstract":"<jats:p>The portability of real high-performance computing (HPC) applications on new platforms is an open and very delicate problem. Especially, the performance portability of the underlying computing kernels is problematic as they need to be tuned for each and every platform the application encounters. This article presents BOAST, a metaprogramming framework dedicated to computing kernels. BOAST allows the description of a kernel and its possible optimizations using a domain-specific language. BOAST runtime will then compare the different versions\u2019performance as well as verify their exactness. BOAST is applied to three use cases: a Laplace kernel in OpenCL and two HPC applications BigDFT (electronic density computation) and SPECFEM3D (seismic and wave propagation).<\/jats:p>","DOI":"10.1177\/1094342017718068","type":"journal-article","created":{"date-parts":[[2017,8,4]],"date-time":"2017-08-04T05:01:51Z","timestamp":1501822911000},"page":"28-44","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":8,"title":["BOAST"],"prefix":"10.1177","volume":"32","author":[{"given":"Brice","family":"Videau","sequence":"first","affiliation":[{"name":"Laboratoire d\u2019Informatique de Grenoble, Saint-Martin-d'H\u00e8res, France"},{"name":"Centre National de la Recherche Scientifique, Paris, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kevin","family":"Pouget","sequence":"additional","affiliation":[{"name":"Laboratoire d\u2019Informatique de Grenoble, Saint-Martin-d'H\u00e8res, France"},{"name":"Universit\u00e9 Grenoble Alpes, Grenoble, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Luigi","family":"Genovese","sequence":"additional","affiliation":[{"name":"Universit\u00e9 Grenoble Alpes, Grenoble, France"},{"name":"CEA\/INAC\/MEM\/L_Sim, Grenoble, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Thierry","family":"Deutsch","sequence":"additional","affiliation":[{"name":"Universit\u00e9 Grenoble Alpes, Grenoble, France"},{"name":"CEA\/INAC\/MEM\/L_Sim, Grenoble, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dimitri","family":"Komatitsch","sequence":"additional","affiliation":[{"name":"Aix Marseille University, Marseille, France"},{"name":"CNRS, Centrale Marseille, Marseille, France"},{"name":"LMA, Marseille, France."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fr\u00e9d\u00e9ric","family":"Desprez","sequence":"additional","affiliation":[{"name":"Laboratoire d\u2019Informatique de Grenoble, Saint-Martin-d'H\u00e8res, France"},{"name":"Inria"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jean-Fran\u00e7ois","family":"M\u00e9haut","sequence":"additional","affiliation":[{"name":"Laboratoire d\u2019Informatique de Grenoble, Saint-Martin-d'H\u00e8res, France"},{"name":"Universit\u00e9 Grenoble Alpes, Grenoble, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2017,8,4]]},"reference":[{"key":"bibr1-1094342017718068","unstructured":"Adeniyi-Jones C. Optimal Compute on ARM Mali GPUs. Available at: http:\/\/www.cs.bris.ac.uk\/home\/simonm\/montblanc\/OpenCL on Mali.pdf (accessed 1 December 2016)."},{"key":"bibr2-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1145\/502874.502897"},{"key":"bibr3-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898719604"},{"key":"bibr4-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3516"},{"key":"bibr5-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1109\/MCSoC.2013.12"},{"key":"bibr6-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1109\/99.660313"},{"key":"bibr7-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.3097"},{"key":"bibr8-1094342017718068","volume-title":"Proceedings of the 6th LACSI Symposium","author":"Djoudi L","year":"2005"},{"key":"bibr9-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1145\/77626.79170"},{"key":"bibr10-1094342017718068","doi-asserted-by":"crossref","unstructured":"Fursin G, Miceli R, Lokhmotov A, (2014) Collective Mind: towards practical and collaborative auto-tuning. Scientific Programming 22(4): 309\u2013329. Available at: https:\/\/hal.inria.fr\/hal-01054763.","DOI":"10.1155\/2014\/797348"},{"key":"bibr11-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1063\/1.2949547"},{"key":"bibr12-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1063\/1.3166140"},{"key":"bibr13-1094342017718068","volume-title":"Compte-Rendu de l\u2019Acad\u00e9mie des Sciences, Calcul Intensif","author":"Genovese L","year":"2010"},{"key":"bibr14-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1016\/0010-4655(93)90057-J"},{"key":"bibr15-1094342017718068","unstructured":"Hartono A, Norris B, Sadayappan P (2009) Annotation-based empirical performance tuning using orio. In: Proceedings of the 23rd IEEE International Parallel & Distributed Processing Symposium, Rome, Italy, Preprint ANL\/MCS-P1556-1008. Available at: http:\/\/www.mcs.anl.gov\/uploads\/cels\/papers\/P1556.pdf"},{"issue":"4","key":"bibr16-1094342017718068","first-page":"1","volume":"28","author":"Hudak P","year":"1996","journal-title":"ACM Computing Surveys"},{"key":"bibr17-1094342017718068","doi-asserted-by":"crossref","unstructured":"Khronos OpenCL consortium. OpenCL: Open Computing Language. Available at: http:\/\/www.khronos.org\/opencl\/ (accessed 1 December 2016).","DOI":"10.1145\/2791321.2791337"},{"key":"bibr18-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1016\/j.crme.2010.11.007"},{"key":"bibr19-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1145\/355841.355847"},{"key":"bibr20-1094342017718068","volume-title":"Ruby Programming Language","author":"Matsumoto Y","year":"2002"},{"key":"bibr21-1094342017718068","unstructured":"MPI (2012) The Message Passing Interface (MPI) Standard. Available at: http:\/\/www.mcs.anl.gov\/research\/projects\/mpi\/ (accessed 1 December 2016)."},{"key":"bibr22-1094342017718068","first-page":"7","volume-title":"Proceeding Dept of Defense HPCMP Users Group Conference","author":"Mucci P","year":"1999"},{"key":"bibr23-1094342017718068","unstructured":"NVIDIA (2011) NVIDIA Compute Unified Device Architecture. Available at: http:\/\/www.nvidia.com\/object\/cudahomenew.html (accessed 1 December 2016)."},{"key":"bibr24-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1177\/1094342004041291"},{"key":"bibr25-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1145\/2499370.2462176"},{"key":"bibr26-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1142\/S0129626409000134"},{"key":"bibr27-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1145\/2581122.2544155"},{"key":"bibr28-1094342017718068","unstructured":"Top500.Org. Top500. Available at: http:\/\/www.top500.org (accessed 1 December 2016)."},{"key":"bibr29-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-40047-6_82"},{"key":"bibr30-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2013.07.013"},{"key":"bibr31-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1002\/spe.626"},{"key":"bibr32-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1109\/71.97902"},{"key":"bibr33-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1007\/BF01407876"},{"key":"bibr34-1094342017718068","volume-title":"Symposium on Application Accelerators in High-Performance Computing","author":"Ye D","year":"2012"},{"key":"bibr35-1094342017718068","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2007.370637"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342017718068","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342017718068","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342017718068","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:15:30Z","timestamp":1777450530000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342017718068"}},"subtitle":["A metaprogramming framework to produce portable and efficient computing kernels for HPC applications"],"short-title":[],"issued":{"date-parts":[[2017,8,4]]},"references-count":35,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2018,1]]}},"alternative-id":["10.1177\/1094342017718068"],"URL":"https:\/\/doi.org\/10.1177\/1094342017718068","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,8,4]]}}}