{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T22:46:18Z","timestamp":1777675578446,"version":"3.51.4"},"reference-count":15,"publisher":"SAGE Publications","issue":"2","license":[{"start":{"date-parts":[[2017,11,7]],"date-time":"2017-11-07T00:00:00Z","timestamp":1510012800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2020,3]]},"abstract":"<jats:p>The ever-increasing computational requirements of HPC and service provider applications are becoming a great challenge for hardware and software designers. These requirements are reaching levels where the isolated development on either computational field is not enough to deal with such challenge. A holistic view of the computational thinking is therefore the only way to success in real scenarios. However, this is not a trivial task as it requires, among others, of hardware\u2013software codesign. In the hardware side, most high-throughput computers are designed aiming for heterogeneity, where accelerators (e.g. Graphics Processing Units (GPUs), Field-Programmable Gate Arrays (FPGAs), etc.) are connected through high-bandwidth bus, such as PCI-Express, to the host CPUs. Applications, either via programmers, compilers, or runtime, should orchestrate data movement, synchronization, and so on among devices with different compute and memory capabilities. This increases the programming complexity and it may reduce the overall application performance. This article evaluates different offloading strategies to leverage heterogeneous systems, based on several cards with the first-generation Xeon Phi coprocessors (Knights Corner). We use a 11-point 3-D Stencil kernel that models heat dissipation as a case study. Our results reveal substantial performance improvements when using several accelerator cards. Additionally, we show that computing of an approximate result by reducing the communication overhead can yield 23% performance gains for double-precision data sets.<\/jats:p>","DOI":"10.1177\/1094342017738352","type":"journal-article","created":{"date-parts":[[2017,12,31]],"date-time":"2017-12-31T10:33:09Z","timestamp":1514716389000},"page":"199-207","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":0,"title":["Offloading strategies for Stencil kernels on the KNC Xeon Phi architecture: Accuracy versus performance"],"prefix":"10.1177","volume":"34","author":[{"given":"Mario","family":"Hern\u00e1ndez","sequence":"first","affiliation":[{"name":"Facultad de Ingenier\u00eda, Universidad Aut\u00f3noma de Guerrero, Guerrero, Mexico"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Juan M","family":"Cebri\u00e1n","sequence":"additional","affiliation":[{"name":"Barcelona Supercomputing Center (BSC), Barcelona, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jos\u00e9 M","family":"Cecilia","sequence":"additional","affiliation":[{"name":"Departamento de Ingenier\u00eda Inform\u00e1tica, Universidad Cat\u00f3lica San Antonio de Murcia, Murcia, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jos\u00e9 M","family":"Garc\u00eda","sequence":"additional","affiliation":[{"name":"Departamento de Ingenier\u00eda y Tecnolog\u00eda de Computadores, Universidad de Murcia, Murcia, Spain"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2017,11,7]]},"reference":[{"key":"bibr1-1094342017738352","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2015.05.004"},{"key":"bibr2-1094342017738352","author":"Chrysos G","year":"2014","journal-title":"Intel\u00ae Xeon Phi\u2122 Coprocessor Architecture"},{"key":"bibr3-1094342017738352","doi-asserted-by":"publisher","DOI":"10.1145\/2324876.2324879"},{"key":"bibr4-1094342017738352","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-802118-7.00020-0"},{"key":"bibr5-1094342017738352","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-410414-3.00007-4"},{"key":"bibr6-1094342017738352","volume-title":"Intel Xeon Phi Coprocessor High-Performance Programming","author":"Jeffers J","year":"2013"},{"key":"bibr7-1094342017738352","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4302-5927-5"},{"key":"bibr8-1094342017738352","unstructured":"Reinders J, Jeffers J (2014) High Performance Parallelism Pearls, Multicore and Many-core Programming Approaches (Characterization and Auto-tuning of 3DFD). Morgan Kaufmann, pp. 377\u2013396."},{"issue":"1","key":"bibr9-1094342017738352","volume-title":"High Performance Parallelism Pearls: Multicore and Many-core Programming Approaches","author":"Reinders J","year":"2015"},{"issue":"2","key":"bibr10-1094342017738352","volume-title":"High Performance Parallelism Pearls: Multicore and Many-core Programming Approaches","author":"Reinders J","year":"2015"},{"key":"bibr11-1094342017738352","doi-asserted-by":"publisher","DOI":"10.1109\/HPEC.2015.7322456"},{"key":"bibr12-1094342017738352","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-41778-3_20"},{"key":"bibr13-1094342017738352","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-06486-4"},{"key":"bibr14-1094342017738352","volume-title":"Heterogeneous System Architecture: A New Compute Platform Infrastructure","author":"Wen-mei WH","year":"2015"},{"key":"bibr15-1094342017738352","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-802118-7.00012-1"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342017738352","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342017738352","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342017738352","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:15:54Z","timestamp":1777450554000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342017738352"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,11,7]]},"references-count":15,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2020,3]]}},"alternative-id":["10.1177\/1094342017738352"],"URL":"https:\/\/doi.org\/10.1177\/1094342017738352","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,11,7]]}}}