{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,1]],"date-time":"2025-10-01T18:15:22Z","timestamp":1759342522453,"version":"3.41.0"},"reference-count":11,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2011,12,19]],"date-time":"2011-12-19T00:00:00Z","timestamp":1324252800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGARCH Comput. Archit. News"],"published-print":{"date-parts":[[2011,12,19]]},"abstract":"<jats:p>Graphics Processing Units (GPUs) are widely used to accelerate scientific applications. Many successes have been reported with speedups of two or three orders of magnitude over serial implementations of the same algorithms. These speedups typically pertain to a specific implementation with fixed parameters mapped to a specific hardware implementation. The implementations are not designed to be easily ported to other GPUs, even from the same manufacturer. When target hardware changes, the application must be re-optimized.<\/jats:p>\n          <jats:p>\n            In this paper we address a different problem. We aim to deliver working, efficient GPU code in a library that is downloaded and run by many different users. The issue is to deliver efficiency independent of the individual user parameters and without a\n            <jats:italic>priori<\/jats:italic>\n            knowledge of the hardware the user will employ. This problem requires a different set of tradeoffs than finding the best runtime for a single solution. Solutions must be adaptable to a range of different parameters both to solve users' problems and to make the best use of the target hardware.\n          <\/jats:p>\n          <jats:p>Another issue is the integration of GPUs into a Problem Solving Environment (PSE) where the use of a GPU is almost invisible from the perspective of the user. Ease of use and smooth interactions with the existing user interface are important to our approach. We illustrate our solution with the incorporation of GPU processing into the Scientific Computing Institute (SCI)Run Biomedical PSE developed at the University of Utah. SCIRun allows scientists to interactively construct many different types of biomedical simulations. We use this environment to demonstrate the effectiveness of the GPU by accelerating time consuming algorithms in the scientist's simulations. Specifically we target the linear solver module, including Conjugate Gradient, Jacobi and MinRes solvers for sparse matrices.<\/jats:p>","DOI":"10.1145\/2082156.2082158","type":"journal-article","created":{"date-parts":[[2011,12,27]],"date-time":"2011-12-27T15:22:22Z","timestamp":1324999342000},"page":"2-7","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["The challenges of writing portable, correct and high performance libraries for GPUs"],"prefix":"10.1145","volume":"39","author":[{"given":"Miriam","family":"Leeser","sequence":"first","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Devon","family":"Yablonski","sequence":"additional","affiliation":[{"name":"Mercury Computer Systems, Chelmsford, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dana","family":"Brooks","sequence":"additional","affiliation":[{"name":"Northeastern University, Boston, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Laurie Smith","family":"King","sequence":"additional","affiliation":[{"name":"College of the Holy Cross, Worcester, MA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2011,12,19]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1080\/17445760802337010"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2049662.2049663"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-13374-9_4"},{"key":"e_1_2_1_4_1","unstructured":"Khronos. The OpenCL Specification. http:\/\/www.khronos.org\/registry\/cl\/specs\/opencl-1.1.pdf June 2010.  Khronos. The OpenCL Specification. http:\/\/www.khronos.org\/registry\/cl\/specs\/opencl-1.1.pdf June 2010."},{"key":"e_1_2_1_5_1","unstructured":"A. Kl\u00f6ckner N. Pinto Y. Lee etal PyCUDA and PyOpenCL: A Scripting-Based Approach to GPU Run-Time Code Generation. http:\/\/arxiv.org\/abs\/0911.3456v2 March 2011.  A. Kl\u00f6ckner N. Pinto Y. Lee et al. PyCUDA and PyOpenCL: A Scripting-Based Approach to GPU Run-Time Code Generation. http:\/\/arxiv.org\/abs\/0911.3456v2 March 2011."},{"key":"e_1_2_1_6_1","unstructured":"Mathworks. MATLAB GPU Computing with NVIDIA CUDA-Enabled GPUs. http:\/\/www.mathworks.com\/discovery\/matlab-gpu.html 2011.  Mathworks. MATLAB GPU Computing with NVIDIA CUDA-Enabled GPUs. http:\/\/www.mathworks.com\/discovery\/matlab-gpu.html 2011."},{"key":"e_1_2_1_7_1","unstructured":"NVIDIA. CUDA 4.0 Math Libraries Performance Boost. http:\/\/developer.nvidia.com\/content\/cuda-40-math-libraries-performance-boost 2011.  NVIDIA. CUDA 4.0 Math Libraries Performance Boost. http:\/\/developer.nvidia.com\/content\/cuda-40-math-libraries-performance-boost 2011."},{"key":"e_1_2_1_8_1","unstructured":"http:\/\/www.scirun.org. SCIRun: A Scientific Computing Problem Solving Environment Scientific Computing and Imaging Institute (SCI).  http:\/\/www.scirun.org. SCIRun: A Scientific Computing Problem Solving Environment Scientific Computing and Imaging Institute (SCI)."},{"key":"e_1_2_1_9_1","unstructured":"The Portland Group. PGI Accelerator Compilers. http:\/\/www.pgroup.com\/resources\/accel.htm 2010.  The Portland Group. PGI Accelerator Compilers. http:\/\/www.pgroup.com\/resources\/accel.htm 2010."},{"key":"e_1_2_1_10_1","unstructured":"N. Whitehead and A. Fit-Florea. Precision & Performance: Floating Point and IEEE 754 Compliance for NVIDIA GPUs. http:\/\/developer.download.nvidia.com\/assets\/cuda\/files\/NVIDIA-CUDA-Floating-Point.pdf 2011.  N. Whitehead and A. Fit-Florea. Precision & Performance: Floating Point and IEEE 754 Compliance for NVIDIA GPUs. http:\/\/developer.download.nvidia.com\/assets\/cuda\/files\/NVIDIA-CUDA-Floating-Point.pdf 2011."},{"key":"e_1_2_1_11_1","unstructured":"D. Yablonski. Numerical Accuracy Differences in CPU and GPGPU Codes. Master's thesis Northeastern University Boston MA 2011. http:\/\/www.coe.neu.edu\/Research\/rcl\/publications.php#theses.  D. Yablonski. Numerical Accuracy Differences in CPU and GPGPU Codes. Master's thesis Northeastern University Boston MA 2011. http:\/\/www.coe.neu.edu\/Research\/rcl\/publications.php#theses."}],"container-title":["ACM SIGARCH Computer Architecture News"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2082156.2082158","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2082156.2082158","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T10:06:41Z","timestamp":1750241201000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2082156.2082158"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,12,19]]},"references-count":11,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2011,12,19]]}},"alternative-id":["10.1145\/2082156.2082158"],"URL":"https:\/\/doi.org\/10.1145\/2082156.2082158","relation":{},"ISSN":["0163-5964"],"issn-type":[{"type":"print","value":"0163-5964"}],"subject":[],"published":{"date-parts":[[2011,12,19]]},"assertion":[{"value":"2011-12-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}