{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,2]],"date-time":"2025-11-02T16:34:11Z","timestamp":1762101251105,"version":"3.41.0"},"reference-count":16,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2011,12,19]],"date-time":"2011-12-19T00:00:00Z","timestamp":1324252800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGARCH Comput. Archit. News"],"published-print":{"date-parts":[[2011,12,19]]},"abstract":"<jats:p>After more than five years since GPUs were first used as accelerators for general scientific computations, the field of General Purpose GPU computing or GPGPU has finally reached mainstream. Developers have now access to a mature hardware and software ecosystem. On the software side, several major open-source packages now support GPU acceleration while on the hardware side cloud-based solutions provide a simple way to access powerful machines with the latest GPUs at low cost. In this context, we look at the GPU acceleration of CAE, with a focus on the matrix solvers. We compare the performance that can be achieved using the open-source solver package PETSc ran on GPU-enabled Amazon EC2 hardware with that of an optimized legacy FEM code ran on a last generation 12-core blade server. Our results show that, although good performance can be achieved, some development is still needed to achieve peak performance.<\/jats:p>","DOI":"10.1145\/2082156.2082161","type":"journal-article","created":{"date-parts":[[2011,12,27]],"date-time":"2011-12-27T15:22:22Z","timestamp":1324999342000},"page":"14-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["GPU accelerated CAE using open solvers and the cloud"],"prefix":"10.1145","volume":"39","author":[{"given":"Serban","family":"Georgescu","sequence":"first","affiliation":[{"name":"Fujitsu Laboratories of Europe Limited, Middlesex, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peter","family":"Chow","sequence":"additional","affiliation":[{"name":"Fujitsu Laboratories of Europe Limited, Middlesex, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2011,12,19]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/MM.2008.57"},{"key":"e_1_2_1_2_1","unstructured":"NVIDIA. NVIDIA Fermi compute architecture whitepaper 2009.  NVIDIA. NVIDIA Fermi compute architecture whitepaper 2009."},{"key":"e_1_2_1_3_1","unstructured":"NVIDIA. CUDA Reference Manual 3.2 edition 2010.  NVIDIA. CUDA Reference Manual 3.2 edition 2010."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1201775.882364"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/882262.882363"},{"volume-title":"18th Symposium Simulationstechnique, Frontiers in Simulation, pages 139--144. SCS Publishing House e.V., 2005. ASIM 2005.","author":"G\u00f6ddeke D.","key":"e_1_2_1_6_1"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/2401945.2401989"},{"volume-title":"15th Annual Conference of the CFD Society of Canada","year":"2007","author":"Menon S.","key":"e_1_2_1_8_1"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-01970-8_90"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-01970-8_87"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1002\/fld.2462"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-010-0112-6"},{"volume-title":"University of California","year":"2004","author":"Vuduc R.","key":"e_1_2_1_13_1"},{"volume-title":"Proc. ACM\/IEEE Conf. Supercomputing (SC09)","year":"2009","author":"Bell N.","key":"e_1_2_1_14_1"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1693453.1693471"},{"volume-title":"Baifukan","year":"2008","author":"Okuda H.","key":"e_1_2_1_16_1"}],"container-title":["ACM SIGARCH Computer Architecture News"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2082156.2082161","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2082156.2082161","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T10:06:41Z","timestamp":1750241201000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2082156.2082161"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,12,19]]},"references-count":16,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2011,12,19]]}},"alternative-id":["10.1145\/2082156.2082161"],"URL":"https:\/\/doi.org\/10.1145\/2082156.2082161","relation":{},"ISSN":["0163-5964"],"issn-type":[{"type":"print","value":"0163-5964"}],"subject":[],"published":{"date-parts":[[2011,12,19]]},"assertion":[{"value":"2011-12-19","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}