{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,29]],"date-time":"2026-03-29T15:15:29Z","timestamp":1774797329346,"version":"3.50.1"},"reference-count":20,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2011,3,29]],"date-time":"2011-03-29T00:00:00Z","timestamp":1301356800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGMETRICS Perform. Eval. Rev."],"published-print":{"date-parts":[[2011,3,29]]},"abstract":"<jats:p>We present the performance analysis of a port of the LU benchmark from the NAS Parallel Benchmark (NPB) suite to NVIDIA's Compute Unified Device Architecture (CUDA), and report on the optimisation efforts employed to take advantage of this platform. Execution times are reported for several different GPUs, ranging from low-end consumergrade products to high-end HPC-grade devices, including the Tesla C2050 built on NVIDIA's Fermi processor.<\/jats:p>\n          <jats:p>We also utilise recently developed performance models of LU to facilitate a comparison between future large-scale distributed clusters of GPU devices and existing clusters built on traditional CPU architectures, including a quad-socket, quad-core AMD Opteron cluster and an IBM BlueGene\/P.<\/jats:p>","DOI":"10.1145\/1964218.1964223","type":"journal-article","created":{"date-parts":[[2011,4,1]],"date-time":"2011-04-01T15:54:25Z","timestamp":1301673265000},"page":"23-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":34,"title":["Performance analysis of a hybrid MPI\/CUDA implementation of the NASLU benchmark"],"prefix":"10.1145","volume":"38","author":[{"given":"S. J.","family":"Pennycook","sequence":"first","affiliation":[{"name":"University of Warwick, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"S. D.","family":"Hammond","sequence":"additional","affiliation":[{"name":"University of Warwick, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"S. A.","family":"Jarvis","sequence":"additional","affiliation":[{"name":"University of Warwick, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"G. R.","family":"Mudalige","sequence":"additional","affiliation":[{"name":"University of Oxford, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2011,3,29]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"http:\/\/www.llnl.gov\/asci_benchmarks\/asci\/limited\/sweep3d\/asci_sweep3d.html","author":"Benchmark ASCI","year":"1995","unstructured":"The ASCI Sweep3D Benchmark . http:\/\/www.llnl.gov\/asci_benchmarks\/asci\/limited\/sweep3d\/asci_sweep3d.html , 1995 . The ASCI Sweep3D Benchmark. http:\/\/www.llnl.gov\/asci_benchmarks\/asci\/limited\/sweep3d\/asci_sweep3d.html, 1995."},{"key":"e_1_2_1_2_1","unstructured":"The Green 500 List : Environmentally Responsible Supercomputing. http:\/\/www.green500.org November 2010.  The Green 500 List : Environmentally Responsible Supercomputing. http:\/\/www.green500.org November 2010."},{"key":"e_1_2_1_3_1","volume-title":"November","author":"Supercomputer Sites 0","year":"2010","unstructured":"Top 50 0 Supercomputer Sites . http:\/\/www.top500.org , November 2010 . Top 500 Supercomputer Sites. http:\/\/www.top500.org, November 2010."},{"key":"e_1_2_1_4_1","volume-title":"Accelerating Data-Serial Applications on GPGPUs: A Systems Approach","author":"Aji A. M.","year":"2008","unstructured":"A. M. Aji and W. C. Feng . Accelerating Data-Serial Applications on GPGPUs: A Systems Approach . Technical Report TR-08-24, Computer Science , Virginia Tech ., 2008 . A. M. Aji and W. C. Feng. Accelerating Data-Serial Applications on GPGPUs: A Systems Approach. Technical Report TR-08-24, Computer Science, Virginia Tech., 2008."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5160984"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-13119-6_36"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/1413370.1413373"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.4108\/ICST.SIMUTOOLS2009.5753"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/850941.852910"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.2514\/6.2010-522"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/360827.360844"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/1815961.1816021"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-9-S2-S10"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2008.4536243"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/BIBE.2008.4696721"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2007.370252"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/648135.748628"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/1345206.1345220"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.9"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the USENIX Workshop on Hot Topics in Parallelism","author":"Vuduc R.","year":"2010","unstructured":"R. Vuduc , A. Chandramowlishwaran , J. Choi , M. E. Guney , and A. Shringarpure . On the Limits of GPU Acceleration . In Proceedings of the USENIX Workshop on Hot Topics in Parallelism , June 2010 . R. Vuduc, A. Chandramowlishwaran, J. Choi, M. E. Guney, and A. Shringarpure. On the Limits of GPU Acceleration. In Proceedings of the USENIX Workshop on Hot Topics in Parallelism, June 2010."}],"container-title":["ACM SIGMETRICS Performance Evaluation Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1964218.1964223","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/1964218.1964223","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:26:47Z","timestamp":1750278407000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/1964218.1964223"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,3,29]]},"references-count":20,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2011,3,29]]}},"alternative-id":["10.1145\/1964218.1964223"],"URL":"https:\/\/doi.org\/10.1145\/1964218.1964223","relation":{},"ISSN":["0163-5999"],"issn-type":[{"value":"0163-5999","type":"print"}],"subject":[],"published":{"date-parts":[[2011,3,29]]},"assertion":[{"value":"2011-03-29","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}