{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,7]],"date-time":"2026-05-07T15:45:35Z","timestamp":1778168735815,"version":"3.51.4"},"reference-count":12,"publisher":"SAGE Publications","issue":"1","license":[{"start":{"date-parts":[[2015,8,17]],"date-time":"2015-08-17T00:00:00Z","timestamp":1439769600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2016,2]]},"abstract":"<jats:p>Graphics processing unit accelerated supercomputers have proved to be very effective, especially with regard to power efficiency, for accelerating compute intensive applications like the high-performance Linpack used in the TOP500 list. This paper presents the details of a CUDA implementation of the high-performance conjugate gradient, a new proposed benchmark that better represents modern application workloads which rely more heavily on memory system and network performance than high-performance Linpack. The results obtained at full scale on the largest graphics processing unit supercomputers in the world, Titan, the Cray XK7 at ORNL and Piz-Daint, the Cray XC30 at CSCS, indicate that graphics processing unit accelerated supercomputers are also very effective for this type of workload. A comparison with other architectures is also presented, showing that graphics processing units, with their high memory bandwidth, are the highest performing devices for this new benchmark.<\/jats:p>","DOI":"10.1177\/1094342015599239","type":"journal-article","created":{"date-parts":[[2015,8,17]],"date-time":"2015-08-17T20:48:41Z","timestamp":1439844521000},"page":"28-38","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":6,"title":["Performance analysis of the high-performance conjugate gradient benchmark on GPUs"],"prefix":"10.1177","volume":"30","author":[{"given":"Everett","family":"Phillips","sequence":"first","affiliation":[{"name":"NVIDIA Corporation, Santa Clara, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Massimiliano","family":"Fatica","sequence":"additional","affiliation":[{"name":"NVIDIA Corporation, Santa Clara, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2015,8,17]]},"reference":[{"key":"bibr1-1094342015599239","doi-asserted-by":"publisher","DOI":"10.1145\/2148600.2148602"},{"key":"bibr2-1094342015599239","doi-asserted-by":"publisher","DOI":"10.1137\/1.9780898719505"},{"key":"bibr3-1094342015599239","first-page":"1","volume-title":"GPU Technology Conference","author":"Cohen J","year":"2012"},{"key":"bibr4-1094342015599239","doi-asserted-by":"crossref","unstructured":"Dongarra J, Heroux MA (2013) Toward a new metric for ranking high-performance computing systems. Sandia Report SAND2013-4744, USA.","DOI":"10.2172\/1089988"},{"key":"bibr5-1094342015599239","unstructured":"Dongarra J, Luszczek P (2005) Introduction to the HPC challenge benchmark suite. ICL Technical Report ICL-UT-05-01 (also appears as CS Department Technical Report UT-CS-05-544)."},{"key":"bibr6-1094342015599239","volume-title":"Matrix Computations, 3rd Edition","author":"Golub GH","year":"1996"},{"key":"bibr7-1094342015599239","doi-asserted-by":"crossref","unstructured":"Heroux MA, Dongarra J, Luszczek P (2013) HPCG technical specification. Sandia Report SAND2013-8752.","DOI":"10.2172\/1113870"},{"key":"bibr8-1094342015599239","doi-asserted-by":"publisher","DOI":"10.1137\/0914041"},{"key":"bibr9-1094342015599239","doi-asserted-by":"publisher","DOI":"10.1137\/0215074"},{"key":"bibr10-1094342015599239","unstructured":"McCalpin JD (1995) Memory bandwidth and machine balance in current high-performance computers. IEEE Computer Society Technical Committee on Computer Architecture (TCCA) Newsletter, December 1995."},{"key":"bibr11-1094342015599239","volume-title":"ASCR HPCG workshop","author":"Park J","year":"2014"},{"key":"bibr12-1094342015599239","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2010.5470394"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342015599239","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/full-xml\/10.1177\/1094342015599239","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342015599239","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:19:34Z","timestamp":1777450774000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342015599239"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,8,17]]},"references-count":12,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2016,2]]}},"alternative-id":["10.1177\/1094342015599239"],"URL":"https:\/\/doi.org\/10.1177\/1094342015599239","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2015,8,17]]}}}