{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T04:51:09Z","timestamp":1750308669811,"version":"3.41.0"},"reference-count":10,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2012,10,8]],"date-time":"2012-10-08T00:00:00Z","timestamp":1349654400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["SIGMETRICS Perform. Eval. Rev."],"published-print":{"date-parts":[[2012,10,8]]},"abstract":"<jats:p>Scalable systems employing a mix of GPUs with CPUs are becoming increasingly prevalent in high-performance computing. The presence of such accelerators introduces significant challenges and complexities to both language developers and end users. This paper provides a close study of efficient coordination mechanisms to handle parallel requests from multiple hosts of control to a GPU under hybrid programming. Using a set of microbenchmarks and applications on a GPU cluster, we show that thread and process-based context hosting have different tradeoffs. Experimental results on application benchmarks suggest that both thread-based context funneling and process-based context switching natively perform similarly on the latest Fermi GPUs, while manually guided context funneling is currently the best way to achieve optimal performance.<\/jats:p>","DOI":"10.1145\/2381056.2381081","type":"journal-article","created":{"date-parts":[[2012,10,11]],"date-time":"2012-10-11T14:55:16Z","timestamp":1349967316000},"page":"119-124","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Towards efficient GPU sharing on multicore processors"],"prefix":"10.1145","volume":"40","author":[{"given":"Lingyuan","family":"Wang","sequence":"first","affiliation":[{"name":"George Washington University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Miaoqing","family":"Huang","sequence":"additional","affiliation":[{"name":"University of Arkansas"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tarek","family":"El-Ghazawi","sequence":"additional","affiliation":[{"name":"George Washington University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,10,8]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"NAS Parallel Benchmarks. http:\/\/www.nas.nasa.gov\/Resources\/Software\/npb.html.  NAS Parallel Benchmarks. http:\/\/www.nas.nasa.gov\/Resources\/Software\/npb.html."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2020373.2020376"},{"key":"e_1_2_1_3_1","first-page":"151","volume-title":"Proc. 23rd International Conference on Languages and Compilers for Parallel Computing (LCPC'10)","author":"Chen L.","year":"2010","unstructured":"Chen , L. , Liu , L. , Tang , S. , Huang , L. , Jing , Z. , Xu , S. , Zhang , D. , and Shou , B . Unified parallel C for GPU clusters: language extensions and compiler implementation . In Proc. 23rd International Conference on Languages and Compilers for Parallel Computing (LCPC'10) ( Oct. 2010 ), pp. 151 -- 165 . Chen, L., Liu, L., Tang, S., Huang, L., Jing, Z., Xu, S., Zhang, D., and Shou, B. Unified parallel C for GPU clusters: language extensions and compiler implementation. In Proc. 23rd International Conference on Languages and Compilers for Parallel Computing (LCPC'10) (Oct. 2010), pp. 151--165."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/762761.762821"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/1504176.1504194"},{"key":"e_1_2_1_6_1","volume-title":"NVIDIA's next generation CUDA computer architecture: Fermi","author":"NVIDIA.","year":"2009","unstructured":"NVIDIA. NVIDIA's next generation CUDA computer architecture: Fermi , 2009 . NVIDIA. NVIDIA's next generation CUDA computer architecture: Fermi, 2009."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5161065"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/HPCSim.2011.5999803"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2016604.2016612"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2010.12"}],"container-title":["ACM SIGMETRICS Performance Evaluation Review"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2381056.2381081","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2381056.2381081","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T20:00:40Z","timestamp":1750276840000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2381056.2381081"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,10,8]]},"references-count":10,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2012,10,8]]}},"alternative-id":["10.1145\/2381056.2381081"],"URL":"https:\/\/doi.org\/10.1145\/2381056.2381081","relation":{},"ISSN":["0163-5999"],"issn-type":[{"type":"print","value":"0163-5999"}],"subject":[],"published":{"date-parts":[[2012,10,8]]},"assertion":[{"value":"2012-10-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}