{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:23:43Z","timestamp":1750220623415,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":45,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,1,15]],"date-time":"2020-01-15T00:00:00Z","timestamp":1579046400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2016YFB1000503"],"award-info":[{"award-number":["2016YFB1000503"]}]},{"name":"National Natural Science Foundation of China","award":["61502019"],"award-info":[{"award-number":["61502019"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,1,15]]},"DOI":"10.1145\/3368474.3368487","type":"proceedings-article","created":{"date-parts":[[2019,12,17]],"date-time":"2019-12-17T13:31:42Z","timestamp":1576589502000},"page":"32-42","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Towards GPU Acceleration of Phonon Computation with ShengBTE"],"prefix":"10.1145","author":[{"given":"Yi","family":"Wei","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xin","family":"You","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hailong","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhongzhi","family":"Luan","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Depei","family":"Qian","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Beihang University, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,1,15]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"2019. Database-almaBTE. http:\/\/www.almabte.eu\/index.php\/database\/  2019. Database-almaBTE. http:\/\/www.almabte.eu\/index.php\/database\/"},{"key":"e_1_3_2_1_2_1","unstructured":"2019. GPUs Now Accelerate Almost 600 HPC Apps. https:\/\/blogs.nvidia.com\/blog\/2018\/11\/16\/gpus-now-accelerate-almost-600-hpc-apps\/  2019. GPUs Now Accelerate Almost 600 HPC Apps. https:\/\/blogs.nvidia.com\/blog\/2018\/11\/16\/gpus-now-accelerate-almost-600-hpc-apps\/"},{"key":"e_1_3_2_1_3_1","unstructured":"2019. Intel Xeon Processor E5-2680 v4 Product Specifications. https:\/\/ark.intel.com\/content\/www\/us\/en\/ark\/products\/91754\/intel-xeon-processor-e5-2680-v4-35m-cache-2-40-ghz.html  2019. Intel Xeon Processor E5-2680 v4 Product Specifications. https:\/\/ark.intel.com\/content\/www\/us\/en\/ark\/products\/91754\/intel-xeon-processor-e5-2680-v4-35m-cache-2-40-ghz.html"},{"key":"e_1_3_2_1_4_1","unstructured":"2019. Materials Project. https:\/\/materialsproject.org\/  2019. Materials Project. https:\/\/materialsproject.org\/"},{"key":"e_1_3_2_1_5_1","unstructured":"2019. Official website of Shengbte. http:\/\/www.shengbte.org\/home  2019. Official website of Shengbte. http:\/\/www.shengbte.org\/home"},{"key":"e_1_3_2_1_6_1","unstructured":"2019. Tesla P100 Data Center Accelerator. https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-p100\/  2019. Tesla P100 Data Center Accelerator. https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-p100\/"},{"key":"e_1_3_2_1_7_1","unstructured":"2019. Tesla V100 Data Center Accelerator. https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-v100\/  2019. Tesla V100 Data Center Accelerator. https:\/\/www.nvidia.com\/en-us\/data-center\/tesla-v100\/"},{"key":"e_1_3_2_1_8_1","volume-title":"Tensorfiow: A system for large-scale machine learning. In 12th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 16). 265--283.","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi , Paul Barham , Jianmin Chen , Zhifeng Chen , Andy Davis , Jeffrey Dean , Matthieu Devin , Sanjay Ghemawat , Geoffrey Irving , Michael Isard , 2016 . Tensorfiow: A system for large-scale machine learning. In 12th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 16). 265--283. Mart\u00edn Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. 2016. Tensorfiow: A system for large-scale machine learning. In 12th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 16). 265--283."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevB.88.144302"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.5555\/2872599.2872609"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.58.1861"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRev.113.1046"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2017.06.023"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2015.01.008"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.45"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevB.67.144304"},{"key":"e_1_3_2_1_17_1","unstructured":"Reinders J VTune Performance Analyzer Essentials. 2005. Measurement and Tuning Techniques for Software Developers.  Reinders J VTune Performance Analyzer Essentials. 2005. Measurement and Tuning Techniques for Software Developers."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1088\/0953-8984\/21\/39\/395502"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2009.07.007"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1002\/jcc.21057"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5194\/gmd-3-415-2010"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.compfluid.2014.12.010"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2654889"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.7554\/eLife.18722"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2010.04.120"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.camwa.2009.08.052"},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2014.02.015"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00371-003-0210-6"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.compfluid.2012.01.018"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.camwa.2011.02.020"},{"volume-title":"CUDA by example: an introduction to general-purpose GPU programming","author":"Sanders Jason","key":"e_1_3_2_1_31_1","unstructured":"Jason Sanders and Edward Kandrot . 2010. CUDA by example: an introduction to general-purpose GPU programming . Addison-Wesley Professional . Jason Sanders and Edward Kandrot. 2010. CUDA by example: an introduction to general-purpose GPU programming. Addison-Wesley Professional."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1175\/BAMS-D-11-00059.1"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jsb.2012.09.006"},{"key":"e_1_3_2_1_34_1","volume-title":"OpenCL: A parallel programming standard for heterogeneous computing systems. Computing in science & engineering 12, 3","author":"Stone John E","year":"2010","unstructured":"John E Stone , David Gohara , and Guochun Shi . 2010. OpenCL: A parallel programming standard for heterogeneous computing systems. Computing in science & engineering 12, 3 ( 2010 ), 66. John E Stone, David Gohara, and Guochun Shi. 2010. OpenCL: A parallel programming standard for heterogeneous computing systems. Computing in science & engineering 12, 3 (2010), 66."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1088\/0953-8984\/26\/22\/225402"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.pepi.2008.10.003"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevB.91.094306"},{"key":"e_1_3_2_1_38_1","unstructured":"Florian Wende Thomas Steinke and Frank Cordes. 2014. Multi-threaded kernel offloading to gpgpu using hyper-q on kepler architecture. (2014).  Florian Wende Thomas Steinke and Frank Cordes. 2014. Multi-threaded kernel offloading to gpgpu using hyper-q on kepler architecture. (2014)."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-32820-6_85"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.2172\/1407078"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNANO.2010.2089638"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1115\/1.1857941"},{"key":"e_1_3_2_1_43_1","volume-title":"Thermal Management of Microelectronic Equipment: Heat Transfer Theory. Analysis Methods, and Design Practices","author":"Yeh LT","year":"2002","unstructured":"LT Yeh and RC Chu . 2002. Thermal Management of Microelectronic Equipment: Heat Transfer Theory. Analysis Methods, and Design Practices , ASME , New York ( 2002 ). LT Yeh and RC Chu. 2002. Thermal Management of Microelectronic Equipment: Heat Transfer Theory. Analysis Methods, and Design Practices, ASME, New York (2002)."},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1039\/C1EE02497C"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.procs.2012.04.104"}],"event":{"name":"HPCAsia2020: International Conference on High Performance Computing in Asia-Pacific Region","sponsor":["IPSJ","SIGHPC ACM Special Interest Group on High Performance Computing, Special Interest Group on High Performance Computing"],"location":"Fukuoka Japan","acronym":"HPCAsia2020"},"container-title":["Proceedings of the International Conference on High Performance Computing in Asia-Pacific Region"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3368474.3368487","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3368474.3368487","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:01:26Z","timestamp":1750197686000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3368474.3368487"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,1,15]]},"references-count":45,"alternative-id":["10.1145\/3368474.3368487","10.1145\/3368474"],"URL":"https:\/\/doi.org\/10.1145\/3368474.3368487","relation":{},"subject":[],"published":{"date-parts":[[2020,1,15]]},"assertion":[{"value":"2020-01-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}