{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T23:41:44Z","timestamp":1783035704292,"version":"3.54.6"},"reference-count":55,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2018,6,12]],"date-time":"2018-06-12T00:00:00Z","timestamp":1528761600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["91530323"],"award-info":[{"award-number":["91530323"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"National Key R8D Plan of China","award":["2016YFB0200603"],"award-info":[{"award-number":["2016YFB0200603"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Archit. Code Optim."],"published-print":{"date-parts":[[2018,6,30]]},"abstract":"<jats:p>High-order stencil computations, frequently found in many applications, pose severe challenges to emerging many-core platforms due to the complexities of hardware architectures as well as the sophisticated computing and data movement patterns. In this article, we tackle the challenges of high-order WENO computations in extreme-scale simulations of 3D gaseous waves on Sunway TaihuLight. We design efficient parallelization algorithms and present effective optimization techniques to fully exploit various parallelisms with reduced memory footprints, enhanced data reuse, and balanced computation load. Test results show the optimized code can scale to 9.98 million cores, solving 12.74 trillion unknowns with 23.12 Pflops double-precision performance.<\/jats:p>","DOI":"10.1145\/3209208","type":"journal-article","created":{"date-parts":[[2018,6,12]],"date-time":"2018-06-12T18:12:29Z","timestamp":1528827149000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Extreme-Scale High-Order WENO Simulations of 3-D Detonation Wave with 10 Million Cores"],"prefix":"10.1145","volume":"15","author":[{"given":"Ying","family":"Cai","sequence":"first","affiliation":[{"name":"Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yulong","family":"Ao","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7426-6248","authenticated-orcid":false,"given":"Chao","family":"Yang","sequence":"additional","affiliation":[{"name":"Peking University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Wenjing","family":"Ma","sequence":"additional","affiliation":[{"name":"Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Haitao","family":"Zhao","sequence":"additional","affiliation":[{"name":"Chinese Academy of Sciences, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2018,6,12]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2017.9"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.103"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevE.72.046307"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0021-9991(02)00032-3"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0045-7825(99)00186-3"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.proeng.2013.08.016"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1137\/070693199"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/1413370.1413375"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the 2016 IEEE 18th International Conference on High Performance Computing and Communications","author":"Dong Wenqian","unstructured":"Wenqian Dong , Letian Kang , Zhe Quan , Kenli Li , Keqin Li , Ziyu Hao , and Xiang-Hui Xie . 2016. Implementing molecular dynamics simulation on sunway taihulight system . In Proceedings of the 2016 IEEE 18th International Conference on High Performance Computing and Communications ; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS\u201916). IEEE , 443--450. Wenqian Dong, Letian Kang, Zhe Quan, Kenli Li, Keqin Li, Ziyu Hao, and Xiang-Hui Xie. 2016. Implementing molecular dynamics simulation on sunway taihulight system. In Proceedings of the 2016 IEEE 18th International Conference on High Performance Computing and Communications; IEEE 14th International Conference on Smart City; IEEE 2nd International Conference on Data Science and Systems (HPCC\/SmartCity\/DSS\u201916). IEEE, 443--450."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.combustflame.2008.06.013"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-03869-3_61"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2017.20"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1086\/422513"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126910"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3126908.3126909"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-016-5588-7"},{"key":"e_1_2_1_17_1","first-page":"8","article-title":"A finite difference method for the computation of discontinuous solutions of the equations of fluid dynamics. Sbornik","volume":"47","author":"Godunov S. K.","year":"1959","unstructured":"S. K. Godunov . 1959 . A finite difference method for the computation of discontinuous solutions of the equations of fluid dynamics. Sbornik : Math. 47 , 8 -- 9 (1959), 357--393. S. K. Godunov. 1959. A finite difference method for the computation of discontinuous solutions of the equations of fluid dynamics. Sbornik: Math. 47, 8--9 (1959), 357--393.","journal-title":"Math."},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1137\/S003614450036757X"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcph.1996.0130"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcph.1999.6207"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2017.51"},{"key":"e_1_2_1_22_1","volume-title":"Scalable graph traversal on sunway taihulight with ten million cores","author":"Lin Heng","unstructured":"Heng Lin , Xiongchao Tang , Bowen Yu , Youwei Zhuo , Wenguang Chen , Jidong Zhai , Wanwang Yin , and Weimin Zheng . 2017. Scalable graph traversal on sunway taihulight with ten million cores . IEEE , 635--645. Heng Lin, Xiongchao Tang, Bowen Yu, Youwei Zhuo, Wenguang Chen, Jidong Zhai, Wanwang Yin, and Weimin Zheng. 2017. Scalable graph traversal on sunway taihulight with ten million cores. IEEE, 635--645."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1006\/jcph.1994.1187"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/2400682.2400718"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11390-016-1696-5"},{"key":"e_1_2_1_26_1","volume-title":"Retrieved","author":"Meuer Hans","year":"2017","unstructured":"Hans Meuer , Erich Strohmaier , Jack Dongarra , Horst Simon , and Meuer Martin . 2017 . Top 500 Supercomputer Lists . Retrieved November 14, 2017 from http:\/\/www.top500.org. Hans Meuer, Erich Strohmaier, Jack Dongarra, Horst Simon, and Meuer Martin. 2017. Top 500 Supercomputer Lists. Retrieved November 14, 2017 from http:\/\/www.top500.org."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2010.2"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2009.5161011"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1063\/1.1649336"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.5555\/3014904.3014911"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.newast.2006.04.007"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.5555\/370049.370403"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/0021-9991(81)90128-5"},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2837476.2837483"},{"key":"e_1_2_1_36_1","volume-title":"40th AIAA Aerospace Sciences Meeting & Exhibit. 773","author":"Shepherd J.","unstructured":"J. Shepherd , F. Pintgen , J. Austin , and C. Eckett . 2002. The structure of the detonation front in gases . In 40th AIAA Aerospace Sciences Meeting & Exhibit. 773 . J. Shepherd, F. Pintgen, J. Austin, and C. Eckett. 2002. The structure of the detonation front in gases. In 40th AIAA Aerospace Sciences Meeting & Exhibit. 773."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2063384.2063388"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1137\/070679065"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10766-016-0454-1"},{"key":"e_1_2_1_40_1","first-page":"345","article-title":"Multi-dimensional detonation wave structure","volume":"15","author":"Strehlow R. A.","year":"1970","unstructured":"R. A. Strehlow . 1970 . Multi-dimensional detonation wave structure . Astronautica Acta 15 (1970), 345 -- 357 . R. A. Strehlow. 1970. Multi-dimensional detonation wave structure. Astronautica Acta 15 (1970), 345--357.","journal-title":"Astronautica Acta"},{"key":"e_1_2_1_41_1","volume-title":"Proceedings of the International Conference on Parallel Processing and Applied Mathematics. Springer, 582--592","author":"Szustak Lukasz","year":"2013","unstructured":"Lukasz Szustak , Krzysztof Rojek , and Pawel Gepner . 2013 . Using Intel Xeon Phi coprocessor to accelerate computations in MPDATA algorithm . In Proceedings of the International Conference on Parallel Processing and Applied Mathematics. Springer, 582--592 . Lukasz Szustak, Krzysztof Rojek, and Pawel Gepner. 2013. Using Intel Xeon Phi coprocessor to accelerate computations in MPDATA algorithm. In Proceedings of the International Conference on Parallel Processing and Applied Mathematics. Springer, 582--592."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1155\/2015\/642705"},{"key":"e_1_2_1_43_1","first-page":"51","article-title":"Toward efficient distribution of MPDATA stencil computation on intel MIC architecture","volume":"14","author":"Szustak Lukasz","year":"2014","unstructured":"Lukasz Szustak , Krzysztof Rojek , Roman Wyrzykowski , and Pawel Gepner . 2014 . Toward efficient distribution of MPDATA stencil computation on intel MIC architecture . Proceedings of HiStencils 14 (2014), 51 -- 56 . Lukasz Szustak, Krzysztof Rojek, Roman Wyrzykowski, and Pawel Gepner. 2014. Toward efficient distribution of MPDATA stencil computation on intel MIC architecture. Proceedings of HiStencils 14 (2014), 51--56.","journal-title":"Proceedings of HiStencils"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2011.11.037"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.2322\/tjsass.47.268"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2015.06.001"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11433-009-0270-3"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.combustflame.2012.10.002"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2011.10.002"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/2929908.2929914"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1002\/prs.11620"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2014.82"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2014.2366754"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.5555\/3014904.3014912"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.5555\/3014904.3014910"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2011.11.020"}],"container-title":["ACM Transactions on Architecture and Code Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3209208","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3209208","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:08:24Z","timestamp":1750208904000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3209208"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,6,12]]},"references-count":55,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2018,6,30]]}},"alternative-id":["10.1145\/3209208"],"URL":"https:\/\/doi.org\/10.1145\/3209208","relation":{},"ISSN":["1544-3566","1544-3973"],"issn-type":[{"value":"1544-3566","type":"print"},{"value":"1544-3973","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,6,12]]},"assertion":[{"value":"2018-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-06-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}