{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T21:05:41Z","timestamp":1784667941448,"version":"3.55.0"},"reference-count":31,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2020,3,9]],"date-time":"2020-03-09T00:00:00Z","timestamp":1583712000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000266","name":"Engineering and Physical Sciences Research Council","doi-asserted-by":"publisher","award":["13220161 and EP\/K026399\/1"],"award-info":[{"award-number":["13220161 and EP\/K026399\/1"]}],"id":[{"id":"10.13039\/501100000266","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Parallel Comput."],"published-print":{"date-parts":[[2020,3,31]]},"abstract":"<jats:p>This article demonstrates the utility and implementation of software prefetching in an unstructured finite volume computational fluid dynamics code of representative size and complexity to an industrial application and across a number of modern processors. We present the benefits of auto-tuning for finding the optimal prefetch distance values across different computational kernels and architectures and demonstrate the importance of choosing the right prefetch destination across the available cache levels for best performance. We discuss the impact of the data layout on the number of prefetch instructions required in kernels with indirect addressing patterns and show how to best implement them in an existing large-scale computational fluid dynamics application. Through this, we show significant full application speed-ups on a range of processors and realistic test cases in both single core\/tile and full socket configurations, such as 1.14\u00d7 on the Intel Xeon Sandy Bridge, 1.09\u00d7 on the Intel Xeon Broadwell, 1.29\u00d7 on the Intel Xeon Skylake, 1.99\u00d7 on the in-order Intel Xeon Phi Knights Corner coprocessor, and 1.51\u00d7 on the out-of-order Intel Xeon Phi Knights Landing many-core processor.<\/jats:p>","DOI":"10.1145\/3380932","type":"journal-article","created":{"date-parts":[[2020,3,9]],"date-time":"2020-03-09T08:20:30Z","timestamp":1583742030000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Software Prefetching for Unstructured Mesh Applications"],"prefix":"10.1145","volume":"7","author":[{"given":"Ioan","family":"Hadade","sequence":"first","affiliation":[{"name":"Oxford Thermofluids Institute, University of Oxford, Oxford, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Timothy M.","family":"Jones","sequence":"additional","affiliation":[{"name":"Computer Laboratory, University of Cambridge, Cambridge, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Feng","family":"Wang","sequence":"additional","affiliation":[{"name":"Oxford Thermofluids Institute, University of Oxford, Oxford, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Luca di","family":"Mare","sequence":"additional","affiliation":[{"name":"Oxford Thermofluids Institute, University of Oxford, Oxford, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,3,9]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"OpenMP Application Program Interface. 2016. OpenMP 4.0 Specifications. Retrieved from http:\/\/www.openmp.org\/wp-content\/uploads\/OpenMP4.0.0.pdf."},{"key":"e_1_2_1_2_1","unstructured":"FUN3D Manual. 2017. FUN3D. Retrieved from https:\/\/fun3d.larc.nasa.gov\/."},{"key":"e_1_2_1_3_1","volume-title":"Proceedings of the IEEE\/ACM International Symposium on Code Generation and Optimization (CGO\u201917)","author":"Ainsworth Sam","unstructured":"Sam Ainsworth and Timothy M. Jones. 2017. Software prefetching for indirect memory accesses. In Proceedings of the IEEE\/ACM International Symposium on Code Generation and Optimization (CGO\u201917). IEEE Press, Piscataway, NJ, 305--317."},{"key":"e_1_2_1_4_1","unstructured":"Intel Vtune Amplifier. 2019. Intel Vtune Amplifier. Retrieved from https:\/\/software.intel.com\/en-us\/vtune."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/125826.125925"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/106972.106979"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1115\/GT2014-25772"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1115\/1.4034600"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of ASME Turbo Expo.","author":"Mare Luca Di","unstructured":"Luca Di Mare, Davendu Y. Kulkarni, Feng Wang, Artyom Romanov, Pandia R. Ramar, and Zacharias I. Zachariadis. 2011. Virtual gas turbines: Geometry and conceptual description. In Proceedings of ASME Turbo Expo."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2018.2826533"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0167-8191(00)00075-2"},{"key":"e_1_2_1_12_1","unstructured":"Ioan Hadade. 2018. Bitbucket. Retrieved from https:\/\/bitbucket.org\/ioanhadade\/au3x-ia3-reproduce\/."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/IA3.2018.00009"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2018.07.001"},{"key":"e_1_2_1_15_1","volume-title":"Numerical Computation of Internal and External Flows","author":"Hirsch Charles","unstructured":"Charles Hirsch. 1990. Numerical Computation of Internal and External Flows. John Wiley and Sons, Chichester, West Sussex, UK."},{"key":"e_1_2_1_16_1","first-page":"248966","article-title":"Intel\u00ae 64 and IA-32 Architectures Optimization Reference Manual","author":"Intel Corporation","year":"2017","unstructured":"Intel Corporation. 2017. Intel\u00ae 64 and IA-32 Architectures Optimization Reference Manual. Number 248966-037.","journal-title":"Number"},{"key":"e_1_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Jim Jeffers James Reinders and Avinash Sodani. 2016. Intel Xeon Phi Processor High Performance Programming. Morgan Kaufmann.","DOI":"10.1016\/B978-0-12-809194-4.00013-2"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2133382.2133384"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1002\/cnm.1160"},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","unstructured":"D. J. Mavriplis. 2003. Revisiting the Least-squares Procedure for Gradient Reconstruction on Unstructured Meshes. Technical Report NASA\/CR-2003-212683. National Aeronautics and Space Administration.","DOI":"10.2514\/6.2003-3986"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/143371.143488"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2015.114"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00348-010-0902-4"},{"key":"e_1_2_1_24_1","unstructured":"Perf. 2019. Perk Wiki. Retrieved from https:\/\/perf.wiki.kernel.org\/index.php\/Main_Page."},{"key":"e_1_2_1_25_1","volume-title":"Technical Report TP-1337. National Aeronautics and Space Administration.","author":"Reid L.","year":"1978","unstructured":"L. Reid and D. Moore. 1978. Design and Overall Performance of Four Highly Loaded, High-speed Inlet Stages for an Advanced High-pressure-ratio Core Compressor. Technical Report TP-1337. National Aeronautics and Space Administration."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1016\/0021-9991(81)90128-5"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1016\/0021-9991(79)90145-1"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/358923.358939"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1115\/GT2016-56227"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.2514\/3.10041"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/216585.216588"}],"container-title":["ACM Transactions on Parallel Computing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3380932","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3380932","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T17:08:07Z","timestamp":1755882487000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3380932"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,3,9]]},"references-count":31,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,3,31]]}},"alternative-id":["10.1145\/3380932"],"URL":"https:\/\/doi.org\/10.1145\/3380932","relation":{},"ISSN":["2329-4949","2329-4957"],"issn-type":[{"value":"2329-4949","type":"print"},{"value":"2329-4957","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,3,9]]},"assertion":[{"value":"2018-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-10-01","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-03-09","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}