{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,30]],"date-time":"2025-10-30T07:14:04Z","timestamp":1761808444350,"version":"3.37.3"},"reference-count":26,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2021,4,28]],"date-time":"2021-04-28T00:00:00Z","timestamp":1619568000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springer.com\/tdm"},{"start":{"date-parts":[[2021,4,28]],"date-time":"2021-04-28T00:00:00Z","timestamp":1619568000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springer.com\/tdm"}],"funder":[{"DOI":"10.13039\/501100013287","name":"Science Challenge Project","doi-asserted-by":"publisher","award":["TZ2016002"],"award-info":[{"award-number":["TZ2016002"]}],"id":[{"id":"10.13039\/501100013287","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2021,11]]},"DOI":"10.1007\/s11227-021-03823-3","type":"journal-article","created":{"date-parts":[[2021,4,28]],"date-time":"2021-04-28T04:04:04Z","timestamp":1619582644000},"page":"13584-13600","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Multilevel parallelism optimization of stencil computations on SIMDlized NUMA architectures"],"prefix":"10.1007","volume":"77","author":[{"given":"Kaifang","family":"Zhang","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Huayou","family":"Su","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yong","family":"Dou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,4,28]]},"reference":[{"key":"3823_CR1","unstructured":"The top-500 list of supercomputer sites. Available from: http:\/\/www.top500.org\/lists\/"},{"key":"3823_CR2","unstructured":"Lameter C (2006) Local and remote memory: memory in a Linux or NUMA System. Linux symposium"},{"key":"3823_CR3","volume-title":"Parallel Programming in OpenMP","author":"R Chandra","year":"2000","unstructured":"Chandra R, Menon R, Dagum L, Kohr D, Maydan D, McDonald J (2000) Parallel Programming in OpenMP. Morgan Kaufman, San Francisco"},{"key":"3823_CR4","volume-title":"Parallel Programming with MPI","author":"PS Pacheco","year":"1997","unstructured":"Pacheco PS (1997) Parallel Programming with MPI. Morgan Kaufman, San Francisco"},{"key":"3823_CR5","volume-title":"Using MPI: Portable Parallel Programming with the Message-Passing Interface","author":"W Gropp","year":"1994","unstructured":"Gropp W, Lusk E, Skjellum A (1994) Using MPI: Portable Parallel Programming with the Message-Passing Interface. MPI Press, Cambridge"},{"key":"3823_CR6","doi-asserted-by":"publisher","first-page":"562","DOI":"10.1016\/j.parco.2011.02.002","volume":"37","author":"Haoqiang Jin","year":"2011","unstructured":"Jin Haoqiang, Jespersen Dennis, Mehrotra Piyush, Biswas Rupak, Huang Lei, Chapman Barbara (2011) High performance computing using MPI and OpenMP on multi-core parallel systems. Parallel Comput. 37:562\u2013575. https:\/\/doi.org\/10.1016\/j.parco.2011.02.002","journal-title":"Parallel Comput."},{"key":"3823_CR7","doi-asserted-by":"crossref","unstructured":"Uday Bondhugula (2013) Compiling affine loop nests for distributed-memory parallel architectures. International Conference for High Performance Computing, Networking, Storage and Analysis, SC. https:\/\/doi.org\/10.1145-2503210.2503289","DOI":"10.1145\/2503210.2503289"},{"key":"3823_CR8","unstructured":"Optimizing applications for NUMA. Development topics and technologies, Intel. Published: 11\/02\/2011, Last Updated: 11\/02\/2011. https:\/\/software.intel.com\/content\/www\/us\/en\/develop\/articles\/optimizing-applications-for-numa.html"},{"issue":"1","key":"3823_CR9","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1137\/070693199","volume":"51","author":"Datta Kaushik","year":"2009","unstructured":"Kaushik Datta, Shoaib Kamil, Samuel Williams, Leonid Oliker, John Shalf, Katherine A (2009) Yelick: optimization and performance modeling of stencil computations on modern microprocessors. SIAM Rev. 51(1):129\u2013159","journal-title":"SIAM Rev."},{"key":"3823_CR10","doi-asserted-by":"publisher","unstructured":"Henretty T, Veras R, Franchetti F, Pouchet L-N, Ramanujam J, Sadayappan P (2013) A stencil compiler for short-vector SIMD architectures. In: Proceedings of the 27th International ACM Conference on International Conference on Supercomputing (ICS'13). Association for Computing Machinery, New York, NY, USA, pp 13\u201324. https:\/\/doi.org\/10.1145\/2464996.2467268","DOI":"10.1145\/2464996.2467268"},{"key":"3823_CR11","doi-asserted-by":"crossref","unstructured":"Tom Henretty (2014) Performance optimization of stencil computations on modern SIMD architectures[Ph.D]","DOI":"10.1145\/2464996.2467268"},{"key":"3823_CR12","doi-asserted-by":"publisher","DOI":"10.1007\/s11227-019-02842-5","author":"Adri\u00e1 Armejach","year":"2019","unstructured":"Armejach Adri\u00e1, Caminal Helena, Cebrian Juan, Langarita Rub\u0117n, Gonz\u00e1lez-Alberquilla Rekai, Adeniyi-Jones Chris, Valero Mateo, Casas Marc, Miquel Moreto (2019) Using Arm\u2019s scalable vector extension on stencil codes. J Supercomput. https:\/\/doi.org\/10.1007\/s11227-019-02842-5","journal-title":"J Supercomput"},{"key":"3823_CR13","doi-asserted-by":"publisher","unstructured":"Jang M, Kim K, Kim K (2011) The performance analysis of ARM NEON technology for mobile platforms. In: RACS'11: Proceedings of the 2011 ACM symposium on research in applied computation, pp 104\u2013106. https:\/\/doi.org\/10.1145\/2103380.2103401","DOI":"10.1145\/2103380.2103401"},{"key":"3823_CR14","unstructured":"Elena L, Arseny T, Dmitry N, Vladimir A (2020) Fast implementation of morphological filtering using ARM NEON extension. arXiv:2002.09474"},{"key":"3823_CR15","first-page":"280","volume-title":"Jens Gustedt","author":"Saied Mariem","year":"2016","unstructured":"Mariem Saied (2016) Jens Gustedt. Automatic Code Generation for Iterative Multi-dimensional Stencil Computations. HiPC, Gilles Muller, pp 280\u2013289"},{"key":"3823_CR16","unstructured":"Saied, Mariem (2018) Automatic code generation and optimization of multi-dimensional stencil computations on distributed-memory architectures"},{"key":"3823_CR17","unstructured":"Grid5000. Available from: https:\/\/www.grid5000.fr\/"},{"key":"3823_CR18","unstructured":"Kaushik Datta, Samuel Williams, Vasily Volkov, et\u00a0al (2009) Auto-tuning the 27-point stencil for multicore[J]. Proc Iwapt the Fourth International Workshop on Automatic Performance Tuning"},{"key":"3823_CR19","doi-asserted-by":"publisher","first-page":"73","DOI":"10.1155\/2001\/712152","volume":"9","author":"Timothy Kaiser","year":"2001","unstructured":"Kaiser Timothy, Baden Scott (2001) Overlapping communication and computation with OpenMP and MPI. Sci. Programm. 9:73\u201381. https:\/\/doi.org\/10.1155\/2001\/712152","journal-title":"Sci. Programm."},{"key":"3823_CR20","doi-asserted-by":"publisher","unstructured":"Dathathri Roshan, Reddy Chandan, Ramashekar Thejas, Bondhugula Uday (2013) Generating efficient data movement code for heterogeneous architectures with distributed-memory. Parallel Architectures and Compilation Techniques-Conference Proceedings, PACT. 375\u2013386. https:\/\/doi.org\/10.1109\/PACT.2013.6618833","DOI":"10.1109\/PACT.2013.6618833"},{"key":"3823_CR21","doi-asserted-by":"crossref","unstructured":"Bondhugula Uday, Baskaran Muthu, Krishnamo Orthy Sriram, Ramanujam J., Rountev Atanas, and Sadayappan Ponnuswamy (2008) Automatic Transformations for Communication-Minimized Parallelization and Locality Optimization in the Polyhedral Model. International Conference on Compiler Construction. 4959. 132-146","DOI":"10.1007\/978-3-540-78791-4_9"},{"key":"3823_CR22","doi-asserted-by":"crossref","unstructured":"Rabenseifner Rolf, Hager Georg, and Jost Gabriele (2009) Hybrid MPI\/OpenMP parallel programming on clusters of multi-core SMP nodes. Euromicro International Conference on Parallel, Distributed and Network-based Processing 427-436","DOI":"10.1109\/PDP.2009.43"},{"issue":"1","key":"3823_CR23","doi-asserted-by":"publisher","first-page":"129","DOI":"10.1137\/070693199","volume":"51","author":"DK Kamil","year":"2009","unstructured":"Kamil DK et al (2009) Optimization and performance modeling of stencil computations on modern microprocessors. SIAM Rev 51(1):129\u2013159","journal-title":"SIAM Rev"},{"key":"3823_CR24","doi-asserted-by":"crossref","unstructured":"Bandishti V., Pananilath I., Bondhugula Uday (2012) Tiling stencil computations to maximize parallelism[C]$$\/\/$$ International Conference on High Performance Computing. IEEE Computer Society Press","DOI":"10.1109\/SC.2012.107"},{"key":"3823_CR25","doi-asserted-by":"crossref","unstructured":"Wellein G., Hager G., Zeiser T., et\u00a0al (2009) Efficient temporal blocking for stencil computations by multicore-aware wavefront parallelization[C]$$\/\/$$ IEEE International Computer Software and Applications Conference. IEEE","DOI":"10.1109\/COMPSAC.2009.82"},{"key":"3823_CR26","doi-asserted-by":"crossref","unstructured":"Datta K, Murphy M, Volkov V et al (2008) Stencil computation optimization and auto-tuning on state-of-the-art multicore architectures. In: Proceedings of the 2008 ACM\/IEEE Conference on Supercomputing (SC'08). IEEE Press, Article 4, pp 1\u201312","DOI":"10.1109\/SC.2008.5222004"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-021-03823-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-021-03823-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-021-03823-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,10,25]],"date-time":"2021-10-25T09:46:26Z","timestamp":1635155186000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-021-03823-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,28]]},"references-count":26,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2021,11]]}},"alternative-id":["3823"],"URL":"https:\/\/doi.org\/10.1007\/s11227-021-03823-3","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"type":"print","value":"0920-8542"},{"type":"electronic","value":"1573-0484"}],"subject":[],"published":{"date-parts":[[2021,4,28]]},"assertion":[{"value":"17 April 2021","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 April 2021","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}