{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,1]],"date-time":"2026-05-01T23:00:55Z","timestamp":1777676455376,"version":"3.51.4"},"reference-count":17,"publisher":"SAGE Publications","issue":"3","license":[{"start":{"date-parts":[[2011,6,29]],"date-time":"2011-06-29T00:00:00Z","timestamp":1309305600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/journals.sagepub.com\/page\/policies\/text-and-data-mining-license"}],"content-domain":{"domain":["journals.sagepub.com"],"crossmark-restriction":true},"short-container-title":["The International Journal of High Performance Computing Applications"],"published-print":{"date-parts":[[2011,8]]},"abstract":"<jats:p>In this paper, we describe the implementation of a multi-graphical processing unit (GPU) fluid flow solver based on the lattice Boltzmann method (LBM). The LBM is a novel approach in computational fluid dynamics, with numerous interesting features from a computational, numerical, and physical standpoint. Our program is based on CUDA and uses POSIX threads to manage multiple computation devices. Using recently released hardware, our solver may therefore run eight GPUs in parallel, which allows us to perform simulations at a rather large scale. Performance and scalability are excellent, the speedup over sequential implementations being at least of two orders of magnitude. In addition, we discuss tiling and communication issues for present and forthcoming implementations.<\/jats:p>","DOI":"10.1177\/1094342011414745","type":"journal-article","created":{"date-parts":[[2011,6,29]],"date-time":"2011-06-29T20:46:44Z","timestamp":1309380404000},"page":"295-303","update-policy":"https:\/\/doi.org\/10.1177\/sage-journals-update-policy","source":"Crossref","is-referenced-by-count":13,"title":["The TheLMA project: Multi-GPU implementation of the lattice Boltzmann method"],"prefix":"10.1177","volume":"25","author":[{"given":"Christian","family":"Obrecht","sequence":"first","affiliation":[{"name":"EDF R&D, D\u00e9partement EnerBAT, Moret-sur-Loing Cedex, France, , Universit\u00e9 de Lyon, Lyon Cedex 07, France, INSA-Lyon, CETHIL, UMR5008, Villeurbanne Cedex, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fr\u00e9d\u00e9ric","family":"Kuznik","sequence":"additional","affiliation":[{"name":"Universit\u00e9 de Lyon, Lyon Cedex 07, France, INSA-Lyon, CETHIL, UMR5008, Villeurbanne Cedex, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bernard","family":"Tourancheau","sequence":"additional","affiliation":[{"name":"Universit\u00e9 de Lyon, Lyon Cedex 07, France, INSA-Lyon, CITI, INRIA, Villeurbanne Cedex, France, Universit\u00e9 Lyon 1, Villeurbanne Cedex, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jean-Jacques","family":"Roux","sequence":"additional","affiliation":[{"name":"Universit\u00e9 de Lyon, Lyon Cedex 07, France, INSA-Lyon, CETHIL, UMR5008, Villeurbanne Cedex, France"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"179","published-online":{"date-parts":[[2011,6,29]]},"reference":[{"key":"atypb1","doi-asserted-by":"publisher","DOI":"10.1088\/1742-6596\/180\/1\/012037"},{"key":"atypb2","volume-title":"18th Euromicro International Conference on Parallel, Distributed and Network-Based Processing (PDP)","author":"Broquedis F."},{"key":"atypb3","volume-title":"Proceedings of the 18th International Symposium on Rarefied Gas Dynamics","author":"Humi\u00e8res D."},{"key":"atypb4","doi-asserted-by":"publisher","DOI":"10.1098\/rsta.2001.0955"},{"key":"atypb5","volume-title":"Proceedings of HPCMP Users Group Conference","author":"Dongarra J."},{"key":"atypb6","volume-title":"Proceedings of the 2004 ACM\/IEEE conference on Supercomputing","author":"Fan Z."},{"key":"atypb7","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.56.1505"},{"key":"atypb8","volume":"27","author":"Kuznik F.","year":"2009","journal-title":"Computers and Mathematics with Applications"},{"key":"atypb9","doi-asserted-by":"publisher","DOI":"10.1103\/PhysRevLett.61.2332"},{"key":"atypb10","volume-title":"Compute Unified Device Architecture Programming Guide version 3.1.1","author":"Nvidia","year":"2010"},{"key":"atypb11","doi-asserted-by":"publisher","DOI":"10.1016\/j.camwa.2010.01.054"},{"key":"atypb12","volume":"61","author":"Obrecht C.","year":"2010","journal-title":"Lecture Notes in Computer Science"},{"key":"atypb13","author":"Papadopoulou M.","year":"2009","journal-title":"Technical report"},{"key":"atypb14","doi-asserted-by":"publisher","DOI":"10.1007\/s00450-009-0087-3"},{"key":"atypb15","volume-title":"Optimizing matrix transpose in CUDA. NVIDIA CUDA SDK Application Note","author":"Ruetsch G.","year":"2009"},{"key":"atypb16","first-page":"1","volume":"13","author":"T\u00f6lke J.","year":"2008","journal-title":"Comput Visualiz Sci"},{"key":"atypb17","doi-asserted-by":"publisher","DOI":"10.1080\/10618560802238275"}],"container-title":["The International Journal of High Performance Computing Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342011414745","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/journals.sagepub.com\/doi\/pdf\/10.1177\/1094342011414745","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T08:18:59Z","timestamp":1777450739000},"score":1,"resource":{"primary":{"URL":"https:\/\/journals.sagepub.com\/doi\/10.1177\/1094342011414745"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,6,29]]},"references-count":17,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2011,8]]}},"alternative-id":["10.1177\/1094342011414745"],"URL":"https:\/\/doi.org\/10.1177\/1094342011414745","relation":{},"ISSN":["1094-3420","1741-2846"],"issn-type":[{"value":"1094-3420","type":"print"},{"value":"1741-2846","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,6,29]]}}}