{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,2]],"date-time":"2026-07-02T01:05:34Z","timestamp":1782954334521,"version":"3.54.5"},"reference-count":28,"publisher":"Springer Science and Business Media LLC","issue":"15","license":[{"start":{"date-parts":[[2022,5,13]],"date-time":"2022-05-13T00:00:00Z","timestamp":1652400000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springer.com\/tdm"},{"start":{"date-parts":[[2022,5,13]],"date-time":"2022-05-13T00:00:00Z","timestamp":1652400000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springer.com\/tdm"}],"funder":[{"name":"the Major Project on the Integration of Industry, Education and Research of Zhongshan","award":["210610173898370"],"award-info":[{"award-number":["210610173898370"]}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2022,10]]},"DOI":"10.1007\/s11227-022-04491-7","type":"journal-article","created":{"date-parts":[[2022,5,13]],"date-time":"2022-05-13T06:02:57Z","timestamp":1652421777000},"page":"17055-17073","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["An effective 3-D fast fourier transform framework for multi-GPU accelerated distributed-memory systems"],"prefix":"10.1007","volume":"78","author":[{"given":"Binbin","family":"Zhou","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lu","family":"Lu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,5,13]]},"reference":[{"key":"4491_CR1","doi-asserted-by":"crossref","unstructured":"Asaadi H, Khaldi D, Chapman B( 2016) A comparative survey of the hpc and big data paradigms: Analysis and experiments. In: 2016 IEEE International Conference on Cluster Computing (CLUSTER), pp. 423\u2013 432 . IEEE","DOI":"10.1109\/CLUSTER.2016.21"},{"key":"4491_CR2","unstructured":"ORNL (Oak Ridge National Laboratory) (2021): Frontier. https:\/\/www.olcf.ornl.gov\/frontier\/. Accessed: 2021-11-01"},{"key":"4491_CR3","unstructured":"Brown WM ( 2011) Gpu acceleration in lammps. In: LAMMPS User\u2019s Workshop and Symposium"},{"key":"4491_CR4","doi-asserted-by":"crossref","unstructured":"Pronk S, P\u00e1ll S, Schulz R, Larsson P, Bjelkmar P, Apostolov R, Shirts MR, Smith JC, Kasson PM, Van Der\u00a0Spoel D, et al ( 2013) Gromacs 4.5: a high-throughput and highly parallel open source molecular simulation toolkit. Bioinformatics 29( 7), 845\u2013 854","DOI":"10.1093\/bioinformatics\/btt055"},{"key":"4491_CR5","doi-asserted-by":"crossref","unstructured":"Salomon-Ferrer R, Gotz AW, Poole D, Le\u00a0Grand S, Walker RC ( 2013) Routine microsecond molecular dynamics simulations with amber on gpus. 2. explicit solvent particle mesh ewald. Journal of chemical theory and computation 9( 9), 3878\u2013 3888","DOI":"10.1021\/ct400314y"},{"key":"4491_CR6","doi-asserted-by":"crossref","unstructured":"Lee M, Malaya N, Moser RD ( 2013) Petascale direct numerical simulation of turbulent channel flow on up to 786k cores. In: SC\u201913: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis, pp. 1\u2013 11 . IEEE","DOI":"10.1145\/2503210.2503298"},{"issue":"1\u20134","key":"4491_CR7","doi-asserted-by":"publisher","first-page":"109","DOI":"10.1016\/S0045-7825(98)00227-8","volume":"172","author":"J-C Michel","year":"1999","unstructured":"Michel J-C, Moulinec H, Suquet P (1999) Effective properties of composite materials with periodic microstructure: a computational approach. Comput Methods Appl Mech Eng 172(1\u20134):109\u2013143","journal-title":"Comput Methods Appl Mech Eng"},{"key":"4491_CR8","doi-asserted-by":"publisher","first-page":"57","DOI":"10.1016\/j.cpc.2015.10.024","volume":"200","author":"J Jung","year":"2016","unstructured":"Jung J, Kobayashi C, Imamura T, Sugita Y (2016) Parallel implementation of 3d fft with volumetric decomposition schemes for efficient molecular dynamics simulations. Comput Phys Commun 200:57\u201365","journal-title":"Comput Phys Commun"},{"key":"4491_CR9","doi-asserted-by":"publisher","first-page":"273","DOI":"10.1016\/j.actamat.2018.05.036","volume":"154","author":"V Tari","year":"2018","unstructured":"Tari V, Lebensohn RA, Pokharel R, Turner TJ, Shade PA, Bernier JV, Rollett AD (2018) Validation of micro-mechanical fft-based simulations using high energy diffraction microscopy on ti-7al. Acta Mater 154:273\u2013283","journal-title":"Acta Mater"},{"issue":"1","key":"4491_CR10","doi-asserted-by":"publisher","first-page":"39","DOI":"10.1088\/0004-637X\/765\/1\/39","volume":"765","author":"AS Almgren","year":"2013","unstructured":"Almgren AS, Bell JB, Lijewski MJ, Luki\u0107 Z, Van Andel E (2013) Nyx: A massively parallel amr code for computational cosmology. Astrophys J 765(1):39","journal-title":"Astrophys J"},{"issue":"8","key":"4491_CR11","doi-asserted-by":"publisher","first-page":"4962","DOI":"10.1021\/acs.chemrev.0c00998","volume":"121","author":"K Kowalski","year":"2021","unstructured":"Kowalski K, Bair R, Bauman NP, Boschen JS, Bylaska EJ, Daily J, de Jong WA, Dunning T Jr, Govind N, Harrison RJ et al (2021) From nwchem to nwchemex: Evolving with the computational chemistry landscape. Chem Rev 121(8):4962\u20134998","journal-title":"Chem Rev"},{"key":"4491_CR12","unstructured":"NVIDIA: cuFFT. https:\/\/docs.nvidia.com\/cuda\/cufft\/index.html"},{"key":"4491_CR13","unstructured":"ROCmSoftwarePlatform (2018) Rocmsoftwareplatform\/ROCFFT: Next generation FFT implementation for ROCM . https:\/\/github.com\/ROCmSoftwarePlatform\/rocFFT"},{"key":"4491_CR14","unstructured":"Gholami A, Hill J, Malhotra D, Biros G (2015) Accfft: a library for distributed-memory fft on cpu and gpu architectures. arXiv preprint arXiv:1506.07933"},{"key":"4491_CR15","unstructured":"Takahashi D (2014) Ffte: A fast fourier transform package. http:\/\/www.ffte.jp\/"},{"key":"4491_CR16","doi-asserted-by":"crossref","unstructured":"Ayala A, Tomov S, Haidar A, Dongarra J ( 2020) heffte: highly efficient fft for exascale. In: International Conference on Computational Science, pp. 262\u2013 275 . Springer","DOI":"10.1007\/978-3-030-50371-0_19"},{"key":"4491_CR17","unstructured":"Barker B ( 2015) Message passing interface (mpi). In: Workshop: High Performance Computing on Stampede, vol. 262"},{"issue":"1","key":"4491_CR18","doi-asserted-by":"publisher","first-page":"46","DOI":"10.1109\/99.660313","volume":"5","author":"L Dagum","year":"1998","unstructured":"Dagum L, Menon R (1998) Openmp: an industry standard API for shared-memory programming. IEEE Comput Sci Eng 5(1):46\u201355","journal-title":"IEEE Comput Sci Eng"},{"issue":"2","key":"4491_CR19","doi-asserted-by":"publisher","first-page":"216","DOI":"10.1109\/JPROC.2004.840301","volume":"93","author":"M Frigo","year":"2005","unstructured":"Frigo M, Johnson SG (2005) The design and implementation of fftw3. Proc IEEE 93(2):216\u2013231","journal-title":"Proc IEEE"},{"key":"4491_CR20","doi-asserted-by":"crossref","unstructured":"Luszczek PR, Bailey DH, Dongarra JJ, Kepner J, Lucas RF, Rabenseifner R, Takahashi D ( 2006) The hpc challenge (hpcc) benchmark suite. In: Proceedings of the 2006 ACM\/IEEE Conference on Supercomputing, vol. 213, pp. 1188455\u2013 1188677","DOI":"10.1145\/1188455.1188677"},{"issue":"10","key":"4491_CR21","doi-asserted-by":"publisher","first-page":"2595","DOI":"10.1109\/TPDS.2013.222","volume":"25","author":"H Wang","year":"2013","unstructured":"Wang H, Potluri S, Bureddy D, Rosales C, Panda DK (2013) Gpu-aware mpi on rdma-enabled clusters: Design, implementation and evaluation. IEEE Trans Parallel Distrib Syst 25(10):2595\u20132605","journal-title":"IEEE Trans Parallel Distrib Syst"},{"key":"4491_CR22","unstructured":"Schroeder TC ( 2011) Peer-to-peer & unified virtual addressing. In: GPU Technology Conference, NVIDIA"},{"key":"4491_CR23","doi-asserted-by":"crossref","unstructured":"Potluri S, Wang H, Bureddy D, Singh AK, Rosales C, Panda DK ( 2012) Optimizing mpi communication on multi-gpu systems using cuda inter-process communication. In: 2012 IEEE 26th International Parallel and Distributed Processing Symposium Workshops & PhD Forum, pp. 1848\u2013 1857 IEEE","DOI":"10.1109\/IPDPSW.2012.228"},{"key":"4491_CR24","unstructured":"ROCmSoftwarePlatform(2018) ROCmSoftwarePlatform\/RCCL: ROCM Communication Collectives Library (RCCL) . https:\/\/github.com\/ROCmSoftwarePlatform\/rccl"},{"key":"4491_CR25","doi-asserted-by":"crossref","unstructured":"Sunitha N, Raju K, Chiplunkar N.N (2017) Performance improvement of cuda applications by reducing cpu-gpu data transfer overhead. In: 2017 international conference on inventive communication and computational technologies (ICICCT), pp 211\u2013 215 . IEEE","DOI":"10.1109\/ICICCT.2017.7975190"},{"issue":"5","key":"4491_CR26","doi-asserted-by":"publisher","first-page":"876","DOI":"10.1007\/s10766-015-0366-5","volume":"43","author":"JL Jodra","year":"2015","unstructured":"Jodra JL, Gurrutxaga I, Muguerza J (2015) Efficient 3d transpositions in graphics processing units. Int J Parallel Prog 43(5):876\u2013891","journal-title":"Int J Parallel Prog"},{"key":"4491_CR27","first-page":"1","volume":"18","author":"G Ruetsch","year":"2009","unstructured":"Ruetsch G, Micikevicius P (2009) Optimizing matrix transpose in Cuda. Nvidia CUDA SDK Appl Note 18:1","journal-title":"Nvidia CUDA SDK Appl Note"},{"key":"4491_CR28","unstructured":"AMD (2021) AMD INSTINCT$$^{\\rm TM}$$ MI100 accelerator | data center GPU | AMD . https:\/\/www.amd.com\/en\/products\/server-accelerators\/instinct-mi100"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-04491-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-022-04491-7\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-022-04491-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,4]],"date-time":"2022-10-04T10:19:19Z","timestamp":1664878759000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-022-04491-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,5,13]]},"references-count":28,"journal-issue":{"issue":"15","published-print":{"date-parts":[[2022,10]]}},"alternative-id":["4491"],"URL":"https:\/\/doi.org\/10.1007\/s11227-022-04491-7","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,5,13]]},"assertion":[{"value":"30 March 2022","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 May 2022","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}