{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:22:05Z","timestamp":1750220525420,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":17,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,6,21]],"date-time":"2021-06-21T00:00:00Z","timestamp":1624233600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Office of Science of theU.S. Department of Energy","award":["DE-AC05-00OR22725"],"award-info":[{"award-number":["DE-AC05-00OR22725"]}]},{"name":"ational Science Founda-tion?s Major Research Instrumentation","award":["725729"],"award-info":[{"award-number":["725729"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,6,21]]},"DOI":"10.1145\/3431379.3460645","type":"proceedings-article","created":{"date-parts":[[2021,6,17]],"date-time":"2021-06-17T04:09:26Z","timestamp":1623902966000},"page":"95-106","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["TEMPI"],"prefix":"10.1145","author":[{"given":"Carl","family":"Pearson","sequence":"first","affiliation":[{"name":"University of Illinois at Urbana-Champaign, Urbana, IL, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kun","family":"Wu","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana-Champaign, Urbana, IL, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"I-Hsin","family":"Chung","sequence":"additional","affiliation":[{"name":"IBM T. J. Watson Research, Yorktown Heights, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinjun","family":"Xiong","sequence":"additional","affiliation":[{"name":"IBM T. J. Watson Research, Yorktown Heights, NY, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wen-Mei","family":"Hwu","sequence":"additional","affiliation":[{"name":"Nvidia Research, Champaign, IL, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,6,21]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/140901.140903"},{"volume-title":"Proceedings of the 25th European MPI Users' Group Meeting","author":"Bienz Amanda","key":"e_1_3_2_1_2_1","unstructured":"Amanda Bienz , William D. Gropp , and Luke N. Olson . 2018. Improving Performance Models for Irregular Point-to-Point Communication . In Proceedings of the 25th European MPI Users' Group Meeting ( Barcelona, Spain) (EuroMPI'18). Association for Computing Machinery, New York, NY, USA, Article 7, 8 pages. https:\/\/doi.org\/10.1145\/3236367.3236368 Amanda Bienz, William D. Gropp, and Luke N. Olson. 2018. Improving Performance Models for Irregular Point-to-Point Communication. In Proceedings of the 25th European MPI Users' Group Meeting (Barcelona, Spain) (EuroMPI'18). Association for Computing Machinery, New York, NY, USA, Article 7, 8 pages. https:\/\/doi.org\/10.1145\/3236367.3236368"},{"key":"e_1_3_2_1_3_1","volume-title":"Modeling Data Movement Performance on Heterogeneous Architectures. arxiv","author":"Bienz Amanda","year":"2010","unstructured":"Amanda Bienz , Luke N. Olson , William D. Gropp , and Shelby Lockhart . 2020. Modeling Data Movement Performance on Heterogeneous Architectures. arxiv : 2010 .10378 [cs.DC] Amanda Bienz, Luke N. Olson, William D. Gropp, and Shelby Lockhart. 2020. Modeling Data Movement Performance on Heterogeneous Architectures. arxiv: 2010.10378 [cs.DC]"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"crossref","unstructured":"C. Chu J. M. Hashmi K. S. Khorassani H. Subramoni and D. K. Panda. 2019. High-Performance Adaptive MPI Derived Datatype Communication for Modern Multi-GPU Systems. In 2019 IEEE 26th International Conference on High Performance Computing Data and Analytics (HiPC). 267--276. https:\/\/doi.org\/10.1109\/HiPC.2019.00041  C. Chu J. M. Hashmi K. S. Khorassani H. Subramoni and D. K. Panda. 2019. High-Performance Adaptive MPI Derived Datatype Communication for Modern Multi-GPU Systems. In 2019 IEEE 26th International Conference on High Performance Computing Data and Analytics (HiPC). 267--276. https:\/\/doi.org\/10.1109\/HiPC.2019.00041","DOI":"10.1109\/HiPC.2019.00041"},{"key":"e_1_3_2_1_5_1","volume-title":"Dynamic Kernel Fusion for Bulk Non-contiguous Data Transfer on GPU Clusters. In 2020 IEEE International Conference on Cluster Computing (CLUSTER). 130--141","author":"Chu C. H.","year":"2020","unstructured":"C. H. Chu , K. S. Khorassani , Q. Zhou , H. Subramoni , and D. K. Panda . 2020 . Dynamic Kernel Fusion for Bulk Non-contiguous Data Transfer on GPU Clusters. In 2020 IEEE International Conference on Cluster Computing (CLUSTER). 130--141 . https:\/\/doi.org\/10.1109\/CLUSTER49012. 2020 .00023 C. H. Chu, K. S. Khorassani, Q. Zhou, H. Subramoni, and D. K. Panda. 2020. Dynamic Kernel Fusion for Bulk Non-contiguous Data Transfer on GPU Clusters. In 2020 IEEE International Conference on Cluster Computing (CLUSTER). 130--141. https:\/\/doi.org\/10.1109\/CLUSTER49012.2020.00023"},{"key":"e_1_3_2_1_6_1","volume-title":"Squyres","author":"Graham Richard L.","year":"2006","unstructured":"Richard L. Graham , Timothy S. Woodall , and Jeffrey M . Squyres . 2006 . Open MPI: A Flexible High Performance MPI. In Parallel Processing and Applied Mathematics,, Roman Wyrzykowski, Jack Dongarra, Norbert Meyer, and Jerzy Wa's niewski (Eds.). Springer Berlin Heidelberg , Berlin, Heidelberg, 228--239. Richard L. Graham, Timothy S. Woodall, and Jeffrey M. Squyres. 2006. Open MPI: A Flexible High Performance MPI. In Parallel Processing and Applied Mathematics,, Roman Wyrzykowski, Jack Dongarra, Norbert Meyer, and Jerzy Wa's niewski (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 228--239."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/648139.749473"},{"key":"e_1_3_2_1_8_1","unstructured":"Jahanzeb Maqbool Hashmi Ching-Hsiang Chu Sourav Chakraborty Mohammadreza Bayatpour Hari Subramoni and Dhabaleswar K Panda. 2020. FALCON-X: Zero-copy MPI derived datatype processing on modern CPU and GPU architectures. J. Parallel and Distrib. Comput. (2020).  Jahanzeb Maqbool Hashmi Ching-Hsiang Chu Sourav Chakraborty Mohammadreza Bayatpour Hari Subramoni and Dhabaleswar K Panda. 2020. FALCON-X: Zero-copy MPI derived datatype processing on modern CPU and GPU architectures. J. Parallel and Distrib. Comput. (2020)."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2013.234"},{"key":"e_1_3_2_1_11_1","unstructured":"Programming Models and Runtime Systems Team. 2021. Yaksa. https:\/\/github.com\/pmodels\/yaksa  Programming Models and Runtime Systems Team. 2021. Yaksa. https:\/\/github.com\/pmodels\/yaksa"},{"key":"e_1_3_2_1_12_1","volume-title":"The MVAPICH project: Transforming research into high-performance MPI library for HPC community. Journal of Computational Science","author":"Panda Dhabaleswar Kumar","year":"2020","unstructured":"Dhabaleswar Kumar Panda , Hari Subramoni , Ching-Hsiang Chu , and Mohammadreza Bayatpour . 2020. The MVAPICH project: Transforming research into high-performance MPI library for HPC community. Journal of Computational Science ( 2020 ), 101208. https:\/\/doi.org\/10.1016\/j.jocs.2020.101208 Dhabaleswar Kumar Panda, Hari Subramoni, Ching-Hsiang Chu, and Mohammadreza Bayatpour. 2020. The MVAPICH project: Transforming research into high-performance MPI library for HPC community. Journal of Computational Science (2020), 101208. https:\/\/doi.org\/10.1016\/j.jocs.2020.101208"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cpc.2017.03.011"},{"volume-title":"Recent Advances in Parallel Virtual Machine and Message Passing Interface","author":"Ross Robert","key":"e_1_3_2_1_14_1","unstructured":"Robert Ross , Robert Latham , William Gropp , Ewing Lusk , and Rajeev Thakur . 2009. Processing MPI Datatypes Outside MPI . In Recent Advances in Parallel Virtual Machine and Message Passing Interface , Matti Ropo, Jan Westerholm, and Jack Dongarra (Eds.). Springer Berlin Heidelberg , Berlin, Heidelberg , 42--53. Robert Ross, Robert Latham, William Gropp, Ewing Lusk, and Rajeev Thakur. 2009. Processing MPI Datatypes Outside MPI. In Recent Advances in Parallel Virtual Machine and Message Passing Interface, Matti Ropo, Jan Westerholm, and Jack Dongarra (Eds.). Springer Berlin Heidelberg, Berlin, Heidelberg, 42--53."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPP.2014.31"},{"key":"e_1_3_2_1_16_1","volume-title":"MPI: A Message-Passing Interface Standard Version 3.1. Technical Report.","author":"MPI","year":"2015","unstructured":"MPI standards committee. 2015 . MPI: A Message-Passing Interface Standard Version 3.1. Technical Report. MPI standards committee. 2015. MPI: A Message-Passing Interface Standard Version 3.1. Technical Report."},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CLUSTER.2011.42"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2907294.2907317"}],"event":{"name":"HPDC '21: The 30th International Symposium on High-Performance Parallel and Distributed Computing","sponsor":["University of Arizona University of Arizona","SIGHPC ACM Special Interest Group on High Performance Computing, Special Interest Group on High Performance Computing","SIGARCH ACM Special Interest Group on Computer Architecture"],"location":"Virtual Event Sweden","acronym":"HPDC '21"},"container-title":["Proceedings of the 30th International Symposium on High-Performance Parallel and Distributed Computing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3431379.3460645","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3431379.3460645","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T21:24:46Z","timestamp":1750195486000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3431379.3460645"}},"subtitle":["An Interposed MPI Library with a Canonical Representation of CUDA-aware Datatypes"],"short-title":[],"issued":{"date-parts":[[2021,6,21]]},"references-count":17,"alternative-id":["10.1145\/3431379.3460645","10.1145\/3431379"],"URL":"https:\/\/doi.org\/10.1145\/3431379.3460645","relation":{},"subject":[],"published":{"date-parts":[[2021,6,21]]},"assertion":[{"value":"2021-06-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}