{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T15:50:28Z","timestamp":1784217028504,"version":"3.55.0"},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2020,6,1]],"date-time":"2020-06-01T00:00:00Z","timestamp":1590969600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"U.S. Department of Energy, Office of Science, Office of Advanced Scientific Computing Research, Applied Mathematics program"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Math. Softw."],"published-print":{"date-parts":[[2020,6,30]]},"abstract":"<jats:p>Our goal is compression of massive-scale grid-structured data, such as the multi-terabyte output of a high-fidelity computational simulation. For such data sets, we have developed a new software package called TuckerMPI, a parallel C++\/MPI software package for compressing distributed data. The approach is based on treating the data as a tensor, i.e., a multidimensional array, and computing its truncated Tucker decomposition, a higher-order analogue to the truncated singular value decomposition of a matrix. The result is a low-rank approximation of the original tensor-structured data. Compression efficiency is achieved by detecting latent global structure within the data, which we contrast to most compression methods that are focused on local structure. In this work, we describe TuckerMPI, our implementation of the truncated Tucker decomposition, including details of the data distribution and in-memory layouts, the parallel and serial implementations of the key kernels, and analysis of the storage, communication, and computational costs. We test the software on 4.5 and 6.7 terabyte data sets distributed across 100 s of nodes (1,000 s of MPI processes), achieving compression ratios between 100 and 200,000\u00d7, which equates to 99--99.999% compression (depending on the desired accuracy) in substantially less time than it would take to even read the same dataset from a parallel file system. Moreover, we show that our method also allows for reconstruction of partial or down-sampled data on a single node, without a parallel computer so long as the reconstructed portion is small enough to fit on a single machine, e.g., in the instance of reconstructing\/visualizing a single down-sampled time step or computing summary statistics. The code is available at https:\/\/gitlab.com\/tensors\/TuckerMPI.<\/jats:p>","DOI":"10.1145\/3378445","type":"journal-article","created":{"date-parts":[[2020,6,1]],"date-time":"2020-06-01T16:12:35Z","timestamp":1591027955000},"page":"1-31","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":44,"title":["TuckerMPI"],"prefix":"10.1145","volume":"46","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1557-8027","authenticated-orcid":false,"given":"Grey","family":"Ballard","sequence":"first","affiliation":[{"name":"Wake Forest University, Winston-Salem, NC"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Alicia","family":"Klinvex","sequence":"additional","affiliation":[{"name":"Sandia National Laboratories, Livermore, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tamara G.","family":"Kolda","sequence":"additional","affiliation":[{"name":"Sandia National Laboratories, Livermore, CA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,6]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of the American Control Conference. 147--152","author":"Afra S.","year":"2014","unstructured":"S. Afra , E. Gildin , and M. Tarrahi . 2014. Heterogeneous reservoir characterization using efficient parameterization through higher order SVD (HOSVD) . In Proceedings of the American Control Conference. 147--152 . DOI:https:\/\/doi.org\/10.1109\/ACC. 2014 .6859246 10.1109\/ACC.2014.6859246 S. Afra, E. Gildin, and M. Tarrahi. 2014. Heterogeneous reservoir characterization using efficient parameterization through higher order SVD (HOSVD). In Proceedings of the American Control Conference. 147--152. DOI:https:\/\/doi.org\/10.1109\/ACC.2014.6859246"},{"key":"e_1_2_1_2_1","volume-title":"Proceedings of the 30th IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201916)","author":"Austin Woody","year":"2016","unstructured":"Woody Austin , Grey Ballard , and Tamara G. Kolda . 2016. Parallel tensor compression for large-scale scientific data . In Proceedings of the 30th IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201916) . 912--922. DOI:https:\/\/doi.org\/10.1109\/IPDPS. 2016 .67 arXiv:1510.06689 10.1109\/IPDPS.2016.67 Woody Austin, Grey Ballard, and Tamara G. Kolda. 2016. Parallel tensor compression for large-scale scientific data. In Proceedings of the 30th IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201916). 912--922. DOI:https:\/\/doi.org\/10.1109\/IPDPS.2016.67 arXiv:1510.06689"},{"key":"e_1_2_1_4_1","volume-title":"TTHRESH: Tensor compression for multidimensional visual data","author":"Ballester-Ripoll Rafael","year":"2019","unstructured":"Rafael Ballester-Ripoll , Peter Lindstrom , and Renato Pajarola . 2019 . TTHRESH: Tensor compression for multidimensional visual data . IEEE Trans. Visual. Comput. Graph . (2019). DOI:https:\/\/doi.org\/10.1109\/TVCG.2019.2904063 10.1109\/TVCG.2019.2904063 Rafael Ballester-Ripoll, Peter Lindstrom, and Renato Pajarola. 2019. TTHRESH: Tensor compression for multidimensional visual data. IEEE Trans. Visual. Comput. Graph. (2019). DOI:https:\/\/doi.org\/10.1109\/TVCG.2019.2904063"},{"key":"e_1_2_1_5_1","volume-title":"Lossy","author":"Ballester-Ripoll Rafael","year":"2015","unstructured":"Rafael Ballester-Ripoll and Renato Pajarola . 2015. Lossy volume compression using Tucker truncation and thresholding. Vis. Comput. 32 ( May 2015 ), 1433--1446. DOI:https:\/\/doi.org\/10.1007\/s00371-015-1130-y 10.1007\/s00371-015-1130-y Rafael Ballester-Ripoll and Renato Pajarola. 2015. Lossy volume compression using Tucker truncation and thresholding. Vis. Comput. 32 (May 2015), 1433--1446. DOI:https:\/\/doi.org\/10.1007\/s00371-015-1130-y"},{"key":"e_1_2_1_6_1","volume-title":"Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201917)","author":"Chakaravarthy V. T.","year":"2017","unstructured":"V. T. Chakaravarthy , J. W. Choi , D. J. Joseph , X. Liu , P. Murali , Y. Sabharwal , and D. Sreedhar . 2017. On optimizing distributed Tucker decomposition for dense tensors . In Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201917) . 1038--1047. DOI:https:\/\/doi.org\/10.1109\/IPDPS. 2017 .86 10.1109\/IPDPS.2017.86 V. T. Chakaravarthy, J. W. Choi, D. J. Joseph, X. Liu, P. Murali, Y. Sabharwal, and D. Sreedhar. 2017. On optimizing distributed Tucker decomposition for dense tensors. In Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201917). 1038--1047. DOI:https:\/\/doi.org\/10.1109\/IPDPS.2017.86"},{"key":"e_1_2_1_7_1","unstructured":"Venkatesan Chakravarthy. 2017. Personal communication.  Venkatesan Chakravarthy. 2017. Personal communication."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.5555\/1285358.1285359"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1088\/1749-4699\/2\/1\/015001"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/SC.2018.00045"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1137\/S0895479896305696"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201916)","author":"Di S.","year":"2016","unstructured":"S. Di and F. Cappello . 2016. Fast error-bounded lossy HPC data compression with SZ . In Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201916) . 730--739. DOI:https:\/\/doi.org\/10.1109\/IPDPS. 2016 .11 10.1109\/IPDPS.2016.11 S. Di and F. Cappello. 2016. Fast error-bounded lossy HPC data compression with SZ. In Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS\u201916). 730--739. DOI:https:\/\/doi.org\/10.1109\/IPDPS.2016.11"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/1066677.1066953"},{"key":"#cr-split#-e_1_2_1_14_1.1","doi-asserted-by":"crossref","unstructured":"A. Garc\u00eda-Magari\u00f1o S. Sor and A. Velazquez. 2016. Data reduction method for droplet deformation experiments based on high order singular value decomposition. Exper. Therm. Fluid Sci. 79 (Dec. 2016) 13--24. DOI:https:\/\/doi.org\/10.1016\/j.expthermflusci.2016.06.017 10.1016\/j.expthermflusci.2016.06.017","DOI":"10.1016\/j.expthermflusci.2016.06.017"},{"key":"#cr-split#-e_1_2_1_14_1.2","doi-asserted-by":"crossref","unstructured":"A. Garc\u00eda-Magari\u00f1o S. Sor and A. Velazquez. 2016. Data reduction method for droplet deformation experiments based on high order singular value decomposition. Exper. Therm. Fluid Sci. 79 (Dec. 2016) 13--24. DOI:https:\/\/doi.org\/10.1016\/j.expthermflusci.2016.06.017","DOI":"10.1016\/j.expthermflusci.2016.06.017"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1017\/S0962492914000087"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jcp.2012.02.007"},{"key":"e_1_2_1_17_1","volume-title":"Proceedings of the 23rd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP\u201918)","author":"Hayashi Koby","unstructured":"Koby Hayashi , Grey Ballard , Yujie Jiang , and Michael J. Tobia . 2018. Shared-memory parallelization of MTTKRP for dense tensors . In Proceedings of the 23rd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP\u201918) . ACM, New York, NY, 393--394. DOI:https:\/\/doi.org\/10.1145\/3178487.3178522 10.1145\/3178487.3178522 Koby Hayashi, Grey Ballard, Yujie Jiang, and Michael J. Tobia. 2018. Shared-memory parallelization of MTTKRP for dense tensors. In Proceedings of the 23rd ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP\u201918). ACM, New York, NY, 393--394. DOI:https:\/\/doi.org\/10.1145\/3178487.3178522"},{"key":"e_1_2_1_18_1","first-page":"2","article-title":"Compression of hyperspectral images using discrete wavelet transform and Tucker decomposition","volume":"5","author":"Karami A.","year":"2012","unstructured":"A. Karami , M. Yazdi , and G. Mercier . 2012 . Compression of hyperspectral images using discrete wavelet transform and Tucker decomposition . IEEE J. Select. Top. Appl. Earth Observ. Remote Sens. 5 , 2 (April 2012), 444--450. DOI:https:\/\/doi.org\/10.1109\/JSTARS.2012.2189200 10.1109\/JSTARS.2012.2189200 A. Karami, M. Yazdi, and G. Mercier. 2012. Compression of hyperspectral images using discrete wavelet transform and Tucker decomposition. IEEE J. Select. Top. Appl. Earth Observ. Remote Sens. 5, 2 (April 2012), 444--450. DOI:https:\/\/doi.org\/10.1109\/JSTARS.2012.2189200","journal-title":"IEEE J. Select. Top. Appl. Earth Observ. Remote Sens."},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the 45th International Conference on Parallel Processing (ICPP\u201916)","author":"Kaya O.","year":"2016","unstructured":"O. Kaya and B. U\u00e7ar . 2016. High performance parallel algorithms for the Tucker decomposition of sparse tensors . In Proceedings of the 45th International Conference on Parallel Processing (ICPP\u201916) . 103--112. DOI:https:\/\/doi.org\/10.1109\/ICPP. 2016 .19 10.1109\/ICPP.2016.19 O. Kaya and B. U\u00e7ar. 2016. High performance parallel algorithms for the Tucker decomposition of sparse tensors. In Proceedings of the 45th International Conference on Parallel Processing (ICPP\u201916). 103--112. DOI:https:\/\/doi.org\/10.1109\/ICPP.2016.19"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1137\/07070111X"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1080\/00102202.2016.1197211"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807591.2807671"},{"key":"e_1_2_1_23_1","volume-title":"Proceedings of the IEEE 5th Symposium on Large Data Analysis and Visualization (LDAV\u201915)","author":"Li S.","year":"2015","unstructured":"S. Li , K. Gruchalla , K. Potter , J. Clyne , and H. Childs . 2015. Evaluating the efficacy of wavelet configurations on turbulent-flow data . In Proceedings of the IEEE 5th Symposium on Large Data Analysis and Visualization (LDAV\u201915) . 81--89. DOI:https:\/\/doi.org\/10.1109\/LDAV. 2015 .7348075 10.1109\/LDAV.2015.7348075 S. Li, K. Gruchalla, K. Potter, J. Clyne, and H. Childs. 2015. Evaluating the efficacy of wavelet configurations on turbulent-flow data. In Proceedings of the IEEE 5th Symposium on Large Data Analysis and Visualization (LDAV\u201915). 81--89. DOI:https:\/\/doi.org\/10.1109\/LDAV.2015.7348075"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2014.2346458"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.combustflame.2014.10.014"},{"key":"e_1_2_1_27_1","volume-title":"Proceedings of the 34th IEEE International Conference on Data Engineering (ICDE\u201918)","author":"Oh S.","year":"2018","unstructured":"S. Oh , N. Park , S. Lee , and U. Kang . 2018. Scalable Tucker factorization for sparse tensors\u2014Algorithms and discoveries . In Proceedings of the 34th IEEE International Conference on Data Engineering (ICDE\u201918) . 1120--1131. DOI:https:\/\/doi.org\/10.1109\/ICDE. 2018 .00104 10.1109\/ICDE.2018.00104 S. Oh, N. Park, S. Lee, and U. Kang. 2018. Scalable Tucker factorization for sparse tensors\u2014Algorithms and discoveries. In Proceedings of the 34th IEEE International Conference on Data Engineering (ICDE\u201918). 1120--1131. DOI:https:\/\/doi.org\/10.1109\/ICDE.2018.00104"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSP.2013.2269903"},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the IEEE 30th International Parallel and Distributed Processing Symposium. 902--911","author":"Smith Shaden","year":"2016","unstructured":"Shaden Smith and George Karypis . 2016 . A medium-grained algorithm for distributed sparse tensor factorization . In Proceedings of the IEEE 30th International Parallel and Distributed Processing Symposium. 902--911 . DOI:https:\/\/doi.org\/10.1109\/IPDPS.2016.113 10.1109\/IPDPS.2016.113 Shaden Smith and George Karypis. 2016. A medium-grained algorithm for distributed sparse tensor factorization. In Proceedings of the IEEE 30th International Parallel and Distributed Processing Symposium. 902--911. DOI:https:\/\/doi.org\/10.1109\/IPDPS.2016.113"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-64203-1_47"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342005051521"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02289464"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1137\/110836067"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the 7th European Conference on Computer Vision (ECCV\u201902)","volume":"2350","author":"Vasilescu M. A. O.","unstructured":"M. A. O. Vasilescu and D. Terzopoulos . 2002. Multilinear analysis of image ensembles: TensorFaces . In Proceedings of the 7th European Conference on Computer Vision (ECCV\u201902) (Lecture Notes in Computer Science) , Vol. 2350 . Springer, 447--460. DOI:https:\/\/doi.org\/10.1007\/3-540-47969-4_30 10.1007\/3-540-47969-4_30 M. A. O. Vasilescu and D. Terzopoulos. 2002. Multilinear analysis of image ensembles: TensorFaces. In Proceedings of the 7th European Conference on Computer Vision (ECCV\u201902) (Lecture Notes in Computer Science), Vol. 2350. Springer, 447--460. DOI:https:\/\/doi.org\/10.1007\/3-540-47969-4_30"}],"container-title":["ACM Transactions on Mathematical Software"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3378445","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3378445","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:45:04Z","timestamp":1750203904000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3378445"}},"subtitle":["A Parallel C++\/MPI Software Package for Large-scale Data Compression via the Tucker Tensor Decomposition"],"short-title":[],"issued":{"date-parts":[[2020,6]]},"references-count":33,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2020,6,30]]}},"alternative-id":["10.1145\/3378445"],"URL":"https:\/\/doi.org\/10.1145\/3378445","relation":{},"ISSN":["0098-3500","1557-7295"],"issn-type":[{"value":"0098-3500","type":"print"},{"value":"1557-7295","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,6]]},"assertion":[{"value":"2019-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-01-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-06-01","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}