{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T11:00:59Z","timestamp":1777546859437,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":41,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,6,25]],"date-time":"2019-06-25T00:00:00Z","timestamp":1561420800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT)","award":["No. 2018R1A5A1060031"],"award-info":[{"award-number":["No. 2018R1A5A1060031"]}]},{"name":"Basic Science Research Program through the National Research Foundation of Korea (NRF) funded by the Ministry of Science, ICT and Future Planning","award":["2017R1E1A1A01077630"],"award-info":[{"award-number":["2017R1E1A1A01077630"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,6,25]]},"DOI":"10.1145\/3299869.3319865","type":"proceedings-article","created":{"date-parts":[[2019,6,18]],"date-time":"2019-06-18T17:41:43Z","timestamp":1560879703000},"page":"759-774","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["DistME"],"prefix":"10.1145","author":[{"given":"Donghyoung","family":"Han","sequence":"first","affiliation":[{"name":"Daegu Gyeongbuk Institute of Science &amp; Technology (DGIST), Daegu, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yoon-Min","family":"Nam","sequence":"additional","affiliation":[{"name":"Daegu Gyeongbuk Institute of Science &amp; Technology (DGIST), Daegu, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jihye","family":"Lee","sequence":"additional","affiliation":[{"name":"Daegu Gyeongbuk Institute of Science &amp; Technology (DGIST), Daegu, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kyongseok","family":"Park","sequence":"additional","affiliation":[{"name":"Korea Institute of Science and Technology Information (KISTI), Daejeon, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hyunwoo","family":"Kim","sequence":"additional","affiliation":[{"name":"Korea Institute of Science and Technology Information (KISTI), Daejeon, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Min-Soo","family":"Kim","sequence":"additional","affiliation":[{"name":"Daegu Gyeongbuk Institute of Science &amp; Technology (DGIST), Daegu, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,6,25]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"et almbox","author":"Anderson E.","year":"1999","unstructured":"E. Anderson , Z. Bai , C. Bischof , L. S. Blackford , J. Demmel , J. Dongarra , J. Du Croz , A. Greenbaum , S. Hammarling , A. McKenney , et almbox . 1999 . LAPACK Users' guide .SIAM. E. Anderson, Z. Bai, C. Bischof, L. S. Blackford, J. Demmel, J. Dongarra, J. Du Croz, A. Greenbaum, S. Hammarling, A. McKenney, et almbox. 1999. LAPACK Users' guide .SIAM."},{"key":"e_1_3_2_1_2_1","volume-title":"http:\/\/hadoop.apache.org\/ Retrieved","year":"2018","unstructured":"Apache. 2011. Hadoop. http:\/\/hadoop.apache.org\/ Retrieved May 2, 2018 from Apache. 2011. Hadoop. http:\/\/hadoop.apache.org\/ Retrieved May 2, 2018 from"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2742797"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1562764.1562783"},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"crossref","unstructured":"S. Blanas J. M. Patel V. Ercegovac J. Rao E. J. Shekita and Y. Tian. 2010. A comparison of join algorithms for log processing in mapreduce. In SIGMOD. ACM 975--986.  S. Blanas J. M. Patel V. Ercegovac J. Rao E. J. Shekita and Y. Tian. 2010. A comparison of join algorithms for log processing in mapreduce. In SIGMOD. ACM 975--986.","DOI":"10.1145\/1807167.1807273"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.14778\/3007263.3007279"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jpdc.2012.04.003"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1807167.1807271"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/582034.582062"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137633"},{"key":"e_1_3_2_1_11_1","unstructured":"S. Chetlur C. Woolley P. Vandermersch J. Cohen J. Tran B. Catanzaro and E. Shelhamer. 2014. cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 (2014).  S. Chetlur C. Woolley P. Vandermersch J. Cohen J. Tran B. Catanzaro and E. Shelhamer. 2014. cudnn: Efficient primitives for deep learning. arXiv preprint arXiv:1410.0759 (2014)."},{"key":"e_1_3_2_1_12_1","volume-title":"Frontiers of Massively Parallel Computation, 1992., Fourth Symposium on the. IEEE, 120--127","author":"Choi J.","unstructured":"J. Choi , J. J. Dongarra , R. Pozo , and D. W. Walker . 1992. ScaLAPACK: A scalable linear algebra library for distributed memory concurrent computers . In Frontiers of Massively Parallel Computation, 1992., Fourth Symposium on the. IEEE, 120--127 . J. Choi, J. J. Dongarra, R. Pozo, and D. W. Walker. 1992. ScaLAPACK: A scalable linear algebra library for distributed memory concurrent computers. In Frontiers of Massively Parallel Computation, 1992., Fourth Symposium on the. IEEE, 120--127."},{"key":"e_1_3_2_1_13_1","volume-title":"Nvidia Multi-Process Service. https:\/\/docs.nvidia.com\/deploy\/mps\/index.html Retrieved","author":"NVIDIA Corporation","year":"2018","unstructured":"NVIDIA Corporation . 2012. Nvidia Multi-Process Service. https:\/\/docs.nvidia.com\/deploy\/mps\/index.html Retrieved May 2, 2018 from NVIDIA Corporation. 2012. Nvidia Multi-Process Service. https:\/\/docs.nvidia.com\/deploy\/mps\/index.html Retrieved May 2, 2018 from"},{"key":"e_1_3_2_1_14_1","volume-title":"cuBLAS. https:\/\/docs.nvidia.com\/cuda\/cublas\/index.html Retrieved","author":"NVIDIA Corporation","year":"2018","unstructured":"NVIDIA Corporation . 2015. cuBLAS. https:\/\/docs.nvidia.com\/cuda\/cublas\/index.html Retrieved May 2, 2018 from NVIDIA Corporation. 2015. cuBLAS. https:\/\/docs.nvidia.com\/cuda\/cublas\/index.html Retrieved May 2, 2018 from"},{"key":"e_1_3_2_1_15_1","volume-title":"Direct methods for sparse linear systems","author":"Davis T. A.","unstructured":"T. A. Davis . 2006. Direct methods for sparse linear systems . Vol. 2 . Siam . T. A. Davis. 2006. Direct methods for sparse linear systems . Vol. 2. Siam."},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/IPDPS.2013.80"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3035937"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDE.2011.5767930"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPDS.2017.2686384"},{"key":"e_1_3_2_1_20_1","first-page":"19","article-title":"The movielens datasets: History and context","volume":"5","author":"Harper F. M.","year":"2016","unstructured":"F. M. Harper and J. A. Konstan . 2016 . The movielens datasets: History and context . ACM TIIS , Vol. 5 , 4 (2016), 19 . F. M. Harper and J. A. Konstan. 2016. The movielens datasets: History and context. ACM TIIS, Vol. 5, 4 (2016), 19.","journal-title":"ACM TIIS"},{"key":"e_1_3_2_1_21_1","unstructured":"M. Kabiljo and A. Ilic. 2015. Recommending items to more than a billion people. https:\/\/code.fb.com\/core-data\/recommending-items-to-more-than-a-billion-people Retrieved May 2 2018 from  M. Kabiljo and A. Ilic. 2015. Recommending items to more than a billion people. https:\/\/code.fb.com\/core-data\/recommending-items-to-more-than-a-billion-people Retrieved May 2 2018 from"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2882903.2915204"},{"key":"e_1_3_2_1_23_1","unstructured":"D. D. Lee and H. S. Seung. 2001. Algorithms for non-negative matrix factorization. In NIPS. 556--562.   D. D. Lee and H. S. Seung. 2001. Algorithms for non-negative matrix factorization. In NIPS. 556--562."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"crossref","unstructured":"S. Luo Z. Gao M. Gubanov L. L. Perez and C. Jermaine. 2018. Scalable linear algebra on a relational database system. IEEE TKDE (2018).  S. Luo Z. Gao M. Gubanov L. L. Perez and C. Jermaine. 2018. Scalable linear algebra on a relational database system. IEEE TKDE (2018).","DOI":"10.1109\/ICDE.2017.108"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"crossref","unstructured":"S. Luo Z. J. Gao M. Gubanov L. L. Perez and C. Jermaine. 2017. Scalable Linear Algebra on a Relational Database System. In ICDE. IEEE 523--534.  S. Luo Z. J. Gao M. Gubanov L. L. Perez and C. Jermaine. 2017. Scalable Linear Algebra on a Relational Database System. In ICDE. IEEE 523--534.","DOI":"10.1109\/ICDE.2017.108"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.5555\/2946645.2946679"},{"key":"e_1_3_2_1_27_1","volume-title":"GPU Technology Conference .","author":"Naumov M.","unstructured":"M. Naumov , LS. Chien , P. Vandermersch , and U. Kapasi . 2010. cuSPARSE library . In GPU Technology Conference . M. Naumov, LS. Chien, P. Vandermersch, and U. Kapasi. 2010. cuSPARSE library. In GPU Technology Conference ."},{"key":"e_1_3_2_1_28_1","volume-title":"R: A Language and Environment for Statistical Computing","author":"Team R Core","year":"2014","unstructured":"R Core Team . 2014 . R: A Language and Environment for Statistical Computing . R Foundation for Statistical Computing , Vienna, Austria . http:\/\/www.R-project.org\/ R Core Team. 2014. R: A Language and Environment for Statistical Computing . R Foundation for Statistical Computing, Vienna, Austria. http:\/\/www.R-project.org\/"},{"key":"e_1_3_2_1_29_1","volume-title":"NIPS Workshop MLSystems .","author":"Schelter S.","unstructured":"S. Schelter , A. Palumbo , S. Quinn , S. Marthi , and A. Musselman . 2016. Samsara: Declarative machine learning on distributed dataflow systems . In NIPS Workshop MLSystems . S. Schelter, A. Palumbo, S. Quinn, S. Marthi, and A. Musselman. 2016. Samsara: Declarative machine learning on distributed dataflow systems. In NIPS Workshop MLSystems ."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CloudCom.2010.17"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"crossref","unstructured":"M. Stonebraker P. Brown A. Poliakov and S. Raman. 2011. The architecture of SciDB. In SSDBM. Springer 1--16.   M. Stonebraker P. Brown A. Poliakov and S. Raman. 2011. The architecture of SciDB. In SSDBM. Springer 1--16.","DOI":"10.1007\/978-3-642-22351-8_1"},{"key":"e_1_3_2_1_32_1","volume-title":"Intel math kernel library developer reference. https:\/\/software.intel.com\/en-us\/articles\/mkl-reference-manual Retrieved","author":"Development Team MKL","year":"2018","unstructured":"MKL Development Team . 2015. Intel math kernel library developer reference. https:\/\/software.intel.com\/en-us\/articles\/mkl-reference-manual Retrieved Oct , 2018 from MKL Development Team. 2015. Intel math kernel library developer reference. https:\/\/software.intel.com\/en-us\/articles\/mkl-reference-manual Retrieved Oct, 2018 from"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.14778\/3275366.3284963"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1002\/(SICI)1096-9128(199704)9:4<255::AID-CPE250>3.0.CO;2-2"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"crossref","unstructured":"R. C. Whaley and J. J. Dongarra. 1998. Automatically tuned linear algebra software. In SC. IEEE 38--38.   R. C. Whaley and J. J. Dongarra. 1998. Automatically tuned linear algebra software. In SC. IEEE 38--38.","DOI":"10.1109\/SC.1998.10004"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-03869-3_82"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"crossref","unstructured":"L. Yu Y. Shao and B. Cui. 2015. Exploiting matrix dependency for efficient distributed matrix computation. In SIGMOD. ACM 93--105.  L. Yu Y. Shao and B. Cui. 2015. Exploiting matrix dependency for efficient distributed matrix computation. In SIGMOD. ACM 93--105.","DOI":"10.1145\/2723372.2723712"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"crossref","unstructured":"Y. Yu M. Tang W. G. Aref Q. M. Malluhi M. M. Abbas and M. Ouzzani. 2017. In-Memory Distributed Matrix Computation Processing and Optimization. In ICDE. IEEE 1047--1058.  Y. Yu M. Tang W. G. Aref Q. M. Malluhi M. M. Abbas and M. Ouzzani. 2017. In-Memory Distributed Matrix Computation Processing and Optimization. In ICDE. IEEE 1047--1058.","DOI":"10.1109\/ICDE.2017.150"},{"key":"e_1_3_2_1_39_1","unstructured":"M. Zaharia M. Chowdhury T. Das A. Dave J. Ma M. McCauley M. J. Franklin S. Shenker and I. Stoica. 2012. Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. In NSDI. USENIX Association 2--2.   M. Zaharia M. Chowdhury T. Das A. Dave J. Ma M. McCauley M. J. Franklin S. Shenker and I. Stoica. 2012. Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing. In NSDI. USENIX Association 2--2."},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934664"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-68880-8_32"}],"event":{"name":"SIGMOD\/PODS '19: International Conference on Management of Data","location":"Amsterdam Netherlands","acronym":"SIGMOD\/PODS '19","sponsor":["SIGMOD ACM Special Interest Group on Management of Data"]},"container-title":["Proceedings of the 2019 International Conference on Management of Data"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3299869.3319865","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3299869.3319865","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:02:16Z","timestamp":1750208536000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3299869.3319865"}},"subtitle":["A Fast and Elastic Distributed Matrix Computation Engine using GPUs"],"short-title":[],"issued":{"date-parts":[[2019,6,25]]},"references-count":41,"alternative-id":["10.1145\/3299869.3319865","10.1145\/3299869"],"URL":"https:\/\/doi.org\/10.1145\/3299869.3319865","relation":{},"subject":[],"published":{"date-parts":[[2019,6,25]]},"assertion":[{"value":"2019-06-25","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}