{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,1]],"date-time":"2026-01-01T10:10:53Z","timestamp":1767262253277,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":39,"publisher":"ACM","license":[{"start":{"date-parts":[[2018,8,13]],"date-time":"2018-08-13T00:00:00Z","timestamp":1534118400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2018,8,13]]},"DOI":"10.1145\/3225058.3225096","type":"proceedings-article","created":{"date-parts":[[2018,8,8]],"date-time":"2018-08-08T19:13:06Z","timestamp":1533755586000},"page":"1-10","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":9,"title":["Matrix Factorization on GPUs with Memory Optimization and Approximate Computing"],"prefix":"10.1145","author":[{"given":"Wei","family":"Tan","sequence":"first","affiliation":[{"name":"Citadel, Chicago, Illinois"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shiyu","family":"Chang","sequence":"additional","affiliation":[{"name":"IBM Research, Yorktown Heights, New York"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liana","family":"Fong","sequence":"additional","affiliation":[{"name":"IBM Research, Yorktown Heights, New York"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cheng","family":"Li","sequence":"additional","affiliation":[{"name":"University of Illinois at Urbana-Champaign, Urbana, Illinois"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zijun","family":"Wang","sequence":"additional","affiliation":[{"name":"IBM Research, Yorktown Heights, New York"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"LiangLiang","family":"Cao","sequence":"additional","affiliation":[{"name":"HelloVera.AI, New York, New York"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2018,8,13]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"Jimmy Ba and Rich Caruana. 2014. Do Deep Nets Really Need to be Deep?. In NIPS. 2654--2662. http:\/\/papers.nips.cc\/paper\/5484-do-deep-nets-really-need-to-be-deep.pdf   Jimmy Ba and Rich Caruana. 2014. Do Deep Nets Really Need to be Deep?. In NIPS. 2654--2662. http:\/\/papers.nips.cc\/paper\/5484-do-deep-nets-really-need-to-be-deep.pdf"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/BigData.2015.7363760"},{"volume-title":"A learning-rate schedule for stochastic gradient methods to matrix factorization","author":"Chin Wei-Sheng","key":"e_1_3_2_1_3_1","unstructured":"Wei-Sheng Chin , Yong Zhuang , Yu-Chin Juan , and Chih-Jen Lin . 2015. A learning-rate schedule for stochastic gradient methods to matrix factorization . In PAKDD. Springer . Wei-Sheng Chin, Yong Zhuang, Yu-Chin Juan, and Chih-Jen Lin. 2015. A learning-rate schedule for stochastic gradient methods to matrix factorization. In PAKDD. Springer."},{"key":"e_1_3_2_1_4_1","unstructured":"Adam Coates Brody Huval Tao Wang David Wu Bryan Catanzaro and Ng Andrew. 2013. Deep learning with COTS HPC systems. In ICML. 1337--1345.   Adam Coates Brody Huval Tao Wang David Wu Bryan Catanzaro and Ng Andrew. 2013. Deep learning with COTS HPC systems. In ICML. 1337--1345."},{"key":"e_1_3_2_1_5_1","volume-title":"Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing.","author":"Cui Henggang","year":"2014","unstructured":"Henggang Cui , James Cipar , Qirong Ho , Jin Kyu Kim , Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing. 2014 . Exploiting Bounded Staleness to Speed Up Big Data Analytics. In USENIX ATC. 37--48. Henggang Cui, James Cipar, Qirong Ho, Jin Kyu Kim, Seunghak Lee, Abhimanu Kumar, Jinliang Wei, Wei Dai, Gregory R. Ganger, Phillip B. Gibbons, Garth A. Gibson, and Eric P. Xing. 2014. Exploiting Bounded Staleness to Speed Up Big Data Analytics. In USENIX ATC. 37--48."},{"key":"e_1_3_2_1_6_1","unstructured":"Gideon Dror Noam Koenigstein Yehuda Koren and Markus Weimer. 2012. The Yahoo! Music Dataset and KDD-Cup '11. In KDD Cup 2011 competition.   Gideon Dror Noam Koenigstein Yehuda Koren and Markus Weimer. 2012. The Yahoo! Music Dataset and KDD-Cup '11. In KDD Cup 2011 competition."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.48"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/BigData.2015.7363811"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/2020408.2020426"},{"volume-title":"Computer architecture: a quantitative approach","author":"Hennessy John L","key":"e_1_3_2_1_10_1","unstructured":"John L Hennessy and David A Patterson . 2011. Computer architecture: a quantitative approach . Elsevier . John L Hennessy and David A Patterson. 2011. Computer architecture: a quantitative approach. Elsevier."},{"volume-title":"Methods of conjugate gradients for solving linear systems","author":"Hestenes Magnus Rudolph","key":"e_1_3_2_1_11_1","unstructured":"Magnus Rudolph Hestenes and Eduard Stiefel . 1952. Methods of conjugate gradients for solving linear systems . Vol. 49 . NBS. Magnus Rudolph Hestenes and Eduard Stiefel. 1952. Methods of conjugate gradients for solving linear systems. Vol. 49. NBS."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2008.22"},{"key":"e_1_3_2_1_13_1","volume-title":"Recommending items to more than a billion people. https:\/\/code.facebook.com\/posts\/861999383875667. (2015). {Online","author":"Kabiljo Maja","year":"2015","unstructured":"Maja Kabiljo and Aleksandar Ilic . 2015. Recommending items to more than a billion people. https:\/\/code.facebook.com\/posts\/861999383875667. (2015). {Online ; accessed 17- Aug- 2015 }. Maja Kabiljo and Aleksandar Ilic. 2015. Recommending items to more than a billion people. https:\/\/code.facebook.com\/posts\/861999383875667. (2015). {Online; accessed 17-Aug-2015}."},{"key":"e_1_3_2_1_14_1","unstructured":"David B Kirk and W Hwu Wen-mei. 2012. Programming massively parallel processors: a hands-on approach. Morgan Kaufmann.   David B Kirk and W Hwu Wen-mei. 2012. Programming massively parallel processors: a hands-on approach. Morgan Kaufmann."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2009.263"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2452376.2452449"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.14778\/2212351.2212354"},{"key":"e_1_3_2_1_18_1","article-title":"MLlib: Machine Learning in Apache Spark","volume":"17","author":"Meng Xiangrui","year":"2016","unstructured":"Xiangrui Meng , Joseph K. Bradley , Burak Yavuz , Evan R. Sparks , Shivaram Venkataraman , Davies Liu , Jeremy Freeman , D. B. Tsai , Manish Amde , Sean Owen , Doris Xin , Reynold Xin , Michael J. Franklin , Reza Zadeh , Matei Zaharia , and Ameet Talwalkar . 2016 . MLlib: Machine Learning in Apache Spark . Journal of Machine Learning Research 17 (2016), 34:1--34:7. http:\/\/jmlr.org\/papers\/v17\/15-237.html Xiangrui Meng, Joseph K. Bradley, Burak Yavuz, Evan R. Sparks, Shivaram Venkataraman, Davies Liu, Jeremy Freeman, D. B. Tsai, Manish Amde, Sean Owen, Doris Xin, Reynold Xin, Michael J. Franklin, Reza Zadeh, Matei Zaharia, and Ameet Talwalkar. 2016. MLlib: Machine Learning in Apache Spark. Journal of Machine Learning Research 17 (2016), 34:1--34:7. http:\/\/jmlr.org\/papers\/v17\/15-237.html","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2893356"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3038228.3038240"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Yusuke Nishioka and Kenjiro Taura. 2015. Scalable Task-Parallel SGD on Matrix Factorization in Multicore Architectures. In ParLearning.  Yusuke Nishioka and Kenjiro Taura. 2015. Scalable Task-Parallel SGD on Matrix Factorization in Multicore Architectures. In ParLearning.","DOI":"10.1109\/IPDPSW.2015.135"},{"key":"e_1_3_2_1_22_1","volume-title":"Wright","author":"Niu Feng","year":"2011","unstructured":"Feng Niu , Benjamin Recht , Christopher Re , and Stephen J . Wright . 2011 . HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent. In NIPS. 693--701. Feng Niu, Benjamin Recht, Christopher Re, and Stephen J. Wright. 2011. HOGWILD!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent. In NIPS. 693--701."},{"key":"e_1_3_2_1_23_1","unstructured":"Nvidia. 2015. cuBLAS. http:\/\/docs.nvidia.com\/cuda\/cublas\/. (2015). {Online; accessed 17-Aug-2015}.  Nvidia. 2015. cuBLAS. http:\/\/docs.nvidia.com\/cuda\/cublas\/. (2015). {Online; accessed 17-Aug-2015}."},{"key":"e_1_3_2_1_24_1","unstructured":"Nvidia. 2016. cuBLAS. http:\/\/docs.nvidia.com\/cuda\/cublas\/#cublas-lt-t-gt-gemmbatched. (2016). {Online; accessed 7-Nov-2016}.  Nvidia. 2016. cuBLAS. http:\/\/docs.nvidia.com\/cuda\/cublas\/#cublas-lt-t-gt-gemmbatched. (2016). {Online; accessed 7-Nov-2016}."},{"key":"e_1_3_2_1_25_1","volume-title":"http:\/\/www.nvidia.com\/object\/nvlink.html. (2016). {Online","author":"Link NVIDIA","year":"2016","unstructured":"Nvidia.2016. NVIDIA NV Link . http:\/\/www.nvidia.com\/object\/nvlink.html. (2016). {Online ; accessed 26- Nov- 2016 }. Nvidia.2016. NVIDIA NVLink. http:\/\/www.nvidia.com\/object\/nvlink.html. (2016). {Online; accessed 26-Nov-2016}."},{"volume-title":"Programming Tensor Cores in CUDA 9. https:\/\/devblogs.nvidia.com\/programming-tensor-cores-cuda-9\/. (2017). {Online","year":"2018","key":"e_1_3_2_1_26_1","unstructured":"Nvidia. 2017. Programming Tensor Cores in CUDA 9. https:\/\/devblogs.nvidia.com\/programming-tensor-cores-cuda-9\/. (2017). {Online ; accessed 21- Jan- 2018 }. Nvidia. 2017. Programming Tensor Cores in CUDA 9. https:\/\/devblogs.nvidia.com\/programming-tensor-cores-cuda-9\/. (2017). {Online; accessed 21-Jan-2018}."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2783258.2783322"},{"key":"e_1_3_2_1_28_1","volume-title":"Glove: Global vectors for word representation. In EMNLP. 1532--1543.","author":"Pennington Jeffrey","year":"2014","unstructured":"Jeffrey Pennington , Richard Socher , and Christopher D Manning . 2014 . Glove: Global vectors for word representation. In EMNLP. 1532--1543. Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014. Glove: Global vectors for word representation. In EMNLP. 1532--1543."},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1864708.1864726"},{"key":"e_1_3_2_1_30_1","volume-title":"NIPS Workshop on Distributed Matrix Computations.","author":"Schelter Sebastian","year":"2014","unstructured":"Sebastian Schelter , Venu Satuluri , and Reza Bosagh Zadeh . 2014 . Factorbird-a Parameter Server Approach to Distributed Matrix Factorization . In NIPS Workshop on Distributed Matrix Computations. Sebastian Schelter, Venu Satuluri, and Reza Bosagh Zadeh. 2014. Factorbird-a Parameter Server Approach to Distributed Matrix Factorization. In NIPS Workshop on Distributed Matrix Computations."},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2907294.2907297"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2012.120"},{"key":"e_1_3_2_1_33_1","unstructured":"Aaron van den Oord Sander Dieleman and Benjamin Schrauwen. 2013. Deep content-based music recommendation. In NIPS. 2643--2651.   Aaron van den Oord Sander Dieleman and Benjamin Schrauwen. 2013. Deep content-based music recommendation. In NIPS. 2643--2651."},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/1498765.1498785"},{"key":"e_1_3_2_1_35_1","unstructured":"Xiaolong Xie Wei Tan Liana L Fong and Yun Liang. 2017. CuMF_SGD: Fast and Scalable Matrix Factorization. In HPDC.  Xiaolong Xie Wei Tan Liana L Fong and Yun Liang. 2017. CuMF_SGD: Fast and Scalable Matrix Factorization. In HPDC."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2012.168"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.14778\/2732967.2732973"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-68880-8_32"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/2507157.2507164"}],"event":{"name":"ICPP 2018: 47th International Conference on Parallel Processing","sponsor":["University of Oregon University of Oregon"],"location":"Eugene OR USA","acronym":"ICPP 2018"},"container-title":["Proceedings of the 47th International Conference on Parallel Processing"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3225058.3225096","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3225058.3225096","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T01:39:07Z","timestamp":1750210747000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3225058.3225096"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,8,13]]},"references-count":39,"alternative-id":["10.1145\/3225058.3225096","10.1145\/3225058"],"URL":"https:\/\/doi.org\/10.1145\/3225058.3225096","relation":{},"subject":[],"published":{"date-parts":[[2018,8,13]]},"assertion":[{"value":"2018-08-13","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}