{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,9]],"date-time":"2026-04-09T13:28:49Z","timestamp":1775741329589,"version":"3.50.1"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"11","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2017,8]]},"abstract":"<jats:p>\n            We present LDA*, a system that has been deployed in one of the largest Internet companies to fulfil their requirements of\n            <jats:italic>\"topic modeling as an internal service\"<\/jats:italic>\n            ---relying on thousands of machines, engineers in different sectors submit their data, some are as large as 1.8TB, to LDA* and get results back in hours. LDA* is motivated by the observation that\n            <jats:italic>none of the existing topic modeling systems is robust enough<\/jats:italic>\n            ---Each of these existing systems is designed for a specific point in the tradeoff space that can be sub-optimal, sometimes by up to 10\u00d7, across workloads.\n          <\/jats:p>\n          <jats:p>\n            Our first contribution is a systematic study of all recently proposed samplers: AliasLDA, F+LDA, LightLDA, and WarpLDA. We discovered a novel system tradeoff among these samplers. Each sampler has different sampling complexity and performs differently, sometimes by 5\u00d7, on documents with different lengths. Based on this tradeoff, we further developed a hybrid sampler that uses different samplers for different types of documents. This hybrid approach works across a wide range of workloads and outperforms the fastest sampler by up to 2x. We then focused on distributed environments in which thousands of workers, each with different performance (due to virtualization and resource sharing), coordinate to train a topic model. Our second contribution is an\n            <jats:italic>asymmetric<\/jats:italic>\n            parameter server architecture that pushes some computation to the parameter server side. This architecture is motivated by the skew of the word frequency distribution and a novel tradeoff we discovered between communication and computation. With this architecture, we outperform the traditional, symmetric architecture by up to 2\u00d7.\n          <\/jats:p>\n          <jats:p>With these two contributions, together with a carefully engineered implementation, our system is able to outperform existing systems by up to 10\u00d7 and has already been running to provide topic modeling services for more than six months.<\/jats:p>","DOI":"10.14778\/3137628.3137649","type":"journal-article","created":{"date-parts":[[2017,9,7]],"date-time":"2017-09-07T13:35:53Z","timestamp":1504791353000},"page":"1406-1417","source":"Crossref","is-referenced-by-count":27,"title":["LDA*"],"prefix":"10.14778","volume":"10","author":[{"given":"Lele","family":"Yut","sequence":"first","affiliation":[{"name":"Peking University and Tencent Inc."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ce","family":"Zhang","sequence":"additional","affiliation":[{"name":"ETH Z\u00fcrich"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yingxia","family":"Shao","sequence":"additional","affiliation":[{"name":"Peking University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bin","family":"Cui","sequence":"additional","affiliation":[{"name":"Peking University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,8]]},"reference":[{"key":"e_1_2_1_1_1","first-page":"27","volume-title":"AUAI","author":"Asuncion A.","year":"2009","unstructured":"A. Asuncion , M. Welling , P. Smyth , and Y. W. Teh . On smoothing and inference for topic models . In AUAI , pages 27 -- 34 , 2009 . A. Asuncion, M. Welling, P. Smyth, and Y. W. Teh. On smoothing and inference for topic models. In AUAI, pages 27--34, 2009."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.14778\/2977797.2977801"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/2487575.2487697"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.0307752101"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2505515.2505642"},{"key":"e_1_2_1_6_1","first-page":"1223","volume-title":"NIPS","author":"Ho Q.","year":"2013","unstructured":"Q. Ho , J. Cipar , H. Cui , S. Lee , J. K. Kim , P. B. Gibbons , G. A. Gibson , G. Ganger , and E. P. Xing . More effective distributed ml via a stale synchronous parallel parameter server . In NIPS , pages 1223 -- 1231 , 2013 . Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P. B. Gibbons, G. A. Gibson, G. Ganger, and E. P. Xing. More effective distributed ml via a stale synchronous parallel parameter server. In NIPS, pages 1223--1231, 2013."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.5555\/2567709.2502622"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3035918.3035933"},{"key":"e_1_2_1_9_1","volume-title":"Angel: a new large-scale machine learning system. National Science Review, page nwx018","author":"Jiang J.","year":"2017","unstructured":"J. Jiang , L. Yu , J. Jiang , Y. Liu , and B. Cui . Angel: a new large-scale machine learning system. National Science Review, page nwx018 , 2017 . J. Jiang, L. Yu, J. Jiang, Y. Liu, and B. Cui. Angel: a new large-scale machine learning system. National Science Review, page nwx018, 2017."},{"key":"e_1_2_1_10_1","volume-title":"Nature","author":"LeCun Y.","year":"2015","unstructured":"Y. LeCun , Y. Bengio , and G. Hinton . Nature , 2015 . Y. LeCun, Y. Bengio, and G. Hinton. Nature, 2015."},{"key":"e_1_2_1_11_1","unstructured":"D. A. Levin Y. Peres and E. L. Wilmer. Markov chains and mixing times. American Mathematical Soc.  D. A. Levin Y. Peres and E. L. Wilmer. Markov chains and mixing times. American Mathematical Soc."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/2623330.2623756"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.5555\/2685048.2685095"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/1961189.1961198"},{"key":"e_1_2_1_15_1","first-page":"3111","volume-title":"Advances in neural information processing systems","author":"Mikolov T.","year":"2013","unstructured":"T. Mikolov , I. Sutskever , K. Chen , G. S. Corrado , and J. Dean . Distributed representations of words and phrases and their compositionality . In Advances in neural information processing systems , pages 3111 -- 3119 , 2013 . T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean. Distributed representations of words and phrases and their compositionality. In Advances in neural information processing systems, pages 3111--3119, 2013."},{"key":"e_1_2_1_16_1","first-page":"1081","volume-title":"NIPS","author":"Newman D.","year":"2007","unstructured":"D. Newman , P. Smyth , M. Welling , and A. U. Asuncion . Distributed inference for latent dirichlet allocation . In NIPS , pages 1081 -- 1088 , 2007 . D. Newman, P. Smyth, M. Welling, and A. U. Asuncion. Distributed inference for latent dirichlet allocation. In NIPS, pages 1081--1088, 2007."},{"key":"e_1_2_1_17_1","first-page":"3102","volume-title":"NIPS","author":"Patterson S.","year":"2013","unstructured":"S. Patterson and Y. W. Teh . Stochastic gradient riemannian langevin dynamics on the probability simplex . In NIPS , pages 3102 -- 3110 , 2013 . S. Patterson and Y. W. Teh. Stochastic gradient riemannian langevin dynamics on the probability simplex. In NIPS, pages 3102--3110, 2013."},{"key":"e_1_2_1_18_1","volume-title":"Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference","author":"Pearl J.","year":"1988","unstructured":"J. Pearl . Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference . Morgan Kaufmann Publishers Inc ., San Francisco, CA, USA, 1988 . J. Pearl. Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 1988."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1401890.1401960"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.14778\/1920841.1920931"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/2031527"},{"key":"e_1_2_1_22_1","first-page":"1353","volume-title":"NIPS","author":"Teh Y. W.","year":"2006","unstructured":"Y. W. Teh , D. Newman , and M. Welling . A collapsed variational bayesian inference algorithm for latent dirichlet allocation . In NIPS , pages 1353 -- 1360 , 2006 . Y. W. Teh, D. Newman, and M. Welling. A collapsed variational bayesian inference algorithm for latent dirichlet allocation. In NIPS, pages 1353--1360, 2006."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/355744.355749"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-02158-9_26"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2015.2472014"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/1557019.1557121"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2736277.2741682"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2736277.2741115"},{"key":"e_1_2_1_30_1","volume-title":"Medlda: maximum margin supervised topic models. Journal of Machine Learning Research, 13(Aug):2237--2278","author":"Zhu J.","year":"2012","unstructured":"J. Zhu , A. Ahmed , and E. P. Xing . Medlda: maximum margin supervised topic models. Journal of Machine Learning Research, 13(Aug):2237--2278 , 2012 . J. Zhu, A. Ahmed, and E. P. Xing. Medlda: maximum margin supervised topic models. Journal of Machine Learning Research, 13(Aug):2237--2278, 2012."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3137628.3137649","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:01:03Z","timestamp":1672221663000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3137628.3137649"}},"subtitle":["a robust and large-scale topic modeling system"],"short-title":[],"issued":{"date-parts":[[2017,8]]},"references-count":29,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2017,8]]}},"alternative-id":["10.14778\/3137628.3137649"],"URL":"https:\/\/doi.org\/10.14778\/3137628.3137649","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2017,8]]}}}