{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,6]],"date-time":"2025-08-06T13:57:24Z","timestamp":1754488644324,"version":"3.41.0"},"reference-count":43,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2021,7,20]],"date-time":"2021-07-20T00:00:00Z","timestamp":1626739200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61832003 and U1811461"],"award-info":[{"award-number":["61832003 and U1811461"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2022,2,28]]},"abstract":"<jats:p>\n            Latent Dirichlet Allocation (LDA) has been widely used for topic modeling, with applications spanning various areas such as natural language processing and information retrieval. While LDA on small and static datasets has been extensively studied, several real-world challenges are posed in practical scenarios where datasets are often huge and are gathered in a streaming fashion. As the state-of-the-art LDA algorithm on streams,\n            <jats:italic>Streaming Variational Bayes<\/jats:italic>\n            (SVB) introduced\n            <jats:italic>Bayesian updating<\/jats:italic>\n            to provide a streaming procedure. However, the utility of SVB is limited in applications since it ignored three challenges of processing real-world streams:\n            <jats:italic>topic evolution<\/jats:italic>\n            ,\n            <jats:italic>data turbulence<\/jats:italic>\n            , and\n            <jats:italic>real-time inference<\/jats:italic>\n            . In this article, we propose a novel distributed LDA algorithm\u2014referred to as\n            <jats:italic>StreamFed-LDA\u2014<\/jats:italic>\n            to deal with challenges on streams. For topic modeling of streaming data, the ability to capture evolving topics is essential for practical online inference. To achieve this goal,\n            <jats:italic>StreamFed-LDA<\/jats:italic>\n            is based on a specialized framework that supports lifelong (continual) learning of evolving topics. On the other hand, data turbulence is commonly present in streams due to real-life events. In that case, the design of\n            <jats:italic>StreamFed-LDA<\/jats:italic>\n            allows the model to learn new characteristics from the most recent data while maintaining the historical information. On massive streaming data, it is difficult and crucial to provide real-time inference results. To increase the throughput and reduce the latency,\n            <jats:italic>StreamFed-LDA<\/jats:italic>\n            introduces additional techniques that substantially reduce both computation and communication costs in distributed systems. Experiments on four real-world datasets show that the proposed framework achieves significantly better performance of online inference compared with the baselines. At the same time,\n            <jats:italic>StreamFed-LDA<\/jats:italic>\n            also reduces the latency by orders of magnitudes in real-world datasets.\n          <\/jats:p>","DOI":"10.1145\/3451528","type":"journal-article","created":{"date-parts":[[2021,7,20]],"date-time":"2021-07-20T21:06:18Z","timestamp":1626815178000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Distributed Latent Dirichlet Allocation on Streams"],"prefix":"10.1145","volume":"16","author":[{"given":"Yunyan","family":"Guo","sequence":"first","affiliation":[{"name":"Harbin Institute of Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianzhong","family":"Li","sequence":"additional","affiliation":[{"name":"Shenzhen Institute of Advanced Technology, Chinese Academy of Sciences, China and Harbin Institute of Technology, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2021,7,20]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.5555\/3026877.3026899"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3219819.3219995"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.5555\/2627435.2638571"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.5555\/2535015"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/290941.290954"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.14778\/3090163.3090168"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1080\/01621459.2017.1285773"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143859"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.5555\/944919.944937"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1137\/16M1080173"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999792.2999805"},{"key":"e_1_2_1_12_1","first-page":"28","article-title":"Apache flink: Stream and batch processing in a single engine","volume":"36","author":"Carbone Paris","year":"2015","unstructured":"Paris Carbone , Asterios Katsifodimos , Stephan Ewen , Volker Markl , Seif Haridi , and Kostas Tzoumas . 2015 . Apache flink: Stream and batch processing in a single engine . Bulletin of the IEEE Computer Society Technical Committee on Data Engineering 36 , 4 (2015), 28 \u2013 38 . http:\/\/sites.computer.org\/debull\/A15dec\/p28.pdf. Paris Carbone, Asterios Katsifodimos, Stephan Ewen, Volker Markl, Seif Haridi, and Kostas Tzoumas. 2015. Apache flink: Stream and batch processing in a single engine. Bulletin of the IEEE Computer Society Technical Committee on Data Engineering 36, 4 (2015), 28\u201338. http:\/\/sites.computer.org\/debull\/A15dec\/p28.pdf.","journal-title":"Bulletin of the IEEE Computer Society Technical Committee on Data Engineering"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence. 5078\u20135085","author":"Chen Xilun","year":"2018","unstructured":"Xilun Chen , K. Sel\u00e7uk Candan , and Maria Luisa Sapino . 2018 . Ims-dtm: Incremental multi-scale dynamic topic models . In Proceedings of the 32nd AAAI Conference on Artificial Intelligence. 5078\u20135085 . Xilun Chen, K. Sel\u00e7uk Candan, and Maria Luisa Sapino. 2018. Ims-dtm: Incremental multi-scale dynamic topic models. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence. 5078\u20135085."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2522968.2522981"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/2997189.2997285"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.5555\/2567709.2502622"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2086737.2086739"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3077136.3084135"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1024940629314"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.5555\/2685048.2685095"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939748"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2736277.2741106"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969239.2969368"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS\u201917)","volume":"54","author":"McMahan H. Brendan","year":"2016","unstructured":"H. Brendan McMahan , Eider Moore , Daniel Ramage , Seth Hampson , Blaise Ag\u00fcera y. Arcas . 2016 . Communication-efficient learning of deep networks from decentralized data . In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS\u201917) , Vol. 54 . 1273\u20131282. H. Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, Blaise Ag\u00fcera y. Arcas. 2016. Communication-efficient learning of deep networks from decentralized data. In Proceedings of the 20th International Conference on Artificial Intelligence and Statistics (AISTATS\u201917), Vol. 54. 1273\u20131282."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.5555\/2946645.2946679"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-014-0808-1"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.5555\/3454287.3455008"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/2999973.2999976"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.5555\/3042817.3042927"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2339530.2339704"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1177\/1094342005051521"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.5555\/3045118.3045384"},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of Machine Learning and Systems 2019 (MLSys\u201919)","author":"Wang Jianyu","year":"2018","unstructured":"Jianyu Wang and Gauri Joshi . 2018 . Adaptive communication strategies to achieve the best error-runtime trade-off in local-update SGD . In Proceedings of Machine Learning and Systems 2019 (MLSys\u201919) , Ameet Talwalkar, Virginia Smith and Matei Zaharia (Eds.). https:\/\/proceedings.mlsys.org\/book\/257.pdf. Jianyu Wang and Gauri Joshi. 2018. Adaptive communication strategies to achieve the best error-runtime trade-off in local-update SGD. In Proceedings of Machine Learning and Systems 2019 (MLSys\u201919), Ameet Talwalkar, Virginia Smith and Matei Zaharia (Eds.). https:\/\/proceedings.mlsys.org\/book\/257.pdf."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2016.2556661"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBDATA.2015.2472014"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2736277.2741682"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2736277.2741115"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.14778\/3137628.3137649"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517349.2522737"},{"key":"e_1_2_1_41_1","volume-title":"Proceedings of the Artificial Intelligence and Statistics. 966\u2013975","author":"Zaheer Manzil","year":"2016","unstructured":"Manzil Zaheer , Michael Wick , Jean-Baptiste Tristan , Alex Smola , and Guy Steele . 2016 . Exponential stochastic cellular automata for massively parallel inference . In Proceedings of the Artificial Intelligence and Statistics. 966\u2013975 . Manzil Zaheer, Michael Wick, Jean-Baptiste Tristan, Alex Smola, and Guy Steele. 2016. Exponential stochastic cellular automata for massively parallel inference. In Proceedings of the Artificial Intelligence and Statistics. 966\u2013975."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2889774"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCCN.2017.8038464"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3451528","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3451528","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:02:59Z","timestamp":1750197779000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3451528"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,7,20]]},"references-count":43,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2022,2,28]]}},"alternative-id":["10.1145\/3451528"],"URL":"https:\/\/doi.org\/10.1145\/3451528","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2021,7,20]]},"assertion":[{"value":"2020-07-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-02-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2021-07-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}