{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,29]],"date-time":"2025-11-29T07:59:48Z","timestamp":1764403188022},"reference-count":64,"publisher":"MIT Press","issue":"4","license":[{"start":{"date-parts":[[2022,8,23]],"date-time":"2022-08-23T00:00:00Z","timestamp":1661212800000},"content-version":"vor","delay-in-days":234,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc-nd\/4.0\/"}],"content-domain":{"domain":["direct.mit.edu"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,12,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>We propose a method that uses neural embeddings to improve the performance of any given LDA-style topic model. Our method, called neural embedding allocation (NEA), deconstructs topic models (LDA or otherwise) into interpretable vector-space embeddings of words, topics, documents, authors, and so on, by learning neural embeddings to mimic the topic model. We demonstrate that NEA improves coherence scores of the original topic model by smoothing out the noisy topics when the number of topics is large. Furthermore, we show NEA\u2019s effectiveness and generality in deconstructing and smoothing LDA, author-topic models, and the recent mixed membership skip-gram topic model and achieve better performance with the embeddings compared to several state-of-the-art models.<\/jats:p>","DOI":"10.1162\/coli_a_00457","type":"journal-article","created":{"date-parts":[[2022,8,23]],"date-time":"2022-08-23T19:35:09Z","timestamp":1661283309000},"page":"1021-1052","update-policy":"http:\/\/dx.doi.org\/10.1162\/mitpressjournals.corrections.policy","source":"Crossref","is-referenced-by-count":5,"title":["Neural Embedding Allocation: Distributed Representations of Topic Models"],"prefix":"10.1162","volume":"48","author":[{"given":"Kamrun Naher","family":"Keya","sequence":"first","affiliation":[{"name":"University of Maryland, Baltimore County Department of Information Systems kkeya1@umbc.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yannis","family":"Papanikolaou","sequence":"additional","affiliation":[{"name":"Healx Department of Research and Development yannis.papanikolaou@healx.io"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"James R.","family":"Foulds","sequence":"additional","affiliation":[{"name":"University of Maryland, Baltimore County Department of Information Systems jfoulds@umbc.edu"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"281","published-online":{"date-parts":[[2022,12,1]]},"reference":[{"key":"2022120515175291300_bib1","doi-asserted-by":"publisher","first-page":"671","DOI":"10.2139\/ssrn.2477899","article-title":"Big data\u2019s disparate impact","volume":"104","author":"Barocas","year":"2016","journal-title":"California Law Review"},{"key":"2022120515175291300_bib2","doi-asserted-by":"publisher","first-page":"610","DOI":"10.1145\/3442188.3445922","article-title":"On the dangers of stochastic parrots: Can language models be too big?","volume-title":"Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency","author":"Bender","year":"2021"},{"issue":"Feb","key":"2022120515175291300_bib3","first-page":"1137","article-title":"A neural probabilistic language model","volume":"3","author":"Bengio","year":"2003","journal-title":"Journal of Machine Learning Research"},{"key":"2022120515175291300_bib4","volume-title":"Adverse Impact and Test Validation: A Practitioner\u2019s Guide to Valid and Defensible Employment Testing","author":"Biddle","year":"2006"},{"issue":"Jan","key":"2022120515175291300_bib5","first-page":"993","article-title":"Latent Dirichlet allocation","volume":"3","author":"Blei","year":"2003","journal-title":"Journal of Machine Learning Research"},{"key":"2022120515175291300_bib6","first-page":"4349","article-title":"Man is to computer programmer as woman is to homemaker? Debiasing word embeddings","volume":"29","author":"Bolukbasi","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022120515175291300_bib7","first-page":"1877","article-title":"Language models are few-shot learners","volume":"33","author":"Brown","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022120515175291300_bib8","doi-asserted-by":"publisher","first-page":"535","DOI":"10.1145\/1150402.1150464","article-title":"Model compression","volume-title":"Proceedings of the 12th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Bucilua\u030c","year":"2006"},{"issue":"6334","key":"2022120515175291300_bib9","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1126\/science.aal4230","article-title":"Semantics derived automatically from language corpora contain human-like biases","volume":"356","author":"Caliskan","year":"2017","journal-title":"Science"},{"issue":"Aug","key":"2022120515175291300_bib10","first-page":"2493","article-title":"Natural language processing (almost) from scratch","volume":"12","author":"Collobert","year":"2011","journal-title":"Journal of Machine Learning Research"},{"key":"2022120515175291300_bib11","doi-asserted-by":"publisher","first-page":"795","DOI":"10.3115\/v1\/P15-1077","article-title":"Gaussian LDA for topic models with word embeddings","volume-title":"Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)","author":"Das","year":"2015"},{"key":"2022120515175291300_bib12","doi-asserted-by":"publisher","first-page":"268","DOI":"10.1145\/3386392.3399569","article-title":"Mitigating demographic bias in AI-based resume filtering","author":"Deshpande","year":"2020","journal-title":"Fairness in User Modeling, Adaptation and Personalization Workshop"},{"key":"2022120515175291300_bib13","first-page":"879","article-title":"Attenuating bias in word vectors","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"Dev","year":"2019"},{"key":"2022120515175291300_bib14","first-page":"4171","article-title":"BERT: Pre-training of deep bidirectional transformers for language understanding","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Devlin","year":"2019"},{"key":"2022120515175291300_bib15","doi-asserted-by":"publisher","first-page":"439","DOI":"10.1162\/tacl_a_00325","article-title":"Topic modeling in embedding spaces","volume":"8","author":"Dieng","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022120515175291300_bib16","doi-asserted-by":"publisher","first-page":"214","DOI":"10.1145\/2090236.2090255","article-title":"Fairness through awareness","volume-title":"Proceedings of the 3rd Innovations in Theoretical Computer Science Conference","author":"Dwork","year":"2012"},{"key":"2022120515175291300_bib17","article-title":"Notes on noise contrastive estimation and negative sampling","author":"Dyer","year":"2014","journal-title":"arXiv preprint arXiv:1410.8251"},{"key":"2022120515175291300_bib18","first-page":"86","article-title":"Mixed membership word embeddings for computational social science","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"Foulds","year":"2018"},{"key":"2022120515175291300_bib19","doi-asserted-by":"publisher","first-page":"1918","DOI":"10.1109\/ICDE48307.2020.00203","article-title":"An intersectional definition of fairness","volume-title":"2020 IEEE 36th International Conference on Data Engineering","author":"Foulds","year":"2020"},{"key":"2022120515175291300_bib20","first-page":"609","article-title":"Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Gonen","year":"2019"},{"issue":"2","key":"2022120515175291300_bib21","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1037\/0033-295X.114.2.211","article-title":"Topics in semantic representation","volume":"114","author":"Griffiths","year":"2007","journal-title":"Psychological Review"},{"key":"2022120515175291300_bib22","article-title":"BERTopic: Neural topic modeling with a class-based TF-IDF procedure","author":"Grootendorst","year":"2022","journal-title":"arXiv preprint arXiv:2203.05794"},{"key":"2022120515175291300_bib23","first-page":"297","article-title":"Noise-contrastive estimation: A new estimation principle for unnormalized statistical models","volume-title":"Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics","author":"Gutmann","year":"2010"},{"issue":"Feb","key":"2022120515175291300_bib24","first-page":"307","article-title":"Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics","volume":"13","author":"Gutmann","year":"2012","journal-title":"Journal of Machine Learning Research"},{"key":"2022120515175291300_bib25","article-title":"Distilling the knowledge in a neural network","author":"Hinton","year":"2015","journal-title":"arXiv preprint arXiv:1503.02531"},{"issue":"8","key":"2022120515175291300_bib26","doi-asserted-by":"publisher","first-page":"1771","DOI":"10.1162\/089976602760128018","article-title":"Training products of experts by minimizing contrastive divergence","volume":"14","author":"Hinton","year":"2002","journal-title":"Neural Computation"},{"key":"2022120515175291300_bib27","first-page":"1","article-title":"Learning distributed representations of concepts","volume-title":"Proceedings of the Eighth Annual Conference of the Cognitive Science Society","author":"Hinton","year":"1986"},{"key":"2022120515175291300_bib28","doi-asserted-by":"publisher","first-page":"2836","DOI":"10.18653\/v1\/N19-1291","article-title":"Scalable collapsed inference for high-dimensional topic models","volume-title":"Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers)","author":"Islam","year":"2019"},{"key":"2022120515175291300_bib29","article-title":"Mitigating demographic biases in social media-based recommender systems","author":"Islam","year":"2019","journal-title":"The 25th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Social Impact Track)"},{"key":"2022120515175291300_bib30","doi-asserted-by":"publisher","first-page":"3779","DOI":"10.1145\/3442381.3449904","article-title":"Debiasing career recommendations with neural fair collaborative filtering","volume-title":"Proceedings of the Web Conference 2021","author":"Islam","year":"2021"},{"key":"2022120515175291300_bib31","doi-asserted-by":"publisher","first-page":"586","DOI":"10.1145\/3461702.3462614","article-title":"Can we obtain fairness for free?","volume-title":"Proceedings of the 2021 AAAI\/ACM Conference on AI, Ethics, and Society","author":"Islam","year":"2021"},{"key":"2022120515175291300_bib32","article-title":"Neural embedding allocation: Distributed representations of words, topics, and documents","volume-title":"Mid-Atlantic Student Colloquium on Speech, Language and Learning","author":"Keya","year":"2018"},{"key":"2022120515175291300_bib33","doi-asserted-by":"publisher","first-page":"190","DOI":"10.1137\/1.9781611976700.22","article-title":"Equitable allocation of healthcare resources with fair survival models","volume-title":"Proceedings of the 2021 SIAM International Conference on Data Mining","author":"Keya","year":"2021"},{"key":"2022120515175291300_bib34","doi-asserted-by":"publisher","first-page":"1746","DOI":"10.3115\/v1\/D14-1181","article-title":"Convolutional neural networks for sentence classification","volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing","author":"Kim","year":"2014"},{"key":"2022120515175291300_bib35","doi-asserted-by":"publisher","first-page":"530","DOI":"10.3115\/v1\/E14-1056","article-title":"Machine reading tea leaves: Automatically evaluating topic coherence and topic model quality","volume-title":"Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics","author":"Lau","year":"2014"},{"key":"2022120515175291300_bib36","first-page":"1188","article-title":"Distributed representations of sentences and documents","volume-title":"International Conference on Machine Learning","author":"Le","year":"2014"},{"key":"2022120515175291300_bib37","first-page":"2177","article-title":"Neural word embedding as implicit matrix factorization","volume":"27","author":"Levy","year":"2014","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022120515175291300_bib38","doi-asserted-by":"publisher","first-page":"891","DOI":"10.1145\/2623330.2623756","article-title":"Reducing the sampling complexity of topic models","volume-title":"Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Li","year":"2014"},{"key":"2022120515175291300_bib39","doi-asserted-by":"publisher","first-page":"2418","DOI":"10.1609\/aaai.v29i1.9522","article-title":"Topical word embeddings","volume-title":"Twenty-ninth AAAI Conference on Artificial Intelligence","author":"Liu","year":"2015"},{"key":"2022120515175291300_bib40","article-title":"RoBERTa: A robustly optimized BERT pretraining approach","author":"Liu","year":"2019","journal-title":"arXiv preprint arXiv:1907.11692"},{"key":"2022120515175291300_bib41","doi-asserted-by":"publisher","first-page":"1908","DOI":"10.1145\/3394486.3403242","article-title":"Hierarchical topic mining via joint spherical tree and text embedding","volume-title":"Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining","author":"Meng","year":"2020"},{"key":"2022120515175291300_bib42","first-page":"1727","article-title":"Neural variational inference for text processing","volume-title":"International Conference on Machine Learning","author":"Miao","year":"2016"},{"key":"2022120515175291300_bib43","article-title":"Efficient estimation of word representations in vector space","author":"Mikolov","year":"2013","journal-title":"arXiv preprint arXiv:1301.3781"},{"key":"2022120515175291300_bib44","first-page":"3111","article-title":"Distributed representations of words and phrases and their compositionality","volume":"26","author":"Mikolov","year":"2013","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022120515175291300_bib45","first-page":"262","article-title":"Optimizing semantic coherence in topic models","volume-title":"Proceedings of the Conference on Empirical Methods in Natural Language Processing","author":"Mimno","year":"2011"},{"key":"2022120515175291300_bib46","first-page":"2265","article-title":"Learning word embeddings efficiently with noise-contrastive estimation","volume":"26","author":"Mnih","year":"2013","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022120515175291300_bib47","article-title":"Mixing Dirichlet topic models and word embeddings to make LDA2vec","author":"Moody","year":"2016","journal-title":"arXiv preprint arXiv:1605.02019"},{"key":"2022120515175291300_bib48","doi-asserted-by":"publisher","first-page":"299","DOI":"10.1162\/tacl_a_00140","article-title":"Improving topic models with latent feature word representations","volume":"3","author":"Nguyen","year":"2015","journal-title":"Transactions of the Association for Computational Linguistics"},{"key":"2022120515175291300_bib49","doi-asserted-by":"publisher","first-page":"7047","DOI":"10.18653\/v1\/2020.acl-main.630","article-title":"tBERT: Topic models and BERT joining forces for semantic similarity detection","volume-title":"Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics","author":"Peinelt","year":"2020"},{"key":"2022120515175291300_bib50","doi-asserted-by":"publisher","first-page":"1532","DOI":"10.3115\/v1\/D14-1162","article-title":"GloVe: Global vectors for word representation","volume-title":"Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing","author":"Pennington","year":"2014"},{"key":"2022120515175291300_bib51","doi-asserted-by":"publisher","first-page":"2227","DOI":"10.18653\/v1\/N18-1202","article-title":"Deep contextualized word representations","volume-title":"Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers)","author":"Peters","year":"2018"},{"key":"2022120515175291300_bib52","article-title":"Improving language understanding by generative pre-training","author":"Radford","year":"2018"},{"issue":"8","key":"2022120515175291300_bib53","first-page":"9","article-title":"Language models are unsupervised multitask learners","volume":"1","author":"Radford","year":"2019","journal-title":"OpenAI Blog"},{"key":"2022120515175291300_bib54","first-page":"487","article-title":"The author-topic model for authors and documents","volume-title":"Proceedings of the 20th Conference on Uncertainty in Artificial Intelligence","author":"Rosen-Zvi","year":"2004"},{"key":"2022120515175291300_bib55","first-page":"33","article-title":"The distributional hypothesis","volume":"20","author":"Sahlgren","year":"2008","journal-title":"Italian Journal of Disability Studies"},{"key":"2022120515175291300_bib56","first-page":"199","article-title":"Effects of age and gender on blogging","volume-title":"AAAI Spring Symposium: Computational Approaches to Analyzing Weblogs","author":"Schler","year":"2006"},{"key":"2022120515175291300_bib57","doi-asserted-by":"publisher","first-page":"375","DOI":"10.1145\/3077136.3080806","article-title":"Jointly learning word embeddings and latent topics","volume-title":"Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval","author":"Shi","year":"2017"},{"key":"2022120515175291300_bib58","doi-asserted-by":"publisher","first-page":"1728","DOI":"10.18653\/v1\/2020.emnlp-main.135","article-title":"Tired of topic models? Clusters of pretrained word embeddings make for fast and good topics too!","volume-title":"Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing","author":"Sia","year":"2020"},{"key":"2022120515175291300_bib59","article-title":"Topic modeling with contextualized word representation clusters","author":"Thompson","year":"2020","journal-title":"arXiv preprint arXiv:2010.12626"},{"key":"2022120515175291300_bib60","first-page":"5998","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2022120515175291300_bib61","first-page":"1387","article-title":"Decoding with large-scale neural language models improves translation","volume-title":"Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing","author":"Vaswani","year":"2013"},{"key":"2022120515175291300_bib62","first-page":"962","article-title":"Fairness constraints: Mechanisms for fair classification","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"Zafar","year":"2017"},{"key":"2022120515175291300_bib63","first-page":"325","article-title":"Learning fair representations","volume-title":"International Conference on Machine Learning","author":"Zemel","year":"2013"},{"key":"2022120515175291300_bib64","doi-asserted-by":"publisher","first-page":"471","DOI":"10.1162\/tacl_a_00326","article-title":"A neural generative model for joint learning topics and topic-specific word embeddings","volume":"8","author":"Zhu","year":"2020","journal-title":"Transactions of the Association for Computational Linguistics"}],"container-title":["Computational Linguistics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/48\/4\/1021\/2061857\/coli_a_00457.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/direct.mit.edu\/coli\/article-pdf\/48\/4\/1021\/2061857\/coli_a_00457.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,5]],"date-time":"2022-12-05T15:19:24Z","timestamp":1670253564000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/coli\/article\/48\/4\/1021\/112769\/Neural-Embedding-Allocation-Distributed"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022]]},"references-count":64,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,12,1]]},"published-print":{"date-parts":[[2022,12,1]]}},"URL":"https:\/\/doi.org\/10.1162\/coli_a_00457","relation":{},"ISSN":["0891-2017","1530-9312"],"issn-type":[{"value":"0891-2017","type":"print"},{"value":"1530-9312","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022]]},"published":{"date-parts":[[2022]]}}}