{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:18:35Z","timestamp":1760242715572,"version":"build-2065373602"},"reference-count":26,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2016,4,13]],"date-time":"2016-04-13T00:00:00Z","timestamp":1460505600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Chinese National Program on Key Basic Research Project","award":["2013CB329304, 2014CB744604"],"award-info":[{"award-number":["2013CB329304, 2014CB744604"]}]},{"DOI":"10.13039\/501100001809","name":"Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61402324, 61272265, 61373122"],"award-info":[{"award-number":["61402324, 61272265, 61373122"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Separating two probability distributions from a mixture model that is made up of the combinations of the two is essential to a wide range of applications. For example, in information retrieval (IR), there often exists a mixture distribution consisting of a relevance distribution that we need to estimate and an irrelevance distribution that we hope to get rid of. Recently, a distribution separation method (DSM) was proposed to approximate the relevance distribution, by separating a seed irrelevance distribution from the mixture distribution. It was successfully applied to an IR task, namely pseudo-relevance feedback (PRF), where the query expansion model is often a mixture term distribution. Although initially developed in the context of IR, DSM is indeed a general mathematical formulation for probability distribution separation. Thus, it is important to further generalize its basic analysis and to explore its connections to other related methods. In this article, we first extend DSM\u2019s theoretical analysis, which was originally based on the Pearson correlation coefficient, to entropy-related measures, including the KL-divergence (Kullback\u2013Leibler divergence), the symmetrized KL-divergence and the JS-divergence (Jensen\u2013Shannon divergence). Second, we investigate the distribution separation idea in a well-known method, namely the mixture model feedback (MMF) approach. We prove that MMF also complies with the linear combination assumption, and then, DSM\u2019s linear separation algorithm can largely simplify the EM algorithm in MMF. These theoretical analyses, as well as further empirical evaluation results demonstrate the advantages of our DSM approach.<\/jats:p>","DOI":"10.3390\/e18040105","type":"journal-article","created":{"date-parts":[[2016,4,13]],"date-time":"2016-04-13T10:08:06Z","timestamp":1460542086000},"page":"105","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Generalized Analysis of a Distribution Separation Method"],"prefix":"10.3390","volume":"18","author":[{"given":"Peng","family":"Zhang","sequence":"first","affiliation":[{"name":"Tianjin Key Laboratory of Cognitive Computing and Application, School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Qian","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Computer Software, Tianjin University, Tianjin 300072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuexian","family":"Hou","sequence":"additional","affiliation":[{"name":"Tianjin Key Laboratory of Cognitive Computing and Application, School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dawei","family":"Song","sequence":"additional","affiliation":[{"name":"Tianjin Key Laboratory of Cognitive Computing and Application, School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"},{"name":"Computing and Communications Department, The Open University, Milton Keynes MK7 6AA, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingfei","family":"Li","sequence":"additional","affiliation":[{"name":"Tianjin Key Laboratory of Cognitive Computing and Application, School of Computer Science and Technology, Tianjin University, Tianjin 300072, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bin","family":"Hu","sequence":"additional","affiliation":[{"name":"Ubiquitous Awareness and Intelligent Solutions Lab, Lanzhou University, Lanzhou 730000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2016,4,13]]},"reference":[{"key":"ref_1","unstructured":"Van Rijsbergen, C.J. (1979). Information Retrieval, Butterworth-Heinemann."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Lavrenko, V., and Croft, W.B. (2001, January 9\u201312). Relevance-Based Language Models. Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, New Orleans, LA, USA.","DOI":"10.1145\/383952.383972"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1145\/248625.248650","article-title":"The effect of accessing non-matching documents on relevance feedback","volume":"15","author":"Dunlop","year":"1997","journal-title":"ACM Trans. Inf. Syst. (TOIS)"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Singhal, A., Mitra, M., and Buckley, C. (1997, January 27\u201331). Learning Routing Queries in a Query Zone. Proceedings of the 20th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Philadelphia, PA, USA.","DOI":"10.1145\/258525.258530"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Wang, X., Fang, H., and Zhai, C. (2008, January 27\u201328). A study of methods for negative relevance feedback. Proceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, University of Maryland, College Park, MD, USA.","DOI":"10.1145\/1390334.1390374"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Zhang, P., Hou, Y., and Song, D. (2009, January 19\u201323). Approximating true relevance distribution from a mixture model based on irrelevance data. Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval, Boston, MA, USA.","DOI":"10.1145\/1571941.1571962"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhai, C., and Lafferty, J. (2001, January 5\u201310). Model-based feedback in the language modeling approach to information retrieval. Proceedings of the Tenth International Conference on Information and Knowledge Management, CIKM 2001, Atlanta, GA, USA.","DOI":"10.1145\/502585.502654"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"59","DOI":"10.2307\/2685263","article-title":"Thirteen Ways to Look at the Correlation Coefficient","volume":"42","author":"Rodgers","year":"1988","journal-title":"Am. Stat."},{"key":"ref_9","unstructured":"Zhai, C. A Note on the Expectation-Maximization (EM) Algorithm. Available online: http:\/\/citeseerx.ist.psu.edu\/viewdoc\/download?doi=10.1.1.149.8289&rep=rep1&type=pdf."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1076","DOI":"10.1016\/j.ipm.2007.12.003","article-title":"Fast exact maximum likelihood estimation for mixture of language model","volume":"44","author":"Zhang","year":"2008","journal-title":"Inf. Process. Manag."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Tao, T., and Zhai, C. (2006, January 6\u201311). Regularized Estimation of Mixture Models for Robust Pseudo-Relevance Feedback. Proceedings of the 29th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Seattle, WA, USA.","DOI":"10.1145\/1148170.1148201"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Chen, K., Liu, S., Chen, B., Jan, E., Wang, H., Hsu, W., and Chen, H. (2014, January 25\u201329). Leveraging Effective Query Modeling Techniques for Speech Recognition and Summarization. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing, EMNLP 2014, Doha, Qatar.","DOI":"10.3115\/v1\/D14-1156"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2071389.2071390","article-title":"A Survey of Automatic Query Expansion in Information Retrieval","volume":"44","author":"Carpineto","year":"2012","journal-title":"ACM Comput. Surv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Lv, Y., and Zhai, C. (2009, January 2\u20136). A comparative study of methods for estimating query language models with pseudo feedback. Proceedings of the 18th ACM Conference on Information and Knowledge Management, CIKM 2009, Hong Kong, China.","DOI":"10.1145\/1645953.1646259"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Miao, J., Huang, J.X., and Ye, Z. (2012, January 12\u201316). Proximity-based rocchio\u2019s model for pseudo relevance. Proceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Retrieval, Portland, OR, USA.","DOI":"10.1145\/2348283.2348356"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Lv, Y., Zhai, C.X., and Chen, W. (2011, January 25\u201329). A boosting approach to improving pseudo-relevance feedback. Proceeding of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2011, Beijing, China.","DOI":"10.1145\/2009916.2009942"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Ye, Z., and Huang, J.X. (2014, January 6\u201311). A simple term frequency transformation model for effective pseudo relevance feedback. Proceedings of the 37th International ACM SIGIR Conference on Research & Development in Information Retrieval, Gold Coast, Australia.","DOI":"10.1145\/2600428.2609636"},{"key":"ref_18","unstructured":"Clinchant, S., and Gaussier, \u00c9. (October, January 29). A Theoretical Analysis of Pseudo-Relevance Feedback Models. Proceedings of the International Conference on the Theory of Information Retrieval, Copenhagen, Denmark."},{"key":"ref_19","unstructured":"Lafferty, J.D., and Zhai, C. (June, January 31). Document language models, query models, and risk minimization for information retrieval. Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval, Carnegie Mellon, PA, USA."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Ogilvie, P., and Callan, J. (2002, January 19\u201322). Experiments using the Lemur toolkit. Proceedings of the ACM 11th Text Retrieval Conference, Gaitherburg, MD, USA.","DOI":"10.6028\/NIST.SP.500-250.cmu-lti"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"365","DOI":"10.1016\/j.ins.2004.07.003","article-title":"A subjective measure of web search quality","volume":"169","author":"Beg","year":"2005","journal-title":"Inf. Sci."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Bennett, P.N., White, R.W., Chu, W., Dumais, S.T., Bailey, P., Borisyuk, F., and Cui, X. (2012, January 12\u201316). Modeling the impact of short-and long-term behaviour on search personalization. Proceedings of the 35th international ACM SIGIR conference on Research and development in information retrieval, Portland, OR, USA.","DOI":"10.1145\/2348283.2348312"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"147","DOI":"10.1145\/1059981.1059982","article-title":"Evaluating implicit measures to improve web search","volume":"23","author":"Fox","year":"2005","journal-title":"ACM Trans. Inf. Syst. (TOIS)"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Gao, J., Yuan, W., Li, X., Deng, K., and Nie, J.Y. (2009, January 19\u201323). Smoothing clickthrough data for web search ranking. Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval, Boston, MA, USA.","DOI":"10.1145\/1571941.1572003"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Dou, Z., Song, R., and Wen, J.R. (2007, January 8\u201312). A large-scale evaluation and analysis of personalized search strategies. Proceedings of the 16th International Conference on World Wide Web, Banff, AB, Canada.","DOI":"10.1145\/1242572.1242651"},{"key":"ref_26","first-page":"1","article-title":"Information retrieval evaluation","volume":"3","author":"Harman","year":"2011","journal-title":"Synth. Lect. Inf. Concepts Retr. Serv."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/18\/4\/105\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T19:22:10Z","timestamp":1760210530000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/18\/4\/105"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,4,13]]},"references-count":26,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2016,4]]}},"alternative-id":["e18040105"],"URL":"https:\/\/doi.org\/10.3390\/e18040105","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2016,4,13]]}}}