{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T23:35:40Z","timestamp":1764977740941,"version":"3.46.0"},"reference-count":31,"publisher":"Walter de Gruyter GmbH","issue":"2","license":[{"start":{"date-parts":[[2017,7,20]],"date-time":"2017-07-20T00:00:00Z","timestamp":1500508800000},"content-version":"unspecified","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by-nc-nd\/3.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,4,24]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>The naive Bayes classifier is a popular classifier, as it is easy to train, requires no cross-validation for parameter tuning, and can be easily extended due to its generative model. Moreover, recently it was shown that the word probabilities (background distribution) estimated from large unlabeled corpora could be used to improve the parameter estimation of naive Bayes. However, previous methods do not explicitly allow to control how much the background distribution can influence the estimation of naive Bayes parameters. In contrast, we investigate an extension of the graphical model of naive Bayes such that a word is either generated from a background distribution or from a class-specific word distribution. We theoretically analyze this model and show the connection to Jelinek-Mercer smoothing. Experiments using four standard text classification data sets show that the proposed method can statistically significantly outperform previous methods that use the same background distribution.<\/jats:p>","DOI":"10.1515\/jisys-2017-0016","type":"journal-article","created":{"date-parts":[[2017,7,20]],"date-time":"2017-07-20T06:01:12Z","timestamp":1500530472000},"page":"259-273","source":"Crossref","is-referenced-by-count":2,"title":["Analysis of the Use of Background Distribution for Naive Bayes Classifiers"],"prefix":"10.1515","volume":"28","author":[{"given":"Daniel","family":"Andrade","sequence":"first","affiliation":[{"name":"Security Research Laboratories, NEC Corporation , Tokyo , Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Akihiro","family":"Tamura","sequence":"additional","affiliation":[{"name":"Graduate School of Science and Engineering , Ehime University , Matsuyama , Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Masaaki","family":"Tsuchida","sequence":"additional","affiliation":[{"name":"AI System Department , DeNA Co., Ltd. , Tokyo , Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"374","published-online":{"date-parts":[[2017,7,20]]},"reference":[{"unstructured":"D. M. Blei, A. Y. Ng and M. I. Jordan, Latent Dirichlet allocation, J. Mach. Learn. Res.3 (2003), 993\u20131022.","key":"2025120523300492680_j_jisys-2017-0016_ref_001_w2aab3b7b1b1b6b1ab1b9b1Aa"},{"doi-asserted-by":"crossref","unstructured":"C. Chemudugunta, P. Smyth and M. Steyvers, Modeling general and specific aspects of documents with a probabilistic topic model, in: NIPS, 19, pp. 241\u2013248, 2006.","key":"2025120523300492680_j_jisys-2017-0016_ref_002_w2aab3b7b1b1b6b1ab1b9b2Aa","DOI":"10.7551\/mitpress\/7503.003.0035"},{"unstructured":"J. Cheng and R. Greiner, Comparing Bayesian network classifiers, in: Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence, Morgan Kaufmann Publishers, Inc., pp. 101\u2013108, 1999.","key":"2025120523300492680_j_jisys-2017-0016_ref_003_w2aab3b7b1b1b6b1ab1b9b3Aa"},{"unstructured":"R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu and P. Kuksa, Natural language processing (almost) from scratch, J. Mach. Learn. Res.12 (2011), 2493\u20132537.","key":"2025120523300492680_j_jisys-2017-0016_ref_004_w2aab3b7b1b1b6b1ab1b9b4Aa"},{"doi-asserted-by":"crossref","unstructured":"H. Daum\u00e9 III and D. Marcu, Bayesian query-focused summarization, in: Proceedings of the 21st International Conference on Computational Linguistics and the 44th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, pp. 305\u2013312, 2006.","key":"2025120523300492680_j_jisys-2017-0016_ref_005_w2aab3b7b1b1b6b1ab1b9b5Aa","DOI":"10.3115\/1220175.1220214"},{"doi-asserted-by":"crossref","unstructured":"A. P. Dempster, N. M. Laird, D. B. Rubin, Maximum likelihood from incomplete data via the EM algorithm, J. R. Stat. Soc.39 (1977), 1\u201338.","key":"2025120523300492680_j_jisys-2017-0016_ref_006_w2aab3b7b1b1b6b1ab1b9b6Aa","DOI":"10.1111\/j.2517-6161.1977.tb01600.x"},{"unstructured":"J. Dem\u0161ar, Statistical comparisons of classifiers over multiple data sets, J. Mach. Learn. Res.7 (2006), 1\u201330.","key":"2025120523300492680_j_jisys-2017-0016_ref_007_w2aab3b7b1b1b6b1ab1b9b7Aa"},{"unstructured":"J. Eisenstein, A. Ahmed and E. P. Xing, Sparse additive generative models of text, in: Proceedings of the 28th International Conference on Machine Learning (ICML-11), pp. 1041\u20131048, 2011.","key":"2025120523300492680_j_jisys-2017-0016_ref_008_w2aab3b7b1b1b6b1ab1b9b8Aa"},{"unstructured":"E. C. Y. He and K. L. J. Zhao, A weakly supervised Bayesian model for violence detection in social media, in: International Joint Conference on Natural Language Processing (IJCNLP), 2013.","key":"2025120523300492680_j_jisys-2017-0016_ref_009_w2aab3b7b1b1b6b1ab1b9b9Aa"},{"doi-asserted-by":"crossref","unstructured":"W. Hersh, C. Buckley, T. J. Leone and D. Hickam, OHSUMED: an interactive retrieval evaluation and new large test collection for research, in: SIGIR, Springer, pp. 192\u2013201, 1994.","key":"2025120523300492680_j_jisys-2017-0016_ref_010_w2aab3b7b1b1b6b1ab1b9c10Aa","DOI":"10.1007\/978-1-4471-2099-5_20"},{"doi-asserted-by":"crossref","unstructured":"R. L. Iman and J. M. Davenport, Approximations of the critical region of the Fbietkan statistic, Commun. Stat. Theory Methods9 (1980), 571\u2013595.10.1080\/03610928008827904","key":"2025120523300492680_j_jisys-2017-0016_ref_011_w2aab3b7b1b1b6b1ab1b9c11Aa","DOI":"10.1080\/03610928008827904"},{"doi-asserted-by":"crossref","unstructured":"L. Jiang, D. Wang and Z. Cai, Discriminatively weighted naive Bayes and its application in text classification, Int. J. Artif. Intell. Tools21 (2012), 1250007.10.1142\/S0218213011004770","key":"2025120523300492680_j_jisys-2017-0016_ref_012_w2aab3b7b1b1b6b1ab1b9c12Aa","DOI":"10.1142\/S0218213011004770"},{"doi-asserted-by":"crossref","unstructured":"L. Jiang, Z. Cai, H. Zhang and D. Wang, naive Bayes text classifiers: a locally weighted learning approach, J. Exp. Theor. Artif. Intell.25 (2013), 273\u2013286.10.1080\/0952813X.2012.721010","key":"2025120523300492680_j_jisys-2017-0016_ref_013_w2aab3b7b1b1b6b1ab1b9c13Aa","DOI":"10.1080\/0952813X.2012.721010"},{"doi-asserted-by":"crossref","unstructured":"L. Jiang, C. Li, S. Wang and L. Zhang, Deep feature weighting for naive Bayes and its application to text classification, Eng. Appl. Artif. Intell.52 (2016), 26\u201339.10.1016\/j.engappai.2016.02.002","key":"2025120523300492680_j_jisys-2017-0016_ref_014_w2aab3b7b1b1b6b1ab1b9c14Aa","DOI":"10.1016\/j.engappai.2016.02.002"},{"doi-asserted-by":"crossref","unstructured":"L. Jiang, S. Wang, C. Li and L. Zhang, Structure extended multinomial naive Bayes, Inf. Sci.329 (2016), 346\u2013356.10.1016\/j.ins.2015.09.037","key":"2025120523300492680_j_jisys-2017-0016_ref_015_w2aab3b7b1b1b6b1ab1b9c15Aa","DOI":"10.1016\/j.ins.2015.09.037"},{"doi-asserted-by":"crossref","unstructured":"Q. Jiang, W. Wang, X. Han, S. Zhang, X. Wang and C. Wang, Deep feature weighting in naive Bayes for Chinese text classification, in: Cloud Computing and Intelligence Systems (CCIS), 2016 4th International Conference on, IEEE, pp. 160\u2013164, 2016.","key":"2025120523300492680_j_jisys-2017-0016_ref_016_w2aab3b7b1b1b6b1ab1b9c16Aa","DOI":"10.1109\/CCIS.2016.7790245"},{"doi-asserted-by":"crossref","unstructured":"T. Joachims, Text categorization with support vector machines: learning with many relevant features, in: European Conference on Machine Learning, 1998.","key":"2025120523300492680_j_jisys-2017-0016_ref_017_w2aab3b7b1b1b6b1ab1b9c17Aa","DOI":"10.1007\/BFb0026683"},{"unstructured":"D. D. Lewis, Y. Yang, T. G. Rose and F. Li, Rcv1: a new benchmark collection for text categorization research, J. Mach. Learn. Res.5 (2004), 361\u2013397.","key":"2025120523300492680_j_jisys-2017-0016_ref_018_w2aab3b7b1b1b6b1ab1b9c18Aa"},{"unstructured":"P. Li, J. Jiang and Y. Wang, Generating templates of entity summaries with an entity-aspect model and pattern mining, in: Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, pp. 640\u2013649, 2010.","key":"2025120523300492680_j_jisys-2017-0016_ref_019_w2aab3b7b1b1b6b1ab1b9c19Aa"},{"unstructured":"M. R. Lucas and D. Downey, Scaling semi-supervised naive Bayes with feature marginals, in: Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, pp. 343\u2013351, 2013.","key":"2025120523300492680_j_jisys-2017-0016_ref_020_w2aab3b7b1b1b6b1ab1b9c20Aa"},{"doi-asserted-by":"crossref","unstructured":"K. Nigam, A. K. McCallum, S. Thrun and T. Mitchell, Text classification from labeled and unlabeled documents using EM, Mach. Learn.39 (2000), 103\u2013134.10.1023\/A:1007692713085","key":"2025120523300492680_j_jisys-2017-0016_ref_021_w2aab3b7b1b1b6b1ab1b9c21Aa","DOI":"10.1023\/A:1007692713085"},{"doi-asserted-by":"crossref","unstructured":"M. Paul and R. Girju, A two-dimensional topic-aspect model for discovering multi-faceted topics, AAAI51 (2010), 61801.","key":"2025120523300492680_j_jisys-2017-0016_ref_022_w2aab3b7b1b1b6b1ab1b9c22Aa","DOI":"10.1609\/aaai.v24i1.7669"},{"unstructured":"J. D. Rennie, L. Shih, J. Teevan and D. R. Karger, Tackling the poor assumptions of naive Bayes text classifiers, in: Proceedings of the International Conference on Machine Learning, 3, pp. 616\u2013623, 2003.","key":"2025120523300492680_j_jisys-2017-0016_ref_023_w2aab3b7b1b1b6b1ab1b9c23Aa"},{"unstructured":"J. Su, J. S. Shirab and S. Matwin, Large scale text classification using semi-supervised multinomial naive Bayes, in: Proceedings of the 28th International Conference on Machine Learning (ICML-11), pp. 97\u2013104, 2011.","key":"2025120523300492680_j_jisys-2017-0016_ref_024_w2aab3b7b1b1b6b1ab1b9c24Aa"},{"unstructured":"H. M. Wallach, D. M. Mimno and A. McCallum, Rethinking LDA: Why priors matter, in: NIPS, 22, pp. 1973\u20131981, 2009.","key":"2025120523300492680_j_jisys-2017-0016_ref_025_w2aab3b7b1b1b6b1ab1b9c25Aa"},{"doi-asserted-by":"crossref","unstructured":"S. Wang, L. Jiang and C. Li, A CFS-Based Feature Weighting Approach to Naive Bayes Text Classifiers, Springer International Publishing, Cham, pp. 555\u2013562, 2014.","key":"2025120523300492680_j_jisys-2017-0016_ref_026_w2aab3b7b1b1b6b1ab1b9c26Aa","DOI":"10.1007\/978-3-319-11179-7_70"},{"doi-asserted-by":"crossref","unstructured":"S. Wang, L. Jiang and C. Li, Adapting naive Bayes tree for text classification, Knowl. Inf. Syst.44 (2015), 77\u201389.10.1007\/s10115-014-0746-y","key":"2025120523300492680_j_jisys-2017-0016_ref_027_w2aab3b7b1b1b6b1ab1b9c27Aa","DOI":"10.1007\/s10115-014-0746-y"},{"doi-asserted-by":"crossref","unstructured":"Y. Yang and X. Liu, A re-examination of text categorization methods, in: ACM SIGIR Conference on Research and Development in Information Retrieval, pp. 42\u201349, 1999.","key":"2025120523300492680_j_jisys-2017-0016_ref_028_w2aab3b7b1b1b6b1ab1b9c28Aa","DOI":"10.1145\/312624.312647"},{"doi-asserted-by":"crossref","unstructured":"C. Zhai and J. Lafferty, A study of smoothing methods for language models applied to ad hoc information retrieval, in: Proceedings of the 24th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, ACM, pp. 334\u2013342, 2001.","key":"2025120523300492680_j_jisys-2017-0016_ref_029_w2aab3b7b1b1b6b1ab1b9c29Aa","DOI":"10.1145\/383952.384019"},{"doi-asserted-by":"crossref","unstructured":"L. Zhang, L. Jiang and C. Li, A new feature selection approach to naive Bayes text classifiers, Int. J. Pattern Recogn. Artif. Intell.30 (2016), 1650003.10.1142\/S0218001416500038","key":"2025120523300492680_j_jisys-2017-0016_ref_030_w2aab3b7b1b1b6b1ab1b9c30Aa","DOI":"10.1142\/S0218001416500038"},{"doi-asserted-by":"crossref","unstructured":"L. Zhang, L. Jiang, C. Li and G. Kong, Two feature weighting approaches for naive Bayes text classifiers, Knowl. Based Syst.100 (2016), 137\u2013144.10.1016\/j.knosys.2016.02.017","key":"2025120523300492680_j_jisys-2017-0016_ref_031_w2aab3b7b1b1b6b1ab1b9c31Aa","DOI":"10.1016\/j.knosys.2016.02.017"}],"container-title":["Journal of Intelligent Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/www.degruyter.com\/view\/j\/jisys.2019.28.issue-2\/jisys-2017-0016\/jisys-2017-0016.xml","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/jisys-2017-0016\/xml","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/jisys-2017-0016\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,5]],"date-time":"2025-12-05T23:30:49Z","timestamp":1764977449000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.degruyterbrill.com\/document\/doi\/10.1515\/jisys-2017-0016\/html"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,7,20]]},"references-count":31,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2017,7,20]]},"published-print":{"date-parts":[[2019,4,24]]}},"alternative-id":["10.1515\/jisys-2017-0016"],"URL":"https:\/\/doi.org\/10.1515\/jisys-2017-0016","relation":{},"ISSN":["2191-026X","0334-1860"],"issn-type":[{"type":"electronic","value":"2191-026X"},{"type":"print","value":"0334-1860"}],"subject":[],"published":{"date-parts":[[2017,7,20]]}}}