{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,30]],"date-time":"2026-06-30T21:37:28Z","timestamp":1782855448135,"version":"3.54.5"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2016,7,20]],"date-time":"2016-07-20T00:00:00Z","timestamp":1468972800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100000038","name":"Natural Sciences and Engineering Research Council of Canada","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100000038","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61432020 and 61472430"],"award-info":[{"award-number":["61432020 and 61472430"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2017,2,28]]},"abstract":"<jats:p>Large and sparse datasets with a lot of missing values are common in the big data era, such as user behaviors over a large number of items. Classification in such datasets is an important topic for machine learning and data mining. Practically, naive Bayes is still a popular classification algorithm for large sparse datasets, as its time and space complexity scales linearly with the size of non-missing values. However, several important questions about the behavior of naive Bayes are yet to be answered. For example, how different mechanisms of data missing, data sparsity, and the number of attributes systematically affect the learning curves and convergence? In this paper, we address several common data missing mechanisms and propose novel data generation methods based on these mechanisms. We generate large and sparse data systematically, and study the entire AUC (Area Under ROC Curve) learning curve and convergence behavior of naive Bayes. We not only have several important experiment observations, but also provide detailed theoretic studies. Finally, we summarize our empirical and theoretic results as an intuitive decision flowchart and a useful guideline for classifying large sparse datasets in practice.<\/jats:p>","DOI":"10.1145\/2948068","type":"journal-article","created":{"date-parts":[[2016,7,21]],"date-time":"2016-07-21T15:13:24Z","timestamp":1469114004000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["The Convergence Behavior of Naive Bayes on Large Sparse Datasets"],"prefix":"10.1145","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9595-8526","authenticated-orcid":false,"given":"Xiang","family":"Li","sequence":"first","affiliation":[{"name":"University of Western Ontario, Ontario, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Charles X.","family":"Ling","sequence":"additional","affiliation":[{"name":"University of Western Ontario, Ontario, Canada"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huaimin","family":"Wang","sequence":"additional","affiliation":[{"name":"National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2016,7,20]]},"reference":[{"key":"e_1_2_1_1_1","volume-title":"Proceedings of KDD Cup and Workshop","author":"Bennett James","year":"2007","unstructured":"James Bennett and Stan Lanning . 2007 . The netflix prize . In Proceedings of KDD Cup and Workshop 2007. 35. James Bennett and Stan Lanning. 2007. The netflix prize. In Proceedings of KDD Cup and Workshop 2007. 35."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2013.10.016"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-7908-2604-3_16"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526806"},{"key":"e_1_2_1_5_1","volume-title":"Proceedings of the 28th AAAI Conference on Artificial Intelligence (AAAI\u201914)","author":"Chen Ning","year":"2014","unstructured":"Ning Chen , Jun Zhu , Jianfei Chen , and Bo Zhang . 2014 . Dropout training for support vector machines . In Proceedings of the 28th AAAI Conference on Artificial Intelligence (AAAI\u201914) . Ning Chen, Jun Zhu, Jianfei Chen, and Bo Zhang. 2014. Dropout training for support vector machines. In Proceedings of the 28th AAAI Conference on Artificial Intelligence (AAAI\u201914)."},{"key":"e_1_2_1_6_1","first-page":"313","article-title":"AUC optimization vs. error rate minimization","volume":"16","author":"Cortes Corinna","year":"2004","unstructured":"Corinna Cortes and Mehryar Mohri . 2004 . AUC optimization vs. error rate minimization . Adv. Neural Inf. Process. Syst. 16 , 16 (2004), 313 -- 320 . Corinna Cortes and Mehryar Mohri. 2004. AUC optimization vs. error rate minimization. Adv. Neural Inf. Process. Syst. 16, 16 (2004), 313--320.","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007413511361"},{"key":"e_1_2_1_8_1","first-page":"275","article-title":"Bayesian nonparametric poisson factorization for recommendation systems","volume":"33","author":"Gopalan Prem","year":"2014","unstructured":"Prem Gopalan , Francisco J. R. Ruiz , Rajesh Ranganath , and David M. Blei . 2014 . Bayesian nonparametric poisson factorization for recommendation systems . Artif. Intell. Stat. 33 (2014), 275 -- 283 . Prem Gopalan, Francisco J. R. Ruiz, Rajesh Ranganath, and David M. Blei. 2014. Bayesian nonparametric poisson factorization for recommendation systems. Artif. Intell. Stat. 33 (2014), 275--283.","journal-title":"Artif. Intell. Stat."},{"key":"e_1_2_1_9_1","first-page":"385","article-title":"Idiot\u2019s Bayes not so stupid after all","volume":"69","author":"Hand David J.","year":"2001","unstructured":"David J. Hand and Keming Yu . 2001 . Idiot\u2019s Bayes not so stupid after all ? Int. Stat. Rev. 69 , 3 (2001), 385 -- 398 . David J. Hand and Keming Yu. 2001. Idiot\u2019s Bayes not so stupid after all? Int. Stat. Rev. 69, 3 (2001), 385--398.","journal-title":"Int. Stat. Rev."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1148\/radiology.143.1.7063747"},{"key":"e_1_2_1_11_1","volume-title":"Proceedings of the 31st International Conference on Machine Learning (ICML\u201914)","author":"Jose M.","year":"2014","unstructured":"Jose M. Hernandez-lobato, Neil Houlsby , and Zoubin Ghahramani . 2014 a. Probabilistic matrix factorization with non-random missing data . In Proceedings of the 31st International Conference on Machine Learning (ICML\u201914) . 1512--1520. Jose M. Hernandez-lobato, Neil Houlsby, and Zoubin Ghahramani. 2014a. Probabilistic matrix factorization with non-random missing data. In Proceedings of the 31st International Conference on Machine Learning (ICML\u201914). 1512--1520."},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the 31st International Conference on Machine Learning (ICML\u201914)","author":"Jose M.","year":"2014","unstructured":"Jose M. Hernandez-lobato, Neil Houlsby , and Zoubin Ghahramani . 2014 b. Stochastic inference for scalable probabilistic modeling of binary matrices . In Proceedings of the 31st International Conference on Machine Learning (ICML\u201914) . 379--387. Jose M. Hernandez-lobato, Neil Houlsby, and Zoubin Ghahramani. 2014b. Stochastic inference for scalable probabilistic modeling of binary matrices. In Proceedings of the 31st International Conference on Machine Learning (ICML\u201914). 379--387."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2008.22"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2010.263"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/5326.661089"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-1-55860-335-6.50023-4"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1089\/big.2013.0037"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2004.12.037"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2006.180"},{"key":"e_1_2_1_20_1","volume-title":"13th International Conference on Machine Learning.","author":"Koller Daphne","year":"1995","unstructured":"Daphne Koller and Mehran Sahami . 1995 . Toward optimal feature selection . In 13th International Conference on Machine Learning. Daphne Koller and Mehran Sahami. 1995. Toward optimal feature selection. In 13th International Conference on Machine Learning."},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2009.263"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2015.53"},{"key":"e_1_2_1_23_1","volume-title":"Ling","author":"Li Xiang","year":"2015","unstructured":"Xiang Li , Huaimin Wang , Bin Gu , and Charles X . Ling . 2015 b. Data sparseness in linear SVM. In Proceedings of the 24th International Conference on Artificial Intelligence (IJCAI). AAAI Press , 3628--3634. Xiang Li, Huaimin Wang, Bin Gu, and Charles X. Ling. 2015b. Data sparseness in linear SVM. In Proceedings of the 24th International Conference on Artificial Intelligence (IJCAI). AAAI Press, 3628--3634."},{"key":"e_1_2_1_24_1","volume-title":"Rubin","author":"Little Roderick J. A.","year":"2014","unstructured":"Roderick J. A. Little and Donald B . Rubin . 2014 . Statistical Analysis with Missing Data. John Wiley & Sons . Roderick J. A. Little and Donald B. Rubin. 2014. Statistical Analysis with Missing Data. John Wiley & Sons."},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the 30th International Conference on Machine Learning (ICML\u201913)","author":"Maaten Laurens","unstructured":"Laurens Maaten , Minmin Chen , Stephen Tyree , and Kilian Q. Weinberger . 2013. Learning with marginalized corrupted features . In Proceedings of the 30th International Conference on Machine Learning (ICML\u201913) . 410--418. Laurens Maaten, Minmin Chen, Stephen Tyree, and Kilian Q. Weinberger. 2013. Learning with marginalized corrupted features. In Proceedings of the 30th International Conference on Machine Learning (ICML\u201913). 410--418."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the 23rd Conference on Uncertainty in Artificial Intelligence (UAI).","author":"Marlin Benjamin M.","year":"2007","unstructured":"Benjamin M. Marlin , Richard S. Zemel , Sam Roweis , and Malcolm Slaney . 2007 . Collaborative filtering and the missing at random assumption . In Proceedings of the 23rd Conference on Uncertainty in Artificial Intelligence (UAI). Benjamin M. Marlin, Richard S. Zemel, Sam Roweis, and Malcolm Slaney. 2007. Collaborative filtering and the missing at random assumption. In Proceedings of the 23rd Conference on Uncertainty in Artificial Intelligence (UAI)."},{"key":"e_1_2_1_27_1","volume-title":"Roweis","author":"Meeds Edward","year":"2006","unstructured":"Edward Meeds , Zoubin Ghahramani , Radford M. Neal , and Sam T . Roweis . 2006 . Modeling dyadic data with binary latent factors. In Proceedings of Advances in Neural Information Processing Systems . 977--984. Edward Meeds, Zoubin Ghahramani, Radford M. Neal, and Sam T. Roweis. 2006. Modeling dyadic data with binary latent factors. In Proceedings of Advances in Neural Information Processing Systems. 977--984."},{"key":"e_1_2_1_29_1","volume-title":"Proceedings ofAdvances in Neural Information Processing Systems (NIPS\u201907)","author":"Mnih Andriy","year":"2007","unstructured":"Andriy Mnih and Ruslan Salakhutdinov . 2007 . Probabilistic matrix factorization . In Proceedings ofAdvances in Neural Information Processing Systems (NIPS\u201907) . 1257--1264. Andriy Mnih and Ruslan Salakhutdinov. 2007. Probabilistic matrix factorization. In Proceedings ofAdvances in Neural Information Processing Systems (NIPS\u201907). 1257--1264."},{"key":"e_1_2_1_31_1","first-page":"841","article-title":"On discriminative vs. generative classifiers: A comparison of logistic regression and Naive","volume":"2","author":"Ng Andrew Y.","year":"2002","unstructured":"Andrew Y. Ng and Michael I. Jordan . 2002 . On discriminative vs. generative classifiers: A comparison of logistic regression and Naive Bayes. Adv. Neural Inf. Process. Syst. 2 (2002), 841 -- 848 . Andrew Y. Ng and Michael I. Jordan. 2002. On discriminative vs. generative classifiers: A comparison of logistic regression and Naive Bayes. Adv. Neural Inf. Process. Syst. 2 (2002), 841--848.","journal-title":"Bayes. Adv. Neural Inf. Process. Syst."},{"key":"e_1_2_1_32_1","volume-title":"Database and Expert Systems Applications","author":"Prinzie Anita","unstructured":"Anita Prinzie and Dirk Van den Poel . 2007. Random multiclass classification: Generalizing random forests to random MNL and random NB . In Database and Expert Systems Applications . Springer , 349--358. Anita Prinzie and Dirk Van den Poel. 2007. Random multiclass classification: Generalizing random forests to random MNL and random NB. In Database and Expert Systems Applications. Springer, 349--358."},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of International Conference on Knowledge Discovery and Data Mining (KDD\u201997)","author":"Provost Foster J.","year":"1997","unstructured":"Foster J. Provost , Tom Fawcett , and others. 1997 . Analysis and visualization of classifier performance: Comparison under imprecise class and cost distributions . In Proceedings of International Conference on Knowledge Discovery and Data Mining (KDD\u201997) . 43--48. Foster J. Provost, Tom Fawcett, and others. 1997. Analysis and visualization of classifier performance: Comparison under imprecise class and cost distributions. In Proceedings of International Conference on Knowledge Discovery and Data Mining (KDD\u201997). 43--48."},{"key":"e_1_2_1_34_1","volume-title":"Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence. AUAI Press, 452--461","author":"Rendle Steffen","year":"2009","unstructured":"Steffen Rendle , Christoph Freudenthaler , Zeno Gantner , and Lars Schmidt-Thieme . 2009 . BPR: Bayesian personalized ranking from implicit feedback . In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence. AUAI Press, 452--461 . Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2009. BPR: Bayesian personalized ranking from implicit feedback. In Proceedings of the 25th Conference on Uncertainty in Artificial Intelligence. AUAI Press, 452--461."},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of Advances in Neural Information Processing Systems. 658--666","author":"Ridgway James","year":"2014","unstructured":"James Ridgway , Pierre Alquier , Nicolas Chopin , and Feng Liang . 2014 . PAC-Bayesian AUC classification and scoring . In Proceedings of Advances in Neural Information Processing Systems. 658--666 . James Ridgway, Pierre Alquier, Nicolas Chopin, and Feng Liang. 2014. PAC-Bayesian AUC classification and scoring. In Proceedings of Advances in Neural Information Processing Systems. 658--666."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1093\/biomet\/63.3.581"},{"key":"e_1_2_1_37_1","volume-title":"Proceedings of International Conference on Knowledge Discovery and Data Mining (KDD\u201996)","author":"Sahami Mehran","year":"1996","unstructured":"Mehran Sahami . 1996 . Learning limited dependence bayesian classifiers . In Proceedings of International Conference on Knowledge Discovery and Data Mining (KDD\u201996) . 335--338. Mehran Sahami. 1996. Learning limited dependence bayesian classifiers. In Proceedings of International Conference on Knowledge Discovery and Data Mining (KDD\u201996). 335--338."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/1390156.1390267"},{"key":"e_1_2_1_39_1","volume-title":"WEBKDD 2000 Workshop. ACM.","author":"Sarwar Badrul M.","unstructured":"Badrul M. Sarwar , George Karypis , Joseph A. Konstan , and John T. Riedl . 2000. Application of Dimensionality Reduction in Recommender System -- A Case Study . WEBKDD 2000 Workshop. ACM. Badrul M. Sarwar, George Karypis, Joseph A. Konstan, and John T. Riedl. 2000. Application of Dimensionality Reduction in Recommender System -- A Case Study. WEBKDD 2000 Workshop. ACM."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-010-0420-4"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2011.181"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/1835804.1835895"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/2487575.2487627"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-3264-1"},{"key":"e_1_2_1_45_1","volume-title":"Liang","author":"Wager Stefan","year":"2014","unstructured":"Stefan Wager , William Fithian , Sida Wang , and Percy S . Liang . 2014 . Altitude training: Strong bounds for single-layer dropout. In Proceedings of Advances in Neural Information Processing Systems . 100--108. Stefan Wager, William Fithian, Sida Wang, and Percy S. Liang. 2014. Altitude training: Strong bounds for single-layer dropout. In Proceedings of Advances in Neural Information Processing Systems. 100--108."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/1060745.1060754"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2948068","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/2948068","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:55:43Z","timestamp":1750222543000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/2948068"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2016,7,20]]},"references-count":44,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2017,2,28]]}},"alternative-id":["10.1145\/2948068"],"URL":"https:\/\/doi.org\/10.1145\/2948068","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"value":"1556-4681","type":"print"},{"value":"1556-472X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2016,7,20]]},"assertion":[{"value":"2015-12-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2016-07-20","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}