{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,19]],"date-time":"2025-09-19T08:57:29Z","timestamp":1758272249589},"reference-count":33,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2019,12,9]]},"abstract":"<jats:p>We introduce the problem of anti-knowledge mining. Our goal is to create an \"anti-knowledge base\" that contains factual mistakes. The resulting data can be used for analysis, training, and benchmarking in the research domain of automated fact checking. Prior data sets feature manually generated fact checks of famous misclaims. Instead, we focus on the long tail of factual mistakes made by Web authors, ranging from erroneous sports results to incorrect capitals.<\/jats:p><jats:p>We mine mistakes automatically, by an unsupervised approach, from Wikipedia updates that correct factual mistakes. Identifying such updates (only a small fraction of the total number of updates) is one of the primary challenges. We mine anti-knowledge by a multi-step pipeline. First, we filter out candidate updates via several simple heuristics. Next, we correlate Wikipedia updates with other statements made on the Web. Using claim occurrence frequencies as input to a probabilistic model, we infer the likelihood of corrections via an iterative expectation-maximization approach. Finally, we extract mistakes in the form of subject-predicate-object triples and rank them according to several criteria. Our end result is a data set containing over 110,000 ranked mistakes with a precision of 85% in the top 1% and a precision of over 60% in the top 25%. We demonstrate that baselines achieve significantly lower precision. Also, we exploit our data to verify several hypothesis on why users make mistakes. We finally show that the AKB can be used to find mistakes on the entire Web.<\/jats:p>","DOI":"10.14778\/3372716.3372727","type":"journal-article","created":{"date-parts":[[2020,1,6]],"date-time":"2020-01-06T20:19:37Z","timestamp":1578341977000},"page":"561-573","source":"Crossref","is-referenced-by-count":15,"title":["Mining an \"anti-knowledge base\" from Wikipedia updates with applications to fact checking and beyond"],"prefix":"10.14778","volume":"13","author":[{"given":"Georgios","family":"Karagiannis","sequence":"first","affiliation":[{"name":"Cornell University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Immanuel","family":"Trummer","sequence":"additional","affiliation":[{"name":"Cornell University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Saehan","family":"Jo","sequence":"additional","affiliation":[{"name":"Cornell University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shubham","family":"Khandelwal","sequence":"additional","affiliation":[{"name":"Cornell University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuezhi","family":"Wang","sequence":"additional","affiliation":[{"name":"Google"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cong","family":"Yu","sequence":"additional","affiliation":[{"name":"Google"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,1,6]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.3115\/v1\/P15-1034"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.5555\/2892753.2892959"},{"key":"e_1_2_1_3_1","first-page":"127","volume-title":"WWW","author":"Bonifati A.","year":"2019","unstructured":"A. Bonifati , W. Martens , and T. Timm . Navigating the maze of wikidata query logs . In WWW , pages 127 -- 138 , 2019 . A. Bonifati, W. Martens, and T. Timm. Navigating the maze of wikidata query logs. In WWW, pages 127--138, 2019."},{"key":"e_1_2_1_4_1","volume-title":"Maximum likelihood estimation of observer error-rates using the em algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1):20--28","author":"Dawid A. P.","year":"1979","unstructured":"A. P. Dawid and A. M. Skene . Maximum likelihood estimation of observer error-rates using the em algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1):20--28 , 1979 . A. P. Dawid and A. M. Skene. Maximum likelihood estimation of observer error-rates using the em algorithm. Journal of the Royal Statistical Society: Series C (Applied Statistics), 28(1):20--28, 1979."},{"key":"e_1_2_1_5_1","first-page":"3097","volume-title":"NIPS","author":"De Sa C. M.","year":"2015","unstructured":"C. M. De Sa , C. Zhang , K. Olukotun , and C. R\u00e9 . Rapidly mixing gibbs sampling for a class of factor graphs using hierarchy width . In NIPS , pages 3097 -- 3105 , 2015 . C. M. De Sa, C. Zhang, K. Olukotun, and C. R\u00e9. Rapidly mixing gibbs sampling for a class of factor graphs using hierarchy width. In NIPS, pages 3097--3105, 2015."},{"key":"e_1_2_1_6_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin J.","year":"2018","unstructured":"J. Devlin , M.-W. Chang , K. Lee , and K. Toutanova . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 , 2018 . J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018."},{"key":"e_1_2_1_7_1","doi-asserted-by":"crossref","first-page":"100","DOI":"10.1145\/988672.988687","volume-title":"WWW","author":"Etzioni O.","year":"2004","unstructured":"O. Etzioni , M. J. Cafarella , D. Downey , S. Kok , A.-M. Popescu , T. Shaked , S. Soderland , D. S. Weld , and A. Yates . Web-scale information extraction in knowitall: (preliminary results) . In WWW , pages 100 -- 110 , 2004 . O. Etzioni, M. J. Cafarella, D. Downey, S. Kok, A.-M. Popescu, T. Shaked, S. Soderland, D. S. Weld, and A. Yates. Web-scale information extraction in knowitall: (preliminary results). In WWW, pages 100--110, 2004."},{"key":"e_1_2_1_8_1","first-page":"1535","volume-title":"EMNLP","author":"Fader A.","year":"2011","unstructured":"A. Fader , S. Soderland , and O. Etzioni . Identifying relations for open information extraction . In EMNLP , pages 1535 -- 1545 , 2011 . A. Fader, S. Soderland, and O. Etzioni. Identifying relations for open information extraction. In EMNLP, pages 1535--1545, 2011."},{"key":"e_1_2_1_9_1","volume-title":"The state of automated factchecking","year":"2016","unstructured":"FullFact. The state of automated factchecking . 2016 . FullFact. The state of automated factchecking. 2016."},{"key":"e_1_2_1_10_1","volume-title":"Proceedings of the 2015 Computation+Journalism Symposium","author":"Hassan N.","year":"2015","unstructured":"N. Hassan , B. Adair , J. T. Hamilton , C. Li , M. Tremayne , J. Yang , and C. Yu . The quest to automate fact-checking . In Proceedings of the 2015 Computation+Journalism Symposium , 2015 . N. Hassan, B. Adair, J. T. Hamilton, C. Li, M. Tremayne, J. Yang, and C. Yu. The quest to automate fact-checking. In Proceedings of the 2015 Computation+Journalism Symposium, 2015."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098131"},{"issue":"12","key":"e_1_2_1_12_1","first-page":"1945","article-title":"Claimbuster: the first-ever end-to-end fact-checking system","volume":"10","author":"Hassan N.","year":"2017","unstructured":"N. Hassan , G. Zhang , F. Arslan , J. Caraballo , D. Jimenez , S. Gawsane , S. Hasan , M. Joseph , A. Kulkarni , A. K. Nayak , Claimbuster: the first-ever end-to-end fact-checking system . VLDB , 10 ( 12 ): 1945 -- 1948 , 2017 . N. Hassan, G. Zhang, F. Arslan, J. Caraballo, D. Jimenez, S. Gawsane, S. Hasan, M. Joseph, A. Kulkarni, A. K. Nayak, et al. Claimbuster: the first-ever end-to-end fact-checking system. VLDB, 10(12):1945--1948, 2017.","journal-title":"VLDB"},{"key":"e_1_2_1_13_1","first-page":"299","volume-title":"SIGMOD","author":"Jo S.","year":"2018","unstructured":"S. Jo , I. Trummer , W. Yu , X. Wang , C. Yu , D. Liu , and N. Mehta . Verifying text summaries of relational data sets . In SIGMOD , pages 299 -- 316 , 2018 . S. Jo, I. Trummer, W. Yu, X. Wang, C. Yu, D. Liu, and N. Mehta. Verifying text summaries of relational data sets. In SIGMOD, pages 299--316, 2018."},{"key":"e_1_2_1_14_1","first-page":"539","volume-title":"WWW","author":"Li L.","year":"2017","unstructured":"L. Li , H. Deng , A. Dong , Y. Chang , R. Baeza-Yates , and H. Zha . Exploring query auto-completion and click logs for contextual-aware web search and query suggestion . In WWW , pages 539 -- 548 , 2017 . L. Li, H. Deng, A. Dong, Y. Chang, R. Baeza-Yates, and H. Zha. Exploring query auto-completion and click logs for contextual-aware web search and query suggestion. In WWW, pages 539--548, 2017."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/2588555.2610509"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/N19-1333"},{"issue":"5","key":"e_1_2_1_17_1","first-page":"635","article-title":"Domain-aware multi-truth discovery from conflicting sources","volume":"11","author":"Lin X.","year":"2018","unstructured":"X. Lin and L. Chen . Domain-aware multi-truth discovery from conflicting sources . VLDB , 11 ( 5 ): 635 -- 647 , 2018 . X. Lin and L. Chen. Domain-aware multi-truth discovery from conflicting sources. VLDB, 11(5):635--647, 2018.","journal-title":"VLDB"},{"key":"e_1_2_1_18_1","first-page":"25","article-title":"Deepdive: Web-scale knowledge-base construction using statistical learning and inference","volume":"12","author":"Niu F.","year":"2012","unstructured":"F. Niu , C. Zhang , C. R\u00e9 , and J. W. Shavlik . Deepdive: Web-scale knowledge-base construction using statistical learning and inference . VLDS , 12 : 25 -- 28 , 2012 . F. Niu, C. Zhang, C. R\u00e9, and J. W. Shavlik. Deepdive: Web-scale knowledge-base construction using statistical learning and inference. VLDS, 12:25--28, 2012.","journal-title":"VLDS"},{"key":"e_1_2_1_19_1","volume-title":"Understand Your World with Bing. https:\/\/blogs.bing.com\/search\/2013\/03\/21\/understand-your-world-with-bing\/","author":"Qian R.","year":"2013","unstructured":"R. Qian . Understand Your World with Bing. https:\/\/blogs.bing.com\/search\/2013\/03\/21\/understand-your-world-with-bing\/ , 2013 . R. Qian. Understand Your World with Bing. https:\/\/blogs.bing.com\/search\/2013\/03\/21\/understand-your-world-with-bing\/, 2013."},{"issue":"3","key":"e_1_2_1_20_1","first-page":"269","article-title":"Rapid training data creation with weak supervision","volume":"11","author":"Ratner A.","year":"2017","unstructured":"A. Ratner , S. H. Bach , H. Ehrenberg , J. Fries , S. Wu , and C. R\u00e9 . Snorkel : Rapid training data creation with weak supervision . VLDB , 11 ( 3 ): 269 -- 282 , 2017 . A. Ratner, S. H. Bach, H. Ehrenberg, J. Fries, S. Wu, and C. R\u00e9. Snorkel: Rapid training data creation with weak supervision. VLDB, 11(3):269--282, 2017.","journal-title":"VLDB"},{"key":"e_1_2_1_21_1","volume-title":"Learning from crowds. JMLR, 11(Apr):1297--1322","author":"Raykar V. C.","year":"2010","unstructured":"V. C. Raykar , S. Yu , L. H. Zhao , G. H. Valadez , C. Florin , L. Bogoni , and L. Moy . Learning from crowds. JMLR, 11(Apr):1297--1322 , 2010 . V. C. Raykar, S. Yu, L. H. Zhao, G. H. Valadez, C. Florin, L. Bogoni, and L. Moy. Learning from crowds. JMLR, 11(Apr):1297--1322, 2010."},{"issue":"11","key":"e_1_2_1_22_1","first-page":"1190","article-title":"Holistic data repairs with probabilistic inference","volume":"10","author":"Rekatsinas T.","year":"2017","unstructured":"T. Rekatsinas , X. Chu , I. F. Ilyas , and C. R\u00e9 . Holoclean : Holistic data repairs with probabilistic inference . VLDB , 10 ( 11 ): 1190 -- 1201 , 2017 . T. Rekatsinas, X. Chu, I. F. Ilyas, and C. R\u00e9. Holoclean: Holistic data repairs with probabilistic inference. VLDB, 10(11):1190--1201, 2017.","journal-title":"VLDB"},{"key":"e_1_2_1_23_1","first-page":"991","volume-title":"WWW","author":"Roitman H.","year":"2016","unstructured":"H. Roitman , S. Hummel , E. Rabinovich , B. Sznajder , N. Slonim , and E. Aharoni . On the retrieval of wikipedia articles containing claims on controversial topics . In WWW , pages 991 -- 996 , 2016 . H. Roitman, S. Hummel, E. Rabinovich, B. Sznajder, N. Slonim, and E. Aharoni. On the retrieval of wikipedia articles containing claims on controversial topics. In WWW, pages 991--996, 2016."},{"key":"e_1_2_1_24_1","volume-title":"https:\/\/schema.org\/ClaimReview","year":"2016","unstructured":"schema.org. ClaimReview. https:\/\/schema.org\/ClaimReview , 2016 . schema.org. ClaimReview. https:\/\/schema.org\/ClaimReview, 2016."},{"key":"e_1_2_1_25_1","first-page":"1033","volume-title":"VLDB","author":"Shen W.","year":"2007","unstructured":"W. Shen , A. Doan , J. F. Naughton , and R. Ramakrishnan . Declarative information extraction using datalog with embedded extraction predicates . VLDB , pages 1033 -- 1044 , 2007 . W. Shen, A. Doan, J. F. Naughton, and R. Ramakrishnan. Declarative information extraction using datalog with embedded extraction predicates. VLDB, pages 1033--1044, 2007."},{"key":"e_1_2_1_26_1","volume-title":"Introducing the Knowledge Graph: things, not strings. https:\/\/googleblog.blogspot.com\/2012\/05\/introducing-knowledge-graph-things-not.html","author":"Singhal A.","year":"2012","unstructured":"A. Singhal . Introducing the Knowledge Graph: things, not strings. https:\/\/googleblog.blogspot.com\/2012\/05\/introducing-knowledge-graph-things-not.html , 2012 . A. Singhal. Introducing the Knowledge Graph: things, not strings. https:\/\/googleblog.blogspot.com\/2012\/05\/introducing-knowledge-graph-things-not.html, 2012."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526794"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2723372.2750548"},{"key":"e_1_2_1_29_1","first-page":"525","volume-title":"Fact Checking Track","author":"Wang X.","year":"2018","unstructured":"X. Wang , C. Yu , S. Baumgartner , and F. Korn . Relevant document discovery for fact-checking articles. In WWW, Journalism, Misinformation , Fact Checking Track , pages 525 -- 533 , 2018 . X. Wang, C. Yu, S. Baumgartner, and F. Korn. Relevant document discovery for fact-checking articles. In WWW, Journalism, Misinformation, Fact Checking Track, pages 525--533, 2018."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1213"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/1281192.1281309"},{"issue":"5","key":"e_1_2_1_32_1","first-page":"541","article-title":"Truth inference in crowdsourcing: Is the problem solved?","volume":"10","author":"Zheng Y.","year":"2017","unstructured":"Y. Zheng , G. Li , Y. Li , C. Shan , and R. Cheng . Truth inference in crowdsourcing: Is the problem solved? VLDB , 10 ( 5 ): 541 -- 552 , 2017 . Y. Zheng, G. Li, Y. Li, C. Shan, and R. Cheng. Truth inference in crowdsourcing: Is the problem solved? VLDB, 10(5):541--552, 2017.","journal-title":"VLDB"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1526709.1526724"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3372716.3372727","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,24]],"date-time":"2023-09-24T19:54:45Z","timestamp":1695585285000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3372716.3372727"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,12,9]]},"references-count":33,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2019,12,9]]}},"alternative-id":["10.14778\/3372716.3372727"],"URL":"https:\/\/doi.org\/10.14778\/3372716.3372727","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2019,12,9]]}}}