{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T15:07:00Z","timestamp":1781104020950,"version":"3.54.1"},"reference-count":73,"publisher":"IGI Global Scientific Publishing","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2012,7,1]]},"abstract":"<p>Researchers have approached knowledge-base construction (KBC) with a wide range of data resources and techniques. The authors present Elementary, a prototype KBC system that is able to combine diverse resources and different KBC techniques via machine learning and statistical inference to construct knowledge bases. Using Elementary, they have implemented a solution to the TAC-KBP challenge with quality comparable to the state of the art, as well as an end-to-end online demonstration that automatically and continuously enriches Wikipedia with structured data by reading millions of webpages on a daily basis. The authors describe several challenges and their solutions in designing, implementing, and deploying Elementary. In particular, the authors first describe the conceptual framework and architecture of Elementary to integrate different data resources and KBC techniques in a principled manner. They then discuss how they address scalability challenges to enable Web-scale deployment. The authors empirically show that this decomposition-based inference approach achieves higher performance than prior inference approaches. To validate the effectiveness of Elementary\u2019s approach to KBC, they experimentally show that its ability to incorporate diverse signals has positive impacts on KBC quality.<\/p>","DOI":"10.4018\/jswis.2012070103","type":"journal-article","created":{"date-parts":[[2013,1,14]],"date-time":"2013-01-14T16:57:39Z","timestamp":1358182659000},"page":"42-73","source":"Crossref","is-referenced-by-count":59,"title":["Elementary"],"prefix":"10.4018","volume":"8","author":[{"given":"Feng","family":"Niu","sequence":"first","affiliation":[{"name":"Computer Sciences Department, University of Wisconsin-Madison, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ce","family":"Zhang","sequence":"additional","affiliation":[{"name":"Computer Sciences Department, University of Wisconsin-Madison, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Christopher","family":"R\u00e9","sequence":"additional","affiliation":[{"name":"Computer Sciences Department, University of Wisconsin-Madison, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jude","family":"Shavlik","sequence":"additional","affiliation":[{"name":"Computer Sciences Department, University of Wisconsin-Madison, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2432","reference":[{"key":"jswis.2012070103-0","doi-asserted-by":"crossref","unstructured":"Ailon, N., Charikar, M., & Newman, A. (2008). Aggregating inconsistent information: Ranking and clustering. Journal of the ACM, 55(23), 1-23-27.","DOI":"10.1145\/1411509.1411513"},{"key":"jswis.2012070103-1","unstructured":"Andrzejewski, D., Livermore, L., Zhu, X., Craven, M., & Recht, B. (2011, July 16-22). A framework for incorporating general domain knowledge into latent Dirichlet allocation using first-order logic. In Proceedings of the International Joint Conference on Artificial Intelligence, Catalonia, Spain (pp. 1171-1177)."},{"key":"jswis.2012070103-2","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-007-0148-y"},{"key":"jswis.2012070103-3","doi-asserted-by":"crossref","unstructured":"Arasu, A., & Garcia-Molina, H. (2003). Extracting structured data from web pages. In Proceedings of the 2003 ACM SIGMOD International Conference on Management of Data, San Diego, CA (pp. 337-348).","DOI":"10.1145\/872757.872799"},{"key":"jswis.2012070103-4","doi-asserted-by":"crossref","unstructured":"Arasu, A., R\u00e9, C., & Suciu, D. (2009, March 29-April 2). Large-scale deduplication with constraints using Dedupalog. In Proceedings of the International Conference on Data Engineering, Shanghai, China (pp. 952-963).","DOI":"10.1109\/ICDE.2009.43"},{"key":"jswis.2012070103-5","author":"D.Bertsekas","year":"1999","journal-title":"Nonlinear programming"},{"key":"jswis.2012070103-6","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511804441","author":"S.Boyd","year":"2004","journal-title":"Convex optimization"},{"key":"jswis.2012070103-7","doi-asserted-by":"crossref","unstructured":"Brin, S. (1999, May 11-14). Extracting patterns and relations from the world wide web. In Proceedings of the International Conference on World Wide Web, Toronto, Canada (pp. 172-183).","DOI":"10.1007\/10704656_11"},{"key":"jswis.2012070103-8","first-page":"347","article-title":"Query answering under expressive entity-relationship schemata.","volume":"1","author":"A.Cal\u00ec","year":"2010","journal-title":"Conceptual Modeling-ER"},{"key":"jswis.2012070103-9","doi-asserted-by":"crossref","unstructured":"Carlson, A., Betteridge, J., Kisiel, B., Settles, B., Hruschka, E., Jr., & Mitchell, T. (2010, July 11-15). Toward an architecture for never-ending language learning. In Proceedings of the Conference on Artificial Intelligence, Atlanta, GA (pp. 1306-1313).","DOI":"10.1609\/aaai.v24i1.7519"},{"key":"jswis.2012070103-10","doi-asserted-by":"crossref","unstructured":"Chen, F., Feng, X., Christopher, R., & Wang, M. (2012, April 1-5). Optimizing statistical information extraction programs over evolving text. In Proceedings of the International Conference on Data Engineering, Washington, DC (pp. 870-881).","DOI":"10.1109\/ICDE.2012.60"},{"key":"jswis.2012070103-11","doi-asserted-by":"publisher","DOI":"10.1145\/320434.320440"},{"key":"jswis.2012070103-12","unstructured":"Chiticariu, L., Krishnamurthy, R., Li, Y., Raghavan, S., Reiss, F., & Vaithyanathan, S. (2010, July 11-16). SystemT: An algebraic approach to declarative information extraction. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Uppsala, Sweden (pp. 128-137)."},{"key":"jswis.2012070103-13","unstructured":"Dean, J., & Ghemawat, S. (2004, October 3). MapReduce: Simplified data processing on large clusters. In Proceeding of the USENIX Symposium on Operating Systems Design and Implementation, San Francisco, CA (pp. 137-150)."},{"key":"jswis.2012070103-14","unstructured":"Derose, P., & Shen, W. Fei. C., Lee, Y., Burdick, D., Doan, A., & Ramakrishnan, R. (2007, January 7-10). DBLife: A community information management platform for the database research community (demo). In Proceedings of the Third Biennial Conference on Innovative Data Systems Research, Asilomar, CA (pp. 169\u2013172)."},{"key":"jswis.2012070103-15","unstructured":"Dredze, M., McNamee, P., Rao, D., Gerber, A., & Finin, T. (2010, July 11-16). Entity disambiguation for knowledge base population. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Uppsala, Sweden (pp. 277-285)."},{"key":"jswis.2012070103-16","doi-asserted-by":"crossref","unstructured":"Fang, L., Sarma, A. D., Yu, C., & Bohannon, P. (2011). Rex: Explaining relationships between pairs. In Proceedings of the Very Large Data Base Endowment, 5(3), 241-252.","DOI":"10.14778\/2078331.2078339"},{"key":"jswis.2012070103-17","doi-asserted-by":"crossref","unstructured":"Finkel, J., Grenager, T., & Manning, C. (2005, June 25-30). Incorporating non-local information into information extraction systems by Gibbs sampling. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Ann Arbor, MI (pp. 363-370).","DOI":"10.3115\/1219840.1219885"},{"key":"jswis.2012070103-18","unstructured":"Friedman, N., Getoor, L., Koller, D., & Pfeffer, A. (1999, July 31-August 6). Learning probabilistic relational models. In Proceedings of the International Joint Conference on Artificial Intelligence, Stockholm, Sweden (pp. 307-333)."},{"key":"jswis.2012070103-19","doi-asserted-by":"crossref","unstructured":"Hearst, M. (1992, August 23-28). Automatic acquisition of hyponyms from large text corpora. In Proceedings of the 14th Conference on Computational Linguistics, Nantes, France (Vol. 2, pp. 539-545).","DOI":"10.3115\/992133.992154"},{"key":"jswis.2012070103-20","unstructured":"Hoffart, J., Yosef, M. A., Bordino, I., F\u00fcrstenau, H., Pin, M., Spaniol, M., et al. (2011). Robust disambiguation of named entities in text. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Edinburgh, UK (pp. 782-792)."},{"key":"jswis.2012070103-21","unstructured":"Hoffmann, R., Zhang, C., Ling, X., Zettlemoyer, L., & Weld, D. (2011, June 19-24). Knowledge-based weak supervision for information extraction of overlapping relations. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Portland, OR (pp. 541-550)."},{"issue":"1","key":"jswis.2012070103-22","first-page":"905","article-title":"Loopy belief propagation: Convergence and effects of message errors.","volume":"6","author":"A.Ihler","year":"2006","journal-title":"Journal of Machine Learning Research"},{"key":"jswis.2012070103-23","unstructured":"Ji, H., Grishman, R., Dang, H., Griftt, K., & Ellis, J. (2010, November 15-16). Overview of the TAC 2010 knowledge base population track. In Text Analysis Conference, Gaithersburg, MD."},{"key":"jswis.2012070103-24","doi-asserted-by":"publisher","DOI":"10.1145\/1519103.1519110"},{"key":"jswis.2012070103-25","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1090\/dimacs\/035\/15","article-title":"A general stochastic approach to solving problems with hard and soft constraints.","volume":"17","author":"H.Kautz","year":"1997","journal-title":"The Satisfiability Problem: Theory and Applications"},{"key":"jswis.2012070103-26","doi-asserted-by":"crossref","unstructured":"Kok, S., & Domingos, P. (2005, August 7-11). Learning the structure of Markov logic networks. In Proceedings of the 22nd International Conference on Machine Learning, Bonn, Germany (pp. 441-448).","DOI":"10.1145\/1102351.1102407"},{"key":"jswis.2012070103-27","author":"D.Koller","year":"2009","journal-title":"Probabilistic Graphical Models: Principles and Techniques"},{"key":"jswis.2012070103-28","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2006.200"},{"key":"jswis.2012070103-29","doi-asserted-by":"crossref","unstructured":"Komodakis, N., & Paragios, N. (2009, June 20-25). Beyond pairwise energies: Efficient optimization for higher-order MRFs. In Proceeding of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL (pp. 2985-2992).","DOI":"10.1109\/CVPR.2009.5206846"},{"key":"jswis.2012070103-30","doi-asserted-by":"crossref","unstructured":"Komodakis, N., Paragios, N., & Tziritas, G. (2007, October 14-21). MRF optimization via dual decomposition: Message-passing revisited. In Proceeding of the IEEE International Conference on Computer Vision, Rio de Janeiro, Brazil (pp. 1-8).","DOI":"10.1109\/ICCV.2007.4408890"},{"key":"jswis.2012070103-31","unstructured":"Lafferty, J., McCallum, A., & Pereira, F. (2001, June 28-July 1). Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proceedings of the International Conference on Machine Learning, Williamstown, MA (pp. 282-289)."},{"key":"jswis.2012070103-32","unstructured":"Lao, N., Mitchell, T., & Cohen, W. (2011, July 27-31). Random walk inference and learning in a large scale knowledge base. In Proceedings of the Conference on Empirical Methods in Natural Language Processing, Edinburgh, UK (pp. 529-539)."},{"key":"jswis.2012070103-33","first-page":"323","article-title":"DIRT-discovery of inference rules from text.","volume":"1","author":"D.Lin","year":"2001","journal-title":"Knowledge Discovery and Data Mining"},{"key":"jswis.2012070103-34","doi-asserted-by":"crossref","unstructured":"Liu, B., Chiticariu, L., Chu, V., Jagadish, H., & Reiss, F. (2010, September 13-17). Automatic rule refinement for information extraction. In Proceedings of the International Conference on Very Large Data Bases, Singapore (pp. 588-597).","DOI":"10.14778\/1920841.1920916"},{"key":"jswis.2012070103-35","doi-asserted-by":"crossref","unstructured":"Lowd, D., & Domingos, P. (2007, September 17-21). Efficient weight learning for Markov logic networks. In Proceedings of the European Conference on Principles of Data Mining and Knowledge Discovery, Warsaw, Poland (pp. 200-211).","DOI":"10.1007\/978-3-540-74976-9_21"},{"key":"jswis.2012070103-36","unstructured":"McCallum, A., Schultz, K., & Singh, S. (2009, December 7-10). Factorie: Probabilistic programming via imperatively defined factor graphs. In Proceedings of the Annual Conference on Neural Information Processing Systems, Vancouver, Canada (pp. 285-292)."},{"key":"jswis.2012070103-37","doi-asserted-by":"crossref","unstructured":"Michelakis, E., Krishnamurthy, R., Haas, P., & Vaithyanathan, S. (2009, June 28-July 2). Uncertainty management in rule-based information extraction systems. In Proceedings of the SIGMOD International Conference on Management of Data, Providence, RI (pp. 101-114).","DOI":"10.1145\/1559845.1559858"},{"key":"jswis.2012070103-38","unstructured":"Milch, B., Marthi, B., Russell, S., Sontag, D., Ong, D., & Kolobov, A. (2005, July 30-August 5). BLOG: Probabilistic models with unknown objects. In Proceedings of the International Joint Conference on Artificial Intelligence, Edinburgh, UK (pp. 1352-1359)."},{"key":"jswis.2012070103-39","unstructured":"Minka, T. (2001, August 2-5). Expectation propagation for approximate Bayesian inference. In Proceeding of the Conference on Uncertainty in Artificial Intelligence, Seattle, WA (pp. 362-369)."},{"key":"jswis.2012070103-40","doi-asserted-by":"crossref","unstructured":"Mintz, M., Bills, S., Snow, R., & Jurafsky, D. (2009, August 2-7). Distant supervision for relation extraction without labeled data. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Singapore (pp. 1003-1011).","DOI":"10.3115\/1690219.1690287"},{"key":"jswis.2012070103-41","unstructured":"Mooney, R. (1999, July 18-22). Relational learning of pattern-match rules for information extraction. In Proceedings of the Sixteenth National Conference on Artificial Intelligence, Orlando, FL (pp. 328-334)."},{"key":"jswis.2012070103-42","doi-asserted-by":"crossref","unstructured":"Nakashole, N., Theobald, M., & Weikum, G. (2011, February 9-12). Scalable knowledge harvesting with high precision and high recall. In Proceedings of the Web Search and Data Mining (pp. 227-236).","DOI":"10.1145\/1935826.1935869"},{"key":"jswis.2012070103-43","unstructured":"Niu, F. (2012). Web-scale Knowledge-base Construction via Statistical Inference and Learning. Unpublished doctoral dissertation, University of Wisconsin-Madison."},{"key":"jswis.2012070103-44","doi-asserted-by":"crossref","unstructured":"Niu, F., R\u00e9, C., Doan, A., & Shavlik, J. (2011, August 28-September 1). Tuffy: Scaling up statistical inference in Markov logic networks using an RDBMS. In Proceedings of the International Conference on Very Large Data Bases, Seattle, WA (pp. 373-384).","DOI":"10.14778\/1978665.1978669"},{"key":"jswis.2012070103-45","doi-asserted-by":"crossref","unstructured":"Niu, F., Zhang, C., R\u00e9, C., & Shavlik, J. (2012, December 10-13). Scaling inference for Markov logic via dual decomposition. In Proceedings of the IEEE International Conference on Data Mining, Brussels, Belgium.","DOI":"10.1109\/ICDM.2012.96"},{"key":"jswis.2012070103-46","unstructured":"Poon, H., & Domingos, P. (2007, July 22-26). Joint inference in information extraction. In Proceedings of the AAAI Conference on Artificial Intelligence, Vancouver, Canada (pp. 913-918)."},{"key":"jswis.2012070103-47","doi-asserted-by":"publisher","DOI":"10.1007\/s10994-006-5833-1"},{"key":"jswis.2012070103-48","unstructured":"Riedel, S. (2008, July 9-12). Improving the accuracy and efficiency of MAP inference for Markov logic. In Proceeding of the Conference on Uncertainty in Artificial Intelligence, Helsinki, Finland (pp. 468-475)."},{"key":"jswis.2012070103-49","unstructured":"Riloff, E. (1993, July 11-15). Automatically constructing a dictionary for information extraction tasks. In Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC (pp. 811-816)."},{"key":"jswis.2012070103-50","unstructured":"Rush, A., Sontag, D., Collins, M., & Jaakkola, T. (2010, October 9-11). On dual decomposition and linear programming relaxations for natural language processing. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (pp. 1-11)."},{"key":"jswis.2012070103-51","doi-asserted-by":"publisher","DOI":"10.1561\/1900000003"},{"key":"jswis.2012070103-52","doi-asserted-by":"crossref","unstructured":"Sen, P., Deshpande, A., & Getoor, L. (2009, August 28). PrDB: Managing and exploiting rich correlations in probabilistic databases. In Proceedings of the International Conference on Very Large Data Bases, Lyon, France (pp. 1065-1090).","DOI":"10.1007\/s00778-009-0153-2"},{"key":"jswis.2012070103-53","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-031-01560-1","author":"B.Settles","year":"2012","journal-title":"Active Learning"},{"key":"jswis.2012070103-54","unstructured":"Shen, W., Doan, A., Naughton, J., & Ramakrishnan, R. (2007, September 23-27). Declarative information extraction using datalog with embedded extraction predicates. In Proceedings of the International Conference on Very Large Data Bases, Vienna, Austria."},{"key":"jswis.2012070103-55","unstructured":"Singh, S., Subramanya, A., Pereira, F., & McCallum, A. (2011, June 19-24). Large-scale cross-document coreference using distributed inference and hierarchical models. In Proceeding of the Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Portland, OR (pp. 793-803)."},{"key":"jswis.2012070103-56","first-page":"1","article-title":"Introduction to dual decomposition for inference.","volume":"1","author":"D.Sontag","year":"2010","journal-title":"Optimization for Machine Learning"},{"key":"jswis.2012070103-57","doi-asserted-by":"crossref","unstructured":"Suchanek, F., Kasneci, G., & Weikum, G. (2007, May 8-12). Yago: A core of semantic knowledge. In Proceedings of the International Conference on World Wide Web, Banff, Canada (pp. 697-706).","DOI":"10.1145\/1242572.1242667"},{"key":"jswis.2012070103-58","doi-asserted-by":"crossref","unstructured":"Suchanek, F., Sozio, M., & Weikum, G. (2009, April 20-24). SOFIE: A self-organizing framework for information extraction. In Proceedings of the International Conference on World Wide Web, Madrid, Spain (pp. 631-640).","DOI":"10.1145\/1526709.1526794"},{"key":"jswis.2012070103-59","unstructured":"Surdeanu, M., & Manning, C. (2010, June 2-4). Ensemble models for dependency parsing: cheap and good? In Proceedings of the Human Language Technologies: Conference of the North American Chapter of the Association of Computational Linguistics, Los Angeles, CA (pp. 649-652)."},{"key":"jswis.2012070103-60","unstructured":"Sutton, C., & McCallum, A. (2004). Collective segmentation and labeling of distant entities in information extraction (No. 04-49). Amherst, MA: Department of Computer Science, University of Massachusetts."},{"key":"jswis.2012070103-61","article-title":"An introduction to conditional random fields for relational learning","author":"C.Sutton","year":"2006","journal-title":"Introduction to Statistical Relational Learning"},{"key":"jswis.2012070103-62","doi-asserted-by":"publisher","DOI":"10.1002\/cpe.938"},{"key":"jswis.2012070103-63","author":"M.Theobald","year":"2010","journal-title":"URDF: Efficient reasoning in uncertain rdf knowledge bases with soft and hard rules (Tech. Rep.)"},{"key":"jswis.2012070103-64","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.2005.856938"},{"key":"jswis.2012070103-65","doi-asserted-by":"publisher","DOI":"10.1561\/2200000001"},{"key":"jswis.2012070103-66","doi-asserted-by":"crossref","unstructured":"Weikum, G., & Theobald, M. (2010, June 6-11). From information to knowledge: Harvesting entities and relationships from web sources. In Proceedings of the ACM Symposium on Principles of Database Systems, Indianapolis, IN (pp. 65-76).","DOI":"10.1145\/1807085.1807097"},{"key":"jswis.2012070103-67","doi-asserted-by":"crossref","unstructured":"Wick, M., McCallum, A., & Miklau, G. (2010, September 13-17). Scalable probabilistic databases with factor graphs and MCMC. In Proceedings of the International Conference on Very Large Data Bases, Singapore (pp. 794-804).","DOI":"10.14778\/1920841.1920942"},{"key":"jswis.2012070103-68","author":"L.Wolsey","year":"1998","journal-title":"Integer Programming"},{"key":"jswis.2012070103-69","unstructured":"Wu, F., & Weld, D. (2010, July 11-16). Open information extraction using Wikipedia. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Uppsala, Sweden (pp. 118-127)."},{"key":"jswis.2012070103-70","unstructured":"Yao, L., Riedel, S., & McCallum, A. (2010, October 9-11). Collective cross-document relation extraction without labelled data. In Proceedings of the Conference on Empirical Methods in Natural Language Processing (pp. 1013-1023)."},{"key":"jswis.2012070103-71","unstructured":"Zhang, C., Niu, F., R\u00e9, C., & Shavlik, J. (2012, July 8-14). Big data versus the crowd: Looking for relationships in all the right places. In Proceedings of the Annual Meeting of the Association for Computational Linguistics, Jeju Island, Korea."},{"key":"jswis.2012070103-72","doi-asserted-by":"crossref","unstructured":"Zhu, J., Nie, Z., Liu, X., Zhang, B., & Wen, J. (2009, April 20-24). Statsnowball: A statistical approach to extracting entity relationships. In Proceedings of the International Conference on World Wide Web, Madrid, Spain (pp. 101-110).","DOI":"10.1145\/1526709.1526724"}],"container-title":["International Journal on Semantic Web and Information Systems"],"original-title":[],"language":"ng","link":[{"URL":"https:\/\/www.igi-global.com\/viewtitle.aspx?TitleId=74339","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,4,29]],"date-time":"2025-04-29T13:08:19Z","timestamp":1745932099000},"score":1,"resource":{"primary":{"URL":"https:\/\/services.igi-global.com\/resolvedoi\/resolve.aspx?doi=10.4018\/jswis.2012070103"}},"subtitle":["Large-Scale Knowledge-Base Construction via Machine Learning and Statistical Inference"],"short-title":[],"issued":{"date-parts":[[2012,7,1]]},"references-count":73,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2012,7]]}},"URL":"https:\/\/doi.org\/10.4018\/jswis.2012070103","relation":{},"ISSN":["1552-6283","1552-6291"],"issn-type":[{"value":"1552-6283","type":"print"},{"value":"1552-6291","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,7,1]]}}}