{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T03:38:43Z","timestamp":1760240323529,"version":"build-2065373602"},"reference-count":45,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2019,5,9]],"date-time":"2019-05-09T00:00:00Z","timestamp":1557360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Coordenacao de Aperfeicoamento de Pessoal de N\\ivel Superior - Brasil (CAPES)","award":["001"],"award-info":[{"award-number":["001"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>One way to increase the understanding of texts by machines is through adding semantic information to lexical items by including metadata tags, a process also called semantic annotation. There are several semantic aspects that can be added to the words, among them the information about the nature of the concept denoted through the association with a category of an ontology. The application of ontologies in the annotation task can span multiple domains. However, this particular research focused its approach on top-level ontologies due to its generalizing characteristic. Considering that annotation is an arduous task that demands time and specialized personnel to perform it, much is done on ways to implement the semantic annotation automatically. The use of machine learning techniques are the most effective approaches in the annotation process. Another factor of great importance for the success of the training process of the supervised learning algorithms is the use of a sufficiently large corpus and able to condense the linguistic variance of the natural language. In this sense, this article aims to present an automatic approach to enrich documents from the American English corpus through a CRF model for semantic annotation of ontologies from Schema.org top-level. The research uses two approaches of the model obtaining promising results for the development of semantic annotation based on top-level ontologies. Although it is a new line of research, the use of top-level ontologies for automatic semantic enrichment of texts can contribute significantly to the improvement of text interpretation by machines.<\/jats:p>","DOI":"10.3390\/info10050171","type":"journal-article","created":{"date-parts":[[2019,5,9]],"date-time":"2019-05-09T11:22:35Z","timestamp":1557400955000},"page":"171","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Ontological Semantic Annotation of an English Corpus Through Condition Random Fields"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9913-3556","authenticated-orcid":false,"given":"Guidson Coelho","family":"de Andrade","sequence":"first","affiliation":[{"name":"Departamento de Inform\u00e1tica\u2014Centro de Ci\u00eancias Exatas e Tecnol\u00f3gicas, Universidade Federal de Vicosa, Vi\u00e7osa MG 36570-900, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9129-8620","authenticated-orcid":false,"given":"Alcione","family":"de Paiva Oliveira","sequence":"additional","affiliation":[{"name":"Departamento de Inform\u00e1tica\u2014Centro de Ci\u00eancias Exatas e Tecnol\u00f3gicas, Universidade Federal de Vicosa, Vi\u00e7osa MG 36570-900, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7459-1657","authenticated-orcid":false,"given":"Alexandra","family":"Moreira","sequence":"additional","affiliation":[{"name":"Departamento de Inform\u00e1tica\u2014Centro de Ci\u00eancias Exatas e Tecnol\u00f3gicas, Universidade Federal de Vicosa, Vi\u00e7osa MG 36570-900, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2019,5,9]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"323","DOI":"10.1590\/S0102-44502000000200005","article-title":"Ling\u00fc\u00edstica de corpus: hist\u00f3rico e problem\u00e1tica","volume":"16","author":"Sardinha","year":"2000","journal-title":"Delta"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"275","DOI":"10.1093\/llc\/8.4.275","article-title":"Corpus annotation schemes","volume":"8","author":"Leech","year":"1993","journal-title":"Lit. Linguist. Comput."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"49","DOI":"10.1016\/j.websem.2004.07.005","article-title":"Semantic annotation, indexing, and retrieval","volume":"2","author":"Kiryakov","year":"2004","journal-title":"Web Semant. Sci. Serv. Agents World Wide Web"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Reeve, L., and Han, H. (2005, January 13\u201317). Survey of semantic annotation platforms. Proceedings of the 2005 ACM Symposium on Applied Computing, Santa Fe, NM, USA.","DOI":"10.1145\/1066677.1067049"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Handschuh, S., and Staab, S. (2003). Annotation for the Semantic Web, IOS Press.","DOI":"10.1109\/MIS.2003.1234768"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"195","DOI":"10.3765\/bls.v13i0.1820","article-title":"Taking: A study in lexical network theory","volume":"Volume 13","author":"Norvig","year":"1987","journal-title":"Annual Meeting of the Berkeley Linguistics Society"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1145\/219717.219748","article-title":"WordNet: a lexical database for English","volume":"38","author":"Miller","year":"1995","journal-title":"Commun. ACM"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Dabrowska, E., and Divjak, D. (2015). Polysemy. Handbook of Cognitive Linguistics, Walter de Gruyter GmbH & Co KG.","DOI":"10.1515\/9783110292022"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Ravin, Y., and Leacock, C. (2000). Polysemy: Theoretical and Computational Approaches, OUP Oxford.","DOI":"10.1093\/oso\/9780198238423.001.0001"},{"key":"ref_10","unstructured":"Firth, J.R. (1957). A synopsis of linguistic theory, 1930\u20131955. Special Volume, Philological Society, Oxford University Press."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"59","DOI":"10.1080\/09500789908666759","article-title":"Judging a word by the company it keeps: The use of concordancing software to explore aspects of the mathematics register","volume":"13","author":"Monaghan","year":"1999","journal-title":"Lang. Educ."},{"key":"ref_12","unstructured":"Pustejovsky, J., and Stubbs, A. (2012). Natural Language Annotation for Machine Learning, O\u2019Reilly Media, Inc."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1006\/knac.1993.1008","article-title":"A translation approach to portable ontology specifications","volume":"5","author":"Gruber","year":"1993","journal-title":"Knowl. Acquis."},{"key":"ref_14","first-page":"81","article-title":"Formal ontology and information systems","volume":"Volume 98","author":"Guarino","year":"1998","journal-title":"Proceedings of FOIS"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1016\/j.websem.2005.10.002","article-title":"Semantic annotation for knowledge management: Requirements and a survey of the state of the art","volume":"4","author":"Uren","year":"2006","journal-title":"Web Semant. Sci. Serv. Agents World Wide Web"},{"key":"ref_16","unstructured":"Guarino, N. (1997). Some organizing principles for a unified top-level ontology. AAAI Spring Symposium on Ontological Engineering, AAAI Press."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1038\/scientificamerican0501-34","article-title":"The semantic web","volume":"284","author":"Hendler","year":"2001","journal-title":"Sci. Am."},{"key":"ref_18","unstructured":"Berners-Lee (2001). Weaving the Web: The Original Design and Ultimate Destiny of the World Wide Web, HarperBusiness."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1109\/5254.920602","article-title":"Ontology learning for the semantic web","volume":"16","author":"Maedche","year":"2001","journal-title":"IEEE Intell. Syst."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"1763","DOI":"10.1007\/s11192-015-1637-z","article-title":"Comparing the topological properties of real and artificially generated scientific manuscripts","volume":"105","author":"Amancio","year":"2015","journal-title":"Scientometrics"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Akimushkin, C., Amancio, D.R., and Oliveira, O.N. (2017). Text authorship identified using the dynamics of word co-occurrence networks. PLoS ONE, 12.","DOI":"10.1371\/journal.pone.0170527"},{"key":"ref_22","unstructured":"Preo\u0163iuc-Pietro, D., Liu, Y., Hopkins, D., and Ungar, L. (August, January 30). Beyond Binary Labels: Political Ideology Prediction of Twitter Users. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vancouver, BC, Canada."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Liu, Y., Zhang, L., Nie, L., Yan, Y., and Rosenblum, D.S. (2016, January 12\u201317). Fortune teller: Predicting your career path. Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.9969"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Estival, D., Nowak, C., and Zschorn, A. (2004, January 25). Towards Ontology-based Natural Language Processing. Proceedings of the Workshop on NLP and XML (NLPXML-2004): RDF\/RDFS and OWL in Language Technology, Barcelona, Spain.","DOI":"10.3115\/1621066.1621075"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Gries, S.T., and Berez, A.L. (2017). Linguistic annotation in\/for corpus linguistics. Handbook of Linguistic Annotation, Springer.","DOI":"10.1007\/978-94-024-0881-2_15"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"FitzGerald, N., T\u00e4ckstr\u00f6m, O., Ganchev, K., and Das, D. (2015, January 17\u201321). Semantic role labeling with neural network factors. Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, Lisbon, Portugal.","DOI":"10.18653\/v1\/D15-1112"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"373","DOI":"10.1515\/9783110199901.373","article-title":"Frame semantics","volume":"34","author":"Fillmore","year":"2006","journal-title":"Cogn. Linguist. Basic Read."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Bergamaschi, S., Cappelli, A., Circiello, A., and Varone, M. (2017, January 17\u201322). Conditional random fields with semantic enhancement for named-entity recognition. Proceedings of the 7th International Conference on Web Intelligence, Mining and Semantics, Amantea, Italy.","DOI":"10.1145\/3102254.3102286"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"148","DOI":"10.1016\/j.jbi.2014.01.012","article-title":"Automatic recognition of disorders, findings, pharmaceuticals and body structures from clinical text: An annotation and machine learning study","volume":"49","author":"Skeppstedt","year":"2014","journal-title":"J. Biomed. Inform."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Pandolfo, L., and Pulina, L. (2017). ADnOTO: A Self-adaptive System for Automatic Ontology-Based Annotation of Unstructured Documents. International Conference on Industrial, Engineering and Other Applications of Applied Intelligent Systems, Springer.","DOI":"10.1007\/978-3-319-60042-0_54"},{"key":"ref_31","unstructured":"Adorni, G., Maratea, M., Pandolfo, L., and Pulina, L. (2015). An Ontology-Based Archive for Historical Research. Description Logics, EUR-WS."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"313","DOI":"10.1016\/j.autcon.2017.02.003","article-title":"Ontology-based semi-supervised conditional random fields for automated information extraction from bridge inspection reports","volume":"81","author":"Liu","year":"2017","journal-title":"Autom. Constr."},{"key":"ref_33","first-page":"1","article-title":"Linked data-the story so far","volume":"5","author":"Bizer","year":"2009","journal-title":"Int. J. Semant. Web Inf. Syst."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/978-3-031-79432-2","article-title":"Linked data: Evolving the web into a global data space","volume":"1","author":"Heath","year":"2011","journal-title":"Synth. Lect. Semant. Web Theory Technol."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Patel-Schneider, P.F. (2014). Analyzing schema.org. International Semantic Web Conference, Springer.","DOI":"10.1007\/978-3-319-11964-9_17"},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1145\/2844544","article-title":"Schema.org: Evolution of structured data on the web","volume":"59","author":"Guha","year":"2016","journal-title":"Commun. ACM"},{"key":"ref_37","first-page":"1","article-title":"HTML5 Microdata and Schema. org","volume":"16","author":"Ronallo","year":"2012","journal-title":"Code4Lib J."},{"key":"ref_38","unstructured":"Ide, N., and Suderman, K. (2004). The American National Corpus First Release. LREC, ELRA."},{"key":"ref_39","unstructured":"Andrade, G.C. (2017). Hybrid Semantic Annotation: Rule-Based and Manual Annotation of the Open American National Corpus with a Top-Level Ontology. [Ph.D. Thesis, Universidade Federal de Vi\u00e7osa]."},{"key":"ref_40","unstructured":"Lafferty, J.D., McCallum, A., and Pereira, F.C.N. (2001). Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data. Proceedings of the Eighteenth International Conference on Machine Learning (ICML \u201901), Morgan Kaufmann Publishers Inc."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1561\/2200000013","article-title":"An introduction to conditional random fields","volume":"4","author":"Sutton","year":"2012","journal-title":"Found. Trends Mach. Learn."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Dietterich, T.G. (2002). Machine learning for sequential data: A review. Joint IAPR International Workshops on Statistical Techniques in Pattern Recognition (SPR) and Structural and Syntactic Pattern Recognition (SSPR), Springer.","DOI":"10.1007\/3-540-70659-3_2"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"F\u00fcrnkranz, J., Scheffer, T., and Spiliopoulou, M. (2006). Efficient Inference in Large Conditional Random Fields. Machine Learning: ECML 2006, Springer.","DOI":"10.1007\/11871842"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"773","DOI":"10.1090\/S0025-5718-1980-0572855-7","article-title":"Updating quasi-Newton matrices with limited storage","volume":"35","author":"Nocedal","year":"1980","journal-title":"Math. Comput."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"301","DOI":"10.1111\/j.1467-9868.2005.00503.x","article-title":"Regularization and variable selection via the elastic net","volume":"67","author":"Zou","year":"2005","journal-title":"J. R. Stat. Soc. Ser. B (Stat. Methodol.)"}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/10\/5\/171\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T12:50:24Z","timestamp":1760187024000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/10\/5\/171"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,5,9]]},"references-count":45,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2019,5]]}},"alternative-id":["info10050171"],"URL":"https:\/\/doi.org\/10.3390\/info10050171","relation":{},"ISSN":["2078-2489"],"issn-type":[{"type":"electronic","value":"2078-2489"}],"subject":[],"published":{"date-parts":[[2019,5,9]]}}}