{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T20:37:34Z","timestamp":1772915854685,"version":"3.50.1"},"reference-count":31,"publisher":"Oxford University Press (OUP)","license":[{"start":{"date-parts":[[2019,11,4]],"date-time":"2019-11-04T00:00:00Z","timestamp":1572825600000},"content-version":"vor","delay-in-days":307,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100004329","name":"Slovenian Research Agency","doi-asserted-by":"publisher","award":["P2-0098"],"award-info":[{"award-number":["P2-0098"]}],"id":[{"id":"10.13039\/501100004329","id-type":"DOI","asserted-by":"publisher"}]},{"name":"European Union\u2019s Horizon 2020","award":["863059"],"award-info":[{"award-number":["863059"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,1,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>The existence of annotated text corpora is essential for the development of public health services and tools based on natural language processing (NLP) and text mining. Recently organized biomedical NLP shared tasks have provided annotated corpora related to different biomedical entities such as genes, phenotypes, drugs, diseases and chemical entities. These are needed to develop named-entity recognition (NER) models that are used for extracting entities from text and finding their relations. However, to the best of our knowledge, there are limited annotated corpora that provide information about food entities despite food and dietary management being an essential public health issue. Hence, we developed a new annotated corpus of food entities, named FoodBase. It was constructed using recipes extracted from Allrecipes, which is currently the largest food-focused social network. The recipes were selected from five categories: \u2018Appetizers and Snacks\u2019, \u2018Breakfast and Lunch\u2019, \u2018Dessert\u2019, \u2018Dinner\u2019 and \u2018Drinks\u2019. Semantic tags used for annotating food entities were selected from the Hansard corpus. To extract and annotate food entities, we applied a rule-based food NER method called FoodIE. Since FoodIE provides a weakly annotated corpus, by manually evaluating the obtained results on 1000 recipes, we created a gold standard of FoodBase. It consists of 12\u2009844 food entity annotations describing 2105 unique food entities. Additionally, we provided a weakly annotated corpus on an additional 21\u2009790 recipes. It consists of 274\u2009053 food entity annotations, 13\u2009079 of which are unique. The FoodBase corpus is necessary for developing corpus-based NER models for food science, as a new benchmark dataset for machine learning tasks such as multi-class classification, multi-label classification and hierarchical multi-label classification. FoodBase can be used for detecting semantic differences\/similarities between food concepts, and after all we believe that it will open a new path for learning food embedding space that can be used in predictive studies.<\/jats:p>","DOI":"10.1093\/database\/baz121","type":"journal-article","created":{"date-parts":[[2019,10,2]],"date-time":"2019-10-02T11:13:53Z","timestamp":1570014833000},"source":"Crossref","is-referenced-by-count":45,"title":["FoodBase corpus: a new resource of annotated food entities"],"prefix":"10.1093","volume":"2019","author":[{"given":"Gorjan","family":"Popovski","sequence":"first","affiliation":[{"name":"Faculty of Computer Science and Engineering, Ss. Cyril and Methodius University, ul.Rudzer Boshkovikj 16, 1000 Skopje, Macedonia"},{"name":"Jo\u017eef Stefan International Postgraduate School, Jamova cesta 39, 1000 Ljubljana, Slovenia"},{"name":"Computer Systems Department, Jo\u017eef Stefan Institute, Jamova cesta 39, 1000 Ljubljana, Slovenia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Barbara Korou\u0161i\u0107","family":"Seljak","sequence":"additional","affiliation":[{"name":"Computer Systems Department, Jo\u017eef Stefan Institute, Jamova cesta 39, 1000 Ljubljana, Slovenia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tome","family":"Eftimov","sequence":"additional","affiliation":[{"name":"Computer Systems Department, Jo\u017eef Stefan Institute, Jamova cesta 39, 1000 Ljubljana, Slovenia"},{"name":"Department of Biomedical Data Science, Stanford University, 450 Serra Mall, Stanford 94305 CA, USA"},{"name":"Center for Population Health Sciences, Stanford University, 450 Serra Mall, Stanford 94305 CA, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2019,11,4]]},"reference":[{"key":"2019110410160093700_ref1","doi-asserted-by":"crossref","first-page":"S3","DOI":"10.1186\/1472-6947-15-S2-S3","article-title":"Using text mining techniques to extract phenotypic information from the PhenoCHF corpus","volume":"15","author":"Alnazzawi","year":"2015,","journal-title":"BMC Med. Inform. Decis. Mak."},{"key":"2019110410160093700_ref2","first-page":"140","article-title":"Mining patents with tmChem, GNormPlus and an ensemble of open systems","author":"Leaman","year":"2015","journal-title":"Proceedings of The fifth BioCreative challenge evaluation workshop"},{"key":"2019110410160093700_ref3","doi-asserted-by":"crossref","first-page":"1633","DOI":"10.1093\/bioinformatics\/bts183","article-title":"ChemSpot: a hybrid system for chemical named entity recognition","volume":"28","author":"Rockt\u00e4schel","year":"2012","journal-title":"Bioinformatics"},{"key":"2019110410160093700_ref4","first-page":"1","article-title":"Overview of BioNLP shared task 2011. In Proceedings of the BioNLP shared task 2011 workshop","author":"Kim","year":"2011,","journal-title":"Association for Computational Linguistics"},{"key":"2019110410160093700_ref5","first-page":"1","article-title":"Overview of BioNLP shared task 2013","author":"N\u00e9dellec","year":"2013","journal-title":"In Proceedings of the BioNLP Shared Task 2013 Workshop"},{"key":"2019110410160093700_ref6","doi-asserted-by":"crossref","DOI":"10.1093\/database\/bau053","article-title":"BioC interoperability track overview","author":"Comeau","year":"2014","journal-title":"Database"},{"key":"2019110410160093700_ref7","doi-asserted-by":"crossref","first-page":"12","DOI":"10.18653\/v1\/W16-3002","article-title":"Overview of the bacteria biotope task at bionlp shared task 2016","author":"Del\u0117ger","year":"2016","journal-title":"In Proceedings of the 4th BioNLP Shared Task Workshop"},{"key":"2019110410160093700_ref8","first-page":"1","article-title":"Overview of the regulatory network of plant seed development (SeeDev) task at the BioNLP shared task 2016","author":"Chaix","year":"2016","journal-title":"Proceedings of the 4th BioNLP Shared Task Workshop"},{"key":"2019110410160093700_ref9","doi-asserted-by":"crossref","first-page":"S3","DOI":"10.1186\/gb-2008-9-s2-s3","article-title":"Overview of BioCreative II gene normalization","volume":"9","author":"Morgan","year":"2008","journal-title":"Genome Biol."},{"key":"2019110410160093700_ref10","doi-asserted-by":"crossref","DOI":"10.1093\/database\/bau039","article-title":"BioCreative-IV virtual issue","author":"Arighi","year":"2014","journal-title":"Database"},{"key":"2019110410160093700_ref11","doi-asserted-by":"crossref","first-page":"S1","DOI":"10.1186\/1471-2105-12-S8-S1","article-title":"Overview of the BioCreative III workshop","volume":"12","author":"Arighi","year":"2011","journal-title":"BMC bioinformatics"},{"key":"2019110410160093700_ref12","first-page":"166","article-title":"Overview of the BioCreative V chemical disease relation (CDR) task","author":"Wei","year":"2015","journal-title":"Proceedings of the fifth BioCreative challenge evaluation workshop"},{"key":"2019110410160093700_ref13","volume-title":"Overview of the BioCreative\/OHNLP Challenge 2018 Task 2: Clinical Semantic Textual Similarity","author":"Wang","year":"2018"},{"key":"2019110410160093700_ref14","first-page":"2370","article-title":"Flow graph corpus from recipe texts","author":"Mori","year":"2014,","journal-title":"In LREC"},{"key":"2019110410160093700_ref15","article-title":"Guide to the carnegie mellon university multimodal activity (cmu-mmac) database","volume":"135","author":"De la Torre","year":"2008","journal-title":"Robotics Institute"},{"key":"2019110410160093700_ref16","volume-title":"The UCREL Semantic Analysis System","author":"Rayson","year":"2004"},{"key":"2019110410160093700_ref17","author":"Hansard corpus"},{"key":"2019110410160093700_ref18","author":"SAMUELS"},{"key":"2019110410160093700_ref19","first-page":"150","article-title":"Grammar and dictionary based named-entity linking for knowledge extraction of evidence-based dietary recommendations. In Proceedings of the 8th International Joint Conference on Knowledge Discovery, Knowledge Engineering","author":"Eftimov","year":"2016,","journal-title":"J. Knowl. Manag."},{"key":"2019110410160093700_ref20","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pone.0179488","article-title":"A rule-based named-entity recognition method for knowledge extraction of evidence-based dietary recommendations","volume":"12","author":"Eftimov","year":"2017","journal-title":"Plos One"},{"key":"2019110410160093700_ref21","first-page":"915","volume-title":"In Proceedings of the 8th International Conference on Pattern Recognition Applications and Methods","author":"Popovski","year":"2019"},{"key":"2019110410160093700_ref22","author":"Allrecipes"},{"key":"2019110410160093700_ref23","doi-asserted-by":"crossref","DOI":"10.1093\/database\/bat064","article-title":"BioC: a minimalist approach to interoperability for biomedical text processing","author":"Comeau","year":"2013","journal-title":"Database."},{"key":"2019110410160093700_ref24","article-title":"NCBO annotator: semantic annotation of biomedical data","volume":"110","author":"Jonquet","year":"2009,","journal-title":"In International Semantic Web Conference, Poster and Demo session"},{"key":"2019110410160093700_ref25","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1038\/s41538-018-0032-6","article-title":"& Hsiao, W. W","volume":"2","author":"Dooley","year":"2018","journal-title":"npj Science of Food"},{"key":"2019110410160093700_ref26","first-page":"279","article-title":"SNOMED-CT: the advanced terminology and coding system for eHealth","volume":"121","author":"Donnelly","year":"2006","journal-title":"Stud. Health Technol. Inform."},{"key":"2019110410160093700_ref27","doi-asserted-by":"crossref","first-page":"D267","DOI":"10.1093\/nar\/gkh061","article-title":"The unified medical language system (UMLS): integrating biomedical terminology","volume":"32","author":"Bodenreider","year":"2004","journal-title":"Nucleic Acids Res."},{"key":"2019110410160093700_ref28","doi-asserted-by":"crossref","first-page":"W170","DOI":"10.1093\/nar\/gkp440","article-title":"BioPortal: ontologies and integrated data resources at the click of a mouse","author":"Noy","year":"2009","journal-title":"Nucleic acids research, 37"},{"key":"2019110410160093700_ref29","first-page":"47","article-title":"Support vector machines with binary tree architecture for multi-class classification","volume":"2","author":"Cheong","year":"2004","journal-title":"Neural Information Processing-Letters and Reviews"},{"key":"2019110410160093700_ref30","doi-asserted-by":"crossref","first-page":"1","DOI":"10.4018\/jdwm.2007070101","article-title":"Multi-label classification: an overview","volume":"3","author":"Tsoumakas","year":"2007","journal-title":"International Journal of Data Warehousing and Mining (IJDWM)"},{"key":"2019110410160093700_ref31","doi-asserted-by":"crossref","first-page":"433","DOI":"10.1007\/978-981-13-2285-3_51","article-title":"Multi-label classification trending challenges and approaches","author":"Pant","year":"2019","journal-title":"In Emerging Trends in Expert Applications and Security"}],"container-title":["Database"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baz121\/30350820\/baz121.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"http:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baz121\/30350820\/baz121.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,11,4]],"date-time":"2019-11-04T15:17:45Z","timestamp":1572880665000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/database\/article\/doi\/10.1093\/database\/baz121\/5611291"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,1,1]]},"references-count":31,"URL":"https:\/\/doi.org\/10.1093\/database\/baz121","relation":{},"ISSN":["1758-0463"],"issn-type":[{"value":"1758-0463","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2019]]},"published":{"date-parts":[[2019,1,1]]},"article-number":"baz121"}}