{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,7,11]],"date-time":"2024-07-11T06:35:35Z","timestamp":1720679735525},"reference-count":33,"publisher":"Springer Science and Business Media LLC","issue":"S4","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2012,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>In the scientific biodiversity community, it is increasingly perceived the need to build a bridge between molecular and traditional biodiversity studies. We believe that the information technology could have a preeminent role in integrating the information generated by these studies with the large amount of molecular data we can find in bioinformatics public databases. This work is primarily aimed at building a bioinformatic infrastructure for the integration of public and private biodiversity data through the development of GIDL, an Intelligent Data Loader coupled with the Molecular Biodiversity Database. The system presented here organizes in an ontological way and locally stores the sequence and annotation data contained in the GenBank primary database.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Methods<\/jats:title>\n            <jats:p>The GIDL architecture consists of a relational database and of an intelligent data loader software. The relational database schema is designed to manage biodiversity information (Molecular Biodiversity Database) and it is organized in four areas: MolecularData, Experiment, Collection and Taxonomy. The MolecularData area is inspired to an established standard in Generic Model Organism Databases, the Chado relational schema. The peculiarity of Chado, and also its strength, is the adoption of an ontological schema which makes use of the Sequence Ontology.<\/jats:p>\n            <jats:p>The Intelligent Data Loader (IDL) component of GIDL is an Extract, Transform and Load software able to parse data, to discover hidden information in the GenBank entries and to populate the Molecular Biodiversity Database. The IDL is composed by three main modules: the Parser, able to parse GenBank flat files; the Reasoner, which automatically builds CLIPS facts mapping the biological knowledge expressed by the Sequence Ontology; the DBFiller, which translates the CLIPS facts into ordered SQL statements used to populate the database. In GIDL Semantic Web technologies have been adopted due to their advantages in data representation, integration and processing.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Results and conclusions<\/jats:title>\n            <jats:p>Entries coming from Virus (814,122), Plant (1,365,360) and Invertebrate (959,065) divisions of GenBank rel.180 have been loaded in the Molecular Biodiversity Database by GIDL. Our system, combining the Sequence Ontology and the Chado schema, allows a more powerful query expressiveness compared with the most commonly used sequence retrieval systems like Entrez or SRS.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-13-s4-s4","type":"journal-article","created":{"date-parts":[[2012,3,28]],"date-time":"2012-03-28T10:46:40Z","timestamp":1332931600000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["GIDL: a rule based expert system for GenBank Intelligent Data Loading into the Molecular Biodiversity database"],"prefix":"10.1186","volume":"13","author":[{"given":"Paolo","family":"Pannarale","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Domenico","family":"Catalano","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Giorgio","family":"De Caro","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Giorgio","family":"Grillo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pietro","family":"Leo","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Graziano","family":"Pappad\u00e0","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Francesco","family":"Rubino","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gaetano","family":"Scioscia","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Flavio","family":"Licciulli","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2012,3,28]]},"reference":[{"issue":"Suppl 14","key":"5101_CR1","doi-asserted-by":"publisher","first-page":"S2","DOI":"10.1186\/1471-2105-10-S14-S2","volume":"10","author":"VS Chavan","year":"2009","unstructured":"Chavan VS, Ingwersen P: Towards a data publishing framework for primary biodiversity data: challenges and potentials for the biodiversity informatics community. BMC Bioinformatics 2009, 10(Suppl 14):S2. 10.1186\/1471-2105-10-S14-S2","journal-title":"BMC Bioinformatics"},{"key":"5101_CR2","doi-asserted-by":"publisher","first-page":"347","DOI":"10.1093\/bib\/bbm037","volume":"8","author":"IN Sarkar","year":"2007","unstructured":"Sarkar IN: Biodiversity informatics: organising and linking across the spectrum of life. Brief Bioinform 2007, 8: 347\u2013357. 10.1093\/bib\/bbm037","journal-title":"Brief Bioinform"},{"issue":"11","key":"5101_CR3","doi-asserted-by":"publisher","first-page":"e1124","DOI":"10.1371\/journal.pone.0001124","volume":"2","author":"C Yesson","year":"2007","unstructured":"Yesson C, Brewer PW, Sutton T, Caithness N, Pahwa JS, Burgess M, Gray WA, White RJ, Jones AC, Bisby FA, Culham A: How global is the global biodiversity information facility? PLoS One 2007, 2(11):e1124. 10.1371\/journal.pone.0001124","journal-title":"PLoS One"},{"key":"5101_CR4","doi-asserted-by":"publisher","first-page":"158","DOI":"10.1186\/1471-2105-8-158","volume":"8","author":"RD Page","year":"2007","unstructured":"Page RD: TBMap: a taxonomic perspective on the phylogenetic database TreeBASE. BMC Bioinformatics 2007, 8: 158. 10.1186\/1471-2105-8-158","journal-title":"BMC Bioinformatics"},{"issue":"Database","key":"5101_CR5","doi-asserted-by":"publisher","first-page":"D38","DOI":"10.1093\/nar\/gkq1172","volume":"39","author":"EW Sayers","year":"2011","unstructured":"Sayers EW, Barrett T, Benson DA, Bolton E, Bryant SH, Canese K, Chetvernin V, Church DM, DiCuccio M, Federhen S, Feolo M, Fingerman IM, Geer LY, Helmberg W, Kapustin Y, Landsman D, Lipman DJ, Lu Z, Madden TL, Madej T, Maglott DR, Marchler-Bauer A, Miller V, Mizrachi I, Ostell J, Panchenko A, Phan L, Pruitt KD, Schuler GD, Sequeira E, et al.: Database resources of the National Center for Biotechnology Information. Nucleic Acids Res 2011, 39(Database):D38\u201351. 10.1093\/nar\/gkq1172","journal-title":"Nucleic Acids Res"},{"key":"5101_CR6","unstructured":"Global Biodiversity Information Facility[http:\/\/www.gbif.org\/]"},{"key":"5101_CR7","first-page":"Unit 1.3","volume":"Chapter 1","author":"G Gibney","year":"2011","unstructured":"Gibney G, Baxevanis AD: Searching NCBI databases using Entrez. Curr Protoc Bioinformatics 2011, Chapter 1: Unit 1.3.","journal-title":"Curr Protoc Bioinformatics"},{"issue":"8","key":"5101_CR8","doi-asserted-by":"publisher","first-page":"1149","DOI":"10.1093\/bioinformatics\/18.8.1149","volume":"18","author":"EM Zdobnov","year":"2002","unstructured":"Zdobnov EM, Lopez R, Apweiler R, Etzold T: The EBI SRS server-new features. Bioinformatics 2002, 18(8):1149\u20131150. 10.1093\/bioinformatics\/18.8.1149","journal-title":"Bioinformatics"},{"issue":"Database","key":"5101_CR9","doi-asserted-by":"publisher","first-page":"D32","DOI":"10.1093\/nar\/gkq1079","volume":"39","author":"DA Benson","year":"2011","unstructured":"Benson DA, Karsch-Mizrachi I, Lipman DJ, Ostell J, Sayers EW: GenBank. Nucleic Acids Res 2011, 39(Database):D32\u201337. 10.1093\/nar\/gkq1079","journal-title":"Nucleic Acids Res"},{"issue":"Database","key":"5101_CR10","doi-asserted-by":"publisher","first-page":"D22","DOI":"10.1093\/nar\/gkq1041","volume":"39","author":"E Kaminuma","year":"2011","unstructured":"Kaminuma E, Kosuge T, Kodama Y, Aono H, Mashima J, Gojobori T, Sugawara H, Ogasawara O, Takagi T, Okubo K, Nakamura Y: DDBJ progress report. Nucleic Acids Res 2011, 39(Database):D22\u201327. 10.1093\/nar\/gkq1041","journal-title":"Nucleic Acids Res"},{"issue":"Database","key":"5101_CR11","doi-asserted-by":"publisher","first-page":"D15","DOI":"10.1093\/nar\/gkq1150","volume":"39","author":"G Cochrane","year":"2011","unstructured":"Cochrane G, Karsch-Mizrachi I, Nakamura Y: The International Nucleotide Sequence Database Collaboration. Nucleic Acids Res 2011, 39(Database):D15\u201318. 10.1093\/nar\/gkq1150","journal-title":"Nucleic Acids Res"},{"key":"5101_CR12","first-page":"422","volume":"2010","author":"MS Hagen","year":"2010","unstructured":"Hagen MS, Lee EK: BIOSPIDA: A Relational Database Translator for NCBI. AMIA Annual Symposium Proceedings 2010, 2010: 422\u2013426.","journal-title":"AMIA Annual Symposium Proceedings"},{"key":"5101_CR13","doi-asserted-by":"publisher","first-page":"34","DOI":"10.1186\/1471-2105-6-34","volume":"6","author":"SP Shah","year":"2005","unstructured":"Shah SP, Huang Y, Xu T, Yuen MM, Ling J, Ouellette BF: Atlas-a data warehouse for integrative bioinformatics. BMC Bioinformatics 2005, 6: 34. 10.1186\/1471-2105-6-34","journal-title":"BMC Bioinformatics"},{"key":"5101_CR14","doi-asserted-by":"publisher","first-page":"170","DOI":"10.1186\/1471-2105-7-170","volume":"7","author":"TJ Lee","year":"2006","unstructured":"Lee TJ, Pouliot Y, Wagner V, Gupta P, Stringer-Calvert DW, Tenenbaum JD, Karp PD: BioWarehouse: a bioinformatics database warehouse toolkit. BMC Bioinformatics 2006, 7: 170. 10.1186\/1471-2105-7-170","journal-title":"BMC Bioinformatics"},{"key":"5101_CR15","unstructured":"Molecular Biodiversity Laboratory[http:\/\/www.mblabproject.it\/]"},{"key":"5101_CR16","doi-asserted-by":"publisher","first-page":"355","DOI":"10.1111\/j.1471-8286.2007.01678.x","volume":"7","author":"S Ratnasingham","year":"2007","unstructured":"Ratnasingham S, Hebert PDN: BOLD: The Barcode of Life Data System ( ). Molecular Ecology Notes 2007, 7: 355\u2013364. http:\/\/www.barcodinglife.org 10.1111\/j.1471-8286.2007.01678.x","journal-title":"Molecular Ecology Notes"},{"key":"5101_CR17","unstructured":"Generic Model Organism Database[http:\/\/www.gmod.org\/]"},{"issue":"13","key":"5101_CR18","doi-asserted-by":"publisher","first-page":"i337","DOI":"10.1093\/bioinformatics\/btm189","volume":"23","author":"CJ Mungall","year":"2007","unstructured":"Mungall CJ, Emmert DB, The FlyBase Consortium: A Chado case study: an ontology-based modular schema for representing genome-associated biological information. Bioinformatics 2007, 23(13):i337-i346. 10.1093\/bioinformatics\/btm189","journal-title":"Bioinformatics"},{"issue":"1","key":"5101_CR19","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1016\/j.jbi.2010.03.002","volume":"44","author":"CJ Mungall","year":"2011","unstructured":"Mungall CJ, Batchelor C, Eilbeck K: Evolution of the Sequence Ontology terms and relationships. J Biomed Inform 2011, 44(1):87\u201393. 10.1016\/j.jbi.2010.03.002","journal-title":"J Biomed Inform"},{"issue":"11","key":"5101_CR20","doi-asserted-by":"publisher","first-page":"1251","DOI":"10.1038\/nbt1346","volume":"25","author":"B Smith","year":"2007","unstructured":"Smith B, Ashburner M, Rosse C, Bard J, Bug W, Ceusters W, Goldberg LJ, Eilbeck K, Ireland A, Mungall CJ, OBI Consortium, Leontis N, Rocca-Serra P, Ruttenberg A, Sansone SA, Scheuermann RH, Shah N, Whetzel PL, Lewis S: The OBO Foundry: coordinated evolution of ontologies to support biomedical data integration. Nat Biotechnol 2007, 25(11):1251\u20131255. 10.1038\/nbt1346","journal-title":"Nat Biotechnol"},{"key":"5101_CR21","doi-asserted-by":"publisher","first-page":"R44","DOI":"10.1186\/gb-2005-6-5-r44","volume":"6","author":"K Eilbeck","year":"2005","unstructured":"Eilbeck K, Lewis SE, Mungall CJ, Yandell M, Stein L, Durbin R, Ashburner M: The Sequence Ontology: a tool for the unification of genome annotations. Genome Biology 2005, 6: R44. 10.1186\/gb-2005-6-5-r44","journal-title":"Genome Biology"},{"key":"5101_CR22","volume-title":"Expert Systems: principles and programming","author":"JC Giarratano","year":"1998","unstructured":"Giarratano JC, Riley G: Expert Systems: principles and programming. Boston: PWS Publishing Company; 1998."},{"key":"5101_CR23","unstructured":"BioJava[http:\/\/biojava.org\/]"},{"key":"5101_CR24","unstructured":"BioPython[http:\/\/www.biopython.org\/]"},{"key":"5101_CR25","unstructured":"BioPerl[http:\/\/www.bioperl.org\/]"},{"key":"5101_CR26","unstructured":"EMBOSS[http:\/\/emboss.sourceforge.net\/]"},{"key":"5101_CR27","doi-asserted-by":"publisher","first-page":"321","DOI":"10.1186\/1471-2105-9-321","volume":"9","author":"TH Lee","year":"2008","unstructured":"Lee TH, Kim YK, Nahm BH: GBParsy: a GenBank flatfile parser library with high speed. BMC Bioinformatics 2008, 9: 321. 10.1186\/1471-2105-9-321","journal-title":"BMC Bioinformatics"},{"key":"5101_CR28","unstructured":"DDBJ\/EMBL\/GenBank Feature Table definition[ftp:\/\/ftp.ncbi.nih.gov\/genbank\/docs\/]"},{"key":"5101_CR29","unstructured":"OWL Web Ontology Language[http:\/\/www.w3.org\/TR\/owl-features\/]"},{"key":"5101_CR30","unstructured":"The OWL API[http:\/\/owlapi.sourceforge.net\/]"},{"key":"5101_CR31","unstructured":"Mapping of the feature table terms and qualifiers to SO[http:\/\/www.sequenceontology.org\/resources\/mapping\/FT_SO.html]"},{"key":"5101_CR32","first-page":"48","volume-title":"Proceedings of Data Integration in the Life Sciences, 4 conf., DILS 2007: 27-29 June 2007, Philadelphia, USA","author":"R Rifaieh","year":"2007","unstructured":"Rifaieh R, Unwin R, Carver J, Miller MA: SWAMI: Integrating Biological Databases and Analysis Tools Within User Friendly Environment. In Proceedings of Data Integration in the Life Sciences, 4 conf., DILS 2007: 27\u201329 June 2007, Philadelphia, USA. Edited by: Sarah Cohen-Boulakia. Val Tannen: Springer; 2007:48\u201358."},{"key":"5101_CR33","doi-asserted-by":"publisher","first-page":"34","DOI":"10.1186\/1471-2105-6-34","volume":"6","author":"SP Shah","year":"2005","unstructured":"Shah SP, Huang Y, Xu T, Yuen MM, Ling J, Ouellette BF: Atlas-a data warehouse for integrative bioinformatics. BMC Bioinformatics 2005, 6: 34. 10.1186\/1471-2105-6-34","journal-title":"BMC Bioinformatics"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-13-S4-S4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T18:44:44Z","timestamp":1630521884000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-13-S4-S4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,3,28]]},"references-count":33,"journal-issue":{"issue":"S4","published-print":{"date-parts":[[2012,12]]}},"alternative-id":["5101"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-13-s4-s4","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,3,28]]},"assertion":[{"value":"28 March 2012","order":1,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"S4"}}