{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,25]],"date-time":"2025-09-25T15:29:45Z","timestamp":1758814185159},"reference-count":48,"publisher":"Springer Science and Business Media LLC","issue":"1","content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2006,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:sec>\n            <jats:title>Background<\/jats:title>\n            <jats:p>Integration of heterogeneous data types is a challenging problem, especially in biology, where the number of databases and data types increase rapidly. Amongst the problems that one has to face are integrity, consistency, redundancy, connectivity, expressiveness and updatability.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Description<\/jats:title>\n            <jats:p>Here we present a system (Biozon) that addresses these problems, and offers biologists a new knowledge resource to navigate through and explore. Biozon unifies multiple biological databases consisting of a variety of data types (such as DNA sequences, proteins, interactions and cellular pathways). It is fundamentally different from previous efforts as it uses a single extensive and tightly connected graph schema wrapped with hierarchical ontology of documents and relations. Beyond warehousing existing data, Biozon computes and stores novel derived data, such as similarity relationships and functional predictions. The integration of similarity data allows propagation of knowledge through inference and fuzzy searches. Sophisticated methods of query that span multiple data types were implemented and first-of-a-kind biological ranking systems were explored and integrated.<\/jats:p>\n          <\/jats:sec>\n          <jats:sec>\n            <jats:title>Conclusion<\/jats:title>\n            <jats:p>The Biozon system is an extensive knowledge resource of heterogeneous biological data. Currently, it holds more than 100 million biological documents and 6.5 billion relations between them. The database is accessible through an advanced web interface that supports complex queries, \"fuzzy\" searches, data materialization and more, online at <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"http:\/\/biozon.org\" ext-link-type=\"uri\">http:\/\/biozon.org<\/jats:ext-link>.<\/jats:p>\n          <\/jats:sec>","DOI":"10.1186\/1471-2105-7-70","type":"journal-article","created":{"date-parts":[[2006,2,17]],"date-time":"2006-02-17T07:45:28Z","timestamp":1140162328000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":67,"title":["BIOZON: a system for unification, management and analysis of heterogeneous biological data"],"prefix":"10.1186","volume":"7","author":[{"given":"Aaron","family":"Birkland","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Golan","family":"Yona","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2006,2,15]]},"reference":[{"key":"809_CR1","doi-asserted-by":"publisher","first-page":"45","DOI":"10.1093\/nar\/28.1.45","volume":"28","author":"A Bairoch","year":"2000","unstructured":"Bairoch A, Apweiler R: The SWISS-PROT protein sequence database and its supplement TrEMBL in 2000. Nucleic Acids Res 2000, 28: 45\u201348.","journal-title":"Nucleic Acids Res"},{"key":"809_CR2","doi-asserted-by":"publisher","first-page":"17","DOI":"10.1093\/nar\/24.1.17","volume":"24","author":"DG George","year":"1996","unstructured":"George DG, Barker WC, Mewes HW, Pfeiffer F, Tsugita A: The PIR-International Protein Sequence Database. Nucleic Acids Research 1996, 24: 17\u201320.","journal-title":"Nucleic Acids Research"},{"key":"809_CR3","doi-asserted-by":"publisher","first-page":"12","DOI":"10.1093\/nar\/27.1.12","volume":"27","author":"DA Benson","year":"1999","unstructured":"Benson DA, Boguski MS, Lipman DJ, Ostell J, Ouellette BFF, Rapp BA, Wheeler DL: GenBank. Nucleic Acids Research 1999, 27: 12\u201317.","journal-title":"Nucleic Acids Research"},{"key":"809_CR4","doi-asserted-by":"publisher","first-page":"242","DOI":"10.1093\/nar\/29.1.242","volume":"29","author":"GD Bader","year":"2001","unstructured":"Bader GD, Donaldson I, Wolting C, Ouellette BFF, Pawson T, Hogue CWV: BIND \u2013 The Biomolecular Interaction Network Database. Nucleic Acids Research 2001, 29: 242\u2013245.","journal-title":"Nucleic Acids Research"},{"key":"809_CR5","doi-asserted-by":"publisher","first-page":"239","DOI":"10.1093\/nar\/29.1.239","volume":"29","author":"I Xenarios","year":"2001","unstructured":"Xenarios I, Fernandez E, Salwinski L, Duan XJ, Thompson MJ, Marcotte EM, Eisenberg D: DIP: The Database of Interacting Proteins: 2001 update. Nucleic Acids Research 2001, 29: 239\u2013241.","journal-title":"Nucleic Acids Research"},{"key":"809_CR6","doi-asserted-by":"publisher","first-page":"56","DOI":"10.1093\/nar\/30.1.56","volume":"30","author":"PD Karp","year":"2002","unstructured":"Karp PD, Riley M, Jr MHS, Paulsen IT, Collado-Vides J, Paley SM, Pellegrini-Toole A, Bonavides C, Gama-Castro S: The EcoCyc Database. Nucleic Acids Research 2002, 30: 56\u201358.","journal-title":"Nucleic Acids Research"},{"key":"809_CR7","doi-asserted-by":"publisher","first-page":"29","DOI":"10.1093\/nar\/27.1.29","volume":"27","author":"H Ogata","year":"1999","unstructured":"Ogata H, Goto S, Sato K, Fujibuchi W, Bono H, Kanehisa M: KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Research 1999, 27: 29\u201334.","journal-title":"Nucleic Acids Research"},{"key":"809_CR8","unstructured":"Nucleic Acids Research, Database Issue1999. [http:\/\/nar.oxfordjournals.org]"},{"issue":"4","key":"809_CR9","doi-asserted-by":"publisher","first-page":"557","DOI":"10.1089\/cmb.1995.2.557","volume":"2","author":"S Davidson","year":"1995","unstructured":"Davidson S, Overton GC, Buneman P: Challenges in Integrating Biological Data Sources. Journal of Computational Biology 1995, 2(4):557\u2013572.","journal-title":"Journal of Computational Biology"},{"key":"809_CR10","doi-asserted-by":"publisher","first-page":"81","DOI":"10.1007\/BF01228708","volume":"1","author":"S Spaccapietra","year":"1992","unstructured":"Spaccapietra S, Parent C, Dupont Y: Model Independent Assertions for Integration of Heterogeneous Schemas. VLDB Journal 1992, 1: 81\u2013126.","journal-title":"VLDB Journal"},{"issue":"3","key":"809_CR11","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1145\/1031570.1031583","volume":"33","author":"T Hernandez","year":"2004","unstructured":"Hernandez T, Kambhampati S: Integration of biological sources: current systems and challenges ahead. SIGMOD Rec 2004, 33(3):51\u201360.","journal-title":"SIGMOD Rec"},{"key":"809_CR12","first-page":"59","volume":"9","author":"T Etzold","year":"1993","unstructured":"Etzold T, Argos P: SRS \u2013 an indexing and retrieval tool for flat file data libraries. Computer Applications in the Biosciences 1993, 9: 59\u201364.","journal-title":"Computer Applications in the Biosciences"},{"key":"809_CR13","doi-asserted-by":"publisher","first-page":"141","DOI":"10.1016\/S0076-6879(96)66012-1","volume":"266","author":"GD Schuler","year":"1996","unstructured":"Schuler GD, Epstein J, Ohkawa H, Kans JA: Entrez: Molecular Biology Database and Retrieval System. Methods in Enzymology 1996, 266: 141\u2013161.","journal-title":"Methods in Enzymology"},{"issue":"2","key":"809_CR14","doi-asserted-by":"publisher","first-page":"489","DOI":"10.1147\/sj.402.0489","volume":"40","author":"L Haas","year":"2001","unstructured":"Haas L, Schwarz P, Kodali P, Kotlar E, Rice J, Swope W: DiscoveryLink: a system for integrated access to life sciences data sources. IBM Systems Journal 2001, 40(2):489\u2013511.","journal-title":"IBM Systems Journal"},{"key":"809_CR15","first-page":"473","volume-title":"Proceedings of the 2001 AMIA Annual Symposium","author":"P Mork","year":"2001","unstructured":"Mork P, Halevy A, Tarczy-Hornoch P: A model for Data Integration Systems of Biomedical Data Applied to Online Genetic Databases. Proceedings of the 2001 AMIA Annual Symposium 2001, 473\u2013477."},{"key":"809_CR16","first-page":"25","volume-title":"Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology","author":"PG Baker","year":"1998","unstructured":"Baker PG, Brass A, Bechhofer S, Goble CA, Paton NW, Stevens R: TAMBIS: Transparent Access to Multiple Bioinformatics Information Sources. Proceedings of the Sixth International Conference on Intelligent Systems for Molecular Biology 1998, 25\u201334."},{"issue":"5","key":"809_CR17","doi-asserted-by":"publisher","first-page":"393","DOI":"10.1016\/0306-4379(95)00021-U","volume":"20","author":"IMA Chen","year":"1995","unstructured":"Chen IMA, Markowitz VM: An Overview of the Object-Protocol Model (OPM) and OPM Data Management Tools. Information Systems 1995, 20(5):393\u2013418.","journal-title":"Information Systems"},{"issue":"2","key":"809_CR18","doi-asserted-by":"publisher","first-page":"512","DOI":"10.1147\/sj.402.0512","volume":"40","author":"SB Davidson","year":"2001","unstructured":"Davidson SB, Crabtree J, Brunk BP, Schug J, Tannen V, Overton GC, Stoeckert CJ: K2\/Kleisli and GUS: Experiments in integrated access to genomic data sources. IBM Systems Journal 2001, 40(2):512\u2013530.","journal-title":"IBM Systems Journal"},{"key":"809_CR19","first-page":"147","volume-title":"Managing Scientific Data","author":"J Chen","year":"2003","unstructured":"Chen J, Chung S, Wong L: The Kleisli Query System as a Backbone for Bioinformatics Data Integration and Analysis. In Managing Scientific Data. Morgan Kaufmann; 2003:147\u2013187."},{"key":"809_CR20","doi-asserted-by":"publisher","first-page":"39","DOI":"10.1109\/SSDM.2000.869777","volume-title":"Intl. Conference on Statistical and Scientific Database Management","author":"A Gupta","year":"2000","unstructured":"Gupta A, Ludascher B, Martone ME: Knowledge-Based Integration of Neuroscience Data Sources. Intl. Conference on Statistical and Scientific Database Management 2000, 39\u201352."},{"issue":"4","key":"809_CR21","doi-asserted-by":"publisher","first-page":"331","DOI":"10.1093\/bib\/3.4.331","volume":"3","author":"MD Wilkinson","year":"2002","unstructured":"Wilkinson MD, Links M: BioMOBY: An Open Source Biological Web Services Proposal. Briefings in Bioinformatics 2002, 3(4):331\u2013341.","journal-title":"Briefings in Bioinformatics"},{"key":"809_CR22","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1109\/BIBE.2000.889583","volume-title":"BIBE '00: Proceedings of the 1st IEEE International Symposium on Bioinformatics and Biomedical Engineering","author":"LM Haas","year":"2000","unstructured":"Haas LM, Kodali P, Rice JE, Schwarz PM, Swope WC: Integrating life sciences data-with a little Garlic. In BIBE '00: Proceedings of the 1st IEEE International Symposium on Bioinformatics and Biomedical Engineering. IEEE Computer Society; 2000:5\u201312."},{"key":"809_CR23","volume-title":"The Collection Programming Language Reference Manual","author":"L Wong","year":"1995","unstructured":"Wong L: The Collection Programming Language Reference Manual.Tech. rep., Institute of Systems Science, Heng Mui Keng Terrace, Singapore 0511; 1995. [http:\/\/citeseer.ist.psu.edu\/wong96collection.html]"},{"key":"809_CR24","volume-title":"AllGenes: a web site providing access to an integrated database of known and predicted human (release 9.0, 2004) and mouse genes. (release 9.0, 2004)","author":"CBIL","year":"2004","unstructured":"CBIL: AllGenes: a web site providing access to an integrated database of known and predicted human (release 9.0, 2004) and mouse genes. (release 9.0, 2004).Center for Bioinformatics, University of Pennsylvania; 2004. [http:\/\/www.allgenes.org\/]"},{"key":"809_CR25","doi-asserted-by":"publisher","first-page":"D235","DOI":"10.1093\/nar\/gkj153","volume":"34","author":"A Birkland","year":"2006","unstructured":"Birkland A, Yona G: BIOZON: a Hub of Heterogeneous Biological Data. Nucleic Acids Research 2006, 34: D235-D242.","journal-title":"Nucleic Acids Research"},{"key":"809_CR26","doi-asserted-by":"publisher","first-page":"25","DOI":"10.1038\/75556","volume":"25","author":"M Ashburner","year":"2000","unstructured":"Ashburner M, Ball CA, Blake JA, Botstein D, Butler H, Cherry JM, Davis AP, Dolinski K, Dwight SS, Eppig JT, Harris MA, Hill DP, Issel-Tarver L, Kasarskis A, Lewis S, Matese JC, Richardson JE, Ringwald M, Sherlock GMRG: Gene Ontology: tool for the unification of biology. Nature Genetics 2000, 25: 25\u201329.","journal-title":"Nature Genetics"},{"key":"809_CR27","doi-asserted-by":"publisher","first-page":"D501","DOI":"10.1093\/nar\/gki025","volume":"33","author":"KD Pruitt","year":"2005","unstructured":"Pruitt KD, Tatusova T, Maglott DR: NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcrips and proteins. Nucleic Acids Research 2005, 33: D501-D504.","journal-title":"Nucleic Acids Research"},{"issue":"4","key":"809_CR28","doi-asserted-by":"publisher","first-page":"334","DOI":"10.1007\/s007780100057","volume":"10","author":"E Rahm","year":"2001","unstructured":"Rahm E, Bernstein PA: A survey of approaches to automatic schema matching. VLDB Journal: Very Large Data Bases 2001, 10(4):334\u2013350.","journal-title":"VLDB Journal: Very Large Data Bases"},{"key":"809_CR29","doi-asserted-by":"publisher","first-page":"37","DOI":"10.1093\/nar\/29.1.37","volume":"29","author":"R Apweiler","year":"2001","unstructured":"Apweiler R, Attwood TK, Bairoch A, Bateman A, Birney E, Biswas M, Bucher P, Cerutti L, Corpet F, Croning MDR, Durbin R, Falquet L, Fleischmann W, Gouzy J, Hermjakob H, Hulo N, Jonassen I, Kahn D, Kanapin A, Karavidopoulou Y, Lopez R, Marx B, Mulder NJ, Oinn TM, Pagni M, Servant F, Sigrist CJA, Zdobnov EM: The InterPro database, an integrated documentation resource for protein families, domains and functional sites. Nucleic Acids Research 2001, 29: 37\u201340.","journal-title":"Nucleic Acids Research"},{"key":"809_CR30","doi-asserted-by":"publisher","first-page":"220","DOI":"10.1093\/nar\/27.1.220","volume":"27","author":"TK Attwood","year":"1999","unstructured":"Attwood TK, Flower DR, Lewis AP, Mabey JE, Morgan SR, Scordis P, Selley JN, Wright W: PRINTS prepares for the new millennium. Nucleic Acids Research 1999, 27: 220\u2013225.","journal-title":"Nucleic Acids Research"},{"key":"809_CR31","doi-asserted-by":"publisher","first-page":"263","DOI":"10.1093\/nar\/27.1.263","volume":"27","author":"F Corpet","year":"1999","unstructured":"Corpet F, Gouzy J, Kahn D: Recent improvements of the ProDom database of protein domain families. Nucleic Acids Research 1999, 27: 263\u2013267.","journal-title":"Nucleic Acids Research"},{"key":"809_CR32","doi-asserted-by":"publisher","first-page":"260","DOI":"10.1093\/nar\/27.1.260","volume":"27","author":"A Bateman","year":"1999","unstructured":"Bateman A, Birney E, Durbin R, Eddy SR, Finn RD, Sonnhammer ELL: Pfam 3.1: 1313 multiple alignments and profile HMMs match the majority of proteins. Nucleic Acids Research 1999, 27: 260\u2013262.","journal-title":"Nucleic Acids Research"},{"issue":"5","key":"809_CR33","doi-asserted-by":"publisher","first-page":"1257","DOI":"10.1006\/jmbi.2001.5293","volume":"315","author":"G Yona","year":"2002","unstructured":"Yona G, Levitt M: Within the Twilight Zone: A Sensitive Profile-Profile Comparison Tool Based on Information Theory. Journal of Molecular Biology 2002, 315(5):1257\u20131275.","journal-title":"Journal of Molecular Biology"},{"key":"809_CR34","volume-title":"Bioinformatics","author":"RY Pinter","year":"2005","unstructured":"Pinter RY, Rokhlenko O, Yeger-Lotem E, Ziv-Ukelson M: Alignment of Metabolic Pathways. Bioinformatics 2005. [To appear]"},{"issue":"3","key":"809_CR35","doi-asserted-by":"publisher","first-page":"281","DOI":"10.1038\/90129","volume":"28","author":"J Brown","year":"2001","unstructured":"Brown J, Douady C, Italia M, Marshall W, Stanhope M: Universal trees based on large combined protein sequence data. Nat Genet 2001, 28(3):281\u2013285.","journal-title":"Nat Genet"},{"issue":"11","key":"809_CR36","doi-asserted-by":"publisher","first-page":"5913","DOI":"10.1073\/pnas.95.11.5913","volume":"95","author":"M Levitt","year":"1998","unstructured":"Levitt M, Gerstein M: A unified statistical framework for sequence comparison and structure comparison. Proc Natl Acad Sci USA 1998, 95(11):5913\u20135920.","journal-title":"Proc Natl Acad Sci USA"},{"issue":"9","key":"809_CR37","doi-asserted-by":"publisher","first-page":"739","DOI":"10.1093\/protein\/11.9.739","volume":"11","author":"IN Shindyalov","year":"1998","unstructured":"Shindyalov IN, Bourne PE: Protein structure alignment by incremental combinatorial extension (CE) of the optimal path. Protein Engineering 1998, 11(9):739\u2013747.","journal-title":"Protein Engineering"},{"key":"809_CR38","doi-asserted-by":"publisher","first-page":"12","DOI":"10.1089\/cmb.2005.12.12","volume":"12","author":"G Yona","year":"2005","unstructured":"Yona G, Kedem K: The URMS-RMS hybrid algorithm for fast and sensitive local protein structure alignment. Journal of Computational Biology 2005, 12: 12\u201332.","journal-title":"Journal of Computational Biology"},{"key":"809_CR39","first-page":"395","volume-title":"Proceedings of the 8th International Conference on Intelligent Systems for Molecular Biology, ISMB'2000 (La Jolla, California, August 16\u201323, 2000)","author":"G Yona","year":"2000","unstructured":"Yona G, Levitt M: Towards a complete map of the protein space based on a unified sequence and structure analysis of all known proteins. Proceedings of the 8th International Conference on Intelligent Systems for Molecular Biology, ISMB'2000 (La Jolla, California, August 16\u201323, 2000) 2000, 395\u2013406."},{"key":"809_CR40","volume-title":"A comprehensive study of the notion of functional link between genes based on microarray data, promoter signals, protein-protein interactions and pathway analysis","author":"W Dirks","year":"2003","unstructured":"Dirks W, Yona G: A comprehensive study of the notion of functional link between genes based on microarray data, promoter signals, protein-protein interactions and pathway analysis. Tech. rep., Computing and Information Science, Cornell University; 2003."},{"key":"809_CR41","volume-title":"The NCBI Handbook, National Center for Biotechnology Information","author":"J Pontius","year":"2003","unstructured":"Pontius J, Wagner L, Schuler G: UniGene: a unified view of the transcriptome. The NCBI Handbook, National Center for Biotechnology Information 2003."},{"key":"809_CR42","volume-title":"BMC bioinformatics","author":"S Tri\u00dfl","year":"2005","unstructured":"Tri\u00dfl S, Rother K, M\u00fcller H, Steinke T, Koch I, Preissner R, Fr\u00f6mmel C, Leser U: Columba: an integrated database of proteins, structures, and annotations. BMC bioinformatics 2005."},{"key":"809_CR43","volume-title":"VLDB","author":"V Hristidis","year":"2002","unstructured":"Hristidis V, Papakonstantinou Y: Discover: Keyword search in relational databases. VLDB 2002."},{"key":"809_CR44","volume-title":"Topology Search over Databases","author":"L Guo","year":"2005","unstructured":"Guo L, Birkland A, Shanmugasundaram J, Yona G: Topology Search over Databases.2005. [http:\/\/biozon.org\/ftp\/data\/papers\/topologies\/]"},{"key":"809_CR45","first-page":"668","volume-title":"Proceedings of the Ninth Annual ACM-SIAM Symposium on Discrete Algorithms","author":"JM Kleinberg","year":"1998","unstructured":"Kleinberg JM: Authoritative sources in a hyperlinked environment. Proceedings of the Ninth Annual ACM-SIAM Symposium on Discrete Algorithms 1998, 668\u2013677."},{"key":"809_CR46","volume-title":"The PageRank Citation Ranking: Bringing Order to the Web","author":"L Page","year":"1998","unstructured":"Page L, Brin S, Motwani R, Winograd T: The PageRank Citation Ranking: Bringing Order to the Web.Tech. rep., Stanford Digital Library Technologies Project; 1998. [http:\/\/citeseer.ist.psu.edu\/page98pagerank.html]"},{"key":"809_CR47","doi-asserted-by":"publisher","first-page":"71","DOI":"10.1186\/1471-2105-7-71","volume":"7","author":"P Shafer","year":"2006","unstructured":"Shafer P, Isganitis T, Yona G: Hubs of Knowledge: using the functional link structure in Biozon to mine for biologically significant entities. BMC Bioinformatics 2006, 7: 71.","journal-title":"BMC Bioinformatics"},{"key":"809_CR48","first-page":"399","volume":"5","author":"M Quist","year":"2004","unstructured":"Quist M, Yona G: Distributional Scaling: An Algorithm for Structure-Preserving Embedding of Metric and Nonmetric Spaces. Journal of Machine Learning Research 2004, 5: 399\u2013420.","journal-title":"Journal of Machine Learning Research"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/1471-2105-7-70.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,9,1]],"date-time":"2021-09-01T03:23:40Z","timestamp":1630466620000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/1471-2105-7-70"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2006,2,15]]},"references-count":48,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2006,12]]}},"alternative-id":["809"],"URL":"https:\/\/doi.org\/10.1186\/1471-2105-7-70","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2006,2,15]]},"assertion":[{"value":"1 July 2005","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 February 2006","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 February 2006","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}],"article-number":"70"}}