{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T01:02:18Z","timestamp":1778634138751,"version":"3.51.4"},"reference-count":39,"publisher":"Oxford University Press (OUP)","issue":"5","license":[{"start":{"date-parts":[[2018,8,7]],"date-time":"2018-08-07T00:00:00Z","timestamp":1533600000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/academic.oup.com\/journals\/pages\/open_access\/funder_policies\/chorus\/standard_publication_model"}],"funder":[{"name":"ERC Advanced","award":["693174"],"award-info":[{"award-number":["693174"]}]},{"name":"Data-Driven Genomic Computing"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,3,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:sec>\n                  <jats:title>Motivation<\/jats:title>\n                  <jats:p>We previously proposed a paradigm shift in genomic data management, based on the Genomic Data Model (GDM) for mediating existing data formats and on the GenoMetric Query Language (GMQL) for supporting, at a high level of abstraction, data extraction and the most common data-driven computations required by tertiary data analysis of Next Generation Sequencing datasets. Here, we present a new GMQL-based system with enhanced accessibility, portability, scalability and performance.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Results<\/jats:title>\n                  <jats:p>The new system has a well-designed modular architecture featuring: (i) an intermediate representation supporting many different implementations (including Spark, Flink and SciDB); (ii) a high-level technology-independent repository abstraction, supporting different repository technologies (e.g., local file system, Hadoop File System, database or others); (iii) several system interfaces, including a user-friendly Web-based interface, a Web Service interface, and a programmatic interface for Python language. Biological use case examples, using public ENCODE, Roadmap Epigenomics and TCGA datasets, demonstrate the relevance of our work.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Availability and implementation<\/jats:title>\n                  <jats:p>The GMQL system is freely available for non-commercial use as open source project at: http:\/\/www.bioinformatics.deib.polimi.it\/GMQLsystem\/.<\/jats:p>\n               <\/jats:sec>\n               <jats:sec>\n                  <jats:title>Supplementary information<\/jats:title>\n                  <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n               <\/jats:sec>","DOI":"10.1093\/bioinformatics\/bty688","type":"journal-article","created":{"date-parts":[[2018,8,6]],"date-time":"2018-08-06T19:12:45Z","timestamp":1533582765000},"page":"729-736","source":"Crossref","is-referenced-by-count":50,"title":["Processing of big heterogeneous genomic datasets for tertiary analysis of Next Generation Sequencing data"],"prefix":"10.1093","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2574-1174","authenticated-orcid":false,"given":"Marco","family":"Masseroli","sequence":"first","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Arif","family":"Canakoglu","sequence":"additional","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Pietro","family":"Pinoli","sequence":"additional","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Abdulrahman","family":"Kaitoua","sequence":"additional","affiliation":[{"name":"The German Research Center for Artificial Intelligence (DFKI), Berlin, Germany"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrea","family":"Gulino","sequence":"additional","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Olha","family":"Horlova","sequence":"additional","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Luca","family":"Nanni","sequence":"additional","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anna","family":"Bernasconi","sequence":"additional","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stefano","family":"Perna","sequence":"additional","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Eirini","family":"Stamoulakatou","sequence":"additional","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Stefano","family":"Ceri","sequence":"additional","affiliation":[{"name":"Dipartimento di Elettronica, Informazione e Bioingegneria, Politecnico di Milano, Milan, Italy"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2018,8,7]]},"reference":[{"key":"2023013107255390900_bty688-B1","doi-asserted-by":"crossref","first-page":"W581","DOI":"10.1093\/nar\/gkw211","article-title":"DeepBlue epigenomic data server: programmatic data retrieval and analysis of epigenome region sets","volume":"44","author":"Albrecht","year":"2016","journal-title":"Nucleic Acids Res"},{"key":"2023013107255390900_bty688-B2","first-page":"325","volume-title":"Proc. 36th Int. Conf. on Conceptual Modeling (ER 2017)","author":"Bernasconi","year":"2017"},{"key":"2023013107255390900_bty688-B3","doi-asserted-by":"crossref","first-page":"1045","DOI":"10.1038\/nbt1010-1045","article-title":"The NIH Roadmap Epigenomics Mapping Consortium","volume":"28","author":"Bernstein","year":"2010","journal-title":"Nat. Biotechnol"},{"key":"2023013107255390900_bty688-B4","first-page":"193","volume-title":"Proc. IEEE Int. Conf. Big Data. IEEE Computer Society","author":"Bertoni","year":"2015"},{"key":"2023013107255390900_bty688-B5","doi-asserted-by":"crossref","first-page":"963","DOI":"10.1145\/1807167.1807271","volume-title":"Proc. 2010 ACM SIGMOD Int. Conf. on Management of Data","author":"Brown","year":"2010"},{"key":"2023013107255390900_bty688-B6","doi-asserted-by":"crossref","first-page":"1113","DOI":"10.1038\/ng.2764","article-title":"The Cancer Genome Atlas Pan-Cancer analysis project","volume":"45","author":"Weinstein","year":"2013","journal-title":"Nat. Genet"},{"key":"2023013107255390900_bty688-B7","volume-title":"Proc. 4th Workshop on Algorithms and Systems on MapReduce and beyond","author":"Cattani","year":"2017"},{"key":"2023013107255390900_bty688-B8","doi-asserted-by":"crossref","first-page":"482","DOI":"10.1007\/978-3-319-60131-1_34","volume-title":"Proc. Int. Conf. Web Eng","author":"Cattani","year":"2017"},{"key":"2023013107255390900_bty688-B9","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1093\/bioinformatics\/btu595","article-title":"BigDataScript: a scripting language for data pipelines","volume":"31","author":"Cingolani","year":"2015","journal-title":"Bioinformatics"},{"key":"2023013107255390900_bty688-B10","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1186\/s12859-016-1419-5","article-title":"TCGA2BED: extracting, extending, integrating, and querying The Cancer Genome Atlas","volume":"18","author":"Cumbo","year":"2017","journal-title":"BMC Bioinformatics"},{"key":"2023013107255390900_bty688-B11","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1145\/1629175.1629198","article-title":"MapReduce: a flexible data processing tool","volume":"53","author":"Dean","year":"2010","journal-title":"Commun. ACM"},{"key":"2023013107255390900_bty688-B12","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1007\/978-1-4939-1720-4_3","article-title":"Choice of next-generation sequencing pipelines","volume":"1231","author":"Del Chierico","year":"2015","journal-title":"Methods Mol. Biol"},{"key":"2023013107255390900_bty688-B13","doi-asserted-by":"crossref","first-page":"57","DOI":"10.1038\/nature11247","article-title":"An integrated encyclopedia of DNA elements in the human genome","volume":"489","year":"2012","journal-title":"Nature"},{"key":"2023013107255390900_bty688-B14","doi-asserted-by":"crossref","first-page":"647","DOI":"10.1038\/nrg2857","article-title":"Computational solutions to large-scale data management and analysis","volume":"11","author":"Eric","year":"2010","journal-title":"Nat. Rev. Genet"},{"key":"2023013107255390900_bty688-B15","doi-asserted-by":"crossref","first-page":"2089","DOI":"10.1093\/bioinformatics\/btw069","article-title":"Integrated genome browser: visual analytics platform for genomics","volume":"32","author":"Freese","year":"2016","journal-title":"Bioinformatics"},{"key":"2023013107255390900_bty688-B16","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1038\/nrg.2016.49","article-title":"Coming of age: ten years of next-generation sequencing technologies","volume":"17","author":"Goodwin","year":"2016","journal-title":"Nat. Rev. Genet"},{"key":"2023013107255390900_bty688-B17","doi-asserted-by":"crossref","first-page":"3081","DOI":"10.1093\/bioinformatics\/btw199","article-title":"GORpipe: a query tool for working with sequence data based on a Genomic Ordered Relational (GOR) architecture","volume":"32","author":"Gu\u00f0bjartsson","year":"2016","journal-title":"Bioinformatics"},{"key":"2023013107255390900_bty688-B18","doi-asserted-by":"crossref","first-page":"115","DOI":"10.1038\/nmeth.3252","article-title":"Orchestrating high-throughput genomic analysis with Bioconductor","volume":"12","author":"Huber","year":"2015","journal-title":"Nat. Methods"},{"key":"2023013107255390900_bty688-B19","doi-asserted-by":"crossref","first-page":"536","DOI":"10.1186\/s12859-017-1945-9","article-title":"Explorative visual analytics on interval-based genomic data and their metadata","volume":"18","author":"Jalili","year":"2017","journal-title":"BMC Bioinformatics"},{"key":"2023013107255390900_bty688-B20","doi-asserted-by":"crossref","first-page":"443","DOI":"10.1109\/TC.2016.2603980","article-title":"Framework for supporting genomic operations","volume":"66","author":"Kaitoua","year":"2017","journal-title":"IEEE Trans. Comput"},{"key":"2023013107255390900_bty688-B21","article-title":"The UCSC Genome Browser","volume":"4","author":"Karolchik","year":"2012","journal-title":"Curr. Protoc. Bioinformatics, Chapter 1(Unit1)"},{"key":"2023013107255390900_bty688-B22","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/bioinformatics\/btt250","article-title":"Using Genome Query Language to uncover genetic variation","volume":"30","author":"Kozanitis","year":"2014","journal-title":"Bioinformatics"},{"key":"2023013107255390900_bty688-B23","doi-asserted-by":"crossref","first-page":"1881","DOI":"10.1093\/bioinformatics\/btv048","article-title":"GenoMetric Query Language: a novel approach to large-scale genomic data management","volume":"31","author":"Masseroli","year":"2015","journal-title":"Bioinformatics"},{"key":"2023013107255390900_bty688-B24","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/j.ymeth.2016.09.002","article-title":"Modeling and interoperability of heterogeneous genomic big data for integrative processing and querying","volume":"111","author":"Masseroli","year":"2016","journal-title":"Methods"},{"key":"2023013107255390900_bty688-B25","doi-asserted-by":"crossref","first-page":"1297","DOI":"10.1101\/gr.107524.110","article-title":"The genome analysis toolkit: mapReduce framework for analyzing next-generation DNA sequencing data","volume":"20","author":"McKenna","year":"2010","journal-title":"Genome Res"},{"key":"2023013107255390900_bty688-B26","doi-asserted-by":"crossref","first-page":"2822","DOI":"10.1093\/bioinformatics\/btu389","article-title":"Cloud4Psi: cloud computing for 3D protein structure similarity searching","volume":"30","author":"Mrozek","year":"2014","journal-title":"Bioinformatics"},{"key":"2023013107255390900_bty688-B27","doi-asserted-by":"crossref","first-page":"77","DOI":"10.1016\/j.ins.2016.02.029","article-title":"HDInsight4PSi: boosting performance of 3D protein structure similarity searching with HDInsight clusters in Microsoft Azure cloud","volume":"349-350","author":"Mrozek","year":"2016","journal-title":"Informat. Sci"},{"key":"2023013107255390900_bty688-B28","doi-asserted-by":"crossref","first-page":"1919","DOI":"10.1093\/bioinformatics\/bts277","article-title":"BEDOPS: high-performance genomic feature operations","volume":"28","author":"Neph","year":"2012","journal-title":"Bioinformatics"},{"key":"2023013107255390900_bty688-B29","doi-asserted-by":"crossref","first-page":"774","DOI":"10.1016\/j.jbi.2013.07.001","article-title":"\u2032Big data\u2032, Hadoop and cloud computing in genomics","volume":"46","author":"O'Driscoll","year":"2013","journal-title":"J. Biomed. Inform"},{"key":"2023013107255390900_bty688-B30","doi-asserted-by":"crossref","first-page":"1099","DOI":"10.1145\/1376616.1376726","volume-title":"Proc. 2008 ACM SIGMOD Int. Conf. on Management of Data.","author":"Olston","year":"2008"},{"key":"2023013107255390900_bty688-B31","doi-asserted-by":"crossref","first-page":"200","DOI":"10.1109\/TCBB.2012.170","article-title":"Genomic Region Operation Kit for extensible processing of deep sequencing data","volume":"10","author":"Ovaska","year":"2013","journal-title":"IEEE\/ACM Trans. Comput. Biol. Bioinform"},{"key":"2023013107255390900_bty688-B32","doi-asserted-by":"crossref","first-page":"841","DOI":"10.1093\/bioinformatics\/btq033","article-title":"BEDTools: a flexible suite of utilities for comparing genomic features","volume":"26","author":"Quinlan","year":"2010","journal-title":"Bioinformatics"},{"key":"2023013107255390900_bty688-B33","doi-asserted-by":"crossref","first-page":"24","DOI":"10.1038\/nbt.1754","article-title":"Integrative genomics viewer","volume":"29","author":"Robinson","year":"2011","journal-title":"Nat. Biotechnol"},{"key":"2023013107255390900_bty688-B34","doi-asserted-by":"crossref","first-page":"187","DOI":"10.1145\/3035918.3064048","volume-title":"Proc. 2017 ACM Int. Conf. on Management of Data.","author":"Roy","year":"2017"},{"key":"2023013107255390900_bty688-B35","doi-asserted-by":"crossref","first-page":"712","DOI":"10.1016\/j.drudis.2017.01.014","article-title":"Next-generation sequencing: big data meets high performance computing","volume":"22","author":"Schmidt","year":"2017","journal-title":"Drug Discov. Today"},{"key":"2023013107255390900_bty688-B36","first-page":"1","volume-title":"Proc. 2010 IEEE 26th Symp. Mass Storage Systems and Technologies (MSST)","author":"Shvachko","year":"2010"},{"key":"2023013107255390900_bty688-B37","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1016\/S0140-6736(14)62453-3","article-title":"UK gears up to decode 100 000 genomes from NHS patients","volume":"385","author":"Siva","year":"2015","journal-title":"Lancet"},{"key":"2023013107255390900_bty688-B38","doi-asserted-by":"crossref","first-page":"2652","DOI":"10.1093\/bioinformatics\/btu343","article-title":"SparkSeq: fast, scalable and cloud-ready tool for the interactive genomic data analysis with nucleotide precision","volume":"30","author":"Wiewi\u00f3rka","year":"2014","journal-title":"Bioinformatics"},{"key":"2023013107255390900_bty688-B39","doi-asserted-by":"crossref","first-page":"749","DOI":"10.1186\/s12864-017-4071-1","article-title":"START: a system for flexible analysis of hundreds of genomic signal tracks in few lines of SQL-like queries","volume":"18","author":"Zhu","year":"2017","journal-title":"BMC Genomics"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/5\/729\/48966068\/bioinformatics_35_5_729.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/5\/729\/48966068\/bioinformatics_35_5_729.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,31]],"date-time":"2023-01-31T10:21:48Z","timestamp":1675160508000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/35\/5\/729\/5067860"}},"subtitle":[],"editor":[{"given":"Inanc","family":"Birol","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2018,8,7]]},"references-count":39,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2019,3,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bty688","relation":{},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2019,3,1]]},"published":{"date-parts":[[2018,8,7]]}}}