{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,12]],"date-time":"2026-03-12T01:05:08Z","timestamp":1773277508105,"version":"3.50.1"},"reference-count":18,"publisher":"Oxford University Press (OUP)","issue":"8","funder":[{"name":"Intramural Research Program of the National Library of Medicine, National Institues of Health, USA"},{"DOI":"10.13039\/100000002","name":"NIH","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000002","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2020,4,15]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Summary<\/jats:title>\n                    <jats:p>Large-scale data analysis in bioinformatics requires pipelined execution of multiple software. Generally each stage in a pipeline takes considerable computing resources and several workflow management systems (WMS), e.g. Snakemake, Nextflow, Common Workflow Language, Galaxy, etc. have been developed to ensure optimum execution of the stages across two invocations of the pipeline. However, when the pipeline needs to be executed with different settings of parameters, e.g. thresholds, underlying algorithms, etc. these WMS require significant scripting to ensure an optimal execution. We developed JUDI on top of DoIt, a Python based WMS, to systematically handle parameter settings based on the principles of database management systems. Using a novel modular approach that encapsulates a parameter database in each task and file associated with a pipeline stage, JUDI simplifies plug-and-play of the pipeline stages. For a typical pipeline with n parameters, JUDI reduces the number of lines of scripting required by a factor of O(n). With properly designed parameter databases, JUDI not only enables reproducing research under published values of parameters but also facilitates exploring newer results under novel parameter settings.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Availability and implementation<\/jats:title>\n                    <jats:p>https:\/\/github.com\/ncbi\/JUDI<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Supplementary information<\/jats:title>\n                    <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/bioinformatics\/btz956","type":"journal-article","created":{"date-parts":[[2019,12,24]],"date-time":"2019-12-24T07:09:47Z","timestamp":1577171387000},"page":"2572-2574","source":"Crossref","is-referenced-by-count":6,"title":["Bioinformatics pipeline using JUDI:\n                    <i>Just Do It!<\/i>"],"prefix":"10.1093","volume":"36","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4840-3944","authenticated-orcid":false,"given":"Soumitra","family":"Pal","sequence":"first","affiliation":[{"name":"National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health , Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6261-277X","authenticated-orcid":false,"given":"Teresa M","family":"Przytycka","sequence":"additional","affiliation":[{"name":"National Center for Biotechnology Information, National Library of Medicine, National Institutes of Health , Bethesda, MD 20894, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2019,12,27]]},"reference":[{"key":"2023013110270460700_btz956-B1","author":"Amstutz","year":"2016"},{"key":"2023013110270460700_btz956-B2","doi-asserted-by":"crossref","first-page":"403","DOI":"10.1186\/gb4161","article-title":"Dissemination of scientific software with Galaxy ToolShed","volume":"15","author":"Blankenberg","year":"2014","journal-title":"Genome Biol"},{"key":"2023013110270460700_btz956-B3","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1093\/bioinformatics\/btu595","article-title":"BigDataScript: a scripting language for data pipelines","volume":"31","author":"Cingolani","year":"2015","journal-title":"Bioinformatics"},{"key":"2023013110270460700_btz956-B4","doi-asserted-by":"crossref","first-page":"284","DOI":"10.1016\/j.future.2017.01.012","article-title":"Scientific workflows for computational reproducibility in the life sciences: status, challenges and opportunities","volume":"75","author":"Cohen-Boulakia","year":"2017","journal-title":"Future Gener. Comp. Syst"},{"key":"2023013110270460700_btz956-B5","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1155\/2005\/128026","article-title":"Pegasus: A Framework for Mapping Complex Scientific Workflows onto Distributed Systems","volume":"13","author":"Deelman","year":"2005","journal-title":"Scientific Programming"},{"key":"2023013110270460700_btz956-B6","doi-asserted-by":"crossref","first-page":"316","DOI":"10.1038\/nbt.3820","article-title":"Nextflow enables reproducible computational workflows","volume":"35","author":"Di Tommaso","year":"2017","journal-title":"Nat. Biotechnol"},{"key":"2023013110270460700_btz956-B7","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1007\/11890850_2","volume-title":"Provenance and Annotation of Data, Lecture Notes in Computer Science","author":"Freire","year":"2006"},{"key":"2023013110270460700_btz956-B8","first-page":"593","volume-title":", SIGMOD \u201912","author":"Freire","year":"2012"},{"key":"2023013110270460700_btz956-B9","doi-asserted-by":"crossref","first-page":"2520","DOI":"10.1093\/bioinformatics\/bts480","article-title":"Snakemake\u2014a scalable bioinformatics workflow engine","volume":"28","author":"K\u00f6ster","year":"2012","journal-title":"Bioinformatics"},{"key":"2023013110270460700_btz956-B10","first-page":"530","article-title":"A review of bioinformatic pipeline frameworks","volume":"18","author":"Leipzig","year":"2017","journal-title":"Brief. Bioinform"},{"key":"2023013110270460700_btz956-B11","doi-asserted-by":"crossref","first-page":"6632","DOI":"10.1093\/nar\/gkz540","article-title":"Co-SELECT reveals sequence non-specific contribution of DNA shape to transcription factor binding in vitro","volume":"47","author":"Pal","year":"2019","journal-title":"Nucleic Acids Res"},{"key":"2023013110270460700_btz956-B12","doi-asserted-by":"crossref","first-page":"751","DOI":"10.1071\/FP08084","article-title":"OpenAlea: a visual programming and component-based software platform for plant modelling","volume":"35","author":"Pradal","year":"2008","journal-title":"Funct. Plant Biol"},{"key":"2023013110270460700_btz956-B13","doi-asserted-by":"crossref","first-page":"81","DOI":"10.1109\/MCSE.2018.05329818","article-title":"Automan: a python-based automation framework for numerical computing","volume":"20","author":"Ramachandran","year":"2018","journal-title":"Comput. Sci. Eng"},{"key":"2023013110270460700_btz956-B14","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1186\/1471-2105-5-40","article-title":"Pegasys: software for executing and integrating analyses of biological sequences","volume":"5","author":"Shah","year":"2004","journal-title":"BMC Bioinformatics"},{"key":"2023013110270460700_btz956-B15","volume-title":"GNU Make: A Program for Directed Recompilation: GNU Make Version 3.81","author":"Stallman","year":"2004"},{"key":"2023013110270460700_btz956-B16","doi-asserted-by":"crossref","first-page":"102","DOI":"10.1186\/1471-2105-13-102","article-title":"Workflows for microarray data processing in the Kepler environment","volume":"13","author":"Stropp","year":"2012","journal-title":"BMC Bioinformatics"},{"key":"2023013110270460700_btz956-B17","doi-asserted-by":"crossref","first-page":"W557","DOI":"10.1093\/nar\/gkt328","article-title":"The Taverna workflow suite: designing and executing workflows of Web Services on the desktop, web or in the cloud","volume":"41","author":"Wolstencroft","year":"2013","journal-title":"Nucleic Acids Res"},{"key":"2023013110270460700_btz956-B18","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1007\/10968987_3","volume-title":"Job Scheduling Strategies for Parallel Processing, Lecture Notes in Computer Science","author":"Yoo","year":"2003"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/academic.oup.com\/bioinformatics\/advance-article-pdf\/doi\/10.1093\/bioinformatics\/btz956\/32350583\/btz956.pdf","content-type":"application\/pdf","content-version":"am","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/36\/8\/2572\/48983591\/btz956.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/36\/8\/2572\/48983591\/btz956.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,1,31]],"date-time":"2023-01-31T15:31:39Z","timestamp":1675179099000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/36\/8\/2572\/5688745"}},"subtitle":[],"editor":[{"given":"Alfonso","family":"Valencia","sequence":"additional","affiliation":[],"role":[{"role":"editor","vocabulary":"crossref"}]}],"short-title":[],"issued":{"date-parts":[[2019,12,27]]},"references-count":18,"journal-issue":{"issue":"8","published-print":{"date-parts":[[2020,4,15]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btz956","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/611764","asserted-by":"object"}]},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2020,4,15]]},"published":{"date-parts":[[2019,12,27]]}}}