{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T08:10:06Z","timestamp":1783152606446,"version":"3.54.6"},"reference-count":39,"publisher":"Oxford University Press (OUP)","issue":"5","license":[{"start":{"date-parts":[[2018,8,8]],"date-time":"2018-08-08T00:00:00Z","timestamp":1533686400000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"The European Commission\u2019s Horizon 2020","award":["654241"],"award-info":[{"award-number":["654241"]}]},{"DOI":"10.13039\/501100001729","name":"Swedish Foundation for Strategic Research","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100001729","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001862","name":"Swedish Research Council FORMAS","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100001862","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100007435","name":"\u00c5ke Wiberg Foundation","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100007435","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2019,3,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Motivation<\/jats:title>\n                    <jats:p>Computational biologists face many challenges related to data size, and they need to manage complicated analyses often including multiple stages and multiple tools, all of which must be deployed to modern infrastructures. To address these challenges and maintain reproducibility of results, researchers need (i) a reliable way to run processing stages in any computational environment, (ii) a well-defined way to orchestrate those processing stages and (iii) a data management layer that tracks data as it moves through the processing pipeline.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>Pachyderm is an open-source workflow system and data management framework that fulfils these needs by creating a data pipelining and data versioning layer on top of projects from the container ecosystem, having Kubernetes as the backbone for container orchestration. We adapted Pachyderm and demonstrated its attractive properties in bioinformatics. A Helm Chart was created so that researchers can use Pachyderm in multiple scenarios. The Pachyderm File System was extended to support block storage. A wrapper for initiating Pachyderm on cloud-agnostic virtual infrastructures was created. The benefits of Pachyderm are illustrated via a large metabolomics workflow, demonstrating that Pachyderm enables efficient and sustainable data science workflows while maintaining reproducibility and scalability.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Availability and implementation<\/jats:title>\n                    <jats:p>Pachyderm is available from https:\/\/github.com\/pachyderm\/pachyderm. The Pachyderm Helm Chart is available from https:\/\/github.com\/kubernetes\/charts\/tree\/master\/stable\/pachyderm. Pachyderm is available out-of-the-box from the PhenoMeNal VRE (https:\/\/github.com\/phnmnl\/KubeNow-plugin) and general Kubernetes environments instantiated via KubeNow. The code of the workflow used for the analysis is available on GitHub (https:\/\/github.com\/pharmbio\/LC-MS-Pachyderm).<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Supplementary information<\/jats:title>\n                    <jats:p>Supplementary data are available at Bioinformatics online.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1093\/bioinformatics\/bty699","type":"journal-article","created":{"date-parts":[[2018,8,7]],"date-time":"2018-08-07T15:12:59Z","timestamp":1533654779000},"page":"839-846","source":"Crossref","is-referenced-by-count":42,"title":["Container-based bioinformatics with Pachyderm"],"prefix":"10.1093","volume":"35","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2187-5426","authenticated-orcid":false,"given":"Jon Ander","family":"Novella","sequence":"first","affiliation":[{"name":"Department of Pharmaceutical Biosciences and Science for Life Laboratory, Uppsala University, Uppsala, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Payam","family":"Emami Khoonsari","sequence":"additional","affiliation":[{"name":"Department of Medical Sciences, Clinical Chemistry, Uppsala University, Uppsala, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stephanie","family":"Herman","sequence":"additional","affiliation":[{"name":"Department of Pharmaceutical Biosciences and Science for Life Laboratory, Uppsala University, Uppsala, Sweden"},{"name":"Department of Medical Sciences, Clinical Chemistry, Uppsala University, Uppsala, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Daniel","family":"Whitenack","sequence":"additional","affiliation":[{"name":"Pachyderm, Inc., San Francisco, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Marco","family":"Capuccini","sequence":"additional","affiliation":[{"name":"Department of Pharmaceutical Biosciences and Science for Life Laboratory, Uppsala University, Uppsala, Sweden"},{"name":"Department of Information Technology, Uppsala University, Uppsala, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Joachim","family":"Burman","sequence":"additional","affiliation":[{"name":"Department of Neuroscience, Uppsala University, Uppsala, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kim","family":"Kultima","sequence":"additional","affiliation":[{"name":"Department of Medical Sciences, Clinical Chemistry, Uppsala University, Uppsala, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8083-2864","authenticated-orcid":false,"given":"Ola","family":"Spjuth","sequence":"additional","affiliation":[{"name":"Department of Pharmaceutical Biosciences and Science for Life Laboratory, Uppsala University, Uppsala, Sweden"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"286","published-online":{"date-parts":[[2018,8,8]]},"reference":[{"key":"2023013107253314900_bty699-B1","doi-asserted-by":"crossref","first-page":"W3","DOI":"10.1093\/nar\/gkw343","article-title":"The galaxy platform for accessible, reproducible and collaborative biomedical analyses: 2016 update","volume":"44","author":"Afgan","year":"2016","journal-title":"Nucleic Acids Res"},{"key":"2023013107253314900_bty699-B2","doi-asserted-by":"crossref","first-page":"142.","DOI":"10.1126\/science.354.6308.142","article-title":"The hard road to reproducibility","volume":"354","author":"Barba","year":"2016","journal-title":"Science"},{"key":"2023013107253314900_bty699-B3","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1161\/CIRCRESAHA.114.303819","article-title":"Reproducibility in science: improving the standard for basic and preclinical research","volume":"116","author":"Begley","year":"2015","journal-title":"Circ. Res"},{"key":"2023013107253314900_bty699-B4","first-page":"108","author":"Burns","year":"2016"},{"key":"2023013107253314900_bty699-B5","first-page":"9:1","volume-title":"2017 Imperial College Computing Student Workshop (ICCSW 2017), Volume 60 of OpenAccess Series in Informatics (OASIcs)","author":"Capuccini","year":"2018"},{"key":"2023013107253314900_bty699-B6","doi-asserted-by":"crossref","first-page":"2580","DOI":"10.1093\/bioinformatics\/btx192","article-title":"Biocontainers: an open-source and community-driven framework for software standardization","volume":"33","author":"da Veiga Leprevost","year":"2017","journal-title":"Bioinformatics"},{"key":"2023013107253314900_bty699-B7","first-page":"e2519","article-title":"A microservice-based portal for x-ray transient and variable sources","volume":"5","author":"D\u2019Agostino","year":"2017","journal-title":"PeerJ Prepr"},{"key":"2023013107253314900_bty699-B8","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1145\/1327452.1327492","article-title":"Mapreduce: simplified data processing on large clusters","volume":"51","author":"Dean","year":"2008","journal-title":"Commun. ACM"},{"key":"2023013107253314900_bty699-B9","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1002\/mas.20108","article-title":"Mass spectrometrybased metabolomics","volume":"26","author":"Dettmer","year":"2007","journal-title":"Mass Spectrom. Rev"},{"key":"2023013107253314900_bty699-B10","doi-asserted-by":"crossref","first-page":"e1273.","DOI":"10.7717\/peerj.1273","article-title":"The impact of docker containers on the performance of genomic pipelines","volume":"3","author":"Di Tommaso","year":"2015","journal-title":"PeerJ"},{"key":"2023013107253314900_bty699-B11","doi-asserted-by":"crossref","first-page":"316","DOI":"10.1038\/nbt.3820","article-title":"Nextflow enables reproducible computational workflows","volume":"35","author":"Di Tommaso","year":"2017","journal-title":"Nat. Biotechnol"},{"key":"2023013107253314900_bty699-B12","first-page":"610","volume-title":"Virtualization vs containerization to support paas","author":"Dua","year":"2014"},{"key":"2023013107253314900_bty699-B13","doi-asserted-by":"crossref","first-page":"12580","DOI":"10.1073\/pnas.1509788112","article-title":"Searching molecular structure databases with tandem mass spectra using CSI: fingerID","volume":"112","author":"Duhrkop","year":"2015","journal-title":"Proc. Natl. Acad. Sci. USA"},{"key":"2023013107253314900_bty699-B14","first-page":"836","author":"Govindan","year":"1996"},{"key":"2023013107253314900_bty699-B15","doi-asserted-by":"crossref","first-page":"631","DOI":"10.1016\/j.cels.2018.03.014","article-title":"Practical computational reproducibility in the life sciences","volume":"6","author":"Gr\u00fcning","year":"2018","journal-title":"Cell Syst"},{"key":"2023013107253314900_bty699-B16","doi-asserted-by":"crossref","first-page":"D781","DOI":"10.1093\/nar\/gks1004","article-title":"MetaboLights\u2014an open-access general-purpose repository for metabolomics studies and associated meta-data","volume":"41","author":"Haug","year":"2013","journal-title":"Nucleic Acids Res"},{"key":"2023013107253314900_bty699-B17","first-page":"213603","article-title":"Interoperable and scalable metabolomics data analysis with microservices","author":"Khoonsari","year":"2017","journal-title":"bioRxiv"},{"key":"2023013107253314900_bty699-B18","doi-asserted-by":"crossref","first-page":"2520","DOI":"10.1093\/bioinformatics\/bts480","article-title":"Snakemake\u2013a scalable bioinformatics workflow engine","volume":"28","author":"K\u00f6ster","year":"2012","journal-title":"Bioinformatics"},{"key":"2023013107253314900_bty699-B19","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1021\/ac202450g","article-title":"CAMERA: an integrated strategy for compound spectra extraction and annotation of liquid chromatography\/mass spectrometry data sets","volume":"84","author":"Kuhl","year":"2012","journal-title":"Anal. Chem"},{"key":"2023013107253314900_bty699-B20","doi-asserted-by":"crossref","first-page":"188.","DOI":"10.1038\/nrd3368","article-title":"Impact of high-throughput screening in biomedical research","volume":"10","author":"Macarron","year":"2011","journal-title":"Nat. Rev. Drug Discov"},{"key":"2023013107253314900_bty699-B21","doi-asserted-by":"crossref","first-page":"255","DOI":"10.1038\/498255a","article-title":"Biology: the big challenges of big data","volume":"498","author":"Marx","year":"2013","journal-title":"Nature"},{"key":"2023013107253314900_bty699-B22","doi-asserted-by":"crossref","first-page":"776","DOI":"10.1093\/bioinformatics\/btw707","article-title":"Cymer: cytometry analysis using knime, docker and r","volume":"33","author":"Muchmore","year":"2016","journal-title":"Bioinformatics"},{"key":"2023013107253314900_bty699-B23","doi-asserted-by":"crossref","first-page":"681.","DOI":"10.1038\/nmeth0910-681","article-title":"Mass spectrometry in high-throughput proteomics: ready for the big time","volume":"7","author":"Nilsson","year":"2010","journal-title":"Nat. Methods"},{"key":"2023013107253314900_bty699-B24","author":"Rajasekar","year":"2007"},{"key":"2023013107253314900_bty699-B25","volume-title":"Kubernetes - Scheduling the Future at Cloud Scale","author":"Rensin","year":"2015"},{"key":"2023013107253314900_bty699-B26","doi-asserted-by":"crossref","first-page":"741.","DOI":"10.1038\/nmeth.3959","article-title":"Openms: a flexible open-source software platform for mass spectrometry data analysis","volume":"13","author":"R\u00f6st","year":"2016","journal-title":"Nat. Methods"},{"key":"2023013107253314900_bty699-B27","doi-asserted-by":"crossref","first-page":"1525","DOI":"10.1093\/bioinformatics\/bts167","article-title":"Bpipe: a tool for running and managing bioinformatics pipelines","volume":"28","author":"Sadedin","year":"2012","journal-title":"Bioinformatics"},{"key":"2023013107253314900_bty699-B28","first-page":"261","article-title":"Strong scaling analysis of a parallel, unstructured, implicit solver and the influence of the operating system interference","volume":"17","author":"Sahni","year":"2009","journal-title":"Sci. Program"},{"key":"2023013107253314900_bty699-B29","doi-asserted-by":"crossref","first-page":"53.","DOI":"10.4103\/2153-3539.197197","article-title":"Use of application containers and workflows for genomic data analysis","volume":"7","author":"Schulz","year":"2016","journal-title":"J. Pathol. Inform"},{"key":"2023013107253314900_bty699-B30","first-page":"38","article-title":"Openstack: toward an open-source solution for cloud computing","volume":"55","author":"Sefraoui","year":"2012","journal-title":"Int. J. Comput. Appl"},{"key":"2023013107253314900_bty699-B31","doi-asserted-by":"crossref","first-page":"1084","DOI":"10.1038\/nbt.2421","article-title":"The expanding scope of dna sequencing","volume":"30","author":"Shendure","year":"2012","journal-title":"Nat. Biotechnol"},{"key":"2023013107253314900_bty699-B32","doi-asserted-by":"crossref","first-page":"173.","DOI":"10.1038\/546173a","article-title":"Software simplified","volume":"546","author":"Silver","year":"2017","journal-title":"Nature News"},{"key":"2023013107253314900_bty699-B33","doi-asserted-by":"crossref","first-page":"31","DOI":"10.1145\/1084805.1084812","article-title":"A survey of data provenance in e-science","volume":"34","author":"Simmhan","year":"2005","journal-title":"SIGMOD Rec"},{"key":"2023013107253314900_bty699-B34","doi-asserted-by":"crossref","first-page":"779","DOI":"10.1021\/ac051437y","article-title":"XCMS: processing mass spectrometry data for metabolite profiling using nonlinear peak alignment, matching, and identification","volume":"78","author":"Smith","year":"2006","journal-title":"Anal. Chem"},{"key":"2023013107253314900_bty699-B35","doi-asserted-by":"crossref","first-page":"3322","DOI":"10.1021\/acs.jproteome.5b00354","article-title":"Analysis of the human adult urinary metabolome variations with age, body mass index, and gender by implementing a comprehensive workflow for univariate and OPLS statistical analyses","volume":"14","author":"Thevenot","year":"2015","journal-title":"J. Proteome Res"},{"key":"2023013107253314900_bty699-B36","doi-asserted-by":"crossref","first-page":"116","DOI":"10.1109\/MS.2015.11","article-title":"Microservices","volume":"32","author":"Th\u00f6nes","year":"2015","journal-title":"IEEE Softw"},{"key":"2023013107253314900_bty699-B37","doi-asserted-by":"crossref","first-page":"1628","DOI":"10.1021\/pr300992u","article-title":"An automated pipeline for high-throughput label-free quantitative proteomics","volume":"12","author":"Weisser","year":"2013","journal-title":"J. Proteome Res"},{"key":"2023013107253314900_bty699-B38","first-page":"95","article-title":"Spark: cluster computing with working sets","volume":"10","author":"Zaharia","year":"2010","journal-title":"HotCloud"},{"key":"2023013107253314900_bty699-B39","first-page":"1","article-title":"Locality-aware scheduling for containers in cloud computing","volume":"99","author":"Zhao","year":"2018","journal-title":"IEEE Trans. Cloud Comput"}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/5\/839\/48965960\/bioinformatics_35_5_839.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/35\/5\/839\/48965960\/bioinformatics_35_5_839.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,4]],"date-time":"2023-09-04T02:34:17Z","timestamp":1693794857000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/35\/5\/839\/5068160"}},"subtitle":[],"editor":[{"given":"Jonathan","family":"Wren","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"editor"}]}],"short-title":[],"issued":{"date-parts":[[2018,8,8]]},"references-count":39,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2019,3,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/bty699","relation":{"has-preprint":[{"id-type":"doi","id":"10.1101\/299032","asserted-by":"object"}]},"ISSN":["1367-4803","1367-4811"],"issn-type":[{"value":"1367-4803","type":"print"},{"value":"1367-4811","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2019,3,1]]},"published":{"date-parts":[[2018,8,8]]}}}