{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,12,30]],"date-time":"2022-12-30T09:54:11Z","timestamp":1672394051806},"reference-count":11,"publisher":"Association for Computing Machinery (ACM)","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2012,8]]},"abstract":"<jats:p>Today large sequencing centers are producing genomic data at the rate of 10 terabytes a day and require complicated processing to transform massive amounts of noisy raw data into biological information. To address these needs, we develop a system for end-to-end processing of genomic data, including alignment of short read sequences, variation discovery, and deep analysis. We also employ a range of quality control mechanisms to improve data quality and parallel processing techniques for performance. In the demo, we will use real genomic data to show details of data transformation through the workflow, the usefulness of end results (ready for use as testable hypotheses), the effects of our quality control mechanisms and improved algorithms, and finally performance improvement.<\/jats:p>","DOI":"10.14778\/2367502.2367534","type":"journal-article","created":{"date-parts":[[2014,6,24]],"date-time":"2014-06-24T12:17:57Z","timestamp":1403612277000},"page":"1906-1909","source":"Crossref","is-referenced-by-count":5,"title":["Massive genomic data processing and deep analysis"],"prefix":"10.14778","volume":"5","author":[{"given":"Abhishek","family":"Roy","sequence":"first","affiliation":[{"name":"University of Massachusetts, Amherst"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yanlei","family":"Diao","sequence":"additional","affiliation":[{"name":"University of Massachusetts, Amherst"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Evan","family":"Mauceli","sequence":"additional","affiliation":[{"name":"Harvard Medical School &amp; Children's Hospital Boston"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yiping","family":"Shen","sequence":"additional","affiliation":[{"name":"Harvard Medical School &amp; Children's Hospital Boston"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Bai-Lin","family":"Wu","sequence":"additional","affiliation":[{"name":"Harvard Medical School &amp; Children's Hospital Boston"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2012,8]]},"reference":[{"issue":"7","key":"e_1_2_1_1_1","doi-asserted-by":"crossref","first-page":"495","DOI":"10.1038\/nmeth0710-495","article-title":"adjusting to data overload","volume":"7","author":"Baker M.","year":"2010","unstructured":"M. Baker . Next-generation sequencing : adjusting to data overload . Nature Method , 7 ( 7 ): 495 -- 499 , 2010 . M. Baker. Next-generation sequencing: adjusting to data overload. Nature Method, 7(7):495--499, 2010.","journal-title":"Nature Method"},{"key":"e_1_2_1_2_1","volume-title":"PAKDD, 47--58","author":"Chui C.","year":"2007","unstructured":"C. Chui , Mining Frequent Itemsets from Uncertain Data . In PAKDD, 47--58 , 2007 . C. Chui, et al. Mining Frequent Itemsets from Uncertain Data. In PAKDD, 47--58, 2007."},{"issue":"3","key":"e_1_2_1_3_1","doi-asserted-by":"crossref","first-page":"R22","DOI":"10.1186\/gb-2012-13-3-r22","article-title":"An integrative probabilistic model for identification of structural variation in sequence data","volume":"13","author":"Sindi S. S.","year":"2012","unstructured":"S. S. Sindi , An integrative probabilistic model for identification of structural variation in sequence data . Genome Biology , 13 ( 3 ): R22 , 2012 . S. S. Sindi, et al. An integrative probabilistic model for identification of structural variation in sequence data. Genome Biology, 13(3):R22, 2012.","journal-title":"Genome Biology"},{"issue":"5","key":"e_1_2_1_4_1","doi-asserted-by":"crossref","first-page":"491","DOI":"10.1038\/ng.806","article-title":"A framework for variation discovery and genotyping using next-generation DNA sequencing data","volume":"43","author":"DePristo M. A.","year":"2011","unstructured":"M. A. DePristo , A framework for variation discovery and genotyping using next-generation DNA sequencing data . Nature Genetics , 43 ( 5 ): 491 -- 498 , 2011 . M. A. DePristo, et al. A framework for variation discovery and genotyping using next-generation DNA sequencing data. Nature Genetics, 43(5):491--498, 2011.","journal-title":"Nature Genetics"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/380995.381002"},{"key":"e_1_2_1_6_1","volume-title":"Searching for SNPs with cloud computing. Genome Biology, 10(11):R134+","author":"Langmead B.","year":"2009","unstructured":"B. Langmead , Searching for SNPs with cloud computing. Genome Biology, 10(11):R134+ , 2009 . B. Langmead, et al. Searching for SNPs with cloud computing. Genome Biology, 10(11):R134+, 2009."},{"issue":"3","key":"e_1_2_1_7_1","doi-asserted-by":"crossref","first-page":"R25","DOI":"10.1186\/gb-2009-10-3-r25","article-title":"Ultrafast and memory-efficient alignment of short DNA sequences to the human genome","volume":"10","author":"Langmead B.","year":"2009","unstructured":"B. Langmead , Ultrafast and memory-efficient alignment of short DNA sequences to the human genome . Genome biology , 10 ( 3 ): R25 , 2009 . B. Langmead, et al. Ultrafast and memory-efficient alignment of short DNA sequences to the human genome. Genome biology, 10(3):R25, 2009.","journal-title":"Genome biology"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1093\/bioinformatics\/btp324"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.artmed.2008.07.008"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1989323.1989370"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1996092.1996106"}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/2367502.2367534","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:50:14Z","timestamp":1672224614000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/2367502.2367534"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,8]]},"references-count":11,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2012,8]]}},"alternative-id":["10.14778\/2367502.2367534"],"URL":"https:\/\/doi.org\/10.14778\/2367502.2367534","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2012,8]]}}}