{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,2]],"date-time":"2026-03-02T05:42:03Z","timestamp":1772430123682,"version":"3.50.1"},"reference-count":27,"publisher":"Oxford University Press (OUP)","license":[{"start":{"date-parts":[[2022,5,11]],"date-time":"2022-05-11T00:00:00Z","timestamp":1652227200000},"content-version":"vor","delay-in-days":130,"URL":"https:\/\/creativecommons.org\/licenses\/by-nc\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000054","name":"National Cancer Institute","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000054","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,5,11]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Cancer is a somatic disease. The lack of Indian-specific reference germline variation resources limits the ability to identify true cancer-associated somatic variants among Indian cancer patients. We integrate two recent studies, the GenomeAsia 100K and the Genomics for Public Health in India (IndiGen) program, describing genome sequence variations across 598 and 1029 healthy individuals of Indian origin, respectively, along with the unique variants generated from our in-house 173 normal germline samples derived from cancer patients to generate the Tata Memorial Centre-SNP database (TMC-SNPdb) 2.0. To show its utility, GATK\/Mutect2-based somatic variant calling was performed on 224 in-house tumor samples to demonstrate a reduction in false-positive somatic variants. In addition to the ethnic-specific variants from GenomeAsia 100K and IndiGenomes databases, 305\u2009132 unique variants generated from 173 in-house normal germline samples derived from cancer patients of Indian origin constitute the Indian specific, TMC-SNPdb 2.0. Of 305\u2009132 unique variants, 11.13% were found in the coding region with missense variants (31.3%) as the most predominant category. Among the non-coding variations, intronic variants (49%) were the highest contributors. The non-synonymous to synonymous SNP ratio was observed to be 1.9, consistent with the previous version of TMC-SNPdb and literature. Using TMC SNPdb 2.0, we analyzed a whole-exome sequence from 224 in-house tumor samples (180 paired and 44 orphans). We show an average depletion of 3.44% variants per paired tumor and significantly higher depletion (P-value\u2009&amp;lt;\u20090.001) for orphan tumors (4.21%), demonstrating the utility of the rare, unique variants found in the ethnic-specific variant datasets in reducing the false-positive somatic mutations. TMC-SNPdb 2.0 is the most exhaustive open-source reference database of germline variants occurring across 1800 Indian individuals to analyze cancer genomes and other genetic disorders. The database and toolkit package is available for download at the following:<\/jats:p><jats:p>Database URL \u00a0http:\/\/www.actrec.gov.in\/pi-webpages\/AmitDutt\/TMCSNPdb2\/TMCSNPdb2.html<\/jats:p>","DOI":"10.1093\/database\/baac029","type":"journal-article","created":{"date-parts":[[2022,4,14]],"date-time":"2022-04-14T11:10:59Z","timestamp":1649934659000},"source":"Crossref","is-referenced-by-count":2,"title":["TMC-SNPdb 2.0: an ethnic-specific database of Indian germline variants"],"prefix":"10.1093","volume":"2022","author":[{"given":"Sanket","family":"Desai","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Rohit","family":"Mishra","sequence":"additional","affiliation":[{"name":"Integrated Cancer Genomics Laboratory, Advanced Centre for Treatment, Research, and Education in Cancer, Kharghar, Navi Mumbai, Maharashtra 410210, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Suhail","family":"Ahmad","sequence":"additional","affiliation":[{"name":"Integrated Cancer Genomics Laboratory, Advanced Centre for Treatment, Research, and Education in Cancer, Kharghar, Navi Mumbai, Maharashtra 410210, India"},{"name":"Homi Bhabha National Institute, Training School Complex, Anushakti Nagar, Mumbai, Maharashtra 400094, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Supriya","family":"Hait","sequence":"additional","affiliation":[{"name":"Integrated Cancer Genomics Laboratory, Advanced Centre for Treatment, Research, and Education in Cancer, Kharghar, Navi Mumbai, Maharashtra 410210, India"},{"name":"Homi Bhabha National Institute, Training School Complex, Anushakti Nagar, Mumbai, Maharashtra 400094, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Asim","family":"Joshi","sequence":"additional","affiliation":[{"name":"Integrated Cancer Genomics Laboratory, Advanced Centre for Treatment, Research, and Education in Cancer, Kharghar, Navi Mumbai, Maharashtra 410210, India"},{"name":"Homi Bhabha National Institute, Training School Complex, Anushakti Nagar, Mumbai, Maharashtra 400094, India"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1119-4774","authenticated-orcid":false,"given":"Amit","family":"Dutt","sequence":"additional","affiliation":[{"name":"Integrated Cancer Genomics Laboratory, Advanced Centre for Treatment, Research, and Education in Cancer, Kharghar, Navi Mumbai, Maharashtra 410210, India"},{"name":"Homi Bhabha National Institute, Training School Complex, Anushakti Nagar, Mumbai, Maharashtra 400094, India"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2022,5,11]]},"reference":[{"key":"2022051115203051200_R1","doi-asserted-by":"crossref","first-page":"285","DOI":"10.1038\/nature19057","article-title":"Analysis of protein-coding genetic variation in 60,706 humans","volume":"536","author":"Lek","year":"2016","journal-title":"Nature"},{"key":"2022051115203051200_R2","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1038\/538161a","article-title":"Genomics is failing on diversity","volume":"538","author":"Popejoy","year":"2016","journal-title":"Nature"},{"key":"2022051115203051200_R3","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1038\/nature15393","article-title":"A global reference for human genetic variation","volume":"526","author":"1000 Genomes Project Consortium","year":"2015","journal-title":"Nature"},{"key":"2022051115203051200_R4","doi-asserted-by":"crossref","first-page":"434","DOI":"10.1038\/s41586-020-2308-7","article-title":"The mutational constraint spectrum quantified from variation in 141,456 humans","volume":"581","author":"Karczewski","year":"2020","journal-title":"Nature"},{"key":"2022051115203051200_R5","doi-asserted-by":"crossref","first-page":"327","DOI":"10.1038\/nature13997","article-title":"The African Genome Variation Project shapes medical genetics in Africa","volume":"517","author":"Gurdasani","year":"2015","journal-title":"Nature"},{"key":"2022051115203051200_R6","doi-asserted-by":"crossref","DOI":"10.1038\/ncomms9018","article-title":"Rare variant discovery by deep whole-genome sequencing of 1,070 Japanese individuals","volume":"6","author":"Nagasaki","year":"2015","journal-title":"Nat. Commun."},{"key":"2022051115203051200_R7","doi-asserted-by":"crossref","first-page":"1071","DOI":"10.1038\/ng.3592","article-title":"Characterization of Greater Middle Eastern genetic variation for enhanced disease gene discovery","volume":"48","author":"Scott","year":"2016","journal-title":"Nat. Genet."},{"key":"2022051115203051200_R8","doi-asserted-by":"crossref","DOI":"10.1016\/j.celrep.2021.110017","article-title":"NyuWa Genome resource: a deep whole-genome sequencing-based variation profile and reference panel for the Chinese population","volume":"37","author":"Zhang","year":"2021","journal-title":"Cell Rep."},{"key":"2022051115203051200_R9","doi-asserted-by":"crossref","first-page":"106","DOI":"10.1038\/s41586-019-1793-z","article-title":"The GenomeAsia 100K Project enables genetic discoveries across Asia","volume":"576","author":"GenomeAsia","year":"2019","journal-title":"Nature"},{"key":"2022051115203051200_R10","first-page":"D1225","article-title":"IndiGenomes: a comprehensive resource of genetic variants from over 1000 Indian genomes","volume":"49","author":"Jain","year":"2021","journal-title":"Nucleic Acids Res."},{"key":"2022051115203051200_R11","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/database\/baw104","article-title":"TMC-SNPdb: an Indian germline variant database derived from whole exome sequences","volume":"2016","author":"Upadhyay","year":"2016","journal-title":"Database (Oxford)"},{"key":"2022051115203051200_R12","doi-asserted-by":"crossref","first-page":"308","DOI":"10.1093\/nar\/29.1.308","article-title":"dbSNP: the NCBI database of genetic variation","volume":"29","author":"Sherry","year":"2001","journal-title":"Nucleic Acids Res."},{"key":"2022051115203051200_R13","doi-asserted-by":"crossref","first-page":"D941","DOI":"10.1093\/nar\/gky1015","article-title":"COSMIC: the catalogue of somatic mutations in cancer","volume":"47","author":"Tate","year":"2019","journal-title":"Nucleic Acids Res."},{"key":"2022051115203051200_R14","doi-asserted-by":"crossref","first-page":"718","DOI":"10.1093\/bioinformatics\/btq671","article-title":"Tabix: fast retrieval of sequence features from generic TAB-delimited files","volume":"27","author":"Li","year":"2011","journal-title":"Bioinformatics"},{"key":"2022051115203051200_R15","doi-asserted-by":"crossref","first-page":"1297","DOI":"10.1101\/gr.107524.110","article-title":"The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data","volume":"20","author":"McKenna","year":"2010","journal-title":"Genome Res."},{"key":"2022051115203051200_R16","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1002\/0471250953.bi1110s43","article-title":"From FastQ data to high confidence variant calls: the Genome Analysis Toolkit best practices pipeline","volume":"43","author":"Van der Auwera","year":"2013","journal-title":"Curr. Protoc. Bioinf."},{"key":"2022051115203051200_R17","first-page":"1","article-title":"Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM","author":"Li","year":"2013"},{"key":"2022051115203051200_R18","doi-asserted-by":"crossref","first-page":"D916","DOI":"10.1093\/nar\/gkaa1087","article-title":"Gencode 2021","volume":"49","author":"Frankish","year":"2021","journal-title":"Nucleic Acids Res."},{"key":"2022051115203051200_R19","doi-asserted-by":"crossref","DOI":"10.1186\/s13059-016-0974-4","article-title":"The ensembl variant effect predictor","volume":"17","author":"McLaren","year":"2016","journal-title":"Genome Biol."},{"key":"2022051115203051200_R20","article-title":"ALFA: allele frequency aggregator","author":"Phan","year":"2020"},{"key":"2022051115203051200_R21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1093\/gigascience\/giab008","article-title":"Twelve years of SAMtools and BCFtools","volume":"10","author":"Danecek","year":"2021","journal-title":"Gigascience"},{"key":"2022051115203051200_R22","doi-asserted-by":"crossref","DOI":"10.1186\/1471-2164-13-194","article-title":"Exome sequencing generates high quality data in non-target regions","volume":"13","author":"Guo","year":"2012","journal-title":"BMC Genomics"},{"key":"2022051115203051200_R23","doi-asserted-by":"crossref","first-page":"W452","DOI":"10.1093\/nar\/gks539","article-title":"SIFT web server: predicting effects of amino acid substitutions on proteins","volume":"40","author":"Sim","year":"2012","journal-title":"Nucleic Acids Res."},{"key":"2022051115203051200_R24","article-title":"Predicting functional effect of human missense mutations using PolyPhen-2","volume":"76","author":"Adzhubei","year":"2013","journal-title":"Curr. Protoc. Hum. Genet."},{"key":"2022051115203051200_R25","doi-asserted-by":"crossref","DOI":"10.1186\/s13073-017-0424-2","article-title":"Analysis of 100,000 human cancer genomes reveals the landscape of tumor mutational burden","volume":"9","author":"Chalmers","year":"2017","journal-title":"Genome Med."},{"key":"2022051115203051200_R26","doi-asserted-by":"crossref","DOI":"10.1186\/s13073-020-00791-w","article-title":"Best practices for variant calling in clinical sequencing","volume":"12","author":"Koboldt","year":"2020","journal-title":"Genome Med."},{"key":"2022051115203051200_R27","doi-asserted-by":"crossref","DOI":"10.1186\/1475-2867-7-2","article-title":"Clinical implications and utility of field cancerization","volume":"7","author":"Dakubo","year":"2007","journal-title":"Cancer Cell Int."}],"container-title":["Database"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baac029\/43679114\/baac029.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/database\/article-pdf\/doi\/10.1093\/database\/baac029\/43679114\/baac029.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,1]],"date-time":"2023-02-01T23:29:02Z","timestamp":1675294142000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/database\/article\/doi\/10.1093\/database\/baac029\/6583650"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,1,1]]},"references-count":27,"URL":"https:\/\/doi.org\/10.1093\/database\/baac029","relation":{},"ISSN":["1758-0463"],"issn-type":[{"value":"1758-0463","type":"electronic"}],"subject":[],"published-other":{"date-parts":[[2022,1,1]]},"published":{"date-parts":[[2022,1,1]]},"article-number":"baac029"}}