{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T22:17:33Z","timestamp":1765232253385},"reference-count":12,"publisher":"Oxford University Press (OUP)","issue":"23","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2015,12,1]]},"abstract":"<jats:title>Abstract<\/jats:title>\n               <jats:p>Motivation: Unambiguous sequence variant descriptions are important in reporting the outcome of clinical diagnostic DNA tests. The standard nomenclature of the Human Genome Variation Society (HGVS) describes the observed variant sequence relative to a given reference sequence. We propose an efficient algorithm for the extraction of HGVS descriptions from two sequences with three main requirements in mind: minimizing the length of the resulting descriptions, minimizing the computation time and keeping the unambiguous descriptions biologically meaningful.<\/jats:p>\n               <jats:p>Results: Our algorithm is able to compute the HGVS descriptions of complete chromosomes or other large DNA strings in a reasonable amount of computation time and its resulting descriptions are relatively small. Additional applications include updating of gene variant database contents and reference sequence liftovers.<\/jats:p>\n               <jats:p>Availability: The algorithm is accessible as an experimental service in the Mutalyzer program suite (https:\/\/mutalyzer.nl). The C++ source code and Python interface are accessible at: https:\/\/github.com\/mutalyzer\/description-extractor.<\/jats:p>\n               <jats:p>Contact: \u00a0j.k.vis@lumc.nl<\/jats:p>","DOI":"10.1093\/bioinformatics\/btv443","type":"journal-article","created":{"date-parts":[[2015,8,1]],"date-time":"2015-08-01T00:40:08Z","timestamp":1438389608000},"page":"3751-3757","source":"Crossref","is-referenced-by-count":27,"title":["An efficient algorithm for the extraction of HGVS variant descriptions from sequences"],"prefix":"10.1093","volume":"31","author":[{"given":"Jonathan K.","family":"Vis","sequence":"first","affiliation":[{"name":"1 Department of Molecular Epidemiology, Leiden University Medical Center, Leiden, The Netherlands,"},{"name":"2 Leiden Institute of Advanced Computer Science, Leiden University, Leiden, The Netherlands,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Martijn","family":"Vermaat","sequence":"additional","affiliation":[{"name":"3 Department of Human Genetics, Leiden University Medical Center, Leiden, The Netherlands,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peter E. M.","family":"Taschner","sequence":"additional","affiliation":[{"name":"3 Department of Human Genetics, Leiden University Medical Center, Leiden, The Netherlands,"},{"name":"4 Generade Center of Expertise Genomics, University of Applied Sciences Leiden, Leiden, The Netherlands and"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joost N.","family":"Kok","sequence":"additional","affiliation":[{"name":"1 Department of Molecular Epidemiology, Leiden University Medical Center, Leiden, The Netherlands,"},{"name":"2 Leiden Institute of Advanced Computer Science, Leiden University, Leiden, The Netherlands,"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jeroen F. J.","family":"Laros","sequence":"additional","affiliation":[{"name":"3 Department of Human Genetics, Leiden University Medical Center, Leiden, The Netherlands,"},{"name":"5 Leiden Genome Technology Center, Leiden University Medical Center, Leiden, The Netherlands"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"286","published-online":{"date-parts":[[2015,7,31]]},"reference":[{"key":"2023020202420587400_btv443-B1","author":"Abbasi","year":"1997"},{"key":"2023020202420587400_btv443-B2","doi-asserted-by":"crossref","first-page":"1731","DOI":"10.1093\/bioinformatics\/btp319","article-title":"Data structures and compression algorithms for genomic sequence data","volume":"25","author":"Brandon","year":"2009","journal-title":"Bioinformatics"},{"key":"2023020202420587400_btv443-B3","doi-asserted-by":"crossref","first-page":"987","DOI":"10.1038\/nbt.2023","article-title":"How to apply de Bruijn graphs to genome assembly","volume":"29","author":"Compeau","year":"2011","journal-title":"Nat. Biotechnol."},{"key":"2023020202420587400_btv443-B4","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1002\/(SICI)1098-1004(200001)15:1<7::AID-HUMU4>3.0.CO;2-N","article-title":"Mutation nomenclature extensions and suggestions to describe complex mutations: a discussion","volume":"15","author":"Den Dunnen","year":"2000","journal-title":"Hum. Mutat."},{"key":"2023020202420587400_btv443-B5","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511574931","volume-title":"Algorithms on Strings, Trees and Sequences: Computer Science and Computational Biology","author":"Gusfield","year":"1997"},{"key":"2023020202420587400_btv443-B6","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1093\/nar\/gku1161","article-title":"The IPD and IMGT\/HLA database: Allele variant databases","volume":"43","author":"Robinson","year":"2015","journal-title":"Nucleic Acids Res."},{"key":"2023020202420587400_btv443-B7","doi-asserted-by":"crossref","first-page":"195","DOI":"10.1016\/0022-2836(81)90087-5","article-title":"Identification of common molecular subsequences","volume":"147","author":"Smith","year":"1981","journal-title":"J. Mol. Biol."},{"key":"2023020202420587400_btv443-B8","doi-asserted-by":"crossref","first-page":"507","DOI":"10.1002\/humu.21427","article-title":"Describing structural changes by extending HGVS sequence variation nomenclature","volume":"32","author":"Taschner","year":"2011","journal-title":"Hum. Mutat."},{"key":"2023020202420587400_btv443-B9","doi-asserted-by":"crossref","first-page":"309","DOI":"10.1145\/357401.357404","article-title":"The string-to-string correction problem with block moves","volume":"2","author":"Tichy","year":"1984","journal-title":"TOCS"},{"key":"2023020202420587400_btv443-B10","doi-asserted-by":"crossref","first-page":"168","DOI":"10.1145\/321796.321811","article-title":"The string-to-string correction problem","volume":"21","author":"Wagner","year":"1974","journal-title":"JACM"},{"key":"2023020202420587400_btv443-B11","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1145\/321879.321880","article-title":"An extension of the string-to-string correction problem","volume":"22","author":"Wagner","year":"1975","journal-title":"JACM"},{"key":"2023020202420587400_btv443-B12","doi-asserted-by":"crossref","first-page":"6","DOI":"10.1002\/humu.20654","article-title":"Improving sequence variant descriptions in mutation databases and literature using the Mutalyzer sequence variation nomenclature checker","volume":"29","author":"Wildeman","year":"2008","journal-title":"Hum. Mutat."}],"container-title":["Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/23\/3751\/49036380\/bioinformatics_31_23_3751.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article-pdf\/31\/23\/3751\/49036380\/bioinformatics_31_23_3751.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,2]],"date-time":"2023-02-02T03:57:17Z","timestamp":1675310237000},"score":1,"resource":{"primary":{"URL":"https:\/\/academic.oup.com\/bioinformatics\/article\/31\/23\/3751\/208405"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2015,7,31]]},"references-count":12,"journal-issue":{"issue":"23","published-print":{"date-parts":[[2015,12,1]]}},"URL":"https:\/\/doi.org\/10.1093\/bioinformatics\/btv443","relation":{},"ISSN":["1367-4811","1367-4803"],"issn-type":[{"value":"1367-4811","type":"electronic"},{"value":"1367-4803","type":"print"}],"subject":[],"published-other":{"date-parts":[[2015,12,1]]},"published":{"date-parts":[[2015,7,31]]}}}