{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2024,4,2]],"date-time":"2024-04-02T07:49:36Z","timestamp":1712044176520},"reference-count":25,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2022,9,13]],"date-time":"2022-09-13T00:00:00Z","timestamp":1663027200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,9,13]],"date-time":"2022-09-13T00:00:00Z","timestamp":1663027200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>Long interspersed element 1 (LINE-1 or L1) retrotransposons are mobile elements that constitute 17\u201320% of the human genome. Strong correlations between abnormal L1 expression and several human diseases have been reported. This has motivated increasing interest in accurate quantification of the number of L1 copies present in any given biologic specimen. A main obstacle toward this aim is that L1s are relatively long DNA segments with regions of high variability, or largely present in the human genome as truncated fragments. These particularities render traditional alignment strategies, such as seed-and-extend inefficient, as the number of segments that are similar to L1s explodes exponentially. This study uses the pattern matching methodology for more accurate identification of L1s. We validate experimentally the superiority of pattern matching for L1 detection over alternative methods and discuss some of its potential applications.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>Pattern matching detected full-length L1 copies with high precision, reasonable computational time, and no prior input information. It also detected truncated and significantly altered copies of L1 with relatively high precision. The method was effectively used to annotate L1s in a target genome and to calculate copy number variation with respect to a reference genome. Crucial to the success of implementation was the selection of a small set of k-mer probes from a set of sequences presenting a stable pattern of distribution in the genome. As in seed-and-extend methods, the pattern matching algorithm sowed these k-mer probes, but instead of using heuristic extensions around the seeds, the analysis was based on distribution patterns within the genome. The desired level of precision could be adjusted, with some loss of recall.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusion<\/jats:title>\n                <jats:p>Pattern matching is more efficient than seed-and-extend methods for the detection of L1 segments whose characterization depends on a finite set of sequences with common areas of low variability. We propose that pattern matching may help establish correlations between L1 copy number and disease states associated with L1 mobilization and evolution.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s12859-022-04907-4","type":"journal-article","created":{"date-parts":[[2022,9,13]],"date-time":"2022-09-13T15:12:54Z","timestamp":1663081974000},"update-policy":"http:\/\/dx.doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Pattern matching for high precision detection of LINE-1s in human genomes"],"prefix":"10.1186","volume":"23","author":[{"given":"Juan O.","family":"Lopez","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jaime","family":"Seguel","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andres","family":"Chamorro","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kenneth S.","family":"Ramos","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,9,13]]},"reference":[{"key":"4907_CR1","doi-asserted-by":"publisher","first-page":"97","DOI":"10.1186\/gm97","volume":"1","author":"VP Belancio","year":"2009","unstructured":"Belancio VP, Deininger PL, Roy-Engel AM. LINE dancing in the human genome: transposable elements and disease. Genome Med. 2009;1:97. https:\/\/doi.org\/10.1186\/gm97.","journal-title":"Genome Med"},{"key":"4907_CR2","doi-asserted-by":"publisher","first-page":"19","DOI":"10.1038\/ng0598-19","volume":"19","author":"HH Kazazian Jr","year":"1998","unstructured":"Kazazian HH Jr, Moran JV. The impact of L1 retrotransposons on the human genome. Nat Genet. 1998;19:19\u201324. https:\/\/doi.org\/10.1038\/ng0598-19.","journal-title":"Nat Genet"},{"key":"4907_CR3","doi-asserted-by":"publisher","first-page":"9","DOI":"10.1186\/s13100-016-0065-9","volume":"7","author":"DC Hancks","year":"2016","unstructured":"Hancks DC, Kazazian HH Jr. Roles for retrotransposon insertions in human disease. Mob DNA. 2016;7:9. https:\/\/doi.org\/10.1186\/s13100-016-0065-9.","journal-title":"Mob DNA"},{"key":"4907_CR4","doi-asserted-by":"publisher","first-page":"498","DOI":"10.1093\/nar\/gki044","volume":"33","author":"T Penzkofer","year":"2004","unstructured":"Penzkofer T, Dandekar T, T Z. L1Base: from functional annotation to prediction of active LINE-1 elements. Nucl Acids Res. 2004;33:498\u2013500. https:\/\/doi.org\/10.1093\/nar\/gki044.","journal-title":"Nucl Acids Res"},{"key":"4907_CR5","unstructured":"L1Base 2. Accessed 7-September-2020. http:\/\/l1base.charite.de\/"},{"key":"4907_CR6","doi-asserted-by":"publisher","DOI":"10.1093\/nar\/gkw925","author":"T Penzkofer","year":"2016","unstructured":"Penzkofer T, J\u00e4ger M, Figlerowicz M, Badge R, Mundlos S, Robinson PN, Zemojtel T. L1Base 2: more retrotransposition-active LINE-1s, more mammalian genomes. Nucl Acids Res. 2016. https:\/\/doi.org\/10.1093\/nar\/gkw925.","journal-title":"Nucl Acids Res"},{"issue":"12","key":"4907_CR7","doi-asserted-by":"publisher","first-page":"350","DOI":"10.1093\/bioinformatics\/btq216","volume":"26","author":"F Hormozdiari","year":"2010","unstructured":"Hormozdiari F, Hajirasouliha I, Dao P, Hach F, Yorukoglu D, Alkan C, Eichler EE, Cenk Sahinalp S. Next-generation VariationHunter: combinatorial algorithms for transposon insertion discovery. Bioinformatics. 2010;26(12):350\u20137. https:\/\/doi.org\/10.1093\/bioinformatics\/btq216.","journal-title":"Bioinformatics"},{"issue":"6097","key":"4907_CR8","doi-asserted-by":"publisher","first-page":"967","DOI":"10.1126\/science.1222077","volume":"337","author":"E Lee","year":"2012","unstructured":"Lee E, Iskow R, Yang L, Gokcumen O, Haseley P, Luquette LJ III, Lohr JG, Harris CC, Ding L, Wilson RK, Wheeler DA, Gibbs RA, Kucherlapati R, Lee C, Kharchenko PV, Park PJ. The cancer genome atlas research network: landscape of somatic retrotransposition in human cancers. Science. 2012;337(6097):967\u201371. https:\/\/doi.org\/10.1126\/science.1222077.","journal-title":"Science"},{"issue":"3","key":"4907_CR9","doi-asserted-by":"publisher","first-page":"389","DOI":"10.1093\/bioinformatics\/bts697","volume":"29","author":"T Keane","year":"2012","unstructured":"Keane T, Wong K, D A. RetroSeq: transposable element discovery from next-generation sequencing data. Bioinformatics. 2012;29(3):389\u201390. https:\/\/doi.org\/10.1093\/bioinformatics\/bts697.","journal-title":"Bioinformatics"},{"key":"4907_CR10","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2164-15-795","author":"J Wu","year":"2014","unstructured":"Wu J, Lee W, Ward A, Walker J, Konkel M, Batzer MGM. Tangram: a comprehensive toolbox for mobile element insertion detection. BMC Genom. 2014. https:\/\/doi.org\/10.1186\/1471-2164-15-795.","journal-title":"BMC Genom"},{"key":"4907_CR11","unstructured":"Steinbiss S. Repeat M. Accessed 25-May-2021. http:\/\/www.repeatmasker.org\/"},{"issue":"3","key":"4907_CR12","doi-asserted-by":"publisher","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","volume":"215","author":"S Altschul","year":"1990","unstructured":"Altschul S, Gish W, Miller W, Myers E, Lipman D. Basic local alignment search tool. J Mol Biol. 1990;215(3):403\u201310. https:\/\/doi.org\/10.1016\/S0022-2836(05)80360-2.","journal-title":"J Mol Biol"},{"issue":"D1","key":"4907_CR13","doi-asserted-by":"publisher","first-page":"854","DOI":"10.1093\/nar\/gkw829","volume":"45","author":"L Clarke","year":"2016","unstructured":"Clarke L, Fairley S, Zheng-Bradley X, Streeter I, Perry E, Lowy E, Tass\u00e9 A-M, Flicek P. The international Genome sample resource (IGSR): a worldwide collection of genome variation incorporating the 1000 Genomes Project data. Nucl Acids Res. 2016;45(D1):854\u20139. https:\/\/doi.org\/10.1093\/nar\/gkw829.","journal-title":"Nucl Acids Res"},{"key":"4907_CR14","doi-asserted-by":"publisher","DOI":"10.1186\/s12859-018-2315-y","author":"A Babaian","year":"2018","unstructured":"Babaian A, Ebou A, et al. bioSyntax: syntax highlighting for computational biology. BMC Bioinform. 2018. https:\/\/doi.org\/10.1186\/s12859-018-2315-y.","journal-title":"BMC Bioinform"},{"key":"4907_CR15","doi-asserted-by":"publisher","first-page":"1061","DOI":"10.1038\/ng.437","volume":"41","author":"C Alkan","year":"2009","unstructured":"Alkan C, Kidd J, Marques-Bonet T, Aksay G, Antonacci F, Hormozdiari F, et al. Personalized copy number and segmental duplication maps using next-generation sequencing. Nat Genet. 2009;41:1061\u20137. https:\/\/doi.org\/10.1038\/ng.437.","journal-title":"Nat Genet"},{"issue":"Suppl 1","key":"4907_CR16","doi-asserted-by":"publisher","first-page":"13","DOI":"10.1186\/1471-2164-14-S1-S13","volume":"14","author":"H Xin","year":"2013","unstructured":"Xin H, Lee D, Hormozdiari F, Yedkar S, Mutlu OCA. Accelerating read mapping with FastHASH. BMC Genom. 2013;14(Suppl 1):13.","journal-title":"BMC Genom"},{"key":"4907_CR17","unstructured":"van Rijsbergen CJ. Evaluation. In: Information retrieval, 2nd ed. Butterworth-Heinemann: Glasgow, Scotland; 1979, pp. 112\u2013140."},{"issue":"14","key":"4907_CR18","doi-asserted-by":"publisher","first-page":"1754","DOI":"10.1093\/bioinformatics\/btp324","volume":"25","author":"H Li","year":"2009","unstructured":"Li H, Durbin R. Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics. 2009;25(14):1754\u201360. https:\/\/doi.org\/10.1093\/bioinformatics\/btp324.","journal-title":"Bioinformatics"},{"issue":"16","key":"4907_CR19","doi-asserted-by":"publisher","first-page":"2078","DOI":"10.1093\/bioinformatics\/btp352","volume":"25","author":"H Li","year":"2009","unstructured":"Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R, Subgroup GPDP. The sequence alignment\/Map format and SAMtools. Bioinformatics. 2009;25(16):2078\u20139. https:\/\/doi.org\/10.1093\/bioinformatics\/btp352.","journal-title":"Bioinformatics"},{"issue":"11","key":"4907_CR20","doi-asserted-by":"publisher","first-page":"1422","DOI":"10.1093\/bioinformatics\/btp163","volume":"25","author":"PJA Cock","year":"2009","unstructured":"Cock PJA, Antao T, Chang JT, Chapman BA, Cox CJ, Dalke A, Friedberg I, Hamelryck T, Kauff F, Wilczynski B, de Hoon MJL. Biopython: freely available Python tools for computational molecular biology and bioinformatics. Bioinformatics. 2009;25(11):1422\u20133. https:\/\/doi.org\/10.1093\/bioinformatics\/btp163.","journal-title":"Bioinformatics"},{"key":"4907_CR21","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-13-238","author":"MJ Chaisson","year":"2012","unstructured":"Chaisson MJ, Tesler G. Mapping single molecule sequencing reads using basic local alignment with successive refinement (BLASR): application and theory. BMC Bioinform. 2012. https:\/\/doi.org\/10.1186\/1471-2105-13-238.","journal-title":"BMC Bioinform"},{"key":"4907_CR22","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1005944","author":"G Mar\u00e7ais","year":"2018","unstructured":"Mar\u00e7ais G, Delcher AL, Phillippy AM, et al. MUMmer4: a fast and versatile genome alignment system. PLOS Comput Biol. 2018. https:\/\/doi.org\/10.1371\/journal.pcbi.1005944.","journal-title":"PLOS Comput Biol"},{"key":"4907_CR23","unstructured":"Steinbiss S. GFF3 Online Validator. Accessed 7-September-2020. http:\/\/genometools.org\/cgi-bin\/gff3validator.cgi"},{"key":"4907_CR24","doi-asserted-by":"publisher","DOI":"10.1186\/1471-2105-16-S17-S3","author":"V Phan","year":"2015","unstructured":"Phan V, Gao S, Tran Q, et al. How genome complexity can explain the difficulty of aligning reads to genomes. BMC Bioinform. 2015. https:\/\/doi.org\/10.1186\/1471-2105-16-S17-S3.","journal-title":"BMC Bioinform"},{"issue":"22","key":"4907_CR25","doi-asserted-by":"publisher","first-page":"4048","DOI":"10.1093\/bioinformatics\/btab408","volume":"37","author":"F Almodaresi","year":"2021","unstructured":"Almodaresi F, Zakeri M, Patro R. PuffAligner: a fast, efficient and accurate aligner based on the Pufferfish index. Bioinformatics. 2021;37(22):4048\u201355. https:\/\/doi.org\/10.1093\/bioinformatics\/btab408.","journal-title":"Bioinformatics"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-022-04907-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-022-04907-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-022-04907-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,9,13]],"date-time":"2022-09-13T15:13:04Z","timestamp":1663081984000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-022-04907-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,13]]},"references-count":25,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2022,12]]}},"alternative-id":["4907"],"URL":"https:\/\/doi.org\/10.1186\/s12859-022-04907-4","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,13]]},"assertion":[{"value":"26 October 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 August 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 September 2022","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"All methods were performed in accordance with the relevant guidelines and regulations.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"375"}}