{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,26]],"date-time":"2026-03-26T12:18:47Z","timestamp":1774527527337,"version":"3.50.1"},"reference-count":30,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,10,19]],"date-time":"2020-10-19T00:00:00Z","timestamp":1603065600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,10,19]],"date-time":"2020-10-19T00:00:00Z","timestamp":1603065600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No.61732009"],"award-info":[{"award-number":["No.61732009"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["No.61772557"],"award-info":[{"award-number":["No.61772557"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Hunan Provincial Science and technology Program","award":["No. 2018wk4001"],"award-info":[{"award-number":["No. 2018wk4001"]}]},{"DOI":"10.13039\/501100013314","name":"111 Project","doi-asserted-by":"crossref","award":["No.B18059"],"award-info":[{"award-number":["No.B18059"]}],"id":[{"id":"10.13039\/501100013314","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n<jats:title>Background<\/jats:title>\n<jats:p>Repetitive sequences account for a large proportion of eukaryotes genomes. Identification of repetitive sequences plays a significant role in many applications, such as structural variation detection and genome assembly. Many existing de novo repeat identification pipelines or tools make use of assembly of the high-frequency <jats:italic>k-mers<\/jats:italic> to obtain repeats. However, a certain degree of sequence coverage is required for assemblers to get the desired assemblies. On the other hand, assemblers cut the reads into shorter <jats:italic>k-mers<\/jats:italic> for assembly, which may destroy the structure of the repetitive regions. For the above reasons, it is difficult to obtain complete and accurate repetitive regions in the genome by using existing tools.<\/jats:p>\n<\/jats:sec><jats:sec>\n<jats:title>Results<\/jats:title>\n<jats:p>In this study, we present a new method called RepAHR for de novo repeat identification by assembly of the high-frequency reads. Firstly, RepAHR scans next-generation sequencing (NGS) reads to find the high-frequency <jats:italic>k-mers<\/jats:italic>. Secondly, RepAHR filters the high-frequency reads from whole NGS reads according to certain rules based on the high-frequency <jats:italic>k-mer<\/jats:italic>. Finally, the high-frequency reads are assembled to generate repeats by using SPAdes, which is considered as an outstanding genome assembler with NGS sequences.<\/jats:p>\n<\/jats:sec><jats:sec>\n<jats:title>Conlusions<\/jats:title>\n<jats:p>We test RepAHR on five data sets, and the experimental results show that RepAHR outperforms RepARK and REPdenovo for detecting repeats in terms of N50, reference alignment ratio, coverage ratio of reference, mask ratio of Repbase and some other metrics.<\/jats:p>\n<\/jats:sec>","DOI":"10.1186\/s12859-020-03779-w","type":"journal-article","created":{"date-parts":[[2020,10,19]],"date-time":"2020-10-19T13:04:25Z","timestamp":1603112665000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["RepAHR: an improved approach for de novo repeat identification by assembly of the high-frequency reads"],"prefix":"10.1186","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0061-1317","authenticated-orcid":false,"given":"Xingyu","family":"Liao","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xin","family":"Gao","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiankai","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Fang-Xiang","family":"Wu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jianxin","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2020,10,19]]},"reference":[{"issue":"6","key":"3779_CR1","doi-asserted-by":"publisher","first-page":"787","DOI":"10.1007\/s10577-011-9230-7","volume":"19","author":"M Janicki","year":"2011","unstructured":"Janicki M, Rooke R, Yang G. Bioinformatics and genomic analysis of transposable elements in eukaryotic genomes. Chromosome Res. 2011;19(6):787. https:\/\/doi.org\/10.1007\/s10577-011-9230-7.","journal-title":"Chromosome Res"},{"issue":"12","key":"3779_CR2","doi-asserted-by":"publisher","first-page":"1002384","DOI":"10.1371\/journal.pgen.1002384","volume":"7","author":"AJ de Koning","year":"2011","unstructured":"de Koning AJ, Gu W, Castoe TA, Batzer MA, Pollock DD. Repetitive elements may comprise over two-thirds of the human genome. PLoS Genet. 2011;7(12):1002384. https:\/\/doi.org\/10.1371\/journal.pgen.1002384.","journal-title":"PLoS Genet"},{"issue":"suppl 1","key":"3779_CR3","doi-asserted-by":"publisher","first-page":"360","DOI":"10.1093\/nar\/gkh099","volume":"32","author":"S Ouyang","year":"2004","unstructured":"Ouyang S, Buell CR. The TIGR plant repeat databases: a collective resource for the identification of repetitive sequences in plants. Nucleic Acids Res. 2004;32(suppl 1):360\u20133. https:\/\/doi.org\/10.1093\/nar\/gkh099.","journal-title":"Nucleic Acids Res"},{"issue":"2","key":"3779_CR4","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1023\/B:GENE.0000040382.48039.a","volume":"121","author":"JP Castro","year":"2004","unstructured":"Castro JP, Carareto CM. Drosophila melanogaster P transposable elements: mechanisms of transposition and regulation. Genetica. 2004;121(2):107\u201318. https:\/\/doi.org\/10.1023\/B:GENE.0000040382.48039.a.","journal-title":"Genetica"},{"issue":"1","key":"3779_CR5","doi-asserted-by":"publisher","first-page":"36","DOI":"10.1038\/nrg3117","volume":"13","author":"TJ Treangen","year":"2012","unstructured":"Treangen TJ, Salzberg SL. Repetitive DNA and next-generation sequencing: computational challenges and solutions. Nat Rev Genet. 2012;13(1):36. https:\/\/doi.org\/10.1038\/nrg3117.","journal-title":"Nat Rev Genet"},{"issue":"1","key":"3779_CR6","doi-asserted-by":"publisher","first-page":"517","DOI":"10.1186\/1471-2164-9-517","volume":"9","author":"S Kurtz","year":"2008","unstructured":"Kurtz S, Narechania A, Stein JC, Ware D. A new method to compute K-mer frequencies and its application to annotate large repetitive plant genomes. BMC Genomics. 2008;9(1):517. https:\/\/doi.org\/10.1186\/1471-2164-9-517.","journal-title":"BMC Genomics"},{"issue":"6","key":"3779_CR7","doi-asserted-by":"publisher","first-page":"792","DOI":"10.1093\/bioinformatics\/btt054","volume":"29","author":"P Nov\u00e1k","year":"2013","unstructured":"Nov\u00e1k P, Neumann P, Pech J, Steinhaisl J, Macas J. RepeatExplorer: a Galaxy-based web server for genome-wide characterization of eukaryotic repetitive elements from next-generation sequence reads. Bioinformatics. 2013;29(6):792\u20133. https:\/\/doi.org\/10.1093\/bioinformatics\/btt054.","journal-title":"Bioinformatics"},{"issue":"9","key":"3779_CR8","doi-asserted-by":"publisher","first-page":"80","DOI":"10.1093\/nar\/gku210","volume":"42","author":"P Koch","year":"2014","unstructured":"Koch P, Platzer M, Downie BR. RepARK-de novo creation of repeat libraries from whole-genome NGS reads. Nucleic Acids Res. 2014;42(9):80. https:\/\/doi.org\/10.1093\/nar\/gku210.","journal-title":"Nucleic Acids Res"},{"issue":"3","key":"3779_CR9","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1186\/1755-8794-8-S3-S5","volume":"8","author":"G Fertin","year":"2015","unstructured":"Fertin G, Jean G, Radulescu A, Rusu I. Hybrid de novo tandem repeat detection using short and long reads. BMC Med Genomics. 2015;8(3):5. https:\/\/doi.org\/10.1186\/1755-8794-8-S3-S5.","journal-title":"BMC Med Genomics"},{"issue":"7","key":"3779_CR10","doi-asserted-by":"publisher","first-page":"1099","DOI":"10.1093\/bioinformatics\/btx717","volume":"34","author":"R Guo","year":"2017","unstructured":"Guo R, Li Y-R, He S, Ou-Yang L, Sun Y, Zhu Z. RepLong: de novo repeat identification using long read sequencing data. Bioinformatics. 2017;34(7):1099\u2013107. https:\/\/doi.org\/10.1093\/bioinformatics\/btx717.","journal-title":"Bioinformatics"},{"issue":"3","key":"3779_CR11","doi-asserted-by":"publisher","first-page":"0150719","DOI":"10.1371\/journal.pone.0150719","volume":"11","author":"C Chu","year":"2016","unstructured":"Chu C, Nielsen R, Wu Y. REPdenovo: inferring de novo repeat motifs from short sequence reads. PLoS ONE. 2016;11(3):0150719. https:\/\/doi.org\/10.1371\/journal.pone.0150719.","journal-title":"PLoS ONE"},{"issue":"6","key":"3779_CR12","doi-asserted-by":"publisher","first-page":"825","DOI":"10.1093\/bioinformatics\/btu762","volume":"31","author":"J Luo","year":"2014","unstructured":"Luo J, Wang J, Zhang Z, Wu F-X, Li M, Pan Y. EPGA: de novo assembly using the distributions of reads and insert size. Bioinformatics. 2014;31(6):825\u201333. https:\/\/doi.org\/10.1093\/bioinformatics\/btu762.","journal-title":"Bioinformatics"},{"issue":"24","key":"3779_CR13","doi-asserted-by":"publisher","first-page":"3988","DOI":"10.1093\/bioinformatics\/btv487","volume":"31","author":"J Luo","year":"2015","unstructured":"Luo J, Wang J, Li W, Zhang Z, Wu F-X, Li M, Pan Y. EPGA2: memory-efficient de novo assembler. Bioinformatics. 2015;31(24):3988\u201390. https:\/\/doi.org\/10.1093\/bioinformatics\/btv487.","journal-title":"Bioinformatics"},{"issue":"5","key":"3779_CR14","doi-asserted-by":"publisher","first-page":"455","DOI":"10.1089\/cmb.2012.0021","volume":"19","author":"A Bankevich","year":"2012","unstructured":"Bankevich A, Nurk S, Antipov D, Gurevich AA, Dvorkin M, Kulikov AS, Lesin VM, Nikolenko SI, Pham S, Prjibelski AD, et al. SPAdes: a new genome assembly algorithm and its applications to single-cell sequencing. J Comput Biol. 2012;19(5):455\u201377. https:\/\/doi.org\/10.1089\/cmb.2012.0021.","journal-title":"J Comput Biol"},{"key":"3779_CR15","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1007\/s40484-019-0166-9","volume":"7","author":"X Liao","year":"2019","unstructured":"Liao X, Li M, Zou Y, Wu F, Pan Y, Wang J. Current challenges and solutions of de novo assembly. Quant Biol. 2019;7:90\u2013109. https:\/\/doi.org\/10.1007\/s40484-019-0166-9.","journal-title":"Quant Biol"},{"key":"3779_CR16","doi-asserted-by":"publisher","unstructured":"Liao X, Zhang X, Wu F, Wang J. de novo repeat detection based on the third generation sequencing reads. In: 2019 IEEE international conference on bioinformatics and biomedicine (2019BIBM). https:\/\/doi.org\/10.1109\/BIBM47256.2019.8982959.","DOI":"10.1109\/BIBM47256.2019.8982959"},{"issue":"4","key":"3779_CR17","doi-asserted-by":"publisher","first-page":"916","DOI":"10.1109\/TCBB.2016.2550433","volume":"14","author":"M Li","year":"2017","unstructured":"Li M, Liao Z, He Y, Wang J, Luo J, Pan Y. ISEA: iterative seed-extension algorithm for de novo assembly using paired-end information and insert size distribution. IEEE\/ACM Trans Comput Biol Bioinform: TCBB. 2017;14(4):916\u201325. https:\/\/doi.org\/10.1109\/TCBB.2016.2550433.","journal-title":"IEEE\/ACM Trans Comput Biol Bioinform: TCBB"},{"issue":"21","key":"3779_CR18","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1158\/0008-5472.CAN-17-0337","volume":"77","author":"JT Robinson","year":"2017","unstructured":"Robinson JT, Thorvaldsd\u00f3ttir H, Wenger AM, Zehir A, Mesirov JP. Variant review with the integrative genomics viewer. Cancer Res. 2017;77(21):31\u20134. https:\/\/doi.org\/10.1158\/0008-5472.CAN-17-0337.","journal-title":"Cancer Res"},{"key":"3779_CR19","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2019.2945761","author":"X Liao","year":"2019","unstructured":"Liao X, Li M, Zou Y, Wu F, Pan Y, Luo F, Wang J. EPGA-SC: a framework for de novo assembly of single-cell sequencing reads. IEEE\/ACM Trans Comput Biol Bioinform. 2019;. https:\/\/doi.org\/10.1109\/TCBB.2019.2945761.","journal-title":"IEEE\/ACM Trans Comput Biol Bioinform"},{"issue":"4","key":"3779_CR20","doi-asserted-by":"publisher","first-page":"357","DOI":"10.1038\/nmeth.1923","volume":"9","author":"B Langmead","year":"2012","unstructured":"Langmead B, Salzberg SL. Fast gapped-read alignment with Bowtie 2. Nat Methods. 2012;9(4):357. https:\/\/doi.org\/10.1038\/nmeth.1923.","journal-title":"Nat Methods"},{"issue":"1","key":"3779_CR21","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1002\/0471250953.bi0410s05","volume":"5","author":"N Chen","year":"2004","unstructured":"Chen N. Using RepeatMasker to identify repetitive elements in genomic sequences. Curr Protoc Bioinform. 2004;5(1):4\u201310. https:\/\/doi.org\/10.1002\/0471250953.bi0410s05.","journal-title":"Curr Protoc Bioinform"},{"issue":"1\u20134","key":"3779_CR22","doi-asserted-by":"publisher","first-page":"462","DOI":"10.1159\/000084979","volume":"110","author":"J Jurka","year":"2005","unstructured":"Jurka J, Kapitonov VV, Pavlicek A, Klonowski P, Kohany O, Walichiewicz J. Repbase update, a database of eukaryotic repetitive elements. Cytogenet Genome Res. 2005;110(1\u20134):462\u20137. https:\/\/doi.org\/10.1159\/000084979.","journal-title":"Cytogenet Genome Res"},{"issue":"3","key":"3779_CR23","doi-asserted-by":"publisher","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","volume":"215","author":"SF Altschul","year":"1990","unstructured":"Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990;215(3):403\u201310. https:\/\/doi.org\/10.1016\/S0022-2836(05)80360-2.","journal-title":"J Mol Biol"},{"issue":"6","key":"3779_CR24","doi-asserted-by":"publisher","first-page":"764","DOI":"10.1093\/bioinformatics\/btr011","volume":"27","author":"G Mar\u00e7ais","year":"2011","unstructured":"Mar\u00e7ais G, Kingsford C. A fast, lock-free approach for efficient parallel counting of occurrences of k-mers. Bioinformatics. 2011;27(6):764\u201370. https:\/\/doi.org\/10.1093\/bioinformatics\/btr011.","journal-title":"Bioinformatics"},{"issue":"10","key":"3779_CR25","doi-asserted-by":"publisher","first-page":"1569","DOI":"10.1093\/bioinformatics\/btv022","volume":"31","author":"S Deorowicz","year":"2015","unstructured":"Deorowicz S, Kokot M, Grabowski S, Debudaj-Grabysz A. KMC 2: fast and resource-frugal k-mer counting. Bioinformatics. 2015;31(10):1569\u201376. https:\/\/doi.org\/10.1093\/bioinformatics\/btv022.","journal-title":"Bioinformatics"},{"issue":"8","key":"3779_CR26","doi-asserted-by":"publisher","first-page":"1916","DOI":"10.1101\/gr.1251803","volume":"13","author":"X Li","year":"2003","unstructured":"Li X, Waterman MS. Estimating the repeat structure and length of DNA sequences using l-tuples. Genome Res. 2003;13(8):1916\u201322. https:\/\/doi.org\/10.1101\/gr.1251803.","journal-title":"Genome Res"},{"issue":"11","key":"3779_CR27","doi-asserted-by":"publisher","first-page":"116","DOI":"10.1186\/gb-2010-11-11-r116","volume":"11","author":"DR Kelley","year":"2010","unstructured":"Kelley DR, Schatz MC, Salzberg SL. Quake: quality-aware detection and correction of sequencing errors. Genome Biol. 2010;11(11):116. https:\/\/doi.org\/10.1186\/gb-2010-11-11-r116.","journal-title":"Genome Biol"},{"key":"3779_CR28","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2018.2861380","author":"X Liao","year":"2018","unstructured":"Liao X, Li M, Zou Y, Wu F, Pan Y, Luo F, Wang J, et al. Improving de novo assembly based on read classification. IEEE\/ACM Trans Comput Biol Bioinform. 2018;. https:\/\/doi.org\/10.1109\/TCBB.2018.2861380.","journal-title":"IEEE\/ACM Trans Comput Biol Bioinform"},{"key":"3779_CR29","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2019.2897558","author":"X Liao","year":"2019","unstructured":"Liao X, Li M, Zou Y, Wu F, Pan Y, Wang J. An efficient trimming algorithm based on multi-feature fusion scoring model for NGS data. IEEE\/ACM Trans Comput Biol Bioinform. 2019;. https:\/\/doi.org\/10.1109\/TCBB.2019.2897558.","journal-title":"IEEE\/ACM Trans Comput Biol Bioinform"},{"key":"3779_CR30","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2018.2876855","author":"B Wu","year":"2018","unstructured":"Wu B, Li M, Liao X, Luo J, Wu F, Pan Y, Wang J. MEC: misassembly error correction in contigs based on distribution of paired-end reads and statistics of gc-contents. IEEE\/ACM Trans Comput Biol Bioinform. 2018;. https:\/\/doi.org\/10.1109\/TCBB.2018.2876855.","journal-title":"IEEE\/ACM Trans Comput Biol Bioinform"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-020-03779-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-020-03779-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-020-03779-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T23:50:11Z","timestamp":1634601011000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-020-03779-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,19]]},"references-count":30,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["3779"],"URL":"https:\/\/doi.org\/10.1186\/s12859-020-03779-w","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,10,19]]},"assertion":[{"value":"10 December 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"24 September 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 October 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Compliance with ethical standards"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"463"}}