{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T05:30:25Z","timestamp":1786512625570,"version":"build-2736575974"},"reference-count":60,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T00:00:00Z","timestamp":1780531200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T00:00:00Z","timestamp":1786492800000},"content-version":"vor","delay-in-days":69,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004281","name":"Narodowe Centrum Nauki","doi-asserted-by":"publisher","award":["2021\/41\/N\/ST6\/01919"],"award-info":[{"award-number":["2021\/41\/N\/ST6\/01919"]}],"id":[{"id":"10.13039\/501100004281","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:sec>\n                    <jats:title>Background<\/jats:title>\n                    <jats:p>Methods for analyzing protein sequence similarity focus mainly on identifying the homologies of proteins with standard amino acid compositions. For motifs with biased compositions, they have already been shown to be suboptimal; thus, we lack dedicated tools to support their analyses. However, motifs with compositional biases also play key roles in protein functions. These domains can be found in transmembrane proteins, bind to RNA through RGG boxes, and may form prions. Nevertheless, many domains remain unknown, as for a long time they were considered nonfunctional and most of the methods mask them to improve homology searches. Therefore, we need better solutions to infer their functions more efficiently.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Results<\/jats:title>\n                    <jats:p>In this research, we developed a new method incorporating three alignment strategies, an algorithm for identifying motifs with similar compositions, 2-mer based filtering, and a new metric for evaluating alignments. These solutions focus mainly on comparing the physicochemical properties of protein sequences rather than their evolutionary relationships. To validate our approach, we compared BLAST with our method in three variants that use local, global\u2013local, and global alignment with the algorithm for identifying compositionally similarities. We used these methods to search for similar transmembrane domains and RGG boxes. We observed that our solutions significantly increased the number of true positives. The greatest increase occurred after we applied our similarity score measure. Compositionally biased motifs frequently consist of two adjacent functionally important motifs; therefore, we also searched for similarities to the K-DE motif of DNA\u2013directed RNA polymerase subunit delta. We found that compared with the other alignment strategies, global\u2013local and global alignment with identifying similar regions included all the submotifs of the query sequence more often.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:title>Conclusion<\/jats:title>\n                    <jats:p>Our method introduces novel strategies that enhance the search for compositionally biased motifs, thereby improving annotation retrieval via sequence matching.<\/jats:p>\n                  <\/jats:sec>","DOI":"10.1186\/s12859-026-06509-w","type":"journal-article","created":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T13:04:02Z","timestamp":1780578242000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["CB-Search: a method for searching for similar protein motifs with biased compositions"],"prefix":"10.1186","volume":"27","author":[{"given":"Patryk","family":"Jarnot","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,6,4]]},"reference":[{"issue":"1","key":"6509_CR1","doi-asserted-by":"publisher","first-page":"676","DOI":"10.1186\/s12864-023-09762-y","volume":"24","author":"L Yang","year":"2023","unstructured":"Yang L, Chen Y, Liu X, Zhang S, Han Q. Genome-wide identification and expression analysis of xyloglucan endotransglucosylase\/hydrolase genes family in Salicaceae during grafting. BMC Genomics. 2023;24(1):676.","journal-title":"BMC Genomics"},{"issue":"1","key":"6509_CR2","doi-asserted-by":"publisher","first-page":"7","DOI":"10.1186\/s12862-023-02107-z","volume":"23","author":"T Litman","year":"2023","unstructured":"Litman T, Stein WD. Ancient lineages of the keratin-associated protein (KRTAP) genes and their co-option in the evolution of the hair follicle. BMC Ecol Evol. 2023;23(1):7.","journal-title":"BMC Ecol Evol"},{"issue":"1","key":"6509_CR3","doi-asserted-by":"publisher","first-page":"323","DOI":"10.1016\/j.cell.2025.10.019","volume":"189","author":"KM Ruff","year":"2026","unstructured":"Ruff KM, King MR, Ying AW, Liu V, Pant A, Lieberman WE, et al. Molecular grammars of predicted intrinsically disordered regions that span the human proteome. Cell. 2026;189(1):323\u201342.","journal-title":"Cell"},{"issue":"7","key":"6509_CR4","doi-asserted-by":"publisher","first-page":"902","DOI":"10.1093\/bioinformatics\/bti070","volume":"21","author":"YK Yu","year":"2005","unstructured":"Yu YK, Altschul SF. The construction of amino acid substitution matrices for the comparison of proteins with non-standard compositions. Bioinformatics. 2005;21(7):902\u201311.","journal-title":"Bioinformatics"},{"issue":"17","key":"6509_CR5","doi-asserted-by":"publisher","first-page":"3389","DOI":"10.1093\/nar\/25.17.3389","volume":"25","author":"SF Altschul","year":"1997","unstructured":"Altschul SF, Madden TL, Sch\u00e4ffer AA, Zhang J, Zhang Z, Miller W, et al. Gapped BLAST and PSI-BLAST: a new generation of protein database search programs. Nucleic Acids Res. 1997;25(17):3389\u2013402.","journal-title":"Nucleic Acids Res"},{"issue":"10","key":"6509_CR6","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1002195","volume":"7","author":"SR Eddy","year":"2011","unstructured":"Eddy SR. Accelerated profile HMM searches. PLoS Comput Biol. 2011;7(10):e1002195.","journal-title":"PLoS Comput Biol"},{"issue":"5","key":"6509_CR7","doi-asserted-by":"publisher","first-page":"bbac299","DOI":"10.1093\/bib\/bbac299","volume":"23","author":"P Jarnot","year":"2022","unstructured":"Jarnot P, Ziemska-Legiecka J, Grynberg M, Gruca A. Insights from analyses of low complexity regions with canonical methods for protein sequence comparison. Brief Bioinform. 2022;23(5):bbac299.","journal-title":"Brief Bioinform"},{"issue":"1","key":"6509_CR8","doi-asserted-by":"publisher","first-page":"3","DOI":"10.1002\/0471250953.bi0301s42","volume":"42","author":"WR Pearson","year":"2013","unstructured":"Pearson WR. An introduction to sequence similarity (\u201chomology\u2019\u2019) searching. Curr Protoc Bioinformatics. 2013;42(1):3\u20131.","journal-title":"Curr Protoc Bioinformatics"},{"issue":"2","key":"6509_CR9","doi-asserted-by":"publisher","first-page":"149","DOI":"10.1016\/0097-8485(93)85006-X","volume":"17","author":"JC Wootton","year":"1993","unstructured":"Wootton JC, Federhen S. Statistics of local complexity in amino acid sequences and sequence databases. Comput Chem. 1993;17(2):149\u201363.","journal-title":"Comput Chem"},{"issue":"10","key":"6509_CR10","doi-asserted-by":"publisher","first-page":"915","DOI":"10.1093\/bioinformatics\/16.10.915","volume":"16","author":"VJ Promponas","year":"2000","unstructured":"Promponas VJ, Enright AJ, Tsoka S, Kreil DP, Leroy C, Hamodrakas S, et al. CAST: an iterative algorithm for the complexity analysis of sequence tracts. Bioinformatics. 2000;16(10):915\u201322.","journal-title":"Bioinformatics"},{"issue":"1","key":"6509_CR11","doi-asserted-by":"publisher","first-page":"476","DOI":"10.1186\/s12859-017-1906-3","volume":"18","author":"PM Harrison","year":"2017","unstructured":"Harrison PM. fLPS: Fast discovery of compositional biases for the protein universe. BMC Bioinformatics. 2017;18(1):476.","journal-title":"BMC Bioinformatics"},{"key":"6509_CR12","doi-asserted-by":"publisher","first-page":"e12363","DOI":"10.7717\/peerj.12363","volume":"9","author":"PM Harrison","year":"2021","unstructured":"Harrison PM. fLPS 2.0: rapid annotation of compositionally-biased regions in biological sequences. PeerJ. 2021;9:e12363.","journal-title":"PeerJ"},{"issue":"2","key":"6509_CR13","doi-asserted-by":"publisher","first-page":"lqab048","DOI":"10.1093\/nargab\/lqab048","volume":"3","author":"SM Cascarina","year":"2021","unstructured":"Cascarina SM, King DC, Osborne Nishimura E, Ross ED. LCD-Composer: an intuitive, composition-centric method enabling the identification and detailed functional mapping of low-complexity domains. NAR Genomics Bioinform. 2021;3(2):lqab048.","journal-title":"NAR Genomics Bioinform"},{"issue":"5","key":"6509_CR14","doi-asserted-by":"publisher","first-page":"672","DOI":"10.1093\/bioinformatics\/18.5.672","volume":"18","author":"MM Alb\u00e0","year":"2002","unstructured":"Alb\u00e0 MM, Laskowski RA, Hancock JM. Detecting cryptically simple protein sequences using the SIMPLE algorithm. Bioinformatics. 2002;18(5):672\u20138.","journal-title":"Bioinformatics"},{"issue":"2","key":"6509_CR15","doi-asserted-by":"publisher","first-page":"160","DOI":"10.1093\/bioinformatics\/bth497","volume":"21","author":"SW Shin","year":"2005","unstructured":"Shin SW, Kim SM. A new algorithm for detecting low-complexity regions in protein sequences. Bioinformatics. 2005;21(2):160\u201370.","journal-title":"Bioinformatics"},{"issue":"1","key":"6509_CR16","doi-asserted-by":"publisher","first-page":"680","DOI":"10.1038\/s41598-023-50991-8","volume":"14","author":"PM Harrison","year":"2024","unstructured":"Harrison PM. Optimizing strategy for the discovery of compositionally-biased or low-complexity regions in proteins. Sci Rep. 2024;14(1):680.","journal-title":"Sci Rep"},{"issue":"2","key":"6509_CR17","doi-asserted-by":"publisher","first-page":"458","DOI":"10.1093\/bib\/bbz007","volume":"21","author":"P Mier","year":"2020","unstructured":"Mier P, Paladin L, Tamana S, Petrosian S, Hajdu-Solt\u00e9sz B, Urbanek A, et al. Disentangling the complexity of low complexity proteins. Brief Bioinform. 2020;21(2):458\u201372.","journal-title":"Brief Bioinform"},{"issue":"20","key":"6509_CR18","doi-asserted-by":"publisher","first-page":"2632","DOI":"10.1093\/bioinformatics\/btp482","volume":"25","author":"J Jorda","year":"2009","unstructured":"Jorda J, Kajava AV. T-REKS: identification of Tandem REpeats in sequences with a K-meanS based algorithm. Bioinformatics. 2009;25(20):2632\u20138.","journal-title":"Bioinformatics"},{"issue":"1","key":"6509_CR19","doi-asserted-by":"publisher","first-page":"382","DOI":"10.1186\/1471-2105-8-382","volume":"8","author":"AM Newman","year":"2007","unstructured":"Newman AM, Cooper JB. XSTREAM: a practical algorithm for identification and architecture modeling of tandem repeats in protein sequences. BMC Bioinform. 2007;8(1):382.","journal-title":"BMC Bioinform"},{"issue":"1","key":"6509_CR20","doi-asserted-by":"publisher","first-page":"883","DOI":"10.1186\/s12864-025-12132-5","volume":"26","author":"E Schumbera","year":"2025","unstructured":"Schumbera E, Dormann D, Walther A, Andrade-Navarro MA. Computational investigation of the sequence context of arginine\/glycine-rich motifs in the human proteome. BMC Genomics. 2025;26(1):883.","journal-title":"BMC Genomics"},{"issue":"2","key":"6509_CR21","doi-asserted-by":"publisher","first-page":"113","DOI":"10.1038\/nsmb895","volume":"12","author":"MD Edwards","year":"2005","unstructured":"Edwards MD, Li Y, Kim S, Miller S, Bartlett W, Black S, et al. Pivotal role of the glycine-rich TM3 helix in gating the MscS mechanosensitive channel. Nature Struct Mol Biol. 2005;12(2):113\u20139.","journal-title":"Nature Struct Mol Biol"},{"issue":"18","key":"6509_CR22","doi-asserted-by":"publisher","first-page":"7128","DOI":"10.1074\/jbc.TM118.001190","volume":"294","author":"TM Franzmann","year":"2019","unstructured":"Franzmann TM, Alberti S. Prion-like low-complexity sequences: key regulators of protein solubility and phase behavior. J Biol Chem. 2019;294(18):7128\u201336.","journal-title":"J Biol Chem"},{"issue":"10","key":"6509_CR23","doi-asserted-by":"publisher","first-page":"1486","DOI":"10.3390\/biom12101486","volume":"12","author":"K Kastano","year":"2022","unstructured":"Kastano K, Mier P, Doszt\u00e1nyi Z, Promponas VJ, Andrade-Navarro MA. Functional tuning of intrinsically disordered regions in human proteins by composition bias. Biomolecules. 2022;12(10):1486.","journal-title":"Biomolecules"},{"issue":"6","key":"6509_CR24","doi-asserted-by":"publisher","first-page":"2399","DOI":"10.1093\/nar\/gkr1078","volume":"40","author":"WH Lin","year":"2012","unstructured":"Lin WH, Kussell E. Evolutionary pressures on simple sequence repeats in prokaryotic coding regions. Nucleic Acids Res. 2012;40(6):2399\u2013413.","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"6509_CR25","doi-asserted-by":"publisher","first-page":"155","DOI":"10.1186\/1471-2148-12-155","volume":"12","author":"N Rad\u00f3-Trilla","year":"2012","unstructured":"Rad\u00f3-Trilla N, Alb\u00e0 M. Dissecting the role of low-complexity regions in the evolution of vertebrate proteins. BMC Evol Biol. 2012;12(1):155.","journal-title":"BMC Evol Biol"},{"issue":"5","key":"6509_CR26","doi-asserted-by":"publisher","first-page":"919","DOI":"10.1016\/j.ajhg.2016.04.001","volume":"98","author":"T Willems","year":"2016","unstructured":"Willems T, Gymrek M, Poznik GD, Tyler-Smith C, Erlich Y. Population-scale sequencing data enable precise estimates of Y-STR mutation rates. Am J Hum Genet. 2016;98(5):919\u201333.","journal-title":"Am J Hum Genet"},{"key":"6509_CR27","doi-asserted-by":"publisher","first-page":"19","DOI":"10.1016\/j.gene.2006.03.023","volume":"378","author":"MA DePristo","year":"2006","unstructured":"DePristo MA, Zilversmit MM, Hartl DL. On the abundance, amino acid composition, and evolutionary dynamics of low-complexity regions in proteins. Gene. 2006;378:19\u201330.","journal-title":"Gene"},{"issue":"34","key":"6509_CR28","doi-asserted-by":"publisher","first-page":"32313","DOI":"10.1074\/jbc.M304709200","volume":"278","author":"M Rasmussen","year":"2003","unstructured":"Rasmussen M, Jacobsson M, Bj\u00f6rck L. Genome-based identification and analysis of collagen-related structural motifs in bacterial and viral proteins. J Biol Chem. 2003;278(34):32313\u20136.","journal-title":"J Biol Chem"},{"issue":"6","key":"6509_CR29","doi-asserted-by":"publisher","first-page":"453","DOI":"10.4161\/viru.25180","volume":"4","author":"AC Doxey","year":"2013","unstructured":"Doxey AC, McConkey BJ. Prediction of molecular mimicry candidates in human pathogenic bacteria. Virulence. 2013;4(6):453\u201366.","journal-title":"Virulence"},{"issue":"12","key":"6509_CR30","doi-asserted-by":"publisher","first-page":"e121","DOI":"10.1093\/nar\/gkt263","volume":"41","author":"J Mistry","year":"2013","unstructured":"Mistry J, Finn RD, Eddy SR, Bateman A, Punta M. Challenges in homology search: HMMER3 and convergent evolution of coiled-coil regions. Nucleic Acids Res. 2013;41(12):e121\u2013e121.","journal-title":"Nucleic Acids Res"},{"issue":"4","key":"6509_CR31","doi-asserted-by":"publisher","first-page":"1727","DOI":"10.3390\/ijms22041727","volume":"22","author":"K Kastano","year":"2021","unstructured":"Kastano K, Mier P, Andrade-Navarro MA. The role of low complexity regions in protein interaction modes: an illustration in huntingtin. Int J Mol Sci. 2021;22(4):1727.","journal-title":"Int J Mol Sci"},{"issue":"3","key":"6509_CR32","doi-asserted-by":"publisher","first-page":"443","DOI":"10.1016\/0022-2836(70)90057-4","volume":"48","author":"SB Needleman","year":"1970","unstructured":"Needleman SB, Wunsch CD. A general method applicable to the search for similarities in the amino acid sequence of two proteins. J Mol Biol. 1970;48(3):443\u201353.","journal-title":"J Mol Biol"},{"issue":"1","key":"6509_CR33","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1016\/0022-2836(81)90087-5","volume":"147","author":"TF Smith","year":"1981","unstructured":"Smith TF, Waterman MS, et al. Identification of common molecular subsequences. J Mol Biol. 1981;147(1):195\u20137.","journal-title":"J Mol Biol"},{"issue":"3","key":"6509_CR34","doi-asserted-by":"publisher","first-page":"403","DOI":"10.1016\/S0022-2836(05)80360-2","volume":"215","author":"SF Altschul","year":"1990","unstructured":"Altschul SF, Gish W, Miller W, Myers EW, Lipman DJ. Basic local alignment search tool. J Mol Biol. 1990;215(3):403\u201310.","journal-title":"J Mol Biol"},{"issue":"23","key":"6509_CR35","doi-asserted-by":"publisher","first-page":"3150","DOI":"10.1093\/bioinformatics\/bts565","volume":"28","author":"L Fu","year":"2012","unstructured":"Fu L, Niu B, Zhu Z, Wu S, Li W. CD-HIT: accelerated for clustering the next-generation sequencing data. Bioinformatics. 2012;28(23):3150\u20132.","journal-title":"Bioinformatics"},{"issue":"2","key":"6509_CR36","doi-asserted-by":"publisher","first-page":"173","DOI":"10.1038\/nmeth.1818","volume":"9","author":"M Remmert","year":"2012","unstructured":"Remmert M, Biegert A, Hauser A, S\u00f6ding J. HHblits: lightning-fast iterative protein sequence searching by HMM-HMM alignment. Nat Methods. 2012;9(2):173\u20135.","journal-title":"Nat Methods"},{"issue":"11","key":"6509_CR37","doi-asserted-by":"publisher","first-page":"1026","DOI":"10.1038\/nbt.3988","volume":"35","author":"M Steinegger","year":"2017","unstructured":"Steinegger M, S\u00f6ding J. MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nat Biotechnol. 2017;35(11):1026\u20138.","journal-title":"Nat Biotechnol"},{"issue":"W1","key":"6509_CR38","doi-asserted-by":"publisher","first-page":"W297","DOI":"10.1093\/nar\/gkab408","volume":"49","author":"G Erd\u0151s","year":"2021","unstructured":"Erd\u0151s G, Pajkos M, Doszt\u00e1nyi Z. IUPred3: prediction of protein disorder enhanced with unambiguous experimental annotation and visualization of evolutionary conservation. Nucleic Acids Res. 2021;49(W1):W297\u2013303.","journal-title":"Nucleic Acids Res"},{"issue":"5","key":"6509_CR39","doi-asserted-by":"publisher","first-page":"btaf297","DOI":"10.1093\/bioinformatics\/btaf297","volume":"41","author":"M Mehdiabadi","year":"2025","unstructured":"Mehdiabadi M, Blum M, Tesei G, Von\u00a0B\u00fclow S, Lindorff-Larsen K, Tosatto SC, et al. MobiDB-lite 4.0: faster prediction of intrinsic protein disorder and structural compactness. Bioinformatics. 2025;41(5):btaf297.","journal-title":"Bioinformatics"},{"issue":"5","key":"6509_CR40","doi-asserted-by":"publisher","first-page":"1027","DOI":"10.1016\/j.jmb.2004.03.016","volume":"338","author":"L K\u00e4ll","year":"2004","unstructured":"K\u00e4ll L, Krogh A, Sonnhammer EL. A combined transmembrane topology and signal peptide prediction method. J Mol Biol. 2004;338(5):1027\u201336.","journal-title":"J Mol Biol"},{"issue":"9","key":"6509_CR41","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1013470","volume":"21","author":"A Lucas","year":"2025","unstructured":"Lucas A, Sch\u00e4ffer DE, Wickramasinghe J, Auslander N. k Mermaid: Ultrafast metagenomic read assignment to protein clusters by hashing of amino acid k-mer frequencies. PLoS Comput Biol. 2025;21(9):e1013470.","journal-title":"PLoS Comput Biol"},{"key":"6509_CR42","doi-asserted-by":"publisher","first-page":"giad101","DOI":"10.1093\/gigascience\/giad101","volume":"12","author":"JM Silva","year":"2023","unstructured":"Silva JM, Qi W, Pinho AJ, Pratas D. AlcoR: alignment-free simulation, mapping, and visualization of low-complexity regions in biological data. GigaScience. 2023;12:giad101.","journal-title":"GigaScience"},{"issue":"1","key":"6509_CR43","doi-asserted-by":"publisher","first-page":"31777","DOI":"10.1038\/s41598-024-82548-8","volume":"14","author":"P Jarnot","year":"2024","unstructured":"Jarnot P. Challenges in adjusting scoring matrices when comparing functional motifs with non-standard compositions. Sci Rep. 2024;14(1):31777.","journal-title":"Sci Rep"},{"key":"6509_CR44","unstructured":"UniProt: the Universal protein knowledgebase in 2025. Nucleic Acids Res. 2025;53(D1):D609\u2013D617."},{"issue":"1","key":"6509_CR45","doi-asserted-by":"publisher","first-page":"1251","DOI":"10.1186\/s12864-024-10960-5","volume":"25","author":"J Ziemska-Legiecka","year":"2024","unstructured":"Ziemska-Legiecka J, Jarnot P, Szyma\u0144ska S, B\u0142aszczyk D, Sta\u015bczak A, Langer-Macio\u0142 H, et al. LCRAnnotationsDB: a database of low complexity regions functional and structural annotations. BMC Genomics. 2024;25(1):1251.","journal-title":"BMC Genomics"},{"issue":"5","key":"6509_CR46","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1003063","volume":"9","author":"AM Schnoes","year":"2013","unstructured":"Schnoes AM, Ream DC, Thorman AW, Babbitt PC, Friedberg I. Biases in the experimental annotations of protein function and their effect on our understanding of protein function space. PLoS Comput Biol. 2013;9(5):e1003063.","journal-title":"PLoS Comput Biol"},{"issue":"12","key":"6509_CR47","doi-asserted-by":"publisher","first-page":"1539","DOI":"10.1002\/prot.26617","volume":"91","author":"A Kryshtafovych","year":"2023","unstructured":"Kryshtafovych A, Schwede T, Topf M, Fidelis K, Moult J. Critical assessment of methods of protein structure prediction (CASP)\u2014Round XV. Proteins: Struct, Funct, Bioinf. 2023;91(12):1539\u201349.","journal-title":"Proteins: Struct, Funct, Bioinf"},{"issue":"1","key":"6509_CR48","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/s13059-019-1835-8","volume":"20","author":"N Zhou","year":"2019","unstructured":"Zhou N, Jiang Y, Bergquist TR, Lee AJ, Kacsoh BZ, Crocker AW, et al. The CAFA challenge reports improved protein function prediction and new functional annotations for hundreds of genes through experimental screens. Genome Biol. 2019;20(1):1\u201323.","journal-title":"Genome Biol"},{"issue":"D1","key":"6509_CR49","doi-asserted-by":"publisher","first-page":"D444","DOI":"10.1093\/nar\/gkae1082","volume":"53","author":"M Blum","year":"2025","unstructured":"Blum M, Andreeva A, Florentino LC, Chuguransky SR, Grego T, Hobbs E, et al. InterPro: the protein sequence classification resource in 2025. Nucleic Acids Res. 2025;53(D1):D444\u201356.","journal-title":"Nucleic Acids Res"},{"issue":"42","key":"6509_CR50","doi-asserted-by":"publisher","first-page":"16817","DOI":"10.1021\/jacs.9b07837","volume":"141","author":"V Kub\u00e1\u0148","year":"2019","unstructured":"Kub\u00e1\u0148 V, Srb P, \u0160t\u00e9gnerov\u00e1 H, Padrta P, Zachrdla M, Jase\u0148\u00e1kov\u00e1 Z, et al. Quantitative conformational analysis of functionally important electrostatic interactions in the intrinsically disordered region of delta subunit of bacterial RNA polymerase. J Am Chem Soc. 2019;141(42):16817\u201328.","journal-title":"J Am Chem Soc"},{"issue":"14","key":"6509_CR51","doi-asserted-by":"publisher","first-page":"4678","DOI":"10.1093\/nar\/gkm414","volume":"35","author":"MG Kann","year":"2007","unstructured":"Kann MG, Sheetlin SL, Park Y, Bryant SH, Spouge JL. The identification of complete domains within protein sequences using accurate E-values for semi-global alignment. Nucleic Acids Res. 2007;35(14):4678\u201385.","journal-title":"Nucleic Acids Res"},{"issue":"1","key":"6509_CR52","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1186\/1752-0509-4-43","volume":"4","author":"A Coletta","year":"2010","unstructured":"Coletta A, Pinney JW, Sol\u00eds DYW, Marsh J, Pettifer SR, Attwood TK. Low-complexity regions within protein sequences have position-dependent roles. BMC Syst Biol. 2010;4(1):43.","journal-title":"BMC Syst Biol"},{"issue":"6","key":"6509_CR53","doi-asserted-by":"publisher","first-page":"2264","DOI":"10.1073\/pnas.87.6.2264","volume":"87","author":"S Karlin","year":"1990","unstructured":"Karlin S, Altschul SF. Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes. Proc Natl Acad Sci. 1990;87(6):2264\u20138.","journal-title":"Proc Natl Acad Sci"},{"key":"6509_CR54","doi-asserted-by":"publisher","first-page":"1764","DOI":"10.1039\/D5TB01841B","volume":"14","author":"F Seco","year":"2026","unstructured":"Seco F, Amador S, Conzuelo F, Baptista AC, Morgado L, Pina AS. Molecular versatility of polyproline II helices: from natural proteins to biomimetic materials. J Mater Chem B. 2026;14:1764\u201382.","journal-title":"J Mater Chem B"},{"issue":"11","key":"6509_CR55","doi-asserted-by":"publisher","first-page":"2150","DOI":"10.1002\/pro.3954","volume":"29","author":"R Trivedi","year":"2020","unstructured":"Trivedi R, Nagarajaram HA. Substitution scoring matrices for proteins-An overview. Protein Sci. 2020;29(11):2150\u201363.","journal-title":"Protein Sci"},{"issue":"3","key":"6509_CR56","doi-asserted-by":"publisher","first-page":"521","DOI":"10.1006\/jmbi.2000.3684","volume":"298","author":"MA Andrade","year":"2000","unstructured":"Andrade MA, Ponting CP, Gibson TJ, Bork P. Homology-based method for identification of protein repeats using statistical significance estimates. J Mol Biol. 2000;298(3):521\u201337.","journal-title":"J Mol Biol"},{"issue":"7","key":"6509_CR57","doi-asserted-by":"publisher","first-page":"1575","DOI":"10.1093\/nar\/30.7.1575","volume":"30","author":"AJ Enright","year":"2002","unstructured":"Enright AJ, Van Dongen S, Ouzounis CA. An efficient algorithm for large-scale detection of protein families. Nucleic Acids Res. 2002;30(7):1575\u201384.","journal-title":"Nucleic Acids Res"},{"issue":"18","key":"6509_CR58","doi-asserted-by":"publisher","first-page":"e121","DOI":"10.1093\/nar\/gkv585","volume":"43","author":"Z Peng","year":"2015","unstructured":"Peng Z, Kurgan L. High-throughput prediction of RNA, DNA and protein binding regions mediated by intrinsic disorder. Nucleic Acids Res. 2015;43(18):e121\u2013e121.","journal-title":"Nucleic Acids Res"},{"key":"6509_CR59","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1186\/1752-0509-4-S1-S3","volume":"4","author":"L Wang","year":"2010","unstructured":"Wang L, Huang C, Yang MQ, Yang JY. BindN+ for accurate prediction of DNA and RNA-binding residues from protein sequence features. BMC Syst Biol. 2010;4:1\u20139.","journal-title":"BMC Syst Biol"},{"issue":"2","key":"6509_CR60","doi-asserted-by":"publisher","first-page":"349","DOI":"10.1016\/j.gpb.2023.04.001","volume":"21","author":"S Wang","year":"2023","unstructured":"Wang S, You R, Liu Y, Xiong Y, Zhu S. NetGO 3.0: protein language model improves large-scale functional annotations. Genomics Proteomics Bioinform. 2023;21(2):349\u201358.","journal-title":"Genomics Proteomics Bioinform"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-026-06509-w","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-026-06509-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-026-06509-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,8,12]],"date-time":"2026-08-12T05:02:55Z","timestamp":1786510975000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1186\/s12859-026-06509-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,4]]},"references-count":60,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2026,12]]}},"alternative-id":["6509"],"URL":"https:\/\/doi.org\/10.1186\/s12859-026-06509-w","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,4]]},"assertion":[{"value":"14 February 2026","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 May 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"4 June 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Not applicable.","order":1,"name":"Ethics","label":"Ethics approval and consent to participate","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"Not applicable.","order":2,"name":"Ethics","label":"Consent for publication","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no competing interests.","order":3,"name":"Ethics","label":"Competing interests","group":{"name":"EthicsHeading","label":"Declarations"}}],"article-number":"173"}}