{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T06:40:18Z","timestamp":1784356818568,"version":"3.55.0"},"reference-count":71,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2021,11,9]],"date-time":"2021-11-09T00:00:00Z","timestamp":1636416000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,11,9]],"date-time":"2021-11-09T00:00:00Z","timestamp":1636416000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"abstract":"<jats:title>Abstract<\/jats:title><jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>Amyloids are insoluble fibrillar aggregates that are highly associated with complex human diseases, such as Alzheimer\u2019s disease, Parkinson\u2019s disease, and type II diabetes. Recently, many studies reported that some specific regions of amino acid sequences may be responsible for the amyloidosis of proteins. It has become very important for elucidating the mechanism of amyloids that identifying the amyloidogenic regions. Accordingly, several computational methods have been put forward to discover amyloidogenic regions. The majority of these methods predicted amyloidogenic regions based on the physicochemical properties of amino acids. In fact, position, order, and correlation of amino acids may also influence the amyloidosis of proteins, which should be also considered in detecting amyloidogenic regions.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>To address this problem, we proposed a novel machine-learning approach for predicting amyloidogenic regions, called ReRF-Pred. Firstly, the pseudo amino acid composition (PseAAC) was exploited to characterize physicochemical properties and correlation of amino acids. Secondly, tripeptides composition (TPC) was employed to represent the order and position of amino acids. To improve the distinguishability of TPC, all possible tripeptides were analyzed by the binomial distribution method, and only those which have significantly different distribution between positive and negative samples remained. Finally, all samples were characterized by PseAAC and TPC of their amino acid sequence, and a random forest-based amyloidogenic regions predictor was trained on these samples. It was proved by validation experiments that the feature set consisted of PseAAC and TPC is the most distinguishable one for detecting amyloidosis. Meanwhile, random forest is superior to other concerned classifiers on almost all metrics. To validate the effectiveness of our model, ReRF-Pred is compared with a series of gold-standard methods on two datasets: Pep-251 and Reg33. The results suggested our method has the best overall performance and makes significant improvements in discovering amyloidogenic regions.<\/jats:p>\n              <\/jats:sec><jats:sec>\n                <jats:title>Conclusions<\/jats:title>\n                <jats:p>The advantages of our method are mainly attributed to that PseAAC and TPC can describe the differences between amyloids and other proteins successfully. The ReRF-Pred server can be accessed at http:\/\/106.12.83.135:8080\/ReRF-Pred\/.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s12859-021-04446-4","type":"journal-article","created":{"date-parts":[[2021,11,9]],"date-time":"2021-11-09T16:02:53Z","timestamp":1636473773000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":17,"title":["ReRF-Pred: predicting amyloidogenic regions of proteins based on their pseudo amino acid composition and tripeptide composition"],"prefix":"10.1186","volume":"22","author":[{"given":"Zhixia","family":"Teng","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zitong","family":"Zhang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Zhen","family":"Tian","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanjuan","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guohua","family":"Wang","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,11,9]]},"reference":[{"issue":"2","key":"4446_CR1","doi-asserted-by":"publisher","first-page":"260","DOI":"10.1016\/j.sbi.2006.03.007","volume":"16","author":"R Nelson","year":"2006","unstructured":"Nelson R, Eisenberg D. Recent atomic models of amyloid fibril structure. Curr Opin Struct Biol. 2006;16(2):260\u20135.","journal-title":"Curr Opin Struct Biol"},{"issue":"7143","key":"4446_CR2","doi-asserted-by":"publisher","first-page":"453","DOI":"10.1038\/nature05695","volume":"447","author":"MR Sawaya","year":"2007","unstructured":"Sawaya MR, Sambashivan S, Nelson R, Ivanova MI, Sievers SA, Apostol MI, Thompson MJ, Balbirnie M, Wiltzius JJW, McFarlane HT, Madsen A, Riekel C, Eisenberg D. Atomic structures of amyloid cross-beta spines reveal varied steric zippers. Nature. 2007;447(7143):453\u20137.","journal-title":"Nature"},{"key":"4446_CR3","doi-asserted-by":"crossref","unstructured":"Selkoe DJ. Alzheimer\u2019s disease: genes, proteins, and therapy. Physiol Rev. 2001;81(2):741\u201366.","DOI":"10.1152\/physrev.2001.81.2.741"},{"key":"4446_CR4","doi-asserted-by":"crossref","unstructured":"Sun Q, Kong W, Mou X, Wang S. Transcriptional regulation analysis of Alzheimer\u2019s disease based on fastnca algorithm. Curr Bioinform. 2019;14(8):771\u201382.","DOI":"10.2174\/1574893614666190919150411"},{"key":"4446_CR5","doi-asserted-by":"crossref","unstructured":"Irwin DJ, Lee VM-Y, Trojanowski JQ. Parkinson\u2019s disease dementia: convergence of -synuclein, tau and amyloid- pathologies. Nat Rev Neurosci. 2013;14(9):626\u201336.","DOI":"10.1038\/nrn3549"},{"key":"4446_CR6","doi-asserted-by":"crossref","unstructured":"Scherzinger E, Sittler A, Schweiger K, Heiser V, Lurz R, Hasenbank R, Bates GP, Lehrach H, Wanker EE. Self-assembly of polyglutamine-containing huntingtin fragments into amyloid-like fibrils: Implications for huntington\u2019s disease pathology. Proc Natl Acad Sci USA. 1999;96(8):4604\u20139.","DOI":"10.1073\/pnas.96.8.4604"},{"issue":"3","key":"4446_CR7","doi-asserted-by":"publisher","first-page":"182","DOI":"10.1016\/j.semarthrit.2007.03.005","volume":"37","author":"Y Berkun","year":"2007","unstructured":"Berkun Y, Padeh S, Reichman B, Zaks N, Rabinovich E, Lidar M, Shainberg B, Livneh A. A single testing of serum amyloid a levels as a tool for diagnosis and treatment dilemmas in familial mediterranean fever. Semin Arthritis Rheum. 2007;37(3):182\u20138.","journal-title":"Semin Arthritis Rheum"},{"issue":"5","key":"4446_CR8","doi-asserted-by":"publisher","first-page":"1059","DOI":"10.1016\/j.bpj.2012.01.039","volume":"102","author":"C-C Lee","year":"2012","unstructured":"Lee C-C, Sun Y, Huang HW. How type ii diabetes-related islet amyloid polypeptide damages lipid bilayers. Biophys J. 2012;102(5):1059\u201368.","journal-title":"Biophys J"},{"key":"4446_CR9","doi-asserted-by":"publisher","first-page":"515","DOI":"10.3389\/fgene.2018.00515","volume":"9","author":"Q Zou","year":"2018","unstructured":"Zou Q, Qu K, Luo Y, Yin D, Ju Y, Tang H. Predicting diabetes mellitus with machine learning techniques. Front Genet. 2018;9:515\u2013515.","journal-title":"Front Genet"},{"issue":"1","key":"4446_CR10","doi-asserted-by":"publisher","first-page":"151","DOI":"10.1016\/j.ymeth.2004.03.012","volume":"34","author":"MR Nilsson","year":"2004","unstructured":"Nilsson MR. Techniques to study amyloid fibril formation in vitro. Methods. 2004;34(1):151\u201360.","journal-title":"Methods"},{"issue":"7","key":"4446_CR11","doi-asserted-by":"publisher","first-page":"1395","DOI":"10.1039\/b706784b","volume":"37","author":"GG Tartaglia","year":"2008","unstructured":"Tartaglia GG, Vendruscolo M. The zyggregator method for predicting protein aggregation propensities. Chem Soc Rev. 2008;37(7):1395\u2013401.","journal-title":"Chem Soc Rev"},{"key":"4446_CR12","doi-asserted-by":"crossref","unstructured":"Conchillo-Sol\u00e9 O, de Groot NS, Avil\u00e9s FX, Vendrell J, Daura X, Ventura S. Aggrescan: a server for the prediction and evaluation of \u201chot spots\u201d of aggregation in polypeptides. BMC Bioinform. 2007;8(1):65\u201365.","DOI":"10.1186\/1471-2105-8-65"},{"issue":"3","key":"4446_CR13","doi-asserted-by":"publisher","first-page":"237","DOI":"10.1038\/nmeth.1432","volume":"7","author":"S Maurer-Stroh","year":"2010","unstructured":"Maurer-Stroh S, Debulpaep M, Kuemmerer N, de la Paz ML, Martins IC, Reumers J, Morris KL, Copland A, Serpell L, Serrano L, Schymkowitz JWH, Rousseau F. Exploring the sequence determinants of amyloid structure using position-specific scoring matrices. Nat Methods. 2010;7(3):237\u201342.","journal-title":"Nat Methods"},{"issue":"1","key":"4446_CR14","doi-asserted-by":"publisher","first-page":"54","DOI":"10.1186\/1471-2105-15-54","volume":"15","author":"P Gasior","year":"2014","unstructured":"Gasior P, Kotulska M. Fish amyloid\u2014a new method for finding amyloidogenic segments in proteins based on site specific co-occurence of aminoacids. BMC Bioinform. 2014;15(1):54\u201354.","journal-title":"BMC Bioinform"},{"key":"4446_CR15","doi-asserted-by":"publisher","first-page":"469","DOI":"10.1093\/nar\/gkp351","volume":"37","author":"C Kim","year":"2009","unstructured":"Kim C, Choi J, Lee SJ, Welsh WJ, Yoon S. Netcssp: web application for predicting chameleon sequences and amyloid fibril formation. Nucleic Acids Res. 2009;37:469\u201373.","journal-title":"Nucleic Acids Res"},{"issue":"10","key":"4446_CR16","doi-asserted-by":"publisher","first-page":"521","DOI":"10.1093\/protein\/gzm042","volume":"20","author":"A Trovato","year":"2007","unstructured":"Trovato A, Seno F, Tosatto SCE. The pasta server for protein aggregation prediction. Protein Eng Des Select. 2007;20(10):521\u20133.","journal-title":"Protein Eng Des Select"},{"issue":"3","key":"4446_CR17","doi-asserted-by":"publisher","first-page":"326","DOI":"10.1093\/bioinformatics\/btp691","volume":"26","author":"SO Garbuzynskiy","year":"2010","unstructured":"Garbuzynskiy SO, Lobanov MY, Galzitskaya OV. Foldamyloid: a method of prediction of amyloidogenic regions from protein sequence. Bioinformatics. 2010;26(3):326\u201332.","journal-title":"Bioinformatics"},{"issue":"1","key":"4446_CR18","doi-asserted-by":"publisher","first-page":"44","DOI":"10.1186\/1472-6807-9-44","volume":"9","author":"KK Frousios","year":"2009","unstructured":"Frousios KK, Iconomidou VA, Karletidi C-M, Hamodrakas SJ. Amyloidogenic determinants are usually not buried. BMC Struct Biol. 2009;9(1):44\u201344.","journal-title":"BMC Struct Biol"},{"key":"4446_CR19","doi-asserted-by":"crossref","unstructured":"Tsolis AC, Papandreou NC, Iconomidou VA, Hamodrakas SJ. A consensus method for the prediction of \u201caggregation-prone\u201d peptides in globular proteins. PLoS ONE. 2013;8(1).","DOI":"10.1371\/journal.pone.0054175"},{"key":"4446_CR20","doi-asserted-by":"crossref","unstructured":"Emily M, Talvas A, Delamarche C. Metamyl: a meta-predictor for amyloid proteins. PLoS ONE. 2013;8(11).","DOI":"10.1371\/journal.pone.0079722"},{"issue":"8","key":"4446_CR21","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1371\/journal.pone.0134679","volume":"10","author":"C Fam\u00edlia","year":"2015","unstructured":"Fam\u00edlia C, Dennison SR, Quintas AL, Phoenix DA. Prediction of peptide and protein propensity for amyloid formation. PLoS ONE. 2015;10(8):1\u201316.","journal-title":"PLoS ONE"},{"issue":"1","key":"4446_CR22","doi-asserted-by":"publisher","first-page":"12961","DOI":"10.1038\/s41598-017-13210-9","volume":"7","author":"M Burdukiewicz","year":"2017","unstructured":"Burdukiewicz M, Sobczyk P, R\u00f6diger S, Duda-Madej A, Mackiewicz P, Kotulska M. Amyloidogenic motifs revealed by n-gram analysis. Sci Rep. 2017;7(1):12961\u201312961.","journal-title":"Sci Rep"},{"key":"4446_CR23","doi-asserted-by":"crossref","unstructured":"Bouziane H, Chouarfia A. Sequence- and structure-based prediction of amyloidogenic regions in proteins. In: Soft Computing, vol. 24, pp 3285\u20133308 (2020)","DOI":"10.1007\/s00500-019-04087-z"},{"key":"4446_CR24","doi-asserted-by":"publisher","first-page":"113362","DOI":"10.1016\/j.ab.2019.113362","volume":"583","author":"C Zhou","year":"2019","unstructured":"Zhou C, Liu S, Zhang S. Identification of amyloidogenic peptides via optimized integrated features space based on physicochemical properties and pssm. Anal Biochem. 2019;583:113362.","journal-title":"Anal Biochem"},{"issue":"1","key":"4446_CR25","doi-asserted-by":"publisher","first-page":"87","DOI":"10.1073\/pnas.2634884100","volume":"101","author":"ML de la Paz","year":"2004","unstructured":"de la Paz ML, Serrano L. Sequence determinants of amyloid fibril formation. Proc Natl Acad Sci USA. 2004;101(1):87\u201392.","journal-title":"Proc Natl Acad Sci USA"},{"issue":"8","key":"4446_CR26","doi-asserted-by":"publisher","first-page":"531","DOI":"10.1093\/protein\/gzp037","volume":"22","author":"PK Teng","year":"2009","unstructured":"Teng PK, Eisenberg D. Short protein segments can drive a non-fibrillizing protein into the amyloid state. Protein Eng Des Select. 2009;22(8):531\u20136.","journal-title":"Protein Eng Des Select"},{"issue":"19","key":"4446_CR27","doi-asserted-by":"publisher","first-page":"7258","DOI":"10.1073\/pnas.0308249101","volume":"101","author":"S Ventura","year":"2004","unstructured":"Ventura S, Zurdo J, Narayanan S, Parre\u00f1o M, Mangues R, Reif B, Chiti F, Giannoni E, Dobson CM, Aviles FX, Serrano L. Short amino acid stretches can mediate amyloid formation in globular proteins: the src homology 3 (sh3) case. Proc Natl Acad Sci USA. 2004;101(19):7258\u201363.","journal-title":"Proc Natl Acad Sci USA"},{"key":"4446_CR28","doi-asserted-by":"crossref","unstructured":"Louros N, Konstantoulea K, Vleeschouwer MD, Ramakers M, Schymkowitz J, Rousseau F. Waltz-db 2.0: an updated database containing structural information of experimentally determined amyloid-forming peptides. Nucleic Acids Res 48 (2020)","DOI":"10.1093\/nar\/gkz758"},{"issue":"20","key":"4446_CR29","doi-asserted-by":"publisher","first-page":"3395","DOI":"10.1093\/bioinformatics\/btv375","volume":"31","author":"PP Wozniak","year":"2015","unstructured":"Wozniak PP, Kotulska M. Amyload: website dedicated to amyloidogenic protein fragments. Bioinformatics. 2015;31(20):3395\u20137.","journal-title":"Bioinformatics"},{"key":"4446_CR30","doi-asserted-by":"crossref","unstructured":"Walsh I, Seno F, Tosatto SCE, Trovato A. Pasta 2.0: an improved server for protein aggregation prediction. Nucleic Acids Res 42, 301\u2013307 (2014)","DOI":"10.1093\/nar\/gku399"},{"issue":"3","key":"4446_CR31","doi-asserted-by":"publisher","first-page":"190","DOI":"10.2174\/1574893614666181212102749","volume":"14","author":"J Zhang","year":"2019","unstructured":"Zhang J, Liu B. A review on the recent developments of sequence-based protein feature extraction methods. Curr Bioinform. 2019;14(3):190\u20139.","journal-title":"Curr Bioinform"},{"key":"4446_CR32","first-page":"6664362","volume":"2021","author":"D Zhang","year":"2021","unstructured":"Zhang D, Chen H-D, Zulfiqar H, Yuan S-S, Huang Q-L, Zhang Z-Y, Deng K-J. iblp: an xgboost-based predictor for identifying bioluminescent proteins. Comput Math Methods Med. 2021;2021:6664362\u20136664362.","journal-title":"Comput Math Methods Med"},{"key":"4446_CR33","doi-asserted-by":"publisher","first-page":"8926750","DOI":"10.1155\/2020\/8926750","volume":"2020","author":"Z Tao","year":"2020","unstructured":"Tao Z, Li Y, Teng Z, Zhao Y. A method for identifying vesicle transport proteins based on libsvm and mrmd. Comput Math Methods Med. 2020;2020:8926750\u20138926750.","journal-title":"Comput Math Methods Med"},{"key":"4446_CR34","doi-asserted-by":"crossref","unstructured":"Lv H, Dao F-Y, Guan Z-X, Yang H, Li Y-W, Lin H. Deep-kcr: accurate detection of lysine crotonylation sites using deep learning method. Briefings in Bioinformatics (2020)","DOI":"10.1093\/bib\/bbaa255"},{"issue":"16","key":"4446_CR35","doi-asserted-by":"publisher","first-page":"4466","DOI":"10.1093\/bioinformatics\/btaa428","volume":"36","author":"T Zhao","year":"2020","unstructured":"Zhao T, Hu Y, Peng J, Cheng L. Deeplgp: a novel deep learning method for prioritizing lncrna target genes. Bioinformatics. 2020;36(16):4466\u201372.","journal-title":"Bioinformatics"},{"issue":"6","key":"4446_CR36","doi-asserted-by":"publisher","first-page":"2185","DOI":"10.1093\/bib\/bbz139","volume":"21","author":"B Liu","year":"2020","unstructured":"Liu B, Zhu Y, Yan K. Fold-ltr-tcp: protein fold recognition based on triadic closure principle. Brief Bioinform. 2020;21(6):2185\u201393.","journal-title":"Brief Bioinform"},{"issue":"21","key":"4446_CR37","doi-asserted-by":"publisher","first-page":"5177","DOI":"10.1093\/bioinformatics\/btaa667","volume":"36","author":"Y-J Tang","year":"2021","unstructured":"Tang Y-J, Pang Y-H, Liu B. Idp-seq2seq: identification of intrinsically disordered regions based on sequence to sequence learning. Bioinformatics. 2021;36(21):5177\u201386.","journal-title":"Bioinformatics"},{"issue":"3","key":"4446_CR38","doi-asserted-by":"publisher","first-page":"246","DOI":"10.1002\/prot.1035","volume":"43","author":"K-C Chou","year":"2001","unstructured":"Chou K-C. Prediction of protein cellular attributes using pseudo- amino acid composition. Proteins. 2001;43(3):246\u201355.","journal-title":"Proteins"},{"issue":"8","key":"4446_CR39","doi-asserted-by":"publisher","first-page":"937","DOI":"10.2174\/1574893615666200129110450","volume":"15","author":"S Naseer","year":"2021","unstructured":"Naseer S, Hussain W, Khan YD, Rasool N. Sequence-based identification of arginine amidation sites in proteins using deep representations of proteins and pseaac. Curr Bioinform. 2021;15(8):937\u201348.","journal-title":"Curr Bioinform"},{"issue":"3","key":"4446_CR40","doi-asserted-by":"publisher","first-page":"235","DOI":"10.2174\/1574893614666191202152328","volume":"15","author":"MAM Hasan","year":"2020","unstructured":"Hasan MAM, Islam KB, Rahman J, Ahmad S. Citrullination site prediction by incorporating sequence coupled effects into pseaac and resolving data imbalance issue. Curr Bioinform. 2020;15(3):235\u201345.","journal-title":"Curr Bioinform"},{"issue":"5","key":"4446_CR41","doi-asserted-by":"publisher","first-page":"396","DOI":"10.2174\/1574893614666190723114923","volume":"15","author":"S Amanat","year":"2020","unstructured":"Amanat S, Ashraf A, Hussain W, Rasool N, Khan YD. Identification of lysine carboxylation sites in proteins by integrating statistical moments and position relative features via general pseaac. Curr Bioinform. 2020;15(5):396\u2013407.","journal-title":"Curr Bioinform"},{"issue":"4","key":"4446_CR42","doi-asserted-by":"publisher","first-page":"1280","DOI":"10.1093\/bib\/bbx165","volume":"20","author":"B Liu","year":"2019","unstructured":"Liu B. Bioseq-analysis: a platform for dna, rna and protein sequence analysis based on machine learning approaches. Brief Bioinform. 2019;20(4):1280\u201394.","journal-title":"Brief Bioinform"},{"issue":"1","key":"4446_CR43","doi-asserted-by":"publisher","first-page":"10","DOI":"10.1093\/bioinformatics\/bth466","volume":"21","author":"K-C Chou","year":"2005","unstructured":"Chou K-C. Using amphiphilic pseudo amino acid composition to predict enzyme subfamily classes. Bioinformatics. 2005;21(1):10\u20139.","journal-title":"Bioinformatics"},{"issue":"1","key":"4446_CR44","doi-asserted-by":"publisher","first-page":"43","DOI":"10.1186\/s12859-020-3388-y","volume":"21","author":"X Zhao","year":"2020","unstructured":"Zhao X, Jiao Q, Li H, Wu Y, Wang H, Huang S, Wang G. Ecfs-dea: an ensemble classifier-based feature selection for differential expression analysis on expression profiles. BMC Bioinform. 2020;21(1):43.","journal-title":"BMC Bioinform"},{"issue":"11","key":"4446_CR45","doi-asserted-by":"publisher","first-page":"1953","DOI":"10.1093\/bioinformatics\/bty002","volume":"34","author":"L Cheng","year":"2018","unstructured":"Cheng L, Hu Y, Sun J, Zhou M, Jiang Q. Dincrna: a comprehensive web-based bioinformatics toolkit for exploring disease associations and ncrna function. Bioinformatics. 2018;34(11):1953\u20136.","journal-title":"Bioinformatics"},{"issue":"4","key":"4446_CR46","doi-asserted-by":"publisher","first-page":"210","DOI":"10.2174\/156652321904191022113307","volume":"19","author":"L Cheng","year":"2019","unstructured":"Cheng L. Computational and biological methods for gene therapy. Curr Gene Ther. 2019;19(4):210\u2013210.","journal-title":"Curr Gene Ther"},{"key":"4446_CR47","doi-asserted-by":"publisher","first-page":"590","DOI":"10.1016\/j.omtn.2019.09.019","volume":"18","author":"L Cheng","year":"2019","unstructured":"Cheng L, Zhao H, Wang P, Zhou W, Luo M, Li T, Han J, Liu S, Jiang Q. Computational methods for identifying similar diseases. Molecular Therapy Nucleic Acids. 2019;18:590\u2013604.","journal-title":"Molecular Therapy Nucleic Acids"},{"issue":"4","key":"4446_CR48","doi-asserted-by":"publisher","first-page":"2466","DOI":"10.3934\/mbe.2019123","volume":"16","author":"JX Tan","year":"2019","unstructured":"Tan JX, Li SH, Zhang ZM, Chen CX, Chen W, Tang H, Lin H. Identification of hormone binding proteins based on machine learning methods. Math Biosci Eng. 2019;16(4):2466\u201380.","journal-title":"Math Biosci Eng"},{"key":"4446_CR49","doi-asserted-by":"publisher","first-page":"787","DOI":"10.1016\/j.knosys.2018.10.007","volume":"163","author":"X-J Zhu","year":"2019","unstructured":"Zhu X-J, Feng C-Q, Lai H-Y, Chen W, Hao L. Predicting protein structural classes for low-similarity sequences by evaluating different features. Knowl Based Syst. 2019;163:787\u201393.","journal-title":"Knowl Based Syst"},{"key":"4446_CR50","doi-asserted-by":"publisher","first-page":"8845133","DOI":"10.1155\/2020\/8845133","volume":"2020","author":"Y Li","year":"2020","unstructured":"Li Y, Zhang Z, Teng Z, Liu X. Predamyl-mlp: prediction of amyloid proteins using multilayer perceptron. Comput Math Methods Med. 2020;2020:8845133.","journal-title":"Comput Math Methods Med"},{"key":"4446_CR51","doi-asserted-by":"crossref","unstructured":"Shida H, Fei G, Quan Z, HuiDing: Mrmd2.0: a python tool for machine learning with feature ranking and reduction. Curr Bioinform 15(10), 1213\u20131221 (2021)","DOI":"10.2174\/1574893615999200503030350"},{"key":"4446_CR52","doi-asserted-by":"crossref","unstructured":"Yang H, Luo Y, Ren X, Wu M, He X, Peng B, Deng K, Yan D, Tang H, Lin H. Risk prediction of diabetes: Big data mining with fusion of multifarious physical examination indicators. Inf Fusion. 2021.","DOI":"10.1016\/j.inffus.2021.02.015"},{"key":"4446_CR53","doi-asserted-by":"publisher","first-page":"1043","DOI":"10.1016\/j.omtn.2020.07.035","volume":"22","author":"M-L Liu","year":"2020","unstructured":"Liu M-L, Su W, Wang J-S, Yang Y-H, Yang H, Lin H. Predicting preference of transcription factors for methylated dna using sequence information. Mol Ther Nucleic acids. 2020;22:1043\u201350.","journal-title":"Mol Ther Nucleic acids"},{"key":"4446_CR54","doi-asserted-by":"crossref","unstructured":"Shao J, Yan K, Liu B. Foldrec-c2c: protein fold recognition by combining cluster-to-cluster model and protein similarity network. Briefings Bioinform. 2020.","DOI":"10.1093\/bib\/bbaa144"},{"key":"4446_CR55","doi-asserted-by":"crossref","unstructured":"Liu B, Gao X, Zhang H. Bioseq-analysis2.0: an updated platform for analyzing dna, rna and protein sequences at sequence level and residue level based on machine learning approaches. Nucleic Acids Res 47(20) (2019)","DOI":"10.1093\/nar\/gkz740"},{"issue":"5","key":"4446_CR56","doi-asserted-by":"publisher","first-page":"1568","DOI":"10.1093\/bib\/bbz123","volume":"21","author":"H Yang","year":"2020","unstructured":"Yang H, Yang W, Dao F-Y, Lv H, Ding H, Chen W, Lin H. A comparison and assessment of computational method for identifying recombination hotspots in saccharomyces cerevisiae. Brief Bioinform. 2020;21(5):1568\u201380.","journal-title":"Brief Bioinform"},{"issue":"1","key":"4446_CR57","doi-asserted-by":"publisher","first-page":"526","DOI":"10.1093\/bib\/bbz177","volume":"22","author":"Z-Y Zhang","year":"2021","unstructured":"Zhang Z-Y, Yang Y-H, Ding H, Wang D, Chen W, Lin H. Design powerful predictor for mrna subcellular location prediction in homo sapiens. Brief Bioinform. 2021;22(1):526\u201335.","journal-title":"Brief Bioinform"},{"key":"4446_CR58","doi-asserted-by":"publisher","first-page":"483","DOI":"10.1007\/s11103-020-01102-y","volume":"105","author":"M Niu","year":"2021","unstructured":"Niu M, Lin Y, Zou Q. sgrnacnn: identifying sgrna on-target activity in four crops using ensembles of convolutional neural networks. Plant Mol Biol. 2021;105:483\u201395.","journal-title":"Plant Mol Biol"},{"issue":"4","key":"4446_CR59","doi-asserted-by":"publisher","first-page":"309","DOI":"10.2174\/1574893614666191202153824","volume":"15","author":"S Nashreen","year":"2020","unstructured":"Nashreen S, Nonita S, Krishna PS, Shobhit V. A sequential ensemble model for communicable disease forecasting. Curr Bioinform. 2020;15(4):309\u201317.","journal-title":"Curr Bioinform"},{"issue":"3","key":"4446_CR60","doi-asserted-by":"publisher","first-page":"184","DOI":"10.2174\/1566523220999200716111502","volume":"20","author":"A Iqubal","year":"2020","unstructured":"Iqubal A, Iqubal MK, Khan A, Ali J, Baboota S, Haque SE. Gene therapy, a novel therapeutic tool for neurological disorders: current progress, challenges and future prospective. Curr Gene Ther. 2020;20(3):184\u201394.","journal-title":"Curr Gene Ther"},{"key":"4446_CR61","doi-asserted-by":"publisher","first-page":"134","DOI":"10.3389\/fbioe.2020.00134","volume":"8","author":"Z Lv","year":"2020","unstructured":"Lv Z, Zhang J, Ding H, Zou Q. Rf-pseu: a random forest predictor for rna pseudouridine sites. Front Bioeng Biotechnol. 2020;8:134.","journal-title":"Front Bioeng Biotechnol"},{"issue":"7","key":"4446_CR62","doi-asserted-by":"publisher","first-page":"2931","DOI":"10.1021\/acs.jproteome.9b00250","volume":"18","author":"X Ru","year":"2019","unstructured":"Ru X, Li L, Zou Q. Incorporating distance-based top-n-gram and random forest to identify electron transport proteins. J Proteome Res. 2019;18(7):2931\u20139.","journal-title":"J Proteome Res"},{"issue":"1","key":"4446_CR63","doi-asserted-by":"publisher","first-page":"44","DOI":"10.2174\/1566523220666200516170137","volume":"20","author":"S Bhakta","year":"2020","unstructured":"Bhakta S, Tsukahara T. Artificial rna editing with adar for gene therapy. Curr Gene Ther. 2020;20(1):44\u201354.","journal-title":"Curr Gene Ther"},{"issue":"1","key":"4446_CR64","doi-asserted-by":"publisher","first-page":"192","DOI":"10.1109\/TCBB.2013.146","volume":"11","author":"L Wei","year":"2014","unstructured":"Wei L, Liao M, Gao Y, Ji R, He Z, Zou Q. Improved and promising identification of human micrornas by incorporating a high-quality negative set. IEEE\/ACM Trans Comput Biol Bioinf. 2014;11(1):192\u2013201.","journal-title":"IEEE\/ACM Trans Comput Biol Bioinf"},{"issue":"384","key":"4446_CR65","doi-asserted-by":"publisher","first-page":"135","DOI":"10.1016\/j.ins.2016.06.026","volume":"384","author":"L Wei","year":"2017","unstructured":"Wei L, Tang J, Zou Q. Local-dpp: an improved dna-binding protein prediction method by exploring local evolutionary information. Inf Sci. 2017;384(384):135\u201344.","journal-title":"Inf Sci"},{"issue":"4","key":"4446_CR66","doi-asserted-by":"publisher","first-page":"1264","DOI":"10.1109\/TCBB.2017.2670558","volume":"16","author":"L Wei","year":"2019","unstructured":"Wei L, Xing P, Shi G, Ji Z, Zou Q. Fast prediction of protein methylation sites using a sequence-based feature selection technique. IEEE\/ACM Trans Comput Biol Bioinf. 2019;16(4):1264\u201373.","journal-title":"IEEE\/ACM Trans Comput Biol Bioinf"},{"key":"4446_CR67","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1016\/j.artmed.2017.03.001","volume":"83","author":"L Wei","year":"2017","unstructured":"Wei L, Xing P, Zeng J, Chen J, Su R, Guo F. Improved prediction of protein-protein interactions using novel negative samples, features, and an ensemble classifier. Artif Intell Med. 2017;83:67\u201374.","journal-title":"Artif Intell Med"},{"key":"4446_CR68","doi-asserted-by":"publisher","first-page":"506681","DOI":"10.3389\/fpls.2021.506681","volume":"12","author":"X Zhao","year":"2021","unstructured":"Zhao X, Wang H, Li H, Wu Y, Wang G. Identifying plant pentatricopeptide repeat proteins using a variable selection method. Front Plant Sci. 2021;12:506681\u2013506681.","journal-title":"Front Plant Sci"},{"key":"4446_CR69","doi-asserted-by":"crossref","unstructured":"Wang G, Luo X, Wang J, Wan J, Xia S, Zhu H, Qian J, Wang Y. Medreaders: a database for transcription factors that bind to methylated dna. Nucleic Acids Res. 2018;46.","DOI":"10.1093\/nar\/gkx1096"},{"key":"4446_CR70","doi-asserted-by":"publisher","first-page":"82","DOI":"10.1016\/j.artmed.2017.02.005","volume":"83","author":"L Wei","year":"2017","unstructured":"Wei L, Wan S, Guo J, Wong KK. A novel hierarchical selective ensemble classifier with bioinformatics application. Artif Intell Med. 2017;83:82\u201390.","journal-title":"Artif Intell Med"},{"issue":"23","key":"4446_CR71","doi-asserted-by":"publisher","first-page":"4007","DOI":"10.1093\/bioinformatics\/bty451","volume":"34","author":"L Wei","year":"2018","unstructured":"Wei L, Zhou C, Chen H, Song J, Su R. Acpred-fl: a sequence-based predictor using effective feature representation to improve the prediction of anti-cancer peptides. Bioinformatics. 2018;34(23):4007\u201316.","journal-title":"Bioinformatics"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04446-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1186\/s12859-021-04446-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-021-04446-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,8]],"date-time":"2023-02-08T18:04:12Z","timestamp":1675879452000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-021-04446-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,11,9]]},"references-count":71,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2021,12]]}},"alternative-id":["4446"],"URL":"https:\/\/doi.org\/10.1186\/s12859-021-04446-4","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,11,9]]},"assertion":[{"value":"30 April 2021","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 October 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"9 November 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no competing interests.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"545"}}