{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T21:12:03Z","timestamp":1784581923823,"version":"3.55.0"},"reference-count":47,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2019,11,6]],"date-time":"2019-11-06T00:00:00Z","timestamp":1572998400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2019,11,6]],"date-time":"2019-11-06T00:00:00Z","timestamp":1572998400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["BMC Bioinformatics"],"published-print":{"date-parts":[[2019,12]]},"abstract":"<jats:title>Abstract<\/jats:title>\n              <jats:sec>\n                <jats:title>Background<\/jats:title>\n                <jats:p>Mass spectra are usually acquired from the Liquid Chromatography-Mass Spectrometry (LC-MS) analysis for isotope labeled proteomics experiments. In such experiments, the mass profiles of labeled (heavy) and unlabeled (light) peptide pairs are represented by isotope clusters (2D or 3D) that provide valuable information about the studied biological samples in different conditions. The core task of quality control in quantitative LC-MS experiment is to filter out low-quality peptides with questionable profiles. The commonly used methods for this problem are the classification approaches. However, the data imbalance problems in previous control methods are often ignored or mishandled. In this study, we introduced a quality control framework based on the extreme gradient boosting machine (XGBoost), and carefully addressed the imbalanced data problem in this framework.<\/jats:p>\n              <\/jats:sec>\n              <jats:sec>\n                <jats:title>Results<\/jats:title>\n                <jats:p>In the XGBoost based framework, we suggest the application of the Synthetic minority over-sampling technique (SMOTE) to re-balance data and use the balanced data to train the boosted trees as the classifier. Then the classifier is applied to other data for the peptide quality assessment. Experimental results show that our proposed framework increases the reliability of peptide heavy-light ratio estimation significantly.<\/jats:p>\n              <\/jats:sec>\n              <jats:sec>\n                <jats:title>Conclusions<\/jats:title>\n                <jats:p>Our results indicate that this framework is a powerful method for the peptide quality assessment. For the feature extraction part, the extracted ion chromatogram (XIC) based features contribute to the peptide quality assessment. To solve the imbalanced data problem, SMOTE brings a much better classification performance. Finally, the XGBoost is capable for the peptide quality control. Overall, our proposed framework provides reliable results for the further proteomics studies.<\/jats:p>\n              <\/jats:sec>","DOI":"10.1186\/s12859-019-3170-1","type":"journal-article","created":{"date-parts":[[2019,11,6]],"date-time":"2019-11-06T12:03:12Z","timestamp":1573041792000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":10,"title":["Quality control of imbalanced mass spectra from isotopic labeling experiments"],"prefix":"10.1186","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9127-8889","authenticated-orcid":false,"given":"Tianjun","family":"Li","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0184-5446","authenticated-orcid":false,"given":"Long","family":"Chen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2756-0054","authenticated-orcid":false,"given":"Min","family":"Gan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2019,11,6]]},"reference":[{"key":"3170_CR1","doi-asserted-by":"publisher","unstructured":"Zhang J, Gao W, Cai J, He S, Zeng R, Chen R. Predicting molecular formulas of fragment ions with isotope patterns in tandem mass spectra. IEEE\/ACM Trans Comput Biol Bioinformatics. 2005; 2(3):217\u201330. \n                    https:\/\/doi.org\/10.1109\/TCBB.2005.43\n                    \n                  .","DOI":"10.1109\/TCBB.2005.43"},{"key":"3170_CR2","doi-asserted-by":"publisher","unstructured":"Chen L, Petritis K, Tegeler T, Petritis B, Haskins WE, Zhang J. Improved quantification of labeled lc-ms. In: 2011 IEEE International Conference on Bioinformatics and Biomedicine: 2011. p. 299\u2013303. \n                    https:\/\/doi.org\/10.1109\/BIBM.2011.75\n                    \n                  .","DOI":"10.1109\/BIBM.2011.75"},{"key":"3170_CR3","doi-asserted-by":"publisher","unstructured":"Cui J, Ma X, Chen L, Zhang J. Scfia: a statistical corresponding feature identification algorithm for lc\/ms. BMC Bioinformatics. 2011; 12:439\u20139. \n                    https:\/\/doi.org\/10.1186\/1471-2105-12-439\n                    \n                  . 1471-2105-12-439[PII].","DOI":"10.1186\/1471-2105-12-439"},{"key":"3170_CR4","doi-asserted-by":"publisher","unstructured":"Yang P, Ma J, Wang P, Zhu Y, Zhou BB, Yang YH. Improving x!tandem on peptide identification from mass spectrometry by self-boosted percolator. IEEE\/ACM Transactions on Computational Biology and Bioinformatics. 2012; 9(5):1273\u201380. \n                    https:\/\/doi.org\/10.1109\/TCBB.2012.86\n                    \n                  .","DOI":"10.1109\/TCBB.2012.86"},{"key":"3170_CR5","doi-asserted-by":"publisher","unstructured":"Keller A, Nesvizhskii AI, Kolker E, Aebersold R. Empirical statistical model to estimate the accuracy of peptide identifications made by ms\/ms and database search. Analytical Chemistry. 2002; 74(20):5383\u201392. \n                    https:\/\/doi.org\/10.1021\/ac025747h\n                    \n                  .","DOI":"10.1021\/ac025747h"},{"key":"3170_CR6","doi-asserted-by":"publisher","unstructured":"Liu Y, Ma B, Zhang K, Lajoie G. An approach for peptide identification by de novo sequencing of mixture spectra. IEEE\/ACM Transactions on Computational Biology and Bioinformatics. 2017; 14(2):326\u201336. \n                    https:\/\/doi.org\/10.1109\/TCBB.2015.2407401\n                    \n                  .","DOI":"10.1109\/TCBB.2015.2407401"},{"key":"3170_CR7","doi-asserted-by":"publisher","first-page":"376","DOI":"10.1074\/mcp.M200025-MCP200","volume":"1","author":"S-E Ong","year":"2002","unstructured":"Ong S-E, Blagoev B, Kratchmarova I, Kristensen DB, Steen H, Pandey A, Mann M. Stable isotope labeling by amino acids in cell culture, silac, as a simple and accurate approach to expression proteomics. Mol Cell Proteome. 2002; 1:376\u201386.","journal-title":"Mol Cell Proteome"},{"key":"3170_CR8","doi-asserted-by":"publisher","unstructured":"Bittremieux W, Tabb DL, Impens F, Staes A, Timmerman E, Martens L, Laukens K. Quality control in mass spectrometry-based proteomics. Mass Spectrom Rev. 2017. \n                    https:\/\/doi.org\/10.1002\/mas.21544\n                    \n                  .","DOI":"10.1002\/mas.21544"},{"key":"3170_CR9","doi-asserted-by":"publisher","unstructured":"Kohlbacher O, Reinert K, Gr\u00f6pl C, Lange E, Pfeifer N, Schulz-Trieglaff O, Sturm M. Topp \u2013 the openms proteomics pipeline. Bioinformatics. 2007; 23(2):191. \n                    https:\/\/doi.org\/10.1093\/bioinformatics\/btl299\n                    \n                  .","DOI":"10.1093\/bioinformatics\/btl299"},{"key":"3170_CR10","doi-asserted-by":"publisher","unstructured":"Cox J., Mann M.Maxquant enables high peptide identification rates, individualized p.p.b.-range mass accuracies and proteome-wide protein quantification. Nat Biotech. 2008; 26(12):1367\u201372. \n                    https:\/\/doi.org\/10.1038\/nbt.1511\n                    \n                  .","DOI":"10.1038\/nbt.1511"},{"key":"3170_CR11","doi-asserted-by":"publisher","unstructured":"Keller A, Eng J, Zhang N, Li X-j, Aebersold R. A uniform proteomics ms\/ms analysis platform utilizing open xml file formats. Mol Syst Biol. 2005; 1(1). \n                    https:\/\/doi.org\/10.1038\/msb4100024\n                    \n                  . \n                    https:\/\/www.embopress.org\/doi\/pdf\/10.1038\/msb4100024\n                    \n                  .","DOI":"10.1038\/msb4100024"},{"issue":"6","key":"3170_CR12","doi-asserted-by":"publisher","first-page":"1150","DOI":"10.1002\/pmic.200900375","volume":"10","author":"EW Deutsch","year":"2010","unstructured":"Deutsch EW, Mendoza L, Shteynberg D, Farrah T, Lam H, Tasman N, Sun Z, Nilsson E, Pratt B, Prazen B, Eng JK, Martin DB, Nesvizhskii A, Aebersold R. A guided tour of the trans-proteomic pipeline. Proteomics. 2010; 10(6):1150\u20139.","journal-title":"Proteomics"},{"key":"3170_CR13","volume-title":"Trans-Proteomic Pipeline: A Pipeline for Proteomic Analysis","author":"PGA Pedrioli","year":"2010","unstructured":"Pedrioli PGA. Trans-Proteomic Pipeline: A Pipeline for Proteomic Analysis. Totowa: Humana Press; 2010, pp. 213\u2013238."},{"key":"3170_CR14","doi-asserted-by":"publisher","unstructured":"Deutsch EW, Lam H, Aebersold R. Data analysis and bioinformatics tools for tandem mass spectrometry in proteomics. Physiol Genomics. 2008; 33(1):18\u201325. \n                    https:\/\/doi.org\/10.1152\/physiolgenomics.00298.2007\n                    \n                  . \n                    https:\/\/www.physiology.org\/doi\/pdf\/10.1152\/physiolgenomics.00298.2007\n                    \n                  .","DOI":"10.1152\/physiolgenomics.00298.2007"},{"key":"3170_CR15","doi-asserted-by":"publisher","unstructured":"Pan C, Kora G, Tabb DL, Pelletier DA, McDonald WH, Hurst GB, Hettich RL, Samatova NF. Robust estimation of peptide abundance ratios and rigorous scoring of their variability and bias in quantitative shotgun proteomics. Anal Chem. 2006; 78(20):7110\u201320. \n                    https:\/\/doi.org\/10.1021\/ac0606554\n                    \n                  .","DOI":"10.1021\/ac0606554"},{"key":"3170_CR16","doi-asserted-by":"publisher","unstructured":"Bakalarski CE, Elias JE, Vill\u00e9n J, Haas W, Gerber SA, Everley PA, Gygi SP. The impact of peptide abundance and dynamic range on stable-isotope-based quantitative proteomic analyses. Journal of Proteome Research. 2008; 7(11):4756\u201365. \n                    https:\/\/doi.org\/10.1021\/pr800333e\n                    \n                  .","DOI":"10.1021\/pr800333e"},{"key":"3170_CR17","doi-asserted-by":"publisher","unstructured":"Sadygov R. G., Zhao Y., Haidacher S. J., Starkey J. M., Tilton R. G., Denner L.Using power spectrum analysis to evaluate 18o-water labeling data acquired from low resolution mass spectrometers. J Proteome Res. 2010; 9(8):4306\u201312. \n                    https:\/\/doi.org\/10.1021\/pr100642q\n                    \n                  .","DOI":"10.1021\/pr100642q"},{"issue":"1","key":"3170_CR18","doi-asserted-by":"publisher","first-page":"144","DOI":"10.1074\/mcp.M500230-MCP200","volume":"5","author":"JC Silva","year":"2006","unstructured":"Silva JC, Gorenstein MV, Li G-Z, Vissers JP, Geromanos SJ. Absolute quantification of proteins by lcmse: a virtue of parallel ms acquisition. Mol Cell Proteomics. 2006; 5(1):144\u201356.","journal-title":"Mol Cell Proteomics"},{"issue":"2","key":"3170_CR19","doi-asserted-by":"publisher","first-page":"137","DOI":"10.1021\/pr0255654","volume":"2","author":"D Anderson","year":"2003","unstructured":"Anderson D, Li W, Payan DG, Noble WS. A new algorithm for the evaluation of shotgun peptide sequencing in proteomics: support vector machine classification of peptide ms\/ms spectra and sequest scores. J Proteome Res. 2003; 2(2):137\u201346.","journal-title":"J Proteome Res"},{"key":"3170_CR20","doi-asserted-by":"publisher","unstructured":"Nefedov AV, Gilski MJ, Sadygov RG. Svm model for quality assessment of medium resolution mass spectra from 18o-water labeling experiments. J Proteome Res. 2011; 10(4):2095\u2013103. \n                    https:\/\/doi.org\/10.1021\/pr1012174\n                    \n                  .","DOI":"10.1021\/pr1012174"},{"key":"3170_CR21","doi-asserted-by":"publisher","unstructured":"Chang C, Zhang J, Han M, Ma J, Zhang W, Wu S, Liu K, Xie H, He F, Zhu Y. Silver: an efficient tool for stable isotope labeling lc-ms data quantitative analysis with quality control methods. Bioinformatics. 2014; 30(4):586\u20137. \n                    https:\/\/doi.org\/10.1093\/bioinformatics\/btt726\n                    \n                  .","DOI":"10.1093\/bioinformatics\/btt726"},{"issue":"10","key":"3170_CR22","doi-asserted-by":"publisher","first-page":"72951","DOI":"10.1371\/journal.pone.0072951","volume":"8","author":"J Cui","year":"2013","unstructured":"Cui J, Petritis K, Tegeler T, Petritis B, Ma X, Jin Y, Gao S-JS, Zhang JM. Accurate lc peak boundary detection for 16o\/18o labeled lc-ms data. PloS one. 2013; 8(10):72951.","journal-title":"PloS one"},{"key":"3170_CR23","doi-asserted-by":"publisher","unstructured":"IZMIRLIAN G. Application of the random forest classification algorithm to a seldi-tof proteomics study in the setting of a cancer prevention trial. Ann N Y Acad Sci. 2004; 1020(1):154\u201374. \n                    https:\/\/doi.org\/10.1196\/annals.1310.015\n                    \n                  .","DOI":"10.1196\/annals.1310.015"},{"key":"3170_CR24","doi-asserted-by":"publisher","unstructured":"Lin X, Wang Q, Yin P, Tang L, Tan Y, Li H, Yan K, Xu G. A method for handling metabonomics data from liquid chromatography\/mass spectrometry: combinational use of support vector machine recursive feature elimination, genetic algorithm and random forest for feature selection. Metabolomics. 2011; 7(4):549\u201358. \n                    https:\/\/doi.org\/10.1007\/s11306-011-0274-7\n                    \n                  .","DOI":"10.1007\/s11306-011-0274-7"},{"key":"3170_CR25","doi-asserted-by":"publisher","unstructured":"Swan AL, Mobasheri A, Allaway D, Liddell S, Bacardit J. Application of machine learning to proteomics data: classification and biomarker identification in postgenomics biology. OMICS. 2013; 17(12):595\u2013610. \n                    https:\/\/doi.org\/10.1089\/omi.2013.0017\n                    \n                  .","DOI":"10.1089\/omi.2013.0017"},{"key":"3170_CR26","unstructured":"Ma C. Deepquality: Mass spectra quality assessment via compressed sensing and deep learning. arXiv preprint arXiv:1710.11430. 2017."},{"key":"3170_CR27","doi-asserted-by":"publisher","unstructured":"Kim M, Eetemadi A, Tagkopoulos I. Deeppep: Deep proteome inference from peptide profiles. PLOS Comput Biol. 2017; 13(9):1\u201317. \n                    https:\/\/doi.org\/10.1371\/journal.pcbi.1005661\n                    \n                  .","DOI":"10.1371\/journal.pcbi.1005661"},{"key":"3170_CR28","doi-asserted-by":"publisher","first-page":"1559","DOI":"10.3389\/fpls.2018.01559","volume":"9","author":"D Zimmer","year":"2018","unstructured":"Zimmer D, Schneider K, Sommer F, Schroda M, M\u00fchlhaus T. Artificial intelligence understands peptide observability and assists with absolute protein quantification. Front Plant Sci. 2018; 9:1559.","journal-title":"Front Plant Sci"},{"issue":"1","key":"3170_CR29","first-page":"321","volume":"16","author":"NV Chawla","year":"2002","unstructured":"Chawla NV, Bowyer KW, Hall LO, Kegelmeyer WP. Smote: Synthetic minority over-sampling technique. J Artif Int Res. 2002; 16(1):321\u201357.","journal-title":"J Artif Int Res"},{"issue":"9","key":"3170_CR30","doi-asserted-by":"publisher","first-page":"1263","DOI":"10.1109\/TKDE.2008.239","volume":"21","author":"H He","year":"2009","unstructured":"He H, Garcia EA. Learning from imbalanced data. IEEE Trans Knowl Data Eng. 2009; 21(9):1263\u201384.","journal-title":"IEEE Trans Knowl Data Eng"},{"key":"3170_CR31","doi-asserted-by":"publisher","unstructured":"Wang S, Yao X. Multiclass imbalance problems: Analysis and potential solutions. IEEE Trans Syst Man Cybern B (Cybernetics). 2012; 42(4):1119\u201330. \n                    https:\/\/doi.org\/10.1109\/TSMCB.2012.2187280\n                    \n                  .","DOI":"10.1109\/TSMCB.2012.2187280"},{"key":"3170_CR32","doi-asserted-by":"publisher","unstructured":"Chen T, Guestrin C. Xgboost: A scalable tree boosting system. In: Proceedings of the 22Nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD \u201916. New York: ACM: 2016. p. 785\u2013794. \n                    https:\/\/doi.org\/10.1145\/2939672.2939785\n                    \n                  .","DOI":"10.1145\/2939672.2939785"},{"issue":"1","key":"3170_CR33","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1023\/A:1010933404324","volume":"45","author":"L Breiman","year":"2001","unstructured":"Breiman L. Random forests. Mach Learn. 2001; 45(1):5\u201332.","journal-title":"Mach Learn"},{"issue":"3","key":"3170_CR34","first-page":"18","volume":"2","author":"A Liaw","year":"2002","unstructured":"Liaw A, Wiener M. Classification and regression by randomforest. R news. 2002; 2(3):18\u201322.","journal-title":"R news"},{"key":"3170_CR35","doi-asserted-by":"publisher","unstructured":"Friedman JH. Greedy function approximation: A gradient boosting machine. Ann Statist. 2001; 29(5):1189\u2013232. \n                    https:\/\/doi.org\/10.1214\/aos\/1013203451\n                    \n                  .","DOI":"10.1214\/aos\/1013203451"},{"key":"3170_CR36","doi-asserted-by":"publisher","unstructured":"Li X-j, Zhang H, Ranish JA, Aebersold R. Automated statistical analysis of protein abundance ratios from data generated by stable-isotope dilution and tandem mass spectrometry. Analytical Chemistry. 2003; 75(23):6648\u201357. \n                    https:\/\/doi.org\/10.1021\/ac034633i\n                    \n                  .","DOI":"10.1021\/ac034633i"},{"key":"3170_CR37","doi-asserted-by":"publisher","unstructured":"Ross SM. Chapter 4 - random variables and expectation In: Ross SM, editor. Introduction to Probability and Statistics for Engineers and Scientists. Fifth edition. Boston: Academic Press: 2014. p. 89\u2013140. \n                    https:\/\/doi.org\/10.1016\/B978-0-12-394811-3.50004-6\n                    \n                  . \n                    http:\/\/www.sciencedirect.com\/science\/article\/pii\/B9780123948113500046\n                    \n                  .","DOI":"10.1016\/B978-0-12-394811-3.50004-6"},{"key":"3170_CR38","unstructured":"Nogueira F. A Python implementation of bayesian global optimization with gaussian processes. \n                    https:\/\/github.com\/fmfn\/BayesianOptimization\n                    \n                  ."},{"key":"3170_CR39","unstructured":"Chen T, He T, Khotilovich V, Xu B, Benesty M, Tang Y. dmlc XGBoost eXtreme Gradient Boosting. \n                    https:\/\/github.com\/dmlc\/xgboost\n                    \n                  ."},{"key":"3170_CR40","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay E. Scikit-learn: Machine learning in Python. J Mach Learn Res. 2011; 12:2825\u201330.","journal-title":"J Mach Learn Res"},{"key":"3170_CR41","unstructured":"Wikipedia contributors. Coefficient of variation \u2014 Wikipedia, The Free Encyclopedia. 2019. \n                    https:\/\/en.wikipedia.org\/w\/index.php?title=Coefficient_of_variation\n                    \n                  ."},{"issue":"4","key":"3170_CR42","first-page":"316","volume":"6","author":"J Canchola","year":"2017","unstructured":"Canchola J, Tang S, Hemyari P, Paxinos E, Marins E. Correct use of percent coefficient of variation (cv) formula for log-transformed data. MOJ Proteomics Bioinform. 2017; 6(4):316\u20137.","journal-title":"MOJ Proteomics Bioinform"},{"key":"3170_CR43","doi-asserted-by":"publisher","unstructured":"Bantscheff M, Schirle M, Sweetman G, Rick J, Kuster B. Quantitative mass spectrometry in proteomics: a critical review. Anal Bioanal Chem. 2007; 389(4):1017\u201331. \n                    https:\/\/doi.org\/10.1007\/s00216-007-1486-6\n                    \n                  .","DOI":"10.1007\/s00216-007-1486-6"},{"key":"3170_CR44","doi-asserted-by":"publisher","unstructured":"Ma L, Fan S. Cure-smote algorithm and hybrid algorithm for feature selection and parameter optimization based on random forests. BMC Bioinformatics. 2017; 18:169. \n                    https:\/\/doi.org\/10.1186\/s12859-017-1578-z\n                    \n                  .","DOI":"10.1186\/s12859-017-1578-z"},{"issue":"4","key":"3170_CR45","doi-asserted-by":"publisher","first-page":"320","DOI":"10.1016\/S1044-0305(99)00157-9","volume":"11","author":"DM Horn","year":"2000","unstructured":"Horn DM, Zubarev RA, McLafferty FW. Automated reduction and interpretation of high resolution electrospray mass spectra of large molecules. Journal of the American Society for Mass Spectrometry. 2000; 11(4):320\u201332.","journal-title":"Journal of the American Society for Mass Spectrometry"},{"issue":"17","key":"3170_CR46","first-page":"1","volume":"18","author":"G Lema\u00eetre","year":"2017","unstructured":"Lema\u00eetre G, Nogueira F, Aridas CK. Imbalanced-learn: A python toolbox to tackle the curse of imbalanced datasets in machine learning. J Mach Learn Res. 2017; 18(17):1\u20135.","journal-title":"J Mach Learn Res"},{"key":"3170_CR47","doi-asserted-by":"publisher","unstructured":"Friedman J, Hastie T, Tibshirani R. Additive logistic regression: a statistical view of boosting. Ann Stat. 2000; 28:337\u2013407. \n                    https:\/\/doi.org\/10.1214\/aos\/1016218223\n                    \n                  .","DOI":"10.1214\/aos\/1016218223"}],"container-title":["BMC Bioinformatics"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-019-3170-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1186\/s12859-019-3170-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1186\/s12859-019-3170-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2020,11,5]],"date-time":"2020-11-05T00:07:54Z","timestamp":1604534874000},"score":1,"resource":{"primary":{"URL":"https:\/\/bmcbioinformatics.biomedcentral.com\/articles\/10.1186\/s12859-019-3170-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,11,6]]},"references-count":47,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2019,12]]}},"alternative-id":["3170"],"URL":"https:\/\/doi.org\/10.1186\/s12859-019-3170-1","relation":{},"ISSN":["1471-2105"],"issn-type":[{"value":"1471-2105","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,11,6]]},"assertion":[{"value":"1 February 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 October 2019","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 November 2019","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"Not applicable.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval and consent to participate"}},{"value":"Not applicable.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"The authors declare that they have no competing interests.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"549"}}