{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T17:36:42Z","timestamp":1782322602073,"version":"3.54.5"},"reference-count":49,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2021,3,30]],"date-time":"2021-03-30T00:00:00Z","timestamp":1617062400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,3,30]],"date-time":"2021-03-30T00:00:00Z","timestamp":1617062400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100006360","name":"Bundesministerium f\u00fcr Wirtschaft und Energie","doi-asserted-by":"publisher","award":["03EGSBW498"],"award-info":[{"award-number":["03EGSBW498"]}],"id":[{"id":"10.13039\/501100006360","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Projekt DEAL"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["SN COMPUT. SCI."],"published-print":{"date-parts":[[2021,5]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Text classification is important to better understand online media. A major problem for creating accurate text classifiers using machine learning is small training sets due to the cost of annotating them. On this basis, we investigated how SVM and NBSVM text classifiers should be designed to achieve high accuracy and how the training sets should be sized to efficiently use annotation labor. We used a four-way repeated-measures full-factorial design of 32 design factor combinations. For each design factor combination 22 training set sizes were examined. These training sets were subsets of seven public text datasets. We study the statistical variance of accuracy estimates by randomly drawing new training sets, resulting in accuracy estimates for 98,560 different experimental runs. Our major contribution is a set of empirically evaluated guidelines for creating online media text classifiers using small training sets. We recommend uni- and bi-gram features as text representation, btc term weighting and a linear-kernel NBSVM. Our results suggest that high classification accuracy can be achieved using a manually annotated dataset of only 300 examples.<\/jats:p>","DOI":"10.1007\/s42979-021-00480-4","type":"journal-article","created":{"date-parts":[[2021,3,30]],"date-time":"2021-03-30T16:02:52Z","timestamp":1617120172000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":12,"title":["Simple Baseline Machine Learning Text Classifiers for Small Datasets"],"prefix":"10.1007","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-4859-6627","authenticated-orcid":false,"given":"Martin","family":"Riekert","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4522-9207","authenticated-orcid":false,"given":"Matthias","family":"Riekert","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Achim","family":"Klein","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2021,3,30]]},"reference":[{"key":"480_CR1","doi-asserted-by":"publisher","first-page":"1","DOI":"10.3390\/info11060314","volume":"11","author":"J Samuel","year":"2020","unstructured":"Samuel J, Ali GGMN, Rahman MM, Esawi E, Samuel Y. COVID-19 public sentiment insights and machine learning for tweets classification. Information. 2020;11:1\u201323.","journal-title":"Information"},{"key":"480_CR2","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/505282.505283","volume":"34","author":"F Sebastiani","year":"2002","unstructured":"Sebastiani F. Machine learning in automated text categorization. ACM Comput Surv. 2002;34:1\u201347.","journal-title":"ACM Comput Surv"},{"key":"480_CR3","unstructured":"Mitchell TM. Machine learning, vol. 45, No. 37. Burr Ridge, IL: McGraw Hill; 1997. p. 870\u20137."},{"key":"480_CR4","unstructured":"Cortes C, Jackel LD, Solla SA, Vapnik V, Denker JS. Learning curves: asymptotic values and rate of convergence. In: 6th International conference on neural information processing system, vol. 6, pp 327\u2013334, 1994"},{"key":"480_CR5","first-page":"2079","volume":"11","author":"GC Cawley","year":"2010","unstructured":"Cawley GC, Talbot NLC. On over-fitting in model selection and subsequent selection bias in performance evaluation. J Mach Learn Res. 2010;11:2079\u2013107.","journal-title":"J Mach Learn Res"},{"key":"480_CR6","doi-asserted-by":"publisher","first-page":"223","DOI":"10.1137\/16M1080173","volume":"60","author":"L Bottou","year":"2016","unstructured":"Bottou L, Curtis FE, Nocedal J. Optimization methods for large-scale machine learning. SIAM Rev. 2016;60:223\u2013311.","journal-title":"SIAM Rev"},{"key":"480_CR7","doi-asserted-by":"publisher","first-page":"1139","DOI":"10.1111\/j.1540-6261.2007.01232.x","volume":"62","author":"PCP Tetlock","year":"2007","unstructured":"Tetlock PCP, Content G, Sentiment I, Role T, Author SM, Source PCT, Journal T. Giving content to investor sentiment: the role of media in the stock market. J Finance. 2007;62:1139\u201368.","journal-title":"J Finance"},{"key":"480_CR8","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1016\/j.ijresmar.2018.09.009","volume":"36","author":"J Hartmann","year":"2019","unstructured":"Hartmann J, Huppertz J, Schamp C, Heitmann M. Comparing automated text classification methods. Int J Res Mark. 2019;36:20\u201338.","journal-title":"Int J Res Mark"},{"key":"480_CR9","doi-asserted-by":"publisher","first-page":"484","DOI":"10.1002\/bs.3830070412","volume":"7","author":"PJ Stone","year":"2007","unstructured":"Stone PJ, Bales RF, Namenwirth JZ, Ogilvie DM. The general inquirer: a computer system for content analysis and retrieval based on the sentence as a unit of information. Behav Sci. 2007;7:484\u201398.","journal-title":"Behav Sci"},{"key":"480_CR10","doi-asserted-by":"publisher","first-page":"363","DOI":"10.1177\/0021943608319388","volume":"45","author":"E Henry","year":"2008","unstructured":"Henry E. Are investors influenced by how earnings press releases are written? J Bus Commun. 2008;45:363\u2013407.","journal-title":"J Bus Commun"},{"key":"480_CR11","doi-asserted-by":"publisher","first-page":"1187","DOI":"10.1111\/1475-679X.12123","volume":"54","author":"T Loughran","year":"2016","unstructured":"Loughran T, McDonald B. Textual analysis in accounting and finance: a survey. J Acc Res. 2016;54:1187\u2013230. https:\/\/doi.org\/10.1111\/1475-679X.12123.","journal-title":"J Acc Res"},{"key":"480_CR12","unstructured":"Wang S, Manning CD. Baselines and bigrams: simple, good sentiment and topic classification. In: Proceedings of the 50th annual meeting of the association for computational linguistics, vol. 2. Jeju, South Korea, pp 90\u201394, 2012"},{"key":"480_CR13","doi-asserted-by":"publisher","first-page":"10760","DOI":"10.1016\/j.eswa.2009.02.063","volume":"36","author":"H Tang","year":"2009","unstructured":"Tang H, Tan S, Cheng X. A survey on sentiment detection of reviews. Expert Syst Appl. 2009;36:10760\u201373.","journal-title":"Expert Syst Appl"},{"key":"480_CR14","doi-asserted-by":"publisher","unstructured":"Zhang X, Zhao J, LeCun Y. Character-level convolutional networks for text classification. In: Proceedings of the 28th International Conference on Neural Information Processing Systems. Cambridge, MA: MIT Press; 2015. p. 649\u201357. https:\/\/doi.org\/10.5555\/2969239.2969312.","DOI":"10.5555\/2969239.2969312"},{"key":"480_CR15","first-page":"321","volume":"320","author":"A Klein","year":"2018","unstructured":"Klein A, Riekert M, Kirilov L, Leukel J. Increasing the explanatory power of investor sentiment analysis for commodities in online media. Lect Notes Bus Inf Process. 2018;320:321\u201332.","journal-title":"Lect Notes Bus Inf Process"},{"key":"480_CR16","unstructured":"Devlin J, Chang M-W, Lee K, Toutanova K. BERT: pre-training of deep bidirectional transformers for language understanding. 2018. arXiv:1810.04805."},{"key":"480_CR17","doi-asserted-by":"crossref","unstructured":"Howard J, Ruder S. Universal language model fine-tuning for text classification. In: 56th Annual Meeting of the Association for Computational Linguistics. 2019. p. 328\u201339. https:\/\/www.aclweb.org\/anthology\/P18-1031\/.","DOI":"10.18653\/v1\/P18-1031"},{"key":"480_CR18","unstructured":"Usherwood P, Smit S. Low-shot classification: a comparison of classical and deep transfer machine learning approaches. 2019. arXiv:1907.07543."},{"key":"480_CR19","unstructured":"B\u00fcy\u00fck\u00f6z B, H\u00fcrriyeto\u011flu A, \u00d6zg\u00fcr A. Analyzing ELMo and DistilBERT on socio-political news classification. In: Proceedings of the workshop on automated extraction of socio-political events from news. 2020, pp. 9\u201318"},{"key":"480_CR20","doi-asserted-by":"publisher","first-page":"105836","DOI":"10.1016\/j.asoc.2019.105836","volume":"86","author":"G Kou","year":"2020","unstructured":"Kou G, Yang P, Peng Y, Xiao F, Chen Y, Alsaadi FE. Evaluation of feature selection methods for text classification with small datasets using multiple criteria decision-making methods. Appl Soft Comput J. 2020;86:105836.","journal-title":"Appl Soft Comput J"},{"key":"480_CR21","doi-asserted-by":"crossref","unstructured":"Abdelwahab O, Bahgat M, Lowrance CJ, Elmaghraby A. Effect of training set size on SVM and Na\u00efve Bayes for Twitter sentiment analysis. In: 2015 IEEE International symposium on signal processing and information technology (ISSPIT). 2016, pp. 46\u201351","DOI":"10.1109\/ISSPIT.2015.7394379"},{"key":"480_CR22","doi-asserted-by":"publisher","first-page":"993","DOI":"10.1007\/s10796-017-9741-7","volume":"19","author":"Y Choi","year":"2017","unstructured":"Choi Y, Lee H. Data properties and the performance of sentiment classification for electronic commerce applications. Inf Syst Front. 2017;19:993\u20131012.","journal-title":"Inf Syst Front"},{"key":"480_CR23","doi-asserted-by":"publisher","first-page":"8","DOI":"10.1186\/1472-6947-12-8","volume":"12","author":"RL Figueroa","year":"2012","unstructured":"Figueroa RL, Zeng-Treitler Q, Kandula S, Ngo LH. Predicting sample size required for classification performance. BMC Med Inform Decis Mak. 2012;12:8.","journal-title":"BMC Med Inform Decis Mak"},{"key":"480_CR24","first-page":"397","volume":"2","author":"C Meek","year":"2002","unstructured":"Meek C, Thiesson B, Heckerman D. The learning-curve sampling method applied to model-based clustering. J Mach Learn Res. 2002;2:397\u2013418.","journal-title":"J Mach Learn Res"},{"key":"480_CR25","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511809071","volume-title":"Introduction to information retrieval","author":"CD Manning","year":"2008","unstructured":"Manning CD, Raghavan P, Sch\u00fctze H. Introduction to information retrieval. Cambridge: Cambridge University Press; 2008."},{"key":"480_CR26","doi-asserted-by":"publisher","first-page":"478","DOI":"10.1007\/s10618-011-0238-6","volume":"24","author":"M Tsytsarau","year":"2011","unstructured":"Tsytsarau M, Palpanas T. Survey on mining subjective data on the web. Data Min Knowl Discov. 2011;24:478\u2013514.","journal-title":"Data Min Knowl Discov"},{"key":"480_CR27","unstructured":"Maas AL, Daly RE, Pham PT, Huang D, Ng AY, Potts C. Learning word vectors for sentiment analysis. In: ACL-HLT 2011 Proceedings of the 49th annual meeting of the association for computational linguistics: human language technologies, vol. 1, 2011, pp. 142\u2013150"},{"key":"480_CR28","unstructured":"Riekert M, Leukel J, Klein A. Online media sentiment: Understanding machine learning-based classifiers. In: 24th European conference on information systems. 2016"},{"key":"480_CR29","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4615-0907-3","volume-title":"Learning to classify text using support vector machines","author":"T Joachims","year":"2002","unstructured":"Joachims T. Learning to classify text using support vector machines. Norwell: Kluwer Academic Publishers; 2002."},{"key":"480_CR30","doi-asserted-by":"publisher","first-page":"110","DOI":"10.1111\/j.1467-8640.2006.00277.x","volume":"22","author":"A Kennedy","year":"2006","unstructured":"Kennedy A, Inkpen D. Sentiment classification of movie reviews using contextual valence shifters. Comput Intell. 2006;22:110\u201325.","journal-title":"Comput Intell"},{"key":"480_CR31","doi-asserted-by":"publisher","first-page":"513","DOI":"10.1016\/0306-4573(88)90021-0","volume":"24","author":"G Salton","year":"1988","unstructured":"Salton G, Buckley C. Term-weighting approaches in automatic text retrieval. Inf Process Manag. 1988;24:513\u201323.","journal-title":"Inf Process Manag"},{"key":"480_CR32","unstructured":"Paltoglou G, Thelwall M. A study of Information Retrieval weighting schemes for sentiment analysis. In: 48th Annual meeting of the association for computational linguistics. 2010, pp. 1386\u20131395"},{"key":"480_CR33","unstructured":"O\u2019Keefe T, Koprinska I. Feature selection and weighting methods in sentiment analysis. In: 14th Australasian document computing symposium. 2009, pp. 67\u201374"},{"key":"480_CR34","doi-asserted-by":"crossref","unstructured":"Pang B, Lee L, Vaithyanathan S. Thumbs up? Sentiment classification using machine learning techniques. In: Proceedings of conference on empirical methods of Nat Lang Process, Philadelphia, PA, USA, 2002, pp. 79\u201386","DOI":"10.3115\/1118693.1118704"},{"key":"480_CR35","volume-title":"Human behavior and the principle of least effort","author":"GK Zipf","year":"1949","unstructured":"Zipf GK. Human behavior and the principle of least effort. Eastford: Martino Publishing; 1949."},{"key":"480_CR36","doi-asserted-by":"publisher","first-page":"503","DOI":"10.1108\/00220410410560582","volume":"60","author":"S Robertson","year":"2004","unstructured":"Robertson S. Understanding inverse document frequency: on theoretical arguments for IDF. J Doc. 2004;60:503\u201320.","journal-title":"J Doc"},{"key":"480_CR37","doi-asserted-by":"crossref","unstructured":"Joachims T. Text categorization with support vector machines: learning with many relevant features. In: Proceedings of 10th European conference on machine learning Chemnitz, Germany, 1998, pp. 137\u2013142","DOI":"10.1007\/BFb0026683"},{"key":"480_CR38","doi-asserted-by":"crossref","unstructured":"Ng V, Dasgupta S, Arifin N. Examining the role of linguistic knowledge sources in the automatic identification and classification of reviews. In: Proceedings of the 21st international conference on computational linguistics and 44th annual meeting of the association for computational linguistics. 2006, pp. 611\u2013618","DOI":"10.3115\/1273073.1273152"},{"key":"480_CR39","doi-asserted-by":"crossref","unstructured":"Boser B, Guyon I, Vapnik V. A training algorithm for optimal margin classifiers. In: 5th Annual ACM workshop on computational learning theory. 1992, pp. 144\u2013152","DOI":"10.1145\/130385.130401"},{"key":"480_CR40","unstructured":"McCallum A, Nigam K. A comparison of event models for naive bayes text classification. In: 15th National conference on artificial intelligence of working, learning and text category. 1998, pp. 41\u201348"},{"key":"480_CR41","doi-asserted-by":"publisher","first-page":"238","DOI":"10.1007\/s12559-019-09669-5","volume":"12","author":"Z Wang","year":"2020","unstructured":"Wang Z, Lin Z. Optimal feature selection for learning-based algorithms for sentiment classification. Cognit Comput. 2020;12:238\u201348.","journal-title":"Cognit Comput"},{"key":"480_CR42","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa F, Grisel O, Weiss R, et al. Scikit-learn: machine learning in Python. J Mach Learn Res. 2011;12:2825\u201330.","journal-title":"J Mach Learn Res"},{"key":"480_CR43","first-page":"1871","volume":"9","author":"R Fan","year":"2008","unstructured":"Fan R, Chang K, Hsieh C. LIBLINEAR: a library for large linear classification. J Mach Learn Res. 2008;9:1871\u20134.","journal-title":"J Mach Learn Res"},{"key":"480_CR44","first-page":"1","volume":"5","author":"R Kohavi","year":"1995","unstructured":"Kohavi R. A study of cross-validation and bootstrap for accuracy estimation and model selection. Int Jt Conf Artif Intell. 1995;5:1\u20137.","journal-title":"Int Jt Conf Artif Intell"},{"key":"480_CR45","doi-asserted-by":"publisher","DOI":"10.1007\/978-0-387-84858-7","volume-title":"The elements of statistical learning","author":"T Hastie","year":"2009","unstructured":"Hastie T, Tibshirani R, Friedman J. The elements of statistical learning. 2nd ed. New York: Springer; 2009.","edition":"2"},{"key":"480_CR46","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun Y, Bengio Y, Hinton G. Deep learning. Nature. 2015;521:436\u201344.","journal-title":"Nature"},{"key":"480_CR47","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1002\/cpe.5604","volume":"32","author":"Z Tang","year":"2020","unstructured":"Tang Z, Li W, Li Y. An improved term weighting scheme for text classification. Concurr Comput. 2020;32:1\u201319.","journal-title":"Concurr Comput"},{"key":"480_CR48","doi-asserted-by":"publisher","first-page":"3797","DOI":"10.1007\/s11042-018-6083-5","volume":"78","author":"X Deng","year":"2019","unstructured":"Deng X, Li Y, Weng J, Zhang J. Feature selection for text classification: a review. Multimed Tools Appl. 2019;78:3797\u2013816.","journal-title":"Multimed Tools Appl"},{"key":"480_CR49","doi-asserted-by":"publisher","first-page":"4","DOI":"10.1186\/1471-2105-10-147","volume":"10","author":"SY Kim","year":"2009","unstructured":"Kim SY. Effects of sample size on robustness and prediction accuracy of a prognostic gene signature. BMC Bioinform. 2009;10:4\u20137. https:\/\/doi.org\/10.1186\/1471-2105-10-147.","journal-title":"BMC Bioinform"}],"container-title":["SN Computer Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42979-021-00480-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s42979-021-00480-4\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s42979-021-00480-4.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,5,13]],"date-time":"2021-05-13T17:36:12Z","timestamp":1620927372000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s42979-021-00480-4"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,30]]},"references-count":49,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2021,5]]}},"alternative-id":["480"],"URL":"https:\/\/doi.org\/10.1007\/s42979-021-00480-4","relation":{},"ISSN":["2662-995X","2661-8907"],"issn-type":[{"value":"2662-995X","type":"print"},{"value":"2661-8907","type":"electronic"}],"subject":[],"published":{"date-parts":[[2021,3,30]]},"assertion":[{"value":"30 September 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 January 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"30 March 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Compliance with Ethical Standards"}},{"value":"Martin Riekert declares that he has no conflict of interest. Matthias Riekert declares that he has no conflict of interest. Achim Klein declares that he has no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interests"}},{"value":"This article does not contain any studies with human participants or animals performed by any of the authors.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}}],"article-number":"178"}}