{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,9]],"date-time":"2026-07-09T15:28:34Z","timestamp":1783610914051,"version":"3.55.0"},"reference-count":77,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2020,8,31]],"date-time":"2020-08-31T00:00:00Z","timestamp":1598832000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2020,8,31]],"date-time":"2020-08-31T00:00:00Z","timestamp":1598832000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["EPJ Data Sci."],"published-print":{"date-parts":[[2020,12]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In this case study, we are extending feature engineering approaches for short text samples by integrating techniques which have been introduced in the context of time series classification and signal processing. The general idea of the presented feature engineering approach is to tokenize the text samples under consideration and map each token to a number, which measures a specific property of the token. Consequently, each text sample becomes a language time series, which is generated from consecutively emitted tokens, and time is represented by the position of the respective token within the text sample. The resulting language time series can be characterised by collections of established time series feature extraction algorithms from time series analysis and signal processing. This approach maps each text sample (irrespective of its original length) to 3970 stylometric features, which can be analysed with standard statistical learning methodologies. The proposed feature engineering technique for short text data is applied to two different corpora: the Federalist Papers data set and the Spooky Books data set. We demonstrate that the extracted language time series features can be successfully combined with standard machine learning approaches for natural language processing and have the potential to improve the classification performance. Furthermore, the suggested feature engineering approach can be used for visualizing differences and commonalities of stylometric features. The presented framework models the systematic feature engineering based on approaches from time series classification and develops a statistical testing methodology for multi-classification problems.<\/jats:p>","DOI":"10.1140\/epjds\/s13688-020-00244-9","type":"journal-article","created":{"date-parts":[[2020,8,31]],"date-time":"2020-08-31T10:03:06Z","timestamp":1598868186000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["Enriching feature engineering for short text samples by language time series analysis"],"prefix":"10.1140","volume":"9","author":[{"given":"Yichen","family":"Tang","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Kelly","family":"Blincoe","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5558-0573","authenticated-orcid":false,"given":"Andreas W.","family":"Kempa-Liehr","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2020,8,31]]},"reference":[{"key":"244_CR1","volume-title":"Fundamentals of speech recognition","author":"LR Rabiner","year":"1993","unstructured":"Rabiner LR, Juang B-H (1993) Fundamentals of speech recognition. Prentice Hall, Englewood Cliffs"},{"issue":"1","key":"244_CR2","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1109\/TIFS.2016.2603960","volume":"12","author":"A Rocha","year":"2017","unstructured":"Rocha A, Scheirer WJ, Forstall CW, Cavalcante T, Theophilo A, Shen B, Carvalho ARB, Stamatatos E (2017) Authorship attribution for social media forensics. IEEE Trans Inf Forensics Secur 12(1):5\u201333. https:\/\/doi.org\/10.1109\/TIFS.2016.2603960","journal-title":"IEEE Trans Inf Forensics Secur"},{"key":"244_CR3","doi-asserted-by":"publisher","first-page":"90","DOI":"10.1016\/j.jbusres.2017.01.010","volume":"74","author":"Z-P Fan","year":"2017","unstructured":"Fan Z-P, Che Y-J, Chen Z-Y (2017) Product sales forecasting using online reviews and historical sales data: a method combining the Bass model and sentiment analysis. J Bus Res 74:90\u2013100. https:\/\/doi.org\/10.1016\/j.jbusres.2017.01.010","journal-title":"J Bus Res"},{"key":"244_CR4","series-title":"Annals of computer science and information systems","doi-asserted-by":"publisher","first-page":"1349","DOI":"10.15439\/2015F230","volume-title":"Proceedings of the federated conference on computer science and information systems","author":"M Skuza","year":"2015","unstructured":"Skuza M, Romanowski A (2015) Sentiment analysis of Twitter data within big data distributed environment for stock prediction. In: Ganzha M, Maciaszek L, Paprzycki M (eds) Proceedings of the federated conference on computer science and information systems. Annals of computer science and information systems, vol\u00a05. Polish Information Processing Society, Warsaw; IEEE, Los Alamitos, pp\u00a01349\u20131354. https:\/\/doi.org\/10.15439\/2015F230"},{"key":"244_CR5","doi-asserted-by":"publisher","first-page":"395","DOI":"10.1038\/nrg3208","volume":"13","author":"PB Jensen","year":"2012","unstructured":"Jensen PB, Jensen LJ, Brunak S (2012) Mining electronic health records: towards better research applications and clinical care. Nat Rev Genet 13:395\u2013405. https:\/\/doi.org\/10.1038\/nrg3208","journal-title":"Nat Rev Genet"},{"issue":"3","key":"244_CR6","doi-asserted-by":"publisher","first-page":"121","DOI":"10.1159\/000050784","volume":"46","author":"T Nakada","year":"2001","unstructured":"Nakada T, Fujii Y, Yoneoka Y, Kwee IL (2001) Planum temporale: where spoken and written language meet. Eur Neurol 46(3):121\u2013125. https:\/\/doi.org\/10.1159\/000050784","journal-title":"Eur Neurol"},{"issue":"3","key":"244_CR7","doi-asserted-by":"publisher","first-page":"225","DOI":"10.1016\/0885-2308(92)90019-Z","volume":"6","author":"J Kupiec","year":"1992","unstructured":"Kupiec J (1992) Robust part-of-speech tagging using a hidden Markov model. Comput Speech Lang 6(3):225\u2013242. https:\/\/doi.org\/10.1016\/0885-2308(92)90019-Z","journal-title":"Comput Speech Lang"},{"key":"244_CR8","series-title":"Lecture notes in morphogenesis","doi-asserted-by":"publisher","first-page":"143","DOI":"10.1007\/978-3-319-24403-7_9","volume-title":"Creativity and universality in language","author":"E Stamatatos","year":"2016","unstructured":"Stamatatos E (2016) Universality of stylistic traits in texts. In: Esposti MD, Altmann EG, Pachet F (eds) Creativity and universality in language. Lecture notes in morphogenesis. Springer, Cham, pp\u00a0143\u2013155. https:\/\/doi.org\/10.1007\/978-3-319-24403-7_9"},{"issue":"3","key":"244_CR9","doi-asserted-by":"crossref","first-page":"538","DOI":"10.1002\/asi.21001","volume":"60","author":"E Stamatatos","year":"2009","unstructured":"Stamatatos E (2009) A survey of modern authorship attribution methods. J Am Soc Inf Sci Technol 60(3):538\u2013556","journal-title":"J Am Soc Inf Sci Technol"},{"key":"244_CR10","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1201\/9781315181080-4","volume-title":"Feature engineering for machine learning and data analytics","author":"BD Fulcher","year":"2018","unstructured":"Fulcher BD (2018) Feature-based time-series analysis. In: Dong G, Liu H (eds) Feature engineering for machine learning and data analytics. Taylor & Francis, Boca Raton, pp\u00a087\u2013116"},{"key":"244_CR11","unstructured":"Christ M, Kempa-Liehr AW, Feindt M (2016) Distributed and parallel time series feature extraction for industrial big data applications. arXiv:1610.07717v1"},{"key":"244_CR12","doi-asserted-by":"crossref","first-page":"72","DOI":"10.1016\/j.neucom.2018.03.067","volume":"307","author":"M Christ","year":"2018","unstructured":"Christ M, Braun N, Neuffer J, Kempa-Liehr AW (2018) Time Series FeatuRe Extraction on basis of Scalable Hypothesis tests (tsfresh\u2014a Python package). Neurocomputing 307:72\u201377","journal-title":"Neurocomputing"},{"issue":"2","key":"244_CR13","doi-asserted-by":"publisher","first-page":"808","DOI":"10.1016\/j.physa.2006.02.042","volume":"370","author":"K Kosmidis","year":"2006","unstructured":"Kosmidis K, Kalampokis A, Argyrakis P (2006) Language time series analysis. Phys A, Stat Mech Appl 370(2):808\u2013816. https:\/\/doi.org\/10.1016\/j.physa.2006.02.042","journal-title":"Phys A, Stat Mech Appl"},{"key":"244_CR14","doi-asserted-by":"publisher","first-page":"257","DOI":"10.1146\/annurev-statistics-041715-033624","volume":"3","author":"J-L Wang","year":"2016","unstructured":"Wang J-L, Chiou J-M, M\u00fcller H-G (2016) Functional data analysis. Annu Rev Stat Appl 3:257\u2013295. https:\/\/doi.org\/10.1146\/annurev-statistics-041715-033624","journal-title":"Annu Rev Stat Appl"},{"issue":"1","key":"244_CR15","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/BF00054024","volume":"30","author":"FJ Tweedie","year":"1996","unstructured":"Tweedie FJ, Singh S, Holmes DI (1996) Neural network applications in stylometry: The Federalist Papers. Comput Humanit 30(1):1\u201310","journal-title":"Comput Humanit"},{"issue":"2","key":"244_CR16","doi-asserted-by":"publisher","first-page":"215","DOI":"10.1093\/llc\/fqq001","volume":"25","author":"ML Jockers","year":"2010","unstructured":"Jockers ML, Witten DM (2010) A comparative study of machine learning methods for authorship attribution. Lit Linguist Comput 25(2):215\u2013223. https:\/\/doi.org\/10.1093\/llc\/fqq001","journal-title":"Lit Linguist Comput"},{"key":"244_CR17","unstructured":"Kaggle (2018) Spooky author identification. https:\/\/www.kaggle.com\/c\/spooky-author-identification\/data"},{"issue":"1","key":"244_CR18","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-015-0015-2","volume":"2","author":"X Fang","year":"2015","unstructured":"Fang X, Zhan J (2015) Sentiment analysis using product review data. J Big Data 2(1):5. https:\/\/doi.org\/10.1186\/s40537-015-0015-2","journal-title":"J Big Data"},{"issue":"10","key":"244_CR19","doi-asserted-by":"publisher","first-page":"2513","DOI":"10.1016\/j.cor.2004.03.016","volume":"32","author":"W Huang","year":"2005","unstructured":"Huang W, Nakamori Y, Wang S-Y (2005) Forecasting stock market movement direction with support vector machine. Comput Oper Res 32(10):2513\u20132522. https:\/\/doi.org\/10.1016\/j.cor.2004.03.016","journal-title":"Comput Oper Res"},{"key":"244_CR20","doi-asserted-by":"publisher","first-page":"915","DOI":"10.1016\/j.asoc.2017.09.027","volume":"62","author":"A Ignatov","year":"2018","unstructured":"Ignatov A (2018) Real-time human activity recognition from accelerometer data using convolutional neural networks. Appl Soft Comput 62:915\u2013922. https:\/\/doi.org\/10.1016\/j.asoc.2017.09.027","journal-title":"Appl Soft Comput"},{"key":"244_CR21","series-title":"Springer series in statistics","doi-asserted-by":"publisher","DOI":"10.1007\/b98888","volume-title":"Functional data analysis","author":"JO Ramsay","year":"2005","unstructured":"Ramsay JO, Silverman BW (2005) Functional data analysis, 2nd edn. Springer series in statistics. Springer, Berlin. https:\/\/doi.org\/10.1007\/b98888","edition":"2"},{"key":"244_CR22","doi-asserted-by":"publisher","first-page":"606","DOI":"10.1007\/s10618-016-0483-9","volume":"31","author":"A Bagnall","year":"2017","unstructured":"Bagnall A, Lines J, Bostrom A, Large J, Keogh E (2017) The great time series classification bake off: a review and experimental evaluation of recent algorithmic advances. Data Min Knowl Discov 31:606\u2013660 https:\/\/doi.org\/10.1007\/s10618-016-0483-9","journal-title":"Data Min Knowl Discov"},{"issue":"4","key":"244_CR23","doi-asserted-by":"crossref","first-page":"451","DOI":"10.1142\/S0218348X02001257","volume":"10","author":"MA Montemurro","year":"2002","unstructured":"Montemurro MA, Pury PA (2002) Long-range fractal correlations in literary corpora. Fractals 10(4):451\u2013461","journal-title":"Fractals"},{"issue":"3","key":"244_CR24","doi-asserted-by":"crossref","DOI":"10.1103\/PhysRevE.86.031108","volume":"86","author":"M Ausloos","year":"2012","unstructured":"Ausloos M (2012) Generalized Hurst exponent and multifractal function of original and translated texts mapped into frequency and length time series. Phys Rev E 86(3):031108","journal-title":"Phys Rev E"},{"issue":"9","key":"244_CR25","doi-asserted-by":"crossref","DOI":"10.1142\/S0218127412502239","volume":"22","author":"M Kalimeri","year":"2012","unstructured":"Kalimeri M, Constantoudis V, Papadimitriou C, Karamanos K, Diakonos FK, Papageorgiou H (2012) Entropy analysis of word-length series of natural language texts: effects of text language and genre. Int J Bifurc Chaos 22(9):1250223","journal-title":"Int J Bifurc Chaos"},{"issue":"11","key":"244_CR26","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pone.0164658","volume":"11","author":"K Tanaka-Ishii","year":"2016","unstructured":"Tanaka-Ishii K, Bunde A (2016) Long-range memory in literary texts: on the universal clustering of the rare words. PLoS ONE 11(11):e0164658","journal-title":"PLoS ONE"},{"issue":"214","key":"244_CR27","doi-asserted-by":"crossref","first-page":"237","DOI":"10.1126\/science.ns-9.214S.237","volume":"9","author":"TC Mendenhall","year":"1887","unstructured":"Mendenhall TC (1887) The characteristic curves of composition. Science 9(214):237\u2013249","journal-title":"Science"},{"issue":"1","key":"244_CR28","first-page":"1","volume":"4","author":"CE Chaski","year":"2005","unstructured":"Chaski CE (2005) Who\u2019s at the keyboard? Authorship attribution in digital evidence investigations. Int J Digit Evid 4(1):1\u201313","journal-title":"Int J Digit Evid"},{"issue":"4","key":"244_CR29","doi-asserted-by":"crossref","first-page":"471","DOI":"10.1162\/089120100750105920","volume":"26","author":"E Stamatatos","year":"2000","unstructured":"Stamatatos E, Fakotakis N, Kokkinakis G (2000) Automatic text categorization in terms of genre and author. Comput Linguist 26(4):471\u2013495","journal-title":"Comput Linguist"},{"issue":"2","key":"244_CR30","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1093\/llc\/19.2.221","volume":"19","author":"G Tambouratzis","year":"2004","unstructured":"Tambouratzis G, Markantonatou S, Hairetakis N, Vassiliou M, Carayannis G, Tambouratzis D (2004) Discriminating the registers and styles in the modern Greek language\u2014part 2: extending the feature vector to optimize author discrimination. Lit Linguist Comput 19(2):221\u2013242","journal-title":"Lit Linguist Comput"},{"issue":"1\u20132","key":"244_CR31","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1023\/A:1023824908771","volume":"19","author":"J Diederich","year":"2003","unstructured":"Diederich J, Kindermann J, Leopold E, Paass G (2003) Authorship attribution with support vector machines. Appl Intell 19(1\u20132):109\u2013123","journal-title":"Appl Intell"},{"issue":"4","key":"244_CR32","doi-asserted-by":"crossref","first-page":"76","DOI":"10.1145\/1121949.1121951","volume":"49","author":"J Li","year":"2006","unstructured":"Li J, Zheng R, Chen H (2006) From fingerprint to writeprint. Commun ACM 49(4):76\u201382","journal-title":"Commun ACM"},{"key":"244_CR33","first-page":"482","volume-title":"Proceedings of the 2006 conference on empirical methods in natural language processing","author":"C Sanderson","year":"2006","unstructured":"Sanderson C, Guenter S (2006) Short text authorship attribution via sequence kernels, Markov chains and author unmasking: an investigation. In: Proceedings of the 2006 conference on empirical methods in natural language processing. Association for Computational Linguistics, Stroudsburg, pp\u00a0482\u2013491"},{"key":"244_CR34","first-page":"969","volume-title":"International conference on natural language processing","author":"\u00d6 Uzuner","year":"2005","unstructured":"Uzuner \u00d6, Katz B (2005) A comparative study of language models for book and author recognition. In: International conference on natural language processing. Springer, Berlin, pp\u00a0969\u2013980"},{"key":"244_CR35","doi-asserted-by":"crossref","first-page":"4551","DOI":"10.1109\/ICMLC.2006.258376","volume-title":"2006 international conference on machine learning and cybernetics","author":"F Khosmood","year":"2006","unstructured":"Khosmood F, Levinson R (2006) Toward unification of source attribution processes and techniques. In: 2006 international conference on machine learning and cybernetics, pp\u00a04551\u20134556"},{"issue":"4","key":"244_CR36","doi-asserted-by":"crossref","first-page":"203","DOI":"10.1093\/llc\/8.4.203","volume":"8","author":"RA Matthews","year":"1993","unstructured":"Matthews RA, Merriam TV (1993) Neural computation in stylometry I: an application to the works of Shakespeare and Fletcher. Lit Linguist Comput 8(4):203\u2013209","journal-title":"Lit Linguist Comput"},{"key":"244_CR37","first-page":"149","volume-title":"Computational linguistics in the Netherlands 2004: selected papers from the fifteenth CLIN meeting","author":"K Luyckx","year":"2005","unstructured":"Luyckx K, Daelemans W (2005) Shallow text analysis and machine learning for authorship attribution. In: Computational linguistics in the Netherlands 2004: selected papers from the fifteenth CLIN meeting. LOT, Utrecht, pp\u00a0149\u2013160"},{"issue":"5","key":"244_CR38","doi-asserted-by":"crossref","first-page":"823","DOI":"10.1142\/S0218213006002965","volume":"15","author":"E Stamatatos","year":"2006","unstructured":"Stamatatos E (2006) Authorship attribution based on feature set subspacing ensembles. Int J Artif Intell Tools 15(5):823\u2013838","journal-title":"Int J Artif Intell Tools"},{"issue":"4","key":"244_CR39","doi-asserted-by":"crossref","first-page":"405","DOI":"10.1093\/llc\/fqm023","volume":"22","author":"G Hirst","year":"2007","unstructured":"Hirst G, Feiguina O (2007) Bigrams of syntactic labels for authorship discrimination of short texts. Lit Linguist Comput 22(4):405\u2013417","journal-title":"Lit Linguist Comput"},{"key":"244_CR40","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-5256-6","volume-title":"Applied Bayesian and classical inference. The case of The Federalist Papers","author":"F Mosteller","year":"1984","unstructured":"Mosteller F, Wallace DL (1984) Applied Bayesian and classical inference. The case of The Federalist Papers, 2nd edn. Springer, New York. https:\/\/doi.org\/10.1007\/978-1-4612-5256-6","edition":"2"},{"key":"244_CR41","doi-asserted-by":"crossref","first-page":"107","DOI":"10.1109\/TCYB.2017.2766189","volume":"49","author":"SH Ding","year":"2017","unstructured":"Ding SH, Fung BC, Iqbal F, Cheung WK (2017) Learning stylometric representations for authorship analysis. IEEE Trans Cybern 49:107\u2013121","journal-title":"IEEE Trans Cybern"},{"key":"244_CR42","doi-asserted-by":"crossref","DOI":"10.3389\/fpsyg.2018.00289","volume":"9","author":"D Kernot","year":"2018","unstructured":"Kernot D, Bossomaier T, Bradbury R (2018) Using Shakespeare\u2019s sotto voce to determine true identity from text. Front Psychol 9:289","journal-title":"Front Psychol"},{"issue":"7","key":"244_CR43","doi-asserted-by":"crossref","first-page":"2429","DOI":"10.1016\/j.physa.2011.12.011","volume":"391","author":"A Mehri","year":"2012","unstructured":"Mehri A, Darooneh AH, Shariati A (2012) The complex networks approach for authorship attribution of books. Phys A, Stat Mech Appl 391(7):2429\u20132437","journal-title":"Phys A, Stat Mech Appl"},{"key":"244_CR44","doi-asserted-by":"crossref","first-page":"49","DOI":"10.1016\/j.physa.2017.12.054","volume":"495","author":"C Akimushkin","year":"2018","unstructured":"Akimushkin C, Amancio DR, Oliveira ON Jr (2018) On the role of words in the network structure of texts: application to authorship attribution. Phys A, Stat Mech Appl 495:49\u201358","journal-title":"Phys A, Stat Mech Appl"},{"issue":"3","key":"244_CR45","doi-asserted-by":"crossref","DOI":"10.1371\/journal.pone.0193703","volume":"13","author":"J Machicao","year":"2018","unstructured":"Machicao J, Corr\u00eaa EA Jr, Miranda GH, Amancio DR, Bruno OM (2018) Authorship attribution based on life-like network automata. PLoS ONE 13(3):e0193703","journal-title":"PLoS ONE"},{"key":"244_CR46","series-title":"Springer proceedings in complexity","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1007\/978-3-319-73198-8_17","volume-title":"Complex networks IX","author":"Y Al Rozz","year":"2018","unstructured":"Al Rozz Y, Menezes R (2018) Author attribution using network motifs. In: Cornelius S, Coronges K, Goncalves B, Sinatra R, Vespignani A (eds) Complex networks IX. Springer proceedings in complexity, pp\u00a0199\u2013207"},{"key":"244_CR47","volume-title":"Working notes of CLEF 2018\u2014conference and labs of the evaluation forum","author":"M Kestemont","year":"2018","unstructured":"Kestemont M, Tschuggnall M, Stamatatos E, Daelemans W, Specht G, Stein B, Potthast M (2018) Overview of the author identification task at PAN-2018: cross-domain authorship attribution and style change detection. In: Working notes of CLEF 2018\u2014conference and labs of the evaluation forum"},{"key":"244_CR48","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1016\/j.physa.2016.03.082","volume":"457","author":"S Martin\u010di\u0107-Ip\u0161i\u0107","year":"2016","unstructured":"Martin\u010di\u0107-Ip\u0161i\u0107 S, Margan D, Me\u0161trovi\u0107 A (2016) Multilayer network of language: a unified framework for structural analysis of linguistic subsystems. Phys A, Stat Mech Appl 457:117\u2013128","journal-title":"Phys A, Stat Mech Appl"},{"issue":"5","key":"244_CR49","doi-asserted-by":"crossref","DOI":"10.1209\/0295-5075\/100\/58002","volume":"100","author":"DR Amancio","year":"2012","unstructured":"Amancio DR, Aluisio SM, Oliveira ON Jr, Costa LdF (2012) Complex networks analysis of language complexity. Europhys Lett 100(5):58002","journal-title":"Europhys Lett"},{"issue":"3","key":"244_CR50","doi-asserted-by":"publisher","first-page":"483","DOI":"10.1162\/coli_a_00325","volume":"44","author":"M Riedl","year":"2018","unstructured":"Riedl M, Biemann C (2018) Using semantics for granularities of tokenization. Comput Linguist 44(3):483\u2013524. https:\/\/doi.org\/10.1162\/coli_a_00325","journal-title":"Comput Linguist"},{"issue":"12","key":"244_CR51","doi-asserted-by":"publisher","first-page":"64","DOI":"10.1145\/2500499","volume":"56","author":"V Dhar","year":"2013","unstructured":"Dhar V (2013) Data science and prediction. Commun ACM 56(12):64\u201373. https:\/\/doi.org\/10.1145\/2500499","journal-title":"Commun ACM"},{"issue":"1","key":"244_CR52","doi-asserted-by":"crossref","first-page":"50","DOI":"10.1214\/aoms\/1177730491","volume":"18","author":"HB Mann","year":"1947","unstructured":"Mann HB, Whitney DR (1947) On a test of whether one of two random variables is stochastically larger than the other. Ann Math Stat 18(1):50\u201360","journal-title":"Ann Math Stat"},{"issue":"2","key":"244_CR53","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1214\/aoms\/1177729639","volume":"22","author":"EL Lehmann","year":"1951","unstructured":"Lehmann EL (1951) Consistency and unbiasedness of certain nonparametric tests. Ann Math Stat 22(2):165\u2013179","journal-title":"Ann Math Stat"},{"key":"244_CR54","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1214\/09-SS051","volume":"4","author":"MP Fay","year":"2010","unstructured":"Fay MP, Proschan MA (2010) Wilcoxon\u2013Mann\u2013Whitney or t-test? On assumptions for hypothesis tests and multiple interpretations of decision rules. Stat Surv 4:1\u201339. https:\/\/doi.org\/10.1214\/09-SS051","journal-title":"Stat Surv"},{"key":"244_CR55","doi-asserted-by":"crossref","first-page":"289","DOI":"10.1111\/j.2517-6161.1995.tb02031.x","volume":"57","author":"Y Benjamini","year":"1995","unstructured":"Benjamini Y, Hochberg Y (1995) Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc, Ser B, Methodol 57:289\u2013300","journal-title":"J R Stat Soc, Ser B, Methodol"},{"key":"244_CR56","doi-asserted-by":"crossref","first-page":"378","DOI":"10.1016\/j.physa.2014.07.063","volume":"414","author":"E Rodriguez","year":"2014","unstructured":"Rodriguez E, Aguilar-Cornejo M, Femat R, Alvarez-Ramirez J (2014) Scale and time dependence of serial correlations in word-length time series of written texts. Phys A, Stat Mech Appl 414:378\u2013386","journal-title":"Phys A, Stat Mech Appl"},{"issue":"11","key":"244_CR57","doi-asserted-by":"crossref","first-page":"7798","DOI":"10.3390\/e17117798","volume":"17","author":"L Guzm\u00e1n-Vargas","year":"2015","unstructured":"Guzm\u00e1n-Vargas L, Obreg\u00f3n-Quintana B, Aguilar-Vel\u00e1zquez D, Hern\u00e1ndez-P\u00e9rez R, Liebovitch LS (2015) Word-length correlations and memory in large texts: a visibility network analysis. Entropy 17(11):7798\u20137810","journal-title":"Entropy"},{"issue":"15","key":"244_CR58","volume":"30","author":"V Constantoudis","year":"2016","unstructured":"Constantoudis V, Kalimeri M, Diakonos F, Karamanos K, Papadimitriou C, Chatzigeorgiou M, Papageorgiou H (2016) Long-range correlations and burstiness in written texts: universal and language-specific aspects. Int J Mod Phys B 30(15):1541005","journal-title":"Int J Mod Phys B"},{"key":"244_CR59","first-page":"73","volume":"4","author":"N Pietraszewska","year":"2015","unstructured":"Pietraszewska N (2015) On the complexity of creole languages: the fractal approach. Acad J Mod Philol 4:73\u201380","journal-title":"Acad J Mod Philol"},{"issue":"34","key":"244_CR60","doi-asserted-by":"crossref","first-page":"3717","DOI":"10.1007\/s11434-011-4752-0","volume":"56","author":"W Deng","year":"2011","unstructured":"Deng W, Wang D, Li W, Wang QA (2011) English and Chinese language frequency time series analysis. Chin Sci Bull 56(34):3717\u20133722","journal-title":"Chin Sci Bull"},{"key":"244_CR61","series-title":"EBook","volume-title":"The Project Gutenberg EBook of The Federalist Papers","author":"A Hamilton","year":"1998","unstructured":"Hamilton A, Jay J, Madison J (1998) The Project Gutenberg EBook of The Federalist Papers. EBook, vol\u00a01404. Project Gutenberg Literary Archive Foundation, Salt Lake City. http:\/\/www.gutenberg.org\/ebooks\/1404"},{"key":"244_CR62","series-title":"EBook","volume-title":"Frankenstein; or, the modern Prometheus","author":"MWG Shelley","year":"2018","unstructured":"Shelley MWG (2018) Frankenstein; or, the modern Prometheus. EBook, vol\u00a084. Project Gutenberg Literary Archive Foundation, Salt Lake City. http:\/\/www.gutenberg.org\/files\/84\/84-h\/84-h.htm"},{"key":"244_CR63","volume-title":"Natural language processing with Python","author":"E Loper","year":"2015","unstructured":"Loper E, Klein E, Bird S (2015) Natural language processing with Python. University of Melbourne, Melbourne. http:\/\/www.nltk.org\/book"},{"key":"244_CR64","first-page":"2825","volume":"12","author":"F Pedregosa","year":"2011","unstructured":"Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, Blondel M, Prettenhofer P, Weiss R, Dubourg V, Vanderplas J, Passos A, Cournapeau D, Brucher M, Perrot M, Duchesnay E (2011) Scikit-learn: machine learning in Python. J Mach Learn Res 12:2825\u20132830","journal-title":"J Mach Learn Res"},{"key":"244_CR65","doi-asserted-by":"publisher","first-page":"195","DOI":"10.1007\/978-981-13-0872-7_15","volume-title":"Logistics, supply chain and financial predictive analytics","author":"P Kumar","year":"2019","unstructured":"Kumar P (2019) Copula functions and applications in engineering. In: Deep K, Jain M, Salhi S (eds) Logistics, supply chain and financial predictive analytics Springer, Singapore, pp\u00a0195\u2013209. https:\/\/doi.org\/10.1007\/978-981-13-0872-7_15"},{"key":"244_CR66","doi-asserted-by":"crossref","first-page":"56","DOI":"10.25080\/Majora-92bf1922-00a","volume-title":"Proceedings of the 9th Python in science conference","author":"W McKinney","year":"2010","unstructured":"McKinney W (2010) Data structures for statistical computing in Python. In: Proceedings of the 9th Python in science conference, pp\u00a056\u201361."},{"key":"244_CR67","doi-asserted-by":"crossref","first-page":"785","DOI":"10.1145\/2939672.2939785","volume-title":"Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining","author":"T Chen","year":"2016","unstructured":"Chen T, Guestrin C (2016) XGBoost: a scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. ACM, New York, pp\u00a0785\u2013794"},{"key":"244_CR68","volume-title":"Working notes of CLEF 2018\u2014conference and labs of the evaluation forum","author":"JE Cust\u00f3dio","year":"2018","unstructured":"Cust\u00f3dio JE, Paraboni I (2018) EACH-USP ensemble cross-domain authorship attribution. In: Working notes of CLEF 2018\u2014conference and labs of the evaluation forum"},{"issue":"2","key":"244_CR69","doi-asserted-by":"crossref","first-page":"573","DOI":"10.1037\/a0029146","volume":"142","author":"JK Kruschke","year":"2013","unstructured":"Kruschke JK (2013) Bayesian estimation supersedes the t test. J Exp Psychol Gen 142(2):573\u2013603","journal-title":"J Exp Psychol Gen"},{"issue":"6","key":"244_CR70","doi-asserted-by":"crossref","first-page":"80","DOI":"10.2307\/3001968","volume":"1","author":"F Wilcoxon","year":"1945","unstructured":"Wilcoxon F (1945) Individual comparisons by ranking methods. Biom Bull 1(6):80\u201383","journal-title":"Biom Bull"},{"key":"244_CR71","doi-asserted-by":"crossref","DOI":"10.7717\/peerj-cs.55","volume":"2","author":"J Salvatier","year":"2016","unstructured":"Salvatier J, Wiecki TV, Fonnesbeck C (2016) Probabilistic programming in Python using PyMC3. PeerJ Comput Sci 2:e55","journal-title":"PeerJ Comput Sci"},{"key":"244_CR72","unstructured":"Wiecki T, Fonnesbeck C (2015) Bayesian estimation supersedes the t-test. https:\/\/docs.pymc.io\/notebooks\/BEST.html"},{"key":"244_CR73","doi-asserted-by":"publisher","first-page":"261","DOI":"10.1038\/s41592-019-0686-2","volume":"17","author":"P Virtanen","year":"2020","unstructured":"Virtanen P, Gommers R, Oliphant TE, Haberland M, Reddy T, Cournapeau D, Burovski E, Peterson P, Weckesser W, Bright J, van der Walt SJ, Brett M, Wilson J, Millman KJ, Mayorov N, Nelson ARJ, Jones E, Kern R, Larson E, Carey CJ, Polat \u0130, Feng Y, Moore EW, VanderPlas J, Laxalde D, Perktold J, Cimrman R, Henriksen I, Quintero EA, Harris CR, Archibald AM, Ribeiro AH, Pedregosa F, van Mulbregt P, Contributors (2020) SciPy 1.0: Fundamental algorithms for scientific computing in Python. Nat Methods 17:261\u2013272. https:\/\/doi.org\/10.1038\/s41592-019-0686-2","journal-title":"Nat Methods"},{"issue":"10","key":"244_CR74","doi-asserted-by":"crossref","first-page":"6567","DOI":"10.1073\/pnas.082099299","volume":"99","author":"R Tibshirani","year":"2002","unstructured":"Tibshirani R, Hastie T, Narasimhan B, Chu G (2002) Diagnosis of multiple cancer types by shrunken centroids of gene expression. Proc Natl Acad Sci USA 99(10):6567\u20136572.","journal-title":"Proc Natl Acad Sci USA"},{"issue":"1","key":"244_CR75","doi-asserted-by":"crossref","first-page":"104","DOI":"10.1214\/ss\/1056397488","volume":"18","author":"R Tibshirani","year":"2003","unstructured":"Tibshirani R, Hastie T, Narasimhan B, Chu G (2003) Class prediction by nearest shrunken centroids, with applications to DNA microarrays. Stat Sci 18(1):104\u2013117","journal-title":"Stat Sci"},{"key":"244_CR76","doi-asserted-by":"crossref","unstructured":"Calvo B, Santafe G (2015) scmamp: statistical comparison of multiple algorithms in multiple problems. R J (accepted for publication)","DOI":"10.32614\/RJ-2016-017"},{"issue":"5","key":"244_CR77","doi-asserted-by":"crossref","first-page":"5443","DOI":"10.1103\/PhysRevE.55.5443","volume":"55","author":"T Schreiber","year":"1997","unstructured":"Schreiber T, Schmitz A (1997) Discrimination power of measures for nonlinearity in a time series. Phys Rev E 55(5):5443\u20135447","journal-title":"Phys Rev E"}],"container-title":["EPJ Data Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1140\/epjds\/s13688-020-00244-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1140\/epjds\/s13688-020-00244-9\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1140\/epjds\/s13688-020-00244-9.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,8,12]],"date-time":"2024-08-12T21:38:57Z","timestamp":1723498737000},"score":1,"resource":{"primary":{"URL":"https:\/\/epjdatascience.springeropen.com\/articles\/10.1140\/epjds\/s13688-020-00244-9"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,31]]},"references-count":77,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2020,12]]}},"alternative-id":["244"],"URL":"https:\/\/doi.org\/10.1140\/epjds\/s13688-020-00244-9","relation":{},"ISSN":["2193-1127"],"issn-type":[{"value":"2193-1127","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,8,31]]},"assertion":[{"value":"28 November 2019","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 August 2020","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 August 2020","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare that they have no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}],"article-number":"26"}}