{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T16:43:51Z","timestamp":1779295431243,"version":"3.51.4"},"reference-count":75,"publisher":"Cambridge University Press (CUP)","issue":"1","license":[{"start":{"date-parts":[[2019,10,10]],"date-time":"2019-10-10T00:00:00Z","timestamp":1570665600000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":["cambridge.org"],"crossmark-restriction":true},"short-container-title":["Nat. Lang. Eng."],"published-print":{"date-parts":[[2021,1]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>We present a method for mining the web for text entered on mobile devices. Using searching, crawling, and parsing techniques, we locate text that can be reliably identified as originating from 300 mobile devices. This includes 341,000 sentences written on iPhones alone. Our data enables a richer understanding of how users type \u201cin the wild\u201d on their mobile devices. We compare text and error characteristics of different device types, such as touchscreen phones, phones with physical keyboards, and tablet computers. Using our mined data, we train language models and evaluate these models on mobile test data. A mixture model trained on our mined data, Twitter, blog, and forum data predicts mobile text better than baseline models. Using phone and smartwatch typing data from 135 users, we demonstrate our models improve the recognition accuracy and word predictions of a state-of-the-art touchscreen virtual keyboard decoder. Finally, we make our language models and mined dataset available to other researchers.<\/jats:p>","DOI":"10.1017\/s1351324919000548","type":"journal-article","created":{"date-parts":[[2019,10,10]],"date-time":"2019-10-10T20:01:29Z","timestamp":1570737689000},"page":"1-33","update-policy":"https:\/\/doi.org\/10.1017\/policypage","source":"Crossref","is-referenced-by-count":11,"title":["Mining, analyzing, and modeling text written on mobile devices"],"prefix":"10.1017","volume":"27","author":[{"given":"K.","family":"Vertanen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"P.O.","family":"Kristensson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2019,10,10]]},"reference":[{"key":"S1351324919000548_ref33","doi-asserted-by":"crossref","unstructured":"Kombrink, S. , Mikolov, T. , Karafi\u00e1t, M. and Burget, L. (2011). Recurrent neural network based language modeling in meeting recognition. In Proceedings of INTERSPEECH. ISCA, vol. 11, pp. 2877\u20132880.","DOI":"10.21437\/Interspeech.2011-720"},{"key":"S1351324919000548_ref55","doi-asserted-by":"publisher","DOI":"10.1145\/2598153.2598157"},{"key":"S1351324919000548_ref59","doi-asserted-by":"crossref","unstructured":"Stolcke, A. (2002). SRILM \u2013 an extensible language modeling toolkit. In Proceedings of INTERSPEECH. ISCA, pp. 901\u2013904.","DOI":"10.21437\/ICSLP.2002-303"},{"key":"S1351324919000548_ref68","unstructured":"Vertanen, K. and Kristensson, P.O. (2011a). The imagination of crowds: conversational AAC language modeling using crowdsourcing and large data sources. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing. Edinburgh, Scotland, UK: Association for Computational Linguistics, pp. 700\u2013711."},{"key":"S1351324919000548_ref26","doi-asserted-by":"publisher","DOI":"10.1080\/19312450709336664"},{"key":"S1351324919000548_ref46","unstructured":"Munro, R. and Manning, C.D. (2012). Short message communications: users, topics, and in-language processing. In: Proceedings of the 2nd ACM Symposium on Computing for Development. ACM."},{"key":"S1351324919000548_ref70","doi-asserted-by":"publisher","DOI":"10.1145\/2555691"},{"key":"S1351324919000548_ref49","doi-asserted-by":"publisher","DOI":"10.1145\/1978942.1979304"},{"key":"S1351324919000548_ref64","doi-asserted-by":"publisher","DOI":"10.1002\/asi.21416"},{"key":"S1351324919000548_ref35","doi-asserted-by":"publisher","DOI":"10.1145\/146370.146380"},{"key":"S1351324919000548_ref58","unstructured":"Stolcke, A. (1998). Entropy-based pruning of backoff language models. In Proceedings of DARPA Broadcast News Transcription and Understanding Workshop. Morgan Kaufmann, pp. 270\u2013274."},{"key":"S1351324919000548_ref34","doi-asserted-by":"publisher","DOI":"10.1145\/2166966.2166972"},{"key":"S1351324919000548_ref5","unstructured":"Brody, S. and Diakopoulos, N. (2011). Cooooooooooooooollllllllllllll!!!!!!!!!!!!!! Using word lengthening to detect sentiment in microblogs. In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing. Edinburgh, Scotland, UK: Association for Computational Linguistics, pp. 562\u2013570."},{"key":"S1351324919000548_ref4","doi-asserted-by":"publisher","DOI":"10.3115\/1075218.1075255"},{"key":"S1351324919000548_ref65","unstructured":"Tong, X. and Evans, D.A. (1996). A statistical approach to automatic OCR error correction in context. In Proceedings of the Fourth Workshop on Very Large Corpora. Association for Computational Linguistics, pp. 88\u2013100."},{"key":"S1351324919000548_ref19","doi-asserted-by":"crossref","unstructured":"Fowler, A. , Partridge, K. , Chelba, C. , Bi, X. , Ouyang, T. and Zhai, S. (2015). Effects of language modeling and its personalization on touchscreen typing performance. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. CHI\u201915. New York, NY, USA: ACM, 649\u2013658.","DOI":"10.1145\/2702123.2702503"},{"key":"S1351324919000548_ref6","doi-asserted-by":"publisher","DOI":"10.1145\/1322391.1322392"},{"key":"S1351324919000548_ref27","unstructured":"Heafield, K. (2011). KenLM: faster and smaller language model queries. In Proceedings of the EMNLP 2011 Sixth Workshop on Statistical Machine Translation. Association for Computational Linguistics, pp. 187\u2013197."},{"key":"S1351324919000548_ref60","unstructured":"Stolcke, A. , Yuret, D. and Madnani, N. (2010). SRILM-FAQ - Frequently Asked Questions About SRI LM Tools. http:\/\/www.speech.sri.com\/projects\/srilm\/manpages\/srilm-faq.7.html."},{"key":"S1351324919000548_ref29","unstructured":"Kalman, Y.M. and Gergle, D. (2009). Letter and punctuation mark repeats as cues in computer-mediated communication. In 95th Annual Meeting of the National Communication Association in Chicago, IL."},{"key":"S1351324919000548_ref67","unstructured":"Vertanen, K. , Fletcher, C. , Gaines, D. , Gould, J. and Kristensson, P.O. (2018). The impact of word, multiple word, and sentence input on virtual keyboard decoding performance. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. CHI\u201918. New York, NY, USA: ACM, pp. 626:1\u2013626:12."},{"key":"S1351324919000548_ref13","first-page":"299","article-title":"Creating a live, public short message service corpus: the NUS SMS corpus","volume":"47","author":"Chen","year":"2013","journal-title":"Language Resources and Evaluation"},{"key":"S1351324919000548_ref14","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-5470-6"},{"key":"S1351324919000548_ref47","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-74889-2_20"},{"key":"S1351324919000548_ref61","unstructured":"Stolcke, A. , Zheng, J. , Wang, W. and Abrash, V. (2011). SRILM at sixteen: update and outlook. In Proceedings of IEEE Automatic Speech Recognition and Understanding Workshop. ASRU\u201911. IEEE, vol. 5."},{"key":"S1351324919000548_ref37","doi-asserted-by":"publisher","DOI":"10.1007\/1-84628-248-9_22"},{"key":"S1351324919000548_ref51","doi-asserted-by":"publisher","DOI":"10.3115\/1628960.1628969"},{"key":"S1351324919000548_ref39","unstructured":"Lui, M. and Baldwin, T. (2012). langid.py: an off-the-shelf language identification tool. In Proceedings of the ACL 2012 System Demonstrations. ACL\u201912. Stroudsburg, PA, USA: Association for Computational Linguistics, pp. 25\u201330."},{"key":"S1351324919000548_ref21","doi-asserted-by":"publisher","DOI":"10.1145\/595576.595578"},{"key":"S1351324919000548_ref57","volume-title":"A USENET Corpus (2005\u20132009)","author":"Shaoul","year":"2009"},{"key":"S1351324919000548_ref71","doi-asserted-by":"publisher","DOI":"10.1145\/2702123.2702135"},{"key":"S1351324919000548_ref17","doi-asserted-by":"publisher","DOI":"10.1016\/j.csl.2014.09.005"},{"key":"S1351324919000548_ref53","doi-asserted-by":"publisher","DOI":"10.1016\/j.chb.2010.07.008"},{"key":"S1351324919000548_ref11","unstructured":"Chen, S.F. , Beeferman, D. and Rosenfeld, R. (1998). Evaluation metrics for language models. In Proceedings of the DARPA Broadcast News Transcription and Understanding Workshop. Morgan Kaufmann, pp. 275\u2013280."},{"key":"S1351324919000548_ref54","unstructured":"Rosenfeld, R. (2000). Two decades of statistical language modeling: where do we go From here? In Proceedings of the IEEE. IEEE, vol. 88, pp. 1270\u20131278."},{"key":"S1351324919000548_ref36","unstructured":"Levenshtein, V.I. (1966). Binary codes capable of correcting deletions, insertions, and reversals. In Soviet Physics Doklady, vol. 10, pp. 707\u2013710. Available at https:\/\/nymity.ch\/sybilhunting\/pdf\/Levenshtein1966a.pdf"},{"key":"S1351324919000548_ref52","unstructured":"Renals, S. (2010). Recognition and understanding of meetings. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics. HLT\u201910. Stroudsburg, PA, USA: Association for Computational Linguistics, pp. 1\u20139."},{"key":"S1351324919000548_ref25","unstructured":"Han, B. and Baldwin, T. (2011). Lexical normalisation of short text messages: makn sens a #twitter. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies - Volume 1. HLT\u201911. Stroudsburg, PA, USA: Association for Computational Linguistics, pp. 368\u2013378."},{"key":"S1351324919000548_ref44","unstructured":"Munro, R. (2011). Subword and spatiotemporal models for identifying actionable information in Haitian Kreyol. In Proceedings of the Fifteenth Conference on Computational Natural Language Learning. CoNLL\u201911. Stroudsburg, PA, USA: Association for Computational Linguistics, pp. 68\u201377."},{"key":"S1351324919000548_ref50","unstructured":"Pauls, A. and Klein, D. (2011). Faster and smaller N-gram language models. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies - Volume 1. HLT\u201911. Stroudsburg, PA, USA: Association for Computational Linguistics, pp. 258\u2013267."},{"key":"S1351324919000548_ref7","unstructured":"Burton, K. , Java, A. and Soboroff, I. (2009). The ICWSM 2009 Spinn3r dataset. In: Proceedings of the 3rd Annual Conference on Weblogs and Social Media. ICWSM\u201909. Palo Alto, California, USA: AAAI."},{"key":"S1351324919000548_ref10","unstructured":"Chen, B. , Kuhn, R. , Foster, G. , Cherry, C. and Huang, F. (2016). Bilingual methods for adaptive training data selection for machine translation. In Proceedings of the Association for Machine Translation in the Americas. AMTA\u201916, pp. 93\u2013103."},{"key":"S1351324919000548_ref9","doi-asserted-by":"crossref","unstructured":"Chelba, C. , Brants, T. , Neveitt, W. and Xu, P. (2010). Study on interaction between entropy pruning and Kneser\u2013Ney smoothing. In Proceedings of INTERSPEECH. ISCA, pp. 2242\u20132245.","DOI":"10.21437\/Interspeech.2010-525"},{"key":"S1351324919000548_ref63","volume-title":"A Corpus Linguistics Study of SMS Text Messaging","author":"Tagg","year":"2009"},{"key":"S1351324919000548_ref74","doi-asserted-by":"publisher","DOI":"10.1016\/B978-012373591-1\/50003-6"},{"key":"S1351324919000548_ref30","doi-asserted-by":"publisher","DOI":"10.1109\/MC.2007.270"},{"key":"S1351324919000548_ref8","doi-asserted-by":"publisher","DOI":"10.3115\/981436.981458"},{"key":"S1351324919000548_ref43","unstructured":"Moore, R.C. and Lewis, W. (2010). Intelligent selection of language model training data. In Proceedings of the ACL 2010 Conference Short Papers. ACLShort\u201910. Stroudsburg, PA, USA: Association for Computational Linguistics, pp. 220\u2013224."},{"key":"S1351324919000548_ref12","doi-asserted-by":"publisher","DOI":"10.3115\/981863.981904"},{"key":"S1351324919000548_ref72","doi-asserted-by":"publisher","DOI":"10.1177\/089443930101900307"},{"key":"S1351324919000548_ref38","volume-title":"The Length of Text Messages and the Use of Predictive Texting: Who Uses It and How Much Do They Have to Say? TESOL","author":"Ling","year":"2007"},{"key":"S1351324919000548_ref69","doi-asserted-by":"crossref","unstructured":"Vertanen, K. and Kristensson, P.O. (2011b). A versatile dataset for text entry evaluations based on genuine mobile emails. In Proceedings of the 13th International Conference on Human Computer Interaction with Mobile Devices and Services. MobileHCI\u201911. New York, NY, USA: ACM, pp. 295\u2013298.","DOI":"10.1145\/2037373.2037418"},{"key":"S1351324919000548_ref66","doi-asserted-by":"publisher","DOI":"10.1145\/2414536.2414577"},{"key":"S1351324919000548_ref1","unstructured":"Baldwin, T. and Chai, J. (2012). Autonomous self-assessment of autocorrections: exploring text message dialogues. In Proceedings of the 2012 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Montr\u00e9al, Canada: Association for Computational Linguistics, pp. 710\u2013719."},{"key":"S1351324919000548_ref62","doi-asserted-by":"crossref","unstructured":"Strik, H. , Cucchiarini, C. and Kessens, J.M. (2001). Comparing the performance of two CSRs: how to determine the significance level of the differences. In Proceedings of INTERSPEECH. ISCA, pp. 2091\u20132094.","DOI":"10.21437\/Eurospeech.2001-493"},{"key":"S1351324919000548_ref20","doi-asserted-by":"publisher","DOI":"10.1145\/2487575.2488202"},{"key":"S1351324919000548_ref56","first-page":"14","article-title":"Do you smile with your nose? Stylistic variation in twitter emoticons","volume":"18","author":"Schnoebelen","year":"2012","journal-title":"University of Pennsylvania Working Papers in Linguistics"},{"key":"S1351324919000548_ref32","unstructured":"Koehn, P. (2004). Statistical significance tests for machine translation evaluation. In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing. Barcelona, Spain: Association for Computational Linguistics, pp. 388\u2013395."},{"key":"S1351324919000548_ref15","doi-asserted-by":"publisher","DOI":"10.3115\/1609067.1609084"},{"key":"S1351324919000548_ref2","doi-asserted-by":"crossref","unstructured":"Bell, P. , Yamamoto, H. , Swietojanski, P. , Wu, Y. , McInnes, F. , Hori, C. and Renals, S. (2013). A lecture transcription system combining neural network acoustic and Language Models. In Proceedings of INTERSPEECH. ISCA, pp. 3087\u20133091.","DOI":"10.21437\/Interspeech.2013-673"},{"key":"S1351324919000548_ref3","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2004.1326009"},{"key":"S1351324919000548_ref18","doi-asserted-by":"crossref","unstructured":"Devlin, J. , Zbib, R. , Huang, Z. , Lamar, T. , Schwartz, R.M. and Makhoul, J. (2014). Fast and robust neural network joint models for statistical machine translation. In Proceedings of the Conference on Computational Linguistics. ACL\u201914. Baltimore, USA: Association for Computational Linguistics, pp. 1370\u20131380.","DOI":"10.3115\/v1\/P14-1129"},{"key":"S1351324919000548_ref75","doi-asserted-by":"crossref","unstructured":"Yao, K. , Zweig, G. , Hwang, M.-Y. , Shi, Y. and Yu, D. (2013). Recurrent neural networks for language understanding. In Proceedings of INTERSPEECH. ISCA, pp. 2524\u20132528.","DOI":"10.21437\/Interspeech.2013-569"},{"key":"S1351324919000548_ref31","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-30115-8_22"},{"key":"S1351324919000548_ref42","doi-asserted-by":"crossref","unstructured":"Mikolov, T. , Karafi\u00e1t, M. , Burget, L. , Cernock\u00fd, J. and Khudanpur, S. (2010). Recurrent neural network based language model. In Proceedings of INTERSPEECH. ISCA, pp. 1045\u20131048.","DOI":"10.21437\/Interspeech.2010-343"},{"key":"S1351324919000548_ref23","doi-asserted-by":"publisher","DOI":"10.1145\/502716.502753"},{"key":"S1351324919000548_ref16","doi-asserted-by":"publisher","DOI":"10.1109\/2.60879"},{"key":"S1351324919000548_ref73","doi-asserted-by":"publisher","DOI":"10.1145\/354401.354427"},{"key":"S1351324919000548_ref24","doi-asserted-by":"publisher","DOI":"10.1145\/642611.642688"},{"key":"S1351324919000548_ref40","doi-asserted-by":"publisher","DOI":"10.1109\/RE.2015.7320414"},{"key":"S1351324919000548_ref41","doi-asserted-by":"crossref","unstructured":"Mikolov, T. , Deoras, A. , Kombrink, S. , Burget, L. and Cernock\u00fd, J. (2011). Empirical evaluation and combination of advanced language modeling techniques. In Proceedings of INTERSPEECH. ISCA, pp. 605\u2013608.","DOI":"10.21437\/Interspeech.2011-242"},{"key":"S1351324919000548_ref28","doi-asserted-by":"publisher","DOI":"10.1016\/0167-6393(90)90008-W"},{"key":"S1351324919000548_ref45","unstructured":"Munro, R. and Manning, C.D. (2010). Subword variation in text message classification. In Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics. Association for Computational Linguistics, pp. 510\u2013518."},{"key":"S1351324919000548_ref22","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.1989.266481"},{"key":"S1351324919000548_ref48","doi-asserted-by":"publisher","DOI":"10.1109\/ICMEW.2013.6618380"}],"container-title":["Natural Language Engineering"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S1351324919000548","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,10,1]],"date-time":"2022-10-01T10:52:07Z","timestamp":1664621527000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S1351324919000548\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,10]]},"references-count":75,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2021,1]]}},"alternative-id":["S1351324919000548"],"URL":"https:\/\/doi.org\/10.1017\/s1351324919000548","relation":{},"ISSN":["1351-3249","1469-8110"],"issn-type":[{"value":"1351-3249","type":"print"},{"value":"1469-8110","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,10,10]]},"assertion":[{"value":"\u00a9 Cambridge University Press 2019","name":"copyright","label":"Copyright","group":{"name":"copyright_and_licensing","label":"Copyright and Licensing"}}]}}