{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T04:05:54Z","timestamp":1760241954586,"version":"build-2065373602"},"reference-count":37,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2018,11,2]],"date-time":"2018-11-02T00:00:00Z","timestamp":1541116800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100002241","name":"Japan Science and Technology Agency","doi-asserted-by":"publisher","award":["JPMJPR14E5"],"award-info":[{"award-number":["JPMJPR14E5"]}],"id":[{"id":"10.13039\/501100002241","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100009028","name":"Research Institute of Science and Technology for Society","doi-asserted-by":"publisher","award":["HITE project 272"],"award-info":[{"award-number":["HITE project 272"]}],"id":[{"id":"10.13039\/501100009028","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Neural language models have drawn a lot of attention for their strong ability to predict natural language text. In this paper, we estimate the entropy rate of natural language with state-of-the-art neural language models. To obtain the estimate, we consider the cross entropy, a measure of the prediction accuracy of neural language models, under the theoretically ideal conditions that they are trained with an infinitely large dataset and receive an infinitely long context for prediction. We empirically verify that the effects of the two parameters, the training data size and context length, on the cross entropy consistently obey a power-law decay with a positive constant for two different state-of-the-art neural language models with different language datasets. Based on the verification, we obtained 1.12 bits per character for English by extrapolating the two parameters to infinity. This result suggests that the upper bound of the entropy rate of natural language is potentially smaller than the previously reported values.<\/jats:p>","DOI":"10.3390\/e20110839","type":"journal-article","created":{"date-parts":[[2018,11,5]],"date-time":"2018-11-05T04:26:39Z","timestamp":1541391999000},"page":"839","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["Cross Entropy of Neural Language Models at Infinity\u2014A New Bound of the Entropy Rate"],"prefix":"10.3390","volume":"20","author":[{"given":"Shuntaro","family":"Takahashi","sequence":"first","affiliation":[{"name":"Graduate School of Frontier Sciences, The University of Tokyo, Chiba 277-8561, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kumiko","family":"Tanaka-Ishii","sequence":"additional","affiliation":[{"name":"Research Center for Advanced Science and Technology, The University of Tokyo, Tokyo 153-0041, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2018,11,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1735","DOI":"10.1162\/neco.1997.9.8.1735","article-title":"Long short-term memory","volume":"9","author":"Hochreiter","year":"1997","journal-title":"Neural Comput."},{"key":"ref_2","unstructured":"Pascanu, R., Mikolov, T., and Bengio, Y. (2013, January 16\u201321). On the Difficulty of Training Recurrent Neural Networks. Proceedings of the 30th International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_3","first-page":"1929","article-title":"Dropout: A Simple Way to Prevent Neural Networks from Overfitting","volume":"15","author":"Srivastava","year":"2014","journal-title":"J. Mach. Learn. Res."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Mikolov, T., Karafi\u00e1t, M., Burget, L., Cernock\u00fd, J., and Khudanpur, S. (2010, January 26\u201330). Recurrent neural network based language model. Proceedings of the 11th Annual Conference of the International Speech Communication Association, Chiba, Japan.","DOI":"10.21437\/Interspeech.2010-343"},{"key":"ref_5","unstructured":"Zilly, J.G., Srivastava, R.K., Koutn\u00edk, J., and Schmidhuber, J. (arXiv, 2016). Recurrent highway networks, arXiv."},{"key":"ref_6","unstructured":"Melis, G., Dyer, C., and Blunsom, P. (May, January 30). On the state of the art of evaluation in neural language models. Proceedings of the 6th International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_7","unstructured":"Merity, S., Keskar, N.S., and Socher, R. (May, January 30). Regularizing and optimizing LSTM language models. Proceedings of the 6th International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_8","unstructured":"Yang, Z., Dai, Z., Salakhutdinov, R., and Cohen, W.W. (May, January 30). Breaking the softmax bottleneck: A high-rank RNN language model. Proceedings of the 6th International Conference on Learning Representations, Vancouver, BC, Canada."},{"key":"ref_9","unstructured":"Merity, S., Keskar, N., and Socher, R. (arXiv, 2018). An Analysis of Neural Language Modeling at Multiple Scales, arXiv."},{"key":"ref_10","unstructured":"Han, Y., Jiao, J., Lee, C.Z., Weissman, T., Wu, Y., and Yu, T. (arXiv, 2018). Entropy Rate Estimation for Markov Chains with Large State Space, arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"701","DOI":"10.1162\/COLI_a_00239","article-title":"Computational linguistics and deep learning","volume":"41","author":"Manning","year":"2015","journal-title":"Comput. Linguist."},{"key":"ref_12","unstructured":"Devlin, J., Chang, M.W., Lee, K., and Toutanova, K. (arXiv, 2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Darmon, D. (2016). Specific Differential Entropy Rate Estimation for Continuous-Valued Time Series. Entropy, 18.","DOI":"10.3390\/e18050190"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Bentz, C., Alikaniotis, D., Cysouw, M., and Ferrer-i Cancho, R. (2017). The Entropy of Words\u2014Learnability and Expressivity across More than 1000 Languages. Entropy, 19.","DOI":"10.20944\/preprints201704.0180.v1"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"50","DOI":"10.1002\/j.1538-7305.1951.tb01366.x","article-title":"Prediction and entropy of printed English","volume":"30","author":"Shannon","year":"1951","journal-title":"Bell Syst. Tech. J."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"413","DOI":"10.1109\/TIT.1978.1055912","article-title":"A Convergent Gambling Estimate of the Entropy of English","volume":"24","author":"Cover","year":"1978","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_17","first-page":"31","article-title":"An Estimate of an Upper Bound for the Entropy of English","volume":"18","author":"Brown","year":"1992","journal-title":"Comput. Linguist."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"414","DOI":"10.1063\/1.166191","article-title":"Entropy estimation of symbol sequences","volume":"6","author":"Grassberger","year":"1996","journal-title":"Chaos"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Takahira, R., Tanaka-Ishii, K., and D\u0119bowski, \u0141. (2016). Entropy Rate Estimates for Natural Language\u2014A New Extrapolation of Compressed Large-Scale Corpora. Entropy, 18.","DOI":"10.3390\/e18100364"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"243","DOI":"10.1515\/FREQ.1990.44.9-10.243","article-title":"Der bekannte Grenzwert der redundanzfreien Information in Texten\u2014eine Fehlinterpretation der Shannonschen Experimente?","volume":"44","author":"Hilberg","year":"1990","journal-title":"Frequenz"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"5903","DOI":"10.3390\/e17085903","article-title":"Maximal Repetitions in Written Texts: Finite Energy Hypothesis vs. Strong Hilberg Conjecture","volume":"17","year":"2015","journal-title":"Entropy"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"1260","DOI":"10.1016\/j.spl.2009.01.016","article-title":"A general definition of conditional information and its application to ergodic decomposition","volume":"79","year":"2009","journal-title":"Stat. Probabil. Lett."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1109\/18.179344","article-title":"Entropy and data compression schemes","volume":"39","author":"Ornstein","year":"1993","journal-title":"IEEE Trans. Inf. Theory"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"191","DOI":"10.1209\/0295-5075\/14\/3\/001","article-title":"Entropy of Symbolic Sequences: The Role of Correlations","volume":"14","author":"Ebeling","year":"1991","journal-title":"Europhys Lett."},{"key":"ref_25","unstructured":"Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M., Yang, Y., and Zhou, Y. (arXiv, 2017). Deep Learning Scaling is Predictable, Empirically, arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Khandelwal, U., He, H., Qi, P., and Jurafsky, D. (2018, January 15\u201320). Sharp Nearby, Fuzzy Far Away: How Neural Language Models Use Context. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia.","DOI":"10.18653\/v1\/P18-1027"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"400","DOI":"10.1109\/TASSP.1987.1165125","article-title":"Estimation of probabilities from sparse data for the language model component of a speech recognizer","volume":"35","author":"Katz","year":"1987","journal-title":"Trans. Acoust. Speech Signal Process."},{"key":"ref_28","unstructured":"Kneser, R., and Ney, H. (1995, January 9\u201312). Improved backing-off for m-gram language modeling. Proceedings of the International Conference on Acoustics, Speech, and Signal Processing, Detroit, MI, USA."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Chelba, C., Mikolov, T., Schuster, M., Ge, Q., Brants, T., Koehn, P., and Robinson, T. (arXiv, 2013). One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling, arXiv.","DOI":"10.21437\/Interspeech.2014-564"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"988","DOI":"10.1109\/72.788640","article-title":"An overview of statistical learning theory","volume":"10","author":"Vapnik","year":"1999","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"605","DOI":"10.1162\/neco.1992.4.4.605","article-title":"Four Types of Learning Curves","volume":"4","author":"Amari","year":"1992","journal-title":"Neural Comput."},{"key":"ref_32","unstructured":"Gal, Y., and Ghahramani, Z. (2016, January 5\u201310). A Theoretically Grounded Application of Dropout in Recurrent Neural Networks. Proceedings of the 30th International Conference on Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_33","unstructured":"Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G.S., Davis, A., Dean, J., and Devin, M. (2018, October 24). TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems. Available online: http:\/\/download.tensorflow.org\/paper\/whitepaper2015.pdf."},{"key":"ref_34","unstructured":"Srivastava, R.K., Greff, K., and Schmidhuber, J. (2015, January 7\u201312). Training Very Deep Networks. Proceedings of the 28th International Conference on Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_35","unstructured":"Sutskever, I., Martens, J., Dahl, G.E., and Hinton, G.E. (2013, January 16\u201321). On the importance of initialization and momentum in deep learning. Proceedings of the 30th International Conference on Machine Learning, Atlanta, GA, USA."},{"key":"ref_36","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017, January 9). Automatic differentiation in PyTorch. Proceedings of the Neural Information Processing Systems Autodiff Workshop, Long Beach, CA, USA."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"838","DOI":"10.1137\/0330046","article-title":"Acceleration of Stochastic Approximation by Averaging","volume":"30","author":"Polyak","year":"1992","journal-title":"SIAM J. Control Optim."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/20\/11\/839\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T15:27:43Z","timestamp":1760196463000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/20\/11\/839"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,11,2]]},"references-count":37,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2018,11]]}},"alternative-id":["e20110839"],"URL":"https:\/\/doi.org\/10.3390\/e20110839","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2018,11,2]]}}}