{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T05:57:48Z","timestamp":1783749468438,"version":"3.55.0"},"reference-count":57,"publisher":"MIT Press","issue":"12","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Neural Computation"],"published-print":{"date-parts":[[2017,12]]},"abstract":"<jats:p>Learning useful information across long time lags is a critical and difficult problem for temporal neural models in tasks such as language modeling. Existing architectures that address the issue are often complex and costly to train. The differential state framework (DSF) is a simple and high-performing design that unifies previously introduced gated neural models. DSF models maintain longer-term memory by learning to interpolate between a fast-changing data-driven representation and a slowly changing, implicitly stable state. Within the DSF framework, a new architecture is presented, the delta-RNN. This model requires hardly any more parameters than a classical, simple recurrent network. In language modeling at the word and character levels, the delta-RNN outperforms popular complex architectures, such as the long short-term memory (LSTM) and the gated recurrent unit (GRU), and, when regularized, performs comparably to several state-of-the-art baselines. At the subword level, the delta-RNN's performance is comparable to that of complex gated architectures.<\/jats:p>","DOI":"10.1162\/neco_a_01017","type":"journal-article","created":{"date-parts":[[2017,9,28]],"date-time":"2017-09-28T20:31:40Z","timestamp":1506630700000},"page":"3327-3352","source":"Crossref","is-referenced-by-count":39,"title":["Learning Simpler Language Models with the Differential State Framework"],"prefix":"10.1162","volume":"29","author":[{"given":"Alexander G.","family":"Ororbia II","sequence":"first","affiliation":[{"name":"College of Information Sciences and Technology, Pennsylvania State University, State College, PA 16802, U.S.A."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tomas","family":"Mikolov","sequence":"additional","affiliation":[{"name":"Facebook, New York, NY 10003, U.S.A."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David","family":"Reitter","sequence":"additional","affiliation":[{"name":"College of Information Sciences and Technology, Pennsylvania State University, State College, PA 16802, U.S.A."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"281","reference":[{"key":"B1","doi-asserted-by":"publisher","DOI":"10.1111\/1467-9280.00063"},{"key":"B2","author":"Ba J. L.","year":"2016","journal-title":"Layer normalization"},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.1002\/0470018860.s00254"},{"key":"B4","author":"Bahdanau D.","year":"2014","journal-title":"Neural machine translation by jointly learning to align and translate"},{"issue":"1","key":"B5","doi-asserted-by":"crossref","DOI":"10.16910\/jemr.2.1.1","volume":"2","author":"Boston M. F.","year":"2008","journal-title":"Journal of Eye Movement Research"},{"key":"B6","author":"Choudhury V.","year":"2015","journal-title":"Thought vectors: Bringing common sense to artificial intelligence"},{"key":"B7","author":"Chung J.","year":"2016","journal-title":"Hierarchical multiscale recurrent neural networks"},{"key":"B8","author":"Chung J.","year":"2014","journal-title":"Empirical evaluation of gated recurrent neural networks on sequence modeling"},{"key":"B9","first-page":"2067","author":"Chung J.","year":"2015","journal-title":"Proceedings of the International Conference on Machine Learning"},{"key":"B10","author":"Cooijmans T.","year":"2016","journal-title":"Recurrent batch normalization"},{"key":"B11","first-page":"14","volume-title":"Proceedings of the 14th Annual Conference of the Cognitive Science Society","author":"Das S.","year":"1992"},{"key":"B12","doi-asserted-by":"publisher","DOI":"10.1207\/s15516709cog1402_1"},{"key":"B13","first-page":"1019","volume-title":"Advances in neural information processing systems, 29","author":"Gal Y.","year":"2016"},{"key":"B14","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2000.861302"},{"key":"B15","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.1991.155350"},{"key":"B16","doi-asserted-by":"publisher","DOI":"10.1023\/A:1010884214864"},{"key":"B17","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1992.4.3.393"},{"key":"B18","doi-asserted-by":"publisher","DOI":"10.1109\/72.286928"},{"key":"B19","author":"Graves A.","year":"2013","journal-title":"Generating sequences with recurrent neural networks"},{"key":"B20","doi-asserted-by":"publisher","DOI":"10.1038\/nature20101"},{"key":"B21","author":"Gulcehre C.","year":"2017","journal-title":"Memory augmented neural networks with wormhole connections"},{"key":"B22","author":"Gulcehre C.","year":"2016","journal-title":"Noisy activation functions"},{"key":"B23","author":"Ha D.","year":"2016","journal-title":"Hypernetworks"},{"key":"B24","doi-asserted-by":"publisher","DOI":"10.3115\/1073336.1073357"},{"key":"B25","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"B26","doi-asserted-by":"publisher","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"B27","first-page":"473","volume-title":"Advances in neural information processing systems","volume":"9","author":"Hochreiter S.","year":"1997"},{"key":"B28","author":"Ioffe S.","year":"2015","journal-title":"Batch normalization: Accelerating deep network training by reducing internal covariate shift"},{"key":"B29","author":"Jernite Y.","year":"2016","journal-title":"Variable computation in recurrent neural networks"},{"key":"B30","first-page":"112","volume-title":"Attractor dynamics and parallelism in a connectionist sequential machine","author":"Jordan M. I.","year":"1990"},{"key":"B31","first-page":"190","volume-title":"Proceedings of the 28th International Conference on Neural Information Processing Systems","author":"Joulin A.","year":"2015"},{"key":"B32","first-page":"2342","author":"Jozefowicz R.","year":"2015","journal-title":"Proceedings of the 32nd International Conference on Machine Learning"},{"key":"B33","author":"Kingma D.","year":"2014","journal-title":"Adam: A method for stochastic optimization"},{"key":"B34","author":"Koutnik J.","year":"2014","journal-title":"A clockwork RNN"},{"key":"B35","author":"Krueger D.","year":"2016","journal-title":"Zoneout: Regularizing rnns by randomly preserving hidden activations"},{"key":"B36","author":"Le Q. V.","year":"2015","journal-title":"A simple way to initialize recurrent networks of rectified linear units"},{"key":"B37","doi-asserted-by":"publisher","DOI":"10.1016\/j.cognition.2007.05.006"},{"key":"B38","volume-title":"Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies","author":"Maas A. L.","year":"2011"},{"issue":"2","key":"B39","first-page":"313","volume":"19","author":"Marcus M. P.","year":"1993","journal-title":"Computational Linguistics"},{"key":"B40","author":"Mikolov T.","year":"2012","journal-title":"Statistical language models based on neural networks"},{"key":"B41","author":"Mikolov T.","year":"2014","journal-title":"Learning longer memory in recurrent neural networks"},{"key":"B42","first-page":"1045","author":"Mikolov T.","year":"2010","journal-title":"Proceedings of the 11th Annual Conference of the International Speech Communication Association"},{"key":"B43","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2011.5947611"},{"key":"B45","volume-title":"Neural net architectures for temporal sequence processing","author":"Mozer M. C.","year":"1993"},{"key":"B46","volume-title":"Bayesian learning for neural networks","author":"Neal R. M.","year":"2012"},{"key":"B47","first-page":"1310","author":"Pascanu R.","year":"2013","journal-title":"Proceedings of the 30th International Conference of Machine Learning"},{"key":"B48","doi-asserted-by":"publisher","DOI":"10.1137\/0330046"},{"key":"B49","author":"Serban I. V.","year":"2016","journal-title":"Piecewise latent variables for neural variational text processing"},{"issue":"1","key":"B50","first-page":"1929","volume":"15","author":"Srivastava N.","year":"2014","journal-title":"Journal of Machine Learning Research"},{"key":"B51","author":"Sukhbaatar S.","year":"2015","journal-title":"End-to-end memory networks"},{"key":"B52","doi-asserted-by":"publisher","DOI":"10.1007\/BFb0054003"},{"key":"B53","author":"Sundermeyer M.","year":"2016","journal-title":"Improvements in language and translation modeling"},{"key":"B54","doi-asserted-by":"publisher","DOI":"10.3115\/1620853.1620921"},{"key":"B55","author":"Wang T.","year":"2015","journal-title":"Larger-context language modelling"},{"key":"B56","author":"Weston J.","year":"2014","journal-title":"Memory networks"},{"key":"B58","author":"Zaremba W.","year":"2014","journal-title":"Recurrent neural network regularization"},{"key":"B59","doi-asserted-by":"publisher","DOI":"10.1007\/s11633-016-1006-2"}],"container-title":["Neural Computation"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mitpressjournals.org\/doi\/pdf\/10.1162\/neco_a_01017","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,8,3]],"date-time":"2022-08-03T16:59:49Z","timestamp":1659545989000},"score":1,"resource":{"primary":{"URL":"https:\/\/direct.mit.edu\/neco\/article\/29\/12\/3327-3352\/8317"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,12]]},"references-count":57,"journal-issue":{"issue":"12","published-print":{"date-parts":[[2017,12]]}},"alternative-id":["10.1162\/neco_a_01017"],"URL":"https:\/\/doi.org\/10.1162\/neco_a_01017","relation":{},"ISSN":["0899-7667","1530-888X"],"issn-type":[{"value":"0899-7667","type":"print"},{"value":"1530-888X","type":"electronic"}],"subject":[],"published":{"date-parts":[[2017,12]]}}}