{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,28]],"date-time":"2026-06-28T05:06:52Z","timestamp":1782623212669,"version":"3.54.5"},"reference-count":38,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2018,3,6]],"date-time":"2018-03-06T00:00:00Z","timestamp":1520294400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U1664264"],"award-info":[{"award-number":["U1664264"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U1509203"],"award-info":[{"award-number":["U1509203"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61174114"],"award-info":[{"award-number":["61174114"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Information"],"abstract":"<jats:p>Recurrent neural networks (RNN) are efficient in modeling sequences for generation and classification, but their training is obstructed by the vanishing and exploding gradient issues. In this paper, we reformulate the RNN unit to learn the residual functions with reference to the hidden state instead of conventional gated mechanisms such as long short-term memory (LSTM) and the gated recurrent unit (GRU). The residual structure has two main highlights: firstly, it solves the gradient vanishing and exploding issues for large time-distributed scales; secondly, the residual structure promotes the optimizations for backward updates. In the experiments, we apply language modeling, emotion classification and polyphonic modeling to evaluate our layer compared with LSTM and GRU layers. The results show that our layer gives state-of-the-art performance, outperforms LSTM and GRU layers in terms of speed, and supports an accuracy competitive with that of the other methods.<\/jats:p>","DOI":"10.3390\/info9030056","type":"journal-article","created":{"date-parts":[[2018,3,6]],"date-time":"2018-03-06T12:16:27Z","timestamp":1520338587000},"page":"56","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":70,"title":["Residual Recurrent Neural Networks for Learning Sequential Representations"],"prefix":"10.3390","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8038-9834","authenticated-orcid":false,"given":"Boxuan","family":"Yue","sequence":"first","affiliation":[{"name":"State Key Laboratory of Industrial Control, Zhejiang University, No. 38 Zheda Road, Hangzhou 310000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junwei","family":"Fu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Industrial Control, Zhejiang University, No. 38 Zheda Road, Hangzhou 310000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Liang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Industrial Control, Zhejiang University, No. 38 Zheda Road, Hangzhou 310000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2018,3,6]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"82","DOI":"10.1109\/MSP.2012.2205597","article-title":"Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups","volume":"29","author":"Hinton","year":"2012","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1109\/TASL.2011.2109382","article-title":"Acoustic modeling using deep belief networks","volume":"20","author":"Mohamed","year":"2012","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"e012012","DOI":"10.1136\/bmjopen-2016-012012","article-title":"Natural language processing to extract symptoms of severe mental illness from clinical text: The clinical record interactive search comprehensive data extraction (cris-code) project","volume":"7","author":"Jackson","year":"2017","journal-title":"BMJ Open"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1016\/j.ijmedinf.2017.02.011","article-title":"Creation of a simple natural language processing tool to support an imaging utilization quality dashboard","volume":"101","author":"Swartz","year":"2017","journal-title":"Int. J. Med. Inform."},{"key":"ref_5","unstructured":"Sawaf, H. (2015). Automatic Machine Translation Using User Feedback. (US20150248457), U.S. Patent."},{"key":"ref_6","unstructured":"Sonoo, S., and Sumita, K. (2017). Machine Translation Apparatus, Machine Translation Method and Computer Program Product. (US20170091177A1), U.S. Patent."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Gallos, L.K., Potiguar, F.Q., Andrade, J.S., and Makse, H.A. (2013). Imdb network revisited: Unveiling fractal and modular properties from a typical small-world network. PLoS ONE, 8.","DOI":"10.1371\/annotation\/7ce29312-158e-49b2-b530-6aca07751cea"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Oghina, A., Breuss, M., Tsagkias, M., and Rijke, M.D. (2012). Predicting imdb movie ratings using social media. Advances in Information Retrieval, Proceedings of the European Conference on IR Research, ECIR 2012, Barcelona, Spain, 1\u20135 April 2012, Springer.","DOI":"10.1007\/978-3-642-28997-2_51"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1207\/s15516709cog1402_1","article-title":"Finding structure in time","volume":"14","author":"Elman","year":"1990","journal-title":"Cogn. Sci."},{"key":"ref_10","unstructured":"Jordan, M.I. (1986). Serial Order: A Parallel Distributed Processing Approach, University of California."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Bengio, Y., Boulanger-Lewandowski, N., and Pascanu, R. (2012, January 25\u201330). Advances in optimizing recurrent networks. Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing, Kyoto, Japan.","DOI":"10.1109\/ICASSP.2013.6639349"},{"key":"ref_12","unstructured":"Bengio, Y., Frasconi, P., and Simard, P. (April, January 28). The problem of learning long-term dependencies in recurrent networks. Proceedings of the IEEE International Conference on Neural Networks, San Francisco, CA, USA."},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1109\/72.279181","article-title":"Learning long-term dependencies with gradient descent is difficult","volume":"5","author":"Bengio","year":"1994","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_14","first-page":"337","article-title":"On the difficulty of training recurrent neural networks","volume":"52","author":"Gustavsson","year":"2013","journal-title":"Comput. Sci."},{"key":"ref_15","unstructured":"Sutskever, I. (2013). Training Recurrent Neural Networks. [Ph.D. Thesis, University of Toronto]."},{"key":"ref_16","first-page":"16","article-title":"Long short-term memory","volume":"9","author":"Sepp","year":"1997","journal-title":"Neural Comput."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cho, K., Merrienboer, B.V., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (arXiv, 2014). Learning phrase representations using rnn encoder-decoder for statistical machine translation, arXiv.","DOI":"10.3115\/v1\/D14-1179"},{"key":"ref_18","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (July, January 26). Deep residual learning for image recognition. Proceedings of the Computer Vision and Pattern Recognition, Caesars Palace, NE, USA."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 11\u201314). Identity mappings in deep residual networks. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","article-title":"Imagenet large scale visual recognition challenge","volume":"115","author":"Russakovsky","year":"2015","journal-title":"Int. J. Comput. Vis."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Ioffe, S., Vanhoucke, V., and Alemi, A. (2017, January 4\u20139). Inception-v4, inception-resnet and the impact of residual connections on learning. Proceedings of the Thirty-First AAAI Conference on Artificial Intelligence, San Francisco, CA, USA.","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"ref_22","unstructured":"Price, P.J. (1992, January 23\u201326). Evaluation of spoken language systems: The atis domain. Proceedings of the Workshop on Speech and Natural Language, Harriman, NY, USA."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"56","DOI":"10.1108\/err.1999.3.5.56.52","article-title":"The internet movie database (imdb)","volume":"3","author":"Lindsay","year":"1999","journal-title":"Electron. Resour. Rev."},{"key":"ref_24","unstructured":"Chung, J., Gulcehre, C., Cho, K.H., and Bengio, Y. (arXiv, 2014). Empirical evaluation of gated recurrent neural networks on sequence modeling, arXiv."},{"key":"ref_25","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 7\u20139). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the 32nd International Conference on Machine Learning, Lille, France."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1704","DOI":"10.1109\/TPAMI.2011.235","article-title":"Aggregating local image descriptors into compact codes","volume":"34","author":"Jegou","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Perronnin, F., and Dance, C. (2007, January 18\u201323). Fisher kernels on visual vocabularies for image categorization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR \u201907), Minneapolis, MN, USA.","DOI":"10.1109\/CVPR.2007.383266"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1109\/TPAMI.2010.57","article-title":"Product quantization for nearest neighbor search","volume":"33","author":"Douze","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_31","unstructured":"Belayadi, A., Ait-Gougam, L., and Mekideche-Chafa, F. (1996). Pattern Recognition and Neural Networks, Cambridge University Press."},{"key":"ref_32","unstructured":"Sermanet, P., Eigen, D., Zhang, X., Mathieu, M., Fergus, R., and Lecun, Y. (arXiv, 2014). Overfeat: Integrated recognition, localization and detection using convolutional networks, arXiv."},{"key":"ref_33","first-page":"524","article-title":"The cascade-correlation learning architecture","volume":"2","author":"Fahlman","year":"1991","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_34","unstructured":"Srivastava, R.K., Greff, K., and Schmidhuber, J. (arXiv, 2015). Highway networks, arXiv."},{"key":"ref_35","unstructured":"Cooijmans, T., Ballas, N., Laurent, C., G\u00fcl\u00e7ehre, \u00c7., and Courville, A. (arXiv, 2016). Recurrent batch normalization, arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Mesnil, G., He, X., Deng, L., and Bengio, Y. (2013, January 25\u201329). Investigation of recurrent-neural-network architectures and learning methods for spoken language understanding. Proceedings of the Interspeech Conference, Lyon, France.","DOI":"10.21437\/Interspeech.2013-596"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Graves, A. (2012). Supervised Sequence Labelling with Recurrent Neural Networks, Springer.","DOI":"10.1007\/978-3-642-24797-2"},{"key":"ref_38","first-page":"3981","article-title":"Modeling temporal dependencies in high-dimensional sequences: Application to polyphonic music generation and transcription","volume":"18","author":"Boulangerlewandowski","year":"2012","journal-title":"Chem. A Eur. J."}],"container-title":["Information"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2078-2489\/9\/3\/56\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T14:57:46Z","timestamp":1760194666000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2078-2489\/9\/3\/56"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,3,6]]},"references-count":38,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2018,3]]}},"alternative-id":["info9030056"],"URL":"https:\/\/doi.org\/10.3390\/info9030056","relation":{},"ISSN":["2078-2489"],"issn-type":[{"value":"2078-2489","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,3,6]]}}}