{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:54:31Z","timestamp":1760237671258,"version":"build-2065373602"},"reference-count":48,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2020,6,10]],"date-time":"2020-06-10T00:00:00Z","timestamp":1591747200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key R\\&amp;D Program of China, Natural Science Foundation of China, Higher Education Innovation Project of Xinjiang Key Science and Technology Project of  Xinjiang,","award":["2017YFB1402101, (61663044, 61761041), No.2016A03007-1, XJEDU2017T002"],"award-info":[{"award-number":["2017YFB1402101, (61663044, 61761041), No.2016A03007-1, XJEDU2017T002"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Modeling the context of a target word is of fundamental importance in predicting the semantic label for slot filling task in Spoken Language Understanding (SLU). Although Recurrent Neural Network (RNN) has shown to successfully achieve the state-of-the-art results for SLU, and Bidirectional RNN is capable of obtaining further improvement by modeling information not only from the past, but also from the future, they only consider limited contextual information of the target word. In order to make the network deeper and hence obtain longer contextual information, we propose to use a multi-layer Time Delay Neural Network (TDNN), which is prevalent in current large vocabulary continuous speech recognition tasks. In particular, we use a TDNN with symmetric time delay offset. To make the stacked TDNN easily trained, residual structures and skip concatenation are adopted. In addition, we further improve the model by introducing ResTDNN-BiLSTM, which combines the advantages of both the residual TDNN and BiLSTM. Experiments on slot filling tasks on the Air Travel Information System (ATIS) and Snips benchmark datasets show the proposed SC-TDNN-C achieves state-of-the-art results without any additional knowledge and data resources. Finally, we review and compare slot filling results by using a variety of existing models and methods.<\/jats:p>","DOI":"10.3390\/sym12060993","type":"journal-article","created":{"date-parts":[[2020,6,16]],"date-time":"2020-06-16T00:50:49Z","timestamp":1592268649000},"page":"993","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Using Deep Time Delay Neural Network for Slot Filling in Spoken Language Understanding"],"prefix":"10.3390","volume":"12","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3269-0953","authenticated-orcid":false,"given":"Zhen","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6604-0951","authenticated-orcid":false,"given":"Hao","family":"Huang","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1617-5804","authenticated-orcid":false,"given":"Kai","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Information Science and Engineering, Xinjiang University, Urumqi 830046, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,6,10]]},"reference":[{"key":"ref_1","first-page":"2993","article-title":"A joint model of intent determination and slot filling for spoken language understanding","volume":"16","author":"Zhang","year":"2016","journal-title":"IJCAI"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Tur, G., and Deng, L. (2011). Intent determination and spoken utterance classification. Spoken Language Understanding: Systems for Extracting Semantic Information from Speech, Wiley.","DOI":"10.1002\/9781119992691"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"1207","DOI":"10.1109\/TASL.2008.2001106","article-title":"An integrative and discriminative technique for spoken utterance classification","volume":"16","author":"Yaman","year":"2008","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Mikolov, T., Karafi\u00e1t, M., Burget, L., \u010cernock\u1ef3, J., and Khudanpur, S. (2010, January 26\u201330). Recurrent neural network based language model. Proceedings of the Eleventh Annual Conference of the International Speech Communication Association, Chiba, Japan.","DOI":"10.21437\/Interspeech.2010-343"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Yao, K., Zweig, G., Hwang, M.-Y., Shi, Y., and Yu, D. (2013). Recurrent neural networks for language understanding. Interspeech, 2524\u20132528.","DOI":"10.21437\/Interspeech.2013-569"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Raymond, C., and Riccardi, G. (2007, January 27\u201331). Generative and discriminative algorithms for spoken language understanding. Proceedings of the Eighth Annual Conference of the International Speech Communication Association, Antwerp, Belgium.","DOI":"10.21437\/Interspeech.2007-448"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Wang, Y.-Y., and Acero, A. (2006, January 17\u201321). Discriminative models for spoken language understanding. Proceedings of the Ninth International Conference on Spoken Language Processing, Jeju, Korea.","DOI":"10.21437\/Interspeech.2006-608"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"30","DOI":"10.1109\/TASL.2011.2134090","article-title":"Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition","volume":"20","author":"Dahl","year":"2011","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Jelinek, F., Lafferty, J.D., and Mercer, R.L. (1992). Basic methods of probabilistic context free grammars. Speech Recognition and Understanding, Springer.","DOI":"10.1007\/978-3-642-76626-8_35"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Peng, B., and Yao, K. (2015). Recurrent neural networks with external memory for language understanding. arXiv.","DOI":"10.1007\/978-3-319-25207-0_3"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Xu, P., and Sarikaya, R. (2013, January 8\u201313). Convolutional neural network based triangular crf for joint intent detection and slot filling. Proceedings of the 2013 IEEE Workshop on Automatic Speech Recognition and Understanding, Olomouc, Czech Republic.","DOI":"10.1109\/ASRU.2013.6707709"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"530","DOI":"10.1109\/TASLP.2014.2383614","article-title":"Using recurrent neural networks for slot filling in spoken language understanding","volume":"23","author":"Mesnil","year":"2014","journal-title":"IEEE\/ACM Trans. Audio Speech Lang. Process."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Yao, K., Peng, B., Zhang, Y., Yu, D., Zweig, G., and Shi, Y. (2014, January 7\u201310). Spoken language understanding using long short-term memory neural networks. Proceedings of the 2014 IEEE Spoken Language Technology Workshop (SLT), South Lake Tahoe, NV, USA.","DOI":"10.1109\/SLT.2014.7078572"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1162\/neco.1989.1.1.39","article-title":"Modular construction of time-delay neural networks for speech recognition","volume":"1","author":"Waibel","year":"1989","journal-title":"Neural Comput."},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"328","DOI":"10.1109\/29.21701","article-title":"Phoneme recognition using time-delay neural networks","volume":"37","author":"Waibel","year":"1989","journal-title":"IEEE Trans. Acoustics Speech Signal Process."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Snyder, D., Garcia-Romero, D., Sell, G., Povey, D., and Khudanpur, S. (2018, January 15\u201320). X-vectors: Robust dnn embeddings for speaker recognition. Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Calgary, AB, Canada.","DOI":"10.1109\/ICASSP.2018.8461375"},{"key":"ref_17","unstructured":"Kershaw, D.J., Robinson, A.J., and Hochberg, M. (1996). Context-dependent classes in a hybrid recurrent network-hmm speech recognition system. Advances in Neural Information Processing Systems, Springer."},{"key":"ref_18","unstructured":"Povey, D., Ghoshal, A., Boulianne, G., Burget, L., Glembek, O., Goel, N., Hannemann, M., Motlicek, P., Qian, Y., and Schwarz, P. (2012). The Kaldi Speech Recognition Toolkit, Idiap."},{"key":"ref_19","unstructured":"Hochreiter, S., Bengio, Y., Frasconi, P., and Schmidhuber, J. (2001). Gradient Flow in Recurrent Nets: The Difficulty of Learning Long-term Dependencies, Wiley Press."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Peddinti, V., Povey, D., and Khudanpur, S. (2015, January 6\u201310). A time delay neural network architecture for efficient modeling of long temporal contexts. Proceedings of the Sixteenth Annual Conference of the International Speech Communication Association, Dresden, Germany.","DOI":"10.21437\/Interspeech.2015-647"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhong, Z., Li, J., Ma, L., Jiang, H., and Zhao, H. (2017, January 23\u201328). Deep residual networks for hyperspectral image classification. Proceedings of the 2017 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Fort Worth, TX, USA.","DOI":"10.1109\/IGARSS.2017.8127330"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Huang, G., Liu, Z., Maaten, L.v., and Weinberger, K.Q. (2017, January 21\u201326). Densely connected convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.243"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Mikolov, T., Kombrink, S., Burget, L., \u010cernock\u1ef3, J., and Khudanpur, S. (2011, January 22\u201327). Extensions of recurrent neural network language model. Proceedings of the 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Prague, Czech Republic.","DOI":"10.1109\/ICASSP.2011.5947611"},{"key":"ref_24","first-page":"79","article-title":"A statistical approach to machine translation","volume":"16","author":"Brown","year":"1990","journal-title":"Comput. Linguist."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Deng, L., Tur, G., He, X., and Hakkani-Tur, D. (2012, January 2\u20135). Use of kernel deep convex networks and end-to-end learning for spoken language understanding. Proceedings of the 2012 IEEE Spoken Language Technology Workshop (SLT), Miami, FL, USA.","DOI":"10.1109\/SLT.2012.6424224"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Tur, G., Deng, L., Hakkani-T\u00fcr, D., and He, X. (2012, January 25\u201330). Towards deeper understanding: Deep convex networks for semantic utterance classification. Proceedings of the 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Kyoto, Japan.","DOI":"10.1109\/ICASSP.2012.6289054"},{"key":"ref_27","first-page":"1137","article-title":"A neural probabilistic language model","volume":"3","author":"Bengio","year":"2003","journal-title":"J. Mach. Learn. Res."},{"key":"ref_28","unstructured":"Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013). Efficient estimation of word representations in vector space. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Collobert, R., and Weston, J. (2008). A unified architecture for natural language processing: Deep neural networks with multitask learning. Proceedings of the 25th International Conference on Machine Learning, ACM.","DOI":"10.1145\/1390156.1390177"},{"key":"ref_30","first-page":"2493","article-title":"Natural language processing (almost) from scratch","volume":"12","author":"Collobert","year":"2011","journal-title":"J. Machine Learn. Res."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Yin, W., and Sch\u00fctze, H. (2015, January 26\u201331). Multigrancnn: An architecture for general matching of text chunks on multiple levels of granularity. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), Beijing, China.","DOI":"10.3115\/v1\/P15-1007"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Kalchbrenner, N., Grefenstette, E., and Blunsom, P. (2014). A convolutional neural network for modelling sentences. arXiv.","DOI":"10.3115\/v1\/P14-1062"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Kim, Y. (2014). Convolutional neural networks for sentence classification. arXiv.","DOI":"10.3115\/v1\/D14-1181"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Hakkani-T\u00fcr, D., Tur, G., Celikyilmaz, A., Chen, Y.N., and Wang, Y.Y. (2016, January 8\u201312). Multi-domain joint semantic frame parsing using bidirectional rnn-lstm. Proceedings of the 17th Annual Meeting of the International Speech Communication Association (INTERSPEECH 2016), San Francisco, CA, USA.","DOI":"10.21437\/Interspeech.2016-402"},{"key":"ref_35","first-page":"685","article-title":"Attention-based recurrent neural network models for joint intent detection and slot filling","volume":"2016","author":"Liu","year":"2016","journal-title":"Interspeech"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Goo, C.-W., Gao, G., Hsu, Y.-K., Huo, C.-L., Chen, T.-C., Hsu, K.-W., and Chen, Y.-N. (2018, January 1\u20136). Slot-gated modeling for joint slot filling and intent prediction. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 2 (Short Papers), New Orleans, LA, USA.","DOI":"10.18653\/v1\/N18-2118"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Zhang, C., Li, Y., Du, N., Fan, W., and Yu, P.S. (2018). Joint slot filling and intent detection via capsule neural networks. arXiv.","DOI":"10.18653\/v1\/P19-1519"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Vukoti, V., and Raymond, C. (2019). Mining polysemous triplets with recurrent neural networks for spoken language understanding. Interspeech, 2019.","DOI":"10.21437\/Interspeech.2019-2977"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Price, P.J. (1990, January 24\u201327). Evaluation of spoken language systems: The atis domain. Proceedings of the third DARPA Speech and Natural Language Workshop, Hidden Valley, PA, USA.","DOI":"10.3115\/116580.116612"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Dahl, G.E., Sainath, T.N., and Hinton, G.E. (2013, January 26\u201331). Improving deep neural networks for lvcsr using rectified linear units and dropout. Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.","DOI":"10.1109\/ICASSP.2013.6639346"},{"key":"ref_41","unstructured":"Ioffe, S., and Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv."},{"key":"ref_42","unstructured":"Hinton, G.E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R.R. (2012). Improving neural networks by preventing co-adaptation of feature detectors. arXiv."},{"key":"ref_43","first-page":"2345","article-title":"Sequence-discriminative training of deep neural networks","volume":"2013","author":"Ghoshal","year":"2013","journal-title":"Interspeech"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1145\/321105.321114","article-title":"A technique for the numerical solution of certain integral equations of the first kind","volume":"9","author":"Phillips","year":"1962","journal-title":"J. ACM (JACM)"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1109\/72.279181","article-title":"Learning long-term dependencies with gradient descent is difficult","volume":"5","author":"Bengio","year":"1994","journal-title":"IEEE Trans. Neural Netw."},{"key":"ref_46","unstructured":"Glorot, X., and Bengio, Y. (2010, January 13\u201315). Understanding the difficulty of training deep feedforward neural networks. Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, Sardinia, Italy."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Hemphill, C.T., Godfrey, J.J., and Doddington, G.R. (1990, January 24\u201327). The atis spoken language systems pilot corpus. Proceedings of the Speech and Natural Language: Proceedings of a Workshop, Hidden Valley, PA, USA.","DOI":"10.3115\/116580.116613"},{"key":"ref_48","unstructured":"Coucke, A., Saade, A., Ball, A., Bluche, T., Caulier, A., Leroy, D., Doumouro, C., Gisselbrecht, T., Caltagirone, F., and Lavril, T. (2018). Snips voice platform: An embedded spoken language understanding system for private-by-design voice interfaces. arXiv."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/6\/993\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:37:33Z","timestamp":1760175453000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/12\/6\/993"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,10]]},"references-count":48,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2020,6]]}},"alternative-id":["sym12060993"],"URL":"https:\/\/doi.org\/10.3390\/sym12060993","relation":{},"ISSN":["2073-8994"],"issn-type":[{"type":"electronic","value":"2073-8994"}],"subject":[],"published":{"date-parts":[[2020,6,10]]}}}