{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:30:01Z","timestamp":1750221001239,"version":"3.41.0"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2019,8,31]],"date-time":"2019-08-31T00:00:00Z","timestamp":1567209600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2019,8,31]]},"abstract":"<jats:p>\n            Due to their special gating schemes, Long Short-Term Memory (LSTM) has shown greater potential to process complex sequential information than the traditional Recurrent Neural Network (RNN). The conventional LSTM, however, fails to take into consideration the impact of salient spatio-temporal dynamics present in the sequential input data. This problem was first addressed by the differential Recurrent Neural Network (dRNN), which uses a differential gating scheme known as Derivative of States (DoS). DoS uses higher orders of internal state derivatives to analyze the change in information gain originated from the salient motions between the successive frames. The weighted combination of several orders of DoS is then used to modulate the gates in dRNN. While each individual order of DoS is good at modeling a certain level of salient spatio-temporal sequences, the sum of all the orders of DoS could distort the detected motion patterns. To address this problem, we propose to control the LSTM gates via individual orders of DoS. To fully utilize the different orders of DoS, we further propose to stack multiple levels of LSTM cells in an increasing order of state derivatives. The proposed model progressively builds up the ability of the LSTM gates to detect salient dynamical patterns in deeper stacked layers modeling higher orders of DoS; thus, the proposed LSTM model is termed deep differential Recurrent Neural Network (\n            <jats:italic>d<\/jats:italic>\n            <jats:sup>2<\/jats:sup>\n            RNN). The effectiveness of the proposed model is demonstrated on three publicly available human activity datasets: NUS-HGA, Violent-Flows, and UCF101. The proposed model outperforms both LSTM and non-LSTM based state-of-the-art algorithms.\n          <\/jats:p>","DOI":"10.1145\/3337928","type":"journal-article","created":{"date-parts":[[2019,9,12]],"date-time":"2019-09-12T14:22:21Z","timestamp":1568298141000},"page":"1-21","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Rethinking the Combined and Individual Orders of Derivative of States for Differential Recurrent Neural Networks"],"prefix":"10.1145","volume":"15","author":[{"given":"Naifan","family":"Zhuang","sequence":"first","affiliation":[{"name":"University of Central Florida, Orlando, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Guo-Jun","family":"Qi","sequence":"additional","affiliation":[{"name":"University of Central Florida, Orlando, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"The Duc","family":"Kieu","sequence":"additional","affiliation":[{"name":"University of the West Indies, Trinidad and Tobago"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kien A.","family":"Hua","sequence":"additional","affiliation":[{"name":"University of Central Florida, Orlando, FL"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2019,9,12]]},"reference":[{"doi-asserted-by":"publisher","key":"e_1_2_1_1_1","DOI":"10.5555\/1889001.1889024"},{"key":"e_1_2_1_2_1","volume-title":"Neural machine translation by jointly learning to align and translate. Arxiv Preprint Arxiv:1409.0473","author":"Bahdanau Dzmitry","year":"2014","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2014. Neural machine translation by jointly learning to align and translate. Arxiv Preprint Arxiv:1409.0473 ( 2014 ). Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural machine translation by jointly learning to align and translate. Arxiv Preprint Arxiv:1409.0473 (2014)."},{"doi-asserted-by":"publisher","key":"e_1_2_1_3_1","DOI":"10.1109\/CVPR.2017.502"},{"doi-asserted-by":"publisher","key":"e_1_2_1_4_1","DOI":"10.1016\/j.neucom.2014.01.019"},{"doi-asserted-by":"publisher","key":"e_1_2_1_5_1","DOI":"10.1142\/S0218001415550071"},{"unstructured":"Fran\u00e7ois Chollet et al. 2015. Keras. Retrieved from https:\/\/github.com\/fchollet\/keras.  Fran\u00e7ois Chollet et al. 2015. Keras. Retrieved from https:\/\/github.com\/fchollet\/keras.","key":"e_1_2_1_6_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_7_1","DOI":"10.1145\/2393347.2396385"},{"doi-asserted-by":"crossref","unstructured":"Manuel P. Cu\u00e9llar Miguel Delgado and M.C. Pegalajar. 2007. An application of non-linear programming to train recurrent neural networks in time series prediction problems. In Enterprise Information Systems VII. Springer 95--102.  Manuel P. Cu\u00e9llar Miguel Delgado and M.C. Pegalajar. 2007. An application of non-linear programming to train recurrent neural networks in time series prediction problems. In Enterprise Information Systems VII. Springer 95--102.","key":"e_1_2_1_8_1","DOI":"10.1007\/978-1-4020-5347-4_11"},{"doi-asserted-by":"publisher","key":"e_1_2_1_9_1","DOI":"10.1109\/CVPR.2015.7298878"},{"volume-title":"An Introduction to Numerical Methods and Analysis","author":"Epperson James F.","unstructured":"James F. Epperson . 2013. An Introduction to Numerical Methods and Analysis . John Wiley 8 Sons. James F. Epperson. 2013. An Introduction to Numerical Methods and Analysis. John Wiley 8 Sons.","key":"e_1_2_1_10_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_11_1","DOI":"10.1006\/jcss.1997.1504"},{"key":"e_1_2_1_12_1","volume-title":"Multi-dimensional recurrent neural networks. CoRR abs\/0705.2011","author":"Graves Alex","year":"2007","unstructured":"Alex Graves , Santiago Fern\u00e1ndez , and J\u00fcrgen Schmidhuber . 2007. Multi-dimensional recurrent neural networks. CoRR abs\/0705.2011 ( 2007 ). arxiv:0705.2011 http:\/\/arxiv.org\/abs\/0705.2011. Alex Graves, Santiago Fern\u00e1ndez, and J\u00fcrgen Schmidhuber. 2007. Multi-dimensional recurrent neural networks. CoRR abs\/0705.2011 (2007). arxiv:0705.2011 http:\/\/arxiv.org\/abs\/0705.2011."},{"doi-asserted-by":"publisher","key":"e_1_2_1_13_1","DOI":"10.1109\/ICASSP.2013.6638947"},{"doi-asserted-by":"publisher","key":"e_1_2_1_14_1","DOI":"10.1109\/TNNLS.2016.2582924"},{"doi-asserted-by":"publisher","key":"e_1_2_1_15_1","DOI":"10.1109\/IJCNN.2013.6706797"},{"doi-asserted-by":"publisher","key":"e_1_2_1_16_1","DOI":"10.1109\/CVPRW.2012.6239348"},{"doi-asserted-by":"publisher","key":"e_1_2_1_17_1","DOI":"10.1109\/CVPR.2016.90"},{"doi-asserted-by":"publisher","key":"e_1_2_1_18_1","DOI":"10.1162\/neco.1997.9.8.1735"},{"key":"e_1_2_1_19_1","volume-title":"International Conference on Machine Learning. 1568--1577","author":"Hu Hao","year":"2017","unstructured":"Hao Hu and Guo-Jun Qi . 2017 . State-frequency memory recurrent neural networks . In International Conference on Machine Learning. 1568--1577 . Hao Hu and Guo-Jun Qi. 2017. State-frequency memory recurrent neural networks. In International Conference on Machine Learning. 1568--1577."},{"doi-asserted-by":"publisher","key":"e_1_2_1_20_1","DOI":"10.1109\/CVPR.2018.00745"},{"doi-asserted-by":"publisher","key":"e_1_2_1_21_1","DOI":"10.1145\/1459359.1459379"},{"doi-asserted-by":"publisher","key":"e_1_2_1_22_1","DOI":"10.3115\/v1\/D14-1080"},{"doi-asserted-by":"publisher","key":"e_1_2_1_23_1","DOI":"10.1109\/TMM.2018.2823900"},{"key":"e_1_2_1_24_1","volume-title":"Hua","author":"Joslyn Kevin","year":"2018","unstructured":"Kevin Joslyn , Naifan Zhuang , and Kien A . Hua . 2018 . Deep segment hash learning for music generation. Arxiv Preprint Arxiv :1805.12176 (2018). Kevin Joslyn, Naifan Zhuang, and Kien A. Hua. 2018. Deep segment hash learning for music generation. Arxiv Preprint Arxiv:1805.12176 (2018)."},{"key":"e_1_2_1_25_1","volume-title":"Grid long short-term memory. Arxiv Preprint Arxiv:1507.01526","author":"Kalchbrenner Nal","year":"2015","unstructured":"Nal Kalchbrenner , Ivo Danihelka , and Alex Graves . 2015. Grid long short-term memory. Arxiv Preprint Arxiv:1507.01526 ( 2015 ). Nal Kalchbrenner, Ivo Danihelka, and Alex Graves. 2015. Grid long short-term memory. Arxiv Preprint Arxiv:1507.01526 (2015)."},{"doi-asserted-by":"publisher","key":"e_1_2_1_26_1","DOI":"10.5244\/C.22.99"},{"doi-asserted-by":"publisher","key":"e_1_2_1_27_1","DOI":"10.1109\/TIP.2016.2624140"},{"key":"e_1_2_1_28_1","volume-title":"Deep collaborative embedding for social image understanding","author":"Li Zechao","year":"2018","unstructured":"Zechao Li , Jinhui Tang , and Tao Mei . 2018. Deep collaborative embedding for social image understanding . IEEE Trans. Pattern Anal. Mach. Intell . ( 2018 ). Zechao Li, Jinhui Tang, and Tao Mei. 2018. Deep collaborative embedding for social image understanding. IEEE Trans. Pattern Anal. Mach. Intell. (2018)."},{"doi-asserted-by":"publisher","key":"e_1_2_1_29_1","DOI":"10.1109\/CVPR.2016.214"},{"volume-title":"2016 IEEE International Conference on Image Processing (ICIP). IEEE, 918--922","author":"Marsden Mark","unstructured":"Mark Marsden , Kevin McGuinness , Suzanne Little , and Noel E . O\u2019Connor. 2016. Holistic features for real-time crowd behaviour anomaly detection . In 2016 IEEE International Conference on Image Processing (ICIP). IEEE, 918--922 . Mark Marsden, Kevin McGuinness, Suzanne Little, and Noel E. O\u2019Connor. 2016. Holistic features for real-time crowd behaviour anomaly detection. In 2016 IEEE International Conference on Image Processing (ICIP). IEEE, 918--922.","key":"e_1_2_1_30_1"},{"doi-asserted-by":"publisher","key":"e_1_2_1_31_1","DOI":"10.1109\/AVSS.2007.4425267"},{"doi-asserted-by":"publisher","key":"e_1_2_1_32_1","DOI":"10.1109\/WACV.2015.27"},{"doi-asserted-by":"publisher","key":"e_1_2_1_33_1","DOI":"10.1109\/ICIP.2015.7351223"},{"key":"e_1_2_1_34_1","volume-title":"2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 1470--1477","author":"Ni Bingbing","year":"2009","unstructured":"Bingbing Ni , Shuicheng Yan , and Ashraf Kassim . 2009 . Recognizing human group activities with localized causalities . In 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 1470--1477 . Bingbing Ni, Shuicheng Yan, and Ashraf Kassim. 2009. Recognizing human group activities with localized causalities. In 2009 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 1470--1477."},{"doi-asserted-by":"publisher","key":"e_1_2_1_35_1","DOI":"10.1145\/2124295.2124363"},{"doi-asserted-by":"publisher","key":"e_1_2_1_36_1","DOI":"10.1109\/ICCV.2017.590"},{"key":"e_1_2_1_37_1","volume-title":"Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition. Arxiv Preprint Arxiv:1402.1128","author":"Sak Ha\u015fim","year":"2014","unstructured":"Ha\u015fim Sak , Andrew Senior , and Fran\u00e7oise Beaufays . 2014. Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition. Arxiv Preprint Arxiv:1402.1128 ( 2014 ). Ha\u015fim Sak, Andrew Senior, and Fran\u00e7oise Beaufays. 2014. Long short-term memory based recurrent neural network architectures for large vocabulary speech recognition. Arxiv Preprint Arxiv:1402.1128 (2014)."},{"doi-asserted-by":"publisher","key":"e_1_2_1_38_1","DOI":"10.1109\/78.650093"},{"doi-asserted-by":"publisher","key":"e_1_2_1_39_1","DOI":"10.1145\/1291233.1291311"},{"doi-asserted-by":"publisher","key":"e_1_2_1_40_1","DOI":"10.1109\/CVPR.2014.285"},{"key":"e_1_2_1_41_1","volume-title":"Hierarchical long short-term concurrent memory for human interaction recognition. Arxiv Preprint Arxiv:1811.00270","author":"Shu Xiangbo","year":"2018","unstructured":"Xiangbo Shu , Jinhui Tang , Guo-Jun Qi , Wei Liu , and Jian Yang . 2018. Hierarchical long short-term concurrent memory for human interaction recognition. Arxiv Preprint Arxiv:1811.00270 ( 2018 ). Xiangbo Shu, Jinhui Tang, Guo-Jun Qi, Wei Liu, and Jian Yang. 2018. Hierarchical long short-term concurrent memory for human interaction recognition. Arxiv Preprint Arxiv:1811.00270 (2018)."},{"unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Advances in Neural Information Processing Systems. 568--576.   Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Advances in Neural Information Processing Systems. 568--576.","key":"e_1_2_1_42_1"},{"key":"e_1_2_1_43_1","volume-title":"Very deep convolutional networks for large-scale image recognition. Arxiv Preprint Arxiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. Arxiv Preprint Arxiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. Arxiv Preprint Arxiv:1409.1556 (2014)."},{"key":"e_1_2_1_44_1","volume-title":"Amir Roshan Zamir, and Mubarak Shah","author":"Soomro Khurram","year":"2012","unstructured":"Khurram Soomro , Amir Roshan Zamir, and Mubarak Shah . 2012 . UCF101: A dataset of 101 human actions classes from videos in the wild. Arxiv Preprint Arxiv :1212.0402 (2012). Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. Arxiv Preprint Arxiv:1212.0402 (2012)."},{"key":"e_1_2_1_45_1","volume-title":"IJCAI","volume":"2","author":"Su Hang","year":"2016","unstructured":"Hang Su , Yinpeng Dong , Jun Zhu , Haibin Ling , and Bo Zhang . 2016 . Crowd scene understanding with coherent recurrent neural networks . In IJCAI , Vol. 2 . 5. Hang Su, Yinpeng Dong, Jun Zhu, Haibin Ling, and Bo Zhang. 2016. Crowd scene understanding with coherent recurrent neural networks. In IJCAI, Vol. 2. 5."},{"key":"e_1_2_1_46_1","volume-title":"Le","author":"Sutskever Ilya","year":"2014","unstructured":"Ilya Sutskever , Oriol Vinyals , and Quoc V . Le . 2014 . Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems . 3104--3112. Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to sequence learning with neural networks. In Advances in Neural Information Processing Systems. 3104--3112."},{"doi-asserted-by":"publisher","key":"e_1_2_1_47_1","DOI":"10.1109\/ICCV.2015.510"},{"doi-asserted-by":"publisher","key":"e_1_2_1_48_1","DOI":"10.1109\/ICCV.2015.460"},{"key":"e_1_2_1_49_1","volume-title":"Translating videos to natural language using deep recurrent neural networks. Arxiv Preprint Arxiv:1412.4729","author":"Venugopalan Subhashini","year":"2014","unstructured":"Subhashini Venugopalan , Huijuan Xu , Jeff Donahue , Marcus Rohrbach , Raymond Mooney , and Kate Saenko . 2014. Translating videos to natural language using deep recurrent neural networks. Arxiv Preprint Arxiv:1412.4729 ( 2014 ). Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko. 2014. Translating videos to natural language using deep recurrent neural networks. Arxiv Preprint Arxiv:1412.4729 (2014)."},{"doi-asserted-by":"publisher","key":"e_1_2_1_50_1","DOI":"10.1109\/CVPR.2015.7299059"},{"doi-asserted-by":"publisher","key":"e_1_2_1_51_1","DOI":"10.1145\/2671188.2749340"},{"key":"e_1_2_1_52_1","volume-title":"Hua","author":"Ye Jun","year":"2018","unstructured":"Jun Ye , Guojun Qi , Naifan Zhuang , Hao Hu , and Kien A . Hua . 2018 . Learning compact features for human activity recognition via probabilistic first-take-all. IEEE Trans. Pattern Anal. Mach. Intell . (2018). Jun Ye, Guojun Qi, Naifan Zhuang, Hao Hu, and Kien A. Hua. 2018. Learning compact features for human activity recognition via probabilistic first-take-all. IEEE Trans. Pattern Anal. Mach. Intell. (2018)."},{"doi-asserted-by":"publisher","key":"e_1_2_1_53_1","DOI":"10.1109\/CVPR.2015.7299101"},{"doi-asserted-by":"publisher","key":"e_1_2_1_54_1","DOI":"10.1109\/ISM.2016.0078"},{"doi-asserted-by":"publisher","key":"e_1_2_1_55_1","DOI":"10.1145\/1291233.1291308"},{"doi-asserted-by":"publisher","key":"e_1_2_1_56_1","DOI":"10.1145\/3097983.3098117"},{"doi-asserted-by":"publisher","key":"e_1_2_1_57_1","DOI":"10.5555\/1950054.1950056"},{"key":"e_1_2_1_58_1","volume-title":"International Conference on Machine Learning. 1604--1612","author":"Zhu Xiaodan","year":"2015","unstructured":"Xiaodan Zhu , Parinaz Sobihani , and Hongyu Guo . 2015 . Long short-term memory over recursive structures . In International Conference on Machine Learning. 1604--1612 . Xiaodan Zhu, Parinaz Sobihani, and Hongyu Guo. 2015. Long short-term memory over recursive structures. In International Conference on Machine Learning. 1604--1612."},{"doi-asserted-by":"publisher","key":"e_1_2_1_59_1","DOI":"10.1109\/JSTSP.2012.2234722"},{"doi-asserted-by":"publisher","key":"e_1_2_1_60_1","DOI":"10.1142\/S1793351X18400196"},{"volume-title":"2016 23rd International Conference on Pattern Recognition (ICPR). IEEE, 3222--3227","author":"Zhuang Naifan","unstructured":"Naifan Zhuang , Jun Ye , and Kien A. Hua . 2016. DLSTM approach to video modeling with hashing for large-scale video retrieval . In 2016 23rd International Conference on Pattern Recognition (ICPR). IEEE, 3222--3227 . Naifan Zhuang, Jun Ye, and Kien A. Hua. 2016. DLSTM approach to video modeling with hashing for large-scale video retrieval. In 2016 23rd International Conference on Pattern Recognition (ICPR). IEEE, 3222--3227.","key":"e_1_2_1_61_1"},{"volume-title":"2017 IEEE International Symposium on Multimedia (ISM). IEEE, 61--68","author":"Zhuang Naifan","unstructured":"Naifan Zhuang , Jun Ye , and Kien A. Hua . 2017. Convolutional DLSTM for crowd scene understanding . In 2017 IEEE International Symposium on Multimedia (ISM). IEEE, 61--68 . Naifan Zhuang, Jun Ye, and Kien A. Hua. 2017. Convolutional DLSTM for crowd scene understanding. In 2017 IEEE International Symposium on Multimedia (ISM). IEEE, 61--68.","key":"e_1_2_1_62_1"},{"volume-title":"2017 12th IEEE International Conference on Automatic Face 8 Gesture Recognition (FG'17)","author":"Zhuang Naifan","unstructured":"Naifan Zhuang , Tuoerhongjiang Yusufu , Jun Ye , and Kien A. Hua . 2017. Group activity recognition with differential recurrent convolutional neural networks . In 2017 12th IEEE International Conference on Automatic Face 8 Gesture Recognition (FG'17) . IEEE, 526--531. Naifan Zhuang, Tuoerhongjiang Yusufu, Jun Ye, and Kien A. Hua. 2017. Group activity recognition with differential recurrent convolutional neural networks. In 2017 12th IEEE International Conference on Automatic Face 8 Gesture Recognition (FG'17). IEEE, 526--531.","key":"e_1_2_1_63_1"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3337928","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3337928","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T00:25:42Z","timestamp":1750206342000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3337928"}},"subtitle":["Deep Differential Recurrent Neural Networks"],"short-title":[],"issued":{"date-parts":[[2019,8,31]]},"references-count":63,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2019,8,31]]}},"alternative-id":["10.1145\/3337928"],"URL":"https:\/\/doi.org\/10.1145\/3337928","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2019,8,31]]},"assertion":[{"value":"2019-01-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-05-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-09-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}