{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T18:05:25Z","timestamp":1785175525318,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":56,"publisher":"ACM","license":[{"start":{"date-parts":[[2019,10,15]],"date-time":"2019-10-15T00:00:00Z","timestamp":1571097600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Engineering Laboratory for Video Technology - Shenzhen Division"},{"name":"Shenzhen Municipal Development and Reform Commission"},{"name":"Aoto-PKUSZ Joint Lab"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2019,10,15]]},"DOI":"10.1145\/3343031.3350876","type":"proceedings-article","created":{"date-parts":[[2019,10,21]],"date-time":"2019-10-21T16:32:26Z","timestamp":1571675546000},"page":"500-509","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":50,"title":["PAN"],"prefix":"10.1145","author":[{"given":"Can","family":"Zhang","sequence":"first","affiliation":[{"name":"Peking University, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuexian","family":"Zou","sequence":"additional","affiliation":[{"name":"Peking University &amp; Peng Cheng Laboratory, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guang","family":"Chen","sequence":"additional","affiliation":[{"name":"Peking University, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Lei","family":"Gan","sequence":"additional","affiliation":[{"name":"Peking University, Shenzhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,10,15]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_1_2_1","volume-title":"Imagenet: A large-scale hierarchical image database.","author":"Deng Jia","year":"2009","unstructured":"Jia Deng , Wei Dong , Richard Socher , Li-Jia Li , Kai Li , and Li Fei-Fei . 2009 . Imagenet: A large-scale hierarchical image database. (2009). Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. (2009)."},{"key":"e_1_3_2_1_3_1","volume-title":"Mohammad Mahdi Arzani, Rahman Yousefzadeh, and Luc Van Gool.","author":"Diba Ali","year":"2017","unstructured":"Ali Diba , Mohsen Fayyaz , Vivek Sharma , Amir Hossein Karami , Mohammad Mahdi Arzani, Rahman Yousefzadeh, and Luc Van Gool. 2017 . Temporal 3d convnets: New architecture and transfer learning for video classification. arXiv preprint arXiv:1711.08200 (2017). Ali Diba, Mohsen Fayyaz, Vivek Sharma, Amir Hossein Karami, Mohammad Mahdi Arzani, Rahman Yousefzadeh, and Luc Van Gool. 2017. Temporal 3d convnets: New architecture and transfer learning for video classification. arXiv preprint arXiv:1711.08200 (2017)."},{"key":"e_1_3_2_1_4_1","volume-title":"Ali Mohammad Pazandeh, and Luc Van Gool","author":"Diba Ali","year":"2016","unstructured":"Ali Diba , Ali Mohammad Pazandeh, and Luc Van Gool . 2016 . Efficient two-stream motion and appearance 3d cnns for video classification. arXiv preprint arXiv:1608.08851 (2016). Ali Diba, Ali Mohammad Pazandeh, and Luc Van Gool. 2016. Efficient two-stream motion and appearance 3d cnns for video classification. arXiv preprint arXiv:1608.08851 (2016)."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298878"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.316"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00630"},{"key":"e_1_3_2_1_8_1","unstructured":"Christoph Feichtenhofer Axel Pinz and Richard Wildes. 2016a. Spatiotemporal residual networks for video action recognition. In Advances in neural information processing systems. 3468--3476.  Christoph Feichtenhofer Axel Pinz and Richard Wildes. 2016a. Spatiotemporal residual networks for video action recognition. In Advances in neural information processing systems. 3468--3476."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.787"},{"key":"e_1_3_2_1_10_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 7844--7853","author":"Feichtenhofer Christoph","year":"2018","unstructured":"Christoph Feichtenhofer , Axel Pinz , Richard P Wildes , and Andrew Zisserman . 2018 . What have we learned from deep representations for action recognition? . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 7844--7853 . Christoph Feichtenhofer, Axel Pinz, Richard P Wildes, and Andrew Zisserman. 2018. What have we learned from deep representations for action recognition?. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 7844--7853."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.213"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46487-9_52"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.106"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.337"},{"key":"e_1_3_2_1_15_1","first-page":"8","article-title":"The","volume":"2","author":"Goyal Raghav","year":"2017","unstructured":"Raghav Goyal , Samira Ebrahimi Kahou , Vincent Michalski , Joanna Materzynska , Susanne Westphal , Heuna Kim , Valentin Haenel , Ingo Fruend , Peter Yianilos , Moritz Mueller-Freitag , and Others. 2017 . The \" Something Something\" Video Database for Learning and Evaluating Visual Common Sense.. In ICCV , Vol. 2. 8 . Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, and Others. 2017. The\" Something Something\" Video Database for Learning and Evaluating Visual Common Sense.. In ICCV, Vol. 2. 8.","journal-title":"In ICCV"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00685"},{"key":"e_1_3_2_1_17_1","volume-title":"Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision & Pattern Recognition .","author":"He Kaiming","year":"2016","unstructured":"Kaiming He , Xiangyu Zhang , Shaoqing Ren , and Sun Jian . 2016 . Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision & Pattern Recognition . Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Sun Jian. 2016. Deep Residual Learning for Image Recognition. In IEEE Conference on Computer Vision & Pattern Recognition ."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2389824"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Fabian Caba Heilbron Victor Escorcia Bernard Ghanem and Juan Carlos Niebles. 2015. ActivityNet: A large-scale video benchmark for human activity understanding. In Computer Vision & Pattern Recognition .  Fabian Caba Heilbron Victor Escorcia Bernard Ghanem and Juan Carlos Niebles. 2015. ActivityNet: A large-scale video benchmark for human activity understanding. In Computer Vision & Pattern Recognition .","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(81)90024-2"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2877936"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.179"},{"key":"e_1_3_2_1_23_1","volume-title":"Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167","author":"Ioffe Sergey","year":"2015","unstructured":"Sergey Ioffe and Christian Szegedy . 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167 ( 2015 ). Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167 (2015)."},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.59"},{"key":"e_1_3_2_1_25_1","volume-title":"The kinetics human action video dataset . arXiv preprint arXiv:1705.06950","author":"Kay Will","year":"2017","unstructured":"Will Kay , Joao Carreira , Karen Simonyan , Brian Zhang , Chloe Hillier , Sudheendra Vijayanarasimhan , Fabio Viola , Tim Green , Trevor Back , Paul Natsev , and Others. 2017. The kinetics human action video dataset . arXiv preprint arXiv:1705.06950 ( 2017 ). Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, and Others. 2017. The kinetics human action video dataset . arXiv preprint arXiv:1705.06950 (2017)."},{"key":"e_1_3_2_1_26_1","volume-title":"International Conference on Neural Information Processing Systems .","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E Hinton . 2012 . ImageNet classification with deep convolutional neural networks . In International Conference on Neural Information Processing Systems . Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet classification with deep convolutional neural networks. In International Conference on Neural Information Processing Systems ."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01249-6_24"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2017.10.011"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00817"},{"key":"e_1_3_2_1_31_1","volume-title":"2018 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1616--1624","author":"Yue-Hei Ng Joe","year":"2018","unstructured":"Joe Yue-Hei Ng , Jonghyun Choi , Jan Neumann , and Larry S Davis . 2018 . Actionflownet: Learning motion representation for action recognition . In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1616--1624 . Joe Yue-Hei Ng, Jonghyun Choi, Jan Neumann, and Larry S Davis. 2018. Actionflownet: Learning motion representation for action recognition. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV). IEEE, 1616--1624."},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.590"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.291"},{"key":"e_1_3_2_1_34_1","volume-title":"On the integration of optical flow and action recognition. arXiv preprint arXiv:1712.08416","author":"Sevilla-Lara Laura","year":"2017","unstructured":"Laura Sevilla-Lara , Yiyi Liao , Fatma Guney , Varun Jampani , Andreas Geiger , and Michael J Black . 2017. On the integration of optical flow and action recognition. arXiv preprint arXiv:1712.08416 ( 2017 ). Laura Sevilla-Lara, Yiyi Liao, Fatma Guney, Varun Jampani, Andreas Geiger, and Michael J Black. 2017. On the integration of optical flow and action recognition. arXiv preprint arXiv:1712.08416 (2017)."},{"key":"e_1_3_2_1_35_1","volume-title":"Action recognition using visual attention. arXiv preprint arXiv:1511.04119","author":"Sharma Shikhar","year":"2015","unstructured":"Shikhar Sharma , Ryan Kiros , and Ruslan Salakhutdinov . 2015. Action recognition using visual attention. arXiv preprint arXiv:1511.04119 ( 2015 ). Shikhar Sharma, Ryan Kiros, and Ruslan Salakhutdinov. 2015. Action recognition using visual attention. arXiv preprint arXiv:1511.04119 (2015)."},{"key":"e_1_3_2_1_36_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014a. Two-stream convolutional networks for action recognition in videos. In Advances in neural information processing systems. 568--576.  Karen Simonyan and Andrew Zisserman. 2014a. Two-stream convolutional networks for action recognition in videos. In Advances in neural information processing systems. 568--576."},{"key":"e_1_3_2_1_37_1","volume-title":"Very deep convolutional networks for large-scale image recognition . arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014b. Very deep convolutional networks for large-scale image recognition . arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014b. Very deep convolutional networks for large-scale image recognition . arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_3_2_1_38_1","volume-title":"Amir Roshan Zamir, and Mubarak Shah","author":"Soomro Khurram","year":"2012","unstructured":"Khurram Soomro , Amir Roshan Zamir, and Mubarak Shah . 2012 . UCF101: A dataset of 101 human actions classes from videos in the wild . arXiv preprint arXiv:1212.0402 (2012). Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild . arXiv preprint arXiv:1212.0402 (2012)."},{"key":"e_1_3_2_1_39_1","volume-title":"International conference on machine learning. 843--852","author":"Srivastava Nitish","year":"2015","unstructured":"Nitish Srivastava , Elman Mansimov , and Ruslan Salakhudinov . 2015 . Unsupervised learning of video representations using lstms . In International conference on machine learning. 843--852 . Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov. 2015. Unsupervised learning of video representations using lstms. In International conference on machine learning. 843--852."},{"key":"e_1_3_2_1_40_1","volume-title":"Secrets of optical flow estimation and their principles. In 2010 IEEE computer society conference on computer vision and pattern recognition","author":"Sun Deqing","unstructured":"Deqing Sun , Stefan Roth , and Michael J Black . 2010. Secrets of optical flow estimation and their principles. In 2010 IEEE computer society conference on computer vision and pattern recognition . IEEE , 2432--2439. Deqing Sun, Stefan Roth, and Michael J Black. 2010. Secrets of optical flow estimation and their principles. In 2010 IEEE computer society conference on computer vision and pattern recognition. IEEE, 2432--2439."},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00931"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00151"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_3_2_1_45_1","volume-title":"Convnet architecture search for spatiotemporal feature learning. arXiv preprint arXiv:1708.05038","author":"Tran Du","year":"2017","unstructured":"Du Tran , Jamie Ray , Zheng Shou , Shih-Fu Chang , and Manohar Paluri . 2017. Convnet architecture search for spatiotemporal feature learning. arXiv preprint arXiv:1708.05038 ( 2017 ). Du Tran, Jamie Ray, Zheng Shou, Shih-Fu Chang, and Manohar Paluri. 2017. Convnet architecture search for spatiotemporal feature learning. arXiv preprint arXiv:1708.05038 (2017)."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00155"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299101"},{"key":"e_1_3_2_1_49_1","volume-title":"Joint pattern recognition symposium","author":"Zach Christopher","unstructured":"Christopher Zach , Thomas Pock , and Horst Bischof . 2007. A duality based approach for realtime TV-L 1 optical flow . In Joint pattern recognition symposium . Springer , 214--223. Christopher Zach, Thomas Pock, and Horst Bischof. 2007. A duality based approach for realtime TV-L 1 optical flow. In Joint pattern recognition symposium. Springer, 214--223."},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"e_1_3_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.297"},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00687"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_49"},{"key":"e_1_3_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2018.8545710"},{"key":"e_1_3_2_1_55_1","volume-title":"Hidden two-stream convolutional networks for action recognition. arXiv preprint arXiv:1704.00389","author":"Zhu Yi","year":"2017","unstructured":"Yi Zhu , Zhenzhong Lan , Shawn Newsam , and Alexander G Hauptmann . 2017. Hidden two-stream convolutional networks for action recognition. arXiv preprint arXiv:1704.00389 ( 2017 ). Yi Zhu, Zhenzhong Lan, Shawn Newsam, and Alexander G Hauptmann. 2017. Hidden two-stream convolutional networks for action recognition. arXiv preprint arXiv:1704.00389 (2017)."},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01216-8_43"}],"event":{"name":"MM '19: The 27th ACM International Conference on Multimedia","location":"Nice France","acronym":"MM '19","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 27th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3343031.3350876","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3343031.3350876","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:13:25Z","timestamp":1750202005000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3343031.3350876"}},"subtitle":["Persistent Appearance Network with an Efficient Motion Cue for Fast Action Recognition"],"short-title":[],"issued":{"date-parts":[[2019,10,15]]},"references-count":56,"alternative-id":["10.1145\/3343031.3350876","10.1145\/3343031"],"URL":"https:\/\/doi.org\/10.1145\/3343031.3350876","relation":{},"subject":[],"published":{"date-parts":[[2019,10,15]]},"assertion":[{"value":"2019-10-15","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}