{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,7]],"date-time":"2026-05-07T16:24:42Z","timestamp":1778171082581,"version":"3.51.4"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"3s","license":[{"start":{"date-parts":[[2020,10,31]],"date-time":"2020-10-31T00:00:00Z","timestamp":1604102400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100004739","name":"Youth Innovation Promotion Association CAS","doi-asserted-by":"crossref","award":["2018497"],"award-info":[{"award-number":["2018497"]}],"id":[{"id":"10.13039\/501100004739","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100012659","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61632019, 61822208"],"award-info":[{"award-number":["61632019, 61822208"]}],"id":[{"id":"10.13039\/501100012659","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2020,10,31]]},"abstract":"<jats:p>In video action recognition, motion is a very crucial clue, which is usually represented by optical flow. However, optical flow is computationally expensive to obtain, which becomes the bottleneck for the efficiency of traditional action recognition algorithms. In this article, we propose a network called MV2Flow to learn motion representation efficiently from the signals in the compressed domain. To learn the network, three losses are defined. First, we select the classical TV-L1 flow as proxy ground truth to guide the learning. Besides, an unsupervised image reconstruction loss is proposed to further refine it. Moreover, toward the task of action recognition, the above two losses are combined with a motion content loss. To evaluate our approach, extensive experiments on two benchmark datasets UCF-101 and HMDB-51 are conducted. The motion representation generated with our MV2Flow has shown comparable classification performance on action recognition with TV-L1 flow, while operating at an over 200\u00d7 faster speed. Based on our MV2Flow and 2D-CNN-based network, we have achieved state-of-the-art performance in the compressed domain. With 3D-CNN-based network, we also achieve comparable accuracy with higher inference speed than methods in the decoded domain setting.<\/jats:p>","DOI":"10.1145\/3422360","type":"journal-article","created":{"date-parts":[[2020,12,31]],"date-time":"2020-12-31T18:43:08Z","timestamp":1609440188000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":17,"title":["MV2Flow"],"prefix":"10.1145","volume":"16","author":[{"given":"Hezhen","family":"Hu","sequence":"first","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wengang","family":"Zhou","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xingze","family":"Li","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ning","family":"Yan","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Houqiang","family":"Li","sequence":"additional","affiliation":[{"name":"University of Science and Technology of China, Institute of Artificial Intelligence, Hefei Comprehensive National Science Center, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,12,31]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2016.7532634"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-540-24673-2_3"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01246-5_22"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.316"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00630"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_2_1_9_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4768--4777","author":"Feichtenhofer Christoph","unstructured":"Christoph Feichtenhofer , Axel Pinz , and Richard P. Wildes . 2017. Spatiotemporal multiplier networks for video action recognition . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4768--4777 . Christoph Feichtenhofer, Axel Pinz, and Richard P. Wildes. 2017. Spatiotemporal multiplier networks for video action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4768--4777."},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.213"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.459"},{"key":"e_1_2_1_12_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2568--2577","author":"Gan Chuang","unstructured":"Chuang Gan , Naiyan Wang , Yi Yang , Dit-Yan Yeung , and Alex G. Hauptmann . 2015. Devnet: A deep event network for multimedia event detection and evidence recounting . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2568--2577 . Chuang Gan, Naiyan Wang, Yi Yang, Dit-Yan Yeung, and Alex G. Hauptmann. 2015. Devnet: A deep event network for multimedia event detection and evidence recounting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2568--2577."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00622"},{"key":"e_1_2_1_14_1","volume-title":"Proceedings of the Asian Conference on Computer Vision. 207--224","author":"G\u00fcney Fatma","year":"2016","unstructured":"Fatma G\u00fcney and Andreas Geiger . 2016 . Deep discrete flow . In Proceedings of the Asian Conference on Computer Vision. 207--224 . Fatma G\u00fcney and Andreas Geiger. 2016. Deep discrete flow. In Proceedings of the Asian Conference on Computer Vision. 207--224."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.imavis.2017.01.010"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1016\/0004-3702(81)90024-2"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2877936"},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the Advances in Neural Information Processing Systems. 2017--2025","author":"Jaderberg Max","year":"2015","unstructured":"Max Jaderberg , Karen Simonyan , Andrew Zisserman , et\u00a0al. 2015 . Spatial transformer networks . In Proceedings of the Advances in Neural Information Processing Systems. 2017--2025 . Max Jaderberg, Karen Simonyan, Andrew Zisserman, et\u00a0al. 2015. Spatial transformer networks. In Proceedings of the Advances in Neural Information Processing Systems. 2017--2025."},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the European Conference on Computer Vision. 3--10","author":"Jason J. Yu","unstructured":"J. Yu Jason , Adam W. Harley , and Konstantinos G. Derpanis . 2016. Back to basics: Unsupervised learning of optical flow via brightness constancy and motion smoothness . In Proceedings of the European Conference on Computer Vision. 3--10 . J. Yu Jason, Adam W. Harley, and Konstantinos G. Derpanis. 2016. Back to basics: Unsupervised learning of optical flow via brightness constancy and motion smoothness. In Proceedings of the European Conference on Computer Vision. 3--10."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of IEEE International Conference on Computer Vision. 2556--2563","author":"Jhuang H.","unstructured":"H. Jhuang , H. Garrote , E. Poggio , T. Serre , and T. Hmdb . 2011. Hmdb: A large video database for human motion recognition . In Proceedings of IEEE International Conference on Computer Vision. 2556--2563 . H. Jhuang, H. Garrote, E. Poggio, T. Serre, and T. Hmdb. 2011. Hmdb: A large video database for human motion recognition. In Proceedings of IEEE International Conference on Computer Vision. 2556--2563."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2012.59"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.332"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_2_1_25_1","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems. 1097--1105","author":"Krizhevsky Alex","unstructured":"Alex Krizhevsky , Ilya Sutskever , and Geoffrey E. Hinton . 2012. Imagenet classification with deep convolutional neural networks . In Proceedings of the Conference on Advances in Neural Information Processing Systems. 1097--1105 . Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Proceedings of the Conference on Advances in Neural Information Processing Systems. 1097--1105."},{"key":"e_1_2_1_26_1","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems. 354--364","author":"Lai Wei-Sheng","year":"2017","unstructured":"Wei-Sheng Lai , Jia-Bin Huang , and Ming-Hsuan Yang . 2017 . Semi-supervised learning for optical flow with generative adversarial networks . In Proceedings of the Conference on Advances in Neural Information Processing Systems. 354--364 . Wei-Sheng Lai, Jia-Bin Huang, and Ming-Hsuan Yang. 2017. Semi-supervised learning for optical flow with generative adversarial networks. In Proceedings of the Conference on Advances in Neural Information Processing Systems. 354--364."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/103085.103090"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33018618"},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence.","author":"Liu Kun","year":"2018","unstructured":"Kun Liu , Wu Liu , Chuang Gan , Mingkui Tan , and Huadong Ma . 2018 . T-C3D: Temporal convolutional 3d network for real-time action recognition . In Proceedings of the AAAI Conference on Artificial Intelligence. Kun Liu, Wu Liu, Chuang Gan, Mingkui Tan, and Huadong Ma. 2018. T-C3D: Temporal convolutional 3d network for real-time action recognition. In Proceedings of the AAAI Conference on Artificial Intelligence."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00817"},{"key":"e_1_2_1_31_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence. 7251--7259","author":"Meister Simon","year":"2018","unstructured":"Simon Meister , Junhwa Hur , and Stefan Roth . 2018 . UnFlow: Unsupervised learning of optical flow with a bidirectional census loss . In Proceedings of the AAAI Conference on Artificial Intelligence. 7251--7259 . Simon Meister, Junhwa Hur, and Stefan Roth. 2018. UnFlow: Unsupervised learning of optical flow with a bidirectional census loss. In Proceedings of the AAAI Conference on Artificial Intelligence. 7251--7259."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.5555\/2318960.2319533"},{"key":"e_1_2_1_33_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 9945--9953","author":"Piergiovanni AJ","unstructured":"AJ Piergiovanni and Michael S. Ryoo . 2019. Representation flow for action recognition . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 9945--9953 . AJ Piergiovanni and Michael S. Ryoo. 2019. Representation flow for action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 9945--9953."},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.590"},{"key":"e_1_2_1_35_1","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4161--4170","author":"Ranjan Anurag","unstructured":"Anurag Ranjan and Michael J. Black . 2017. Optical flow estimation using a spatial pyramid network . In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4161--4170 . Anurag Ranjan and Michael J. Black. 2017. Optical flow estimation using a spatial pyramid network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 4161--4170."},{"key":"e_1_2_1_36_1","volume-title":"Proceedings of the AAAI Conference on Artificial Intelligence. 1495--1501","author":"Ren Zhe","year":"2017","unstructured":"Zhe Ren , Junchi Yan , Bingbing Ni , Bin Liu , Xiaokang Yang , and Hongyuan Zha . 2017 . Unsupervised deep learning for optical flow estimation . In Proceedings of the AAAI Conference on Artificial Intelligence. 1495--1501 . Zhe Ren, Junchi Yan, Bingbing Ni, Bin Liu, Xiaokang Yang, and Hongyuan Zha. 2017. Unsupervised deep learning for optical flow estimation. In Proceedings of the AAAI Conference on Artificial Intelligence. 1495--1501."},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_2_1_38_1","volume-title":"Proceedings of the German Conference on Pattern Recognition. 281--297","author":"Sevilla-Lara Laura","unstructured":"Laura Sevilla-Lara , Yiyi Liao , Fatma G\u00fcney , Varun Jampani , Andreas Geiger , and Michael J. Black . 2018. On the integration of optical flow and action recognition . In Proceedings of the German Conference on Pattern Recognition. 281--297 . Laura Sevilla-Lara, Yiyi Liao, Fatma G\u00fcney, Varun Jampani, Andreas Geiger, and Michael J. Black. 2018. On the integration of optical flow and action recognition. In Proceedings of the German Conference on Pattern Recognition. 281--297."},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Zheng Shou Zhicheng Yan Yannis Kalantidis Laura Sevilla-Lara Marcus Rohrbach Xudong Lin and Shih-Fu Chang. 2019. DMC-Net: Generating discriminative motion cues for fast compressed video action recognition. Retrieved from https:\/\/Arxiv:1901.03460.  Zheng Shou Zhicheng Yan Yannis Kalantidis Laura Sevilla-Lara Marcus Rohrbach Xudong Lin and Shih-Fu Chang. 2019. DMC-Net: Generating discriminative motion cues for fast compressed video action recognition. Retrieved from https:\/\/Arxiv:1901.03460.","DOI":"10.1109\/CVPR.2019.00136"},{"key":"e_1_2_1_40_1","volume-title":"Proceedings of the Conference on Advances in Neural Information Processing Systems. 568--576","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014 . Two-stream convolutional networks for action recognition in videos . In Proceedings of the Conference on Advances in Neural Information Processing Systems. 568--576 . Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In Proceedings of the Conference on Advances in Neural Information Processing Systems. 568--576."},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2019.2896029"},{"key":"e_1_2_1_42_1","volume-title":"Amir Roshan Zamir, and Mubarak Shah","author":"Soomro Khurram","year":"2012","unstructured":"Khurram Soomro , Amir Roshan Zamir, and Mubarak Shah . 2012 . UCF101: A dataset of 101 human actions classes from videos in the wild. CRCV-TR- 12-01 (2012), 2, 5, 6, 7. Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. CRCV-TR-12-01 (2012), 2, 5, 6, 7."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00931"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.510"},{"key":"e_1_2_1_45_1","unstructured":"Du Tran Jamie Ray Zheng Shou Shih-Fu Chang and Manohar Paluri. 2017. Convnet architecture search for spatiotemporal feature learning. Retrieved from https:\/\/Arxiv:1708.05038.  Du Tran Jamie Ray Zheng Shou Shih-Fu Chang and Manohar Paluri. 2017. Convnet architecture search for spatiotemporal feature learning. Retrieved from https:\/\/Arxiv:1708.05038."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00675"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2018.2830102"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.441"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00155"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.175"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00631"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2008.927112"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299101"},{"key":"e_1_2_1_56_1","unstructured":"Christopher Zach Thomas Pock and Horst Bischof. 2007. A duality-based approach for realtime TV-L1 optical flow. In Pattern Recognition. 214--223.  Christopher Zach Thomas Pock and Horst Bischof. 2007. A duality-based approach for realtime TV-L1 optical flow. In Pattern Recognition. 214--223."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.297"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.219"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3422360","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3422360","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:03:21Z","timestamp":1750197801000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3422360"}},"subtitle":["Learning Motion Representation for Fast Compressed Video Action Recognition"],"short-title":[],"issued":{"date-parts":[[2020,10,31]]},"references-count":58,"journal-issue":{"issue":"3s","published-print":{"date-parts":[[2020,10,31]]}},"alternative-id":["10.1145\/3422360"],"URL":"https:\/\/doi.org\/10.1145\/3422360","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,10,31]]},"assertion":[{"value":"2019-10-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-07-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2020-12-31","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}