{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,18]],"date-time":"2026-07-18T22:24:28Z","timestamp":1784413468643,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":73,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T00:00:00Z","timestamp":1602460800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,10,12]]},"DOI":"10.1145\/3394171.3413502","type":"proceedings-article","created":{"date-parts":[[2020,10,12]],"date-time":"2020-10-12T13:10:44Z","timestamp":1602508244000},"page":"1142-1151","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":89,"title":["Depth Guided Adaptive Meta-Fusion Network for Few-shot Video Recognition"],"prefix":"10.1145","author":[{"given":"Yuqian","family":"Fu","sequence":"first","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Li","family":"Zhang","sequence":"additional","affiliation":[{"name":"University of Oxford, Oxford, United Kingdom"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Junke","family":"Wang","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yanwei","family":"Fu","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu-Gang","family":"Jiang","sequence":"additional","affiliation":[{"name":"Fudan University, Shanghai, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,10,12]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition. arXiv preprint","author":"Bishay Mina","year":"2019","unstructured":"Mina Bishay , Georgios Zoumpourlis , and Ioannis Patras . 2019 . TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition. arXiv preprint (2019). Mina Bishay, Georgios Zoumpourlis, and Ioannis Patras. 2019. TARN: Temporal Attentive Relation Network for Few-Shot and Zero-Shot Action Recognition. arXiv preprint (2019)."},{"key":"e_1_3_2_2_2_1","unstructured":"Zeyd Boukhers Kimiaki Shirahama and Marcin Grzegorzek. [n.d.]. Example-based 3d trajectory extraction of objects from 2d videos. TCSVT ([n. d.]).  Zeyd Boukhers Kimiaki Shirahama and Marcin Grzegorzek. [n.d.]. Example-based 3d trajectory extraction of objects from 2d videos. TCSVT ([n. d.])."},{"key":"e_1_3_2_2_3_1","doi-asserted-by":"crossref","unstructured":"Kaidi Cao Jingwei Ji Zhangjie Cao Chien-Yi Chang and Juan Carlos Niebles. 2020. Few-shot video classification via temporal alignment. In CVPR.  Kaidi Cao Jingwei Ji Zhangjie Cao Chien-Yi Chang and Juan Carlos Niebles. 2020. Few-shot video classification via temporal alignment. In CVPR.","DOI":"10.1109\/CVPR42600.2020.01063"},{"key":"e_1_3_2_2_4_1","doi-asserted-by":"crossref","unstructured":"Joao Carreira and Andrew Zisserman. 2017. Quo vadis action recognition? a new model and the kinetics dataset. In CVPR.  Joao Carreira and Andrew Zisserman. 2017. Quo vadis action recognition? a new model and the kinetics dataset. In CVPR.","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_3_2_2_5_1","volume-title":"Kuan-Ying Lee, and Winston Hsu.","author":"Chang Ya-Liang","year":"2019","unstructured":"Ya-Liang Chang , Zhe Yu Liu , Kuan-Ying Lee, and Winston Hsu. 2019 . Learnable gated temporal shift module for deep video inpainting. arXiv preprint (2019). Ya-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, and Winston Hsu. 2019. Learnable gated temporal shift module for deep video inpainting. arXiv preprint (2019)."},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"crossref","unstructured":"Zitian Chen Yanwei Fu Kaiyu Chen and Yu-Gang Jiang. 2019 a. Image block augmentation for one-shot learning. In AAAI.  Zitian Chen Yanwei Fu Kaiyu Chen and Yu-Gang Jiang. 2019 a. Image block augmentation for one-shot learning. In AAAI.","DOI":"10.1609\/aaai.v33i01.33013379"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Zitian Chen Yanwei Fu Yu-Xiong Wang Lin Ma Wei Liu and Martial Hebert. 2018. Image deformation meta-network for one-shot learning. In CVPR.  Zitian Chen Yanwei Fu Yu-Xiong Wang Lin Ma Wei Liu and Martial Hebert. 2018. Image deformation meta-network for one-shot learning. In CVPR.","DOI":"10.1109\/CVPR.2019.00888"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Zitian Chen Yanwei Fu Yu-Xiong Wang Lin Ma Wei Liu and Martial Hebert. 2019 b. Image deformation meta-networks for one-shot learning. In CVPR.  Zitian Chen Yanwei Fu Yu-Xiong Wang Lin Ma Wei Liu and Martial Hebert. 2019 b. Image deformation meta-networks for one-shot learning. In CVPR.","DOI":"10.1109\/CVPR.2019.00888"},{"key":"e_1_3_2_2_9_1","unstructured":"Jinwoo Choi Chen Gao Joseph Messou and Jia-Bin Huang. 2019. Why Can-t I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition. In NeurIPS.  Jinwoo Choi Chen Gao Joseph Messou and Jia-Bin Huang. 2019. Why Can-t I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition. In NeurIPS."},{"key":"e_1_3_2_2_10_1","volume-title":"Imagenet: A large-scale hierarchical image database. In CVPR.","author":"Deng Jia","year":"2009","unstructured":"Jia Deng , Wei Dong , Richard Socher , Li-Jia Li , Kai Li , and Li Fei-Fei . 2009 . Imagenet: A large-scale hierarchical image database. In CVPR. Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In CVPR."},{"key":"e_1_3_2_2_11_1","volume-title":"Viola","author":"Matsakis Nicholas E.","year":"2000","unstructured":"Nicholas E. Matsakis Erik G. Miller and Paul A . Viola . 2000 . Learning from one example through shared densities on transforms. In CVPR. Nicholas E. Matsakis Erik G. Miller and Paul A. Viola. 2000. Learning from one example through shared densities on transforms. In CVPR."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"crossref","unstructured":"Christoph Feichtenhofer Haoqi Fan Jitendra Malik and Kaiming He. 2019. Slowfast networks for video recognition. In CVPR.  Christoph Feichtenhofer Haoqi Fan Jitendra Malik and Kaiming He. 2019. Slowfast networks for video recognition. In CVPR.","DOI":"10.1109\/ICCV.2019.00630"},{"key":"e_1_3_2_2_13_1","unstructured":"Chelsea Finn Pieter Abbeel and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In ICML.  Chelsea Finn Pieter Abbeel and Sergey Levine. 2017. Model-agnostic meta-learning for fast adaptation of deep networks. In ICML."},{"key":"e_1_3_2_2_14_1","unstructured":"Yuqian Fu Chengrong Wang Yanwei Fu Yu-Xiong Wang Cong Bai Xiangyang Xue and Yu-Gang Jiang. 2019. Embodied One-Shot Video Recognition: Learning from Actions of a Virtual Embodied Agent. In ACM MM.  Yuqian Fu Chengrong Wang Yanwei Fu Yu-Xiong Wang Cong Bai Xiangyang Xue and Yu-Gang Jiang. 2019. Embodied One-Shot Video Recognition: Learning from Actions of a Virtual Embodied Agent. In ACM MM."},{"key":"e_1_3_2_2_15_1","unstructured":"Hang Gao Zheng Shou Alireza Zareian Hanwang Zhang and Shih-Fu Chang. 2018. Low-shot learning via covariance-preserving adversarial augmentation networks. In NeurIPS.  Hang Gao Zheng Shou Alireza Zareian Hanwang Zhang and Shih-Fu Chang. 2018. Low-shot learning via covariance-preserving adversarial augmentation networks. In NeurIPS."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"crossref","unstructured":"Leon A Gatys Alexander S Ecker and Matthias Bethge. 2016. Image style transfer using convolutional neural networks. In CVPR.  Leon A Gatys Alexander S Ecker and Matthias Bethge. 2016. Image style transfer using convolutional neural networks. In CVPR.","DOI":"10.1109\/CVPR.2016.265"},{"key":"e_1_3_2_2_17_1","doi-asserted-by":"crossref","unstructured":"Andreas Geiger Philip Lenz and Raquel Urtasun. 2012. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR.  Andreas Geiger Philip Lenz and Raquel Urtasun. 2012. Are we ready for autonomous driving? the kitti vision benchmark suite. In CVPR.","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"e_1_3_2_2_18_1","volume":"201","author":"Godard Cl\u00e9ment","unstructured":"Cl\u00e9ment Godard , Oisin Mac Aodha , Michael Firman , and Gabriel J Brostow. 201 9. Digging into self-supervised monocular depth estimation. In ICCV. Cl\u00e9ment Godard, Oisin Mac Aodha, Michael Firman, and Gabriel J Brostow. 2019. Digging into self-supervised monocular depth estimation. In ICCV.","journal-title":"Gabriel J Brostow."},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"crossref","unstructured":"Abhinav Gupta and Larry S Davis. 2007. Objects in action: An approach for combining action understanding and object perception. In CVPR.  Abhinav Gupta and Larry S Davis. 2007. Objects in action: An approach for combining action understanding and object perception. In CVPR.","DOI":"10.1109\/CVPR.2007.383331"},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"crossref","unstructured":"Bharath Hariharan and Ross Girshick. 2017. Low-shot visual recognition by shrinking and hallucinating features. In ICCV.  Bharath Hariharan and Ross Girshick. 2017. Low-shot visual recognition by shrinking and hallucinating features. In ICCV.","DOI":"10.1109\/ICCV.2017.328"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"crossref","unstructured":"K. He Y. Fu W. Zhang C. Wang Y.-G. Jiang F. Huang and X. Xue. 2018. Harnessing Synthesized Abstraction Images to Improve Facial Attribute Recognition.  K. He Y. Fu W. Zhang C. Wang Y.-G. Jiang F. Huang and X. Xue. 2018. Harnessing Synthesized Abstraction Images to Improve Facial Attribute Recognition.","DOI":"10.24963\/ijcai.2018\/102"},{"key":"e_1_3_2_2_22_1","unstructured":"Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR.  Kaiming He Xiangyu Zhang Shaoqing Ren and Jian Sun. 2016. Deep residual learning for image recognition. In CVPR."},{"key":"e_1_3_2_2_23_1","volume-title":"Weakly-supervised compositional featureaggregation for few-shot recognition. arXiv preprint","author":"Hu Ping","year":"2019","unstructured":"Ping Hu , Ximeng Sun , Kate Saenko , and Stan Sclaroff . 2019. Weakly-supervised compositional featureaggregation for few-shot recognition. arXiv preprint ( 2019 ). Ping Hu, Ximeng Sun, Kate Saenko, and Stan Sclaroff. 2019. Weakly-supervised compositional featureaggregation for few-shot recognition. arXiv preprint (2019)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Xun Huang and Serge Belongie. 2017. Arbitrary style transfer in real-time with adaptive instance normalization. In ICCV.  Xun Huang and Serge Belongie. 2017. Arbitrary style transfer in real-time with adaptive instance normalization. In ICCV.","DOI":"10.1109\/ICCV.2017.167"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"crossref","unstructured":"Xun Huang Ming-Yu Liu Serge Belongie and Jan Kautz. 2018. Multimodal Unsupervised Image-to-image Translation. In ECCV.  Xun Huang Ming-Yu Liu Serge Belongie and Jan Kautz. 2018. Multimodal Unsupervised Image-to-image Translation. In ECCV.","DOI":"10.1007\/978-3-030-01219-9_11"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"crossref","unstructured":"Nazli Ikizler-Cinbis and Stan Sclaroff. 2010. Object scene and actions: Combining multiple features for human action recognition. In ECCV.  Nazli Ikizler-Cinbis and Stan Sclaroff. 2010. Object scene and actions: Combining multiple features for human action recognition. In ECCV.","DOI":"10.1007\/978-3-642-15549-9_36"},{"key":"e_1_3_2_2_27_1","unstructured":"Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML.  Sergey Ioffe and Christian Szegedy. 2015. Batch normalization: Accelerating deep network training by reducing internal covariate shift. In ICML."},{"key":"e_1_3_2_2_28_1","volume-title":"Jan C Van Gemert, and Cees GM Snoek","author":"Jain Mihir","year":"2015","unstructured":"Mihir Jain , Jan C Van Gemert, and Cees GM Snoek . 2015 . What do 15,000 object categories tell us about classifying and localizing actions?. In CVPR. Mihir Jain, Jan C Van Gemert, and Cees GM Snoek. 2015. What do 15,000 object categories tell us about classifying and localizing actions?. In CVPR."},{"key":"e_1_3_2_2_29_1","volume-title":"Informative joints based human action recognition using skeleton contexts","author":"Jiang Min","year":"2015","unstructured":"Min Jiang , Jun Kong , George Bebis , and Hongtao Huo . 2015. Informative joints based human action recognition using skeleton contexts . Signal Processing : Image Communication ( 2015 ). Min Jiang, Jun Kong, George Bebis, and Hongtao Huo. 2015. Informative joints based human action recognition using skeleton contexts. Signal Processing: Image Communication (2015)."},{"key":"e_1_3_2_2_30_1","volume-title":"et almbox","author":"Kay Will","year":"2017","unstructured":"Will Kay , Joao Carreira , Karen Simonyan , Brian Zhang , Chloe Hillier , Sudheendra Vijayanarasimhan , Fabio Viola , Tim Green , Trevor Back , Paul Natsev , et almbox . 2017 . The kinetics human action video dataset. arXiv preprint (2017). Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et almbox. 2017. The kinetics human action video dataset. arXiv preprint (2017)."},{"key":"e_1_3_2_2_31_1","volume-title":"ICML workshops.","author":"Koch Gregory","year":"2015","unstructured":"Gregory Koch , Richard Zemel , and Ruslan Salakhutdinov . 2015 . Siamese neural networks for one-shot image recognition . In ICML workshops. Gregory Koch, Richard Zemel, and Ruslan Salakhutdinov. 2015. Siamese neural networks for one-shot image recognition. In ICML workshops."},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"crossref","unstructured":"Hildegard Kuehne Hueihan Jhuang Est'ibaliz Garrote Tomaso Poggio and Thomas Serre. 2011. HMDB: a large video database for human motion recognition. In ICCV.  Hildegard Kuehne Hueihan Jhuang Est'ibaliz Garrote Tomaso Poggio and Thomas Serre. 2011. HMDB: a large video database for human motion recognition. In ICCV.","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"e_1_3_2_2_33_1","volume-title":"Lake and Ruslan Salakhutdinov","author":"Brenden","year":"2013","unstructured":"Brenden M. Lake and Ruslan Salakhutdinov . 2013 . One-shot learning by inverting a compositional causal process. In NeurIPS. Brenden M. Lake and Ruslan Salakhutdinov. 2013. One-shot learning by inverting a compositional causal process. In NeurIPS."},{"key":"e_1_3_2_2_34_1","volume-title":"Tenenbaum","author":"Lake Brenden M.","year":"2011","unstructured":"Brenden M. Lake , Ruslan Salakhutdinov , Jason Gross , and Joshua B . Tenenbaum . 2011 . One shot learning of simple visual concepts. In CogSci . Brenden M. Lake, Ruslan Salakhutdinov, Jason Gross, and Joshua B. Tenenbaum. 2011. One shot learning of simple visual concepts. In CogSci."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"crossref","unstructured":"Li-Jia Li and Li Fei-Fei. 2007. What where and who? classifying events by scene and object recognition. In ICCV.  Li-Jia Li and Li Fei-Fei. 2007. What where and who? classifying events by scene and object recognition. In ICCV.","DOI":"10.1109\/ICCV.2007.4408872"},{"key":"e_1_3_2_2_36_1","volume-title":"Resound: Towards action recognition without representation. In ECCV.","author":"Li Yingwei","year":"2018","unstructured":"Yingwei Li , Yi Li , and Nuno Vasconcelos . 2018 . Resound: Towards action recognition without representation. In ECCV. Yingwei Li, Yi Li, and Nuno Vasconcelos. 2018. Resound: Towards action recognition without representation. In ECCV."},{"key":"e_1_3_2_2_37_1","volume-title":"Tsm: Temporal shift module for efficient video understanding. In ICCV.","author":"Lin Ji","year":"2019","unstructured":"Ji Lin , Chuang Gan , and Song Han . 2019 . Tsm: Temporal shift module for efficient video understanding. In ICCV. Ji Lin, Chuang Gan, and Song Han. 2019. Tsm: Temporal shift module for efficient video understanding. In ICCV."},{"key":"e_1_3_2_2_38_1","volume-title":"CVPR workshops.","author":"Liu Chen","year":"2020","unstructured":"Chen Liu , Chengming Xu , Yikai Wang , Li Zhang , and Yanwei Fu . 2020 . An Embarrassingly Simple Baseline to One-Shot Learning . In CVPR workshops. Chen Liu, Chengming Xu, Yikai Wang, Li Zhang, and Yanwei Fu. 2020. An Embarrassingly Simple Baseline to One-Shot Learning. In CVPR workshops."},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"crossref","unstructured":"Marcin Marszalek Ivan Laptev and Cordelia Schmid. 2009. Actions in context. In CVPR.  Marcin Marszalek Ivan Laptev and Cordelia Schmid. 2009. Actions in context. In CVPR.","DOI":"10.1109\/CVPR.2009.5206557"},{"key":"e_1_3_2_2_40_1","doi-asserted-by":"crossref","unstructured":"Taesung Park Ming-Yu Liu Ting-Chun Wang and Jun-Yan Zhu. 2019. Semantic Image Synthesis With Spatially-Adaptive Normalization. In CVPR.  Taesung Park Ming-Yu Liu Ting-Chun Wang and Jun-Yan Zhu. 2019. Semantic Image Synthesis With Spatially-Adaptive Normalization. In CVPR.","DOI":"10.1109\/CVPR.2019.00244"},{"key":"e_1_3_2_2_41_1","volume-title":"Egocentric Action Recognition by Video Attention and Temporal Context. arXiv preprint","author":"Perez-Rua Juan-Manuel","year":"2020","unstructured":"Juan-Manuel Perez-Rua , Antoine Toisoul , Brais Martinez , Victor Escorcia , Li Zhang , Xiatian Zhu , and Tao Xiang . 2020. Egocentric Action Recognition by Video Attention and Temporal Context. arXiv preprint ( 2020 ). Juan-Manuel Perez-Rua, Antoine Toisoul, Brais Martinez, Victor Escorcia, Li Zhang, Xiatian Zhu, and Tao Xiang. 2020. Egocentric Action Recognition by Video Attention and Temporal Context. arXiv preprint (2020)."},{"key":"e_1_3_2_2_42_1","volume-title":"Long-Term Cloth-Changing Person Re-identification. arXiv preprint arXiv:2005.12633","author":"Qian Xuelin","year":"2020","unstructured":"Xuelin Qian , Wenxuan Wang , Li Zhang , Fangrui Zhu , Yanwei Fu , Tao Xiang , Yu-Gang Jiang , and Xiangyang Xue . 2020. Long-Term Cloth-Changing Person Re-identification. arXiv preprint arXiv:2005.12633 ( 2020 ). Xuelin Qian, Wenxuan Wang, Li Zhang, Fangrui Zhu, Yanwei Fu, Tao Xiang, Yu-Gang Jiang, and Xiangyang Xue. 2020. Long-Term Cloth-Changing Person Re-identification. arXiv preprint arXiv:2005.12633 (2020)."},{"key":"e_1_3_2_2_43_1","unstructured":"Eli Schwartz Leonid Karlinsky Joseph Shtok Sivan Harary Mattias Marder Abhishek Kumar Rogerio Feris Raja Giryes and Alex Bronstein. 2018. Delta-encoder: an effective sample synthesis method for few-shot object recognition. In NeurIPS.  Eli Schwartz Leonid Karlinsky Joseph Shtok Sivan Harary Mattias Marder Abhishek Kumar Rogerio Feris Raja Giryes and Alex Bronstein. 2018. Delta-encoder: an effective sample synthesis method for few-shot object recognition. In NeurIPS."},{"key":"e_1_3_2_2_44_1","volume-title":"Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV.","author":"Selvaraju Ramprasaath R","year":"2017","unstructured":"Ramprasaath R Selvaraju , Michael Cogswell , Abhishek Das , Ramakrishna Vedantam , Devi Parikh , and Dhruv Batra . 2017 . Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV. Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In ICCV."},{"key":"e_1_3_2_2_45_1","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In NeurIPS.  Karen Simonyan and Andrew Zisserman. 2014. Two-stream convolutional networks for action recognition in videos. In NeurIPS."},{"key":"e_1_3_2_2_46_1","unstructured":"Jake Snell Kevin Swersky and Richard Zemel. 2017b. Prototypical networks for few-shot learning. In NeurIPS.  Jake Snell Kevin Swersky and Richard Zemel. 2017b. Prototypical networks for few-shot learning. In NeurIPS."},{"key":"e_1_3_2_2_47_1","volume-title":"Zemeln","author":"Snell Jake","year":"2017","unstructured":"Jake Snell , Kevin Swersky , and Richard S . Zemeln . 2017 a. Prototypical networks for few-shot learning. In NeurIPS. Jake Snell, Kevin Swersky, and Richard S. Zemeln. 2017a. Prototypical networks for few-shot learning. In NeurIPS."},{"key":"e_1_3_2_2_48_1","volume-title":"Amir Roshan Zamir, and Mubarak Shah","author":"Soomro Khurram","year":"2012","unstructured":"Khurram Soomro , Amir Roshan Zamir, and Mubarak Shah . 2012 . UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint (2012). Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint (2012)."},{"key":"e_1_3_2_2_49_1","volume-title":"Hospedales","author":"Sung Flood","year":"2018","unstructured":"Flood Sung , Yongxin Yang , Li Zhang , Tao Xiang , Philip H.S. Torr , and Timothy M . Hospedales . 2018 . Learning to Compare : Relation Network for Few-Shot Learning. In CVPR. Flood Sung, Yongxin Yang, Li Zhang, Tao Xiang, Philip H.S. Torr, and Timothy M. Hospedales. 2018. Learning to Compare: Relation Network for Few-Shot Learning. In CVPR."},{"key":"e_1_3_2_2_50_1","volume-title":"Learning to learn: Meta-critic networks for sample efficient learning. arXiv preprint","author":"Sung Flood","year":"2017","unstructured":"Flood Sung , Li Zhang , Tao Xiang , Timothy Hospedales , and Yongxin Yang . 2017. Learning to learn: Meta-critic networks for sample efficient learning. arXiv preprint ( 2017 ). Flood Sung, Li Zhang, Tao Xiang, Timothy Hospedales, and Yongxin Yang. 2017. Learning to learn: Meta-critic networks for sample efficient learning. arXiv preprint (2017)."},{"key":"e_1_3_2_2_51_1","volume-title":"Learning from Web Data with Memory Module. arXiv preprint","author":"Tu Yi","year":"2019","unstructured":"Yi Tu , Li Niu , Junjie Chen , Dawei Cheng , and Liqing Zhang . 2019. Learning from Web Data with Memory Module. arXiv preprint ( 2019 ). Yi Tu, Li Niu, Junjie Chen, Dawei Cheng, and Liqing Zhang. 2019. Learning from Web Data with Memory Module. arXiv preprint (2019)."},{"key":"e_1_3_2_2_52_1","doi-asserted-by":"crossref","unstructured":"Dmitry Ulyanov Andrea Vedaldi and Victor Lempitsky. 2017. Improved texture networks: Maximizing quality and diversity in feed-forward stylization and texture synthesis. In CVPR.  Dmitry Ulyanov Andrea Vedaldi and Victor Lempitsky. 2017. Improved texture networks: Maximizing quality and diversity in feed-forward stylization and texture synthesis. In CVPR.","DOI":"10.1109\/CVPR.2017.437"},{"key":"e_1_3_2_2_53_1","unstructured":"Oriol Vinyals Charles Blundell Timothy Lillicrap Koray Kavukcuoglu and Daan Wierstra. 2016. Matching networks for one shot learning. In NeurIPS.  Oriol Vinyals Charles Blundell Timothy Lillicrap Koray Kavukcuoglu and Daan Wierstra. 2016. Matching networks for one shot learning. In NeurIPS."},{"key":"e_1_3_2_2_54_1","doi-asserted-by":"crossref","unstructured":"Chunyu Wang Yizhou Wang and Alan L Yuille. 2013. An approach to pose-based action recognition. In CVPR.  Chunyu Wang Yizhou Wang and Alan L Yuille. 2013. An approach to pose-based action recognition. In CVPR.","DOI":"10.1109\/CVPR.2013.123"},{"key":"e_1_3_2_2_55_1","doi-asserted-by":"crossref","unstructured":"Limin Wang Yuanjun Xiong Zhe Wang Yu Qiao Dahua Lin Xiaoou Tang and Luc Van Gool. 2016. Temporal segment networks: Towards good practices for deep action recognition. In ECCV.  Limin Wang Yuanjun Xiong Zhe Wang Yu Qiao Dahua Lin Xiaoou Tang and Luc Van Gool. 2016. Temporal segment networks: Towards good practices for deep action recognition. In ECCV.","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"e_1_3_2_2_56_1","doi-asserted-by":"crossref","unstructured":"Xiaolong Wang and Abhinav Gupta. 2018. Videos as space-time region graphs. In ECCV.  Xiaolong Wang and Abhinav Gupta. 2018. Videos as space-time region graphs. In ECCV.","DOI":"10.1007\/978-3-030-01228-1_25"},{"key":"e_1_3_2_2_57_1","doi-asserted-by":"crossref","unstructured":"Yikai Wang Chengming Xu Chen Liu Li Zhang and Yanwei Fu. 2020 a. Instance Credibility Inference for Few-Shot Learning. In CVPR.  Yikai Wang Chengming Xu Chen Liu Li Zhang and Yanwei Fu. 2020 a. Instance Credibility Inference for Few-Shot Learning. In CVPR.","DOI":"10.1109\/CVPR42600.2020.01285"},{"key":"e_1_3_2_2_58_1","volume-title":"2020 b. How to trust unlabeled data? Instance Credibility Inference for Few-Shot Learning. arXiv preprint","author":"Wang Yikai","year":"2020","unstructured":"Yikai Wang , Li Zhang , Yuan Yao , and Yanwei Fu . 2020 b. How to trust unlabeled data? Instance Credibility Inference for Few-Shot Learning. arXiv preprint ( 2020 ). Yikai Wang, Li Zhang, Yuan Yao, and Yanwei Fu. 2020 b. How to trust unlabeled data? Instance Credibility Inference for Few-Shot Learning. arXiv preprint (2020)."},{"key":"e_1_3_2_2_59_1","doi-asserted-by":"crossref","unstructured":"Yu-Xiong Wang Ross Girshick Martial Hebert and Bharath Hariharan. 2018. Low-shot learning from imaginary data. In CVPR.  Yu-Xiong Wang Ross Girshick Martial Hebert and Bharath Hariharan. 2018. Low-shot learning from imaginary data. In CVPR.","DOI":"10.1109\/CVPR.2018.00760"},{"key":"e_1_3_2_2_60_1","doi-asserted-by":"crossref","unstructured":"Yu-Xiong Wang and Martial Hebert. 2016. Learning to learn: Model regression networks for easy small sample learning. In ECCV.  Yu-Xiong Wang and Martial Hebert. 2016. Learning to learn: Model regression networks for easy small sample learning. In ECCV.","DOI":"10.1007\/978-3-319-46466-4_37"},{"key":"e_1_3_2_2_61_1","unstructured":"Yu-Xiong Wang Deva Ramanan and Martial Hebert. 2017. Learning to model the tail. In NeurIPS.  Yu-Xiong Wang Deva Ramanan and Martial Hebert. 2017. Learning to model the tail. In NeurIPS."},{"key":"e_1_3_2_2_62_1","unstructured":"Jianxin Wu Adebola Osuntogun Tanzeem Choudhury Matthai Philipose and James M Rehg. 2007. A scalable approach to activity recognition based on object use. In ICCV.  Jianxin Wu Adebola Osuntogun Tanzeem Choudhury Matthai Philipose and James M Rehg. 2007. A scalable approach to activity recognition based on object use. In ICCV."},{"key":"e_1_3_2_2_63_1","doi-asserted-by":"crossref","unstructured":"Z. Wu Y. Fu Y.-G. Jiang and L. Sigal. 2016. Harnessing Object and Scene Semantics for Large-Scale Video Understanding. In CVPR.  Z. Wu Y. Fu Y.-G. Jiang and L. Sigal. 2016. Harnessing Object and Scene Semantics for Large-Scale Video Understanding. In CVPR.","DOI":"10.1109\/CVPR.2016.339"},{"key":"e_1_3_2_2_64_1","volume-title":"Shared control of a robotic arm using non-invasive brain-computer interface and computer vision guidance. Robotics and Autonomous Systems","author":"Xu Yang","year":"2019","unstructured":"Yang Xu , Cheng Ding , Xiaokang Shu , Kai Gui , Yulia Bezsudnova , Xinjun Sheng , and Dingguo Zhang . 2019. Shared control of a robotic arm using non-invasive brain-computer interface and computer vision guidance. Robotics and Autonomous Systems ( 2019 ). Yang Xu, Cheng Ding, Xiaokang Shu, Kai Gui, Yulia Bezsudnova, Xinjun Sheng, and Dingguo Zhang. 2019. Shared control of a robotic arm using non-invasive brain-computer interface and computer vision guidance. Robotics and Autonomous Systems (2019)."},{"key":"e_1_3_2_2_65_1","doi-asserted-by":"crossref","unstructured":"An Yan Yali Wang Zhifeng Li and Yu Qiao. 2019. PA3D: Pose-action 3D machine for video recognition. In CVPR.  An Yan Yali Wang Zhifeng Li and Yu Qiao. 2019. PA3D: Pose-action 3D machine for video recognition. In CVPR.","DOI":"10.1109\/CVPR.2019.00811"},{"key":"e_1_3_2_2_66_1","doi-asserted-by":"crossref","unstructured":"Sijie Yan Yuanjun Xiong and Dahua Lin. 2018. Spatial temporal graph convolutional networks for skeleton-based action recognition. In AAAI.  Sijie Yan Yuanjun Xiong and Dahua Lin. 2018. Spatial temporal graph convolutional networks for skeleton-based action recognition. In AAAI.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"e_1_3_2_2_67_1","doi-asserted-by":"crossref","unstructured":"Weilong Yang Yang Wang and Greg Mori. 2010. Recognizing human actions from still images with latent poses. In CVPR.  Weilong Yang Yang Wang and Greg Mori. 2010. Recognizing human actions from still images with latent poses. In CVPR.","DOI":"10.1109\/CVPR.2010.5539879"},{"key":"e_1_3_2_2_68_1","volume-title":"Philip HS Torr, and Piotr Koniusz","author":"Zhang Hongguang","year":"2020","unstructured":"Hongguang Zhang , Li Zhang , Xiaojuan Qi , Hongdong Li , Philip HS Torr, and Piotr Koniusz . 2020 . Few-shot Action Recognition with Permutation-invariant Attention. In ECCV. Hongguang Zhang, Li Zhang, Xiaojuan Qi, Hongdong Li, Philip HS Torr, and Piotr Koniusz. 2020. Few-shot Action Recognition with Permutation-invariant Attention. In ECCV."},{"key":"e_1_3_2_2_69_1","volume-title":"NeurIPS workshops.","author":"Zhang Li","year":"2017","unstructured":"Li Zhang , Flood Sung , Feng Liu , Tao Xiang , Shaogang Gong , Yongxin Yang , and Timothy M Hospedales . 2017 a. Actor-critic sequence training for image captioning . In NeurIPS workshops. Li Zhang, Flood Sung, Feng Liu, Tao Xiang, Shaogang Gong, Yongxin Yang, and Timothy M Hospedales. 2017a. Actor-critic sequence training for image captioning. In NeurIPS workshops."},{"key":"e_1_3_2_2_70_1","doi-asserted-by":"crossref","unstructured":"Li Zhang Tao Xiang and Shaogang Gong. 2017b. Learning a deep embedding model for zero-shot learning. In CVPR.  Li Zhang Tao Xiang and Shaogang Gong. 2017b. Learning a deep embedding model for zero-shot learning. In CVPR.","DOI":"10.1109\/CVPR.2017.321"},{"key":"e_1_3_2_2_71_1","doi-asserted-by":"crossref","unstructured":"Zhun Zhong Liang Zheng Guoliang Kang Shaozi Li and Yi Yang. 2020. Random erasing data augmentation. In AAAI.  Zhun Zhong Liang Zheng Guoliang Kang Shaozi Li and Yi Yang. 2020. Random erasing data augmentation. In AAAI.","DOI":"10.1609\/aaai.v34i07.7000"},{"key":"e_1_3_2_2_72_1","doi-asserted-by":"crossref","unstructured":"Bolei Zhou Aditya Khosla Agata Lapedriza Aude Oliva and Antonio Torralba. 2016. Learning deep features for discriminative localization. In CVPR.  Bolei Zhou Aditya Khosla Agata Lapedriza Aude Oliva and Antonio Torralba. 2016. Learning deep features for discriminative localization. In CVPR.","DOI":"10.1109\/CVPR.2016.319"},{"key":"e_1_3_2_2_73_1","unstructured":"Linchao Zhu and Yi Yang. 2018. Compound memory networks for few-shot video classification. In ECCV.  Linchao Zhu and Yi Yang. 2018. Compound memory networks for few-shot video classification. In ECCV."}],"event":{"name":"MM '20: The 28th ACM International Conference on Multimedia","location":"Seattle WA USA","acronym":"MM '20","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 28th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413502","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3394171.3413502","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:47:13Z","timestamp":1750193233000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3394171.3413502"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,12]]},"references-count":73,"alternative-id":["10.1145\/3394171.3413502","10.1145\/3394171"],"URL":"https:\/\/doi.org\/10.1145\/3394171.3413502","relation":{},"subject":[],"published":{"date-parts":[[2020,10,12]]},"assertion":[{"value":"2020-10-12","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}