{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T05:25:04Z","timestamp":1755926704138,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":48,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Chinese National Science Funding","award":["62132006"],"award-info":[{"award-number":["62132006"]}]},{"name":"MoE-China Mobile Research Fund Project","award":["MCM20180702"],"award-info":[{"award-number":["MCM20180702"]}]},{"name":"Shanghai Key Laboratory of Digital Media Processing and Transmissions"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548011","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:42:35Z","timestamp":1665416555000},"page":"5201-5209","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Multi-Scale Coarse-to-Fine Transformer for Frame Interpolation"],"prefix":"10.1145","author":[{"given":"Chen","family":"Li","sequence":"first","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Li","family":"Song","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xueyi","family":"Zou","sequence":"additional","affiliation":[{"name":"Huawei Noah's Ark Lab, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiaming","family":"Guo","sequence":"additional","affiliation":[{"name":"Huawei Noah's Ark Lab, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Youliang","family":"Yan","sequence":"additional","affiliation":[{"name":"Huawei Noah's Ark Lab, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenjun","family":"Zhang","sequence":"additional","affiliation":[{"name":"Shanghai Jiao Tong University, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00382"},{"key":"e_1_3_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2941941"},{"key":"e_1_3_2_2_3_1","volume-title":"Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation. arXiv preprint arXiv:2105.05537","author":"Cao Hu","year":"2021","unstructured":"Hu Cao , Yueyue Wang , Joy Chen , Dongsheng Jiang , Xiaopeng Zhang , Qi Tian , and Manning Wang . 2021b. Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation. arXiv preprint arXiv:2105.05537 ( 2021 ). Hu Cao, Yueyue Wang, Joy Chen, Dongsheng Jiang, Xiaopeng Zhang, Qi Tian, and Manning Wang. 2021b. Swin-Unet: Unet-like Pure Transformer for Medical Image Segmentation. arXiv preprint arXiv:2105.05537 (2021)."},{"key":"e_1_3_2_2_4_1","volume-title":"Video Super-Resolution Transformer. arXiv preprint arXiv:2106.06847","author":"Cao Jiezhang","year":"2021","unstructured":"Jiezhang Cao , Yawei Li , Kai Zhang , and Luc Van Gool . 2021a. Video Super-Resolution Transformer. arXiv preprint arXiv:2106.06847 ( 2021 ). Jiezhang Cao, Yawei Li, Kai Zhang, and Luc Van Gool. 2021a. Video Super-Resolution Transformer. arXiv preprint arXiv:2106.06847 (2021)."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01212"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6634"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58583-9_7"},{"key":"e_1_3_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV48630.2021.00064"},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6693"},{"key":"e_1_3_2_2_11_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly etal 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020).  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00180"},{"key":"e_1_3_2_2_13_1","volume-title":"RIFE: Real-Time Intermediate Flow Estimation for Video Frame Interpolation. arXiv preprint arXiv:2011.06294","author":"Huang Zhewei","year":"2020","unstructured":"Zhewei Huang , Tianyuan Zhang , Wen Heng , Boxin Shi , and Shuchang Zhou . 2020 . RIFE: Real-Time Intermediate Flow Estimation for Video Frame Interpolation. arXiv preprint arXiv:2011.06294 (2020). Zhewei Huang, Tianyuan Zhang, Wen Heng, Boxin Shi, and Shuchang Zhou. 2020. RIFE: Real-Time Intermediate Flow Estimation for Video Frame Interpolation. arXiv preprint arXiv:2011.06294 (2020)."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00938"},{"key":"e_1_3_2_2_15_1","volume-title":"Flavr: Flow-agnostic video representations for fast frame interpolation. arXiv preprint arXiv:2012.08512","author":"Kalluri Tarun","year":"2020","unstructured":"Tarun Kalluri , Deepak Pathak , Manmohan Chandraker , and Du Tran . 2020 . Flavr: Flow-agnostic video representations for fast frame interpolation. arXiv preprint arXiv:2012.08512 (2020). Tarun Kalluri, Deepak Pathak, Manmohan Chandraker, and Du Tran. 2020. Flavr: Flow-agnostic video representations for fast frame interpolation. arXiv preprint arXiv:2012.08512 (2020)."},{"key":"e_1_3_2_2_16_1","unstructured":"Angelos Katharopoulos Apoorv Vyas Nikolaos Pappas and Francc ois Fleuret. 2020. Transformers are rnns: Fast autoregressive transformers with linear attention. In Int'l Conf. on Machine Learning. PMLR 5156--5165.  Angelos Katharopoulos Apoorv Vyas Nikolaos Pappas and Francc ois Fleuret. 2020. Transformers are rnns: Fast autoregressive transformers with linear attention. In Int'l Conf. on Machine Learning. PMLR 5156--5165."},{"key":"e_1_3_2_2_17_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba . 2014 . Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014). Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00536"},{"key":"e_1_3_2_2_19_1","volume-title":"Localvit: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707","author":"Li Yawei","year":"2021","unstructured":"Yawei Li , Kai Zhang , Jiezhang Cao , Radu Timofte , and Luc Van Gool . 2021 . Localvit: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707 (2021). Yawei Li, Kai Zhang, Jiezhang Cao, Radu Timofte, and Luc Van Gool. 2021. Localvit: Bringing locality to vision transformers. arXiv preprint arXiv:2104.05707 (2021)."},{"key":"e_1_3_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW54120.2021.00210"},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-66823-5_3"},{"key":"e_1_3_2_2_22_1","volume-title":"Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030","author":"Liu Ze","year":"2021","unstructured":"Ze Liu , Yutong Lin , Yue Cao , Han Hu , Yixuan Wei , Zheng Zhang , Stephen Lin , and Baining Guo . 2021. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030 ( 2021 ). Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. arXiv preprint arXiv:2103.14030 (2021)."},{"key":"e_1_3_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.478"},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00059"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.35"},{"key":"e_1_3_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00183"},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00548"},{"key":"e_1_3_2_2_28_1","volume-title":"Proceedings of the IEEE Conf. on Computer Vision and Pattern Recognition. 670--679","author":"Niklaus Simon","year":"2017","unstructured":"Simon Niklaus , Long Mai , and Feng Liu . 2017 . Video frame interpolation via adaptive convolution . In Proceedings of the IEEE Conf. on Computer Vision and Pattern Recognition. 670--679 . Simon Niklaus, Long Mai, and Feng Liu. 2017. Video frame interpolation via adaptive convolution. In Proceedings of the IEEE Conf. on Computer Vision and Pattern Recognition. 670--679."},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58568-6_7"},{"key":"e_1_3_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01427"},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.85"},{"volume-title":"Int'l Conf. on Pattern Recognition","author":"Plizzari Chiara","key":"e_1_3_2_2_32_1","unstructured":"Chiara Plizzari , Marco Cannici , and Matteo Matteucci . 2021. Spatial temporal transformer network for skeleton-based action recognition . In Int'l Conf. on Pattern Recognition . Springer , 694--701. Chiara Plizzari, Marco Cannici, and Matteo Matteucci. 2021. Spatial temporal transformer network for skeleton-based action recognition. In Int'l Conf. on Pattern Recognition. Springer, 694--701."},{"key":"e_1_3_2_2_33_1","volume-title":"Stand-alone self-attention in vision models. arXiv preprint arXiv:1906.05909","author":"Ramachandran Prajit","year":"2019","unstructured":"Prajit Ramachandran , Niki Parmar , Ashish Vaswani , Irwan Bello , Anselm Levskaya , and Jonathon Shlens . 2019. Stand-alone self-attention in vision models. arXiv preprint arXiv:1906.05909 ( 2019 ). Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jonathon Shlens. 2019. Stand-alone self-attention in vision models. arXiv preprint arXiv:1906.05909 (2019)."},{"key":"e_1_3_2_2_34_1","volume-title":"Proceedings of the Asian Conf. on Computer Vision.","author":"Shi Lei","year":"2020","unstructured":"Lei Shi , Yifan Zhang , Jian Cheng , and Hanqing Lu . 2020 . Decoupled Spatial-Temporal Attention Network for Skeleton-Based Action-Gesture Recognition . In Proceedings of the Asian Conf. on Computer Vision. Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. 2020. Decoupled Spatial-Temporal Attention Network for Skeleton-Based Action-Gesture Recognition. In Proceedings of the Asian Conf. on Computer Vision."},{"key":"e_1_3_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01696"},{"key":"e_1_3_2_2_36_1","volume-title":"Amir Roshan Zamir, and Mubarak Shah","author":"Soomro Khurram","year":"2012","unstructured":"Khurram Soomro , Amir Roshan Zamir, and Mubarak Shah . 2012 . UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012). Khurram Soomro, Amir Roshan Zamir, and Mubarak Shah. 2012. UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402 (2012)."},{"key":"e_1_3_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.33"},{"key":"e_1_3_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00881"},{"key":"e_1_3_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.514"},{"key":"e_1_3_2_2_40_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez Lukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008."},{"key":"e_1_3_2_2_41_1","volume-title":"Uformer: A General U-Shaped Transformer for Image Restoration. arXiv preprint arXiv:2106.03106","author":"Wang Zhendong","year":"2021","unstructured":"Zhendong Wang , Xiaodong Cun , Jianmin Bao , and Jianzhuang Liu . 2021 . Uformer: A General U-Shaped Transformer for Image Restoration. arXiv preprint arXiv:2106.03106 (2021). Zhendong Wang, Xiaodong Cun, Jianmin Bao, and Jianzhuang Liu. 2021. Uformer: A General U-Shaped Transformer for Image Restoration. arXiv preprint arXiv:2106.03106 (2021)."},{"key":"e_1_3_2_2_42_1","volume-title":"Visual transformers: Token-based image representation and processing for computer vision. arXiv preprint arXiv:2006.03677","author":"Wu Bichen","year":"2020","unstructured":"Bichen Wu , Chenfeng Xu , Xiaoliang Dai , Alvin Wan , Peizhao Zhang , Zhicheng Yan , Masayoshi Tomizuka , Joseph Gonzalez , Kurt Keutzer , and Peter Vajda . 2020. Visual transformers: Token-based image representation and processing for computer vision. arXiv preprint arXiv:2006.03677 ( 2020 ). Bichen Wu, Chenfeng Xu, Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Zhicheng Yan, Masayoshi Tomizuka, Joseph Gonzalez, Kurt Keutzer, and Peter Vajda. 2020. Visual transformers: Token-based image representation and processing for computer vision. arXiv preprint arXiv:2006.03677 (2020)."},{"key":"e_1_3_2_2_43_1","first-page":"1647","article-title":"Quadratic Video Interpolation","volume":"32","author":"Xu Xiangyu","year":"2019","unstructured":"Xiangyu Xu , Li Siyao , Wenxiu Sun , Qian Yin , and Ming-Hsuan Yang . 2019 . Quadratic Video Interpolation . Advances in Neural Information Processing Systems , Vol. 32 (2019), 1647 -- 1656 . Xiangyu Xu, Li Siyao, Wenxiu Sun, Qian Yin, and Ming-Hsuan Yang. 2019. Quadratic Video Interpolation. Advances in Neural Information Processing Systems, Vol. 32 (2019), 1647--1656.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-018-01144-2"},{"key":"e_1_3_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58517-4_31"},{"key":"e_1_3_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2940510"},{"key":"e_1_3_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00681"},{"key":"e_1_3_2_2_48_1","volume-title":"Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159","author":"Zhu Xizhou","year":"2020","unstructured":"Xizhou Zhu , Weijie Su , Lewei Lu , Bin Li , Xiaogang Wang , and Jifeng Dai . 2020. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159 ( 2020 ). Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2020. Deformable detr: Deformable transformers for end-to-end object detection. arXiv preprint arXiv:2010.04159 (2020)."}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Lisboa Portugal","acronym":"MM '22"},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548011","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548011","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:02:29Z","timestamp":1750186949000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548011"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":48,"alternative-id":["10.1145\/3503161.3548011","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548011","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}