{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T16:33:09Z","timestamp":1780763589500,"version":"3.54.1"},"publisher-location":"New York, NY, USA","reference-count":40,"publisher":"ACM","license":[{"start":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T00:00:00Z","timestamp":1665360000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"SZSTC Grant","award":["No.JCYJ20190809172201639 and WDZC20200820200655001"],"award-info":[{"award-number":["No.JCYJ20190809172201639 and WDZC20200820200655001"]}]},{"name":"Shen- zhen Key Laboratory","award":["ZDSYS20210623092001004"],"award-info":[{"award-number":["ZDSYS20210623092001004"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2022,10,10]]},"DOI":"10.1145\/3503161.3548395","type":"proceedings-article","created":{"date-parts":[[2022,10,10]],"date-time":"2022-10-10T15:43:12Z","timestamp":1665416592000},"page":"779-789","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":15,"title":["DeViT"],"prefix":"10.1145","author":[{"given":"Jiayin","family":"Cai","sequence":"first","affiliation":[{"name":"Kuaishou Technology, BeiJing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Changlin","family":"Li","sequence":"additional","affiliation":[{"name":"Kuaishou Technology, ShenZhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xin","family":"Tao","sequence":"additional","affiliation":[{"name":"Kuaishou Technology, ShenZhen, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chun","family":"Yuan","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yu-Wing","family":"Tai","sequence":"additional","affiliation":[{"name":"Kuaishou Technology, Hong Kong, Hong Kong"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,10,10]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/83.935036"},{"key":"e_1_3_2_1_2_1","volume-title":"Kuan-Ying Lee, and Winston Hsu.","author":"Chang Ya-Liang","year":"2019","unstructured":"Ya-Liang Chang , Zhe Yu Liu , Kuan-Ying Lee, and Winston Hsu. 2019 . Learnable gated temporal shift module for deep video inpainting. arXiv preprint arXiv:1907.01131 (2019). Ya-Liang Chang, Zhe Yu Liu, Kuan-Ying Lee, and Winston Hsu. 2019. Learnable gated temporal shift module for deep video inpainting. arXiv preprint arXiv:1907.01131 (2019)."},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.89"},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185578"},{"key":"e_1_3_2_1_5_1","volume-title":"Fuseformer: Fusing fine-grained information in transformers for video inpainting. In ICCV.","author":"Liu R","year":"2021","unstructured":"Liu R et al. 2021 . Fuseformer: Fusing fine-grained information in transformers for video inpainting. In ICCV. Liu R et al. 2021. Fuseformer: Fusing fine-grained information in transformers for video inpainting. In ICCV."},{"key":"e_1_3_2_1_6_1","volume-title":"Flow-edge guided video completion","author":"Gao Chen","unstructured":"Chen Gao , Ayush Saraf , Jia-Bin Huang , and Johannes Kopf . 2020. Flow-edge guided video completion . In ECCV. Springer , 713--729. Chen Gao, Ayush Saraf, Jia-Bin Huang, and Johannes Kopf. 2020. Flow-edge guided video completion. In ECCV. Springer, 713--729."},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"crossref","unstructured":"Leon A Gatys Alexander S Ecker and Matthias Bethge. 2016. Image style transfer using convolutional neural networks. In CVPR. 2414--2423.  Leon A Gatys Alexander S Ecker and Matthias Bethge. 2016. Image style transfer using convolutional neural networks. In CVPR. 2414--2423.","DOI":"10.1109\/CVPR.2016.265"},{"key":"e_1_3_2_1_8_1","volume-title":"James Tompkin, Jan Kautz, and Christian Theobalt.","author":"Granados Miguel","year":"2012","unstructured":"Miguel Granados , Kwang In Kim , James Tompkin, Jan Kautz, and Christian Theobalt. 2012 . Background inpainting for videos with dynamic objects and a free-moving camera. In ECCV. Springer , 682--695. Miguel Granados, Kwang In Kim, James Tompkin, Jan Kautz, and Christian Theobalt. 2012. Background inpainting for videos with dynamic objects and a free-moving camera. In ECCV. Springer, 682--695."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1587\/transinf.2020EDP7194"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2980179.2982398","article-title":"Temporally coherent completion of dynamic video","volume":"35","author":"Huang Jia-Bin","year":"2016","unstructured":"Jia-Bin Huang , Sing Bing Kang , Narendra Ahuja , and Johannes Kopf . 2016 . Temporally coherent completion of dynamic video . ACM Transactions on Graphics (TOG) 35 , 6 (2016), 1 -- 11 . Jia-Bin Huang, Sing Bing Kang, Narendra Ahuja, and Johannes Kopf. 2016. Temporally coherent completion of dynamic video. ACM Transactions on Graphics (TOG) 35, 6 (2016), 1--11.","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_3_2_1_11_1","unstructured":"Max Jaderberg Karen Simonyan Andrew Zisserman and Koray Kavukcuoglu. 2015. Spatial transformer networks. (2015).  Max Jaderberg Karen Simonyan Andrew Zisserman and Koray Kavukcuoglu. 2015. Spatial transformer networks. (2015)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"crossref","unstructured":"Dahun Kim SanghyunWoo Joon-Young Lee and In So Kweon. 2019. Deep video inpainting. In CVPR. 5792--5801.  Dahun Kim SanghyunWoo Joon-Young Lee and In So Kweon. 2019. Deep video inpainting. In CVPR. 5792--5801.","DOI":"10.1109\/CVPR.2019.00594"},{"key":"e_1_3_2_1_14_1","unstructured":"Wei-Sheng Lai Jia-Bin Huang Oliver Wang Eli Shechtman Ersin Yumer and Ming-Hsuan Yang. 2018. Learning blind video temporal consistency. In ECCV. 170--185.  Wei-Sheng Lai Jia-Bin Huang Oliver Wang Eli Shechtman Ersin Yumer and Ming-Hsuan Yang. 2018. Learning blind video temporal consistency. In ECCV. 170--185."},{"key":"e_1_3_2_1_15_1","volume-title":"DaeYeun Won, and Seon Joo Kim.","author":"Lee Sungho","year":"2019","unstructured":"Sungho Lee , Seoung Wug Oh , DaeYeun Won, and Seon Joo Kim. 2019 . Copy-andpaste networks for deep video inpainting. In ICCV. 4413--4421. Sungho Lee, Seoung Wug Oh, DaeYeun Won, and Seon Joo Kim. 2019. Copy-andpaste networks for deep video inpainting. In ICCV. 4413--4421."},{"key":"e_1_3_2_1_16_1","first-page":"305","article-title":"Learning How to Inpaint from Global Image Statistics","volume":"1","author":"Levin Anat","year":"2003","unstructured":"Anat Levin , Assaf Zomet , and Yair Weiss . 2003 . Learning How to Inpaint from Global Image Statistics .. In ICCV , Vol. 1. 305 -- 312 . Anat Levin, Assaf Zomet, and Yair Weiss. 2003. Learning How to Inpaint from Global Image Statistics.. In ICCV, Vol. 1. 305--312.","journal-title":"ICCV"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58548-8_42"},{"key":"e_1_3_2_1_18_1","volume-title":"Tsm: Temporal shift module for efficient video understanding. In CVPR. 7083--7093.","author":"Lin Ji","year":"2019","unstructured":"Ji Lin , Chuang Gan , and Song Han . 2019 . Tsm: Temporal shift module for efficient video understanding. In CVPR. 7083--7093. Ji Lin, Chuang Gan, and Song Han. 2019. Tsm: Temporal shift module for efficient video understanding. In CVPR. 7083--7093."},{"key":"e_1_3_2_1_19_1","unstructured":"Guilin Liu Fitsum A Reda Kevin J Shih Ting-ChunWang Andrew Tao and Bryan Catanzaro. 2018. Image in painting for irregular holes using partial convolutions. In ECCV. 85--100.  Guilin Liu Fitsum A Reda Kevin J Shih Ting-ChunWang Andrew Tao and Bryan Catanzaro. 2018. Image in painting for irregular holes using partial convolutions. In ECCV. 85--100."},{"key":"e_1_3_2_1_20_1","volume-title":"Decoupled Spatial-Temporal Transformer for Video Inpainting. arXiv preprint arXiv:2104.06637","author":"Liu Rui","year":"2021","unstructured":"Rui Liu , Hanming Deng , Yangyi Huang , Xiaoyu Shi , Lewei Lu , Wenxiu Sun , Xiaogang Wang , Jifeng Dai , and Hongsheng Li. 2021. Decoupled Spatial-Temporal Transformer for Video Inpainting. arXiv preprint arXiv:2104.06637 ( 2021 ). Rui Liu, Hanming Deng, Yangyi Huang, Xiaoyu Shi, Lewei Lu, Wenxiu Sun, Xiaogang Wang, Jifeng Dai, and Hongsheng Li. 2021. Decoupled Spatial-Temporal Transformer for Video Inpainting. arXiv preprint arXiv:2104.06637 (2021)."},{"key":"e_1_3_2_1_21_1","unstructured":"Loshchilov Ilya Hutter and Frank. 2018. Decoupled Weight Decay Regularization. ICLR.  Loshchilov Ilya Hutter and Frank. 2018. Decoupled Weight Decay Regularization. ICLR."},{"key":"e_1_3_2_1_22_1","volume-title":"Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957","author":"Miyato Takeru","year":"2018","unstructured":"Takeru Miyato , Toshiki Kataoka , Masanori Koyama , and Yuichi Yoshida . 2018. Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957 ( 2018 ). Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. 2018. Spectral normalization for generative adversarial networks. arXiv preprint arXiv:1802.05957 (2018)."},{"key":"e_1_3_2_1_23_1","volume-title":"Edgeconnect: Generative image inpainting with adversarial edge learning. arXiv preprint arXiv:1901.00212","author":"Nazeri Kamyar","year":"2019","unstructured":"Kamyar Nazeri , Eric Ng , Tony Joseph , Faisal Z Qureshi , and Mehran Ebrahimi . 2019 . Edgeconnect: Generative image inpainting with adversarial edge learning. arXiv preprint arXiv:1901.00212 (2019). Kamyar Nazeri, Eric Ng, Tony Joseph, Faisal Z Qureshi, and Mehran Ebrahimi. 2019. Edgeconnect: Generative image inpainting with adversarial edge learning. arXiv preprint arXiv:1901.00212 (2019)."},{"key":"e_1_3_2_1_24_1","volume-title":"Video inpainting of complex scenes. Siam journal on imaging sciences 7, 4","author":"Newson Alasdair","year":"2014","unstructured":"Alasdair Newson , Andr\u00e9s Almansa , Matthieu Fradet , Yann Gousseau , and Patrick P\u00e9rez . 2014. Video inpainting of complex scenes. Siam journal on imaging sciences 7, 4 ( 2014 ), 1993--2019. Alasdair Newson, Andr\u00e9s Almansa, Matthieu Fradet, Yann Gousseau, and Patrick P\u00e9rez. 2014. Video inpainting of complex scenes. Siam journal on imaging sciences 7, 4 (2014), 1993--2019."},{"key":"e_1_3_2_1_25_1","unstructured":"Seoung Wug Oh Sungho Lee Joon-Young Lee and Seon Joo Kim. 2019. Onionpeel networks for deep video completion. In ICCV. 4403--4412.  Seoung Wug Oh Sungho Lee Joon-Young Lee and Seon Joo Kim. 2019. Onionpeel networks for deep video completion. In ICCV. 4403--4412."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.278"},{"key":"e_1_3_2_1_27_1","volume-title":"Markus Gross, and Alexander Sorkine-Hornung.","author":"Perazzi Federico","year":"2016","unstructured":"Federico Perazzi , Jordi Pont-Tuset , Brian McWilliams , Luc Van Gool , Markus Gross, and Alexander Sorkine-Hornung. 2016 . A benchmark dataset and evaluation methodology for video object segmentation. In CVPR. 724--732. Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. 2016. A benchmark dataset and evaluation methodology for video object segmentation. In CVPR. 724--732."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2008.4587842"},{"key":"e_1_3_2_1_29_1","unstructured":"Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008.  Ashish Vaswani Noam Shazeer Niki Parmar Jakob Uszkoreit Llion Jones Aidan N Gomez \u0141ukasz Kaiser and Illia Polosukhin. 2017. Attention is all you need. In Advances in neural information processing systems. 5998--6008."},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33015232"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2836316"},{"key":"e_1_3_2_1_32_1","volume-title":"Video-to-video synthesis. arXiv preprint arXiv:1808.06601","author":"Wang Ting-Chun","year":"2018","unstructured":"Ting-Chun Wang , Ming-Yu Liu , Jun-Yan Zhu , Guilin Liu , Andrew Tao , Jan Kautz , and Bryan Catanzaro . 2018. Video-to-video synthesis. arXiv preprint arXiv:1808.06601 ( 2018 ). Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, and Bryan Catanzaro. 2018. Video-to-video synthesis. arXiv preprint arXiv:1808.06601 (2018)."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2004.1315022"},{"key":"e_1_3_2_1_34_1","volume-title":"Align-and-attend network for globally and locally coherent video in painting. arXiv preprint arXiv:1905.13066","author":"Woo Sanghyun","year":"2019","unstructured":"Sanghyun Woo , Dahun Kim , KwanYong Park , Joon-Young Lee , and In So Kweon . 2019. Align-and-attend network for globally and locally coherent video in painting. arXiv preprint arXiv:1905.13066 ( 2019 ). Sanghyun Woo, Dahun Kim, KwanYong Park, Joon-Young Lee, and In So Kweon. 2019. Align-and-attend network for globally and locally coherent video in painting. arXiv preprint arXiv:1905.13066 (2019)."},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"crossref","unstructured":"Rui Xu Xiaoxiao Li Bolei Zhou and Chen Change Loy. 2019. Deep flow-guided video inpainting. In CVPR. 3723--3732.  Rui Xu Xiaoxiao Li Bolei Zhou and Chen Change Loy. 2019. Deep flow-guided video inpainting. In CVPR. 3723--3732.","DOI":"10.1109\/CVPR.2019.00384"},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"crossref","unstructured":"Fuzhi Yang Huan Yang Jianlong Fu Hongtao Lu and Baining Guo. 2020. Learning Texture Transformer Network for Image Super-Resolution. In CVPR. 5791--5800.  Fuzhi Yang Huan Yang Jianlong Fu Hongtao Lu and Baining Guo. 2020. Learning Texture Transformer Network for Image Super-Resolution. In CVPR. 5791--5800.","DOI":"10.1109\/CVPR42600.2020.00583"},{"key":"e_1_3_2_1_37_1","unstructured":"Jiahui Yu Zhe Lin Jimei Yang Xiaohui Shen Xin Lu and Thomas S Huang. 2019. Free-form image inpainting with gated convolution. In ICCV. 4471--4480.  Jiahui Yu Zhe Lin Jimei Yang Xiaohui Shen Xin Lu and Thomas S Huang. 2019. Free-form image inpainting with gated convolution. In ICCV. 4471--4480."},{"key":"e_1_3_2_1_38_1","volume-title":"Learning Joint Spatial-Temporal Transformations for Video In painting","author":"Zeng Yanhong","unstructured":"Yanhong Zeng , Jianlong Fu , and Hongyang Chao . 2020. Learning Joint Spatial-Temporal Transformations for Video In painting . In ECCV. Springer , 528--543. Yanhong Zeng, Jianlong Fu, and Hongyang Chao. 2020. Learning Joint Spatial-Temporal Transformations for Video In painting. In ECCV. Springer, 528--543."},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Yanhong Zeng Jianlong Fu Hongyang Chao and Baining Guo. 2019. Learning pyramid-context encoder network for high-quality image in painting. In CVPR. 1486--1494.  Yanhong Zeng Jianlong Fu Hongyang Chao and Baining Guo. 2019. Learning pyramid-context encoder network for high-quality image in painting. In CVPR. 1486--1494.","DOI":"10.1109\/CVPR.2019.00158"},{"key":"e_1_3_2_1_40_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 16448--16457","author":"Zou Xueyan","year":"2021","unstructured":"Xueyan Zou , Linjie Yang , Ding Liu , and Yong Jae Lee . 2021 . Progressive temporal feature alignment network for video in painting . In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 16448--16457 . Xueyan Zou, Linjie Yang, Ding Liu, and Yong Jae Lee. 2021. Progressive temporal feature alignment network for video in painting. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 16448--16457."}],"event":{"name":"MM '22: The 30th ACM International Conference on Multimedia","location":"Lisboa Portugal","acronym":"MM '22","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 30th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548395","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3503161.3548395","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T19:00:44Z","timestamp":1750186844000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3503161.3548395"}},"subtitle":["Deformed Vision Transformers in Video Inpainting"],"short-title":[],"issued":{"date-parts":[[2022,10,10]]},"references-count":40,"alternative-id":["10.1145\/3503161.3548395","10.1145\/3503161"],"URL":"https:\/\/doi.org\/10.1145\/3503161.3548395","relation":{},"subject":[],"published":{"date-parts":[[2022,10,10]]},"assertion":[{"value":"2022-10-10","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}