{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,14]],"date-time":"2026-08-14T16:31:40Z","timestamp":1786725100982,"version":"3.56.0"},"reference-count":57,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2020,8,12]],"date-time":"2020-08-12T00:00:00Z","timestamp":1597190400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100011199","name":"European Research Council","doi-asserted-by":"publisher","award":["realFlow (StG- 2015-637014)"],"award-info":[{"award-number":["realFlow (StG- 2015-637014)"]}],"id":[{"id":"10.13039\/100011199","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100005156","name":"Alexander von Humboldt-Stiftung","doi-asserted-by":"crossref","award":["Sofja Kovalevskaja Award"],"award-info":[{"award-number":["Sofja Kovalevskaja Award"]}],"id":[{"id":"10.13039\/100005156","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2020,8,31]]},"abstract":"<jats:p>\n            Our work explores temporal self-supervision for GAN-based video generation tasks. While adversarial training successfully yields generative models for a variety of areas, temporal relationships in the generated data are much less explored. Natural temporal changes are crucial for sequential generation tasks, e.g. video super-resolution and unpaired video translation. For the former, state-of-the-art methods often favor simpler norm losses such as\n            <jats:italic toggle=\"yes\">L<\/jats:italic>\n            <jats:sup>2<\/jats:sup>\n            over adversarial training. However, their averaging nature easily leads to temporally smooth results with an undesirable lack of spatial detail. For unpaired video translation, existing approaches modify the generator networks to form spatio-temporal cycle consistencies. In contrast, we focus on improving learning objectives and propose a temporally self-supervised algorithm. For both tasks, we show that temporal adversarial learning is key to achieving temporally coherent solutions without sacrificing spatial detail. We also propose a novel Ping-Pong loss to improve the long-term temporal consistency. It effectively prevents recurrent networks from accumulating artifacts temporally without depressing detailed features. Additionally, we propose a first set of metrics to quantitatively evaluate the accuracy as well as the perceptual quality of the temporal evolution. A series of user studies confirm the rankings computed with these metrics. Code, data, models, and results are provided at https:\/\/github.com\/thunil\/TecoGAN.\n          <\/jats:p>","DOI":"10.1145\/3386569.3392457","type":"journal-article","created":{"date-parts":[[2020,8,12]],"date-time":"2020-08-12T11:44:27Z","timestamp":1597232667000},"update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":199,"title":["Learning temporal coherence via self-supervision for GAN-based video generation"],"prefix":"10.1145","volume":"39","author":[{"given":"Mengyu","family":"Chu","sequence":"first","affiliation":[{"name":"Technical University of Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"You","family":"Xie","sequence":"additional","affiliation":[{"name":"Technical University of Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jonas","family":"Mayer","sequence":"additional","affiliation":[{"name":"Technical University of Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Laura","family":"Leal-Taix\u00e9","sequence":"additional","affiliation":[{"name":"Technical University of Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nils","family":"Thuerey","sequence":"additional","affiliation":[{"name":"Technical University of Munich, Germany"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2020,8,12]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"Recycle-GAN: Unsupervised Video Retargeting. In The European Conference on Computer Vision (ECCV).","author":"Bansal Aayush","year":"2018","unstructured":"Aayush Bansal, Shugao Ma, Deva Ramanan, and Yaser Sheikh. 2018. Recycle-GAN: Unsupervised Video Retargeting. In The European Conference on Computer Vision (ECCV)."},{"key":"e_1_2_2_2_1","volume-title":"Unsupervised video-to-video translation. arXiv preprint arXiv:1806.03698","author":"Bashkirova Dina","year":"2018","unstructured":"Dina Bashkirova, Ben Usman, and Kate Saenko. 2018. Unsupervised video-to-video translation. arXiv preprint arXiv:1806.03698 (2018)."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00652"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.2307\/2334029"},{"key":"e_1_2_2_5_1","volume-title":"Large Scale GAN Training for High Fidelity Natural Image Synthesis. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=B1xsqj09Fm","author":"Brock Andrew","year":"2019","unstructured":"Andrew Brock, Jeff Donahue, and Karen Simonyan. 2019. Large Scale GAN Training for High Fidelity Natural Image Synthesis. In International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=B1xsqj09Fm"},{"key":"e_1_2_2_6_1","first-page":"7","article-title":"Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation","volume":"1","author":"Caballero Jose","year":"2017","unstructured":"Jose Caballero, Christian Ledig, Andrew P Aitken, Alejandro Acosta, Johannes Totz, Zehan Wang, and Wenzhe Shi. 2017. Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation.. In CVPR, Vol. 1. 7.","journal-title":"CVPR"},{"key":"e_1_2_2_7_1","volume-title":"Tears of Steel. https:\/\/mango.blender.org\/. Online","author":"Blender Foundation","year":"2018","unstructured":"(CC) Blender Foundation | mango.blender.org. 2011. Tears of Steel. https:\/\/mango.blender.org\/. Online; accessed 15 Nov. 2018."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.126"},{"key":"e_1_2_2_9_1","volume-title":"Mocycle-GAN: Unpaired Video-to-Video Translation. arXiv preprint arXiv:1908.09514 (August","author":"Chen Yang","year":"2019","unstructured":"Yang Chen, Yingwei Pan, Ting Yao, Xinmei Tian, and Tao Mei. 2019. Mocycle-GAN: Unpaired Video-to-Video Translation. arXiv preprint arXiv:1908.09514 (August 2019)."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.316"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.13511"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.5555\/1763974.1764031"},{"key":"e_1_2_2_13_1","unstructured":"Gustav Theodor Fechner and Wilhelm Max Wundt. 1889. Elemente der Psychophysik: erster Theil. Breitkopf & H\u00e4rtel."},{"key":"e_1_2_2_14_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in neural information processing systems. 2672--2680."},{"key":"e_1_2_2_15_1","unstructured":"Ishaan Gulrajani Faruk Ahmed Martin Arjovsky Vincent Dumoulin and Aaron C Courville. 2017. Improved training of wasserstein gans. In Advances in neural information processing systems. 5767--5777."},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00402"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201365"},{"key":"e_1_2_2_18_1","volume-title":"Image-To-Image Translation With Conditional Adversarial Networks. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Isola Phillip","unstructured":"Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. 2017. Image-To-Image Translation With Conditional Adversarial Networks. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3323006"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00340"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356557"},{"key":"e_1_2_2_23_1","volume-title":"Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196","author":"Karras Tero","year":"2017","unstructured":"Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2017. Progressive growing of gans for improved quality, stability, and variation. arXiv preprint arXiv:1710.10196 (2017)."},{"key":"e_1_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3272127.3275065"},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356560"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.182"},{"key":"e_1_2_2_27_1","volume-title":"IEEE Conference on Computer Vision and Pattern Recognition","volume":"2","author":"Lai Wei-Sheng","year":"2017","unstructured":"Wei-Sheng Lai, Jia-Bin Huang, Narendra Ahuja, and Ming-Hsuan Yang. 2017. Deep laplacian pyramid networks for fast and accurate superresolution. In IEEE Conference on Computer Vision and Pattern Recognition, Vol. 2. 5."},{"key":"e_1_2_2_28_1","doi-asserted-by":"crossref","unstructured":"Christian Ledig Lucas Theis Ferenc Husz\u00e1r Jose Caballero Andrew Cunningham Alejandro Acosta Andrew Aitken Alykhan Tejani Johannes Totz Zehan Wang et al. 2016. Photo-realistic single image super-resolution using a generative adversarial network. arXiv:1609.04802 (2016).","DOI":"10.1109\/CVPR.2017.19"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.68"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995614"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.274"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00470"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.304"},{"key":"e_1_2_2_34_1","volume-title":"NTIRE 2019 Challenge on Video Deblurring and Super-Resolution: Dataset and Study. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops.","author":"Nah Seungjun","year":"2019","unstructured":"Seungjun Nah, Sungyong Baik, Seokil Hong, Gyeongsik Moon, Sanghyun Son, Radu Timofte, and Kyoung Mu Lee. 2019. NTIRE 2019 Challenge on Video Deblurring and Super-Resolution: Dataset and Study. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops."},{"key":"e_1_2_2_35_1","volume-title":"Time reversal as self-supervision. arXiv preprint arXiv:1810.01128","author":"Nair Suraj","year":"2018","unstructured":"Suraj Nair, Mohammad Babaeizadeh, Chelsea Finn, Sergey Levine, and Vikash Kumar. 2018. Time reversal as self-supervision. arXiv preprint arXiv:1810.01128 (2018)."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350864"},{"key":"e_1_2_2_37_1","volume-title":"Michael Hirsch, and Bernhard Sch\u00f6lkopf.","author":"P\u00e9rez-Pellitero Eduardo","year":"2018","unstructured":"Eduardo P\u00e9rez-Pellitero, Mehdi SM Sajjadi, Michael Hirsch, and Bernhard Sch\u00f6lkopf. 2018. Photorealistic Video Super Resolution. arXiv preprint arXiv:1807.07930 (2018)."},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00194"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-45886-1_3"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.481"},{"key":"e_1_2_2_41_1","volume-title":"Frame-Recurrent Video Super-Resolution. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR","author":"Sajjadi Mehdi SM","year":"2018","unstructured":"Mehdi SM Sajjadi, Raviteja Vemulapalli, and Matthew Brown. 2018. Frame-Recurrent Video Super-Resolution. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2018)."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.207"},{"key":"e_1_2_2_43_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201333"},{"key":"e_1_2_2_45_1","volume-title":"Detail-Revealing Deep Video Super-Resolution. In The IEEE International Conference on Computer Vision (ICCV).","author":"Tao Xin","year":"2017","unstructured":"Xin Tao, Hongyun Gao, Renjie Liao, Jue Wang, and Jiaya Jia. 2017. Detail-Revealing Deep Video Super-Resolution. In The IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073633"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2836316"},{"key":"e_1_2_2_48_1","unstructured":"Ting-Chun Wang Ming-Yu Liu Jun-Yan Zhu Guilin Liu Andrew Tao Jan Kautz and Bryan Catanzaro. 2018a. Video-to-Video Synthesis. In Advances in Neural Information Processing Systems (NIPS)."},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2019.00247"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00267"},{"key":"e_1_2_2_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3323024"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3322956"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201304"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00824"},{"key":"e_1_2_2_55_1","volume-title":"The unreasonable effectiveness of deep features as a perceptual metric. arXiv preprint","author":"Zhang Richard","year":"2018","unstructured":"Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. 2018. The unreasonable effectiveness of deep features as a perceptual metric. arXiv preprint (2018)."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.244"},{"key":"e_1_2_2_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00953"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3386569.3392457","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3386569.3392457","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,25]],"date-time":"2025-06-25T05:40:15Z","timestamp":1750830015000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3386569.3392457"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,8,12]]},"references-count":57,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2020,8,31]]}},"alternative-id":["10.1145\/3386569.3392457"],"URL":"https:\/\/doi.org\/10.1145\/3386569.3392457","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,8,12]]},"assertion":[{"value":"2020-08-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}