{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,6]],"date-time":"2026-06-06T14:37:21Z","timestamp":1780756641989,"version":"3.54.1"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2019,11,8]],"date-time":"2019-11-08T00:00:00Z","timestamp":1573171200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2019,12,31]]},"abstract":"<jats:p>\n            Automatic generation of a high-quality video from a single image remains a challenging task despite the recent advances in deep generative models. This paper proposes a method that can create a high-resolution, long-term animation using convolutional neural networks (CNNs) from a single landscape image where we mainly focus on skies and waters. Our key observation is that the\n            <jats:italic>motion<\/jats:italic>\n            (e.g., moving clouds) and\n            <jats:italic>appearance<\/jats:italic>\n            (e.g., time-varying colors in the sky) in natural scenes have different time scales. We thus learn them separately and predict them with decoupled control while handling future uncertainty in both predictions by introducing latent codes. Unlike previous methods that infer output frames directly, our CNNs predict spatially-smooth intermediate data, i.e., for motion, flow fields for warping, and for appearance, color transfer maps, via self-supervised learning, i.e., without explicitly-provided ground truth. These intermediate data are applied not to each previous output frame, but to the input image only once for each output frame. This design is crucial to alleviate error accumulation in long-term predictions, which is the essential problem in previous recurrent approaches. The output frames can be looped like cinemagraph, and also be controlled directly by specifying latent codes or indirectly via visual annotations. We demonstrate the effectiveness of our method through comparisons with the state-of-the-arts on video prediction as well as appearance manipulation. Resultant videos, codes, and datasets will be available at http:\/\/www.cgg.cs.tsukuba.ac.jp\/~endo\/projects\/AnimatingLandscape.\n          <\/jats:p>","DOI":"10.1145\/3355089.3356523","type":"journal-article","created":{"date-parts":[[2019,11,8]],"date-time":"2019-11-08T20:27:58Z","timestamp":1573244878000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":32,"title":["Animating landscape"],"prefix":"10.1145","volume":"38","author":[{"given":"Yuki","family":"Endo","sequence":"first","affiliation":[{"name":"University of Tsukuba &amp; Toyohashi University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yoshihiro","family":"Kanamori","sequence":"additional","affiliation":[{"name":"University of Tsukuba"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shigeru","family":"Kuriyama","sequence":"additional","affiliation":[{"name":"Toyohashi University of Technology"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,11,8]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"Stochastic Variational Video Prediction. (4","author":"Babaeizadeh Mohammad","year":"2018","unstructured":"Mohammad Babaeizadeh, Chelsea Finn, Dumitru Erhan, Roy H. Campbell, and Sergey Levine. 2018. Stochastic Variational Video Prediction. (4 2018)."},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185562"},{"key":"e_1_2_2_3_1","volume-title":"Proceedings, Part XVI. 781--797","author":"Byeon Wonmin","year":"2018","unstructured":"Wonmin Byeon, Qin Wang, Rupesh Kumar Srivastava, and Petros Koumoutsakos. 2018. ContextVP: Fully Context-Aware Video Prediction. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8--14, 2018, Proceedings, Part XVI. 781--797."},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1073204.1073273"},{"key":"e_1_2_2_5_1","volume-title":"Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017","author":"Emily","year":"2017","unstructured":"Emily L. Denton and Vighnesh Birodkar. 2017. Unsupervised Learning of Disentangled Representations from Video. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, 4--9 December 2017, Long Beach, CA, USA. 4417--4426."},{"key":"e_1_2_2_6_1","volume-title":"FlowNet: Learning Optical Flow with Convolutional Networks. In 2015 IEEE International Conference on Computer Vision, ICCV 2015","author":"Dosovitskiy Alexey","year":"2015","unstructured":"Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip H\u00e4usser, Caner Hazirbas, Vladimir Golkov, Patrick van der Smagt, Daniel Cremers, and Thomas Brox. 2015. FlowNet: Learning Optical Flow with Convolutional Networks. In 2015 IEEE International Conference on Computer Vision, ICCV 2015, Santiago, Chile, December 7--13, 2015. 2758--2766."},{"key":"e_1_2_2_7_1","volume-title":"Im2Flow: Motion Hallucination from Static Images for Action Recognition. CoRR abs\/1712.04109","author":"Gao Ruohan","year":"2017","unstructured":"Ruohan Gao, Bo Xiong, and Kristen Grauman. 2017. Im2Flow: Motion Hallucination from Static Images for Action Recognition. CoRR abs\/1712.04109 (2017). arXiv:1712.04109"},{"key":"e_1_2_2_8_1","volume-title":"Image Style Transfer Using Convolutional Neural Networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016","author":"Gatys Leon A.","year":"2016","unstructured":"Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. 2016. Image Style Transfer Using Convolutional Neural Networks. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27--30, 2016. 2414--2423."},{"key":"e_1_2_2_9_1","volume-title":"Controllable Video Generation With Sparse Trajectories. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018","author":"Hao Zekun","year":"2018","unstructured":"Zekun Hao, Xun Huang, and Serge J. Belongie. 2018. Controllable Video Generation With Sparse Trajectories. In 2018 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2018, Salt Lake City, UT, USA, June 18--22, 2018. 7854--7863."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/1276377.1276382"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2389824"},{"key":"e_1_2_2_12_1","volume-title":"Proceedings, Part II. 694--711","author":"Johnson Justin","year":"2016","unstructured":"Justin Johnson, Alexandre Alahi, and Li Fei-Fei. 2016. Perceptual Losses for Real-Time Style Transfer and Super-Resolution. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part II. 694--711."},{"key":"e_1_2_2_13_1","volume-title":"Manipulating Attributes of Natural Scenes via Hallucination. CoRR abs\/1808.07413","author":"Karacan Levent","year":"2018","unstructured":"Levent Karacan, Zeynep Akata, Aykut Erdem, and Erkut Erdem. 2018. Manipulating Attributes of Natural Scenes via Hallucination. CoRR abs\/1808.07413 (2018). arXiv:1808.07413"},{"key":"e_1_2_2_14_1","volume-title":"ICLR","author":"Karras Tero","year":"2018","unstructured":"Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2018. Progressive Growing of GANs for Improved Quality, Stability, and Variation. In ICLR 2018."},{"key":"e_1_2_2_15_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. CoRR abs\/1412.6980 (2014). arXiv:1412.6980"},{"key":"e_1_2_2_16_1","volume-title":"Kingma and Max Welling","author":"Diederik","year":"2013","unstructured":"Diederik P. Kingma and Max Welling. 2013. Auto-Encoding Variational Bayes. CoRR abs\/1312.6114 (2013). arXiv:1312.6114 http:\/\/arxiv.org\/abs\/1312.6114"},{"key":"e_1_2_2_17_1","volume-title":"Hinton","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. In Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3--6, 2012, Lake Tahoe, Nevada, United States. 1106--1114."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601101"},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-70139-4"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01240-3_37"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818061"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461950"},{"key":"e_1_2_2_23_1","volume-title":"Cox","author":"Lotter William","year":"2017","unstructured":"William Lotter, Gabriel Kreiman, and David D. Cox. 2017. Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning. (4 2017)."},{"key":"e_1_2_2_24_1","volume-title":"Deep Photo Style Transfer. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017","author":"Luan Fujun","year":"2017","unstructured":"Fujun Luan, Sylvain Paris, Eli Shechtman, and Kavita Bala. 2017. Deep Photo Style Transfer. In 2017 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA, July 21--26, 2017, 6997--7005."},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766903"},{"key":"e_1_2_2_26_1","volume-title":"ICLR'06","author":"Mathieu Michael","year":"2016","unstructured":"Michael Mathieu, Camille Couprie, and Yann Lecun. 2016. Deep multi-scale video prediction beyond mean square error. In ICLR'06."},{"key":"e_1_2_2_27_1","volume-title":"Photorealistic Style Transfer with Screened Poisson Equation. In British Machine Vision Conference 2017, BMVC 2017","author":"Mechrez Roey","year":"2017","unstructured":"Roey Mechrez, Eli Shechtman, and Lihi Zelnik-Manor. 2017. Photorealistic Style Transfer with Screened Poisson Equation. In British Machine Vision Conference 2017, BMVC 2017, London, UK, September 4--7, 2017."},{"key":"e_1_2_2_28_1","volume-title":"So Kweon, and Sing Bing Kang. 2017. Personalized Cinemagraphs Using Semantic Understanding and Collaborative Learning. In IEEE International Conference on Computer Vision, ICCV 2017","author":"Oh Tae-Hyun","year":"2017","unstructured":"Tae-Hyun Oh, Kyungdon Joo, Neel Joshi, Baoyuan Wang, In So Kweon, and Sing Bing Kang. 2017. Personalized Cinemagraphs Using Semantic Understanding and Collaborative Learning. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22--29, 2017. 5170--5179."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2009.01408.x"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8659.2011.02062.x"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00371-016-1337-6"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12940"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.291"},{"key":"e_1_2_2_34_1","volume-title":"a baseline for generative models of natural videos. CoRR abs\/1412.6604","author":"Ranzato Marc'Aurelio","year":"2014","unstructured":"Marc'Aurelio Ranzato, Arthur Szlam, Joan Bruna, Micha\u00ebl Mathieu, Ronan Collobert, and Sumit Chopra. 2014. Video (language) modeling: a baseline for generative models of natural videos. CoRR abs\/1412.6604 (2014). arXiv:1412.6604"},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/38.946629"},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.5555\/3298239.3298457"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0908-3"},{"key":"e_1_2_2_38_1","volume-title":"U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI) (LNCS)","volume":"9351","author":"Ronneberger O.","year":"2015","unstructured":"O. Ronneberger, P. Fischer, and T. Brox. 2015. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI) (LNCS), Vol. 9351. Springer, 234--241. http:\/\/lmb.informatik.uni-freiburg.de\/Publications\/2015\/RFB15a (available on arXiv:1505.04597 [cs.CV])."},{"key":"e_1_2_2_39_1","volume-title":"Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH 2000","author":"Sch\u00f6dl Arno","year":"2000","unstructured":"Arno Sch\u00f6dl, Richard Szeliski, David Salesin, and Irfan A. Essa. 2000. Video textures. In Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, SIGGRAPH 2000, New Orleans, LA, USA, July 23--28, 2000. 489--498."},{"key":"e_1_2_2_40_1","volume-title":"Recognizing Human Actions: A Local SVM Approach. In 17th International Conference on Pattern Recognition, ICPR 2004","author":"Sch\u00fcldt Christian","year":"2004","unstructured":"Christian Sch\u00fcldt, Ivan Laptev, and Barbara Caputo. 2004. Recognizing Human Actions: A Local SVM Approach. In 17th International Conference on Pattern Recognition, ICPR 2004, Cambridge, UK, August 23--26, 2004. 32--36."},{"key":"e_1_2_2_41_1","volume-title":"Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems","author":"Shi Xingjian","year":"2015","unstructured":"Xingjian Shi, Zhourong Chen, Hao Wang, Dit-Yan Yeung, Wai-Kin Wong, and Wangchun Woo. 2015. Convolutional LSTM Network: A Machine Learning Approach for Precipitation Nowcasting. In Advances in Neural Information Processing Systems 28: Annual Conference on Neural Information Processing Systems 2015, December 7--12, 2015, Montreal, Quebec, Canada. 802--810."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/2508363.2508419"},{"key":"e_1_2_2_43_1","unstructured":"K. Simonyan and A. Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR abs\/1409.1556 (2014)."},{"key":"e_1_2_2_44_1","volume-title":"Proceedings of the 32nd International Conference on Machine Learning, ICML 2015","author":"Srivastava Nitish","year":"2015","unstructured":"Nitish Srivastava, Elman Mansimov, and Ruslan Salakhutdinov. 2015. Unsupervised Learning of Video Representations using LSTMs. In Proceedings of the 32nd International Conference on Machine Learning, ICML 2015, Lille, France, 6--11 July 2015. 843--852."},{"key":"e_1_2_2_45_1","volume-title":"2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2005","author":"Tai Yu-Wing","year":"2005","unstructured":"Yu-Wing Tai, Jiaya Jia, and Chi-Keung Tang. 2005. Local Color Transfer via Probabilistic Segmentation by Expectation-Maximization. In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR 2005), 20--26 June 2005, San Diego, CA, USA. 747--754."},{"key":"e_1_2_2_46_1","volume-title":"CoRR abs\/1711.01558","author":"Tolstikhin Ilya O.","year":"2017","unstructured":"Ilya O. Tolstikhin, Olivier Bousquet, Sylvain Gelly, and Bernhard Sch\u00f6lkopf. 2017. Wasserstein Auto-Encoders. CoRR abs\/1711.01558 (2017). arXiv:1711.01558 http:\/\/arxiv.org\/abs\/1711.01558"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1145\/2897824.2925942"},{"key":"e_1_2_2_48_1","volume-title":"Generating Videos with Scene Dynamics. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016","author":"Vondrick Carl","year":"2016","unstructured":"Carl Vondrick, Hamed Pirsiavash, and Antonio Torralba. 2016. Generating Videos with Scene Dynamics. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5--10, 2016, Barcelona, Spain. 613--621."},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.281"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00917"},{"key":"e_1_2_2_51_1","volume-title":"Occlusion Aware Unsupervised Learning of Optical Flow. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).","author":"Wang Yang","year":"2018","unstructured":"Yang Wang, Yi Yang, Zhenheng Yang, Liang Zhao, Peng Wang, and Wei Xu. 2018b. Occlusion Aware Unsupervised Learning of Optical Flow. In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR)."},{"key":"e_1_2_2_52_1","volume-title":"DeepFlow: Large Displacement Optical Flow with Deep Matching. In IEEE International Conference on Computer Vision, ICCV 2013","author":"Weinzaepfel Philippe","year":"2013","unstructured":"Philippe Weinzaepfel, J\u00e9r\u00f4me Revaud, Za\u00efd Harchaoui, and Cordelia Schmid. 2013. DeepFlow: Large Displacement Optical Flow with Deep Matching. In IEEE International Conference on Computer Vision, ICCV 2013, Sydney, Australia, December 1--8, 2013. 1385--1392."},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.12008"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00251"},{"key":"e_1_2_2_55_1","volume-title":"Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016","author":"Xue Tianfan","year":"2016","unstructured":"Tianfan Xue, Jiajun Wu, Katherine L. Bouman, and Bill Freeman. 2016. Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks. In Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5--10, 2016, Barcelona, Spain. 91--99."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_2_57_1","volume-title":"Proceedings, Part VIII. 262--277","author":"Zhou Yipin","unstructured":"Yipin Zhou and Tamara L. Berg. 2016. Learning Temporal Transformations from Time-Lapse Videos. In Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11--14, 2016, Proceedings, Part VIII. 262--277."},{"key":"e_1_2_2_58_1","volume-title":"Toward Multimodal Image-to-Image Translation. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017","author":"Zhu Jun-Yan","year":"2017","unstructured":"Jun-Yan Zhu, Richard Zhang, Deepak Pathak, Trevor Darrell, Alexei A. Efros, Oliver Wang, and Eli Shechtman. 2017. Toward Multimodal Image-to-Image Translation. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, 4--9 December 2017, Long Beach, CA, USA. 465--476."}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3355089.3356523","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3355089.3356523","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:44:41Z","timestamp":1750203881000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3355089.3356523"}},"subtitle":["self-supervised learning of decoupled motion and appearance for single-image video synthesis"],"short-title":[],"issued":{"date-parts":[[2019,11,8]]},"references-count":58,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2019,12,31]]}},"alternative-id":["10.1145\/3355089.3356523"],"URL":"https:\/\/doi.org\/10.1145\/3355089.3356523","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,11,8]]},"assertion":[{"value":"2019-11-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}