{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T20:44:27Z","timestamp":1761597867488,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":37,"publisher":"ACM","license":[{"start":{"date-parts":[[2017,10,23]],"date-time":"2017-10-23T00:00:00Z","timestamp":1508716800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Guangdong Science and Technology Project","award":["2014B010117007"],"award-info":[{"award-number":["2014B010117007"]}]},{"name":"Shenzhen Peacock Plan","award":["20130408-183003656"],"award-info":[{"award-number":["20130408-183003656"]}]},{"name":"Shenzhen Key Laboratory for Intelligent Multimedia and Virtual Reality","award":["ZDSYS201703031405467"],"award-info":[{"award-number":["ZDSYS201703031405467"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2017,10,23]]},"DOI":"10.1145\/3123266.3123349","type":"proceedings-article","created":{"date-parts":[[2017,10,20]],"date-time":"2017-10-20T13:04:26Z","timestamp":1508504666000},"page":"1503-1512","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":18,"title":["Learning Object-Centric Transformation for Video Prediction"],"prefix":"10.1145","author":[{"given":"Xiongtao","family":"Chen","sequence":"first","affiliation":[{"name":"Peking University, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenmin","family":"Wang","sequence":"additional","affiliation":[{"name":"Peking University, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jinzhuo","family":"Wang","sequence":"additional","affiliation":[{"name":"Peking University, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Weimian","family":"Li","sequence":"additional","affiliation":[{"name":"Peking University, Shenzhen, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2017,10,23]]},"reference":[{"volume-title":"Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467","year":"2016","author":"Abadi Mart\u00e9n","key":"e_1_3_2_1_1_1"},{"volume-title":"Multiple object recognition with visual attention. arXiv preprint arXiv:1412.7755","year":"2014","author":"Ba Jimmy","key":"e_1_3_2_1_2_1"},{"key":"e_1_3_2_1_3_1","unstructured":"Bert De Brabandere Xu Jia Tinne Tuytelaars and Luc Van Gool. 2016. Dynamic filter networks. In Neural Information Processing Systems (NIPS).  Bert De Brabandere Xu Jia Tinne Tuytelaars and Luc Van Gool. 2016. Dynamic filter networks. In Neural Information Processing Systems (NIPS)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.316"},{"key":"e_1_3_2_1_5_1","unstructured":"Chelsea Finn Ian Goodfellow and Sergey Levine. 2016. Unsuper- vised learning for physical interaction through video prediction. In Advances In Neural Information Processing Systems. 64--72.  Chelsea Finn Ian Goodfellow and Sergey Levine. 2016. Unsuper- vised learning for physical interaction through video prediction. In Advances In Neural Information Processing Systems. 64--72."},{"volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 5515--5524","year":"2016","author":"Flynn John","key":"e_1_3_2_1_6_1"},{"key":"e_1_3_2_1_7_1","unstructured":"Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in neural information processing systems. 2672--2680.   Ian Goodfellow Jean Pouget-Abadie Mehdi Mirza Bing Xu David Warde-Farley Sherjil Ozair Aaron Courville and Yoshua Bengio. 2014. Generative adversarial nets. In Advances in neural information processing systems. 2672--2680."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10578-9_23"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-013-0683-3"},{"key":"e_1_3_2_1_10_1","unstructured":"Max Jaderberg Karen Simonyan Andrew Zisserman and others. 2015. Spatial transformer networks. In Advances in Neural Information Processing Systems. 2017--2025.   Max Jaderberg Karen Simonyan Andrew Zisserman and others. 2015. Spatial transformer networks. In Advances in Neural Information Processing Systems. 2017--2025."},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.223"},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33765-9_15"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10578-9_45"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.147"},{"volume-title":"Deep multi-scale video prediction beyond mean square error. arXiv preprint arXiv:1511.05440","year":"2015","author":"Mathieu Michael","key":"e_1_3_2_1_15_1"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298747"},{"key":"e_1_3_2_1_17_1","unstructured":"Volodymyr Mnih Nicolas Heess Alex Graves and others. 2014. Recurrent models of visual attention. In Advances in neural information processing systems. 2204--2212.   Volodymyr Mnih Nicolas Heess Alex Graves and others. 2014. Recurrent models of visual attention. In Advances in neural information processing systems. 2204--2212."},{"key":"e_1_3_2_1_18_1","unstructured":"Junhyuk Oh Xiaoxiao Guo Honglak Lee Richard L Lewis and Satinder Singh. 2015. Action-conditional video prediction using deep networks in atari games. In Advances in Neural Information Processing Systems. 2863--2871.   Junhyuk Oh Xiaoxiao Guo Honglak Lee Richard L Lewis and Satinder Singh. 2015. Action-conditional video prediction using deep networks in atari games. In Advances in Neural Information Processing Systems. 2863--2871."},{"volume-title":"Spatio-temporal video autoencoder with differentiable memory. arXiv preprint arXiv:1511.06309","year":"2015","author":"Patraucean Viorica","key":"e_1_3_2_1_19_1"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10578-9_12"},{"volume-title":"a baseline for generative models of natural videos. arXiv preprint arXiv:1412.6604","year":"2014","author":"Ranzato MarcAurelio","key":"e_1_3_2_1_21_1"},{"volume-title":"One-shot generalization in deep generative models. arXiv preprint arXiv:1603.05106","year":"2016","author":"Rezende Danilo Jimenez","key":"e_1_3_2_1_22_1"},{"volume-title":"Amir Roshan Zamir, and Mubarak Shah","year":"2012","author":"Soomro Khurram","key":"e_1_3_2_1_23_1"},{"key":"e_1_3_2_1_24_1","unstructured":"Nitish Srivastava Elman Mansimov and Ruslan Salakhutdinov. 2015. Unsupervised Learning of Video Representations using LSTMs.. In ICML. 843--852.   Nitish Srivastava Elman Mansimov and Ruslan Salakhutdinov. 2015. Unsupervised Learning of Video Representations using LSTMs.. In ICML. 843--852."},{"volume-title":"Arthur Szlam, Du Tran, and Soumith Chintala.","year":"2017","author":"van Amersfoort Joost","key":"e_1_3_2_1_25_1"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126359"},{"key":"e_1_3_2_1_27_1","unstructured":"Carl Vondrick Hamed Pirsiavash and Antonio Torralba. 2016. Generating videos with scene dynamics. In Advances In Neural Information Processing Systems. 613--621.  Carl Vondrick Hamed Pirsiavash and Antonio Torralba. 2016. Generating videos with scene dynamics. In Advances In Neural Information Processing Systems. 613--621."},{"volume-title":"One-Step Time- Dependent Future Video Frame Prediction with a Convolutional Encoder-Decoder Neural Network. arXiv preprint arX- iv:1702.04125","year":"2017","author":"Vukotic Vedran","key":"e_1_3_2_1_28_1"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46478-7_51"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.281"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995407"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF00992696"},{"key":"e_1_3_2_1_34_1","first-page":"77","article-title":"Show, Attend and Tell: Neural Image Caption Generation with Visual Attention","volume":"14","author":"Xu Kelvin","year":"2015","journal-title":"ICML"},{"key":"e_1_3_2_1_35_1","unstructured":"Tianfan Xue Jiajun Wu Katherine Bouman and Bill Freeman. 2016. Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks. In Advances in Neural Information Processing Systems. 91--99.  Tianfan Xue Jiajun Wu Katherine Bouman and Bill Freeman. 2016. Visual dynamics: Probabilistic future frame synthesis via cross convolutional networks. In Advances in Neural Information Processing Systems. 91--99."},{"key":"e_1_3_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.293"},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.5555\/1888028.1888082"}],"event":{"name":"MM '17: ACM Multimedia Conference","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Mountain View California USA","acronym":"MM '17"},"container-title":["Proceedings of the 25th ACM international conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3123266.3123349","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3123266.3123349","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T03:39:29Z","timestamp":1750217969000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3123266.3123349"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2017,10,23]]},"references-count":37,"alternative-id":["10.1145\/3123266.3123349","10.1145\/3123266"],"URL":"https:\/\/doi.org\/10.1145\/3123266.3123349","relation":{},"subject":[],"published":{"date-parts":[[2017,10,23]]},"assertion":[{"value":"2017-10-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}