{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T13:02:23Z","timestamp":1784552543692,"version":"3.55.0"},"publisher-location":"New York, NY, USA","reference-count":53,"publisher":"ACM","license":[{"start":{"date-parts":[[2021,10,17]],"date-time":"2021-10-17T00:00:00Z","timestamp":1634428800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000161","name":"National Institute of Standards and Technology","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000161","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100004351","name":"Cisco Systems","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100004351","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2021,10,17]]},"DOI":"10.1145\/3474085.3475596","type":"proceedings-article","created":{"date-parts":[[2021,10,18]],"date-time":"2021-10-18T20:56:12Z","timestamp":1634590572000},"page":"974-982","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":22,"title":["Cross-View Exocentric to Egocentric Video Synthesis"],"prefix":"10.1145","author":[{"given":"Gaowen","family":"Liu","sequence":"first","affiliation":[{"name":"Cisco Systems, San Jose, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hao","family":"Tang","sequence":"additional","affiliation":[{"name":"ETH, Zurich, Switzerland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hugo M.","family":"Latapie","sequence":"additional","affiliation":[{"name":"Cisco Systems, San Jose, CA, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jason J.","family":"Corso","sequence":"additional","affiliation":[{"name":"Stevens Institute for Artificial Intelligence, Hoboken, NJ, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yan","family":"Yan","sequence":"additional","affiliation":[{"name":"Illinois Institute of Technology, Chicago, IL, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2021,10,17]]},"reference":[{"key":"e_1_3_2_1_1_1","unstructured":"[n.d.]. https:\/\/gopro.com\/en\/us\/.  [n.d.]. https:\/\/gopro.com\/en\/us\/."},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995731"},{"key":"e_1_3_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Shervin Ardeshir and Ali Borji. 2016. Ego2top: Matching viewers in egocentric and top-view videos. In ECCV.  Shervin Ardeshir and Ali Borji. 2016. Ego2top: Matching viewers in egocentric and top-view videos. In ECCV.","DOI":"10.1007\/978-3-319-46454-1_16"},{"key":"e_1_3_2_1_4_1","volume-title":"Recycle-gan: Unsupervised video retargeting. In ECCV.","author":"Bansal Aayush","year":"2018","unstructured":"Aayush Bansal , Shugao Ma , Deva Ramanan , and Yaser Sheikh . 2018 . Recycle-gan: Unsupervised video retargeting. In ECCV. Aayush Bansal, Shugao Ma, Deva Ramanan, and Yaser Sheikh. 2018. Recycle-gan: Unsupervised video retargeting. In ECCV."},{"key":"e_1_3_2_1_5_1","unstructured":"Andrew Brock Jeff Donahue and Karen Simonyan. 2019. Large scale GAN training for high fidelity natural image synthesis. In ICLR.  Andrew Brock Jeff Donahue and Karen Simonyan. 2019. Large scale GAN training for high fidelity natural image synthesis. In ICLR."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"crossref","unstructured":"Caroline Chan Shiry Ginosar Tinghui Zhou and Alexei A Efros. 2019. Everybody dance now. In ICCV.  Caroline Chan Shiry Ginosar Tinghui Zhou and Alexei A Efros. 2019. Everybody dance now. In ICCV.","DOI":"10.1109\/ICCV.2019.00603"},{"key":"e_1_3_2_1_7_1","volume-title":"Chen Change Loy, and Dahua Lin","author":"Chen Kai","year":"2018","unstructured":"Kai Chen , Jiaqi Wang , Shuo Yang , Xingcheng Zhang , Yuanjun Xiong , Chen Change Loy, and Dahua Lin . 2018 . Optimizing video object detection via a scale-time lattice. In CVPR. Kai Chen, Jiaqi Wang, Shuo Yang, Xingcheng Zhang, Yuanjun Xiong, Chen Change Loy, and Dahua Lin. 2018. Optimizing video object detection via a scale-time lattice. In CVPR."},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"crossref","unstructured":"Yunjey Choi Youngjung Uh Jaejun Yoo and Jung-Woo Ha. 2020. Stargan v2: Diverse image synthesis for multiple domains. In CVPR.  Yunjey Choi Youngjung Uh Jaejun Yoo and Jung-Woo Ha. 2020. Stargan v2: Diverse image synthesis for multiple domains. In CVPR.","DOI":"10.1109\/CVPR42600.2020.00821"},{"key":"e_1_3_2_1_9_1","unstructured":"Emily Denton and Rob Fergus. 2018. Stochastic video generation with a learned prior. In ICML.  Emily Denton and Rob Fergus. 2018. Stochastic video generation with a learned prior. In ICML."},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"crossref","unstructured":"Bin Duan Wei Wang Hao Tang Hugo Latapie and Yan Yan. 2021. Cascade attention guided residue learning gan for cross-modal translation. In ICPR.  Bin Duan Wei Wang Hao Tang Hugo Latapie and Yan Yan. 2021. Cascade attention guided residue learning gan for cross-modal translation. In ICPR.","DOI":"10.1109\/ICPR48806.2021.9412890"},{"key":"e_1_3_2_1_11_1","unstructured":"Mohamed Elfeki Krishna Regmi Shervin Ardeshir and Ali Borji. 2019. From third person to first person: Dataset and baselines for synthesis and retrieval. In CVPR.  Mohamed Elfeki Krishna Regmi Shervin Ardeshir and Ali Borji. 2019. From third person to first person: Dataset and baselines for synthesis and retrieval. In CVPR."},{"key":"e_1_3_2_1_12_1","volume-title":"Yong Jae Lee, David J Crandall, and Michael S Ryoo.","author":"Fan Chenyou","year":"2017","unstructured":"Chenyou Fan , Jangwon Lee , Mingze Xu , Krishna Kumar Singh , Yong Jae Lee, David J Crandall, and Michael S Ryoo. 2017 . Identifying first-person camera wearers in third-person videos. In CVPR. Chenyou Fan, Jangwon Lee, Mingze Xu, Krishna Kumar Singh, Yong Jae Lee, David J Crandall, and Michael S Ryoo. 2017. Identifying first-person camera wearers in third-person videos. In CVPR."},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126269"},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2354936"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-33718-5_23"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/2969033.2969125"},{"key":"e_1_3_2_1_17_1","unstructured":"Minyoung Huh Shao-Hua Sun and Ning Zhang. 2019. Feedback adversarial learning: Spatial feedback for improving generative adversarial networks. In CVPR.  Minyoung Huh Shao-Hua Sun and Ning Zhang. 2019. Feedback adversarial learning: Spatial feedback for improving generative adversarial networks. In CVPR."},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"crossref","unstructured":"Phillip Isola Jun-Yan Zhu Tinghui Zhou and Alexei A Efros. 2017. Image-toimage translation with conditional adversarial networks. In CVPR.  Phillip Isola Jun-Yan Zhu Tinghui Zhou and Alexei A Efros. 2017. Image-toimage translation with conditional adversarial networks. In CVPR.","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2012.2200554"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Tero Karras Samuli Laine and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. In CVPR.  Tero Karras Samuli Laine and Timo Aila. 2019. A style-based generator architecture for generative adversarial networks. In CVPR.","DOI":"10.1109\/CVPR.2019.00453"},{"key":"e_1_3_2_1_21_1","first-page":"1228","article-title":"Refinenet: Multi-path refinement networks for dense prediction","volume":"42","author":"Lin Guosheng","year":"2019","unstructured":"Guosheng Lin , Fayao Liu , Anton Milan , Chunhua Shen , and Ian Reid . 2019 . Refinenet: Multi-path refinement networks for dense prediction . IEEE TPAMI 42 , 5 (2019), 1228 -- 1242 . Guosheng Lin, Fayao Liu, Anton Milan, Chunhua Shen, and Ian Reid. 2019. Refinenet: Multi-path refinement networks for dense prediction. IEEE TPAMI 42, 5 (2019), 1228--1242.","journal-title":"IEEE TPAMI"},{"key":"e_1_3_2_1_22_1","volume-title":"Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In CVPR.","author":"Lin Guosheng","year":"2017","unstructured":"Guosheng Lin , Anton Milan , Chunhua Shen , and Ian Reid . 2017 . Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In CVPR. Guosheng Lin, Anton Milan, Chunhua Shen, and Ian Reid. 2017. Refinenet: Multi-path refinement networks for high-resolution semantic segmentation. In CVPR."},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"crossref","unstructured":"Gaowen Liu Hao Tang Hugo Latapie and Yan Yan. 2020. Exocentric to egocentric image generation via parallel generative adversarial network. In ICASSP.  Gaowen Liu Hao Tang Hugo Latapie and Yan Yan. 2020. Exocentric to egocentric image generation via parallel generative adversarial network. In ICASSP.","DOI":"10.1109\/ICASSP40776.2020.9053957"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.5555\/3157096.3157149"},{"key":"e_1_3_2_1_25_1","unstructured":"Michael Mathieu Camille Couprie and Yann LeCun. 2016. Deep multi-scale video prediction beyond mean square error. In ICLR.  Michael Mathieu Camille Couprie and Yann LeCun. 2016. Deep multi-scale video prediction beyond mean square error. In ICLR."},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2012.6239188"},{"key":"e_1_3_2_1_27_1","unstructured":"Junting Pan Chengyu Wang Xu Jia Jing Shao Lu Sheng Junjie Yan and Xiaogang Wang. 2019. Video generation from single semantic label map. In CVPR.  Junting Pan Chengyu Wang Xu Jia Jing Shao Lu Sheng Junjie Yan and Xiaogang Wang. 2019. Video generation from single semantic label map. In CVPR."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"crossref","unstructured":"Eunbyung Park Jimei Yang Ersin Yumer Duygu Ceylan and Alexander C Berg. 2017. Transformation-grounded image generation network for novel 3d view synthesis. In CVPR.  Eunbyung Park Jimei Yang Ersin Yumer Duygu Ceylan and Alexander C Berg. 2017. Transformation-grounded image generation network for novel 3d view synthesis. In CVPR.","DOI":"10.1109\/CVPR.2017.82"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.5555\/2354409.2355089"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.325"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"crossref","unstructured":"Krishna Regmi and Ali Borji. 2018. Cross-view image synthesis using conditional gans. In CVPR.  Krishna Regmi and Ali Borji. 2018. Cross-view image synthesis using conditional gans. In CVPR.","DOI":"10.1109\/CVPR.2018.00369"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2019.07.008"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_3_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.352"},{"key":"e_1_3_2_1_35_1","doi-asserted-by":"crossref","unstructured":"Masaki Saito Eiichi Matsumoto and Shunta Saito. 2017. Temporal generative adversarial nets with singular value clipping. In ICCV.  Masaki Saito Eiichi Matsumoto and Shunta Saito. 2017. Temporal generative adversarial nets with singular value clipping. In ICCV.","DOI":"10.1109\/ICCV.2017.308"},{"key":"e_1_3_2_1_36_1","volume-title":"Singan: Learning a generative model from a single natural image. In ICCV.","author":"Shaham Tamar Rott","year":"2019","unstructured":"Tamar Rott Shaham , Tali Dekel , and Tomer Michaeli . 2019 . Singan: Learning a generative model from a single natural image. In ICCV. Tamar Rott Shaham, Tali Dekel, and Tomer Michaeli. 2019. Singan: Learning a generative model from a single natural image. In ICCV."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"crossref","unstructured":"Firas Shama Roey Mechrez Alon Shoshan and Lihi Zelnik-Manor. 2019. Adversarial feedback loop. In ICCV.  Firas Shama Roey Mechrez Alon Shoshan and Lihi Zelnik-Manor. 2019. Adversarial feedback loop. In ICCV.","DOI":"10.1109\/ICCV.2019.00330"},{"key":"e_1_3_2_1_38_1","first-page":"8916","article-title":"Unified generative adversarial networks for controllable image-to-image translation","volume":"29","author":"Tang Hao","year":"2020","unstructured":"Hao Tang , Hong Liu , and Nicu Sebe . 2020 . Unified generative adversarial networks for controllable image-to-image translation . IEEE TIP 29 (2020), 8916 -- 8929 . Hao Tang, Hong Liu, and Nicu Sebe. 2020. Unified generative adversarial networks for controllable image-to-image translation. IEEE TIP 29 (2020), 8916--8929.","journal-title":"IEEE TIP"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240704"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3343031.3350980"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"crossref","unstructured":"Hao Tang Dan Xu Nicu Sebe Yanzhi Wang Jason J Corso and Yan Yan. 2019. Multi-channel attention selection gan with cascaded semantic guidance for crossview image translation. In CVPR.  Hao Tang Dan Xu Nicu Sebe Yanzhi Wang Jason J Corso and Yan Yan. 2019. Multi-channel attention selection gan with cascaded semantic guidance for crossview image translation. In CVPR.","DOI":"10.1109\/CVPR.2019.00252"},{"key":"e_1_3_2_1_42_1","volume-title":"Philip HS Torr, and Nicu Sebe","author":"Tang Hao","year":"2020","unstructured":"Hao Tang , Dan Xu , Yan Yan , Philip HS Torr, and Nicu Sebe . 2020 . Local classspecific and global image-level generative adversarial networks for semanticguided scene generation. In CVPR. Hao Tang, Dan Xu, Yan Yan, Philip HS Torr, and Nicu Sebe. 2020. Local classspecific and global image-level generative adversarial networks for semanticguided scene generation. In CVPR."},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126462"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"crossref","unstructured":"Maxim Tatarchenko Alexey Dosovitskiy and Thomas Brox. 2016. Multi-view 3d models from single images with a convolutional network. In ECCV.  Maxim Tatarchenko Alexey Dosovitskiy and Thomas Brox. 2016. Multi-view 3d models from single images with a convolutional network. In ECCV.","DOI":"10.1007\/978-3-319-46478-7_20"},{"key":"e_1_3_2_1_45_1","volume-title":"Mocogan: Decomposing motion and content for video generation. In CVPR.","author":"Tulyakov Sergey","year":"2018","unstructured":"Sergey Tulyakov , Ming-Yu Liu , Xiaodong Yang , and Jan Kautz . 2018 . Mocogan: Decomposing motion and content for video generation. In CVPR. Sergey Tulyakov, Ming-Yu Liu, Xiaodong Yang, and Jan Kautz. 2018. Mocogan: Decomposing motion and content for video generation. In CVPR."},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/3326943.3327049"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"crossref","unstructured":"Xiaolong Wang Allan Jabri and Alexei A Efros. 2019. Learning correspondence from the cycle-consistency of time. In CVPR.  Xiaolong Wang Allan Jabri and Alexei A Efros. 2019. Learning correspondence from the cycle-consistency of time. In CVPR.","DOI":"10.1109\/CVPR.2019.00267"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_3_2_1_49_1","first-page":"2984","article-title":"Egocentric daily activity recognition via multitask clustering","volume":"24","author":"Yan Yan","year":"2015","unstructured":"Yan Yan , Elisa Ricci , Gaowen Liu , and Nicu Sebe . 2015 . Egocentric daily activity recognition via multitask clustering . IEEE TIP 24 , 10 (2015), 2984 -- 2995 . Yan Yan, Elisa Ricci, Gaowen Liu, and Nicu Sebe. 2015. Egocentric daily activity recognition via multitask clustering. IEEE TIP 24, 10 (2015), 2984--2995.","journal-title":"IEEE TIP"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"crossref","unstructured":"Ryo Yonetani Kris M Kitani and Yoichi Sato. 2016. Visual motif discovery via first-person vision. In ECCV.  Ryo Yonetani Kris M Kitani and Yoichi Sato. 2016. Visual motif discovery via first-person vision. In ECCV.","DOI":"10.1007\/978-3-319-46475-6_12"},{"key":"e_1_3_2_1_51_1","unstructured":"Han Zhang Ian Goodfellow Dimitris Metaxas and Augustus Odena. 2019. Selfattention generative adversarial networks. In ICML.  Han Zhang Ian Goodfellow Dimitris Metaxas and Augustus Odena. 2019. Selfattention generative adversarial networks. In ICML."},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"crossref","unstructured":"Tinghui Zhou Shubham Tulsiani Weilun Sun Jitendra Malik and Alexei A Efros. 2016. View synthesis by appearance flow. In ECCV.  Tinghui Zhou Shubham Tulsiani Weilun Sun Jitendra Malik and Alexei A Efros. 2016. View synthesis by appearance flow. In ECCV.","DOI":"10.1007\/978-3-319-46493-0_18"},{"key":"e_1_3_2_1_53_1","unstructured":"Jun-Yan Zhu Taesung Park Phillip Isola and Alexei A Efros. 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV.  Jun-Yan Zhu Taesung Park Phillip Isola and Alexei A Efros. 2017. Unpaired image-to-image translation using cycle-consistent adversarial networks. In ICCV."}],"event":{"name":"MM '21: ACM Multimedia Conference","location":"Virtual Event China","acronym":"MM '21","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 29th ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475596","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3475596","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3474085.3475596","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T20:48:23Z","timestamp":1750193303000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3474085.3475596"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,17]]},"references-count":53,"alternative-id":["10.1145\/3474085.3475596","10.1145\/3474085"],"URL":"https:\/\/doi.org\/10.1145\/3474085.3475596","relation":{},"subject":[],"published":{"date-parts":[[2021,10,17]]},"assertion":[{"value":"2021-10-17","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}