{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,17]],"date-time":"2026-07-17T06:09:56Z","timestamp":1784268596124,"version":"3.55.0"},"reference-count":65,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2019,11,8]],"date-time":"2019-11-08T00:00:00Z","timestamp":1573171200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by-sa\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2019,12,31]]},"abstract":"<jats:p>In order to provide an immersive visual experience, modern displays require head mounting, high image resolution, low latency, as well as high refresh rate. This poses a challenging computational problem. On the other hand, the human visual system can consume only a tiny fraction of this video stream due to the drastic acuity loss in the peripheral vision. Foveated rendering and compression can save computations by reducing the image quality in the peripheral vision. However, this can cause noticeable artifacts in the periphery, or, if done conservatively, would provide only modest savings. In this work, we explore a novel foveated reconstruction method that employs the recent advances in generative adversarial neural networks. We reconstruct a plausible peripheral video from a small fraction of pixels provided every frame. The reconstruction is done by finding the closest matching video to this sparse input stream of pixels on the learned manifold of natural videos. Our method is more efficient than the state-of-the-art foveated rendering, while providing the visual experience with no noticeable quality degradation. We conducted a user study to validate our reconstruction method and compare it against existing foveated rendering and video compression techniques. Our method is fast enough to drive gaze-contingent head-mounted displays in real time on modern hardware. We plan to publish the trained network to establish a new quality bar for foveated rendering and compression as well as encourage follow-up research.<\/jats:p>","DOI":"10.1145\/3355089.3356557","type":"journal-article","created":{"date-parts":[[2019,11,8]],"date-time":"2019-11-08T20:27:58Z","timestamp":1573244878000},"page":"1-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":119,"title":["DeepFovea"],"prefix":"10.1145","volume":"38","author":[{"given":"Anton S.","family":"Kaplanyan","sequence":"first","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Anton","family":"Sochenov","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas","family":"Leimk\u00fchler","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mikhail","family":"Okunev","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Todd","family":"Goodall","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Gizem","family":"Rufo","sequence":"additional","affiliation":[{"name":"Facebook Reality Labs"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,11,8]]},"reference":[{"key":"e_1_2_2_1_1","volume-title":"YouTube-8M: A Large-Scale Video Classification Benchmark. CoRR abs\/1609.08675","author":"Abu-El-Haija Sami","year":"2016","unstructured":"Sami Abu-El-Haija, Nisarg Kothari, Joonseok Lee, Paul Natsev, George Toderici, Balakrishnan Varadarajan, and Sudheendra Vijayanarasimhan. 2016. YouTube-8M: A Large-Scale Video Classification Benchmark. CoRR abs\/1609.08675 (2016)."},{"key":"e_1_2_2_2_1","volume-title":"Proceedings of the 34th International Conference on Machine Learning (Proceedings of Machine Learning Research), Doina Precup and Yee Whye Teh (Eds.)","volume":"70","author":"Arjovsky Martin","year":"2017","unstructured":"Martin Arjovsky, Soumith Chintala, and L\u00e9on Bottou. 2017. Wasserstein Generative Adversarial Networks. In Proceedings of the 34th International Conference on Machine Learning (Proceedings of Machine Learning Research), Doina Precup and Yee Whye Teh (Eds.), Vol. 70. PMLR, 214--223."},{"key":"e_1_2_2_3_1","volume-title":"Hinton","author":"Ba Lei Jimmy","year":"2016","unstructured":"Lei Jimmy Ba, Ryan Kiros, and Geoffrey E. Hinton. 2016. Layer Normalization. CoRR abs\/1607.06450 (2016). arXiv:1607.06450 http:\/\/arxiv.org\/abs\/1607.06450"},{"key":"e_1_2_2_4_1","volume-title":"Bovik","author":"Bampis Christos","year":"2018","unstructured":"Christos Bampis, Zhi Li, Ioannis Katsavounidis, Te-Yuan Huang, Chaitanya Ekanadham, and Alan C. Bovik. 2018. Towards Perceptually Optimized End-to-end Adaptive Video Streaming. arXiv preprint arXiv:1808.03898 (2018)."},{"key":"e_1_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01228-1_8"},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1167\/14.12.22"},{"key":"e_1_2_2_7_1","article-title":"Interactive Reconstruction of Monte Carlo Image Sequences Using a Recurrent Denoising Autoencoder","volume":"36","author":"Alla Chaitanya Chakravarty R.","year":"2017","unstructured":"Chakravarty R. Alla Chaitanya, Anton S. Kaplanyan, Christoph Schied, Marco Salvi, Aaron Lefohn, Derek Nowrouzezahrai, and Timo Aila. 2017. Interactive Reconstruction of Monte Carlo Image Sequences Using a Recurrent Denoising Autoencoder. ACM Trans. Graph. (Proc. SIGGRAPH) 36, 4, Article 98 (2017), 98:1--98:12 pages.","journal-title":"ACM Trans. Graph. (Proc. SIGGRAPH)"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.image.2018.06.009"},{"key":"e_1_2_2_9_1","volume-title":"Conference on Learning Representations, ICLR abs\/1511","author":"Clevert Djork-Arn\u00e9","year":"2016","unstructured":"Djork-Arn\u00e9 Clevert, Thomas Unterthiner, and Sepp Hochreiter. 2016. Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs). Conference on Learning Representations, ICLR abs\/1511.07289 (2016)."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/7529.8927"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1002\/cne.902920402"},{"key":"e_1_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.89.20.9666"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1146\/annurev.psych.58.110405.085632"},{"key":"e_1_2_2_14_1","volume-title":"Perry","author":"Geisler Wilson S.","year":"1998","unstructured":"Wilson S. Geisler and Jeffrey S. Perry. 1998. Real-time foveated multiresolution system for low-bandwidth video communication., 3299 - 3299 - 12 pages."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.81"},{"key":"e_1_2_2_16_1","volume-title":"Generative Adversarial Networks. arXiv e-prints","author":"Goodfellow Ian J.","year":"2014","unstructured":"Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014. Generative Adversarial Networks. arXiv e-prints (2014), arXiv:1406.2661. https:\/\/arxiv.org\/abs\/1406.2661"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2366145.2366183"},{"key":"e_1_2_2_18_1","unstructured":"Lars Haglund. 2006. The SVT High Definition Multi Format Test Set. (2006). https:\/\/media.xiph.org\/video\/derf\/vqeg.its.bldrdoc.gov\/HDTV\/SVT_MultiFormat\/SVT_MultiFormat_v10.pdf"},{"key":"e_1_2_2_19_1","volume-title":"Deep Residual Learning for Image Recognition. IEEE Conference on Computer Vision and Pattern Recognition","author":"He Kaiming","year":"2016","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. IEEE Conference on Computer Vision and Pattern Recognition (2016), 770--778."},{"key":"e_1_2_2_20_1","article-title":"Extending the Graphics Pipeline with Adaptive","volume":"33","author":"He Yong","year":"2014","unstructured":"Yong He, Yan Gu, and Kayvon Fatahalian. 2014. Extending the Graphics Pipeline with Adaptive, Multi-rate Shading. ACM Trans. Graph. (Proc. SIGGRAPH) 33, 4, Article 142 (2014), 142:1--142:12 pages.","journal-title":"Multi-rate Shading. ACM Trans. Graph. (Proc. SIGGRAPH)"},{"key":"e_1_2_2_21_1","volume-title":"Long short-term memory. Neural computation 9, 8","author":"Hochreiter Sepp","year":"1997","unstructured":"Sepp Hochreiter and J\u00fcrgen Schmidhuber. 1997. Long short-term memory. Neural computation 9, 8 (1997), 1735--1780."},{"key":"e_1_2_2_22_1","volume-title":"Proc. Conf. Computer Vision and Pattern Recognition. http:\/\/lmb.informatik.uni-freiburg.de\/\/Publications\/2017\/IMKDB17","author":"Ilg E.","unstructured":"E. Ilg, N. Mayer, T. Saikia, M. Keuper, A. Dosovitskiy, and T. Brox. 2017. FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks. In Proc. Conf. Computer Vision and Pattern Recognition. http:\/\/lmb.informatik.uni-freiburg.de\/\/Publications\/2017\/IMKDB17"},{"key":"e_1_2_2_23_1","volume-title":"Proc. Conf. Computer Vision and Pattern Recognition","author":"Isola Phillip","year":"2017","unstructured":"Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. 2017. Image-to-Image Translation with Conditional Adversarial Networks. Proc. Conf. Computer Vision and Pattern Recognition (2017), 5967--5976."},{"key":"e_1_2_2_24_1","volume-title":"International Conference on Learning Representations.","author":"Karras Tero","year":"2018","unstructured":"Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2018. Progressive Growing of GANs for Improved Quality, Stability, and Variation. In International Conference on Learning Representations."},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1364\/JOSAA.1.000107"},{"key":"e_1_2_2_26_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. CoRR (2014)."},{"key":"e_1_2_2_27_1","volume-title":"Hinton","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. 2012. ImageNet Classification with Deep Convolutional Neural Networks. Neural Information Processing Systems 25 (01 2012)."},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2015.7351227"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_2_2_30_1","volume-title":"Proc. Conf. Computer Vision and Pattern Recognition. 105--114","author":"Ledig C.","unstructured":"C. Ledig, L. Theis, F. Husz\u00e1r, J. Caballero, A. Cunningham, A. Acosta, A. Aitken, A. Tejani, J. Totz, Z. Wang, and W. Shi. 2017. Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network. In Proc. Conf. Computer Vision and Pattern Recognition. 105--114."},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/83.931092"},{"key":"e_1_2_2_32_1","volume-title":"Precomputed Real-Time Texture Synthesis with Markovian Generative Adversarial Networks. CoRR abs\/1604.04382","author":"Li Chuan","year":"2016","unstructured":"Chuan Li and Michael Wand. 2016. Precomputed Real-Time Texture Synthesis with Markovian Generative Adversarial Networks. CoRR abs\/1604.04382 (2016)."},{"key":"e_1_2_2_33_1","volume-title":"Image Inpainting for Irregular Holes Using Partial Convolutions. arXiv preprint arXiv:1804.07723","author":"Liu Guilin","year":"2018","unstructured":"Guilin Liu, Fitsum A. Reda, Kevin Shih, Ting-Chun Wang, Andrew Tao, and Bryan Catanzaro. 2018. Image Inpainting for Irregular Holes Using Partial Convolutions. arXiv preprint arXiv:1804.07723 (2018)."},{"key":"e_1_2_2_34_1","volume-title":"Visual quality assessment: recent developments, coding applications and future trends. Transactions on Signal and Information Processing 2","author":"Liu Tsung-Jung","year":"2013","unstructured":"Tsung-Jung Liu, Yu-Chieh Lin, Weisi Lin, and C-C Jay Kuo. 2013. Visual quality assessment: recent developments, coding applications and future trends. Transactions on Signal and Information Processing 2 (2013)."},{"key":"e_1_2_2_35_1","volume-title":"Allan G. Rempel, and Wolfgang Heidrich.","author":"Mantiuk Rafat","year":"2011","unstructured":"Rafat Mantiuk, Kil Joong Kim, Allan G. Rempel, and Wolfgang Heidrich. 2011. HDR-VDP-2: a calibrated visual metric for visibility and quality predictions in all luminance conditions. In ACM Transactions on graphics (TOG), Vol. 30. ACM, 40."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/127719.122736"},{"key":"e_1_2_2_37_1","volume-title":"Spectral Normalization for Generative Adversarial Networks. CoRR abs\/1802.05957","author":"Miyato Takeru","year":"2018","unstructured":"Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. 2018. Spectral Normalization for Generative Adversarial Networks. CoRR abs\/1802.05957 (2018)."},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.278"},{"key":"e_1_2_2_39_1","article-title":"Towards Foveated Rendering for Gazetracked Virtual Reality","volume":"35","author":"Patney Anjul","year":"2016","unstructured":"Anjul Patney, Marco Salvi, Joohwan Kim, Anton Kaplanyan, Chris Wyman, Nir Benty, David Luebke, and Aaron Lefohn. 2016. Towards Foveated Rendering for Gazetracked Virtual Reality. ACM Trans. Graph. (Proc. SIGGRAPH Asia) 35, 6, Article 179 (2016), 179:1--179:12 pages.","journal-title":"ACM Trans. Graph. (Proc. SIGGRAPH Asia)"},{"key":"e_1_2_2_40_1","volume-title":"Photorealistic Video Super Resolution. CoRR abs\/1807.07930","author":"P\u00e9rez-Pellitero Eduardo","year":"2018","unstructured":"Eduardo P\u00e9rez-Pellitero, Mehdi S. M. Sajjadi, Michael Hirsch, and Bernhard Sch\u00f6lkopf. 2018. Photorealistic Video Super Resolution. CoRR abs\/1807.07930 (2018)."},{"key":"e_1_2_2_41_1","unstructured":"E. P\u00e9rez-Pellitero M. S. M. Sajjadi M. Hirsch and B. Sch\u00f6lkopf. 2018. Photorealistic Video Super Resolution."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBC.2004.834028"},{"key":"e_1_2_2_43_1","volume-title":"International Conference on Systems, Signals and Image Processing","author":"Rimac-Drlje S.","year":"2011","unstructured":"S. Rimac-Drlje, G. Martinovi\u0107, and B. Zovko-Cihlar. 2011. Foveation-based content Adaptive Structural Similarity index. International Conference on Systems, Signals and Image Processing (2011), 1--4."},{"key":"e_1_2_2_44_1","doi-asserted-by":"crossref","unstructured":"Oren Rippel Sanjay Nair Carissa Lew Steve Branson Alexander G. Anderson and Lubomir Bourdev. 2018. Learned Video Compression. (2018).","DOI":"10.1109\/ICCV.2019.00355"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1364\/JOSA.56.001141"},{"key":"e_1_2_2_46_1","first-page":"234","article-title":"U-Net: Convolutional Networks for Biomedical Image Segmentation","volume":"9351","author":"Ronneberger O.","year":"2015","unstructured":"O. Ronneberger, P. Fischer, and T. Brox. 2015. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Medical Image Computing and Computer-Assisted Intervention (MICCAI) (LNCS), Vol. 9351. 234--241.","journal-title":"Medical Image Computing and Computer-Assisted Intervention (MICCAI) (LNCS)"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1068\/p130665"},{"key":"e_1_2_2_48_1","volume-title":"The statistics of natural images. Network: computation in neural systems 5, 4","author":"Ruderman Daniel L","year":"1994","unstructured":"Daniel L Ruderman. 1994. The statistics of natural images. Network: computation in neural systems 5, 4 (1994), 517--548."},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neunet.2014.09.003"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2010.2042111"},{"key":"e_1_2_2_51_1","volume-title":"Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR abs\/1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. CoRR abs\/1409.1556 (2014)."},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2012.2214933"},{"key":"e_1_2_2_53_1","volume-title":"Adaptive Image-Space Sampling for Gaze-Contingent Real-time Rendering. Computer Graphics Forum (Proc. of Eurographics Symposium on Rendering) 35","author":"Stengel Michael","year":"2016","unstructured":"Michael Stengel, Steve Grogorick, Martin Eisemann, and Marcus Magnor. 2016. Adaptive Image-Space Sampling for Gaze-Contingent Real-time Rendering. Computer Graphics Forum (Proc. of Eurographics Symposium on Rendering) 35, 4 (2016), 129--139."},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3130800.3130807"},{"key":"e_1_2_2_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/2931002.2931011"},{"key":"e_1_2_2_56_1","volume-title":"Human Vision, Visual Processing, and Digital Display IV","author":"Ulichney Robert A","unstructured":"Robert A Ulichney. 1993. Void-and-cluster method for dither array generation. In Human Vision, Visual Processing, and Digital Display IV, Vol. 1913. International Society for Optics and Photonics, 332--343."},{"key":"e_1_2_2_57_1","unstructured":"Alex Vlachos. 2015. Advanced VR Rendering. http:\/\/media.steampowered.com\/apps\/valve\/2015\/Alex_Vlachos_Advanced_VR_Rendering_GDC2015.pdf Game Developers Conference Talk."},{"key":"e_1_2_2_58_1","unstructured":"Ting-Chun Wang Ming-Yu Liu Jun-Yan Zhu Guilin Liu Andrew Tao Jan Kautz and Bryan Catanzaro. 2018. Video-to-Video Synthesis. In Neural Information Processing Systems."},{"key":"e_1_2_2_59_1","volume-title":"Proc. SPIE 4472","author":"Wang Zhou","year":"2001","unstructured":"Zhou Wang, Alan Conrad Bovik, Ligang Lu, and Jack L Kouloheris. 2001. Foveated wavelet image quality index. Proc. SPIE 4472 (2001), 42--53."},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.819861"},{"key":"e_1_2_2_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2003.809015"},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.5555\/3151666.3151696"},{"key":"e_1_2_2_63_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.13150"},{"key":"e_1_2_2_64_1","unstructured":"Y. Ye E. Alshina and J. Boyce. 2017. Algorithm descriptions of projection format conversion and video quality metrics in 360Lib. Joint Video Exploration Team of ITU-T SG 16 (2017)."},{"key":"e_1_2_2_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00068"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3355089.3356557","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3355089.3356557","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:44:41Z","timestamp":1750203881000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3355089.3356557"}},"subtitle":["neural reconstruction for foveated rendering and video compression using learned statistics of natural videos"],"short-title":[],"issued":{"date-parts":[[2019,11,8]]},"references-count":65,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2019,12,31]]}},"alternative-id":["10.1145\/3355089.3356557"],"URL":"https:\/\/doi.org\/10.1145\/3355089.3356557","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,11,8]]},"assertion":[{"value":"2019-11-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}