{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,18]],"date-time":"2026-01-18T13:05:52Z","timestamp":1768741552192,"version":"3.49.0"},"reference-count":61,"publisher":"Association for Computing Machinery (ACM)","issue":"CoNEXT1","license":[{"start":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T00:00:00Z","timestamp":1711584000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Netw."],"published-print":{"date-parts":[[2024,3,28]]},"abstract":"<jats:p>As mobile devices become increasingly popular for video streaming, it is crucial to optimize the streaming experience for these devices. Although deep learning-based video enhancement techniques are gaining attention, most of them cannot support real-time enhancement on mobile devices. Additionally, many of these techniques are focused solely on super-resolution and cannot handle partial or complete loss or corruption of video frames, which is common in the Internet and wireless networks.<\/jats:p>\n          <jats:p>To overcome these challenges, we present NERVE, a novel approach in this paper. NERVE consists of (i) a novel video frame recovery scheme, (ii) a new super-resolution algorithm, and (iii) an enhancement-aware video bit rate adaptation algorithm. We implement NERVE on an iPhone 12, and it can support 30 frames per second (FPS). We evaluate NERVE in various networks such as 3G, 4G, 5G, and WiFi networks. Our evaluation shows that NERVE enables real-time video recovery and enhancement, and results in 24% - 83% increase in video Quality of Experience (QoE) in our video streaming system.<\/jats:p>","DOI":"10.1145\/3649472","type":"journal-article","created":{"date-parts":[[2024,3,28]],"date-time":"2024-03-28T12:07:53Z","timestamp":1711627673000},"page":"1-19","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["NERVE: Real-Time Neural Video Recovery and Enhancement on Mobile Devices"],"prefix":"10.1145","volume":"2","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-5057-0314","authenticated-orcid":false,"given":"Zhaoyuan","family":"He","sequence":"first","affiliation":[{"name":"The University of Texas at Austin, Austin, Texas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5481-2851","authenticated-orcid":false,"given":"Yifan","family":"Yang","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1590-9749","authenticated-orcid":false,"given":"Lili","family":"Qiu","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin &amp; Microsoft Research Asia, Austin, Texas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5028-2697","authenticated-orcid":false,"given":"Kyoungjun","family":"Park","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin, Austin, Texas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3518-5212","authenticated-orcid":false,"given":"Yuqing","family":"Yang","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Shanghai, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2024,3,28]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"aioquic 2019. https:\/\/github.com\/aiortc\/aioquic."},{"key":"e_1_2_1_2_1","volume-title":"Futuregan: Anticipating the future \u00a8 frames of video sequences using spatio-temporal 3d convolutions in progressively growing autoencoder GANs. arXiv:1810.01325","author":"Aigner S.","year":"2018","unstructured":"S. Aigner and M. Korner. Futuregan: Anticipating the future \u00a8 frames of video sequences using spatio-temporal 3d convolutions in progressively growing autoencoder GANs. arXiv:1810.01325, 2018."},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1145\/3230543.3230558"},{"key":"e_1_2_1_4_1","article-title":"Deepstream: Video streaming enhancements using compressed deep neural networks","author":"Amirpour H.","year":"2022","unstructured":"H. Amirpour, M. Ghanbari, and C. Timmerer. Deepstream: Video streaming enhancements using compressed deep neural networks. IEEE Transactions on Circuits and Systems for Video Technology, 2022.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.161"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00491"},{"key":"e_1_2_1_7_1","volume-title":"Computer Vision and Pattern Recognition","author":"Chan K. C.","year":"2021","unstructured":"K. C. Chan, X. Wang, K. Yu, C. Dong, and C. C. Loy. Basicvsr: The search for essential components in video superresolution and beyond. Computer Vision and Pattern Recognition, 2021."},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386290.3396929"},{"key":"e_1_2_1_9_1","volume-title":"Proc. of Conference Record of 35th Asilomar Conference on Signals, Systems and Computers","author":"Chin S. K.","year":"2001","unstructured":"S. K. Chin and R. Braun. A survey of udp packet loss characteristics. In Proc. of Conference Record of 35th Asilomar Conference on Signals, Systems and Computers, 2001."},{"key":"e_1_2_1_10_1","unstructured":"Chrome is deploying http\/3 and ietf quic. https:\/\/blog.chromium.org\/2020\/10\/chrome-is-deploying-http3-and-ietfquic. html."},{"key":"e_1_2_1_11_1","volume-title":"Learning temporal coherence via self-supervision for gan-based video generation. ACM Transactions on Graphics (TOG), 39(4):75--1","author":"Chu M.","year":"2020","unstructured":"M. Chu, Y. Xie, J. Mayer, L. Leal-Taix\u00e9, and N. Thuerey. Learning temporal coherence via self-supervision for gan-based video generation. ACM Transactions on Graphics (TOG), 39(4):75--1, 2020."},{"key":"e_1_2_1_12_1","unstructured":"Cirp. https:\/\/cirpapple.substack.com\/p\/iphone-14-pro-and-pro-max-soar."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM41043.2020.9155477"},{"key":"e_1_2_1_14_1","volume-title":"Efficient video super-resolution through recurrent latent space propagation. arXiv: Image and Video Processing","author":"Fuoli D.","year":"2019","unstructured":"D. Fuoli, S. Gu, and R. Timofte. Efficient video super-resolution through recurrent latent space propagation. arXiv: Image and Video Processing, 2019."},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3301293.3302373"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2619239.2626296"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2342468.2342470"},{"key":"e_1_2_1_18_1","volume-title":"Proc. of IMC","author":"Jiang H.","year":"2012","unstructured":"H. Jiang, Y. Wang, K. Lee, and I. Rhee. Tackling bufferbloat in 3g\/4g mobile networks. In Proc. of IMC, 2012."},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/2413176.2413189"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19784-0_22"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01219-9_7"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/3098822.3098842"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3548370"},{"key":"e_1_2_1_24_1","first-page":"335","volume-title":"Proceedings, Part X 16","author":"Li W.","year":"2020","unstructured":"W. Li, X. Tao, T. Guo, L. Qi, J. Lu, and J. Jia. Mucan: Multi-correspondence aggregation network for video superresolution. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part X 16, pages 335--351. Springer, 2020."},{"key":"e_1_2_1_25_1","volume-title":"Sensors","author":"Lorincz J.","year":"2021","unstructured":"J. Lorincz, Z. Klarin. A comprehensive overview of tcp congestion control in 5g networks: Research challenges and future perspectives. Sensors, 2021."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3098822.3098843"},{"key":"e_1_2_1_27_1","unstructured":"Medium report about 'top 10 most popular types of videos on youtube'. https:\/\/mag.octoly.com\/here-are-the-top-10- most-popular-types-of-videos-on-youtube-4ea1e1a192ac."},{"key":"e_1_2_1_28_1","unstructured":"Mobile rrn. https:\/\/github.com\/MediaTek-NeuroPilot\/mai22-real-time-video-sr."},{"key":"e_1_2_1_29_1","unstructured":"Nearly 60% of americans now stream video daily on smartphones tablets and computers. https:\/\/www.nexttv.com\/ news\/nearly-60-of-americans-now-stream-video-daily-on-smart-phones-tablets-and-computers."},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2019.00251"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3452296.3472923"},{"key":"e_1_2_1_32_1","unstructured":"Capture network log. chrome:\/\/net-export\/."},{"key":"e_1_2_1_33_1","unstructured":"Proximal policy optimization (ppo). https:\/\/openai.com\/blog\/openai-baselines-ppo\/."},{"key":"e_1_2_1_34_1","unstructured":"Psnr. https:\/\/en.wikipedia.org\/wiki\/Peak_signal-to-noise_ratio."},{"key":"e_1_2_1_35_1","volume-title":"Zero-shot text-to-image generation. arXiv: Computer Vision and Pattern Recognition","author":"Ramesh A.","year":"2021","unstructured":"A. Ramesh, M. Pavlov, G. Goh, S. Gray, C. Voss, A. Radford, M. Chen, and I. Sutskever. Zero-shot text-to-image generation. arXiv: Computer Vision and Pattern Recognition, 2021."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.291"},{"key":"e_1_2_1_37_1","volume-title":"Polynomial codes over certain finite fields. Journal of the society for industrial and applied mathematics, 8(2):300--304","author":"Reed I. S.","year":"1960","unstructured":"I. S. Reed and G. Solomon. Polynomial codes over certain finite fields. Journal of the society for industrial and applied mathematics, 8(2):300--304, 1960."},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00693"},{"key":"e_1_2_1_39_1","first-page":"20378","article-title":"Movement pruning: Adaptive sparsity by fine-tuning","volume":"33","author":"Sanh V.","year":"2020","unstructured":"V. Sanh, T. Wolf, and A. Rush. Movement pruning: Adaptive sparsity by fine-tuning. Advances in Neural Information Processing Systems, 33:20378--20389, 2020.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2018.8451090"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11760-020-01671-x"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.207"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNET.2020.2996964"},{"key":"e_1_2_1_44_1","unstructured":"Ssim. https:\/\/en.wikipedia.org\/wiki\/Structural_similarity."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00507"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00186"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00165"},{"key":"e_1_2_1_48_1","volume-title":"Proc. of NeurIPS","author":"Vondrick C.","year":"2016","unstructured":"C. Vondrick, H. Pirsiavash, and A. Torralba. Generating videos with scene dynamics. In Proc. of NeurIPS, 2016."},{"key":"e_1_2_1_49_1","first-page":"514","volume-title":"Computer Vision--ACCV 2018:  14th Asian Conference on Computer Vision, Perth, Australia, December 2--6","author":"Wang L.","year":"2018","unstructured":"L. Wang, Y. Guo, Z. Lin, X. Deng, and W. An. Learning for video super-resolution through hr optical flow estimation. In Computer Vision--ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2--6, 2018, Revised Selected Papers, Part I 14, pages 514--529. Springer, 2019."},{"key":"e_1_2_1_50_1","volume-title":"Proc. of ICLR","author":"Wang Y.","year":"2019","unstructured":"Y. Wang, L. Jiang, M.-H. Yang, L.-J. Li, M. Longand, and L. Fei-Fei. In Proc. of ICLR, 2019."},{"key":"e_1_2_1_51_1","unstructured":"Wowza's dash bitrate recommendation. https:\/\/www.wowza.com\/docs\/how-to-encode-source-video-for-wowzastreaming- cloud."},{"key":"e_1_2_1_52_1","volume-title":"Online video super-resolution with convolutional kernel bypass graft","author":"Xiao J.","year":"2022","unstructured":"J. Xiao, X. Jiang, N. Zheng, H. Yang, Y. Yang, Y. Yang, D. Li, and K.-M. Lam. Online video super-resolution with convolutional kernel bypass graft. 2022."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3387514.3405882"},{"key":"e_1_2_1_54_1","volume-title":"Videogpt: Video generation using vq-vae and transformers. arXiv: Computer Vision and Pattern Recognition","author":"Yan W.","year":"2021","unstructured":"W. Yan, Y. Zhang, P. Abbeel, and A. Srinivas. Videogpt: Video generation using vq-vae and transformers. arXiv: Computer Vision and Pattern Recognition, 2021."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3372224.3419185"},{"key":"e_1_2_1_56_1","first-page":"645","volume-title":"13th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 18)","author":"Yeo H.","year":"2018","unstructured":"H. Yeo, Y. Jung, J. Kim, J. Shin, and D. Han. Neural adaptive content-aware internet video delivery. In 13th {USENIX} Symposium on Operating Systems Design and Implementation ({OSDI} 18), pages 645--661, 2018."},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00320"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2785956.2787486"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2019.00048"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/505202.505228"},{"key":"e_1_2_1_61_1","article-title":"Neural-enhanced adaptive streaming of vbr-encoded videos with selective prefetching","author":"Zhou G.","year":"2022","unstructured":"G. Zhou, Z. Luo, M. Hu, and D. Wu. Presr: Neural-enhanced adaptive streaming of vbr-encoded videos with selective prefetching. IEEE Transactions on Broadcasting, 2022.","journal-title":"IEEE Transactions on Broadcasting"}],"container-title":["Proceedings of the ACM on Networking"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3649472","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3649472","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T20:31:19Z","timestamp":1755981079000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3649472"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,3,28]]},"references-count":61,"journal-issue":{"issue":"CoNEXT1","published-print":{"date-parts":[[2024,3,28]]}},"alternative-id":["10.1145\/3649472"],"URL":"https:\/\/doi.org\/10.1145\/3649472","relation":{},"ISSN":["2834-5509"],"issn-type":[{"value":"2834-5509","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,28]]},"assertion":[{"value":"2024-03-28","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}