{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,7]],"date-time":"2026-03-07T18:32:39Z","timestamp":1772908359154,"version":"3.50.1"},"publisher-location":"New York, NY, USA","reference-count":33,"publisher":"ACM","license":[{"start":{"date-parts":[[2023,9,10]],"date-time":"2023-09-10T00:00:00Z","timestamp":1694304000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key Research and Development Program of China","award":["No. 2021YFF0900504"],"award-info":[{"award-number":["No. 2021YFF0900504"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61971382"],"award-info":[{"award-number":["61971382"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62262018"],"award-info":[{"award-number":["62262018"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2023,9,10]]},"DOI":"10.1145\/3609395.3610597","type":"proceedings-article","created":{"date-parts":[[2023,9,26]],"date-time":"2023-09-26T21:19:55Z","timestamp":1695763195000},"page":"28-33","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["LiveAE: Attention-based and Edge-assisted Viewport Prediction for Live 360\u00b0 Video Streaming"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-8647-4731","authenticated-orcid":false,"given":"Zipeng","family":"Pan","sequence":"first","affiliation":[{"name":"Communication University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3783-7974","authenticated-orcid":false,"given":"Yuan","family":"Zhang","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1170-636X","authenticated-orcid":false,"given":"Tao","family":"Lin","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4153-313X","authenticated-orcid":false,"given":"Jinyao","family":"Yan","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Media Convergence and Communication, Communication University of China, Beijing, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,9,26]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2018.8486606"},{"key":"e_1_3_2_1_2_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin , Ming-Wei Chang , Kenton Lee , and Kristina Toutanova . 2018 . Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018). Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_3_2_1_3_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly etal 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020).  Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_3_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2021.3067686"},{"key":"e_1_3_2_1_5_1","volume-title":"Proceedings of the 12th ACM Multimedia Systems Conference. 132--145","author":"Feng Xianglong","year":"2021","unstructured":"Xianglong Feng , Weitian Li , and Sheng Wei . 2021 . LiveROI: region of interest analysis for viewport prediction in live mobile virtual reality streaming . In Proceedings of the 12th ACM Multimedia Systems Conference. 132--145 . Xianglong Feng, Weitian Li, and Sheng Wei. 2021. LiveROI: region of interest analysis for viewport prediction in live mobile virtual reality streaming. In Proceedings of the 12th ACM Multimedia Systems Conference. 132--145."},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/VR46266.2020.00104"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3328914"},{"key":"e_1_3_2_1_8_1","volume-title":"large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677","author":"Goyal Priya","year":"2017","unstructured":"Priya Goyal , Piotr Doll\u00e1r , Ross Girshick , Pieter Noordhuis , Lukasz Wesolowski , Aapo Kyrola , Andrew Tulloch , Yangqing Jia , and Kaiming He. 2017. Accurate , large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677 ( 2017 ). Priya Goyal, Piotr Doll\u00e1r, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017. Accurate, large minibatch sgd: Training imagenet in 1 hour. arXiv preprint arXiv:1706.02677 (2017)."},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00065"},{"key":"e_1_3_2_1_10_1","volume-title":"Retrieved","author":"Apple Inc.","year":"2023","unstructured":"Apple Inc. 2023 . Core ML . Retrieved June 19, 2023 from https:\/\/developer.apple.com\/documentation\/coreml Apple Inc. 2023. Core ML. Retrieved June 19, 2023 from https:\/\/developer.apple.com\/documentation\/coreml"},{"key":"e_1_3_2_1_11_1","volume-title":"Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101","author":"Loshchilov Ilya","year":"2017","unstructured":"Ilya Loshchilov and Frank Hutter . 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 ( 2017 ). Ilya Loshchilov and Frank Hutter. 2017. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017)."},{"key":"e_1_3_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.07.028"},{"key":"e_1_3_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413751"},{"key":"e_1_3_2_1_14_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10480--10490","author":"Miles Roy","year":"2023","unstructured":"Roy Miles , Mehmet Kerim Yucel , Bruno Manganelli , and Albert Sa\u00e0-Garriga . 2023 . MobileVOS: Real-Time Video Object Segmentation Contrastive Learning meets Knowledge Distillation . In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10480--10490 . Roy Miles, Mehmet Kerim Yucel, Bruno Manganelli, and Albert Sa\u00e0-Garriga. 2023. MobileVOS: Real-Time Video Object Segmentation Contrastive Learning meets Knowledge Distillation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 10480--10490."},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.image.2018.05.005"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/3386290.3396934"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/AIVR.2018.00033"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/3241539.3241565"},{"key":"e_1_3_2_1_19_1","first-page":"12116","article-title":"Do vision transformers see like convolutional neural networks","volume":"34","author":"Raghu Maithra","year":"2021","unstructured":"Maithra Raghu , Thomas Unterthiner , Simon Kornblith , Chiyuan Zhang , and Alexey Dosovitskiy . 2021 . Do vision transformers see like convolutional neural networks ? Advances in Neural Information Processing Systems 34 (2021), 12116 -- 12128 . Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy. 2021. Do vision transformers see like convolutional neural networks? Advances in Neural Information Processing Systems 34 (2021), 12116--12128.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_20_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman . 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 ( 2014 ). Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1145\/3339825.3391856"},{"key":"e_1_3_2_1_22_1","volume-title":"Live 360 Degree Video Delivery based on User Collaboration in a Streaming Flock","author":"Sun Liyang","year":"2022","unstructured":"Liyang Sun , Yixiang Mao , Tongyu Zong , Yong Liu , and Yao Wang . 2022. Live 360 Degree Video Delivery based on User Collaboration in a Streaming Flock . IEEE Transactions on Multimedia ( 2022 ). Liyang Sun, Yixiang Mao, Tongyu Zong, Yong Liu, and Yao Wang. 2022. Live 360 Degree Video Delivery based on User Collaboration in a Streaming Flock. IEEE Transactions on Multimedia (2022)."},{"key":"e_1_3_2_1_23_1","volume-title":"Attention is all you need. Advances in neural information processing systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani , Noam Shazeer , Niki Parmar , Jakob Uszkoreit , Llion Jones , Aidan N Gomez , \u0141ukasz Kaiser , and Illia Polosukhin . 2017. Attention is all you need. Advances in neural information processing systems 30 ( 2017 ). Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)."},{"key":"e_1_3_2_1_24_1","volume-title":"Proceedings of the 28th Annual International Conference on Mobile Computing And Networking. 542--555","author":"Wang Shibo","year":"2022","unstructured":"Shibo Wang , Shusen Yang , Hailiang Li , Xiaodan Zhang , Chen Zhou , Chenren Xu , Feng Qian , Nanbin Wang , and Zongben Xu . 2022 . SalientVR: saliency-driven mobile 360-degree video streaming with gaze information . In Proceedings of the 28th Annual International Conference on Mobile Computing And Networking. 542--555 . Shibo Wang, Shusen Yang, Hailiang Li, Xiaodan Zhang, Chen Zhou, Chenren Xu, Feng Qian, Nanbin Wang, and Zongben Xu. 2022. SalientVR: saliency-driven mobile 360-degree video streaming with gaze information. In Proceedings of the 28th Annual International Conference on Mobile Computing And Networking. 542--555."},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413999"},{"key":"e_1_3_2_1_26_1","volume-title":"arXiv preprint arXiv:1907.10701","author":"Wang Yu Emma","year":"2019","unstructured":"Yu Emma Wang , Gu-Yeon Wei , and David Brooks . 2019. Benchmarking TPU, GPU , and CPU platforms for deep learning. arXiv preprint arXiv:1907.10701 ( 2019 ). Yu Emma Wang, Gu-Yeon Wei, and David Brooks. 2019. Benchmarking TPU, GPU, and CPU platforms for deep learning. arXiv preprint arXiv:1907.10701 (2019)."},{"key":"e_1_3_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123266.3123291"},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3240508.3240556"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3359989.3365413"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00559"},{"key":"e_1_3_2_1_31_1","volume-title":"2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1--6.","author":"Zhang Lei","year":"2022","unstructured":"Lei Zhang , Weizhen Xu , Donghuan Lu , Laizhong Cui , and Jiangchuan Liu . 2022 . MFVP: Mobile-Friendly Viewport Prediction for Live 360-Degree Video Streaming . In 2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1--6. Lei Zhang, Weizhen Xu, Donghuan Lu, Laizhong Cui, and Jiangchuan Liu. 2022. MFVP: Mobile-Friendly Viewport Prediction for Live 360-Degree Video Streaming. In 2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1--6."},{"key":"e_1_3_2_1_32_1","volume-title":"Efficient, Economical, and Quality-of-Experience-Driven VR Video System Based on MPEG OMAF","author":"Zhang Qi","year":"2022","unstructured":"Qi Zhang , Jianchao Wei , Shanshe Wang , Siwei Ma , and Wen Gao . 2022. RealVR : Efficient, Economical, and Quality-of-Experience-Driven VR Video System Based on MPEG OMAF . IEEE Transactions on Multimedia ( 2022 ). Qi Zhang, Jianchao Wei, Shanshe Wang, Siwei Ma, and Wen Gao. 2022. RealVR: Efficient, Economical, and Quality-of-Experience-Driven VR Video System Based on MPEG OMAF. IEEE Transactions on Multimedia (2022)."},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2019.8737361"}],"event":{"name":"EMS '23: 2023 Workshop on Emerging Multimedia Systems","location":"New York NY USA","acronym":"EMS '23","sponsor":["SIGCOMM ACM Special Interest Group on Data Communication"]},"container-title":["Proceedings of the 2023 Workshop on Emerging Multimedia Systems"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3609395.3610597","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3609395.3610597","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:46:23Z","timestamp":1750178783000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3609395.3610597"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,10]]},"references-count":33,"alternative-id":["10.1145\/3609395.3610597","10.1145\/3609395"],"URL":"https:\/\/doi.org\/10.1145\/3609395.3610597","relation":{},"subject":[],"published":{"date-parts":[[2023,9,10]]},"assertion":[{"value":"2023-09-26","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}