{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T20:39:11Z","timestamp":1768336751501,"version":"3.49.0"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities of China","doi-asserted-by":"crossref","award":["TN2216010"],"award-info":[{"award-number":["TN2216010"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"name":"\u2018Jie Bang Gua Shuai\u2019 Science and Technology Major Project of Liaoning Province in 2022","award":["2022JH1\/10400025"],"award-info":[{"award-number":["2022JH1\/10400025"]}]},{"name":"National Key Research and Development Program of China","award":["2018YFB1702000"],"award-info":[{"award-number":["2018YFB1702000"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2026,1,31]]},"abstract":"<jats:p>\n                    The application of animation models in facial video compression has yielded significant coding gains, particularly at ultra-low bitrates. Despite notable advancements, research on portrait video scenes, especially in half-body and full-body contexts, remains underexplored. The mapping motion features and dynamic regions is often imprecise, leading to inaccurate motion parameters. Additionally, the reconstruction of occluded background is frequently suboptimal, as occluded regions lack ground-truth data. To this end, we propose a learning-based portrait video compression (LPVC) framework for half-body and full-body portrait videos with static backgrounds. Specifically, we first design a keypoint detector driven by human parsing to integrate semantic properties into the animation model. This facilitates richer features and enhances the motion features for dynamic areas. We further devise a background incremental coding scheme to reconstruct high-quality occluded backgrounds. The scheme incorporates background occlusion calculation and compensation modules to process the newly revealed background pixels between successive frames, eliminating redundant background transmission. Finally, a spatio-temporal portrait background fusion generative adversarial network (SPBF-GAN) is devised to learn spatio-temporal differential representations of videos. Experimental results demonstrate that our proposed scheme achieves satisfying perceptual performance at ultra-low bitrates and exhibits higher semantic fidelity in semantic analysis tasks such as human parsing. Code is available at:\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" ext-link-type=\"uri\" xlink:href=\"https:\/\/github.com\/Chen8023\/LPVC\">https:\/\/github.com\/Chen8023\/LPVC<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1145\/3772085","type":"journal-article","created":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T14:47:41Z","timestamp":1761576461000},"page":"1-23","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Portrait Video Compression with Semantic-guided Animation Model and Background Incremental Coding"],"prefix":"10.1145","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6026-8020","authenticated-orcid":false,"given":"Xinyi","family":"Chen","sequence":"first","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1877-7355","authenticated-orcid":false,"given":"Weimin","family":"Lei","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7446-4025","authenticated-orcid":false,"given":"Wei","family":"Zhang","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-7945-9804","authenticated-orcid":false,"given":"Wenhui","family":"Ye","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5471-6296","authenticated-orcid":false,"given":"Yanwen","family":"Wang","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,1,13]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"Miko\u0142aj Bi\u0144kowski Danica J. Sutherland Michael Arbel and Arthur Gretton. 2018. Demystifying MMD GANs. arXiv:1801.01401. Retrieved from https:\/\/arxiv.org\/abs\/1801.01401"},{"key":"e_1_3_1_3_2","unstructured":"Gisle Bjontegaard. 2001. Calculation of Average PSNR Differences between RD-Curves. ITU SG16 Doc. VCEG-M33. 2001."},{"key":"e_1_3_1_4_2","unstructured":"B. Bross J. Chen S. Liu and Y. K. Wang. 2020. Versatile Video Coding (Draft 10). Jvet-S2001."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/DCC52660.2022.00009"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3271130"},{"key":"e_1_3_1_7_2","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1109\/PCS.2018.8456249","volume-title":"Proceedings of the 2018 Picture Coding Symposium (PCS)","author":"Chen Yue","year":"2018","unstructured":"Yue Chen, Debargha Murherjee, Jingning Han, Adrian Grange, Yaowu Xu, Zoe Liu, Sarah Parker, Cheng Chen, Hui Su, Urvang Joshi, et al. 2018. An overview of core coding tools in the AV1 video codec. In Proceedings of the 2018 Picture Coding Symposium (PCS), 41\u201345."},{"issue":"5","key":"e_1_3_1_8_2","first-page":"2567","article-title":"Image quality assessment: Unifying structure and texture similarity","volume":"44","author":"Ding Keyan","year":"2020","unstructured":"Keyan Ding, Kede Ma, Shiqi Wang, and Eero P. Simoncelli. 2020. Image quality assessment: Unifying structure and texture similarity. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 5 (2020), 2567\u20132581.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3419394.3423658"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3130754"},{"key":"e_1_3_1_11_2","article-title":"GANs trained by a two time-scale update rule converge to a local nash equilibrium","volume":"30","author":"Heusel Martin","year":"2017","unstructured":"Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017. GANs trained by a two time-scale update rule converge to a local nash equilibrium. In Advances in Neural Information Processing Systems, Vol. 30 (2017).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cogr.2021.08.003"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2024.3350643"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.632"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i1.19999"},{"key":"e_1_3_1_17_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_1_18_2","volume-title":"Proceedings of the IEEE International Conference on Image Processing (ICIP)","author":"Konuko Goluck","year":"2022","unstructured":"Goluck Konuko, St\u00e9phane Lathuili\u00e8re, and Giuseppe Valenzise. 2022. A hybrid deep animation codec for low-bitrate video conferencing. In Proceedings of the IEEE International Conference on Image Processing (ICIP)."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP49359.2023.10222205"},{"key":"e_1_3_1_20_2","first-page":"4210","volume-title":"Proceedings of the ICASSP 2021\u20132021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)","author":"Konuko Goluck","year":"2021","unstructured":"Goluck Konuko, Giuseppe Valenzise, and St\u00e9phane Lathuili\u00e8re. 2021. Ultra-low bitrate video conferencing using deep image animation. In Proceedings of the ICASSP 2021\u20132021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 4210\u20134214."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3581783.3612530"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2020.3048039"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3091909"},{"key":"e_1_3_1_24_2","first-page":"3546","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Lin Jianping","year":"2020","unstructured":"Jianping Lin, Dong Liu, Houqiang Li, and Feng Wu. 2020. M-LVC: Multiple frames prediction for learned video compression. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3546\u20133554."},{"key":"e_1_3_1_25_2","first-page":"11006","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Lu Guo","year":"2019","unstructured":"Guo Lu, Wanli Ouyang, Dong Xu, Xiaoyun Zhang, Chunlei Cai, and Zhiyong Gao. 2019. DVC: An end-to-end deep video compression framework. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 11006\u201311015."},{"key":"e_1_3_1_26_2","first-page":"3139","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV)","author":"Misra Diganta","year":"2021","unstructured":"Diganta Misra, Trikay Nalamada, Ajay Uppili Arasanipalai, and Qibin Hou. 2021. Rotate to attend: Convolutional triplet attention module. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), 3139\u20133148."},{"key":"e_1_3_1_27_2","doi-asserted-by":"crossref","first-page":"390","DOI":"10.1109\/PCS.2013.6737765","volume-title":"Proceedings of the 2013 Picture Coding Symposium (PCS)","author":"Mukherjee Debargha","year":"2013","unstructured":"Debargha Mukherjee, Jim Bankoski, Adrian Grange, Jingning Han, John Koleszar, Paul Wilkins, Yaowu Xu, and Ronald Bultje. 2013. The latest open-source video codec VP9\u2013An overview and preliminary results. In Proceedings of the 2013 Picture Coding Symposium (PCS). IEEE, 390\u2013393."},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-950"},{"key":"e_1_3_1_29_2","first-page":"2388","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Oquab Maxime","year":"2021","unstructured":"Maxime Oquab, Pierre Stock, Daniel Haziza, Tao Xu, Peizhao Zhang, Onur Celebi, Yana Hasson, Patrick Labatut, Bobo Bose-Kolanu, Thibault Peyronel, et al. 2021. Low bandwidth video-chat compression using deep generative models. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2388\u20132397."},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-24574-4_28"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2011.6126544"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.jvcir.2022.103683"},{"key":"e_1_3_1_33_2","article-title":"First order motion model for image animation","volume":"32","author":"Siarohin Aliaksandr","year":"2019","unstructured":"Aliaksandr Siarohin, St\u00e9phane Lathuili\u00e8re, Sergey Tulyakov, Elisa Ricci, and Nicu Sebe. 2019. First order motion model for image animation. In Advances in Neural Information Processing Systems, Vol. 32 (2019).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_34_2","first-page":"13653","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Siarohin Aliaksandr","year":"2021","unstructured":"Aliaksandr Siarohin, Oliver J. Woodford, Jian Ren, Menglei Chai, and Sergey Tulyakov. 2021. Motion representations for articulated animation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 13653\u201313662."},{"key":"e_1_3_1_35_2","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"e_1_3_1_36_2","unstructured":"Sri Srinivasan. 2020. Cisco Webex: Supporting Customers during this Unprecedented Time. Retrieved August 2020 from https:\/\/gblogs.cisco.com\/la\/cl-pmarrone-en-cisco-webex-supporting-customers-during-this-unprecedented-time\/"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2012.2221191"},{"key":"e_1_3_1_38_2","first-page":"1","volume-title":"Proceedings of the 2022 IEEE International Conference on Multimedia and Expo (ICME)","author":"Tang Anni","year":"2022","unstructured":"Anni Tang, Yan Huang, Jun Ling, Zhiyu Zhang, Yiwei Zhang, Rong Xie, and Li Song. 2022. Generative compression for face video: A hybrid scheme. In Proceedings of the 2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1\u20136."},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3661311"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3715144"},{"key":"e_1_3_1_41_2","first-page":"1738","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Volokitin Anna","year":"2022","unstructured":"Anna Volokitin, Stefan Brugger, Ali Benlalah, Sebastian Martin, Brian Amberg, and Michael Tschannen. 2022. Neural face video compression using multiple views. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 1738\u20131742."},{"key":"e_1_3_1_42_2","first-page":"1","volume-title":"Proceedings of the 2022 IEEE International Conference on Multimedia and Expo (ICME)","author":"Wang Ruofan","year":"2022","unstructured":"Ruofan Wang, Qi Mao, Shiqi Wang, Chuanmin Jia, Ronggang Wang, and Siwei Ma. 2022. Disentangled visual representations for extreme human body video compression. In Proceedings of the 2022 IEEE International Conference on Multimedia and Expo (ICME). IEEE, 1\u20136."},{"key":"e_1_3_1_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3063165"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP46576.2022.9897729"},{"key":"e_1_3_1_45_2","first-page":"1","volume-title":"Proceedings of the 2021 IEEE International Conference on Multimedia & Expo Workshops (ICMEW)","author":"Wieckowski Adam","year":"2021","unstructured":"Adam Wieckowski, Jens Brandenburg, Tobias Hinz, Christian Bartnik, Valeri George, Gabriel Hege, Christian Helmrich, Anastasia Henkel, Christian Lehmann, Christian Stoffers, et al. 2021. VVenC: An open and optimized VVC encoder implementation. In Proceedings of the 2021 IEEE International Conference on Multimedia & Expo Workshops (ICMEW). IEEE, 1\u20132."},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2003.815165"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/214762.214771"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_1"},{"key":"e_1_3_1_49_2","first-page":"1","volume-title":"Proceedings of the 2020 IEEE International Symposium on Circuits and Systems (ISCAS)","author":"Wu Yaojun","year":"2020","unstructured":"Yaojun Wu, Tianyu He, and Zhibo Chen. 2020. Memorize, then recall: A generative framework for low bit-rate surveillance video compression. In Proceedings of the 2020 IEEE International Symposium on Circuits and Systems (ISCAS). IEEE, 1\u20135."},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1145\/3636510"},{"key":"e_1_3_1_51_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3195904"},{"key":"e_1_3_1_52_2","first-page":"6628","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Yang Ren","year":"2020","unstructured":"Ren Yang, Fabian Mentzer, Luc Van Gool, and Radu Timofte. 2020. Learning for video compression with hierarchical quality and recurrent enhancement. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6628\u20136637."},{"key":"e_1_3_1_53_2","unstructured":"Polina Zablotskaia Aliaksandr Siarohin Bo Zhao and Leonid Sigal. 2019. DwNet: Dense warp-based network for pose-guided human video generation. arXiv:1910.09139. Retrieved from https:\/\/arxiv.org\/abs\/1910.09139"},{"key":"e_1_3_1_54_2","first-page":"1","volume-title":"Proceedings of the 2019 Picture Coding Symposium (PCS)","author":"Zhang Jiaqi","year":"2019","unstructured":"Jiaqi Zhang, Chuanmin Jia, Meng Lei, Shanshe Wang, Siwei Ma, and Wen Gao. 2019. Recent development of AVS video coding standard: AVS3. In Proceedings of the 2019 Picture Coding Symposium (PCS). IEEE, 1\u20135."},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_3_1_56_2","first-page":"3657","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhao Jian","year":"2022","unstructured":"Jian Zhao and Hui Zhang. 2022. Thin-plate spline motion model for image animation. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3657\u20133666."},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2022.3169951"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3563699"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2024.3420435"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3772085","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,13]],"date-time":"2026-01-13T14:20:13Z","timestamp":1768314013000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3772085"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,13]]},"references-count":58,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,1,31]]}},"alternative-id":["10.1145\/3772085"],"URL":"https:\/\/doi.org\/10.1145\/3772085","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,13]]},"assertion":[{"value":"2025-04-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-10-12","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-13","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}