{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,2]],"date-time":"2026-04-02T17:50:51Z","timestamp":1775152251484,"version":"3.50.1"},"reference-count":54,"publisher":"MDPI AG","issue":"6","license":[{"start":{"date-parts":[[2025,6,10]],"date-time":"2025-06-10T00:00:00Z","timestamp":1749513600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012226","name":"Fundamental Research Funds for the Central Universities of China","doi-asserted-by":"crossref","award":["TN2216010"],"award-info":[{"award-number":["TN2216010"]}],"id":[{"id":"10.13039\/501100012226","id-type":"DOI","asserted-by":"crossref"}]},{"name":"\u2018Jie Bang Gua Shuai\u2019 Science and Technology Major Project of Liaoning Province in 2022","award":["2022JH1\/10400025"],"award-info":[{"award-number":["2022JH1\/10400025"]}]},{"name":"National Key Research and Development Program of China","award":["2018YFB1702000"],"award-info":[{"award-number":["2018YFB1702000"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Deep neural video compression codecs have shown great promise in recent years. However, there are still considerable challenges for ultra-low bitrate video coding. Inspired by recent diffusion models for image and video compression attempts, we attempt to leverage diffusion models for ultra-low bitrate portrait video compression. In this paper, we propose a predictive portrait video compression method that leverages the temporal prediction capabilities of diffusion models. Specifically, we develop a temporal diffusion predictor based on a conditional latent diffusion model, with the predicted results serving as decoded frames. We symmetrically integrate a temporal diffusion predictor at the encoding and decoding side, respectively. When the perceptual quality of the predicted results in encoding end falls below a predefined threshold, a new frame sequence is employed for prediction. While the predictor at the decoding side directly generates predicted frames as reconstruction based on the evaluation results. This symmetry ensures that the prediction frames generated at the decoding end are consistent with those at the encoding end. We also design an adaptive coding strategy that incorporates frame quality assessment and adaptive keyframe control. To ensure consistent quality of subsequent predicted frames and achieve high perceptual reconstruction, this strategy dynamically evaluates the visual quality of the predicted results during encoding, retains the predicted frames that meet the quality threshold, and adaptively adjusts the length of the keyframe sequence based on motion complexity. The experimental results demonstrate that, compared with the traditional video codecs and other popular methods, the proposed scheme provides superior compression performance at ultra-low bitrates while maintaining competitiveness in visual effects, achieving more than 24% bitrate savings compared with VVC in terms of perceptual distortion.<\/jats:p>","DOI":"10.3390\/sym17060913","type":"journal-article","created":{"date-parts":[[2025,6,10]],"date-time":"2025-06-10T05:11:49Z","timestamp":1749532309000},"page":"913","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Ultra-Low Bitrate Predictive Portrait Video Compression with Diffusion Models"],"prefix":"10.3390","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6026-8020","authenticated-orcid":false,"given":"Xinyi","family":"Chen","sequence":"first","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang 110169, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1877-7355","authenticated-orcid":false,"given":"Weimin","family":"Lei","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang 110169, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7446-4025","authenticated-orcid":false,"given":"Wei","family":"Zhang","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang 110169, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-5471-6296","authenticated-orcid":false,"given":"Yanwen","family":"Wang","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang 110169, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-7070-8707","authenticated-orcid":false,"given":"Mingxin","family":"Liu","sequence":"additional","affiliation":[{"name":"The School of Computer Science and Engineering, Northeastern University, Shenyang 110169, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,6,10]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"560","DOI":"10.1109\/TCSVT.2003.815165","article-title":"Overview of the H. 264\/AVC video coding standard","volume":"13","author":"Wiegand","year":"2003","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1649","DOI":"10.1109\/TCSVT.2012.2221191","article-title":"Overview of the high efficiency video coding (HEVC) standard","volume":"22","author":"Sullivan","year":"2012","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"3736","DOI":"10.1109\/TCSVT.2021.3101953","article-title":"Overview of the versatile video coding (VVC) standard and its applications","volume":"31","author":"Bross","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"145","DOI":"10.1016\/j.cogr.2021.08.003","article-title":"Recent trending on learning based video compression: A survey","volume":"1","author":"Hoang","year":"2021","journal-title":"Cogn. Robot."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Lu, G., Ouyang, W., Xu, D., Zhang, X., Cai, C., and Gao, Z. (2019, January 15\u201320). Dvc: An end-to-end deep video compression framework. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01126"},{"key":"ref_6","first-page":"18114","article-title":"Deep contextual video compression","volume":"34","author":"Li","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_7","first-page":"6840","article-title":"Denoising diffusion probabilistic models","volume":"33","author":"Ho","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_8","first-page":"23371","article-title":"Mcvd-masked conditional video diffusion for prediction, generation, and interpolation","volume":"35","author":"Voleti","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Yang, R., Srivastava, P., and Mandt, S. (2023). Diffusion probabilistic modeling for video generation. Entropy, 25.","DOI":"10.3390\/e25101469"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Hu, J., Cheng, W., Paudel, D., and Yang, J. (2024, January 16\u201322). Extdm: Distribution extrapolation diffusion model for video prediction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01827"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B. (2022, January 18\u201324). High-resolution image synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"ref_12","unstructured":"Pan, Z., Zhou, X., and Tian, H. (2022). Extreme generative image compression by learning text embedding from diffusion models. arXiv."},{"key":"ref_13","first-page":"64971","article-title":"Lossy image compression with conditional diffusion models","volume":"36","author":"Yang","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Chen, L., Li, Z., Lin, B., Zhu, B., Wang, Q., Yuan, S., Zhou, X., Cheng, X., and Yuan, L. (2024). Od-vae: An omni-dimensional video compressor for improving latent video diffusion model. arXiv.","DOI":"10.1109\/ICME59968.2025.11210036"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Ma, W., and Chen, Z. (2025). Diffusion-based perceptual neural video compression with temporal diffusion information reuse. arXiv.","DOI":"10.1145\/3761815"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Li, B., Liu, Y., Niu, X., Bait, B., Han, W., Deng, L., and Gunduz, D. (2024, January 24\u201326). Extreme Video Compression with Prediction Using Pre-trained Diffusion Models. Proceedings of the 2024 16th International Conference on Wireless Communications and Signal Processing (WCSP), Hefei, China.","DOI":"10.1109\/WCSP62071.2024.10827256"},{"key":"ref_17","unstructured":"Yu, L., Lezama, J., Gundavarapu, N.B., Versari, L., Sohn, K., Minnen, D., Cheng, Y., Birodkar, V., Gupta, A., and Gu, X. (2023). Language Model Beats Diffusion\u2013Tokenizer is Key to Visual Generation. arXiv."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2020","DOI":"10.1109\/TPAMI.2024.3515454","article-title":"Bevformer: Learning bird\u2019s-eye-view representation from lidar-camera via spatiotemporal transformers","volume":"47","author":"Li","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Mukherjee, D., Bankoski, J., Grange, A., Han, J., Koleszar, J., Wilkins, P., Xu, Y., and Bultje, R. (2013, January 8\u201311). The latest open-source video codec VP9-an overview and preliminary results. Proceedings of the 2013 Picture Coding Symposium (PCS), San Jose, CA, USA.","DOI":"10.1109\/PCS.2013.6737765"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Chen, Y., Murherjee, D., Han, J., Grange, A., Xu, Y., Liu, Z., Parker, S., Chen, C., Su, H., and Joshi, U. (2018, January 24\u201327). An overview of core coding tools in the AV1 video codec. Proceedings of the 2018 Picture Coding Symposium (PCS), San Francisco, CA, USA.","DOI":"10.1109\/PCS.2018.8456249"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Zhang, J., Jia, C., Lei, M., Wang, S., Ma, S., and Gao, W. (2019, January 11\u201315). Recent development of AVS video coding standard: AVS3. Proceedings of the 2019 Picture Coding Symposium (PCS), Ningbo, China.","DOI":"10.1109\/PCS48520.2019.8954503"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Wieckowski, A., Brandenburg, J., Hinz, T., Bartnik, C., George, V., Hege, G., Helmrich, C., Henkel, A., Lehmann, C., and Stoffers, C. (2021, January 5\u20139). VVenC: An open and optimized VVC encoder implementation. Proceedings of the 2021 IEEE International Conference on Multimedia & Expo Workshops (ICMEW), Shenzhen, China.","DOI":"10.1109\/ICMEW53276.2021.9455944"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"1363","DOI":"10.1007\/s11760-021-02088-w","article-title":"Improved intra-subpartition coding mode for versatile video coding","volume":"16","author":"Akbulut","year":"2022","journal-title":"Signal Image Video Process"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"2777","DOI":"10.1007\/s11042-021-11678-2","article-title":"Machine Learning-Based approaches to reduce HEVC intra coding unit partition decision complexity","volume":"81","author":"Amna","year":"2022","journal-title":"Multimed. Tools Appl."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"1481","DOI":"10.1109\/TCSVT.2022.3195904","article-title":"DFCE: Decoder-friendly chrominance enhancement for HEVC intra coding","volume":"33","author":"Yang","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"6411","DOI":"10.1109\/TMM.2022.3208516","article-title":"Efficient VVC intra prediction based on deep feature fusion and probability estimation","volume":"25","author":"Zhao","year":"2022","journal-title":"IEEE Trans. Multimed."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"826","DOI":"10.1109\/TCSVT.2021.3063165","article-title":"Neural network-based enhancement to inter prediction for video coding","volume":"32","author":"Wang","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"103683","DOI":"10.1016\/j.jvcir.2022.103683","article-title":"Low complexity inter coding scheme for Versatile Video Coding (VVC)","volume":"90","author":"Shang","year":"2023","journal-title":"J. Vis. Commun. Image Represent."},{"key":"ref_29","first-page":"1901","article-title":"Convolutional neural network-based arithmetic coding for HEVC intra-predicted residues","volume":"30","author":"Ma","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"7222","DOI":"10.1109\/TIP.2022.3221278","article-title":"Deformable Wiener Filter for Future Video Coding","volume":"31","author":"Meng","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"12092","DOI":"10.1109\/TCSVT.2024.3420435","article-title":"Neural Network Based Multi-Level In-Loop Filtering for Versatile Video Coding","volume":"34","author":"Zhu","year":"2024","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Hu, Z., Lu, G., and Xu, D. (2021, January 20\u201325). FVC: A New Framework towards Deep Video Compression in Feature Space. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00155"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"673","DOI":"10.1109\/LSP.2023.3277343","article-title":"Enhanced motion compensation for deep video compression","volume":"30","author":"Guo","year":"2023","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"17541","DOI":"10.1109\/ACCESS.2024.3350643","article-title":"HDVC: Deep video compression with hyperprior-based entropy coding","volume":"12","author":"Hu","year":"2024","journal-title":"IEEE Access"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Li, J., Li, B., and Lu, Y. (2022, January 10\u201314). Hybrid spatial-temporal entropy modelling for neural video compression. Proceedings of the 30th ACM International Conference on Multimedia, Lisbon, Portugal.","DOI":"10.1145\/3503161.3547845"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Li, J., Li, B., and Lu, Y. (2023, January 18\u201322). Neural video compression with diverse contexts. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02166"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"4579","DOI":"10.1109\/TPAMI.2024.3356548","article-title":"Vnvc: A versatile neural video coding framework for efficient human-machine vision","volume":"46","author":"Sheng","year":"2024","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, J., Li, B., and Lu, Y. (2024, January 16\u201322). Neural video compression with feature modulation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.02466"},{"key":"ref_39","first-page":"139","article-title":"Generative adversarial nets","volume":"27","author":"Goodfellow","year":"2014","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Konuko, G., Valenzise, G., and Lathuili\u00e8re, S. (2021, January 6\u201311). Ultra-low bitrate video conferencing using deep image animation. Proceedings of the ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada.","DOI":"10.1109\/ICASSP39728.2021.9414731"},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Konuko, G., Lathuili\u00e8re, S., and Valenzise, G. (2022, January 16\u201319). A hybrid deep animation codec for low-bitrate video conferencing. Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP), Bordeaux, France.","DOI":"10.1109\/ICIP46576.2022.10458867"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Konuko, G., Lathuili\u00e8re, S., and Valenzise, G. (2023, January 8\u201311). Predictive coding for animation-based video compression. Proceedings of the 2023 IEEE International Conference on Image Processing (ICIP), Kuala Lumpur, Malaysia.","DOI":"10.1109\/ICIP49359.2023.10222205"},{"key":"ref_43","unstructured":"Kingma, D.P., and Welling, M. (2014, January 14\u201316). Auto-Encoding Variational Bayes. Proceedings of the International Conference on Learning Representations, Banff, AB, Canada."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhang, R., Isola, P., Efros, A.A., Shechtman, E., and Wang, O. (2018, January 18\u201323). The unreasonable effectiveness of deep features as a perceptual metric. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"520","DOI":"10.1145\/214762.214771","article-title":"Arithmetic coding for data compression","volume":"30","author":"Witten","year":"1987","journal-title":"Commun. ACM"},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"311","DOI":"10.1007\/s00530-024-01499-2","article-title":"Model-based portrait video compression with spatial constraint and adaptive pose processing","volume":"30","author":"Chen","year":"2024","journal-title":"Multimed. Syst."},{"key":"ref_47","unstructured":"Gu, X., Wen, C., Ye, W., Song, J., and Gao, Y. (2023). Seer: Language instructed video prediction with latent diffusion models. arXiv."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Blattmann, A., Rombach, R., Ling, H., Dockhorn, T., Kim, S.W., Fidler, S., and Kreis, K. (2023, January 17\u201324). Align your latents: High-resolution video synthesis with latent diffusion models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.02161"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Zhang, W., Zhu, M., and Derpanis, K.G. (2013, January 1\u20138). From actemes to action: A strongly-supervised representation for detailed action understanding. Proceedings of the IEEE International Conference on Computer Vision, Sydney, Australia.","DOI":"10.1109\/ICCV.2013.280"},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Zhao, L., Peng, X., Tian, Y., Kapadia, M., and Metaxas, D. (2018, January 8\u201314). Learning to forecast and refine residual motion for image-to-video generation. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_24"},{"key":"ref_51","first-page":"7137","article-title":"First order motion model for image animation","volume":"32","author":"Siarohin","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_52","unstructured":"Zablotskaia, P., Siarohin, A., Zhao, B., and Sigal, L. (2019). Dwnet: Dense warp-based network for pose-guided human video generation. arXiv."},{"key":"ref_53","first-page":"2567","article-title":"Image quality assessment: Unifying structure and texture similarity","volume":"44","author":"Ding","year":"2020","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_54","unstructured":"Song, J., Meng, C., and Ermon, S. (2020). Denoising diffusion implicit models. arXiv."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/6\/913\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,2]],"date-time":"2026-04-02T17:02:44Z","timestamp":1775149364000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/6\/913"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,10]]},"references-count":54,"journal-issue":{"issue":"6","published-online":{"date-parts":[[2025,6]]}},"alternative-id":["sym17060913"],"URL":"https:\/\/doi.org\/10.3390\/sym17060913","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,6,10]]}}}