{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T09:56:15Z","timestamp":1781690175624,"version":"3.54.5"},"reference-count":61,"publisher":"Wiley","issue":"7","license":[{"start":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T00:00:00Z","timestamp":1780444800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T00:00:00Z","timestamp":1780444800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100008043","name":"Zhejiang Sci-Tech University","doi-asserted-by":"publisher","award":["22062338\u2010Y"],"award-info":[{"award-number":["22062338\u2010Y"]}],"id":[{"id":"10.13039\/501100008043","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Expert Systems"],"published-print":{"date-parts":[[2026,7]]},"abstract":"<jats:title>ABSTRACT<\/jats:title>\n                  <jats:p>\n                    The proliferation of unmanned aerial vehicles (UAVs) in the low\u2010altitude economy has driven a surge in demand for high\u2010fidelity video transmission; however, existing video codec methods struggle to balance the conflicting requirements of resource\u2010constrained onboard platforms and high\u2010quality downstream visual perception. While conventional macroscopic codecs offer computational efficiency, they suffer from artefacts and bandwidth congestion in complex UAV environments. Conversely, microscopic neural video compression (NVC) models achieve superior rate\u2010distortion performance but incur prohibitive computational overhead. To reconcile these limitations, we propose the Generative Content\u2010steering Video Codec (GCVC), a novel framework designed from a mesoscopic perspective. GCVC synergises lightweight architectural design with generative semantic recovery by integrating streamlined encoding with a Stable Diffusion\u2010based (SD\u2010based) frame predictor. By employing a strategic frame\u2010skipping protocol to alleviate transmission bottlenecks and utilising the generative prior to synthesise missing scenes, our framework achieves state\u2010of\u2010the\u2010art reconstruction fidelity at significantly lower bitrates. Extensive experiments on public and self\u2010collected UAV datasets demonstrate that GCVC outperforms both traditional standards and cutting\u2010edge NVC approaches, establishing a new benchmark for efficient, perception\u2010aware video compression in intelligent UAV systems. The implementation codes will be publicly available at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/ZSTU-CV-Lab\/GCVC\">https:\/\/github.com\/ZSTU\u2010CV\u2010Lab\/GCVC<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1111\/exsy.70308","type":"journal-article","created":{"date-parts":[[2026,6,3]],"date-time":"2026-06-03T23:40:43Z","timestamp":1780530043000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Mesoscopic Insights: Generative Content\u2010Steering Video Codec for\n                    <scp>UAVs<\/scp>"],"prefix":"10.1111","volume":"43","author":[{"given":"Qincheng","family":"Wen","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology (School of Artificial Intelligence) Zhejiang Sci\u2010Tech University  Hangzhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yi","family":"Shen","sequence":"additional","affiliation":[{"name":"Department of Mathematics Zhejiang Sci\u2010Tech University  Hangzhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chenyu","family":"Liu","sequence":"additional","affiliation":[{"name":"Department of Mathematics Zhejiang Sci\u2010Tech University  Hangzhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jun","family":"Xu","sequence":"additional","affiliation":[{"name":"Department of Mathematics Zhejiang Sci\u2010Tech University  Hangzhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yande","family":"Li","sequence":"additional","affiliation":[{"name":"Centre for Perceptual and Interactive Intelligence The Chinese University of Hong Kong  Hong Kong China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7346-1110","authenticated-orcid":false,"given":"Mingjie","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Mathematics Zhejiang Sci\u2010Tech University  Hangzhou China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2026,6,3]]},"reference":[{"key":"e_1_2_9_2_1","doi-asserted-by":"crossref","unstructured":"Agustsson E. D.Minnen N.Johnston J.Balle S. J.Hwang andG.Toderici.2020.\u201cScale\u2010Space Flow for End\u2010to\u2010End Optimized Video Compression.\u201dInProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 8503\u20138512.","DOI":"10.1109\/CVPR42600.2020.00853"},{"key":"e_1_2_9_3_1","unstructured":"Ball\u00e9 J. V.Laparra andE. P.Simoncelli.2016.\u201cEnd\u2010to\u2010End Optimized Image Compression.\u201darXiv Preprint arXiv:1611.01704."},{"key":"e_1_2_9_4_1","unstructured":"Ball\u00e9 J. D.Minnen S.Singh S. J.Hwang andN.Johnston.2018.\u201cVariational Image Compression With a Scale Hyperprior.\u201darXiv Preprint arXiv:1802.01436."},{"key":"e_1_2_9_5_1","doi-asserted-by":"publisher","DOI":"10.1111\/exsy.13398"},{"key":"e_1_2_9_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10586-022-03627-x"},{"key":"e_1_2_9_7_1","doi-asserted-by":"publisher","DOI":"10.3390\/rs14030620"},{"key":"e_1_2_9_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2023.3247433"},{"key":"e_1_2_9_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSAC.2023.3310097"},{"key":"e_1_2_9_10_1","doi-asserted-by":"crossref","unstructured":"Cheng Z. H.Sun M.Takeuchi andJ.Katto.2020.\u201cLearned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention Modules.\u201dInProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 7939\u20137948.","DOI":"10.1109\/CVPR42600.2020.00796"},{"key":"e_1_2_9_11_1","doi-asserted-by":"crossref","unstructured":"Du D. Y.Qi H.Yu et\u00a0al.2018.\u201cThe Unmanned Aerial Vehicle Benchmark: Object Detection and Tracking.\u201dInProceedings of the European Conference on Computer Vision (Eccv) 370\u2013386.","DOI":"10.1007\/978-3-030-01249-6_23"},{"key":"e_1_2_9_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAES.2021.3087821"},{"key":"e_1_2_9_13_1","doi-asserted-by":"crossref","unstructured":"Fu T. H.Zhang F.Mu andH.Chen.2019.\u201cFast Cu Partitioning Algorithm for h. 266\/Vvc Intra\u2010Frame Coding.\u201dIn2019 IEEE International Conference on Multimedia and Expo (Icme) 55\u201360.","DOI":"10.1109\/ICME.2019.00018"},{"key":"e_1_2_9_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3422622"},{"key":"e_1_2_9_15_1","doi-asserted-by":"crossref","unstructured":"Habibian A. T. v.Rozendaal J. M.Tomczak andT. S.Cohen.2019.\u201cVideo Compression With Rate\u2010Distortion Autoencoders.\u201dInProceedings of the Ieee\/Cvf International Conference on Computer Vision 7033\u20137042.","DOI":"10.1109\/ICCV.2019.00713"},{"key":"e_1_2_9_16_1","doi-asserted-by":"crossref","unstructured":"He D. Z.Yang W.Peng R.Ma H.Qin andY.Wang.2022.\u201cElic: Efficient Learned Image Compression With Unevenly Grouped Space\u2010Channel Contextual Adaptive Coding.\u201dInProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 5718\u20135727.","DOI":"10.1109\/CVPR52688.2022.00563"},{"key":"e_1_2_9_17_1","doi-asserted-by":"crossref","unstructured":"He D. Y.Zheng B.Sun Y.Wang andH.Qin.2021.\u201cCheckerboard Context Model for Efficient Learned Image Compression.\u201dInProceedings of the Ieee\/Cvf Conference on Computer Vision and Pattern Recognition 14771\u201314780.","DOI":"10.1109\/CVPR46437.2021.01453"},{"key":"e_1_2_9_18_1","first-page":"6840","article-title":"Denoising Diffusion Probabilistic Models","volume":"33","author":"Ho J.","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_9_19_1","doi-asserted-by":"crossref","unstructured":"Hu Z. G.Lu J.Guo S.Liu W.Jiang andD.Xu.2022.\u201cCoarse\u2010to\u2010Fine Deep Video Coding With Hyperprior\u2010Guided Mode Prediction.\u201dInProceedings of the Ieee\/Cvf Conference on Computer Vision and Pattern Recognition 5921\u20135930.","DOI":"10.1109\/CVPR52688.2022.00583"},{"key":"e_1_2_9_20_1","unstructured":"Khanpour A. T.Wang A.Vahidi\u2010Shams et\u00a0al.2025.\u201cUav\u2010Based Intelligent Traffic Surveillance System: Real\u2010Time Vehicle Detection Classification Tracking and Behavioral Analysis.\u201darXiv Preprint arXiv:2509.04624."},{"key":"e_1_2_9_21_1","unstructured":"Kingma D. P. andM.Welling.2013.\u201cAuto\u2010Encoding Variational Bayes.\u201darXiv Preprint arXiv: 1312.6114."},{"key":"e_1_2_9_22_1","first-page":"18114","article-title":"Deep Contextual Video Compression","volume":"34","author":"Li J.","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_9_23_1","doi-asserted-by":"crossref","unstructured":"Li J. B.Li andY.Lu.2022.\u201cHybrid Spatial\u2010Temporal Entropy Modelling for Neural Video Compression.\u201dInProceedings of the 30th Acm International Conference on Multimedia 1503\u20131511.","DOI":"10.1145\/3503161.3547845"},{"key":"e_1_2_9_24_1","doi-asserted-by":"crossref","unstructured":"Li J. B.Li andY.Lu.2023.\u201cNeural Video Compression With Diverse Contexts.\u201dInProceedings of the Ieee\/Cvf Conference on Computer Vision and Pattern Recognition 22616\u201322626.","DOI":"10.1109\/CVPR52729.2023.02166"},{"key":"e_1_2_9_25_1","doi-asserted-by":"crossref","unstructured":"Li J. B.Li andY.Lu.2024.\u201cNeural Video Compression With Feature Modulation.\u201dInProceedings of the Ieee\/Cvf Conference on Computer Vision and Pattern Recognition 26099\u201326108.","DOI":"10.1109\/CVPR52733.2024.02466"},{"key":"e_1_2_9_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TFUZZ.2024.3400861"},{"key":"e_1_2_9_27_1","doi-asserted-by":"crossref","unstructured":"Lin J. D.Liu H.Li andF.Wu.2020.\u201cM\u2010Lvc: Multiple Frames Prediction for Learned Video Compression.\u201dInProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition 3546\u20133554.","DOI":"10.1109\/CVPR42600.2020.00360"},{"key":"e_1_2_9_28_1","doi-asserted-by":"crossref","unstructured":"Liu J. H.Sun andJ.Katto.2023.\u201cLearned Image Compression With Mixed Transformer\u2010Cnn Architectures.\u201dInProceedings of the Ieee\/Cvf Conference on Computer Vision and Pattern Recognition 14388\u201314397.","DOI":"10.1109\/CVPR52729.2023.01383"},{"key":"e_1_2_9_29_1","doi-asserted-by":"crossref","unstructured":"Liu S. X.Li H.Lu andY.He.2022.\u201cMulti\u2010Object Tracking Meets Moving Uav.\u201dInProceedings of the Ieee\/Cvf Conference on Computer Vision and Pattern Recognition 8876\u20138885.","DOI":"10.1109\/CVPR52688.2022.00867"},{"key":"e_1_2_9_30_1","first-page":"1","article-title":"Diffusion\u2010Based Perceptual Neural Video Compression With Temporal Diffusion Information Reuse","volume":"21","author":"Ma W.","year":"2025","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications, Diffusion\u2010Based Perceptual Neural Video Compression With Temporal Diffusion Information Reuse"},{"key":"e_1_2_9_31_1","doi-asserted-by":"crossref","unstructured":"Ma W. andZ.Chen.2025b.\u201cDiffvc\u2010Osd: One\u2010Step Diffusion\u2010Based Perceptual Neural Video Compression Framework.\u201darXiv Preprint arXiv:2508.07682.","DOI":"10.1109\/VCIP67698.2025.11396819"},{"key":"e_1_2_9_32_1","first-page":"10794","article-title":"Joint Autoregressive and Hierarchical Priors for Learned Image Compression","author":"Minnen D.","year":"2018","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_9_33_1","doi-asserted-by":"publisher","DOI":"10.3390\/drones6060147"},{"key":"e_1_2_9_34_1","doi-asserted-by":"publisher","DOI":"10.1111\/exsy.70033"},{"key":"e_1_2_9_35_1","doi-asserted-by":"publisher","DOI":"10.1109\/TBC.2016.2580920"},{"key":"e_1_2_9_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2015.2428571"},{"key":"e_1_2_9_37_1","unstructured":"Podell D. Z.English K.Lacey et\u00a0al.2023.\u201cSdxl: Improving Latent Diffusion Models for High\u2010Resolution Image Synthesis.\u201darXiv Preprint arXiv:2307.01952."},{"key":"e_1_2_9_38_1","doi-asserted-by":"crossref","unstructured":"Ranasinghe Y. N. G.Nair W. G. C.Bandara andV. M.Patel.2024.\u201cCrowddiff: Multi\u2010Hypothesis Crowd Density Estimation Using Diffusion Models.\u201dInProceedings of the Ieee\/Cvf Conference on Computer Vision and Pattern Recognition 12809\u201312819.","DOI":"10.1109\/CVPR52733.2024.01217"},{"key":"e_1_2_9_39_1","doi-asserted-by":"crossref","unstructured":"Rombach R. A.Blattmann D.Lorenz P.Esser andB.Ommer.2022.\u201cHigh\u2010Resolution Image Synthesis With Latent Diffusion Models.\u201dInProceedings of the Ieee\/Cvf Conference on Computer Vision and Pattern Recognition 10684\u201310695.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_9_40_1","unstructured":"Salimans T. andJ.Ho.2022.\u201cProgressive Distillation for Fast Sampling of Diffusion Models.\u201darXiv Preprint arXiv:2202.00512."},{"key":"e_1_2_9_41_1","doi-asserted-by":"crossref","unstructured":"Sauer A. D.Lorenz A.Blattmann andR.Rombach.2024.\u201cAdversarial Diffusion Distillation.\u201dInEuropean Conference on Computer Vision 87\u2013103.","DOI":"10.1007\/978-3-031-73016-0_6"},{"key":"e_1_2_9_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3072202"},{"key":"e_1_2_9_43_1","doi-asserted-by":"crossref","unstructured":"Sheng X. J.Li B.Li L.Li D.Liu andY.Lu.2022.\u201cTemporal Context Mining for Learned Video Compression.\u201dIEEE Transactions on Multimedia.","DOI":"10.1109\/TMM.2022.3220421"},{"key":"e_1_2_9_44_1","doi-asserted-by":"publisher","DOI":"10.1080\/19475705.2016.1238852"},{"key":"e_1_2_9_45_1","unstructured":"Song J. C.Meng andS.Ermon.2020.\u201cDenoising Diffusion Implicit Models.\u201darXiv Preprint arXiv:2010.02502."},{"key":"e_1_2_9_46_1","unstructured":"Song Y. P.Dhariwal M.Chen andI.Sutskever.2023.\u201cConsistency Models.\u201dInInternational Conference on Machine Learning."},{"key":"e_1_2_9_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10489-021-02893-3"},{"key":"e_1_2_9_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.coastaleng.2016.03.011"},{"key":"e_1_2_9_49_1","doi-asserted-by":"crossref","unstructured":"Veerabadran V. R.Pourreza A.Habibian andT. S.Cohen.2020.\u201cAdversarial Distortion for Learned Video Compression.\u201dInProceedings of the Ieee\/Cvf Conference on Computer Vision and Pattern Recognition Workshops 168\u2013169.","DOI":"10.1109\/CVPRW50498.2020.00092"},{"key":"e_1_2_9_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2022.3140306"},{"key":"e_1_2_9_51_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2003.815165"},{"key":"e_1_2_9_52_1","first-page":"1623","volume-title":"Proceedings of Machine Learning Research","author":"Wu J.","year":"2024"},{"key":"e_1_2_9_53_1","doi-asserted-by":"publisher","DOI":"10.1111\/exsy.70004"},{"key":"e_1_2_9_54_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-018-01144-2"},{"issue":"2","key":"e_1_2_9_55_1","first-page":"396","article-title":"A Review on Marine Search and Rescue Operations Using Unmanned Aerial Vehicles","volume":"9","author":"Yeong S.","year":"2015","journal-title":"International Journal of Marine and Environmental Sciences"},{"key":"e_1_2_9_56_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2022.109019"},{"key":"e_1_2_9_57_1","first-page":"47455","article-title":"Improved Distribution Matching Distillation for Fast Image Synthesis","volume":"37","author":"Yin T.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_9_58_1","doi-asserted-by":"publisher","DOI":"10.1111\/exsy.70070"},{"key":"e_1_2_9_59_1","doi-asserted-by":"crossref","unstructured":"Zhang R. P.Isola A. A.Efros E.Shechtman andO.Wang.2018.\u201cThe Unreasonable Effectiveness of Deep Features as a Perceptual Metric.\u201dInProceedings of the IEEE Conference on Computer Vision and Pattern Recognition 586\u2013595.","DOI":"10.1109\/CVPR.2018.00068"},{"issue":"1","key":"e_1_2_9_60_1","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3580499","article-title":"A Universal Optimization Framework for Learning\u2010Based Image Codec","volume":"20","author":"Zhao J.","year":"2023","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_2_9_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2021.3119563"},{"key":"e_1_2_9_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2017.2745538"}],"container-title":["Expert Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/exsy.70308","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/exsy.70308","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/exsy.70308","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,17]],"date-time":"2026-06-17T09:24:16Z","timestamp":1781688256000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/exsy.70308"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,3]]},"references-count":61,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2026,7]]}},"alternative-id":["10.1111\/exsy.70308"],"URL":"https:\/\/doi.org\/10.1111\/exsy.70308","archive":["Portico"],"relation":{},"ISSN":["0266-4720","1468-0394"],"issn-type":[{"value":"0266-4720","type":"print"},{"value":"1468-0394","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,3]]},"assertion":[{"value":"2026-03-17","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-05-15","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-03","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70308"}}