{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,11]],"date-time":"2025-09-11T22:32:55Z","timestamp":1757629975703,"version":"3.44.0"},"reference-count":39,"publisher":"Association for Computing Machinery (ACM)","issue":"9","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62111540272"],"award-info":[{"award-number":["62111540272"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,9,30]]},"abstract":"<jats:p>Since the quality of the reference frames is critical for VVC inter-coding, the neural network (NN)-based reference frame generation aims to generate a better quality reference frame from the two decoded frames in the decoded picture buffer (DPB) and inserts it into the reference picture list (RPL) for inter-prediction. However, it is difficult to directly apply video frame interpolation or extrapolation to the reference frame generation because compression artifacts severely degrade the quality of the generated reference frame. In this article, we propose a multiscale reference frame generation network for VVC inter-coding, named MRFGNet. To reduce the compression artifacts, MRFGNet adopts the high-performance operation point (HOP) network, which has been released by JVET, as preprocessing for frame enhancement. Unlike the HOP in-loop filter that takes multiple inputs, MRFGNet only takes the reconstructed frame and quantization parameter (QP)\u00a0map as input for the HOP network. Moreover, a frame generation network is proposed to conduct accurate motion estimation based on directional optical flow. Thus, in both RA and LDB configurations, MRFGNet has the same architecture to achieve both bidirectional and unidirectional predictions. MRFGNet estimates optical flow at multiple scales with the same dimension, thereby leveraging scale-independent bidirectional optical flow prediction. An optical flow warping and fusion module is designed to two cascaded U-nets for flow feature maps and get two frames, and finally, the two frames are fused to generate a reference frame. Furthermore, a novel training strategy based on QP distance is utilized to optimize MRFGNet by taking compressed data with higher quality as label for training. Experimental results show that MRFGNet achieves average BD rate changes of {\u22125.10% (Y), \u221212.14% (U), \u221211.51% (V)} and {\u22125.98% (Y), \u221215.32% (U), \u221214.22% (V)} over VTM_11.0-NNVC_4.0 anchor under RA and LDB configurations, respectively, as well as outperforms the state-of-the art method for reference frame generation, i.e., JVET-AD0160, by {0.78% (Y), 1.42% (U), 1.44% (V)} and {2.82% (Y), 4.92% (U), 6.62 %(V)} under RA and LDB configurations, respectively.<\/jats:p>","DOI":"10.1145\/3750049","type":"journal-article","created":{"date-parts":[[2025,7,23]],"date-time":"2025-07-23T16:19:08Z","timestamp":1753287548000},"page":"1-20","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["MRFGNet: Multiscale Reference Frame Generation Network for VVC Inter-Coding"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-5664-5449","authenticated-orcid":false,"given":"Pengyu","family":"Li","sequence":"first","affiliation":[{"name":"School of Electronic Engineering, Xidian University, Xi\u2019an,\u00a0China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0299-7206","authenticated-orcid":false,"given":"Cheolkon","family":"Jung","sequence":"additional","affiliation":[{"name":"School of Electronic Engineering, Xidian University, Xi\u2019an,\u00a0China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,9,10]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"E. Alshina and F. Galpin. 2023. AHG11: BoG report on NN-filter design unification. JVET-AD0380 Joint Video Exploration Team (JVET)."},{"key":"e_1_3_1_3_2","volume-title":"EE1-5.1: Deep Reference Frame Generation for Inter Prediction Enhancement","author":"Bao W.","year":"2023","unstructured":"W. Bao, W. Meng, J. Jia, Y. Zhang, H. Wang, Z. Chen, Z. Liu, X. Xu, and S. Liu. 2023. EE1-5.1: Deep Reference Frame Generation for Inter Prediction Enhancement. Technical Report JVET-AE0112. Wuhan University, Tencent."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3101953"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2021.3126593"},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3252911"},{"key":"e_1_3_1_7_2","unstructured":"F. Galpin S. Eadie Y. Li L. Wang Z. Xie D. Rusanovskyy Y. Li R. Chang J. Li and E. Alshina. 2023. AhG11: EE1-0 high operation point model. JVET-AE0191 Joint Video Exploration Team (JVET)."},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2024.3364536"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCAS45731.2020.9180452"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528173"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2020.2995243"},{"key":"e_1_3_1_12_2","first-page":"5754","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Hur Junhwa","year":"2019","unstructured":"Junhwa Hur and Stefan Roth. 2019. Iterative residual refinement for joint optical flow and occlusion estimation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 5754\u20135763."},{"key":"e_1_3_1_13_2","volume-title":"AHG11: Deep Reference Frame Generation for Inter Prediction Enhancement","author":"Jia J.","year":"2023","unstructured":"J. Jia, Y. Zhang, H. Zhu, Z. Chen, Z. Liu, X. Xu, and S. Liu. 2023. AHG11: Deep Reference Frame Generation for Inter Prediction Enhancement. Technical Report JVET-AC0114. Wuhan University, Tencent."},{"issue":"5","key":"e_1_3_1_14_2","first-page":"3111","article-title":"Deep reference frame generation method for VVC inter prediction enhancement","volume":"34","author":"Jia Jianghao","year":"2023","unstructured":"Jianghao Jia, Yuantong Zhang, Han Zhu, Zhenzhong Chen, Zizheng Liu, Xiaozhong Xu, and Shan Liu. 2023. Deep reference frame generation method for VVC inter prediction enhancement. IEEE Transactions on Circuits and Systems for Video Technology 34, 5 (2023), 3111\u20133124.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_15_2","volume-title":"EE1-2.1: Deep Reference Frame Generation for Inter Prediction Enhancement","author":"Jia J.","year":"2023","unstructured":"J. Jia, Y. Zhang, H. Zhu, Z. Chen, Z. Liu, X. Xu, and S. Liu. 2023. EE1-2.1: Deep Reference Frame Generation for Inter Prediction Enhancement. Technical Report JVET-AD0160. Wuhan University."},{"key":"e_1_3_1_16_2","unstructured":"Joint Video Experts Team (JVET). 2015. VTM-11.0-NNVC. Retrieved from https:\/\/vcgit.hhi.fraunhofer.de\/jvet-ahg-nnvc\/VVCSoftware\u02d9VTM\/-\/tree\/VTM-11.0\u02d9nnvc"},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00201"},{"key":"e_1_3_1_18_2","first-page":"85","volume-title":"Proceedings of the IEEE Winter Conference on Applications of Computer Vision","author":"LaBonte Tyler","year":"2023","unstructured":"Tyler LaBonte, Yale Song, Xin Wang, Vibhav Vineet, and Neel Joshi. 2023. Scaling novel object detection with weakly supervised detection transformers. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision, 85\u201396."},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.23919\/APSIPA.2018.8659611"},{"key":"e_1_3_1_20_2","unstructured":"J. Li K. Zhang L. Zhang and M. Wang. 2023. AHG11: Swin-transformer based in-loop filter for natural and screen contents. JVET-AC0179 Joint Video Exploration Team (JVET)."},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2022.3181116"},{"key":"e_1_3_1_22_2","doi-asserted-by":"crossref","unstructured":"Yue Li Li Zhang Kai Zhang Yuwen He and Jizheng Xu. 2020. AHG11: Convolutional neural networks-based in-loop filter. JVET-T0088.","DOI":"10.1109\/ICIP42928.2021.9506027"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/VCIP.2018.8698615"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2019.2961504"},{"key":"e_1_3_1_25_2","unstructured":"Shan Liu Andrew Segall Elena Alshina and Ru-Ling Liao. 2021. JVET common test conditions and evaluation procedures for neural network-based video coding technology. JVET 24th Meeting Document JVET-X2016."},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.37"},{"key":"e_1_3_1_27_2","first-page":"250","volume-title":"Proceedings of the European Conference on Computer Vision","author":"Reda Fitsum","year":"2022","unstructured":"Fitsum Reda, Janne Kontkanen, Eric Tabellion, Deqing Sun, Caroline Pantofaru, and Brian Curless. 2022. Film: Frame interpolation for large motion. In Proceedings of the European Conference on Computer Vision. Springer, 250\u2013266."},{"key":"e_1_3_1_28_2","unstructured":"Jay N. Shingala Ajay Shyam Ajat Suneja Siddarth P. Badya Tong Shao Arjun Arora Peng Yin Fangjun Pu Taoran Lu and Sean McCarthy. 2023. EE1-1.10: Complexity reduction on neural-network loop filter. JVET-AC0106."},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2012.2221191"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00931"},{"key":"e_1_3_1_31_2","unstructured":"Hongtao Wang Marta Karczewicz Jianle Chen and An Meher Kotra. 2020. AHG11: Neural network-based in-loop filter. JVET-T0079."},{"key":"e_1_3_1_32_2","doi-asserted-by":"crossref","unstructured":"Liqiang Wang Xiaozhong Xu Shan Liu and Franck Galpin. 2022. EE1-1.2: Neural network based in-loop filter with a single model. JVET-Z0091 Joint Video Exploration Team (JVET).","DOI":"10.1109\/ICME52920.2022.9859910"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2003.815165"},{"key":"e_1_3_1_34_2","unstructured":"R. Yang M. Santamaria N. Zou F. Cricri R. G. Youvalari J. Lainema H. Zhang and M. M. Hannuksela. 2023. AHG11: Neural network loop filter. JVET-AD0109 Joint Video Exploration Team (JVET)."},{"key":"e_1_3_1_35_2","first-page":"7008","volume-title":"Proceedings of the IEEE Conference on Computer Vision","author":"Yin Yufei","year":"2023","unstructured":"Yufei Yin, Jiajun Deng, Wengang Zhou, Li Li, and Houqiang Li. 2023. Cyclic-bootstrap labeling for weakly supervised object detection. In Proceedings of the IEEE Conference on Computer Vision, 7008\u20137018."},{"key":"e_1_3_1_36_2","unstructured":"Hao Zhang Cheolkon Jung Yang Liu and Ming Li. 2022. EE1-1.3 related: Lightweight and efficient CNN In-loop filter. JVET-AB0090 Joint Video Exploration Team (JVET)."},{"key":"e_1_3_1_37_2","unstructured":"Li Zhang Hongtao Wang Muhammed Coban Anand Meher Kotra Marta Karczewicz Franck Galpin D. Liu R. Sj\u00f6berg K. Andersson and J. Str\u00f6m. 2022. EE1-1.6: Deep in-loop filter with fixed point implementation. JVET-AA0111 Joint Video Exploration Team (JVET)."},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01655"},{"key":"e_1_3_1_39_2","unstructured":"Chuan Zhou Zhuoyi Lv and Jinrong Zhang. 2023. EE1-1.8: QP-based loss function design for NN-based in-loop filter. JVET-AC0118 Joint Video Exploration Team (JVET)."},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3282980"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3750049","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,9,10]],"date-time":"2025-09-10T16:01:41Z","timestamp":1757520101000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3750049"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,9,10]]},"references-count":39,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2025,9,30]]}},"alternative-id":["10.1145\/3750049"],"URL":"https:\/\/doi.org\/10.1145\/3750049","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"type":"print","value":"1551-6857"},{"type":"electronic","value":"1551-6865"}],"subject":[],"published":{"date-parts":[[2025,9,10]]},"assertion":[{"value":"2024-09-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-07-13","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-09-10","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}