{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T02:31:22Z","timestamp":1782268282813,"version":"3.54.5"},"reference-count":66,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,12,23]],"date-time":"2024-12-23T00:00:00Z","timestamp":1734912000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"Institute of Information and Communications Technology Planning and Evaluation","award":["RS-2022-00167169"],"award-info":[{"award-number":["RS-2022-00167169"]}]},{"name":"NRF","award":["NRF-2022R1A2C4002052"],"award-info":[{"award-number":["NRF-2022R1A2C4002052"]}]},{"name":"Ministry of Culture, Sports and Tourism in 2024","award":["RS-2024-00439534"],"award-info":[{"award-number":["RS-2024-00439534"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2025,1,31]]},"abstract":"<jats:p>In this article, we propose an efficient reference-based deep in-loop filtering method for video coding. Existing reference-based in-loop filters often face challenges in improving coding efficiency due to the difficulty in capturing relevant textures from the reference frames. Our method accurately predicts the texture of a reference block and uses this information to restore the current block. To achieve this, we develop a reference-to-current feature estimation module that conveys high-quality information from previously coded frames in the feature domain, thereby preventing loss of detail due to inaccurate prediction. Although a neural network is trained to restore a coded video frame to be similar to the current frame, their performance can significantly degrade when operating with various quantization parameters (QPs) and managing different levels of distortion. This problem becomes further severe in the reference-to-current feature estimation, in which QP values are applied differently to video frames. We address this problem by developing a QP-aware convolution layer with a small number of learnable parameters to generate reliable features and adapt to fine-grained adaptive QPs among consecutive frames. The proposed method is implemented into the versatile video coding (VVC) reference software, VTM version 10.0. Experimental results demonstrate that the proposed method improves coding performance significantly in VVC.<\/jats:p>","DOI":"10.1145\/3702643","type":"journal-article","created":{"date-parts":[[2024,11,1]],"date-time":"2024-11-01T09:08:25Z","timestamp":1730452105000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["Reference-based In-loop Filter with Robust Neural Feature Transfer for Video Coding"],"prefix":"10.1145","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7235-3209","authenticated-orcid":false,"given":"Nayoung","family":"Kim","sequence":"first","affiliation":[{"name":"NAVER WEBTOON AI, Seoul, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0227-3419","authenticated-orcid":false,"given":"Jung-Kyung","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Electronic and Electrical Engineering, Ewha W. University, Seoul, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1637-9479","authenticated-orcid":false,"given":"Je-Won","family":"Kang","sequence":"additional","affiliation":[{"name":"Department of Electronic and Electrical Engineering and Graduate Program in Smart Factory, Ewha W. University, Seoul, Korea"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,12,23]]},"reference":[{"key":"e_1_3_1_2_2","unstructured":"2019. Convolutional neural network loop filter. Recommendation ITU-T SG 16 WP3 and ISO\/IEC JTC 1\/SC 29\/WG 11."},{"key":"e_1_3_1_3_2","volume-title":"Proceedings of the JVET Meeting","author":"Boyce Jill M.","year":"2018","unstructured":"Jill M. Boyce, Karsten Suehring, Xiang Li, and Vadim Seregin. 2018. JVET common test conditions and software reference configurations for SDR video. In Proceedings of the JVET Meeting."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00382"},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/PCS50896.2021.9477457"},{"key":"e_1_3_1_6_2","unstructured":"Wei Jiang Cheung Auyeung Wei Wang. 2020. AhG: A case study to reduce computation of neural network based in-loop filter by pruning (JVET-T0057). ISO\/IEC\/JTC1\/SC29\/WG11."},{"key":"e_1_3_1_7_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-51811-4_3"},{"key":"e_1_3_1_8_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i07.6697"},{"key":"e_1_3_1_9_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3134465"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2023.3260266"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3092949"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.73"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-10593-2_13"},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2012.2221529"},{"key":"e_1_3_1_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2944806"},{"key":"e_1_3_1_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_17_2","unstructured":"Y.-L. Hsiao O. Chubach C. Chen T. Chuang C. Hsu Y. Huang and S. Lei. 2019. Convolutional neural network loop filter (JVET-O0056). ISO\/IEC\/JTC1\/SC29\/WG11."},{"key":"e_1_3_1_18_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3084345"},{"key":"e_1_3_1_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3089498"},{"key":"e_1_3_1_20_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-11021-5_20"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3329949"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/VCIP.2017.8305149"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3460820"},{"key":"e_1_3_1_24_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4842-2766-4_12"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2020.3035281"},{"key":"e_1_3_1_26_2","unstructured":"Yejin Kim Manri Cheon and Junwoo Lee. 2020. Texture transform attention for realistic image inpainting. arXiv:2012.04242. Retrieved from https:\/\/arxiv.org\/abs\/2012.04242"},{"key":"e_1_3_1_27_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv:1412.6980. Retrieved from https:\/\/arxiv.org\/pdf\/1412.6980"},{"key":"e_1_3_1_28_2","doi-asserted-by":"publisher","DOI":"10.23919\/APSIPA.2018.8659611"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2020.2993566"},{"key":"e_1_3_1_30_2","doi-asserted-by":"publisher","DOI":"10.5573\/IEIESPC.2023.12.2.122"},{"key":"e_1_3_1_31_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00191"},{"key":"e_1_3_1_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/DCC.2019.00035"},{"key":"e_1_3_1_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2019.2921877"},{"key":"e_1_3_1_34_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58607-2_20"},{"key":"e_1_3_1_35_2","unstructured":"Y. Li L. Zhang and K. Zhang. 2021. AHG11: Deep in-loop filter with adaptive model selection (JVET-V0100). ISO\/IEC\/JTC1\/SC29\/WG11."},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3578585"},{"issue":"4","key":"e_1_3_1_37_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3502723","article-title":"NR-CNN: Nested-residual guided CNN in-loop filtering for video coding","volume":"18","author":"Lin Kai","year":"2022","unstructured":"Kai Lin, Chuanmin Jia, Xinfeng Zhang, Shanshe Wang, Siwei Ma, and Wen Gao. 2022. NR-CNN: Nested-residual guided CNN in-loop filtering for video coding. ACM Transactions on Multimedia Computing, Communications, and Applications 18, 4 (2022), 1\u201322.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2022.3142414"},{"key":"e_1_3_1_39_2","doi-asserted-by":"publisher","DOI":"10.1109\/DCC47342.2020.00027"},{"key":"e_1_3_1_40_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2020.3043064"},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/VCIP49819.2020.9301884"},{"key":"e_1_3_1_42_2","doi-asserted-by":"crossref","unstructured":"Fatemeh Nasiri Wassim Hamidouche Luce Morin Nicolas Dhollande and Gildas Cocherel. 2021. A CNN-based prediction-aware quality enhancement framework for VVC. arXiv:2105.05658. Retrieved from https:\/\/arxiv.org\/abs\/2105.05658","DOI":"10.1109\/OJSP.2021.3092598"},{"key":"e_1_3_1_43_2","doi-asserted-by":"crossref","unstructured":"Fatemeh Nasiri Wassim Hamidouche Luce Morin Nicolas Dhollande and Gildas Cocherel. 2021. Model selection CNN-based VVC quality enhancement. arXiv:2105.03338. Retrieved from https:\/\/arxiv.org\/pdf\/2105.03338","DOI":"10.1109\/PCS50896.2021.9477473"},{"key":"e_1_3_1_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2012.2223053"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.2982534"},{"key":"e_1_3_1_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00206"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.291"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2018.8451589"},{"key":"e_1_3_1_49_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00931"},{"key":"e_1_3_1_50_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2013.2271974"},{"key":"e_1_3_1_51_2","unstructured":"Versatile Video Coding Test Model (VTM) 10.0 [Online]. 2020. Retrieved from https:\/\/vcgit.hhi.fraunhofer.de\/jvet\/VVCSoftware_VTM\/-\/releases\/VTM-10.0"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2021.3068638"},{"key":"e_1_3_1_53_2","unstructured":"H. Wang J. Chen A. Kotra K. Reuze and M. Karczewicz. 2021. Neural network-based in-loop filter with no deblocking filtering stage (JVET-V0114). ISO\/IEC\/JTC1\/SC29\/WG11."},{"key":"e_1_3_1_54_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2019.2944473"},{"key":"e_1_3_1_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2019.00247"},{"key":"e_1_3_1_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/MMSP.2019.8901772"},{"key":"e_1_3_1_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/VCIP.2018.8698740"},{"key":"e_1_3_1_58_2","doi-asserted-by":"publisher","DOI":"10.1109\/DCC50243.2021.00010"},{"key":"e_1_3_1_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00583"},{"key":"e_1_3_1_60_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICME.2019.00098"},{"key":"e_1_3_1_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00697"},{"key":"e_1_3_1_62_2","unstructured":"Junru Li Li Zhang Yue Li and Kai Zhang. 2022. EE1-1.6: Deep in-loop filter with fixed point implementation (JVET-AA0111). ISO\/IEC\/JTC1\/SC29\/EE1."},{"issue":"7","key":"e_1_3_1_63_2","first-page":"1888","article-title":"Recursive residual convolutional neural network-based in-loop filtering for intra frames","volume":"30","author":"Zhang Shufang","year":"2019","unstructured":"Shufang Zhang, Zenghui Fan, Nam Ling, and Minqiang Jiang. 2019. Recursive residual convolutional neural network-based in-loop filtering for intra frames. IEEE Transactions on Circuits and Systems for Video Technology 30, 7 (2019), 1888\u20131900.","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"e_1_3_1_64_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_18"},{"key":"e_1_3_1_65_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2815841"},{"key":"e_1_3_1_66_2","unstructured":"Yulun Zhang Yapeng Tian Yu Kong Bineng Zhong and Yun Fu. 2020. Residual dense network for image restoration. arXiv:1812.10477. Retrieved from https:\/\/arxiv.org\/abs\/1812.10477"},{"key":"e_1_3_1_67_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_6"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3702643","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3702643","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:03Z","timestamp":1750295883000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3702643"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,23]]},"references-count":66,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2025,1,31]]}},"alternative-id":["10.1145\/3702643"],"URL":"https:\/\/doi.org\/10.1145\/3702643","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,23]]},"assertion":[{"value":"2023-12-02","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-10-24","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-12-23","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}