{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:41:29Z","timestamp":1760150489484,"version":"build-2065373602"},"reference-count":42,"publisher":"MDPI AG","issue":"23","license":[{"start":{"date-parts":[[2023,11,27]],"date-time":"2023-11-27T00:00:00Z","timestamp":1701043200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100010418","name":"Institute of Information &amp; Communications Technology Planning &amp; Evaluation (IITP) grant funded by the Korea government (MSIT)","doi-asserted-by":"publisher","award":["2017-0-00072"],"award-info":[{"award-number":["2017-0-00072"]}],"id":[{"id":"10.13039\/501100010418","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>After the development of the Versatile Video Coding (VVC) standard, research on neural network-based video coding technologies continues as a potential approach for future video coding standards. Particularly, neural network-based intra prediction is receiving attention as a solution to mitigate the limitations of traditional intra prediction performance in intricate images with limited spatial redundancy. This study presents an intra prediction method based on coarse-to-fine networks that employ both convolutional neural networks and fully connected layers to enhance VVC intra prediction performance. The coarse networks are designed to adjust the influence on prediction performance depending on the positions and conditions of reference samples. Moreover, the fine networks generate refined prediction samples by considering continuity with adjacent reference samples and facilitate prediction through upscaling at a block size unsupported by the coarse networks. The proposed networks are integrated into the VVC test model (VTM) as an additional intra prediction mode to evaluate the coding performance. The experimental results show that our coarse-to-fine network architecture provides an average gain of 1.31% Bj\u00f8ntegaard delta-rate (BD-rate) saving for the luma component compared with VTM 11.0 and an average of 0.47% BD-rate saving compared with the previous related work.<\/jats:p>","DOI":"10.3390\/s23239452","type":"journal-article","created":{"date-parts":[[2023,11,28]],"date-time":"2023-11-28T07:41:32Z","timestamp":1701157292000},"page":"9452","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Coarse-to-Fine Network-Based Intra Prediction in Versatile Video Coding"],"prefix":"10.3390","volume":"23","author":[{"given":"Dohyeon","family":"Park","sequence":"first","affiliation":[{"name":"Department of Electronics and Information Engineering, Korea Aerospace University, Goyang 10540, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Gihwa","family":"Moon","sequence":"additional","affiliation":[{"name":"Department of Electronics and Information Engineering, Korea Aerospace University, Goyang 10540, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1437-2422","authenticated-orcid":false,"given":"Byung Tae","family":"Oh","sequence":"additional","affiliation":[{"name":"Department of Electronics and Information Engineering, Korea Aerospace University, Goyang 10540, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3686-4786","authenticated-orcid":false,"given":"Jae-Gon","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of Electronics and Information Engineering, Korea Aerospace University, Goyang 10540, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,11,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"171","DOI":"10.1109\/TMM.2017.2713642","article-title":"264 and H.265 video bandwidth prediction","volume":"20","author":"Kalampogia","year":"2018","journal-title":"IEEE Trans. Multimed."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"1649","DOI":"10.1109\/TCSVT.2012.2221191","article-title":"Overview of the high efficiency video coding (HEVC) standard","volume":"22","author":"Sullivan","year":"2012","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"3736","DOI":"10.1109\/TCSVT.2021.3101953","article-title":"Overview of the Versatile Video Coding (VVC) standard and its applications","volume":"31","author":"Bross","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"3818","DOI":"10.1109\/TCSVT.2021.3088134","article-title":"Block partitioning structure in the VVC standard","volume":"31","author":"Huang","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"3603","DOI":"10.1109\/TCSVT.2020.3040291","article-title":"Geometric partitioning mode in Versatile Video Coding: Algorithm review and analysis","volume":"31","author":"Gao","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"3834","DOI":"10.1109\/TCSVT.2021.3072430","article-title":"Intra prediction and mode coding in VVC","volume":"31","author":"Pfaff","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3848","DOI":"10.1109\/TCSVT.2021.3101212","article-title":"Motion vector coding and block merging in the Versatile Video Coding standard","volume":"31","author":"Chien","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"3878","DOI":"10.1109\/TCSVT.2021.3087706","article-title":"Transform coding in the VVC standard","volume":"31","author":"Zhao","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"3891","DOI":"10.1109\/TCSVT.2021.3072202","article-title":"Quantization and entropy coding in the Versatile Video Coding (VVC) standard","volume":"31","author":"Schwarz","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"3907","DOI":"10.1109\/TCSVT.2021.3072297","article-title":"VVC in-loop filters","volume":"31","author":"Karczewicz","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_11","unstructured":"Pfaff, J., Stallenberger, B., Schafer, M., Merkle, P., Helle, P., Hinz, T., Schwarz, H., Marpe, D., and Wiegand, T. (, January March). Affine linear weighted intra prediction. Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO\/IEC JTC 1\/SC 29, doc. JVET-N0217. Proceedings of the 14th Meeting, Geneva, Switzerland."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Koo, M., Salehifar, M., Lim, J., and Kim, S. (2019, January 12\u201315). Low Frequency Non-Separable Transform (LFNST). Proceedings of the 2019 Picture Coding Symposium (PCS), Ningbo, China.","DOI":"10.1109\/PCS48520.2019.8954507"},{"key":"ref_13","unstructured":"Alshina, E., Galpin, F., Li, Y., Santamaria, M., Wang, H., Wang, L., and Xie, Z. (, January October). EE1: Summary of exploration experiments on neural network-based video coding. Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO\/IEC JTC 1\/SC 29, doc. JVET-AB0023. Proceedings of the 28th Meeting, Mainz, Germany."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Li, J., Li, Y., Lin, C., Zhang, K., and Zhang, L. (2022, January 18\u201324). A neural network enhanced video coding framework beyond VVC. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPRW56347.2022.00191"},{"key":"ref_15","unstructured":"Pfaff, J., Helle, P., Merkle, P., Sch\u00e4fer, M., Stallenberger, B., Hinz, T., Schwarz, H., Marpe, D., and Wiegand, T. (2020). Data-driven intra-prediction modes in the development of the Versatile Video Coding standard. ICT Discov., 3."},{"key":"ref_16","unstructured":"Chen, J., Ye, Y., and Kim, S. (, January October). Algorithm description for Versatile Video Coding and Test Model 11 (VTM 11). Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO\/IEC JTC 1\/SC 29, doc. JVET-T2002. Proceedings of the 20th Meeting, by Teleconference (Online)."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"2764","DOI":"10.1109\/TMM.2019.2963620","article-title":"Enhanced intra prediction for video coding by using multiple neural networks","volume":"22","author":"Sun","year":"2020","journal-title":"IEEE Trans. Multimed."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"3024","DOI":"10.1109\/TMM.2019.2920603","article-title":"Progressive spatial recurrent neural network for intra prediction","volume":"21","author":"Hu","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"679","DOI":"10.1109\/TIP.2019.2934565","article-title":"Context-adaptive neural network-based prediction for image compression","volume":"29","author":"Dumas","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"354","DOI":"10.1109\/JSTSP.2020.3034768","article-title":"Intra-frame coding using a conditional autoencoder","volume":"15","author":"Brand","year":"2021","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Dumas, T., Galpin, F., and Bordes, P. (July, January 29). Combined Neural Network-based Intra Prediction and Transform Selection. Proceedings of the 2021 Picture Coding Symposium (PCS), Bristol, UK.","DOI":"10.1109\/PCS50896.2021.9477455"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"366","DOI":"10.1109\/JSTSP.2020.3044482","article-title":"Attention-based neural networks for chroma intra prediction in video coding","volume":"15","author":"Blanch","year":"2021","journal-title":"IEEE J. Sel. Top. Signal Process."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"111052","DOI":"10.1109\/ACCESS.2022.3215163","article-title":"Machine Learning-Based Early Skip Decision for Intra Prediction in VVC","volume":"10","author":"Park","year":"2022","journal-title":"IEEE Access"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Yu, J., Lin, Z., Yang, J., Shen, X., Lu, X., and Huang, T. (2018, January 18\u201322). Generative image inpainting with contextual attention. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00577"},{"key":"ref_25","unstructured":"Yu, J., Lin, Z., Yang, J., Shen, X., Lu, X., and Huang, T. (November, January 27). Free-form image inpainting with gated convolution. Proceedings of the International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Jin, X., Chen, Z., Liu, S., and Zhou, W. (2018, January 23\u201326). Augmented Coarse-to-Fine Video Frame Synthesis with Semantic Loss. Proceedings of the PRCV 2018: Pattern Recognition and Computer Vision, Guangzhou, China.","DOI":"10.1007\/978-3-030-03398-9_38"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"543","DOI":"10.1109\/LSP.2022.3147441","article-title":"Coarse-to-Fine Spatio-Temporal Information Fusion for Compressed Video Quality Enhancement","volume":"29","author":"Luo","year":"2022","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_28","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201322). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"303","DOI":"10.1007\/BF02551274","article-title":"Approximation by superpositions of a sigmoidal function","volume":"2","author":"Cybenke","year":"1989","journal-title":"Math. Control Signals Syst."},{"key":"ref_31","unstructured":"Arora, R., Basu, A., Mianjy, P., and Mukherjee, A. (2016). Understanding deep neural networks with rectified linear units. arXiv."},{"key":"ref_32","unstructured":"Maas, A., Hannun, A., and Ng, A.Y. (2013, January 16\u201321). Rectifier nonlinearities improve neural network acoustic models. Proceedings of the 30th International Conference on Machine Learning (ICML), Atlanta, GA, USA."},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"3847","DOI":"10.1109\/TMM.2021.3108943","article-title":"BVI-DVC: A Training database for deep video compression","volume":"24","author":"Ma","year":"2021","journal-title":"IEEE Trans. Multimed."},{"key":"ref_34","unstructured":"Lu, X., Liu, S., and Li, Z. (2021). Tencent Video Dataset (TVD): A video dataset for learning-based visual data compression and analysis. arXiv."},{"key":"ref_35","unstructured":"(2020, December 12). VVC Test Model. Available online: https:\/\/vcgit.hhi.fraunhofer.de\/jvet\/VVCSoftware_VTM\/-\/tree\/VTM-11.0."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Paul, M., Antony, A., and Sreelekha, G. (2014, January 24\u201327). Performance improvement of HEVC using adaptive quantization. Proceedings of the 2014 International Conference on Advances in Computing, Communications and Informatics (ICACCI), Delhi, India.","DOI":"10.1109\/ICACCI.2014.6968266"},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"1661","DOI":"10.1109\/83.869177","article-title":"A mathematical analysis of the DCT coefficient distributions for images","volume":"9","author":"Lam","year":"2000","journal-title":"IEEE Trans. Image Process."},{"key":"ref_38","unstructured":"Kingma, D., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Abdoli, M., Guionnet, T., Raulet, M., Kulupana, F., and Blasi, S. (2020, January 6\u201310). Decoder-side intra mode derivation for next generation video coding. Proceedings of the 2020 IEEE International Conference on Multimedia and Expo (ICME), London, UK.","DOI":"10.1109\/ICME46284.2020.9102799"},{"key":"ref_40","unstructured":"Li, Y., Wang, H., Wang, L., Galpin, F., and Str\u00f6m, J. (, January October). Algorithm description for neural network-based video coding 3 (NNVC 3). Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO\/IEC JTC 1\/SC 29, doc. JVET-AB2019. Proceedings of the 28th Meeting, Mainz, Germany."},{"key":"ref_41","unstructured":"Bossen, F., Boyce, J., Suehring, K., Li, X., and Seregin, V. (, January October). VTM common test conditions and software reference configurations for SDR video. Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO\/IEC JTC 1\/SC 29, doc. JVET-T2010. Proceedings of the 20th Meeting, by Teleconference (Online)."},{"key":"ref_42","unstructured":"Bj\u00d8ntegaard, G. (, January April). Calculation of average PSNR differences between RD curves. Video Coding Experts Group (VCEG) of ITU-T SG 16 WP 3, doc. VCEG-M33. Proceedings of the 13th Meeting, Austin, TX, USA."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/23\/9452\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:31:52Z","timestamp":1760131912000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/23\/9452"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,27]]},"references-count":42,"journal-issue":{"issue":"23","published-online":{"date-parts":[[2023,12]]}},"alternative-id":["s23239452"],"URL":"https:\/\/doi.org\/10.3390\/s23239452","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2023,11,27]]}}}