{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T16:22:30Z","timestamp":1775665350226,"version":"3.50.1"},"reference-count":44,"publisher":"MDPI AG","issue":"5","license":[{"start":{"date-parts":[[2023,2,27]],"date-time":"2023-02-27T00:00:00Z","timestamp":1677456000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Institute for Information &amp; communications Technology Planning &amp; Evaluation (IITP)","award":["2017-0-00072"],"award-info":[{"award-number":["2017-0-00072"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>As the demands of various network-dependent services such as Internet of things (IoT) applications, autonomous driving, and augmented and virtual reality (AR\/VR) increase, the fifthgeneration (5G) network is expected to become a key communication technology. The latest video coding standard, versatile video coding (VVC), can contribute to providing high-quality services by achieving superior compression performance. In video coding, inter bi-prediction serves to improve the coding efficiency significantly by producing a precise fused prediction block. Although block-wise methods, such as bi-prediction with CU-level weight (BCW), are applied in VVC, it is still difficult for the linear fusion-based strategy to represent diverse pixel variations inside a block. In addition, a pixel-wise method called bi-directional optical flow (BDOF) has been proposed to refine bi-prediction block. However, the non-linear optical flow equation in BDOF mode is applied under assumptions, so this method is still unable to accurately compensate various kinds of bi-prediction blocks. In this paper, we propose an attention-based bi-prediction network (ABPN) to substitute for the whole existing bi-prediction methods. The proposed ABPN is designed to learn efficient representations of the fused features by utilizing an attention mechanism. Furthermore, the knowledge distillation (KD)- based approach is employed to compress the size of the proposed network while keeping comparable output as the large model. The proposed ABPN is integrated into the VTM-11.0 NNVC-1.0 standard reference software. When compared with VTM anchor, it is verified that the BD-rate reduction of the lightweighted ABPN can be up to 5.89% and 4.91% on Y component under random access (RA) and low delay B (LDB), respectively.<\/jats:p>","DOI":"10.3390\/s23052631","type":"journal-article","created":{"date-parts":[[2023,2,28]],"date-time":"2023-02-28T02:01:51Z","timestamp":1677549711000},"page":"2631","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":7,"title":["Attention-Based Bi-Prediction Network for Versatile Video Coding (VVC) over 5G Network"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-7982-5192","authenticated-orcid":false,"given":"Young-Ju","family":"Choi","sequence":"first","affiliation":[{"name":"Department of IT Engineering, Sookmyung Women\u2019s University, Seoul 04310, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9011-0921","authenticated-orcid":false,"given":"Young-Woon","family":"Lee","sequence":"additional","affiliation":[{"name":"Department of Computer Engineering, Sunmoon University, Asan 31460, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1256-0649","authenticated-orcid":false,"given":"Jongho","family":"Kim","sequence":"additional","affiliation":[{"name":"Media Coding Research Section, Electronics and Telecommunications Research Institute, Daejeon 34129, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1675-4814","authenticated-orcid":false,"given":"Se Yoon","family":"Jeong","sequence":"additional","affiliation":[{"name":"Media Coding Research Section, Electronics and Telecommunications Research Institute, Daejeon 34129, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4297-5327","authenticated-orcid":false,"given":"Jin Soo","family":"Choi","sequence":"additional","affiliation":[{"name":"Media Coding Research Section, Electronics and Telecommunications Research Institute, Daejeon 34129, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6555-3464","authenticated-orcid":false,"given":"Byung-Gyu","family":"Kim","sequence":"additional","affiliation":[{"name":"Department of IT Engineering, Sookmyung Women\u2019s University, Seoul 04310, Republic of Korea"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,2,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Ding, A.Y., and Janssen, M. (2018, January 8\u20139). Opportunities for applications using 5G networks: Requirements, challenges, and outlook. Proceedings of the Seventh International Conference on Telecommunications and Remote Sensing, New York, NY, USA.","DOI":"10.1145\/3278161.3278166"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"105895","DOI":"10.1016\/j.compag.2020.105895","article-title":"A survey on the 5G network and its impact on agriculture: Challenges and opportunities","volume":"180","author":"Tang","year":"2021","journal-title":"Comput. Electron. Agric."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"163","DOI":"10.1007\/s00779-017-1058-5","article-title":"Context-aware block-based motion estimation algorithm for multimedia internet of things (IoT) platform","volume":"22","author":"Saha","year":"2018","journal-title":"Pers. Ubiquitous Comput."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"1063","DOI":"10.1007\/s11227-016-1730-y","article-title":"Fast coding unit (CU) determination algorithm for high-efficiency video coding (HEVC) in smart surveillance application","volume":"73","author":"Kim","year":"2017","journal-title":"J. Supercomput."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1089\/big.2020.0274","article-title":"Residual-based graph convolutional network for emotion recognition in conversation for smart Internet of Things","volume":"9","author":"Choi","year":"2021","journal-title":"Big Data"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2","DOI":"10.23919\/ETR.2020.9905504","article-title":"Versatile Video Coding explained\u2014The Future of Video in a 5G World","volume":"2020","author":"Litwic","year":"2020","journal-title":"Ericsson Technol. Rev."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Tang, G., Hu, Y., Xiao, H., Zheng, L., She, X., and Qin, N. (2021, January 9\u201313). Design of Real-time video transmission system based on 5G network. Proceedings of the IEEE Conference on Industrial Electronics and Applications (ICIEA), Kristiansand, Norway.","DOI":"10.1109\/ICIEA51954.2021.9516191"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"560","DOI":"10.1109\/TCSVT.2003.815165","article-title":"Overview of the H. 264\/AVC video coding standard","volume":"13","author":"Wiegand","year":"2003","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"1649","DOI":"10.1109\/TCSVT.2012.2221191","article-title":"Overview of the high efficiency video coding (HEVC) standard","volume":"22","author":"Sullivan","year":"2012","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"3736","DOI":"10.1109\/TCSVT.2021.3101953","article-title":"Overview of the versatile video coding (VVC) standard and its applications","volume":"31","author":"Bross","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"3848","DOI":"10.1109\/TCSVT.2021.3101212","article-title":"Motion Vector Coding and Block Merging in the Versatile Video Coding Standard","volume":"31","author":"Chien","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"3862","DOI":"10.1109\/TCSVT.2021.3100744","article-title":"Subblock-Based Motion Derivation and Inter Prediction Refinement in the Versatile Video Coding Standard","volume":"31","author":"Yang","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_13","unstructured":"Liu, S., Segall, A., Alshina, E., and Liao, R.L. (2021, January 6\u201315). JVET common test conditions and evaluation procedures for neural network-based video coding technology. Proceedings of the Document JVET-X2016, 24th JVET Meeting, Vitual."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Dong, C., Loy, C.C., He, K., and Tang, X. (2014, January 6\u201312). Learning a deep convolutional network for image super-resolution. Proceedings of the European Conference on Computer Vision (ECCV), Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10593-2_13"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chen, Y.S., Wang, Y.C., Kao, M.H., and Chuang, Y.Y. (2018, January 18\u201323). Deep photo enhancer: Unpaired learning for image enhancement from photographs with gans. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00660"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"4474","DOI":"10.1109\/TIP.2020.2972118","article-title":"Learning a deep dual attention network for video super-resolution","volume":"29","author":"Li","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Choi, Y.J., Lee, Y.W., and Kim, B.G. (2021, January 10\u201315). Wavelet Attention Embedding Networks for Video Super-Resolution. Proceedings of the IEEE International Conference on Pattern Recognition (ICPR), Milan, Italy.","DOI":"10.1109\/ICPR48806.2021.9412623"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"3343","DOI":"10.1109\/TIP.2019.2896489","article-title":"Content-aware convolutional neural network for in-loop filtering in high efficiency video coding","volume":"28","author":"Jia","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1871","DOI":"10.1109\/TCSVT.2019.2935508","article-title":"A switchable deep learning approach for in-loop filtering in video coding","volume":"30","author":"Ding","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"145214","DOI":"10.1109\/ACCESS.2019.2944473","article-title":"Attention-based dual-scale CNN in-loop filter for Versatile Video Coding","volume":"7","author":"Wang","year":"2019","journal-title":"IEEE Access"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"2039","DOI":"10.1109\/TCSVT.2018.2867568","article-title":"Enhancing quality for HEVC compressed videos","volume":"29","author":"Yang","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_22","unstructured":"Guan, Z., Xing, Q., Xu, M., Yang, R., Liu, T., and Wang, Z. (2019). MFQE 2.0: A new approach for multi-frame quality enhancement on compressed video. arXiv."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"158","DOI":"10.1016\/j.neucom.2018.12.090","article-title":"Post-processing for intra coding through perceptual adversarial learning and progressive refinement","volume":"394","author":"Jin","year":"2020","journal-title":"Neurocomputing"},{"key":"ref_24","first-page":"1816","article-title":"CNN-based intra-prediction for lossless HEVC","volume":"30","author":"Schiopu","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1109\/TMM.2019.2924591","article-title":"Generative adversarial network-based intra prediction for video coding","volume":"22","author":"Zhu","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"840","DOI":"10.1109\/TCSVT.2018.2816932","article-title":"Convolutional neural network-based fractional-pixel motion compensation","volume":"29","author":"Yan","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"3923","DOI":"10.1109\/TCSVT.2021.3107135","article-title":"Deep Affine Motion Compensation Network for Inter Prediction in VVC","volume":"32","author":"Jin","year":"2021","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"3291","DOI":"10.1109\/TCSVT.2018.2876399","article-title":"Enhanced bi-prediction with convolutional neural network for high-efficiency video coding","volume":"29","author":"Zhao","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_29","first-page":"1856","article-title":"Convolutional neural network based bi-prediction utilizing spatial and temporal information in video coding","volume":"30","author":"Mao","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lim, B., Son, S.H., Kim, H.W., Nah, S.J., and Lee, K.M. (2017, January 21\u201326). Enhanced deep residual networks for single image super-resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Honolulu, HI, USA.","DOI":"10.1109\/CVPRW.2017.151"},{"key":"ref_31","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the knowledge in a neural network. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Huo, S., Liu, D., Wu, F., and Li, H. (2018, January 27\u201330). Convolutional neural network-based motion compensation refinement for video coding. Proceedings of the IEEE International Symposium on Circuits and Systems (ISCAS), Florence, Italy.","DOI":"10.1109\/ISCAS.2018.8351609"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Wang, Y., Fan, X., Jia, C., Zhao, D., and Gao, W. (2018, January 23\u201327). Neural network based inter prediction for HEVC. Proceedings of the IEEE International Conference on Multimedia and Expo (ICME), San Diego, CA, USA.","DOI":"10.1109\/ICME.2018.8486600"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"4832","DOI":"10.1109\/TIP.2019.2913545","article-title":"Enhanced motion-compensated video coding with deep virtual reference frame generation","volume":"28","author":"Zhao","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"2497","DOI":"10.1109\/TMM.2019.2961504","article-title":"Deep reference generation with multi-domain hierarchical constraints for inter prediction","volume":"22","author":"Liu","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"1178","DOI":"10.1109\/TCSVT.2020.2995243","article-title":"Deep network-based frame extrapolation with reference frame alignment","volume":"31","author":"Huo","year":"2020","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"3321","DOI":"10.1109\/TIP.2021.3060803","article-title":"Affine transformation-based deep frame prediction","volume":"30","author":"Choi","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_38","unstructured":"Alshin, A., and Alshina, E. (April, January 29). Bi-directional pptical flow for future video codec. Proceedings of the Data Compression Conference (DCC), Snowbird, UT, USA."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Lai, W.S., Huang, J.B., Ahuja, N., and Yang, M.H. (2017, January 21\u201326). Deep laplacian pyramid networks for fast and accurate super-resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.618"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Ma, D., Zhang, F., and Bull, D. (2021). BVI-DVC: A training database for deep video compression. arXiv.","DOI":"10.1109\/TMM.2021.3108943"},{"key":"ref_41","unstructured":"Kingma, D., and Ba, J. (May, January 30). Adam: A method for stochastic optimization. Proceedings of the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada."},{"key":"ref_42","unstructured":"Loshchilov, I., and Hutter, F. (2016). Sgdr: Stochastic gradient descent with warm restarts. arXiv."},{"key":"ref_43","unstructured":"Bj\u00f8ntegaard, G. (2001, January 2\u20134). Calculation of average PSNR differences between RD-curves. Proceedings of the VCEG-M33, Austin, TX, USA."},{"key":"ref_44","unstructured":"Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J. (2016). Pruning convolutional neural networks for resource efficient inference. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/5\/2631\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:44:09Z","timestamp":1760121849000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/5\/2631"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,2,27]]},"references-count":44,"journal-issue":{"issue":"5","published-online":{"date-parts":[[2023,3]]}},"alternative-id":["s23052631"],"URL":"https:\/\/doi.org\/10.3390\/s23052631","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,2,27]]}}}