{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:49:19Z","timestamp":1760147359057,"version":"build-2065373602"},"reference-count":45,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2023,1,29]],"date-time":"2023-01-29T00:00:00Z","timestamp":1674950400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key R&amp;D Program of China","doi-asserted-by":"publisher","award":["2019YFB1405902"],"award-info":[{"award-number":["2019YFB1405902"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Ultra-high-definition (UHD) video has brought new challenges to objective video quality assessment (VQA) due to its high resolution and high frame rate. Most existing VQA methods are designed for non-UHD videos\u2014when they are employed to deal with UHD videos, the processing speed will be slow and the global spatial features cannot be fully extracted. In addition, these VQA methods usually segment the video into multiple segments, predict the quality score of each segment, and then average the quality score of each segment to obtain the quality score of the whole video. This breaks the temporal correlation of the video sequences and is inconsistent with the characteristics of human visual perception. In this paper, we present a no-reference VQA method, aiming to effectively and efficiently predict quality scores for UHD videos. First, we construct a spatial distortion feature network based on a super-resolution model (SR-SDFNet), which can quickly extract the global spatial distortion features of UHD videos. Then, to aggregate the spatial distortion features of each UHD frame, we propose a time fusion network based on a reinforcement learning model (RL-TFNet), in which the actor network continuously combines multiple frame features extracted by SR-SDFNet and outputs an action to adjust the current quality score to approximate the subjective score, and the critic network outputs action values to optimize the quality perception of the actor network. Finally, we conduct large-scale experiments on UHD VQA databases and the results reveal that, compared to other state-of-the-art VQA methods, our method achieves competitive quality prediction performance with a shorter runtime and fewer model parameters.<\/jats:p>","DOI":"10.3390\/s23031511","type":"journal-article","created":{"date-parts":[[2023,1,30]],"date-time":"2023-01-30T02:28:34Z","timestamp":1675045714000},"page":"1511","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Blind Video Quality Assessment for Ultra-High-Definition Video Based on Super-Resolution and Deep Reinforcement Learning"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3931-1239","authenticated-orcid":false,"given":"Zefeng","family":"Ying","sequence":"first","affiliation":[{"name":"School of Information and Communication Engineering, Communication University of China, Beijing 100024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Da","family":"Pan","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Communication University of China, Beijing 100024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ping","family":"Shi","sequence":"additional","affiliation":[{"name":"School of Information and Communication Engineering, Communication University of China, Beijing 100024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,1,29]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"7446","DOI":"10.1109\/TIP.2021.3106801","article-title":"St-greed: Space-time generalized entropic differences for frame rate dependent video quality prediction","volume":"30","author":"Madhusudana","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"2738","DOI":"10.1109\/TMM.2019.2908377","article-title":"Quality assessment for video with degradation along salient trajectories","volume":"21","author":"Wu","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"5027","DOI":"10.1109\/TIP.2019.2914950","article-title":"Study of subjective quality and objective blind quality prediction of stereoscopic videos","volume":"28","author":"Appina","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Korhonen, J., Su, Y., and You, J. (2020, January 12\u201316). Blind natural video quality prediction via statistical temporal features and deep spatial features. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413845"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"193020","DOI":"10.1109\/ACCESS.2020.3032080","article-title":"Multi-model standard for bitstream-, pixel-based and hybrid video quality assessment of uhd\/4k: Itu-t p.1204","volume":"8","author":"Raake","year":"2020","journal-title":"IEEE Access"},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"31842","DOI":"10.1109\/ACCESS.2021.3059932","article-title":"Modular framework and instances of pixel-based video quality models for uhd-1\/4k","volume":"9","author":"Rao","year":"2021","journal-title":"IEEE Access"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"5944","DOI":"10.1109\/TCSVT.2022.3164467","article-title":"Blindly assess quality of in-the-wild videos via quality-aware pre-training and motion perception","volume":"32","author":"Li","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1109\/LSP.2021.3136487","article-title":"No-reference video quality assessment using voxel-wise fmri models of the visual cortex","volume":"29","author":"Mahankali","year":"2022","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"3500","DOI":"10.1109\/TCSVT.2021.3114509","article-title":"Spatiotemporal representation learning for blind video quality assessment","volume":"32","author":"Liu","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"425","DOI":"10.1109\/OJSP.2021.3090333","article-title":"Rapique: Rapid and accurate video quality prediction of user generated content","volume":"2","author":"Tu","year":"2021","journal-title":"IEEE Open J. Signal Process."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Zheng, Q., Tu, Z., Fan, Y., Zeng, X., and Bovik, A.C. (2022, January 23\u201327). No-reference quality assessment of variable frame-rate videos using temporal bandpass statistics. Proceedings of the ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Singapore.","DOI":"10.1109\/ICASSP43922.2022.9746997"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"2244","DOI":"10.1109\/TCSVT.2018.2868063","article-title":"Blind video quality assessment with weakly supervised learning and resampling strategy","volume":"29","author":"Zhang","year":"2019","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, C., Su, L., and Huang, Q. (2017, January 21\u201323). Cnn-mr for no reference video quality assessment. Proceedings of the 2017 4th International Conference on Information Science and Control Engineering (ICISCE), Changsha, China.","DOI":"10.1109\/ICISCE.2017.56"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Wang, C., Su, L., and Zhang, W. (2018, January 10\u201312). Come for no-reference video quality assessment. Proceedings of the 2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), Miami, FL, USA.","DOI":"10.1109\/MIPR.2018.00056"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Chen, P., Li, L., Ma, L., Wu, J., and Shi, G. (2020, January 12\u201316). Rirnet: Recurrent-in-recurrent network for video quality assessment. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3413717"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"5612","DOI":"10.1109\/TIP.2020.2984879","article-title":"No-reference video quality assessment using natural spatiotemporal scene statistics","volume":"29","author":"Dendi","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"8059","DOI":"10.1109\/TIP.2021.3112055","article-title":"Chipqa: No-reference video quality prediction via space-time chips","volume":"30","author":"Ebenezer","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"1903","DOI":"10.1109\/TCSVT.2021.3088505","article-title":"Learning generalized spatial-temporal deep feature representation for no-reference video quality assessment","volume":"32","author":"Chen","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Xu, J., Li, J., Zhou, X., Zhou, W., Wang, B., and Chen, Z. (2021, January 20\u201324). Perceptual quality assessment of internet videos. Proceedings of the 29th ACM International Conference on Multimedia, Virtual Event.","DOI":"10.1145\/3474085.3475486"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Varga, D. (2022). No-reference video quality assessment using the temporal statistics of global and local image features. Sensors, 22.","DOI":"10.3390\/s22249696"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"3081","DOI":"10.1007\/s11042-022-13383-0","article-title":"Study on no-reference video quality assessment method incorporating dual deep learning networks","volume":"82","author":"Li","year":"2022","journal-title":"Multim. Tools Appl."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Lin, L., Zheng, Y., Chen, W., Lan, C., and Zhao, T. (2023). Saliency-aware spatio-temporal artifact detection for compressed video quality assessment. arXiv.","DOI":"10.1109\/LSP.2023.3283541"},{"key":"ref_23","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M.A. (2013). Playing atari with deep reinforcement learning. arXiv."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Wang, Z., Zhang, J., Lin, M., Wang, J., Luo, P., and Ren, J. (2020, January 13\u201319). Learning a reinforced agent for flexible exposure bracketing selection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00189"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Sun, M., Xiao, J., and Lim, E.G. (2021, January 20\u201325). Iterative shrinking for referring expression grounding using deep reinforcement learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01384"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Nauata, N., Hosseini, S., Chang, K.-H., Chu, H., Cheng, C.-Y., and Furukawa, Y. (2021, January 20\u201325). House-gan++: Generative adversarial layout refinement network towards intelligent computational agent for professional architects. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01342"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Wang, Y., Dong, M., Shen, J., Wu, Y., Cheng, S., and Pantic, M. (2020, January 13\u201319). Dynamic face video segmentation via reinforcement learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00699"},{"key":"ref_28","unstructured":"Lu, Y., Fu, J., Li, X., Zhou, W., Liu, S., Zhang, X., Jia, C., Liu, Y., and Chen, Z. (2022). Medical Image Computing and Computer Assisted Intervention\u2013MICCAI 2022: 25th International Conference, Singapore, 18\u201322 September 2022, Proceedings, Part I, Springer Nature."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"102427","DOI":"10.1016\/j.media.2022.102427","article-title":"Image quality assessment for machine learning tasks using meta-reinforcement learning","volume":"78","author":"Saeed","year":"2022","journal-title":"Med. Image Anal."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Wang, X., Xie, L., Dong, C., and Shan, Y. (2021, January 11\u201317). Real-esrgan: Training real-world blind super-resolution with pure synthetic data. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV) Workshops, Montreal, BC, Canada.","DOI":"10.1109\/ICCVW54120.2021.00217"},{"key":"ref_31","unstructured":"Netix (2021, January 12). Netix Vmaf. Available online: https:\/\/github.com\/Netflix\/vmaf."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"430","DOI":"10.1109\/TIP.2005.859378","article-title":"Image information and visual quality","volume":"15","author":"Sheikh","year":"2006","journal-title":"IEEE Trans. Image Process."},{"key":"ref_33","unstructured":"Lillicrap, T., Hunt, J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1467","DOI":"10.1109\/TCSVT.2017.2683504","article-title":"Subjective and objective quality assessment of compressed 4k uhd videos for immersive experience","volume":"28","author":"Cheon","year":"2018","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_35","unstructured":"Rao, R.R.R., G\u00f6ring, S., Robitza, W., Feiten, B., and Raake, A. (2019, January 9\u201311). Avt-vqdb-uhd-1: A large scale video quality database for uhd-1. Proceedings of the 2019 IEEE International Symposium on Multimedia (ISM), San Diego, CA, USA."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"2579","DOI":"10.1109\/TIP.2015.2426416","article-title":"A feature-enriched completely blind image quality evaluator","volume":"24","author":"Zhang","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_37","doi-asserted-by":"crossref","first-page":"3951","DOI":"10.1109\/TIP.2017.2708503","article-title":"dipiq: Blind image quality assessment by learning-to-rank discriminable image pairs","volume":"26","author":"Ma","year":"2017","journal-title":"IEEE Trans. Image Process."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, D., Jiang, T., and Jiang, M. (2019, January 21\u201325). Quality assessment of in-the-wild videos. Proceedings of the 27th ACM International Conference on Multimedia, Nice, France.","DOI":"10.1145\/3343031.3351028"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"5923","DOI":"10.1109\/TIP.2019.2923051","article-title":"Two-level approach for no-reference consumer video quality assessment","volume":"28","author":"Korhonen","year":"2019","journal-title":"IEEE Trans. Image Process."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"4449","DOI":"10.1109\/TIP.2021.3072221","article-title":"Ugc-vqa: Benchmarking blind video quality assessment for user generated content","volume":"30","author":"Tu","year":"2021","journal-title":"IEEE Trans. Image Process."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Hosu, V., Hahn, F., Jenadeleh, M., Lin, H., Men, H., Szir\u00e1nyi, T., Li, S., and Saupe, D. (June, January 31). The konstanz natural video database (konvid-1k). Proceedings of the 2017 Ninth International Conference on Quality of Multimedia Experience (QoMEX), Erfurt, Germany.","DOI":"10.1109\/QoMEX.2017.7965673"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_43","unstructured":"Howard, A., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Li, K., Li, K., Wang, L., Zhong, B., and Fu, Y.R. (2018, January 8\u201314). Image Super-Resolution Using Very Deep Residual Channel Attention Networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_18"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1007\/s11760-019-01510-8","article-title":"No-reference video quality assessment via pretrained cnn and lstm networks","volume":"13","author":"Varga","year":"2019","journal-title":"Signal Image Video Process."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1511\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T18:19:21Z","timestamp":1760120361000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/3\/1511"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,29]]},"references-count":45,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2023,2]]}},"alternative-id":["s23031511"],"URL":"https:\/\/doi.org\/10.3390\/s23031511","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2023,1,29]]}}}