{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:29:34Z","timestamp":1760236174571,"version":"build-2065373602"},"reference-count":44,"publisher":"MDPI AG","issue":"11","license":[{"start":{"date-parts":[[2021,10,25]],"date-time":"2021-10-25T00:00:00Z","timestamp":1635120000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"The National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["61871272,62001300"],"award-info":[{"award-number":["61871272,62001300"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"The Natural Science Foundation of Guangdong Province","award":["2021A1515011807,2021A1515011679"],"award-info":[{"award-number":["2021A1515011807,2021A1515011679"]}]},{"name":"The RD Program of Shenzhen","award":["JCYJ20180508152204044"],"award-info":[{"award-number":["JCYJ20180508152204044"]}]},{"name":"The Shenzhen Fundamental Research Program","award":["JCYJ20190808173617147"],"award-info":[{"award-number":["JCYJ20190808173617147"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Convolutional Neural Networks (CNNs) have been widely used in video super-resolution (VSR). Most existing VSR methods focus on how to utilize the information of multiple frames, while neglecting the feature correlations of the intermediate features, thus limiting the feature expression of the models. To address this problem, we propose a novel SAA network, that is, Scale-and-Attention-Aware Networks, to apply different attention to different temporal-length streams, while further exploring both spatial and channel attention on separate streams with a newly proposed Criss-Cross Channel Attention Module (C3AM). Experiments on public VSR datasets demonstrate the superiority of our method over other state-of-the-art methods in terms of both quantitative and qualitative metrics.<\/jats:p>","DOI":"10.3390\/e23111398","type":"journal-article","created":{"date-parts":[[2021,10,25]],"date-time":"2021-10-25T21:40:21Z","timestamp":1635198021000},"page":"1398","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":4,"title":["S2A: Scale-Attention-Aware Networks for Video Super-Resolution"],"prefix":"10.3390","volume":"23","author":[{"given":"Taian","family":"Guo","sequence":"first","affiliation":[{"name":"College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tao","family":"Dai","sequence":"additional","affiliation":[{"name":"College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ling","family":"Liu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zexuan","family":"Zhu","sequence":"additional","affiliation":[{"name":"College of Computer Science and Software Engineering, Shenzhen University, Shenzhen 518060, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shu-Tao","family":"Xia","sequence":"additional","affiliation":[{"name":"Tsinghua Shenzhen International Graduate School, Tsinghua University, Beijing 100084, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,10,25]]},"reference":[{"key":"ref_1","unstructured":"Huang, Y., Wang, W., and Wang, L. (2015, January 7\u201312). Bidirectional recurrent convolutional networks for multi-frame super-resolution. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Tao, X., Gao, H., Liao, R., Wang, J., and Jia, J. (2017, January 22\u201329). Detail-revealing deep video super-resolution. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.479"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Sajjadi, M.S., Vemulapalli, R., and Brown, M. (2018, January 18\u201323). Frame-recurrent video super-resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00693"},{"key":"ref_4","unstructured":"Yi, P., Wang, Z., Jiang, K., Jiang, J., and Ma, J. (November, January 29). Progressive fusion video super-resolution network via exploiting non-local spatio-temporal correlations. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Isobe, T., Li, S., Jia, X., Yuan, S., Slabaugh, G., Xu, C., Li, Y.L., Wang, S., and Tian, Q. (2020, January 16\u201318). Video super-resolution with temporal group attention. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00803"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Xiao, Z., Fu, X., Huang, J., Cheng, Z., and Xiong, Z. (2021, January 16\u201325). Space-time distillation for video super-resolution. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00215"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1109\/TCI.2016.2532323","article-title":"Video super-resolution with convolutional neural networks","volume":"2","author":"Kappeler","year":"2016","journal-title":"IEEE Trans. Comput. Imaging"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Caballero, J., Ledig, C., Aitken, A., Acosta, A., Totz, J., Wang, Z., and Shi, W. (2017, January 21\u201326). Real-time video super-resolution with spatio-temporal networks and motion compensation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.304"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Haris, M., Shakhnarovich, G., and Ukita, N. (2019, January 16\u201320). Recurrent Back-Projection Network for Video Super-Resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00402"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Li, S., He, F., Du, B., Zhang, L., Xu, Y., and Tao, D. (2019). Fast Spatio-Temporal Residual Network for Video Super-Resolution. arXiv.","DOI":"10.1109\/CVPR.2019.01077"},{"key":"ref_11","unstructured":"Liu, H., Ruan, Z., Zhao, P., Dong, C., Shang, F., Liu, Y., and Yang, L. (2020). Video super resolution based on deep learning: A comprehensive survey. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1109\/TPAMI.2015.2439281","article-title":"Image super-resolution using deep convolutional networks","volume":"38","author":"Dong","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","unstructured":"Kim, J., Kwon Lee, J., and Mu Lee, K. (July, January 26). Deeply-recursive convolutional network for image super-resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Tai, Y., Yang, J., and Liu, X. (2017, January 21\u201326). Image super-resolution via deep recursive residual network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.298"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Lai, W.S., Huang, J.B., Ahuja, N., and Yang, M.H. (2017, January 21\u201326). Deep laplacian pyramid networks for fast and accurate super-resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.618"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Ledig, C., Theis, L., Husz\u00e1r, F., Caballero, J., Cunningham, A., Acosta, A., Aitken, A., Tejani, A., Totz, J., and Wang, Z. (2017, January 21\u201326). Photo-realistic single image super-resolution using a generative adversarial network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.19"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Tai, Y., Yang, J., Liu, X., and Xu, C. (2017, January 22\u201329). Memnet: A persistent memory network for image restoration. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.486"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Tian, Y., Kong, Y., Zhong, B., and Fu, Y. (2018, January 18\u201323). Residual dense network for image super-resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00262"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Li, K., Li, K., Wang, L., Zhong, B., and Fu, Y. (2018, January 8\u201314). Image super-resolution using very deep residual channel attention networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_18"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Dai, T., Cai, J., Zhang, Y., Xia, S.T., and Zhang, L. (2019, January 16\u201320). Second-order Attention Network for Single Image Super-Resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01132"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Soh, J.W., Park, G.Y., Jo, J., and Cho, N.I. (2019, January 16\u201320). Natural and Realistic Single Image Super-Resolution with Explicit Natural Manifold Discrimination. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00831"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Dong, C., Loy, C.C., and Tang, X. (2016). Accelerating the super-resolution convolutional neural network. European Conference on Computer Vision, Springer.","DOI":"10.1007\/978-3-319-46475-6_25"},{"key":"ref_23","unstructured":"Shi, W., Caballero, J., Husz\u00e1r, F., Totz, J., Aitken, A.P., Bishop, R., Rueckert, D., and Wang, Z. (July, January 26). Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_24","unstructured":"Kim, J., Kwon Lee, J., and Mu Lee, K. (July, January 26). Accurate image super-resolution using very deep convolutional networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Tong, T., Li, G., Liu, X., and Gao, Q. (2017, January 21\u201326). Image super-resolution using dense skip connections. Proceedings of the IEEE International Conference on Computer Vision, Honolulu, HI, USA.","DOI":"10.1109\/ICCV.2017.514"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Wang, X., Yu, K., Wu, S., Gu, J., Liu, Y., Dong, C., Qiao, Y., and Change Loy, C. (2018, January 8\u201314). Esrgan: Enhanced super-resolution generative adversarial networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-11021-5_5"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Haris, M., Shakhnarovich, G., and Ukita, N. (2018, January 18\u201323). Deep Back-Projection Networks for Super-Resolution. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00179"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Lv, Y., Dai, T., Chen, B., Lu, J., Xia, S.T., and Cao, J. (2021, January 18\u201319). HOCA: Higher-Order Channel Attention for Single Image Super-Resolution. Proceedings of the 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Paris, France.","DOI":"10.1109\/ICASSP39728.2021.9414892"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Liao, R., Tao, X., Li, R., Ma, Z., and Jia, J. (2015, January 11\u201318). Video super-resolution via deep draft-ensemble learning. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.68"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Liu, D., Wang, Z., Fan, Y., Liu, X., Wang, Z., Chang, S., and Huang, T. (2017, January 22\u201329). Robust video super-resolution with learned temporal dynamics. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.274"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Chan, K.C., Wang, X., Yu, K., Dong, C., and Loy, C.C. (2021, January 16\u201325). BasicVSR: The search for essential components in video super-resolution and beyond. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00491"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Liu, S., Zheng, C., Lu, K., Gao, S., Wang, N., Wang, B., Zhang, D., Zhang, X., and Xu, T. (2021, January 16\u201325). Evsrnet: Efficient video super-resolution with neural architecture search. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPRW53098.2021.00281"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Jo, Y., Wug Oh, S., Kang, J., and Joo Kim, S. (2018, January 18\u201323). Deep video super-resolution network using dynamic upsampling filters without explicit motion compensation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00340"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201323). Non-local neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_35","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_36","unstructured":"Buades, A., Coll, B., and Morel, J.M. (2005, January 20\u201325). A non-local algorithm for image denoising. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_37","unstructured":"Huang, Z., Wang, X., Huang, L., Huang, C., Wei, Y., and Liu, W. (November, January 27). Ccnet: Criss-cross attention for semantic segmentation. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., and Sun, G. (2018, January 18\u201323). Squeeze-and-excitation networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"1106","DOI":"10.1007\/s11263-018-01144-2","article-title":"Video enhancement with task-oriented flow","volume":"127","author":"Xue","year":"2019","journal-title":"Int. J. Comput. Vis."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Liu, C., and Sun, D. (2011, January 20\u201325). A Bayesian approach to adaptive video super resolution. Proceedings of the 24th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2011, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995614"},{"key":"ref_41","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_42","unstructured":"Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017, January 4\u20139). Automatic differentiation in pytorch. Proceedings of the 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA."},{"key":"ref_43","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1109\/TPAMI.2013.127","article-title":"On Bayesian adaptive video super resolution","volume":"36","author":"Liu","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_44","unstructured":"Tian, Y., Zhang, Y., Fu, Y., and Xu, C. (2018). Tdan: Temporally deformable alignment network for video super-resolution. arXiv."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/11\/1398\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T07:23:07Z","timestamp":1760167387000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/23\/11\/1398"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,10,25]]},"references-count":44,"journal-issue":{"issue":"11","published-online":{"date-parts":[[2021,11]]}},"alternative-id":["e23111398"],"URL":"https:\/\/doi.org\/10.3390\/e23111398","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2021,10,25]]}}}