{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,7]],"date-time":"2026-02-07T19:45:05Z","timestamp":1770493505140,"version":"3.49.0"},"reference-count":50,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2023,8,23]],"date-time":"2023-08-23T00:00:00Z","timestamp":1692748800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62002283"],"award-info":[{"award-number":["62002283"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62311530046"],"award-info":[{"award-number":["62311530046"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62066020"],"award-info":[{"award-number":["62066020"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["2022GXLH-01-24"],"award-info":[{"award-number":["2022GXLH-01-24"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["xhj032021017"],"award-info":[{"award-number":["xhj032021017"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Key Research and Development Program of Shaanxi","award":["62002283"],"award-info":[{"award-number":["62002283"]}]},{"name":"Key Research and Development Program of Shaanxi","award":["62311530046"],"award-info":[{"award-number":["62311530046"]}]},{"name":"Key Research and Development Program of Shaanxi","award":["62066020"],"award-info":[{"award-number":["62066020"]}]},{"name":"Key Research and Development Program of Shaanxi","award":["2022GXLH-01-24"],"award-info":[{"award-number":["2022GXLH-01-24"]}]},{"name":"Key Research and Development Program of Shaanxi","award":["xhj032021017"],"award-info":[{"award-number":["xhj032021017"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["62002283"],"award-info":[{"award-number":["62002283"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["62311530046"],"award-info":[{"award-number":["62311530046"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["62066020"],"award-info":[{"award-number":["62066020"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["2022GXLH-01-24"],"award-info":[{"award-number":["2022GXLH-01-24"]}]},{"name":"Fundamental Research Funds for the Central Universities","award":["xhj032021017"],"award-info":[{"award-number":["xhj032021017"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Remote Sensing"],"abstract":"<jats:p>Hyperspectral image (HSI) super-resolution is a practical and challenging task as it requires the reconstruction of a large number of spectral bands. Achieving excellent reconstruction results can greatly benefit subsequent downstream tasks. The current mainstream hyperspectral super-resolution methods mainly utilize 3D convolutional neural networks (3D CNN) for design. However, the commonly used small kernel size in 3D CNN limits the model\u2019s receptive field, preventing it from considering a wider range of contextual information. Though the receptive field could be expanded by enlarging the kernel size, it results in a dramatic increase in model parameters. Furthermore, the popular vision transformers designed for natural images are not suitable for processing HSI. This is because HSI exhibits sparsity in the spatial domain, which can lead to significant computational resource waste when using self-attention. In this paper, we design a hybrid architecture called HyFormer, which combines the strengths of CNN and transformer for hyperspectral super-resolution. The transformer branch enables intra-spectra interaction to capture fine-grained contextual details at each specific wavelength. Meanwhile, the CNN branch facilitates efficient inter-spectra feature extraction among different wavelengths while maintaining a large receptive field. Specifically, in the transformer branch, we propose a novel Grouping-Aggregation transformer (GAT), comprising grouping self-attention (GSA) and aggregation self-attention (ASA). The GSA is employed to extract diverse fine-grained features of targets, while the ASA facilitates interaction among heterogeneous textures allocated to different channels. In the CNN branch, we propose a Wide-Spanning Separable 3D Attention (WSSA) to enlarge the receptive field while keeping a low parameter number. Building upon WSSA, we construct a wide-spanning CNN module to efficiently extract inter-spectra features. Extensive experiments demonstrate the superior performance of our HyFormer.<\/jats:p>","DOI":"10.3390\/rs15174131","type":"journal-article","created":{"date-parts":[[2023,8,23]],"date-time":"2023-08-23T08:01:21Z","timestamp":1692777681000},"page":"4131","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["HyFormer: Hybrid Grouping-Aggregation Transformer and Wide-Spanning CNN for Hyperspectral Image Super-Resolution"],"prefix":"10.3390","volume":"15","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4146-2661","authenticated-orcid":false,"given":"Yantao","family":"Ji","sequence":"first","affiliation":[{"name":"School of Software Engineering, Xi\u2019an Jiaotong University, Xi\u2019an 710049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jingang","family":"Shi","sequence":"additional","affiliation":[{"name":"School of Software Engineering, Xi\u2019an Jiaotong University, Xi\u2019an 710049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yaping","family":"Zhang","sequence":"additional","affiliation":[{"name":"School of Software Engineering, Xi\u2019an Jiaotong University, Xi\u2019an 710049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haokun","family":"Yang","sequence":"additional","affiliation":[{"name":"School of Software Engineering, Xi\u2019an Jiaotong University, Xi\u2019an 710049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuan","family":"Zong","sequence":"additional","affiliation":[{"name":"Key Laboratory of Child Development and Learning Science, Southeast University, Nanjing 211189, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ling","family":"Xu","sequence":"additional","affiliation":[{"name":"School of Human Settlements and Civil Engineering, Xi\u2019an Jiaotong University, Xi\u2019an 710049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,8,23]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1038\/35017638","article-title":"Detection of preinvasive cancer cells","volume":"406","author":"Backman","year":"2000","journal-title":"Nature"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"010901","DOI":"10.1117\/1.JBO.19.1.010901","article-title":"Medical hyperspectral imaging: A review","volume":"19","author":"Lu","year":"2014","journal-title":"J. Biomed. Opt."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1145\/2185520.2185534","article-title":"3D imaging spectroscopy for measuring hyperspectral patterns on solid objects","volume":"31","author":"Kim","year":"2012","journal-title":"ACM Trans. Graph."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1016\/S0169-1368(99)00007-4","article-title":"Remote sensing for mineral exploration","volume":"14","author":"Sabins","year":"1999","journal-title":"Ore Geol. Rev."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"599","DOI":"10.1080\/2150704X.2022.2057824","article-title":"Self-paced collaborative representation with manifold weighting for hyperspectral anomaly detection","volume":"13","author":"Ji","year":"2022","journal-title":"Remote Sens. Lett."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"80","DOI":"10.1186\/s13007-017-0233-z","article-title":"Hyperspectral image analysis techniques for the detection and classification of the early onset of plant disease and stress","volume":"13","author":"Lowe","year":"2017","journal-title":"Plant Methods"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"5345","DOI":"10.1109\/TNNLS.2018.2798162","article-title":"Deep hyperspectral image sharpening","volume":"29","author":"Dian","year":"2018","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"8059","DOI":"10.1109\/TGRS.2020.2986313","article-title":"Hyperspectral pansharpening using deep prior and dual attention residual network","volume":"58","author":"Zheng","year":"2020","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"8028","DOI":"10.1109\/TIP.2020.3009830","article-title":"A truncated matrix decomposition for hyperspectral image super-resolution","volume":"29","author":"Liu","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"2672","DOI":"10.1109\/TNNLS.2018.2885616","article-title":"Learning a low tensor-train rank representation for hyperspectral image super-resolution","volume":"30","author":"Dian","year":"2019","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"1860","DOI":"10.1109\/TIP.2005.854479","article-title":"Super-resolution reconstruction of hyperspectral images","volume":"14","author":"Akgun","year":"2005","journal-title":"IEEE Trans. Image Process."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1250","DOI":"10.1109\/LGRS.2016.2579661","article-title":"Hyperspectral image super-resolution by spectral mixture analysis and spatial\u2013spectral group sparsity","volume":"13","author":"Li","year":"2016","journal-title":"IEEE Geosci. Remote Sens. Lett."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Wang, Y., Chen, X., Han, Z., and He, S. (2017). Hyperspectral image super-resolution via nonlocal low-rank tensor approximation and total variation regularization. Remote Sens., 9.","DOI":"10.3390\/rs9121286"},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"295","DOI":"10.1109\/TPAMI.2015.2439281","article-title":"Image super-resolution using deep convolutional networks","volume":"38","author":"Dong","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Li, K., Li, K., Wang, L., Zhong, B., and Fu, Y. (2018, January 8\u201314). Image super-resolution using very deep residual channel attention networks. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01234-2_18"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2980","DOI":"10.1109\/TIP.2018.2813163","article-title":"Hallucinating face image by regularization models in high-resolution feature space","volume":"27","author":"Shi","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Niu, B., Wen, W., Ren, W., Zhang, X., Yang, L., Wang, S., Zhang, K., Cao, X., and Shen, H. (2020, January 23\u201328). Single image super-resolution via a holistic attention network. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Part XII 16.","DOI":"10.1007\/978-3-030-58610-2_12"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"2223","DOI":"10.1109\/TMM.2019.2898752","article-title":"Face hallucination via coarse-to-fine recursive kernel regression structure","volume":"21","author":"Shi","year":"2019","journal-title":"IEEE Trans. Multimed."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Tian, C., Zhang, Y., Zuo, W., Lin, C.W., Zhang, D., and Yuan, Y. (2022). A heterogeneous group CNN for image super-resolution. IEEE Trans. Neural Netw. Learn. Syst., 1\u201313.","DOI":"10.1109\/TNNLS.2022.3210433"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Zhou, L., Cai, H., Gu, J., Li, Z., Liu, Y., Chen, X., Qiao, Y., and Dong, C. (2022, January 23\u201327). Efficient image super-resolution using vast-receptive-field attention. Proceedings of the Computer Vision\u2013ECCV 2022 Workshops, Tel Aviv, Israel. Part II.","DOI":"10.1007\/978-3-031-25063-7_16"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Liu, T., Cheng, J., and Tan, S. (2023, January 18\u201322). Spectral Bayesian Uncertainty for Image Super-Resolution. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01742"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Mei, S., Yuan, X., Ji, J., Zhang, Y., Wan, S., and Du, Q. (2017). Hyperspectral image spatial super-resolution via 3D full convolutional neural network. Remote Sens., 9.","DOI":"10.3390\/rs9111139"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Yang, J., Zhao, Y.Q., Chan, J.C.W., and Xiao, L. (2019). A multi-scale wavelet 3D-CNN for hyperspectral image super-resolution. Remote Sens., 11.","DOI":"10.3390\/rs11131557"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Li, Q., Wang, Q., and Li, X. (2020). Mixed 2D\/3D convolutional network for hyperspectral image super-resolution. Remote Sens., 12.","DOI":"10.3390\/rs12101660"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"8693","DOI":"10.1109\/TGRS.2020.3047363","article-title":"Exploring the relationship between 2D\/3D convolution for hyperspectral image super-resolution","volume":"59","author":"Li","year":"2021","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zhang, J., Shao, M., Wan, Z., and Li, Y. (2021). Multi-scale feature mapping network for hyperspectral image super-resolution. Remote Sens., 13.","DOI":"10.3390\/rs13204180"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"85","DOI":"10.1016\/j.neucom.2021.10.041","article-title":"Hyperspectral image super-resolution via multi-domain feature learning","volume":"472","author":"Li","year":"2022","journal-title":"Neurocomputing"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Tang, Z., Xu, Q., Wu, P., Shi, Z., and Pan, B. (2022). Feedback Refined Local-Global Network for Super-Resolution of Hyperspectral Imagery. Remote Sens., 14.","DOI":"10.3390\/rs14081944"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Zhang, J., Zheng, R., Chen, X., Hong, Z., Li, Y., and Lu, R. (2023). Spectral Correlation and Spatial High\u2013Low Frequency Information of Hyperspectral Image Super-Resolution Network. Remote Sens., 15.","DOI":"10.3390\/rs15092472"},{"key":"ref_30","first-page":"4905","article-title":"Understanding the effective receptive field in deep convolutional neural networks","volume":"29","author":"Luo","year":"2016","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_31","first-page":"6000","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_32","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 11\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Online.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_34","first-page":"9355","article-title":"Twins: Revisiting the design of spatial attention in vision transformers","volume":"34","author":"Chu","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Shi, J., Wang, Y., Dong, S., Hong, X., Yu, Z., Wang, F., Wang, C., and Gong, Y. (2022, January 23\u201329). Idpt: Interconnected dual pyramid transformer for face super-resolution. Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), Messe Wien, Austria.","DOI":"10.24963\/ijcai.2022\/182"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Cheng, B., Misra, I., Schwing, A.G., Kirillov, A., and Girdhar, R. (2022, January 18\u201324). Masked-attention mask transformer for universal image segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00135"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Shi, J., Wang, Y., Yu, Z., Li, G., Hong, X., Wang, F., and Gong, Y. (2023). Exploiting Multi-scale Parallel Self-attention and Local Variation via Dual-branch transformer-CNN Structure for Face Super-resolution. IEEE Trans. Multimed., 1\u201314.","DOI":"10.1109\/TMM.2023.3301225"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Li, Y., Mao, H., Girshick, R., and He, K. (2022, January 23\u201327). Exploring plain vision transformer backbones for object detection. Proceedings of the Computer Vision\u2013ECCV 2022: 17th European Conference, Tel Aviv, Israel. Part IX.","DOI":"10.1007\/978-3-031-20077-9_17"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Li, J., Yu, Z., and Shi, J. (2023, January 7\u201314). Learning motion-robust remote photoplethysmography through arbitrary resolution videos. Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA.","DOI":"10.1609\/aaai.v37i1.25217"},{"key":"ref_40","unstructured":"Pan, Z., Cai, J., and Zhuang, B. (2022). Fast vision transformers with hilo attention. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Ji, Y., Jiang, P., Shi, J., Guo, Y., Zhang, R., and Wang, F. (2022, January 16\u201319). Information-Growth Swin transformer Network for Image Super-Resolution. Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP), Bordeaux, France.","DOI":"10.1109\/ICIP46576.2022.9897359"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Xie, S., Sun, C., Huang, J., Tu, Z., and Murphy, K. (2018, January 8\u201314). Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"ref_43","unstructured":"Hendrycks, D., and Gimpel, K. (2016). Gaussian error linear units (gelus). arXiv."},{"key":"ref_44","unstructured":"Ba, J.L., Kiros, J.R., and Hinton, G.E. (2016). Layer normalization. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"2241","DOI":"10.1109\/TIP.2010.2046811","article-title":"Generalized assorted pixel camera: Postcapture control of resolution, dynamic range, and spectrum","volume":"19","author":"Yasuma","year":"2010","journal-title":"IEEE Trans. Image Process."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Chakrabarti, A., and Zickler, T. (2011, January 20\u201325). Statistics of real-world hyperspectral images. Proceedings of the CVPR 2011, Colorado Springs, CO, USA.","DOI":"10.1109\/CVPR.2011.5995660"},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"39","DOI":"10.1016\/j.visres.2015.07.005","article-title":"Spatial distributions of local illumination color in natural scenes","volume":"120","author":"Nascimento","year":"2016","journal-title":"Vis. Res."},{"key":"ref_48","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"1082","DOI":"10.1109\/TCI.2020.2996075","article-title":"Learning spatial-spectral prior for super-resolution of hyperspectral imagery","volume":"6","author":"Jiang","year":"2020","journal-title":"IEEE Trans. Comput. Imaging"},{"key":"ref_50","first-page":"5541416","article-title":"A Group-Based Embedding Learning and Integration Network for Hyperspectral Image Super-Resolution","volume":"60","author":"Wang","year":"2022","journal-title":"IEEE Trans. Geosci. Remote Sens."}],"container-title":["Remote Sensing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/17\/4131\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:40:43Z","timestamp":1760128843000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2072-4292\/15\/17\/4131"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,8,23]]},"references-count":50,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2023,9]]}},"alternative-id":["rs15174131"],"URL":"https:\/\/doi.org\/10.3390\/rs15174131","relation":{},"ISSN":["2072-4292"],"issn-type":[{"value":"2072-4292","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,8,23]]}}}