{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,19]],"date-time":"2026-06-19T03:59:48Z","timestamp":1781841588235,"version":"3.54.5"},"reference-count":34,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2025,8,29]],"date-time":"2025-08-29T00:00:00Z","timestamp":1756425600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Scientific Research Project of Education Department of Jilin Province","award":["JJKH20240573KJ"],"award-info":[{"award-number":["JJKH20240573KJ"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Symmetry"],"abstract":"<jats:p>Recognizing Manchu words can be challenging due to their complex character variations, subtle differences between similar characters, and homographic polysemy. Most studies rely on character segmentation techniques for character recognition or use convolutional neural networks (CNNs) to encode word images for word recognition. However, these methods can lead to segmentation errors or a loss of semantic information, which reduces the accuracy of word recognition. To address the limitations in the long-range dependency modeling of CNNs and enhance semantic coherence, we propose a hybrid architecture to fuse the spatial features of original images and spectral features. Specifically, we first leverage the Short-Time Fourier Transform (STFT) to preprocess the raw input images and thereby obtain their multi-view spectral features. Then, we leverage a primary CNN block and a pair of symmetric CNN blocks to construct a symmetric spectral enhancement module, which is used to encode the raw input features and the multi-view spectral features. Subsequently, we design a feature fusion module via Swin Transformer to fuse multi-view spectral embedding and thereby concat it with the raw input embedding. Finally, we leverage a Transformer decoder to obtain the target output. We conducted extensive experiments on Manchu words benchmark datasets to evaluate the effectiveness of our proposed framework. The experimental results demonstrated that our framework performs robustly in word recognition tasks and exhibits excellent generalization capabilities. Additionally, our model outperformed other baseline methods in multiple writing-style font-recognition tasks.<\/jats:p>","DOI":"10.3390\/sym17091408","type":"journal-article","created":{"date-parts":[[2025,8,29]],"date-time":"2025-08-29T12:25:57Z","timestamp":1756470357000},"page":"1408","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Research on Multi-Path Feature Fusion Manchu Recognition Based on Swin Transformer"],"prefix":"10.3390","volume":"17","author":[{"given":"Yu","family":"Zhou","sequence":"first","affiliation":[{"name":"School of Mathematics and Computer Science, Jilin Normal University, Siping 136000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mingyan","family":"Li","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Jilin Normal University, Siping 136000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hang","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Jilin Normal University, Siping 136000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jinchi","family":"Yu","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Jilin Normal University, Siping 136000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Mingchen","family":"Sun","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Jilin University, Changchun 130012, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Dadong","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Jilin Normal University, Siping 136000, China"},{"name":"Jilin Provincial Key Laboratory for Numerical Simulation, Jilin Normal University, Siping 136000, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2025,8,29]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1685","DOI":"10.11834\/jig.240015","article-title":"Survey on text analysis and recognition for multiethnic scripts","volume":"29","author":"Wang","year":"2024","journal-title":"J. Image Graph."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Zhang, G.-y., Li, J., and Wang, A. (2006, January 13\u201316). A New Recognition Method for the Handwritten Manchu Character Unit. Proceedings of the 2006 International Conference on Machine Learning and Cybernetics, Dalian, China.","DOI":"10.1109\/ICMLC.2006.258471"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"523","DOI":"10.1080\/09720529.2016.1177963","article-title":"A New Method for Baseline Extraction of Manchu Word","volume":"19","author":"Zheng","year":"2016","journal-title":"J. Discret. Math. Sci. Cryptogr."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Huang, D., Li, M., Zheng, R.R., Xu, S., and Bi, J.J. (2017, January 21\u201323). Synthetic data and DAG-SVM classifier for segmentation-free Manchu word recognition. Proceedings of the 2017 International Conference on Computing Intelligence and Information System (CIIS), Nanjing, China.","DOI":"10.1109\/CIIS.2017.15"},{"key":"ref_5","unstructured":"Sile, H., Jabu, Q.J., and Tao, X. (2025, August 24). Information Technology Manchu Nominal Characters, Presentation Characters, and Use Rules of Controlling Characters, Available online: https:\/\/openstd.samr.gov.cn\/bzgk\/gb\/newGbInfo?hcno=67DA394E47B970F80BBABE5511B9AAE2."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Li, M., Zheng, R., Xu, S., Feng, Y., and Hou, D. (2018, January 13\u201315). Manchu Word Recognition Based on Convolutional Neural Network with Spatial Pyramid Pooling. Proceedings of the 2018 11th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), Beijing, China.","DOI":"10.1109\/CISP-BMEI.2018.8633131"},{"key":"ref_7","unstructured":"Cheng, C. (2023). Hafumbuk\u016b: A Text Recognition Model for the Manchu Script. [Bachelor\u2019s Thesis, Harvard College]. Available online: https:\/\/dash.harvard.edu\/server\/api\/core\/bitstreams\/074e5d47-1251-4d40-babf-15d76a357a75\/content."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Li, M., Lv, T., Chen, J., Cui, L., Lu, Y., Florencio, D., Zhang, C., Li, Z., and Wei, F. (2023, January 7\u201314). TrOCR: Transformer-based optical character recognition with pre-trained models. Proceedings of the AAAI Conference on Artificial Intelligence, Washington, DC, USA.","DOI":"10.1609\/aaai.v37i11.26538"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"43","DOI":"10.1080\/09720529.2016.1177965","article-title":"Manchu Character Segmentation and Recognition Method","volume":"20","author":"Xu","year":"2016","journal-title":"J. Discret. Math. Sci. Cryptogr."},{"key":"ref_10","first-page":"2347","article-title":"Off-line Manchu character recognition based on multi-classifier ensemble with combination features","volume":"33","author":"Wei","year":"2012","journal-title":"Comput. Eng. Des."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"93","DOI":"10.1007\/s10032-009-0106-8","article-title":"Multi-font printed Mongolian document recognition system","volume":"13","author":"Peng","year":"2010","journal-title":"Int. J. Doc. Anal. Recognit."},{"key":"ref_12","first-page":"536","article-title":"An improved Manchu character recognition method","volume":"39","author":"Xu","year":"2016","journal-title":"J. Mech. Eng. Res. Dev."},{"key":"ref_13","first-page":"5520338","article-title":"OCR with the deep CNN model for ligature script-based languages like Manchu","volume":"2021","author":"Zhang","year":"2021","journal-title":"Sci. Program."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zheng, R., Liu, M., Hao, J., Bao, J., and Wang, B. (2018, January 13\u201315). Segmentation-Free Multi-Font Printed Manchu Word Recognition Using Deep Convolutional Features and Data Augmentation. Proceedings of the 2018 11th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), Beijing, China.","DOI":"10.1109\/CISP-BMEI.2018.8633208"},{"key":"ref_15","first-page":"80","article-title":"Manchu Script Letters Dataset Creation and Labeling","volume":"22","author":"Snowberger","year":"2024","journal-title":"J. Inf. Commun. Converg. Eng."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2298","DOI":"10.1109\/TPAMI.2016.2646371","article-title":"An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition","volume":"39","author":"Shi","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Ren, Q.-D.-E.-J., Wang, L., Ma, Z., and Barintag, S. (2024). Offline Mongolian handwriting recognition based on data augmentation and improved ECA-Net. Electronics, 13.","DOI":"10.3390\/electronics13050835"},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1007\/s10032-021-00388-y","article-title":"An end-to-end network for irregular printed Mongolian recognition","volume":"25","author":"Cui","year":"2022","journal-title":"Int. J. Doc. Anal. Recognit."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Zhang, H., Wei, H., Bao, F., and Gao, G. (2017, January 9\u201315). Segmentation-free printed traditional Mongolian OCR using sequence to sequence with attention model. Proceedings of the 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR), Kyoto, Japan.","DOI":"10.1109\/ICDAR.2017.101"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Sun, S., Wang, H., and Wang, Y. (2023, January 20\u201323). A Hybrid Approach Using Convolution and Transformer for Mongolian Ancient Documents Recognition. Proceedings of the Neural Information Processing: 30th International Conference, ICONIP 2023, Proceedings, Part XIII, Changsha, China.","DOI":"10.1007\/978-981-99-8178-6_13"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Ren, S., Zhou, D., He, S., Feng, J., and Wang, X. (2022, January 21\u201324). Shunted self-attention via multi-scale token aggregation. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01058"},{"key":"ref_22","unstructured":"Gu, X., Wang, L., Dai, Z., Chen, Y., Han, X., and Zhang, Y. (2023). AdaFuse: Adaptive Medical Image Fusion Based on Spatial-Frequential Cross Attention. arXiv."},{"key":"ref_23","unstructured":"Li, K., and Li, Y. (2024). Lightweight Single-Image Super-Resolution Network Based on Dual Paths. arXiv."},{"key":"ref_24","first-page":"5512913","article-title":"Dual-View Spectral and Global Spatial Feature Fusion Network for Hyperspectral Image Classification","volume":"61","author":"Tan","year":"2023","journal-title":"IEEE Trans. Geosci. Remote Sens."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., and Zhang, Z. (2021, January 11\u201317). Swin Transformer: Hierarchical vision transformer using shifted windows. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Zeng, C., and Sun, C. (2022, January 29\u201331). Swin Transformer with Feature Pyramid Networks for Scene Text Detection of the Secondary Circuit Cabinet Wiring. Proceedings of the 2022 IEEE 4th International Conference on Power, Intelligent Computing and Systems (ICPICS), Shenyang, China.","DOI":"10.1109\/ICPICS55264.2022.9873542"},{"key":"ref_27","unstructured":"Zhang, B., Chen, J., and Wen, Q. (2022). Single Image Super-Resolution Using Lightweight Networks Based on Swin Transformer. arXiv."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Chen, C.-F.R., Fan, Q., and Panda, R. (2021, January 11\u201317). CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification. Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision (ICCV), Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00041"},{"key":"ref_29","unstructured":"Huang, Y., Chen, Y., Wang, C., Xu, H., Shi, B., and Wang, H. (2024). Image Super-Resolution Reconstruction Network Based on Enhanced Swin Transformer via Alternating Aggregation of Local\u2013Global Features. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Jing, T., Liu, C., and Chen, Y. (2025). A Lightweight Single-Image Super-Resolution Method Based on the Parallel Connection of Convolution and Swin Transformer Blocks. Appl. Sci., 15.","DOI":"10.3390\/app15041806"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"599","DOI":"10.1016\/j.isprsjprs.2023.07.001","article-title":"An Attention-Based Multiscale Transformer Network for Remote Sensing Image Change Detection","volume":"202","author":"Li","year":"2023","journal-title":"ISPRS J. Photogramm. Remote Sens."},{"key":"ref_32","unstructured":"(2025, June 11). Manchu Dataset. Available online: https:\/\/deepwiki.com\/tyotakuki\/ManchuOCR."},{"key":"ref_33","unstructured":"Wang, Z., Liu, S., Wang, M., Wang, X., and Qu, Y. (2022, January 22\u201326). AMRE: An Attention-Based CRNN for Manchu Word Recognition on a Woodblock-Printed Dataset. Proceedings of the Neural Information Processing: 29th International Conference, ICONIP 2022, Proceedings, Part II, Virtual."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"129374","DOI":"10.1016\/j.eswa.2025.129374","article-title":"SSC3: A novel structure-connected cognition cube network for Manchu word recognition","volume":"297","author":"Bi","year":"2025","journal-title":"Expert Syst. Appl."}],"container-title":["Symmetry"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/9\/1408\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:35:05Z","timestamp":1760034905000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2073-8994\/17\/9\/1408"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,8,29]]},"references-count":34,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2025,9]]}},"alternative-id":["sym17091408"],"URL":"https:\/\/doi.org\/10.3390\/sym17091408","relation":{},"ISSN":["2073-8994"],"issn-type":[{"value":"2073-8994","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,8,29]]}}}