{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T01:33:43Z","timestamp":1760060023501,"version":"build-2065373602"},"reference-count":49,"publisher":"MDPI AG","issue":"8","license":[{"start":{"date-parts":[[2025,7,25]],"date-time":"2025-07-25T00:00:00Z","timestamp":1753401600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>Leveraging the ability of Vision Transformers (ViTs) to model contextual information across spatial patches, Masked Image Modeling (MIM) has emerged as a successful pre-training paradigm for visual representation learning by masking parts of the input and reconstructing the original image. However, this characteristic of ViTs has led many existing MIM methods to focus primarily on spatial patch reconstruction, overlooking the importance of semantic continuity in the channel dimension. Therefore, we propose a novel Masked Channel Modeling (MCM) pre-training paradigm, which reconstructs masked channel features using the contextual information from unmasked channels, thereby enhancing the model\u2019s understanding of images from the perspective of channel semantic continuity. Considering that traditional RGB reconstruction targets lack sufficient semantic attributes in the channel dimension, MCM introduces advanced features extracted by the CLIP image encoder as reconstruction targets. This guides the model to better capture semantic continuity across feature channels. Extensive experiments on downstream tasks, including image classification, object detection, and semantic segmentation, demonstrate the effectiveness and superiority of MCM. Our code will be available later.<\/jats:p>","DOI":"10.3390\/e27080794","type":"journal-article","created":{"date-parts":[[2025,7,25]],"date-time":"2025-07-25T14:40:02Z","timestamp":1753454402000},"page":"794","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Masked Channel Modeling Enables Vision Transformers to Learn Better Semantics"],"prefix":"10.3390","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-1718-5600","authenticated-orcid":false,"given":"Jiayi","family":"Chen","sequence":"first","affiliation":[{"name":"School of Telecommunications Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8472-1475","authenticated-orcid":false,"given":"Yanbiao","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-7023-3966","authenticated-orcid":false,"given":"Wei","family":"Dai","sequence":"additional","affiliation":[{"name":"School of Telecommunications Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7119-3215","authenticated-orcid":false,"given":"Zhihao","family":"Li","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,7,25]]},"reference":[{"key":"ref_1","unstructured":"Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2018). Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_2","unstructured":"Bao, H., Dong, L., Piao, S., and Wei, F. (2021). Beit: Bert pre-training of image transformers. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., and Girshick, R. (2022, January 18\u201322). Masked autoencoders are scalable vision learners. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_5","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ranftl, R. (2021, January 10\u201317). Vision transformers for dense prediction. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01196"},{"key":"ref_7","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2021, January 18\u201324). Training data-efficient image transformers & distillation through attention. Proceedings of the International Conference on Machine Learning (PMLR), Virtual."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Wei, C., Fan, H., Xie, S., Wu, C., Yuille, A., and Feichtenhofer, C. (2022, January 18\u201324). Masked feature prediction for self-supervised visual pre-training. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01426"},{"key":"ref_9","unstructured":"Radford, A., Kim, J., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning (PMLR), Virtual."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"2493","DOI":"10.1007\/s11263-024-01983-2","article-title":"Geometric prior guided feature representation learning for long-tailed classification","volume":"132","author":"Ma","year":"2024","journal-title":"Int. J. Comput. Vis."},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"54611","DOI":"10.1109\/ACCESS.2025.3554583","article-title":"Exploring Beyond Logits: Hierarchical Dynamic Labeling Based on Embeddings for Semi-Supervised Classification","volume":"13","author":"Chen","year":"2025","journal-title":"IEEE Access"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Dai, W., Ma, Y., Chen, J., Chen, X., and Li, S. (2025). Tradeoffs Between Richness and Bias of Augmented Data in Long-Tail Recognition. Entropy, 27.","DOI":"10.3390\/e27020201"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"3394","DOI":"10.1109\/TPAMI.2025.3534435","article-title":"Predicting and enhancing the fairness of DNNs with the curvature of perceptual manifolds","volume":"47","author":"Ma","year":"2025","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","doi-asserted-by":"crossref","first-page":"760","DOI":"10.1007\/s11263-024-02204-6","article-title":"Masked Channel Modeling for Bootstrapping Visual Pre-training","volume":"133","author":"Liu","year":"2025","journal-title":"Int. J. Comput. Vis."},{"key":"ref_15","unstructured":"Pham, C., Caicedo, J.C., and Plummer, B.A. (2025). ChA-MAEViT: Unifying Channel-Aware Masked Autoencoders and Multi-Channel Vision Transformers for Improved Cross-Channel Learning. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"1546","DOI":"10.1007\/s11263-023-01898-4","article-title":"Mimic before reconstruct: Enhancing masked autoencoders with feature mimicking","volume":"132","author":"Gao","year":"2024","journal-title":"Int. J. Comput. Vis."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft COCO: Common objects in context. Proceedings of the Computer Vision\u2014ECCV 2014: 13th European Conference, Zurich, Switzerland. Proceedings, Part V.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"He, K., Gkioxari, G., Doll\u00e1r, P., and Girshick, R. (2017, January 22\u201329). Mask R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.322"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Lin, T.-Y., Doll\u00e1r, P., Girshick, R., He, K., Hariharan, B., and Belongie, S. (2017, January 21\u201326). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.106"},{"key":"ref_21","unstructured":"Wu, Y., Kirillov, A., Massa, F., Lo, W.-Y., and Girshick, R. (November, January 27). Detectron2. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Li, Y., Mao, H., Girshick, R., and He, K. (2022, January 23\u201327). Exploring plain vision transformer backbones for object detection. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20077-9_17"},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"302","DOI":"10.1007\/s11263-018-1140-0","article-title":"Semantic understanding of scenes through the ade20k dataset","volume":"127","author":"Zhou","year":"2019","journal-title":"Int. J. Comput. Vis."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Xiao, T., Liu, Y., Zhou, B., Jiang, Y., and Sun, J. (2018, January 8). Unified perceptual parsing for scene understanding. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01228-1_26"},{"key":"ref_25","unstructured":"MMSegmentation Contributors (2025, July 14). MMSegmentation: OpenMMLab Semantic Segmentation Toolbox and Benchmark. Available online: https:\/\/github.com\/open-mmlab\/mmsegmentation."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Chen, X., Xie, S., and He, K. (2021, January 10\u201317). An empirical study of training self-supervised vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00950"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Caron, M., Touvron, H., Misra, I., J\u00e9gou, H., Mairal, J., Bojanowski, P., and Joulin, A. (2021, January 10\u201317). Emerging properties in self-supervised vision transformers. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Xie, Z., Zhang, Z., Cao, Y., Lin, Y., Bao, J., Yao, Z., Dai, Q., and Hu, H. (2022, January 18\u201324). Simmim: A simple framework for masked image modeling. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00943"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Wang, H., Tang, Y., Wang, Y., Guo, J., Deng, Z., and Han, K. (2023, January 17\u201324). Masked Image Modeling with Local Multi-Scale Reconstruction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00211"},{"key":"ref_30","first-page":"14290","article-title":"Semmae: Semantic-guided masking for learning masked autoencoders","volume":"35","author":"Li","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_31","unstructured":"Liu, Y., Zhang, S., Chen, J., Chen, K., and Lin, D. (2023). Pixmim: Rethinking pixel reconstruction in masked image modeling. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wei, L., Xie, L., Zhou, W., Li, H., and Tian, Q. (2022, January 23\u201327). Mvp: Multimodality-guided visual pre-training. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20056-4_20"},{"key":"ref_33","unstructured":"Hou, Z., Sun, F., Chen, Y., Xie, Y., and Kung, S. (2022). Milan: Masked image pretraining on language assisted representation. arXiv."},{"key":"ref_34","unstructured":"Zhou, J., Wei, C., Wang, H., Shen, W., Xie, C., Yuille, A., and Kong, T. (2021). ibot: Image bert pre-training with online tokenizer. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Dong, X., Bao, J., Zhang, T., Chen, D., Zhang, W., Yuan, L., Chen, D., Wen, F., and Yu, N. (2022, January 23\u201327). Bootstrapped masked autoencoders for vision BERT pretraining. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20056-4_15"},{"key":"ref_36","first-page":"35632","article-title":"Mcmae: Masked convolution meets masked autoencoders","volume":"35","author":"Gao","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_37","unstructured":"Ma, Y., Jiao, L., Liu, F., Yang, S., Liu, X., and Li, L. (November, January 29). Orthogonal uncertainty representation of data manifold for robust long-tailed learning. Proceedings of the 31st ACM International Conference on Multimedia (ACM MM), Ottawa, ON, Canada."},{"key":"ref_38","unstructured":"Ramesh, A., Pavlov, M., Goh, G., Gray, S., Voss, C., Radford, A., Chen, M., and Sutskever, I. (2021, January 18\u201324). Zero-shot text-to-image generation. Proceedings of the International Conference on Machine Learning (PMLR), Virtual."},{"key":"ref_39","unstructured":"Rolfe, J. (2016). Discrete variational autoencoders. arXiv."},{"key":"ref_40","unstructured":"Glorot, X., and Bengio, Y. (2010, January 13\u201315). Understanding the difficulty of training deep feedforward neural networks. Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics, Sardinia, Italy."},{"key":"ref_41","unstructured":"Goyal, P., Doll\u00e1r, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K. (2017). Accurate, large minibatch SGD: Training ImageNet in 1 hour. arXiv."},{"key":"ref_42","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_43","unstructured":"Chen, M., Radford, A., Child, R., Wu, J., Jun, H., Luan, D., and Sutskever, I. (2020, January 13\u201318). Generative pretraining from pixels. Proceedings of the International Conference on Machine Learning (PMLR), Virtual."},{"key":"ref_44","unstructured":"Loshchilov, I., and Hutter, F. (2016). Sgdr: Stochastic gradient descent with warm restarts. arXiv."},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Cubuk, E., Zoph, B., Shlens, J., and Le, Q. (2020, January 14\u201319). Randaugment: Practical automated data augmentation with a reduced search space. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, Seattle, WA, USA.","DOI":"10.1109\/CVPRW50498.2020.00359"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z. (2016, January 27\u201330). Rethinking the inception architecture for computer vision. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.308"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Zhang, H., Cisse, M., Dauphin, Y., and Lopez-Paz, D. (2017). mixup: Beyond empirical risk minimization. arXiv.","DOI":"10.1007\/978-1-4899-7687-1_79"},{"key":"ref_48","unstructured":"Yun, S., Han, D., Oh, S., Chun, S., Choe, J., and Yoo, Y. (November, January 27). Cutmix: Regularization strategy to train strong classifiers with localizable features. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_49","unstructured":"Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. (2016, January 11\u201314). Deep networks with stochastic depth. Proceedings of the Computer Vision\u2014ECCV 2016: 14th European Conference, Amsterdam, The Netherlands. Proceedings, Part IV 14."}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/8\/794\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:16:16Z","timestamp":1760033776000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/8\/794"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,25]]},"references-count":49,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2025,8]]}},"alternative-id":["e27080794"],"URL":"https:\/\/doi.org\/10.3390\/e27080794","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2025,7,25]]}}}