{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T00:51:24Z","timestamp":1760057484393,"version":"build-2065373602"},"reference-count":51,"publisher":"MDPI AG","issue":"2","license":[{"start":{"date-parts":[[2025,2,14]],"date-time":"2025-02-14T00:00:00Z","timestamp":1739491200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"China Postdoctoral Science Foundation (CPSF)","award":["GZC20232033","2023M742738","62406231"],"award-info":[{"award-number":["GZC20232033","2023M742738","62406231"]}]},{"name":"National Natural Science Foundation of China","award":["GZC20232033","2023M742738","62406231"],"award-info":[{"award-number":["GZC20232033","2023M742738","62406231"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>In long-tail scenarios, models have a very high demand for high-quality data. Information augmentation, as an important class of data-centric methods, has been proposed to improve model performance by expanding the richness and quantity of samples in tail classes. However, the underlying mechanisms behind the effectiveness of information augmentation methods remain underexplored. This has led to reliance on empirical and intricate fine-tuning in the use of information augmentation for long-tail recognition tasks. In this work, we simultaneously consider the richness gain and distribution shift introduced by information augmentation methods and propose effective information gain (EIG) to explore the mechanisms behind the effectiveness of these methods. We find that when the value of the effective information gain appropriately balances the richness gain and distribution shift, the performance of information augmentation methods is fully realized. Comprehensive experiments on long-tail benchmark datasets CIFAR-10-LT, CIFAR-100-LT, and ImageNet-LT demonstrate that using effective information gain to filter augmented data can further enhance model performance without any modifications to the model\u2019s architecture. Therefore, in addition to proposing new model architectures, data-centric approaches also hold significant potential in the field of long-tail recognition.<\/jats:p>","DOI":"10.3390\/e27020201","type":"journal-article","created":{"date-parts":[[2025,2,14]],"date-time":"2025-02-14T04:48:51Z","timestamp":1739508531000},"page":"201","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Tradeoffs Between Richness and Bias of Augmented Data in Long-Tail Recognition"],"prefix":"10.3390","volume":"27","author":[{"ORCID":"https:\/\/orcid.org\/0009-0004-7023-3966","authenticated-orcid":false,"given":"Wei","family":"Dai","sequence":"first","affiliation":[{"name":"School of Telecommunications Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8472-1475","authenticated-orcid":false,"given":"Yanbiao","family":"Ma","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-1718-5600","authenticated-orcid":false,"given":"Jiayi","family":"Chen","sequence":"additional","affiliation":[{"name":"School of Telecommunications Engineering, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xiaohua","family":"Chen","sequence":"additional","affiliation":[{"name":"Department of Automation, Tsinghua University, Beijing 100190, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuo","family":"Li","sequence":"additional","affiliation":[{"name":"School of Artificial Intelligence, Xidian University, Xi\u2019an 710071, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,2,14]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"1263","DOI":"10.1109\/TKDE.2008.239","article-title":"Learning from Imbalanced Data","volume":"21","author":"He","year":"2009","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"10795","DOI":"10.1109\/TPAMI.2023.3268118","article-title":"Deep long-tailed learning: A survey","volume":"45","author":"Zhang","year":"2023","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"6852","DOI":"10.1109\/TKDE.2022.3177226","article-title":"Distantly-Supervised Long-Tailed Relation Extraction Using Constraint Graphs","volume":"35","author":"Liang","year":"2023","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_4","first-page":"1761","article-title":"Learning Relation Prototype From Unlabeled Texts for Long-Tail Relation Extraction","volume":"35","author":"Cao","year":"2023","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"194","DOI":"10.1109\/TKDE.2019.2924656","article-title":"On Both Cold-Start and Long-Tail Recommendation with Social Data","volume":"33","author":"Li","year":"2021","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"2213","DOI":"10.1109\/TKDE.2023.3315284","article-title":"Tackling Long-tailed Distribution Issue in Graph Neural Networks via Normalization","volume":"36","author":"Liang","year":"2023","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"103901","DOI":"10.1016\/j.cviu.2023.103901","article-title":"Improving rare relation inferring for scene graph generation using bipartite graph network","volume":"239","author":"Lu","year":"2024","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"104146","DOI":"10.1016\/j.cviu.2024.104146","article-title":"Class Probability Space Regularization for semi-supervised semantic segmentation","volume":"249","author":"Yin","year":"2024","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"103206","DOI":"10.1016\/j.cviu.2021.103206","article-title":"Class knowledge overlay to visual feature learning for zero-shot image classification","volume":"207","author":"Xie","year":"2021","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_10","unstructured":"Ma, Y., Dai, W., and Chen, J. (2025, January 24\u201328). Pursuing Better Decision Boundaries for Long-Tailed Object Detection via Category Information Amount. Proceedings of the The Thirteenth International Conference on Learning Representations, Singapore."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Ren, N., Li, X., Wu, Y., and Fu, Y. (2024). Mixed Mutual Transfer for Long-Tailed Image Classificationy. Entropy, 26.","DOI":"10.3390\/e26100839"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, Z., Miao, Z., Zhan, X., Wang, J., Gong, B., and Yu, S.X. (2019, January 15\u201320). Large-Scale Long-Tailed Recognition in an Open World. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00264"},{"key":"ref_13","doi-asserted-by":"crossref","first-page":"6963","DOI":"10.1109\/TNNLS.2022.3213522","article-title":"Label-aware distribution calibration for long-tailed classification","volume":"35","author":"Wang","year":"2022","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_14","unstructured":"Kang, B., Xie, S., Rohrbach, M., Yan, Z., Gordo, A., Feng, J., and Kalantidis, Y. (2019). Decoupling representation and classifier for long-tailed recognition. arXiv."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Alshammari, S., Wang, Y.X., Ramanan, D., and Kong, S. (2022, January 18\u201324). Long-tailed recognition via weight balancing. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00677"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wang, T., Li, Y., Kang, B., Li, J., Liew, J., Tang, S., Hoi, S., and Feng, J. (2020, January 23\u201328). The devil is in classification: A simple framework for long-tail instance segmentation. Proceedings of the Computer Vision\u2014ECCV 2020: 16th European Conference, Glasgow, UK. Part XIV 16.","DOI":"10.1007\/978-3-030-58568-6_43"},{"key":"ref_17","unstructured":"Ma, Y., Jiao, L., Liu, F., Li, Y., Yang, S., and Liu, X. (2023, January 1\u20135). Delving into Semantic Scale Imbalance. Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Ma, Y., Jiao, L., Liu, F., Yang, S., Liu, X., and Li, L. (2023, January 18\u201322). Curvature-Balanced Feature Manifold Learning for Long-Tailed Classification. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01519"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Cui, Y., Jia, M., Lin, T.Y., Song, Y., and Belongie, S. (2019, January 15\u201320). Class-balanced loss based on effective number of samples. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00949"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Goyal, P., Girshick, R., He, K., and Doll\u00e1r, P. (2017, January 22\u201329). Focal loss for dense object detection. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.324"},{"key":"ref_21","first-page":"4175","article-title":"Balanced meta-softmax for long-tailed visual recognition","volume":"33","author":"Ren","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_22","unstructured":"Sinha, S., Ohashi, H., and Nakamura, K. (December, January 30). Class-wise difficulty-balanced loss for solving class-imbalance. Proceedings of the Asian Conference on Computer Vision, Kyoto, Japan."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Tan, J., Wang, C., Li, B., Li, Q., Ouyang, W., Yin, C., and Yan, J. (2020, January 13\u201319). Equalization loss for long-tailed object recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01168"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"214","DOI":"10.1109\/TKDE.2018.2826011","article-title":"Adaptive cost-sensitive online classification","volume":"31","author":"Zhao","year":"2018","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Chu, P., Bian, X., Liu, S., and Ling, H. (2020, January 23\u201328). Feature space augmentation for long-tailed data. Proceedings of the Computer Vision\u2014ECCV 2020: 16th European Conference, Glasgow, UK. Part XXIX 16.","DOI":"10.1007\/978-3-030-58526-6_41"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Park, S., Hong, Y., Heo, B., Yun, S., and Choi, J.Y. (2022, January 18\u201324). The majority can help the minority: Context-rich minority oversampling for long-tailed classification. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00676"},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"2772","DOI":"10.1109\/TMM.2023.3303697","article-title":"Feature distribution representation learning based on knowledge transfer for long-tailed classification","volume":"26","author":"Ma","year":"2023","journal-title":"IEEE Trans. Multimed."},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"669","DOI":"10.1038\/s42256-022-00516-1","article-title":"Advances, challenges and opportunities in creating data for trustworthy AI","volume":"4","author":"Liang","year":"2022","journal-title":"Nat. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., and Ringshia, P. (2021). Dynabench: Rethinking benchmarking in NLP. arXiv.","DOI":"10.18653\/v1\/2021.naacl-main.324"},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"100336","DOI":"10.1016\/j.patter.2021.100336","article-title":"Data and its (dis) contents: A survey of dataset development and use in machine learning research","volume":"2","author":"Paullada","year":"2021","journal-title":"Patterns"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"559","DOI":"10.1038\/s41592-018-0083-2","article-title":"Optimal experimental design","volume":"15","author":"Smucker","year":"2018","journal-title":"Nat. Methods"},{"key":"ref_32","first-page":"115","article-title":"A survey on bias and fairness in machine learning","volume":"54","author":"Mehrabi","year":"2021","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_33","unstructured":"Gontijo-Lopes, R., Smullin, S., Cubuk, E.D., and Dyer, E. (2021, January 3\u20137). Tradeoffs in Data Augmentation: An Empirical Study. Proceedings of the International Conference on Learning Representations, Virtual."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1140","DOI":"10.1137\/21M1437056","article-title":"Model-centric data manifold: The data through the eyes of the model","volume":"15","author":"Grementieri","year":"2022","journal-title":"SIAM J. Imaging Sci."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"103270","DOI":"10.1016\/j.cviu.2021.103270","article-title":"Learning to teach and learn for semi-supervised few-shot image classification","volume":"212","author":"Li","year":"2021","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"103582","DOI":"10.1016\/j.cviu.2022.103582","article-title":"Balanced softmax cross-entropy for incremental learning with and without memory","volume":"225","author":"Jodelet","year":"2022","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Yin, X., Yu, X., Sohn, K., Liu, X., and Chandraker, M. (2019, January 15\u201320). Feature transfer learning for face recognition with under-represented data. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00585"},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Liu, J., Sun, Y., Han, C., Dou, Z., and Li, W. (2020, January 13\u201319). Deep representation learning on long-tailed data: A learnable embedding augmentation perspective. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00304"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Liu, B., Li, H., Kang, H., Hua, G., and Vasconcelos, N. (2021, January 11\u201317). Gistnet: A geometric structure transfer network for long-tailed recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00810"},{"key":"ref_40","unstructured":"Ahn, S., Ko, J., and Yun, S.Y. (2023, January 1\u20135). CUDA: Curriculum of Data Augmentation for Long-tailed Recognition. Proceedings of the The Eleventh International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Zhong, Z., Cui, J., Liu, S., and Jia, J. (2021, January 20\u201325). Improving calibration for long-tailed recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01622"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Chou, H.P., Chang, S.C., Pan, J.Y., Wei, W., and Juan, D.C. (2020, January 23\u201328). Remix: Rebalanced mixup. Proceedings of the Computer Vision\u2013ECCV 2020 Workshops, Glasgow, UK. Part VI 16.","DOI":"10.1007\/978-3-030-65414-6_9"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Wei, X.S., Zhou, B., and Wu, J. (2021, January 2\u20139). Bag of tricks for long-tailed visual recognition with deep convolutional neural networks. Proceedings of the AAAI Conference on Artificial Intelligence, Virtual.","DOI":"10.1609\/aaai.v35i4.16458"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Zang, Y., Huang, C., and Loy, C.C. (2021, January 11\u201317). Fasa: Feature augmentation and sampling adaptation for long-tailed instance segmentation. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00344"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Li, S., Gong, K., Liu, C.H., Wang, Y., Qiao, F., and Cheng, X. (2021, January 20\u201325). Metasaug: Meta semantic augmentation for long-tailed visual recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00517"},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Kim, J., Jeong, J., and Shin, J. (2020, January 13\u201319). M2m: Imbalanced classification via major-to-minor translation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01391"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Xie, C., Tan, M., Gong, B., Wang, J., Yuille, A.L., and Le, Q.V. (2020, January 13\u201319). Adversarial examples improve image recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00090"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Lu, Y., Qian, P., Huang, G., and Wang, H. (2023, January 4\u201310). Personalized Federated Learning on Long-Tailed Data via Adversarial Feature Augmentation. Proceedings of the ICASSP 2023\u20142023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece.","DOI":"10.1109\/ICASSP49357.2023.10097083"},{"key":"ref_49","unstructured":"Zada, S., Benou, I., and Irani, M. (2022, January 17\u201323). Pure noise to the rescue of insufficient data: Improving imbalanced classification by training on random noise images. Proceedings of the International Conference on Machine Learning, PMLR, Baltimore, MD, USA."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Ma, Y., Jiao, L., Liu, F., Yang, S., Liu, X., and Li, L. (2023). Orthogonal Uncertainty Representation of Data Manifold for Robust Long-Tailed Learning. arXiv.","DOI":"10.1145\/3581783.3611698"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"361","DOI":"10.1016\/j.eng.2019.09.010","article-title":"A geometric understanding of deep learning","volume":"6","author":"Lei","year":"2020","journal-title":"Engineering"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/2\/201\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T16:34:21Z","timestamp":1760027661000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/27\/2\/201"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,14]]},"references-count":51,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["e27020201"],"URL":"https:\/\/doi.org\/10.3390\/e27020201","relation":{},"ISSN":["1099-4300"],"issn-type":[{"type":"electronic","value":"1099-4300"}],"subject":[],"published":{"date-parts":[[2025,2,14]]}}}