{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T23:59:01Z","timestamp":1768867141359,"version":"3.49.0"},"reference-count":29,"publisher":"Association for Computing Machinery (ACM)","issue":"1","funder":[{"name":"NSF","award":["2231623"],"award-info":[{"award-number":["2231623"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2026,2,28]]},"abstract":"<jats:p>In this article, a novel hierarchical deep neural network (DNN) is introduced that augments an input DNN to significantly enhance its image classification accuracy while reducing the inference time and the hardware overhead. The architecture comprises a hybrid framework that combines binary classifiers based on Convolutional Neural Networks (CNNs) with refined classifiers employing Vision Transformers (ViTs). A distinctive training approach is employed, where embedded models are designed and trained based on image distributions processed by binary classifiers, enhancing the system\u2019s precision and efficiency. An algorithm determines the optimal inclusion of components within a cascading structure, enabling the construction, training, and deployment of specialized deep-learning networks. Additionally, two algorithms are introduced to optimize the architecture for multi-GPU systems. Extensive experimentation across multiple baseline DNNs, including both CNNs and ViTs, and diverse datasets demonstrates the versatility and superiority of our proposed structure over traditional methods, with particularly strong improvements observed on larger and more complex datasets such as ImageNet.<\/jats:p>","DOI":"10.1145\/3779421","type":"journal-article","created":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T14:49:49Z","timestamp":1765205389000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Improved Image Classification using Lightweight Deep Neural Network Enhancements"],"prefix":"10.1145","volume":"17","author":[{"ORCID":"https:\/\/orcid.org\/0009-0003-2930-7911","authenticated-orcid":false,"given":"Vasilis","family":"Pentsos","sequence":"first","affiliation":[{"name":"School of Electrical, Computer and Biomedical Engineering, Southern Illinois University Carbondale, Carbondale, Illinois, USA and Electrical and Computer Engineering and Computer Science, University of Detroit Mercy, Detroit, Michigan, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2575-3588","authenticated-orcid":false,"given":"Spyros","family":"Tragoudas","sequence":"additional","affiliation":[{"name":"School of Electrical, Computer and Biomedical Engineering, Southern Illinois University Carbondale, Carbondale, Illinois, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-1348-2618","authenticated-orcid":false,"given":"Kiriti","family":"Nagesh Gowda","sequence":"additional","affiliation":[{"name":"ML Computer Vision Group, Advanced Micro Devices, Inc., Santa Clara, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-9852-764X","authenticated-orcid":false,"given":"Mike","family":"Schmit","sequence":"additional","affiliation":[{"name":"ML Computer Vision Group, Advanced Micro Devices, Inc., Santa Clara, California, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,1,19]]},"reference":[{"key":"e_1_3_1_2_2","volume-title":"End-to-End Text Recognition with Deep Learning Architectures","author":"Alsharif Ouais","year":"2015","unstructured":"Ouais Alsharif. 2015. End-to-End Text Recognition with Deep Learning Architectures. McGill University (Canada)."},{"key":"e_1_3_1_3_2","volume-title":"Proceedings of the IEEE International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE Conference on Robotics, Automation and Mechatronics (RAM)","author":"Bu Wei","year":"2017","unstructured":"Wei Bu, Jiangjian Xiao, Chuanhong Zhou, Minmin Yang, and Chengbin Peng. 2017. A cascade framework for masked face detection. In Proceedings of the IEEE International Conference on Cybernetics and Intelligent Systems (CIS) and IEEE Conference on Robotics, Automation and Mechatronics (RAM)."},{"key":"e_1_3_1_4_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00644"},{"key":"e_1_3_1_5_2","first-page":"32","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"Cao Kaidi","year":"2019","unstructured":"Kaidi Cao, Colin Wei, Adrien Gaidon, Nuno Arechiga, and Tengyu Ma. 2019. Learning imbalanced datasets with label-distribution-aware margin loss. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 32."},{"key":"e_1_3_1_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2022.3219832"},{"key":"e_1_3_1_7_2","first-page":"95","volume-title":"European Conference on Computer Vision (ECCV)","author":"Chou Yi-Lin","year":"2020","unstructured":"Yi-Lin Chou, Chih-Yuan Chuang, and Antonio Torralba. 2020. Remix: Rebalanced mixup. In European Conference on Computer Vision (ECCV). Springer, 95\u2013110."},{"key":"e_1_3_1_8_2","unstructured":"Charles Corbi\u00e8re Nicolas Thome Avner Bar-Hen Matthieu Cord and Patrick P\u00e9rez. 2019. Addressing failure prediction by learning model confidence. arXiv:1910.04851. Retrieved from https:\/\/arxiv.org\/abs\/1910.04851"},{"key":"e_1_3_1_9_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16\u2009\u00d7\u200916 words: Transformers for image recognition at scale. arXiv:2010.11929. Retrieved from https:\/\/arxiv.org\/abs\/2010.11929"},{"key":"e_1_3_1_10_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.476"},{"key":"e_1_3_1_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_1_12_2","volume-title":"Proceedings of International Conference on Learning Representations","author":"Hendrycks Dan","year":"2017","unstructured":"Dan Hendrycks and Kevin Gimpel. 2017. A baseline for detecting misclassified and out-of-distribution examples in neural networks. In Proceedings of International Conference on Learning Representations."},{"key":"e_1_3_1_13_2","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW)","author":"Hong Yaqian","year":"2023","unstructured":"Yaqian Hong, Zhaohui Li, Zhenhua Gu, Xu Xie, and Jae-Hyun Choi. 2023. Enhancing classification accuracy on limited data via unconditional GAN. In Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops (ICCVW)."},{"key":"e_1_3_1_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.01048"},{"key":"e_1_3_1_15_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Kang Bingyi","year":"2020","unstructured":"Bingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan, Albert Gordo, Jiashi Feng, and Yannis Kalantidis. 2020. Decoupling representation and classifier for long-tailed recognition. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_1_16_2","volume-title":"Advances in Neural Information Processing Systems (NeurIPS)","author":"LaBonte Lucas","year":"2023","unstructured":"Lucas LaBonte, Ziyang Li, Lu Yuan, Baoyuan Yang, and Jianfeng Wang. 2023. SELF: Selective fine-tuning for fairness and robustness. In Advances in Neural Information Processing Systems (NeurIPS)."},{"key":"e_1_3_1_17_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7299170"},{"key":"e_1_3_1_18_2","volume-title":"Proceedings of the 4th International Conference on Learning Representations (ICLR)","author":"Mishkin Dmytro","year":"2016","unstructured":"Dmytro Mishkin and Jiri Matas. 2016. All you need is a good init. In Proceedings of the 4th International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_1_19_2","first-page":"988","volume-title":"Proceedings of the 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA)","author":"Pentsos Vasileios","year":"2021","unstructured":"Vasileios Pentsos, Bijay Raj Paudel, Spyros Tragoudas, Kiriti Nagesh Gowda, and Mike Schmit. 2021. Improved CNN classification accuracy with the addition of shallow cascading CNNs. In Proceedings of the 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA). IEEE, 988\u2013991."},{"issue":"19","key":"e_1_3_1_20_2","article-title":"A multitask cascading CNN with multiscale infrared optical flow feature fusion-based abnormal crowd behavior monitoring UAV","volume":"20","author":"Shao Yanhua","year":"2020","unstructured":"Yanhua Shao, Wenfeng Li, Hongyu Chu, Zhiyuan Chang, Xiaoqiang Zhang, and Huayi Zhan. 2020. A multitask cascading CNN with multiscale infrared optical flow feature fusion-based abnormal crowd behavior monitoring UAV. Sensors 20, 19 (2020), 5550.","journal-title":"Sensors"},{"key":"e_1_3_1_21_2","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv:1409.1556. Retrieved from https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"e_1_3_1_22_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Advances in Neural Information Processing Systems, Vol. 30.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_1_23_2","unstructured":"Qin Xu Sitong Li Jiahui Wang Bo Jiang and Jinhui Tang. 2024. Context-semantic quality awareness network for fine-grained visual categorization. arXiv:2403.10298. Retrieved from https:\/\/arxiv.org\/abs\/2403.10298"},{"key":"e_1_3_1_24_2","volume-title":"Proceedings of the IEEE International Conference on Computer Vision","author":"Yan Zhicheng","year":"2015","unstructured":"Zhicheng Yan, Hao Zhang, Robinson Piramuthu, Vignesh Jagadeesh, Dennis DeCoste, Wei Di, and Yizhou Yu. 2015. HD-CNN: Hierarchical deep convolutional neural networks for large scale visual recognition. In Proceedings of the IEEE International Conference on Computer Vision."},{"key":"e_1_3_1_25_2","unstructured":"Jinghan Yang Anupam Pani and Yunchao Zhang. 2025. How to achieve higher accuracy with less training points? arXiv:2504.13586. Retrieved from https:\/\/arxiv.org\/abs\/2504.13586"},{"key":"e_1_3_1_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.557"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00515"},{"key":"e_1_3_1_28_2","volume-title":"Proceedings of the International Conference on Learning Representations (ICLR)","author":"Zhou Linyi","year":"2024","unstructured":"Linyi Zhou, Sihui Liu, Shixiang Zhai, and Chen Sun. 2024. Implicit adversarial data augmentation for robust learning under distribution shifts. In Proceedings of the International Conference on Learning Representations (ICLR)."},{"key":"e_1_3_1_29_2","first-page":"6898","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Zhou Wenxuan","year":"2022","unstructured":"Wenxuan Zhou, Shuo Wang, Lu Jiang, and Yi Yang. 2022. Balanced-mixup for long-tailed image classification. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 6898\u20136907."},{"key":"e_1_3_1_30_2","unstructured":"Xinqi Zhu and Michael Bain. 2017. B-CNN: Branch convolutional neural network for hierarchical classification. arXiv:1709.09890. Retrieved from https:\/\/arxiv.org\/abs\/1709.09890"}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3779421","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,1,19]],"date-time":"2026-01-19T11:29:22Z","timestamp":1768822162000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3779421"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,1,19]]},"references-count":29,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,2,28]]}},"alternative-id":["10.1145\/3779421"],"URL":"https:\/\/doi.org\/10.1145\/3779421","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,1,19]]},"assertion":[{"value":"2024-09-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-11-05","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-01-19","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}