{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T20:59:02Z","timestamp":1775595542291,"version":"3.50.1"},"reference-count":138,"publisher":"Association for Computing Machinery (ACM)","issue":"1","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. ACM Manag. Data"],"published-print":{"date-parts":[[2026,4,2]]},"abstract":"<jats:p>Modern knowledge and large volumes of data are increasingly encoded within neural networks, making the task of simplifying their structures and reducing the number of parameters especially relevant, both to improve efficiency and to facilitate deployment in resource-constrained environments. This paper presents a novel approach to neural network compression that addresses redundancy at both the filter and architectural levels through a unified framework grounded in information flow analysis. Building upon the concept of tensor flow divergence, which quantifies how information transforms across network layers, we develop a two-stage optimization process. The first stage employs iterative divergence-aware pruning to identify and remove redundant filters while preserving critical information pathways. The second stage extends this principle to higher-level architecture optimization by analyzing layer-wise contributions to information propagation and selectively eliminating entire layers that demonstrate minimal impact on network performance. The proposed method naturally adapts to diverse architectures, including convolutional networks, transformers, and hybrid designs, providing a consistent metric for comparing the structural importance across different layer types. Experimental validation across multiple modern architectures and datasets reveals that this combined approach achieves substantial model compression while maintaining competitive accuracy. The presented approach achieves parameter reduction results that are globally comparable to state-of-the-art solutions and outperform them across a wide range of modern neural network architectures, from convolutional models to transformers. The results demonstrate how flow divergence serves as an effective guiding principle for both filter-level and layer-level optimization, offering practical benefits for deployment in resource-constrained environments.<\/jats:p>","DOI":"10.1145\/3786659","type":"journal-article","created":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T17:54:13Z","timestamp":1775584453000},"page":"1-28","source":"Crossref","is-referenced-by-count":0,"title":["IDAP++: Advancing Divergence-Aware Pruning with Joint Filter and Layer Optimization"],"prefix":"10.1145","volume":"4","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6839-3558","authenticated-orcid":false,"given":"Aleksei","family":"Samarin","sequence":"first","affiliation":[{"name":"Wayy LLC, Miami, USA and ITMO University, St. Petersburg, Russian Federation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-8758-9383","authenticated-orcid":false,"given":"Artem","family":"Nazarenko","sequence":"additional","affiliation":[{"name":"Wayy LLC, Miami, USA and ITMO University, St. Petersburg, Russian Federation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-9835-0584","authenticated-orcid":false,"given":"Egor","family":"Kotenko","sequence":"additional","affiliation":[{"name":"Wayy LLC, Miami, USA and Saint Petersburg State University, St. Petersburg, Russian Federation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-8348-4731","authenticated-orcid":false,"given":"Alexander","family":"Savelev","sequence":"additional","affiliation":[{"name":"Wayy LLC, Miami, USA and ITMO University, St. Petersburg, Russian Federation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-7555-9446","authenticated-orcid":false,"given":"Aleksei","family":"Toropov","sequence":"additional","affiliation":[{"name":"Wayy LLC, Miami, USA and ITMO University, St. Petersburg, Russian Federation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4241-4298","authenticated-orcid":false,"given":"Alexandr","family":"Motyko","sequence":"additional","affiliation":[{"name":"Wayy LLC, Miami, Russian Federation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4508-2527","authenticated-orcid":false,"given":"Valentin","family":"Malykh","sequence":"additional","affiliation":[{"name":"Wayy LLC, Miami, USA and ITMO University, St. Petersburg, Russian Federation"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2026,4,7]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Shipeng Bai Jun Chen Xintian Shen Yixuan Qian and Yong Liu. 2023. Unified Data-Free Compression: Pruning and Quantization without Fine-Tuning. arXiv:2308.07209 [cs.LG] https:\/\/arxiv.org\/abs\/2308.07209"},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","unstructured":"Luciano Baresi and Giovanni Quattrocchi. 2022. Training and Serving Machine Learning Models at Scale. 669-683. doi:10.1007\/978-3-031-20984-0_48","DOI":"10.1007\/978-3-031-20984-0_48"},{"key":"e_1_2_1_3_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-021-82543-3"},{"key":"e_1_2_1_4_1","first-page":"446","volume-title":"Switzerland","author":"Bossard Lukas","year":"2014","unstructured":"Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. 2014. Food-101-mining discriminative components with random forests. In Computer vision-ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part VI 13. Springer, 446-461."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-025-92586-5"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","unstructured":"Tom Brown Benjamin Mann Nick Ryder Melanie Subbiah Jared Kaplan Prafulla Dhariwal Arvind Neelakantan Pranav Shyam Girish Sastry Amanda Askell Sandhini Agarwal Ariel Herbert-Voss Gretchen Krueger Tom Henighan Rewon Child Aditya Ramesh Daniel Ziegler Jeffrey Wu Clemens Winter and Dario Amodei. 2020. Language Models are Few-Shot Learners. doi:10.48550\/arXiv.2005.14165","DOI":"10.48550\/arXiv.2005.14165"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.70470\/KHWARIZMIA\/2023\/010"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00132"},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3486618"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","unstructured":"James Cameron Alexandra Sala Georgios Antoniou Paul Brennan Holly Butler Justin Conn Siobhan Connal Tom Curran Mark Hegarty Rose McHardy Daniel Orringer David Palmer Benjamin Smith and Matthew Baker. 2022. Multi-cancer early detection with a spectroscopic liquid biopsy platform. doi:10.21203\/rs.3.rs-1696059\/v1","DOI":"10.21203\/rs.3.rs-1696059\/v1"},{"key":"e_1_2_1_11_1","unstructured":"Pierre-Luc Carrier and Aaron Courville. 2013. FER-2013 Dataset. https:\/\/www.kaggle.com\/datasets\/msambare\/fer2013."},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3435937"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3447085"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","unstructured":"Ting-Wu Chin Ruizhou Ding Cha Zhang and Diana Marculescu. 2020. Towards Efficient Model Compression via Learned Global Ranking. 1515-1525. doi:10.1109\/CVPR42600.2020.00159","DOI":"10.1109\/CVPR42600.2020.00159"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.52783\/jes.3220"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","unstructured":"Tim Dettmers and Luke Zettlemoyer. 2022. The case for 4-bit precision: k-bit Inference Scaling Laws. doi:10.48550\/arXiv.2212.09720","DOI":"10.48550\/arXiv.2212.09720"},{"key":"e_1_2_1_18_1","unstructured":"Prafulla Dhariwal and Alex Nichol. 2021. Diffusion Models Beat GANs on Image Synthesis. arXiv:2105.05233 [cs.LG] https:\/\/arxiv.org\/abs\/2105.05233"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4471-6419-7_15"},{"key":"e_1_2_1_20_1","volume-title":"Jia Yu, Chi Wang, Jaeyoung Do, Yinan Li, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David Lomet, and Tim Kraska.","author":"Ding Jialin","year":"2021","unstructured":"Jialin Ding, Umar Farooq Minhas, Jia Yu, Chi Wang, Jaeyoung Do, Yinan Li, Hantian Zhang, Badrish Chandramouli, Johannes Gehrke, Donald Kossmann, David Lomet, and Tim Kraska. 2021. ALEX: An Updatable Adaptive Learned Index. (01 2021)."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the 9th International Conference on Learning Representations (ICLR). https:\/\/openreview.net\/forum?id=YicbFdNTTy","author":"Dosovitskiy Alexey","year":"2021","unstructured":"Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of the 9th International Conference on Learning Representations (ICLR). https:\/\/openreview.net\/forum?id=YicbFdNTTy"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01268"},{"key":"e_1_2_1_23_1","first-page":"4475","article-title":"Optimal brain compression: A framework for accurate post-training quantization and pruning","volume":"35","author":"Frantar Elias","year":"2022","unstructured":"Elias Frantar and Dan Alistarh. 2022. Optimal brain compression: A framework for accurate post-training quantization and pruning. Advances in Neural Information Processing Systems, Vol. 35 (2022), 4475-4488.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_24_1","unstructured":"Elias Frantar and Dan Alistarh. 2023. SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot. arXiv:2301.00774 [cs.LG] https:\/\/arxiv.org\/abs\/2301.00774"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","unstructured":"Shangqian Gao Feihu Huang Yanfu Zhang and Heng Huang. 2022. Disentangled Differentiable Network Pruning. 328-345. doi:10.1007\/978-3-031-20083-0_20","DOI":"10.1007\/978-3-031-20083-0_20"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","unstructured":"Amir Gholami Sehoon Kim Dong Zhen Zhewei Yao Michael Mahoney and Kurt Keutzer. 2022. A Survey of Quantization Methods for Efficient Neural Network Inference. 291-326. doi:10.1201\/9781003162810-13","DOI":"10.1201\/9781003162810-13"},{"key":"e_1_2_1_27_1","unstructured":"Aaron Grattafiori Abhimanyu Dubey Abhinav Jauhri et al. 2024. The Llama 3 Herd of Models. arXiv:2407.21783 [cs.AI] https:\/\/arxiv.org\/abs\/2407.21783"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2016.2582924"},{"key":"e_1_2_1_29_1","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Hao Yongchang","year":"2024","unstructured":"Yongchang Hao, Yanshuai Cao, and Lili Mou. 2024. FLORA: low-rank adapters are secretly gradient compressors. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML'24). JMLR.org, Article 700, 18 pages."},{"key":"e_1_2_1_30_1","volume-title":"Vanillakd: Revisit the power of vanilla knowledge distillation from small scale to large scale. arXiv preprint arXiv:2305.15781","author":"Hao Zhiwei","year":"2023","unstructured":"Zhiwei Hao, Jianyuan Guo, Kai Han, Han Hu, Chang Xu, and Yunhe Wang. 2023. Vanillakd: Revisit the power of vanilla knowledge distillation from small scale to large scale. arXiv preprint arXiv:2305.15781 (2023)."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICNN.1993.298572"},{"key":"e_1_2_1_32_1","volume-title":"Deep Residual Learning for Image Recognition. CoRR","author":"He Kaiming","year":"2015","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2015. Deep Residual Learning for Image Recognition. CoRR, Vol. abs\/1512.03385 (2015). arXiv:1512.03385 http:\/\/arxiv.org\/abs\/1512.03385"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.155"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","unstructured":"Benjamin Hilprecht Andreas Schmidt Moritz Kulessa Alejandro Molina Kristian Kersting and Carsten Binnig. 2019. DeepDB: Learn from Data not from Queries! doi:10.48550\/arXiv.1909.00607","DOI":"10.48550\/arXiv.1909.00607"},{"key":"e_1_2_1_35_1","unstructured":"Jonathan Ho Ajay Jain and Pieter Abbeel. 2020. Denoising Diffusion Probabilistic Models. arXiv:2006.11239 [cs.LG] https:\/\/arxiv.org\/abs\/2006.11239"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00140"},{"key":"e_1_2_1_37_1","unstructured":"Edward J. Hu Yelong Shen Phillip Wallis Zeyuan Allen-Zhu Yuanzhi Li Shean Wang Lu Wang and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685 [cs.CL] https:\/\/arxiv.org\/abs\/2106.09685"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3677052.3698696"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-25066-8_3"},{"key":"e_1_2_1_41_1","unstructured":"Albert Q. Jiang Alexandre Sablayrolles Antoine Roux Arthur Mensch Blanche Savary Chris Bamford Devendra Singh Chaplot Diego de las Casas Emma Bou Hanna Florian Bressand Gianna Lengyel Guillaume Bour Guillaume Lample L\u00e9lio Renard Lavaud Lucile Saulnier Marie-Anne Lachaux Pierre Stock Sandeep Subramanian Sophia Yang Szymon Antoniak Teven Le Scao Th\u00e9ophile Gervet Thibaut Lavril Thomas Wang Timoth\u00e9e Lacroix and William El Sayed. 2024. Mixtral of Experts. arXiv:2401.04088 [cs.LG] https:\/\/arxiv.org\/abs\/2401.04088"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CAC48633.2019.8997079"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-021-03819-2"},{"key":"e_1_2_1_44_1","volume-title":"Varshney","author":"Kafle Swatantra","year":"2025","unstructured":"Swatantra Kafle, Geethu Joseph, and Pramod K. Varshney. 2025. One-bit Compressed Sensing using Generative Models. arXiv:2502.12762 [cs.LG] https:\/\/arxiv.org\/abs\/2502.12762"},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-023-16740-9"},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00453"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3060406"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.compfluid.2013.06.001"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1109\/MCSE.2010.88"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCVW.2013.77"},{"key":"e_1_2_1_51_1","unstructured":"Alex Krizhevsky Geoffrey Hinton et al. 2009. Learning multiple layers of features from tiny images. (2009)."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.emnlp-main.279"},{"key":"e_1_2_1_53_1","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Kurtic Eldar","year":"2023","unstructured":"Eldar Kurtic, Elias Frantar, and Dan Alistarh. 2023. ZipLM: inference-aware structured pruning of language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS '23). Curran Associates Inc., Red Hook, NY, USA, Article 2863, 21 pages."},{"key":"e_1_2_1_54_1","volume-title":"Optimal brain damage. Advances in neural information processing systems","author":"LeCun Yann","year":"1989","unstructured":"Yann LeCun, John Denker, and Sara Solla. 1989. Optimal brain damage. Advances in neural information processing systems, Vol. 2 (1989)."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3587095"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/BigData55660.2022.10020252"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3581784.3607034"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ins.2024.121418"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","unstructured":"Hao Li Asim Kadav Igor Durdanovic Hanan Samet and H.P. Graf. 2016. Pruning Filters for Efficient ConvNets. (08 2016). doi:10.48550\/arXiv.1608.08710","DOI":"10.48550\/arXiv.1608.08710"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","unstructured":"Yang Lin Tianyu Zhang Peiqin Sun Zheng Li and Shuchang Zhou. 2021. FQ-ViT: Fully Quantized Vision Transformer without Retraining. doi:10.48550\/arXiv.2111.13824","DOI":"10.48550\/arXiv.2111.13824"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2024.3361989"},{"key":"e_1_2_1_62_1","volume-title":"Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models. arXiv:2402.17177 [cs.CV] https:\/\/arxiv.org\/abs\/2402.17177","author":"Liu Yixin","year":"2024","unstructured":"Yixin Liu, Kai Zhang, Yuan Li, Zhiling Yan, Chujie Gao, Ruoxi Chen, Zhengqing Yuan, Yue Huang, Hanchi Sun, Jianfeng Gao, Lifang He, and Lichao Sun. 2024b. Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models. arXiv:2402.17177 [cs.CV] https:\/\/arxiv.org\/abs\/2402.17177"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"e_1_2_1_64_1","unstructured":"Zhenhua Liu Yunhe Wang Kai Han Siwei Ma and Wen Gao. 2021. Post-Training Quantization for Vision Transformer. arXiv:2106.14156 [cs.CV] https:\/\/arxiv.org\/abs\/2106.14156"},{"key":"e_1_2_1_65_1","unstructured":"Zechun Liu Changsheng Zhao Forrest Iandola Chen Lai Yuandong Tian Igor Fedorov Yunyang Xiong Ernie Chang Yangyang Shi Raghuraman Krishnamoorthi Liangzhen Lai and Vikas Chandra. 2024c. MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases. arXiv:2402.14905 [cs.LG] https:\/\/arxiv.org\/abs\/2402.14905"},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","unstructured":"Artur Lopes and Rafael Ruggiero. 2021. Nonequilibrium in Thermodynamic Formalism: the Second Law gases and Information Geometry. doi:10.48550\/arXiv.2103.08333","DOI":"10.48550\/arXiv.2103.08333"},{"key":"e_1_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01264-9_8"},{"key":"e_1_2_1_68_1","unstructured":"Xinyin Ma Gongfan Fang and Xinchao Wang. 2023. LLM-Pruner: On the Structural Pruning of Large Language Models. arXiv:2305.11627 [cs.CL] https:\/\/arxiv.org\/abs\/2305.11627"},{"key":"e_1_2_1_69_1","unstructured":"Bennert Machenhauer and E. Rasmussen. 1972. On the integration of the spectral hydrodynamical equations by a transform method. (01 1972)."},{"key":"e_1_2_1_70_1","doi-asserted-by":"publisher","unstructured":"Nantia Makrynioti Ruy Ley-Wild and Vasilis Vassalos. 2021. Machine learning in SQL by translation to TensorFlow. 1-11. doi:10.1145\/3462462.3468879","DOI":"10.1145\/3462462.3468879"},{"key":"e_1_2_1_71_1","unstructured":"Gary Marcus Ernest Davis and Scott Aaronson. 2022. A very preliminary analysis of DALL-E 2. arXiv:2204.13807 [cs.CV] https:\/\/arxiv.org\/abs\/2204.13807"},{"key":"e_1_2_1_72_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2004.03814"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.14778\/3342263.3342644"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-020-2679-9"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06735-9"},{"key":"e_1_2_1_76_1","doi-asserted-by":"publisher","DOI":"10.1039\/d4dd90010c"},{"key":"e_1_2_1_77_1","unstructured":"Ari S. Morcos David G. T. Barrett Neil C. Rabinowitz and Matthew Botvinick. 2018. On the importance of single directions for generalization. arXiv:1803.06959 [stat.ML] https:\/\/arxiv.org\/abs\/1803.06959"},{"key":"e_1_2_1_78_1","unstructured":"Yao Mu Shoufa Chen Mingyu Ding Jianyu Chen Runjian Chen and Ping Luo. 2022. CtrlFormer: Learning Transferable State Representation for Visual Control via Transformer. arXiv:2206.08883 [cs.CV] https:\/\/arxiv.org\/abs\/2206.08883"},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","unstructured":"Parfait Munezero Mattias Villani and Robert Kohn. 2021. Dynamic Mixture of Experts Models for Online Prediction. doi:10.48550\/arXiv.2109.11449","DOI":"10.48550\/arXiv.2109.11449"},{"key":"e_1_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICVGIP.2008.47"},{"key":"e_1_2_1_81_1","unstructured":"OpenAI Josh Achiam Steven Adler Sandhini Agarwal Lama Ahmad Ilge Akkaya Florencia Aleman Diogo Almeida Janko Altenschmidt Sam Altman Shyamal Anadkat Red Avila Igor Babuschkin Suchir Balaji Valerie Balcom Paul Baltescu Haiming Bao Mohammad Bavarian Jeff Belgum and Barret Zoph. 2023. GPT-4 Technical Report. doi:10.48550\/arXiv.2303.08774"},{"key":"e_1_2_1_82_1","unstructured":"Maxime Oquab Timoth\u00e9e Darcet Th\u00e9o Moutakanni Huy Vo Marc Szafraniec Vasil Khalidov Pierre Fernandez Daniel Haziza Francisco Massa Alaaeldin El-Nouby Mahmoud Assran Nicolas Ballas Wojciech Galuba Russell Howes Po-Yao Huang Shang-Wen Li Ishan Misra Michael Rabbat Vasu Sharma Gabriel Synnaeve Hu Xu Herv\u00e9 Jegou Julien Mairal Patrick Labatut Armand Joulin and Piotr Bojanowski. 2024. DINOv2: Learning Robust Visual Features without Supervision. arXiv:2304.07193 [cs.CV] https:\/\/arxiv.org\/abs\/2304.07193"},{"key":"e_1_2_1_83_1","doi-asserted-by":"publisher","unstructured":"Joshua Osondu. 2025. Red AI vs. Green AI in Education: How Educational Institutions and Students Can Lead Environmentally Sustainable Artificial Intelligence Practices. doi:10.13140\/RG.2.2.27929.12644","DOI":"10.13140\/RG.2.2.27929.12644"},{"key":"e_1_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-20083-0_18"},{"key":"e_1_2_1_85_1","doi-asserted-by":"crossref","unstructured":"Omkar M Parkhi Andrea Vedaldi et al. 2012. Cats and dogs. CVPR (2012). https:\/\/www.robots.ox.ac.uk\/ vgg\/data\/pets\/","DOI":"10.1109\/CVPR.2012.6248092"},{"key":"e_1_2_1_86_1","doi-asserted-by":"publisher","unstructured":"David Patterson Joseph Gonzalez Urs H\u00f6lzle Quoc Le Chen Liang Llu\u00eds-Miquel Mungu\u00eda Daniel Rothchild David So Maud Texier and Jeffrey Dean. 2022. The Carbon Footprint of Machine Learning Training Will Plateau Then Shrink. doi:10.36227\/techrxiv.19139645.v2","DOI":"10.36227\/techrxiv.19139645.v2"},{"key":"e_1_2_1_87_1","unstructured":"Federico Nicol\u00e1s Peccia and Oliver Bringmann. 2024. Embedded Distributed Inference of Deep Neural Networks: A Systematic Review. arXiv:2405.03360 [cs.DC] https:\/\/arxiv.org\/abs\/2405.03360"},{"key":"e_1_2_1_88_1","unstructured":"Baolin Peng Chunyuan Li Pengcheng He Michel Galley and Jianfeng Gao. 2023. Instruction Tuning with GPT-4. arXiv:2304.03277 [cs.CL] https:\/\/arxiv.org\/abs\/2304.03277"},{"key":"e_1_2_1_89_1","doi-asserted-by":"publisher","DOI":"10.1063\/5.0152213"},{"key":"e_1_2_1_90_1","unstructured":"Aditya Ramesh Prafulla Dhariwal Alex Nichol Casey Chu and Mark Chen. 2022. Hierarchical Text-Conditional Image Generation with CLIP Latents. arXiv:2204.06125 [cs.CV] https:\/\/arxiv.org\/abs\/2204.06125"},{"key":"e_1_2_1_91_1","volume-title":"Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Doll\u00e1r, and Christoph Feichtenhofer.","author":"Ravi Nikhila","year":"2024","unstructured":"Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R\u00e4dle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Doll\u00e1r, and Christoph Feichtenhofer. 2024. SAM 2: Segment Anything in Images and Videos. arXiv:2408.00714 [cs.CV] https:\/\/arxiv.org\/abs\/2408.00714"},{"key":"e_1_2_1_92_1","doi-asserted-by":"publisher","unstructured":"Albert Reuther Peter Michaleas Michael Jones Vijay Gadepally Siddharth Samsi and Jeremy Kepner. 2021. AI Accelerator Survey and Trends. 1-9. doi:10.1109\/HPEC49654.2021.9622867","DOI":"10.1109\/HPEC49654.2021.9622867"},{"key":"e_1_2_1_93_1","unstructured":"Danilo Jimenez Rezende and Shakir Mohamed. 2016. Variational Inference with Normalizing Flows. arXiv:1505.05770 [stat.ML] https:\/\/arxiv.org\/abs\/1505.05770"},{"key":"e_1_2_1_94_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00181-022-02229-1"},{"key":"e_1_2_1_95_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2013.6638949"},{"key":"e_1_2_1_96_1","doi-asserted-by":"publisher","DOI":"10.1109\/lra.2023.3246839"},{"key":"e_1_2_1_97_1","volume-title":"IDAP: Advancing Divergence-Aware Pruning with Joint Filter and Layer Optimization. https:\/\/github.com\/user852154\/divergence_aware_pruning.","author":"Samarin Aleksei","year":"2025","unstructured":"Aleksei Samarin, Artem Nazarenko, Egor Kotenko, Alexander Savelev, Aleksei Toropov, Alexandr Motyko, and Valentin Malykh. 2025. IDAP: Advancing Divergence-Aware Pruning with Joint Filter and Layer Optimization. https:\/\/github.com\/user852154\/divergence_aware_pruning."},{"key":"e_1_2_1_98_1","doi-asserted-by":"publisher","unstructured":"Victor Sanh Lysandre Debut Julien Chaumond and Thomas Wolf. 2019. DistilBERT a distilled version of BERT: smaller faster cheaper and lighter. doi:10.48550\/arXiv.1910.01108","DOI":"10.48550\/arXiv.1910.01108"},{"key":"e_1_2_1_99_1","doi-asserted-by":"crossref","unstructured":"Andrew Saxe Yamini Bansal Joel Dapello Madhu Advani Artemy Kolchinsky Brendan Tracey and David Cox. 2018. On the information bottleneck theory of deep learning.","DOI":"10.1088\/1742-5468\/ab3985"},{"key":"e_1_2_1_100_1","doi-asserted-by":"publisher","unstructured":"Divya Saxena Jiannong Cao Jiahao Xu and Tarun Kulshrestha. 2024. RG-GAN: dynamic regenerative pruning for data-efficient generative adversarial networks. In Proceedings of the Thirty-Eighth AAAI Conference on Artificial Intelligence and Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence and Fourteenth Symposium on Educational Advances in Artificial Intelligence (AAAI'24\/IAAI'24\/EAAI'24). AAAI Press Article 523 9 pages. doi:10.1609\/aaai.v38i5.28271","DOI":"10.1609\/aaai.v38i5.28271"},{"key":"e_1_2_1_101_1","unstructured":"Ruoxi Shi Hansheng Chen Zhuoyang Zhang Minghua Liu Chao Xu Xinyue Wei Linghao Chen Chong Zeng and Hao Su. 2023. Zero123: a Single Image to Consistent Multi-view Diffusion Base Model. arXiv:2310.15110 [cs.CV] https:\/\/arxiv.org\/abs\/2310.15110"},{"key":"e_1_2_1_102_1","doi-asserted-by":"crossref","unstructured":"Kyuhong Shim Iksoo Choi Wonyong Sung and Jungwook Choi. 2021. Layer-wise Pruning of Transformer Attention Heads for Efficient Language Modeling. arXiv:2110.03252 [cs.CL] https:\/\/arxiv.org\/abs\/2110.03252","DOI":"10.1109\/ISOCC53507.2021.9613933"},{"key":"e_1_2_1_103_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cie.2018.03.039"},{"key":"e_1_2_1_104_1","doi-asserted-by":"publisher","unstructured":"Hayk Shoukourian Torsten Wilde Detlef Labrenz and Arndt Bode. 2017. Using Machine Learning for Data Center Cooling Infrastructure Efficiency Prediction. 954-963. doi:10.1109\/IPDPSW.2017.25","DOI":"10.1109\/IPDPSW.2017.25"},{"key":"e_1_2_1_105_1","unstructured":"Ravid Shwartz-Ziv. 2022. Information Flow in Deep Neural Networks. arXiv:2202.06749 [cs.LG] https:\/\/arxiv.org\/abs\/2202.06749"},{"key":"e_1_2_1_106_1","doi-asserted-by":"publisher","DOI":"10.3390\/electronics8060661"},{"key":"e_1_2_1_107_1","volume-title":"Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556","author":"Simonyan Karen","year":"2014","unstructured":"Karen Simonyan and Andrew Zisserman. 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 (2014)."},{"key":"e_1_2_1_108_1","doi-asserted-by":"publisher","unstructured":"Varun Sundar and Rajat Dwaraknath. 2021. [Reproducibility Report] Rigging the Lottery: Making All Tickets Winners. doi:10.48550\/arXiv.2103.15767","DOI":"10.48550\/arXiv.2103.15767"},{"key":"e_1_2_1_109_1","doi-asserted-by":"publisher","unstructured":"Mingxing Tan and Quoc Le. 2019a. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. doi:10.48550\/arXiv.1905.11946","DOI":"10.48550\/arXiv.1905.11946"},{"key":"e_1_2_1_110_1","volume-title":"Le","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc V. Le. 2019b. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. ArXiv, Vol. abs\/1905.11946 (2019). https:\/\/api.semanticscholar.org\/CorpusID:167217261"},{"key":"e_1_2_1_111_1","volume-title":"Proceedings of the 36th International Conference on Neural Information Processing Systems","author":"Tang Yehui","year":"2022","unstructured":"Yehui Tang, Kai Han, Jianyuan Guo, Chang Xu, Chao Xu, and Yunhe Wang. 2022. GhostNetV2: enhance cheap operation with long-range attention. In Proceedings of the 36th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS '22). Curran Associates Inc., Red Hook, NY, USA, Article 724, 14 pages."},{"key":"e_1_2_1_112_1","unstructured":"Gemini Team et al. 2024. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv:2403.05530 [cs.CL] https:\/\/arxiv.org\/abs\/2403.05530"},{"key":"e_1_2_1_113_1","volume-title":"Gemini: A Family of Highly Capable Multimodal Models. arXiv:2312.11805 [cs.CL] https:\/\/arxiv.org\/abs\/2312.11805","author":"Gemini Team","year":"2025","unstructured":"Gemini Team et al., 2025. Gemini: A Family of Highly Capable Multimodal Models. arXiv:2312.11805 [cs.CL] https:\/\/arxiv.org\/abs\/2312.11805"},{"key":"e_1_2_1_114_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2312.11805"},{"key":"e_1_2_1_115_1","doi-asserted-by":"publisher","DOI":"10.3390\/jimaging8030064"},{"key":"e_1_2_1_116_1","volume-title":"Contrastive representation distillation. arXiv preprint arXiv:1910.10699","author":"Tian Yonglong","year":"2019","unstructured":"Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2019. Contrastive representation distillation. arXiv preprint arXiv:1910.10699 (2019)."},{"key":"e_1_2_1_117_1","doi-asserted-by":"publisher","DOI":"10.1088\/1361-6404\/aadf9b"},{"key":"e_1_2_1_118_1","volume-title":"Bensen","author":"Tripp Charles Edison","year":"2024","unstructured":"Charles Edison Tripp, Jordan Perr-Sauer, Jamil Gafur, Amabarish Nag, Avi Purkayastha, Sagi Zisman, and Erik A. Bensen. 2024. Measuring the Energy Consumption and Efficiency of Deep Neural Networks: An Empirical Analysis and Design Recommendations. arXiv:2403.08151 [cs.LG] https:\/\/arxiv.org\/abs\/2403.08151"},{"key":"e_1_2_1_119_1","doi-asserted-by":"publisher","unstructured":"Shikhar Tuli and N.K. Jha. 2023. AccelTran: A Sparsity-Aware Accelerator for Dynamic Inference with Transformers. doi:10.48550\/arXiv.2302.14705","DOI":"10.48550\/arXiv.2302.14705"},{"key":"e_1_2_1_120_1","doi-asserted-by":"publisher","unstructured":"Kalim Ullah Hisham Alghamdi Ghulam Hafeez Imran Khan Safeer Ullah and Sadia Murawwat. 2024. A Swarm Intelligence-Based Approach for Multi-Objective Optimization Considering Renewable Energy in Smart Grid. 1-7. doi:10.1109\/ICECET61485.2024.10698431","DOI":"10.1109\/ICECET61485.2024.10698431"},{"key":"e_1_2_1_121_1","unstructured":"Aaron Van Den Oord Oriol Vinyals et al. 2017. Neural discrete representation learning. Advances in neural information processing systems Vol. 30 (2017)."},{"key":"e_1_2_1_122_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00914"},{"key":"e_1_2_1_123_1","doi-asserted-by":"publisher","DOI":"10.63471\/jbvada24001"},{"key":"e_1_2_1_124_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-025-98935-8"},{"key":"e_1_2_1_125_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00881"},{"key":"e_1_2_1_126_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCCN.2022.3149092"},{"key":"e_1_2_1_127_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41586-023-06887-8"},{"key":"e_1_2_1_128_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2111.00364"},{"key":"e_1_2_1_129_1","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2202.06258"},{"key":"e_1_2_1_130_1","doi-asserted-by":"publisher","unstructured":"Guangxuan Xiao Ji Lin Mickael Seznec Julien Demouth and Song Han. 2022. SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models. doi:10.48550\/arXiv.2211.10438","DOI":"10.48550\/arXiv.2211.10438"},{"key":"e_1_2_1_131_1","volume-title":"Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747","author":"Xiao Han","year":"2017","unstructured":"Han Xiao, Kashif Rasul, and Roland Vollgraf. 2017. Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747 (2017)."},{"key":"e_1_2_1_132_1","unstructured":"Hui Yang Sifu Yue and Yunzhong He. 2023. Auto-GPT for Online Decision Making: Benchmarks and Additional Opinions. arXiv:2306.02224 [cs.AI] https:\/\/arxiv.org\/abs\/2306.02224"},{"key":"e_1_2_1_133_1","volume-title":"Diffusion Models: A Comprehensive Survey of Methods and Applications. arXiv:2209.00796 [cs.LG] https:\/\/arxiv.org\/abs\/2209.00796","author":"Yang Ling","year":"2024","unstructured":"Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2024. Diffusion Models: A Comprehensive Survey of Methods and Applications. arXiv:2209.00796 [cs.LG] https:\/\/arxiv.org\/abs\/2209.00796"},{"key":"e_1_2_1_134_1","volume-title":"International Conference on Machine Learning. PMLR, 11875-11886","author":"Yao Zhewei","year":"2021","unstructured":"Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael Mahoney, et al., 2021. Hawq-v3: Dyadic neural network quantization. In International Conference on Machine Learning. PMLR, 11875-11886."},{"key":"e_1_2_1_135_1","unstructured":"Fanghua Yu Jinjin Gu Jinfan Hu Zheyuan Li and Chao Dong. 2025. UniCon: Unidirectional Information Flow for Effective Control of Large-Scale Diffusion Models. arXiv:2503.17221 [cs.CV] https:\/\/arxiv.org\/abs\/2503.17221"},{"key":"e_1_2_1_136_1","doi-asserted-by":"publisher","unstructured":"Ofir Zafrir Ariel Larey Guy Boudoukh Haihao Shen and Moshe Wasserblat. 2021. Prune Once for All: Sparse Pre-Trained Language Models. doi:10.48550\/arXiv.2111.05754","DOI":"10.48550\/arXiv.2111.05754"},{"key":"e_1_2_1_137_1","doi-asserted-by":"publisher","DOI":"10.31661\/gmj.v13i.3332"},{"key":"e_1_2_1_138_1","doi-asserted-by":"publisher","unstructured":"Jinjin Zhang Qiuyu Huang Junjie Liu Xiefan Guo and di Huang. 2025. Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models. doi:10.48550\/arXiv.2503.18352","DOI":"10.48550\/arXiv.2503.18352"}],"container-title":["Proceedings of the ACM on Management of Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3786659","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,7]],"date-time":"2026-04-07T19:59:14Z","timestamp":1775591954000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3786659"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,2]]},"references-count":138,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,4,2]]}},"alternative-id":["10.1145\/3786659"],"URL":"https:\/\/doi.org\/10.1145\/3786659","relation":{},"ISSN":["2836-6573"],"issn-type":[{"value":"2836-6573","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,2]]}}}