{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,17]],"date-time":"2026-03-17T05:32:12Z","timestamp":1773725532143,"version":"3.50.1"},"reference-count":30,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2022,9,30]],"date-time":"2022-09-30T00:00:00Z","timestamp":1664496000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Embed. Comput. Syst."],"published-print":{"date-parts":[[2022,9,30]]},"abstract":"<jats:p>\n            We propose a framework for Class-aware Personalized Neural Network Inference (CAP\u2019NN), which prunes an already-trained neural network model based on the preferences of individual users. Specifically, by adapting to the subset of output classes that each user is expected to encounter, CAP\u2019NN is able to prune not only ineffectual neurons but also\n            <jats:italic>miseffectual<\/jats:italic>\n            neurons that confuse classification, without the need to retrain the network. CAP\u2019NN also exploits the similarities among pruning requests from different users to minimize the timing overheads of pruning the network. To achieve this, we propose a clustering algorithm that groups similar classes in the network based on the firing rates of neurons for each class and then implement a lightweight cache architecture to store and reuse information from previously pruned networks. In our experiments with VGG-16, AlexNet, and ResNet-152 networks, CAP\u2019NN achieves, on average, up to 47% model size reduction while actually\n            <jats:italic>improving<\/jats:italic>\n            the top-1(5) classification accuracy by up to 3.9%(3.4%) when the user only encounters a subset of the trained classes in these networks.\n          <\/jats:p>","DOI":"10.1145\/3520126","type":"journal-article","created":{"date-parts":[[2022,3,21]],"date-time":"2022-03-21T12:37:55Z","timestamp":1647866275000},"page":"1-24","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["CAP\u2019NN: A Class-aware Framework for Personalized Neural Network Inference"],"prefix":"10.1145","volume":"21","author":[{"given":"Maedeh","family":"Hemmat","sequence":"first","affiliation":[{"name":"University of Wisconsin\u2013Madison, Madison, Wisconsin, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Joshua San","family":"Miguel","sequence":"additional","affiliation":[{"name":"University of Wisconsin\u2013Madison, Madison, Wisconsin, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Azadeh","family":"Davoodi","sequence":"additional","affiliation":[{"name":"University of Wisconsin\u2013Madison, Madison, Wisconsin, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,12,9]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/MICRO.2012.26"},{"key":"e_1_3_2_3_2","article-title":"A clustering technique based on Elbow method and k-means in WSN","volume":"105","author":"Bholowalia P.","year":"2014","unstructured":"P. Bholowalia and A. Kumar. 2014. A clustering technique based on Elbow method and k-means in WSN. Int. J. Comput. Appl. 105, 9 (2014), 17\u201324.","journal-title":"Int. J. Comput. Appl."},{"issue":"1","key":"e_1_3_2_4_2","first-page":"L3","article-title":"Understanding the limitations of existing energy-efficient design approaches for deep neural networks","volume":"2","author":"Chen Yu-Hsin","year":"2018","unstructured":"Yu-Hsin Chen, Tien-Ju Yang, Joel Emer, and Vivienne Sze. 2018. Understanding the limitations of existing energy-efficient design approaches for deep neural networks. Energy 2, L1 (2018), L3.","journal-title":"Energy"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3316781.3317792"},{"key":"e_1_3_2_6_2","first-page":"598","volume-title":"Proceedings of the Conference on Neural Information Processing Systems","author":"Cun Y. Le","year":"1990","unstructured":"Y. Le Cun, J. S. Denker, and S. A. Solla. 1990. Optimal brain damage. In Proceedings of the Conference on Neural Information Processing Systems. 598\u2013605."},{"key":"e_1_3_2_7_2","first-page":"1269","volume-title":"Proceedings of the Conference on Neural Information Processing Systems","author":"Denton E. L.","year":"2014","unstructured":"E. L. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus. 2014. Exploiting linear structure within convolutional networks for efficient evaluation. In Proceedings of the Conference on Neural Information Processing Systems. 1269\u20131277."},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCAD.2012.2185930"},{"key":"e_1_3_2_9_2","first-page":"113","volume-title":"Proceedings of the Computer Vision and Pattern Recognition Conference","author":"Guo J.","year":"2017","unstructured":"J. Guo and M. Potkonjak. 2017. Pruning ConvNets online for efficient specialist model pruning. In Proceedings of the Computer Vision and Pattern Recognition Conference. 113\u2013120."},{"key":"e_1_3_2_10_2","first-page":"1135","volume-title":"Proceedings of the Conference on Neural Information Processing Systems","author":"Han S.","year":"2015","unstructured":"S. Han, J. Pool, J. Tran, and W. Dally. 2015. Learning both weights and connections for efficient neural network. In Proceedings of the Conference on Neural Information Processing Systems. 1135\u20131143."},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICNN.1993.298572"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.155"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1109\/DAC18072.2020.9218741"},{"key":"e_1_3_2_14_2","unstructured":"LeaderGPU. [n.d.]. Retrieved from https:\/\/www.leadergpu.com\/catalog\/tensorflow."},{"key":"e_1_3_2_15_2","unstructured":"Hengyuan Hu Rui Peng Y. W. Tai and Ch. Tang. 2016. Network trimming: A data-driven neuron pruning approach towards efficient deep architectures. Retrieved from https:\/\/arXiv:1607.03250."},{"key":"e_1_3_2_16_2","first-page":"1061","volume-title":"Proceedings of the International Conference on Acoustics, Speech, and Signal Processing","author":"Hu Y. H.","year":"1991","unstructured":"Y. H. Hu, Qiuzhen Xue, and W. J. Tompkins. 1991. Structural simplification of a feed-forward, multi-layer perceptron artificial neural network. In Proceedings of the International Conference on Acoustics, Speech, and Signal Processing. 1061\u20131064."},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3079856.3080246"},{"key":"e_1_3_2_18_2","unstructured":"Y. D. Kim Eunhyeok Park Sungjoo Yoo Taelim Choi Lu Yang and Dongjun Shin. 2015. Compression of deep convolutional neural networks for fast and low power mobile applications. Retrieved from https:\/\/arXiv:1511.06530."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.1991.155331"},{"key":"e_1_3_2_20_2","unstructured":"Hao Li Asim Kadav Igor Durdanovic Hanan Samet and Hans Peter Graf. 2016. Pruning filters for efficient convnets. Retrieved from https:\/\/arXiv:1608.08710."},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.541"},{"key":"e_1_3_2_22_2","unstructured":"Pavlo Molchanov Stephen Tyree Tero Karras Timo Aila and Jan Kautz. 2017. Pruning convolutional neural networks for resource efficient inference. Retrieved from https:\/\/arXiv:1611.06440."},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/3287624.3287722"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3287624.3287643"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.23919\/DATE.2017.7927280"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/ISCA.2016.12"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1145\/3287624.3287663"},{"key":"e_1_3_2_28_2","first-page":"2074","volume-title":"Proceedings of the Conference on Neural Information Processing Systems","author":"Wen Wei","year":"2016","unstructured":"Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li. 2016. Learning structured sparsity in deep neural networks. In Proceedings of the Conference on Neural Information Processing Systems. 2074\u20132082."},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/GLOCOM.2018.8647616"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/JETCAS.2018.2833383"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1145\/2684746.2689060"}],"container-title":["ACM Transactions on Embedded Computing Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3520126","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3520126","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T18:10:32Z","timestamp":1750183832000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3520126"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,30]]},"references-count":30,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2022,9,30]]}},"alternative-id":["10.1145\/3520126"],"URL":"https:\/\/doi.org\/10.1145\/3520126","relation":{},"ISSN":["1539-9087","1558-3465"],"issn-type":[{"value":"1539-9087","type":"print"},{"value":"1558-3465","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,30]]},"assertion":[{"value":"2021-06-29","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-02-19","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2022-12-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}