{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,1]],"date-time":"2025-11-01T05:48:04Z","timestamp":1761976084350,"version":"3.35.0"},"reference-count":53,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"2","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Fundamentals"],"published-print":{"date-parts":[[2025,2,1]]},"DOI":"10.1587\/transfun.2023eap1163","type":"journal-article","created":{"date-parts":[[2024,8,4]],"date-time":"2024-08-04T22:13:53Z","timestamp":1722809633000},"page":"149-159","source":"Crossref","is-referenced-by-count":1,"title":["Accelerating CNN Inference with an Adaptive Quantization Method Using Computational Complexity-Aware Regularization"],"prefix":"10.1587","volume":"E108.A","author":[{"given":"Kengo","family":"NAKATA","sequence":"first","affiliation":[{"name":"Kioxia Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Daisuke","family":"MIYASHITA","sequence":"additional","affiliation":[{"name":"Kioxia Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"DEGUCHI","sequence":"additional","affiliation":[{"name":"Kioxia Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryuichi","family":"FUJIMOTO","sequence":"additional","affiliation":[{"name":"Kioxia Corporation"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] K. Simonyan and A. Zisserman, \u201cVery deep convolutional networks for large-scale image recognition,\u201d International Conference on Learning Representations (ICLR), 2015."},{"key":"2","doi-asserted-by":"crossref","unstructured":"[2] K. He, X. Zhang, S. Ren, and J. Sun, \u201cDeep residual learning for image recognition,\u201d IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.770-778, 2016. 10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] K. He, G. Gkioxari, P. Dollar, and R. Girshick, \u201cMask R-CNN,\u201d IEEE\/CVF International Conference on Computer Vision (ICCV), pp.2980-2988, 2017. 10.1109\/iccv.2017.322","DOI":"10.1109\/ICCV.2017.322"},{"key":"4","unstructured":"[4] M. Courbariaux, I. Hubara, D. Soudry, R. El-Yaniv, and Y. Bengio, \u201cBinarized neural networks: Training deep neural networks with weights and activations constrained to +1 or -1,\u201d arXiv:1602.02830, 2016. 10.48550\/arXiv.1602.02830"},{"key":"5","doi-asserted-by":"crossref","unstructured":"[5] D. Zhang, J. Yang, D. Ye, and G. Hua, \u201cLQ-Nets: Learned quantization for highly accurate and compact deep neural networks,\u201d European Conference on Computer Vision (ECCV), pp.373-390, 2018. 10.1007\/978-3-030-01237-3_23","DOI":"10.1007\/978-3-030-01237-3_23"},{"key":"6","doi-asserted-by":"publisher","unstructured":"[6] Y. Zhou, S. Moosavi-Dezfooli, N. Cheung, and P. Frossard, \u201cAdaptive quantization for deep neural network,\u201d AAAI Conference on Artificial Intelligence, pp.4596-4604, 2018. 10.1609\/aaai.v32i1.11623","DOI":"10.1609\/aaai.v32i1.11623"},{"key":"7","unstructured":"[7] H. Wu, P. Judd, X. Zhang, M. Isaev, and P. Micikevicius, \u201cInteger quantization for deep learning inference: Principles and empirical evaluation,\u201d arXiv:2004.09602, 2020. 10.48550\/arXiv.2004.09602"},{"key":"8","doi-asserted-by":"publisher","unstructured":"[8] J. Choquette, W. Gandhi, O. Giroux, N. Stam, and R. Krashinsky, \u201cNVIDIA A100 tensor core GPU: Performance and innovation,\u201d IEEE Micro, vol.41, no.2, pp.29-35, 2021. 10.1109\/mm.2021.3061394","DOI":"10.1109\/MM.2021.3061394"},{"key":"9","doi-asserted-by":"publisher","unstructured":"[9] J. Lee, C. Kim, S. Kang, D. Shin, S. Kim, and H. Yoo, \u201cUNPU: An energy-efficient deep neural network accelerator with fully variable weight bit precision,\u201d IEEE J. Solid-State Circuits, vol.54, no.1, pp.173-185, 2019. 10.1109\/jssc.2018.2865489","DOI":"10.1109\/JSSC.2018.2865489"},{"key":"10","doi-asserted-by":"crossref","unstructured":"[10] J.S. Park, C. Park, S. Kwon, H.S. Kim, T. Jeon, Y. Kang, H. Lee, D. Lee, J. Kim, Y. Lee, S. Park, J.W. Jang, S. Ha, M. Kim, J. Bang, S.H. Lim, and I. Kang, \u201cA multi-mode 8K-MAC HW-utilization-aware neural processing unit with a unified multi-precision datapath in 4\u2006nm flagship mobile SoC,\u201d IEEE International Solid- State Circuits Conference (ISSCC), pp.246-248, 2022. 10.1109\/isscc42614.2022.9731639","DOI":"10.1109\/ISSCC42614.2022.9731639"},{"key":"11","doi-asserted-by":"crossref","unstructured":"[11] S. Sasaki, A. Maki, D. Miyashita, and J. Deguchi, \u201cPost training weight compression with distribution-based filter-wise quantization step,\u201d IEEE Symposium in Low-Power and High-Speed Chips (COOL CHIPS), pp.1-3, 2019. 10.1109\/coolchips.2019.8721356","DOI":"10.1109\/CoolChips.2019.8721356"},{"key":"12","doi-asserted-by":"crossref","unstructured":"[12] L. Zeng, Z. Wang, and X. Tian, \u201cKCNN: Kernel-wise quantization to remarkably decrease multiplications in convolutional neural network,\u201d International Joint Conference on Artificial Intelligence (IJCAI), pp.4234-4242, 2019. 10.24963\/ijcai.2019\/588","DOI":"10.24963\/ijcai.2019\/588"},{"key":"13","doi-asserted-by":"crossref","unstructured":"[13] Y. Cai, Z. Yao, Z. Dong, A. Gholami, M.W. Mahoney, and K. Keutzer, \u201cZeroQ: A novel zero shot quantization framework,\u201d IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.13166-13175, 2020. 10.1109\/cvpr42600.2020.01318","DOI":"10.1109\/CVPR42600.2020.01318"},{"key":"14","doi-asserted-by":"crossref","unstructured":"[14] Z. Dong, Z. Yao, A. Gholami, M.W. Mahoney, and K. Keutzer, \u201cHAWQ: Hessian AWare Quantization of neural networks with mixed-precision,\u201d IEEE\/CVF International Conference on Computer Vision (ICCV), pp.293-302, 2019. 10.1109\/iccv.2019.00038","DOI":"10.1109\/ICCV.2019.00038"},{"key":"15","unstructured":"[15] Z. Dong, Z. Yao, D. Arfeen, A. Gholami, M.W. Mahoney, and K. Keutzer, \u201cHAWQ-V2: Hessian aware trace-weighted quantization of neural networks,\u201d Advances in Neural Information Processing Systems (NeurIPS), pp.18518-18529, 2020."},{"key":"16","doi-asserted-by":"crossref","unstructured":"[16] K. Wang, Z. Liu, Y. Lin, J. Lin, and S. Han, \u201cHAQ: Hardware-aware automated quantization with mixed precision,\u201d IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.8604-8612, 2019. 10.1109\/cvpr.2019.00881","DOI":"10.1109\/CVPR.2019.00881"},{"key":"17","unstructured":"[17] Q. Lou, F. Guo, M. Kim, L. Liu, and L. Jiang., \u201cAutoQ: Automated kernel-wise neural network quantization,\u201d International Conference on Learning Representations (ICLR), 2020."},{"key":"18","unstructured":"[18] S.K. Esser, J.L. McKinstry, D. Bablani, R. Appuswamy, and D.S. Modha, \u201cLearned step size quantization,\u201d International Conference on Learning Representations (ICLR), 2020."},{"key":"19","unstructured":"[19] S. Jain, A. Gural, M. Wu, and C. Dick, \u201cTrained quantization thresholds for accurate and efficient fixed-point inference of deep neural networks,\u201d Proc. Machine Learning and Systems (MLSys), pp.112-128, 2020."},{"key":"20","unstructured":"[20] S. Uhlich, L. Mauch, F. Cardinaux, K. Yoshiyama, J.A. Garcia, S. Tiedemann, T. Kemp, and A. Nakamura, \u201cMixed precision DNNs: All you need is a good parametrization,\u201d International Conference on Learning Representations (ICLR), 2020."},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] K. Nakata, D. Miyashita, J. Deguchi, and R. Fujimoto, \u201cAdaptive quantization method for CNN with computational-complexity-aware regularization,\u201d IEEE International Symposium on Circuits and Systems (ISCAS), pp.1-5, 2021. 10.1109\/iscas51556.2021.9401657","DOI":"10.1109\/ISCAS51556.2021.9401657"},{"key":"22","unstructured":"[22] R. Banner, Y. Nahshan, and D. Soudry, \u201cPost training 4-bit quantization of convolutional networks for rapid-deployment,\u201d Advances in Neural Information Processing Systems (NeurIPS), 2019."},{"key":"23","doi-asserted-by":"publisher","unstructured":"[23] A.T. Elthakeb, P. Pilligundla, F. Mireshghallah, A. Yazdanbakhsh, and H. Esmaeilzadeh, \u201cReLeQ: A reinforcement learning approach for automatic deep quantization of neural networks,\u201d IEEE Micro, vol.40, no.5, pp.37-45, 2020. 10.1109\/mm.2020.3009475","DOI":"10.1109\/MM.2020.3009475"},{"key":"24","doi-asserted-by":"publisher","unstructured":"[24] A. Maki, D. Miyashita, S. Sasaki, K. Nakata, F. Tachibana, T. Suzuki, J. Deguchi, and R. Fujimoto, \u201cWeight compression MAC accelerator for effective inference of deep learning,\u201d IEICE Trans. Electron., vol.E103-C, no.10, pp.514-523, Oct. 2020. 10.1587\/transele.2019ctp0007","DOI":"10.1587\/transele.2019CTP0007"},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] K. Nakata, D. Miyashita, A. Maki, F. Tachibana, S. Sasaki, J. Deguchi, and R. Fujimoto, \u201cQuantization strategy for pareto-optimally low-cost and accurate CNN,\u201d IEEE International Conference on Artificial Intelligence Circuits and Systems (AICAS), pp.1-4, 2021. 10.1109\/aicas51828.2021.9458452","DOI":"10.1109\/AICAS51828.2021.9458452"},{"key":"26","doi-asserted-by":"crossref","unstructured":"[26] A. Maki, D. Miyashita, K. Nakata, F. Tachibana, T. Suzuki, and J. Deguchi, \u201cFPGA-based CNN processor with filter-wise-optimized bit precision,\u201d IEEE Asian Solid-State Circuits Conference (A-SSCC), pp.47-50, 2018. 10.1109\/asscc.2018.8579342","DOI":"10.1109\/ASSCC.2018.8579342"},{"key":"27","doi-asserted-by":"crossref","unstructured":"[27] K. He, X. Zhang, S. Ren, and J. Sun, \u201cIdentity mappings in deep residual networks,\u201d arXiv:1603.05027, 2016. 10.48550\/arXiv.1603.05027","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"28","doi-asserted-by":"crossref","unstructured":"[28] J. Deng, W. Dong, R. Socher, L.J. Li, K. Li, and L. Fei-Fei, \u201cImageNet: A large-scale hierarchical image database,\u201d IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.248-255, 2009. 10.1109\/cvpr.2009.5206848","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"29","unstructured":"[29] B. Wu, Y. Wang, P. Zhang, Y. Tian, P. Vajda, and K. Keutzer, \u201cMixed precision quantization of convnets via differentiable neural architecture search,\u201d arXiv:1812.00090, 2018. 10.48550\/arXiv.1812.00090"},{"key":"30","doi-asserted-by":"crossref","unstructured":"[30] Z. Cai and N. Vasconcelos, \u201cRethinking differentiable search for mixed-precision neural networks,\u201d Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.2346-2355, 2020. 10.1109\/cvpr42600.2020.00242","DOI":"10.1109\/CVPR42600.2020.00242"},{"key":"31","unstructured":"[31] M. van Baalen, C. Louizos, M. Nagel, R.A. Amjad, Y. Wang, T. Blankevoort, and M. Welling, \u201cBayesian bits: Unifying quantization and pruning,\u201d Advances in Neural Information Processing Systems (NeurIPS), pp.5741-5752, 2020."},{"key":"32","doi-asserted-by":"publisher","unstructured":"[32] L. Yang and Q. Jin, \u201cFracbits: Mixed precision quantization via fractional bit-widths,\u201d AAAI Conference on Artificial Intelligence, pp.10612-10620, 2021. 10.1609\/aaai.v35i12.17269","DOI":"10.1609\/aaai.v35i12.17269"},{"key":"33","doi-asserted-by":"publisher","unstructured":"[33] C. Baskin, N. Liss, E. Schwartz, E. Zheltonozhskii, R. Giryes, A.M. Bronstein, and A. Mendelson, \u201cUNIQ: Uniform noise injection for non-uniform quantization of neural networks,\u201d ACM Trans. Comput. Syst., vol.37, no.1-4, pp.1-15, 2021. 10.1145\/3444943","DOI":"10.1145\/3444943"},{"key":"34","doi-asserted-by":"publisher","unstructured":"[34] B. Hawks, J. Duarte, N.J. Fraser, A. Pappalardo, N. Tran, and Y. Umuroglu, \u201cPs and Qs: Quantization-aware pruning for efficient low latency neural network inference,\u201d Frontiers in Artificial Intelligence, vol.4, 2021. 10.3389\/frai.2021.676564","DOI":"10.3389\/frai.2021.676564"},{"key":"35","doi-asserted-by":"crossref","unstructured":"[35] Z. Liu, Y. Wang, K. Han, S. Ma, and W. Gao, \u201cInstance-aware dynamic neural network quantization,\u201d Proc. IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.12434-12443, 2022. 10.1109\/cvpr52688.2022.01211","DOI":"10.1109\/CVPR52688.2022.01211"},{"key":"36","doi-asserted-by":"crossref","unstructured":"[36] C. Hong, S. Baik, H. Kim, S. Nah, and K.M. Lee, \u201cCADyQ: Content-aware dynamic quantization for image super-resolution,\u201d European Conference on Computer Vision (ECCV), pp.367-383, 2022. 10.1007\/978-3-031-20071-7_22","DOI":"10.1007\/978-3-031-20071-7_22"},{"key":"37","doi-asserted-by":"crossref","unstructured":"[37] C. Tang, K. Ouyang, Z. Chai, Y. Bai, Y. Meng, Z. Wang, and W. Zhu, \u201cSEAM: Searching transferable mixed-precision quantization policy through large margin regularization,\u201d ACM International Conference on Multimedia, pp.7971-7980, 2023. 10.1145\/3581783.3611975","DOI":"10.1145\/3581783.3611975"},{"key":"38","doi-asserted-by":"crossref","unstructured":"[38] Y. Umuroglu, L. Rasnayake, and M. Sjalander, \u201cBISMO: A scalable bit-serial matrix multiplication overlay for reconfigurable computing,\u201d International Conference on Field Programmable Logic and Applications (FPL), pp.307-3077, 2018. 10.1109\/fpl.2018.00059","DOI":"10.1109\/FPL.2018.00059"},{"key":"39","unstructured":"[39] Y. Wang, G. Wei, and D. Brooks, \u201cBenchmarking TPU, GPU, and CPU platforms for deep learning,\u201d arXiv:1907.10701, 2019. 10.48550\/arXiv.1907.10701"},{"key":"40","doi-asserted-by":"crossref","unstructured":"[40] N.P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, S. Bates, S. Bhatia, N. Boden, A. Borchers, R. Boyle, P.-L. Cantin, C. Chao, C. Clark, J. Coriell, M. Daley, M. Dau, J. Dean, B. Gelb, T.V. Ghaemmaghami, R. Gottipati, W. Gulland, R. Hagmann, C.R. Ho, D. Hogberg, J. Hu, R. Hundt, D. Hurt, J. Ibarz, A. Jaffey, A. Jaworski, A. Kaplan, H. Khaitan, D. Killebrew, A. Koch, N. Kumar, S. Lacy, J. Laudon, J. Law, D. Le, C. Leary, Z. Liu, K. Lucke, A. Lundin, G. MacKean, A. Maggiore, M. Mahony, K. Miller, R. Nagarajan, R. Narayanaswami, R. Ni, K. Nix, T. Norrie, M. Omernick, N. Penukonda, A. Phelps, J. Ross, M. Ross, A. Salek, E. Samadiani, C. Severn, G. Sizikov, M. Snelham, J. Souter, D. Steinberg, A. Swing, M. Tan, G. Thorson, B. Tian, H. Toma, E. Tuttle, V. Vasudevan, R. Walter, W. Wang, E. Wilcox, and D.H. Yoon, \u201cIn-datacenter performance analysis of a tensor processing unit,\u201d ACM\/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), pp.1-12, 2017. 10.1145\/3079856.3080246","DOI":"10.1145\/3079856.3080246"},{"key":"41","doi-asserted-by":"crossref","unstructured":"[41] A. Castell\u00f3, M.F. Dolz, E.S. Quintana-Ort\u0131\u0301, and J. Duato, \u201cTheoretical scalability analysis of distributed deep convolutional neural networks,\u201d IEEE\/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGRID), pp.534-541, 2019. 10.1109\/CCGRID.2019.00068","DOI":"10.1109\/CCGRID.2019.00068"},{"key":"42","unstructured":"[42] A. Canziani, A. Paszke, and E. Culurciello, \u201cAn analysis of deep neural network models for practical applications,\u201d arXiv:1605.07678, 2016. 10.48550\/arXiv.1605.07678"},{"key":"43","doi-asserted-by":"publisher","unstructured":"[43] S. Bianco, R. Cadene, L. Celona, and P. Napoletano, \u201cBenchmark analysis of representative deep neural network architectures,\u201d IEEE Access, vol.6, pp.64270-64277, 2018. 10.1109\/access.2018.2877890","DOI":"10.1109\/ACCESS.2018.2877890"},{"key":"44","unstructured":"[44] H. Cai, C. Gan, T. Wang, Z. Zhang, and S. Han, \u201cOnce for all: Train one network and specialize it for efficient deployment,\u201d International Conference on Learning Representations (ICLR), 2020."},{"key":"45","doi-asserted-by":"crossref","unstructured":"[45] Z. Guo, X. Zhang, H. Mu, W. Heng, Z. Liu, Y. Wei, and J. Sun, \u201cSingle path one-shot neural architecture search with uniform sampling,\u201d European Conference on Computer Vision (ECCV), pp.544-560, 2020. 10.1007\/978-3-030-58517-4_32","DOI":"10.1007\/978-3-030-58517-4_32"},{"key":"46","doi-asserted-by":"crossref","unstructured":"[46] Z. Liu, J. Li, Z. Shen, G. Huang, S. Yan, and C. Zhang, \u201cLearning efficient convolutional networks through network slimming,\u201d IEEE International Conference on Computer Vision (ICCV), pp.2755-2763, 2017. 10.1109\/iccv.2017.298","DOI":"10.1109\/ICCV.2017.298"},{"key":"47","unstructured":"[47] H. Li, A. Kadav, I. Durdanovic, H. Samet, and H.P. Graf, \u201cPruning filters for efficient convnets,\u201d International Conference on Learning Representations (ICLR), 2017."},{"key":"48","unstructured":"[48] A. Renda, J. Frankle, and M. Carbin, \u201cComparing rewinding and fine-tuning in neural network pruning,\u201d International Conference on Learning Representations (ICLR), 2020."},{"key":"49","doi-asserted-by":"crossref","unstructured":"[49] M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L. Chen, \u201cMobileNetV2: Inverted residuals and linear bottlenecks,\u201d IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp.4510-4520, 2018. 10.1109\/cvpr.2018.00474","DOI":"10.1109\/CVPR.2018.00474"},{"key":"50","unstructured":"[50] Y. Bengio, N. L\u00e9onard, and A.C. Courville, \u201cEstimating or propagating gradients through stochastic neurons for conditional computation,\u201d arXiv:1308.3432, 2013. 10.48550\/arXiv.1308.3432"},{"key":"51","unstructured":"[51] A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, \u201cAutomatic differentiation in PyTorch,\u201d NIPS 2017 Workshop on Autodiff, 2017."},{"key":"52","unstructured":"[52] A. Coates, A. Ng, and H. Lee, \u201cAn analysis of single-layer networks in unsupervised feature Learning,\u201d International Conference on Artificial Intelligence and Statistics (AISTATS), pp.215-223, 2011."},{"key":"53","unstructured":"[53] I. Loshchilov and F. Hutter, \u201cSGDR: Stochastic gradient descent with restarts,\u201d arXiv:1608.03983, 2016. 10.48550\/arXiv.1608.03983"}],"container-title":["IEICE Transactions on Fundamentals of Electronics, Communications and Computer Sciences"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transfun\/E108.A\/2\/E108.A_2023EAP1163\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,1]],"date-time":"2025-02-01T03:29:18Z","timestamp":1738380558000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transfun\/E108.A\/2\/E108.A_2023EAP1163\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,2,1]]},"references-count":53,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025]]}},"URL":"https:\/\/doi.org\/10.1587\/transfun.2023eap1163","relation":{},"ISSN":["0916-8508","1745-1337"],"issn-type":[{"type":"print","value":"0916-8508"},{"type":"electronic","value":"1745-1337"}],"subject":[],"published":{"date-parts":[[2025,2,1]]},"article-number":"2023EAP1163"}}