{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,24]],"date-time":"2026-07-24T08:04:06Z","timestamp":1784880246546,"version":"3.55.0"},"reference-count":53,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2022,2,28]],"date-time":"2022-02-28T00:00:00Z","timestamp":1646006400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["71971127"],"award-info":[{"award-number":["71971127"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Entropy"],"abstract":"<jats:p>A critical problem in large neural networks is over parameterization with a large number of weight parameters, which limits their use on edge devices due to prohibitive computational power and memory\/storage requirements. To make neural networks more practical on edge devices and real-time industrial applications, they need to be compressed in advance. Since edge devices cannot train or access trained networks when internet resources are scarce, the preloading of smaller networks is essential. Various works in the literature have shown that the redundant branches can be pruned strategically in a fully connected network without sacrificing the performance significantly. However, majority of these methodologies need high computational resources to integrate weight training via the back-propagation algorithm during the process of network compression. In this work, we draw attention to the optimization of the network structure for preserving performance despite compression by pruning aggressively. The structure optimization is performed using the simulated annealing algorithm only, without utilizing back-propagation for branch weight training. Being a heuristic-based, non-convex optimization method, simulated annealing provides a globally near-optimal solution to this NP-hard problem for a given percentage of branch pruning. Our simulation results have shown that simulated annealing can significantly reduce the complexity of a fully connected network while maintaining the performance without the help of back-propagation.<\/jats:p>","DOI":"10.3390\/e24030348","type":"journal-article","created":{"date-parts":[[2022,2,28]],"date-time":"2022-02-28T20:09:57Z","timestamp":1646078997000},"page":"348","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":38,"title":["Neural Network Structure Optimization by Simulated Annealing"],"prefix":"10.3390","volume":"24","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8762-7708","authenticated-orcid":false,"given":"Chun Lin","family":"Kuo","sequence":"first","affiliation":[{"name":"Tsinghua-Berkeley Shenzhen Institute, Nanshan, Shenzhen 518071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ercan Engin","family":"Kuruoglu","sequence":"additional","affiliation":[{"name":"Tsinghua-Berkeley Shenzhen Institute, Nanshan, Shenzhen 518071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7202-1922","authenticated-orcid":false,"given":"Wai Kin Victor","family":"Chan","sequence":"additional","affiliation":[{"name":"Tsinghua-Berkeley Shenzhen Institute, Nanshan, Shenzhen 518071, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,28]]},"reference":[{"key":"ref_1","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.C. (2018, January 18\u201323). MobileNetV2: Inverted Residuals and Linear Bottlenecks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2015, January 7\u201313). Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.123"},{"key":"ref_4","first-page":"249","article-title":"Understanding the difficulty of training deep feedforward neural networks","volume":"Volume 9","author":"Teh","year":"2010","journal-title":"Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics"},{"key":"ref_5","unstructured":"(1989, January 18\u201322). Theory of the backpropagation neural network. Proceedings of the International 1989 Joint Conference on Neural Networks, Washington, DC, USA."},{"key":"ref_6","first-page":"2121","article-title":"Adaptive subgradient methods for online learning and stochastic optimization","volume":"12","author":"Duchi","year":"2011","journal-title":"J. Mach. Learn. Res."},{"key":"ref_7","first-page":"1504","article-title":"RMSProp and equilibrated adaptive learning rates for non-convex optimization","volume":"28","author":"Dauphin","year":"2015","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_8","unstructured":"Kingma, D.P., and Ba, J. (2015, January 7\u20139). Adam: A Method for Stochastic Optimization. Proceedings of the 3rd International Conference for Learning Representations, San Diego, CA, USA."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Wang, K., Liu, Z., Lin, Y., Lin, J., and Han, S. (2019, January 15\u201320). HAQ: Hardware- aware automated quantization with mixed precision. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00881"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Zhang, X., Zou, J., Ming, X., He, K., and Sun, J. (2015, January 7\u201312). Efficient and Accurate Approximations of Nonlinear Convolutional Networks. Proceedings of the 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298809"},{"key":"ref_11","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2014). Distilling the Knowledge in a Neural Network. arXiv."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Liu, Z., Mu, H., Zhang, X., Guo, Z., Yang, X., Tim Cheng, K.T., and Sun, J. (November, January 27). MetaPruning: Meta Learning for Automatic Neural Network Channel Pruning. Proceedings of the 2019 IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00339"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Lin, M., Ji, R., Wang, Y., Zhang, Y., Zhang, B., Tian, Y., and Shao, L. (2020, January 13\u201319). HRank: Filter Pruning using High-Rank Feature Map. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00160"},{"key":"ref_14","unstructured":"Zoph, B., and Le, Q.V. (2017). Neural Architecture Search with Reinforcement Learning. arXiv."},{"key":"ref_15","unstructured":"Cai, H., Zhu, L., and Han, S. (2019). ProxylessNAS: Direct Neural Architecture Search on Target Task and Hardware. arXiv."},{"key":"ref_16","unstructured":"Han, S., Mao, H., and Dally, W.J. (2015). Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding. arXiv."},{"key":"ref_17","first-page":"13986","article-title":"Directional Pruning of Deep Neural Networks","volume":"33","author":"Zhanyu","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_18","unstructured":"Shulman, Y. (2020). DiffPrune: Neural Network Pruning with Deterministic Approximate Binary Gates and L0 Regularization. arXiv."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Ye, X., Dai, P., Luo, J., Guo, X., Qi, Y., Yang, J., and Chen, Y. (2020, January 23\u201328). Accelerating CNN Training by Pruning Activation Gradients. Proceedings of the European Conference on Computer Vision (ECCV), Glasgow, UK.","DOI":"10.1007\/978-3-030-58595-2_20"},{"key":"ref_20","first-page":"20378","article-title":"Movement Pruning: Adaptive Sparsity by Fine-Tuning","volume":"33","author":"Victor","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"485","DOI":"10.1109\/JPROC.2020.2976475","article-title":"Model Compression and Hardware Acceleration for Neural Networks: A Comprehensive Survey","volume":"108","author":"Deng","year":"2020","journal-title":"Proc. IEEE"},{"key":"ref_22","unstructured":"Krishnan, G., Du, X., and Cao, Y. (2019). Structural Pruning in Deep Neural Networks: A Small-World Approach. arXiv."},{"key":"ref_23","unstructured":"Crowley, E.J., Turner, J., Storkey, A.J., and O\u2019Boyle, M.F.P. (2018). Pruning neural networks: Is it time to nip it in the bud?. arXiv."},{"key":"ref_24","unstructured":"Cortes, C., Lawrence, N., Lee, D., Sugiyama, M., and Garnett, R. (2015). Learning both Weights and Connections for Efficient Neural Network. Advances in Neural Information Processing Systems, Curran Associates, Inc."},{"key":"ref_25","unstructured":"Louizos, C., Welling, M., and Kingma, D.P. (2018). Learning Sparse Neural Networks through L0 Regularization. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Cho, M., Joshi, A., and Hegde, C. (2021, January 5\u20136). ESPN: Extremely Sparse Pruned Networks. Proceedings of the 2021 IEEE Data Science and Learning Workshop (DSLW), Toronto, ON, Canada.","DOI":"10.1109\/DSLW51110.2021.9523404"},{"key":"ref_27","first-page":"1","article-title":"Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks","volume":"22","author":"Hoefler","year":"2021","journal-title":"J. Mach. Learn. Res."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"He, T., Fan, Y., Qian, Y., Tan, T., and Yu, K. (2014, January 4\u20139). Reshaping deep neural network for fast decoding by node-pruning. Proceedings of the 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Florence, Italy.","DOI":"10.1109\/ICASSP.2014.6853595"},{"key":"ref_29","first-page":"6869","article-title":"Quantized neural networks: Training neural networks with low precision weights and activations","volume":"18","author":"Hubara","year":"2017","journal-title":"J. Mach. Learn. Res."},{"key":"ref_30","unstructured":"Xu, C., Yao, J., Lin, Z., Ou, W., Cao, Y., Wang, Z., and Zha, H. (2018). Alternating multi-bit quantization for recurrent neural networks. arXiv."},{"key":"ref_31","unstructured":"Choi, Y., El-Khamy, M., and Jungwon, L. (2017). Towards the Limit of Network Quantization. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Haase, P., Schwarz, H., Kirchhoffer, H., Wiedemann, S., Marinc, T., Marban, A., Muller, K., Samek, W., Marpe, D., and Wiegand, T. (2020, January 25\u201328). Dependent Scalar Quantization for Neural Network Compression. Proceedings of the 2020 IEEE International Conference on Image Processing (ICIP), Abu Dhabi, United Arab Emirates.","DOI":"10.1109\/ICIP40778.2020.9190955"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"772","DOI":"10.1109\/TNNLS.2019.2910073","article-title":"Compact and Computationally Efficient Representation of Deep Neural Networks","volume":"31","author":"Wiedemann","year":"2020","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Lin, M., Ji, R., Zhang, Y., Zhang, B., Wu, Y., and Tian, Y. (2020). Channel Pruning via Automatic Structure Search. arXiv.","DOI":"10.24963\/ijcai.2020\/94"},{"key":"ref_35","unstructured":"Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H.P. (2017). Pruning Filters for Efficient ConvNets. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"He, Y., Ding, Y., Liu, P., Zhu, L., Zhang, H., and Yang, Y. (2020, January 13\u201319). Learning Filter Pruning Criteria for Deep Convolutional Neural Networks Acceleration. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00208"},{"key":"ref_37","first-page":"589","article-title":"Optimal Brain Damage","volume":"2","author":"LeCun","year":"1989","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_38","first-page":"164","article-title":"Second order derivatives for network pruning: Optimal Brain Surgeon","volume":"5","author":"Hassibi","year":"1992","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_39","unstructured":"Chen, X., Zhu, J., Jiang, J., and Tsui, C.Y. (2021). Tight Compression: Compressing CNN Through Fine-Grained Pruning and Weight Permutation for Efficient Implementation. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1016\/j.procs.2015.12.114","article-title":"Simulated Annealing Algorithm for Deep Learning","volume":"72","author":"Rere","year":"2015","journal-title":"Procedia Comput. Sci."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"133","DOI":"10.1007\/s11036-018-1196-7","article-title":"Applying Improved Convolutional Neural Network in Image Classification","volume":"25","author":"Hu","year":"2020","journal-title":"Mob. Netw. Appl."},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Ayumi, V., Rere, L.M.R., Fanany, M.I., and Arymurthy, A.M. (2016, January 15\u201316). Optimization of convolutional neural network using microcanonical annealing algorithm. Proceedings of the 2016 International Conference on Advanced Computer Science and Information Systems (ICACSIS), Malang, Indonesia.","DOI":"10.1109\/ICACSIS.2016.7872787"},{"key":"ref_43","unstructured":"Han, F., Tu, J., and Zhan, Y. (2010, January 22\u201324). A Neural Network Pruning Method Optimized with PSO Algorithm. Proceedings of the 2010 Second International Conference on Computer Modeling and Simulation, Sanya, China."},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Wu, W. (2012, January 18\u201320). Neural network structure optimization based on improved genetic algorithm. Proceedings of the 2012 IEEE Fifth International Conference on Advanced Computational Intelligence (ICACI), Nanjing, China.","DOI":"10.1109\/ICACI.2012.6463299"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"1073","DOI":"10.1007\/s00521-016-2619-7","article-title":"Topology optimization of neural networks based on a coupled genetic algorithm and particle swarm optimization techniques (c-GA\u2013PSO-NN)","volume":"29","author":"Marjani","year":"2018","journal-title":"Neural Comput. Appl."},{"key":"ref_46","first-page":"5","article-title":"Neural Network Structure Optimization Algorithm","volume":"12","author":"Grzegorz","year":"2018","journal-title":"J. Autom. Mob. Robot. Intell. Syst."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Otten, R., and van Ginneken, L. (1989). The Annealing Algorithm. Engineering and Computer Science Free Previewcover, Springer.","DOI":"10.1007\/978-1-4613-1627-5"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"227","DOI":"10.1016\/j.jtbi.2017.01.046","article-title":"The information capacity of the genetic code: Is the natural code optimal","volume":"419","author":"Kuruoglu","year":"2017","journal-title":"J. Theor. Biol."},{"key":"ref_49","unstructured":"Kuruoglu, E.E., and Ayanoglu, E. (1993, January 17\u201322). Design of finite-state machines for quantization using simulated annealing. Proceedings of the 1993 IEEE International Symposium on Information Theory, San Antonio, TX, USA."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"310","DOI":"10.1016\/j.neucom.2021.09.003","article-title":"Simulated annealing for optimization of graphs and sequences","volume":"465","author":"Liu","year":"2021","journal-title":"Neurocomputing"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"1087","DOI":"10.1063\/1.1699114","article-title":"Equation of State Calculations by Fast Computing Machines","volume":"21","author":"Metropolis","year":"1953","journal-title":"J. Chem. Phys."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Laarhoven, P.J.M., and Aarts, E.H.L. (1987). Simulated Annealing: Theory and Applications. Mathematics and Its Applications, Springer.","DOI":"10.1007\/978-94-015-7744-1"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Vasudevan, A., Anderson, A., and Gregg, D. (2017, January 10\u201312). Parallel Multi Channel convolution using General Matrix Multiplication. Proceedings of the 2017 IEEE 28th International Conference on Application-Specific Systems, Architectures and Processors (ASAP), Seattle, WA, USA.","DOI":"10.1109\/ASAP.2017.7995254"}],"container-title":["Entropy"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/3\/348\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:29:13Z","timestamp":1760135353000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1099-4300\/24\/3\/348"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,28]]},"references-count":53,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["e24030348"],"URL":"https:\/\/doi.org\/10.3390\/e24030348","relation":{},"ISSN":["1099-4300"],"issn-type":[{"value":"1099-4300","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,2,28]]}}}