{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,20]],"date-time":"2025-10-20T10:29:10Z","timestamp":1760956150654,"version":"build-2065373602"},"reference-count":39,"publisher":"MDPI AG","issue":"4","license":[{"start":{"date-parts":[[2022,2,18]],"date-time":"2022-02-18T00:00:00Z","timestamp":1645142400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2019YFB2204303"],"award-info":[{"award-number":["2019YFB2204303"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["U20A20205","U21A20504"],"award-info":[{"award-number":["U20A20205","U21A20504"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Strategic Priority Research Program of the Chinese Academy of Science","award":["XDB32050200"],"award-info":[{"award-number":["XDB32050200"]}]},{"name":"Key Research Program of the Chinese Academy of Sciences","award":["XDPB22"],"award-info":[{"award-number":["XDPB22"]}]},{"name":"Youth Innovation Promotion Association Program Chinese Academy of Sciences","award":["2021109"],"award-info":[{"award-number":["2021109"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Siamese networks have been extensively studied in recent years. Most of the previous research focuses on improving accuracy, while merely a few recognize the necessity of reducing parameter redundancy and computation load. Even less work has been done to optimize the runtime memory cost when designing networks, making the Siamese-network-based tracker difficult to deploy on edge devices. In this paper, we present SiamMixer, a lightweight and hardware-friendly visual object-tracking network. It uses patch-by-patch inference to reduce memory use in shallow layers, where each small image region is processed individually. It merges and globally encodes feature maps in deep layers to enhance accuracy. Benefiting from these techniques, SiamMixer demonstrates a comparable accuracy to other large trackers with only 286 kB parameters and 196 kB extra memory use for feature maps. Additionally, we verify the impact of various activation functions and replace all activation functions with ReLU in SiamMixer. This reduces the cost when deploying on mobile devices.<\/jats:p>","DOI":"10.3390\/s22041585","type":"journal-article","created":{"date-parts":[[2022,2,21]],"date-time":"2022-02-21T08:34:47Z","timestamp":1645432487000},"page":"1585","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["SiamMixer: A Lightweight and Hardware-Friendly Visual Object-Tracking Network"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-1963-8234","authenticated-orcid":false,"given":"Li","family":"Cheng","sequence":"first","affiliation":[{"name":"State Key Laboratory of Superlattices and Microstructures, Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China"},{"name":"Center of Materials Science and Optoelectronics Engineering, University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xuemin","family":"Zheng","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Superlattices and Microstructures, Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China"},{"name":"Center of Materials Science and Optoelectronics Engineering, University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5148-4218","authenticated-orcid":false,"given":"Mingxin","family":"Zhao","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Superlattices and Microstructures, Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China"},{"name":"Center of Materials Science and Optoelectronics Engineering, University of Chinese Academy of Sciences, Beijing 100049, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Runjiang","family":"Dou","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Superlattices and Microstructures, Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shuangming","family":"Yu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Superlattices and Microstructures, Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Nanjian","family":"Wu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Superlattices and Microstructures, Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China"},{"name":"Center of Materials Science and Optoelectronics Engineering, University of Chinese Academy of Sciences, Beijing 100049, China"},{"name":"The Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Liyuan","family":"Liu","sequence":"additional","affiliation":[{"name":"State Key Laboratory of Superlattices and Microstructures, Institute of Semiconductors, Chinese Academy of Sciences, Beijing 100083, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,2,18]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"438","DOI":"10.1049\/el.2014.0033","article-title":"Heterogeneous vision chip and LBP-based algorithm for high-speed tracking","volume":"50","author":"Yang","year":"2014","journal-title":"Electron. Lett."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Li, B., Wu, W., Wang, Q., Zhang, F., Xing, J., and Yan, J. (2019, January 15\u201320). SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00441"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"771","DOI":"10.1007\/978-3-030-58589-1_46","article-title":"Ocean: Object-Aware Anchor-Free Tracking","volume":"12366","author":"Vedaldi","year":"2020","journal-title":"Computer Vision\u2014ECCV 2020"},{"key":"ref_4","first-page":"12549","article-title":"SiamFC++: Towards Robust and Accurate Visual Tracking with Target Estimation Guidelines","volume":"34","author":"Xu","year":"2020","journal-title":"Proc. AAAI Conf. Artif. Intell."},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S.E., Anguelov, D., Erhan, D., Vanhoucke, V., and Rabinovich, A. (2015, January 7\u201312). Going deeper with convolutions. Proceedings of the 2015 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"ref_6","first-page":"1097","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Danelljan, M., Bhat, G., Khan, F.S., and Felsberg, M. (2019, January 15\u201320). ATOM: Accurate Tracking by Overlap Maximization. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00479"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Bhat, G., Danelljan, M., Gool, L.V., and Timofte, R. (2019, January 27\u201328). Learning Discriminative Model Prediction for Tracking. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea.","DOI":"10.1109\/ICCV.2019.00628"},{"key":"ref_9","unstructured":"Iandola, F.N., Han, S., Moskewicz, M.W., Ashraf, K., Dally, W.J., and Keutzer, K. (2016). SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5 MB model size. arXiv."},{"key":"ref_10","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. arXiv."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Sandler, M., Howard, A.G., Zhu, M., Zhmoginov, A., and Chen, L. (2018, January 18\u201322). MobileNetV2: Inverted Residuals and Linear Bottlenecks. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition 2018, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00474"},{"key":"ref_12","first-page":"850","article-title":"Fully-Convolutional Siamese Networks for Object Tracking","volume":"9914","author":"Hua","year":"2016","journal-title":"Computer Vision \u2013 ECCV 2016"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Li, B., Yan, J., Wu, W., Zhu, Z., and Hu, X. (2018, January 18\u201323). High Performance Visual Tracking With Siamese Region Proposal Network. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00935"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhang, Z., and Peng, H. (2019, January 15\u201320). Deeper and Wider Siamese Networks for Real-Time Visual Tracking. Proceedings of the 2019 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00472"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Guo, Q., Feng, W., Zhou, C., Huang, R., Wan, L., and Wang, S. (2017, January 22\u201329). Learning Dynamic Siamese Network for Visual Object Tracking. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.196"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Yang, T., Xu, P., Hu, R., Chai, H., and Chan, A.B. (2020, January 14\u201319). ROAM: Recurrently Optimizing Tracking Model. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00675"},{"key":"ref_17","unstructured":"Han, S., Mao, H., and Dally, W.J. (2016, January 2\u20134). Deep Compression: Compressing Deep Neural Network with Pruning, Trained Quantization and Huffman Coding. Proceedings of the 4th International Conference on Learning Representations, San Juan, Puerto Rico."},{"key":"ref_18","unstructured":"Hinton, G., Vinyals, O., and Dean, J. (2015). Distilling the Knowledge in a Neural Network. arXiv."},{"key":"ref_19","unstructured":"Esser, S.K., McKinstry, J.L., Bablani, D., Appuswamy, R., and Modha, D.S. (2020, January 26\u201330). Learned Step Size quantization. Proceedings of the 8th International Conference on Learning Representations, Addis Ababa, Ethiopia."},{"key":"ref_20","first-page":"1","article-title":"A comprehensive survey of neural architecture search: Challenges and solutions","volume":"54","author":"Ren","year":"2021","journal-title":"ACM Comput. Surv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Yan, B., Peng, H., Wu, K., Wang, D., Fu, J., and Lu, H. (2021, January 20\u201325). LightTrack: Finding Lightweight Neural Networks for Object Tracking via One-Shot Architecture Search. Proceedings of the 2021 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.01493"},{"key":"ref_22","unstructured":"Lin, J., Chen, W.M., Cai, H., Gan, C., and Han, S. (2021). MCUNetV2: Memory-Efficient Patch-based Inference for Tiny Deep Learning. arXiv."},{"key":"ref_23","first-page":"1","article-title":"MLP-mixer: An all-mlp architecture for vision","volume":"34","author":"Tolstikhin","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_24","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3\u20137). An Image is Worth 16 \u00d7 16 Words: Transformers for Image Recognition at Scale. Proceedings of the 9th International Conference on Learning Representations, Virtual Event, Austria."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"472","DOI":"10.1007\/978-3-030-01261-8_28","article-title":"Triplet Loss in Siamese Network for Object Tracking","volume":"11217","author":"Ferrari","year":"2018","journal-title":"Computer Vision\u2014ECCV 2018"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"1834","DOI":"10.1109\/TPAMI.2014.2388226","article-title":"Object Tracking Benchmark","volume":"37","author":"Wu","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"445","DOI":"10.1007\/978-3-319-46448-0_27","article-title":"A Benchmark and Simulator for UAV Tracking","volume":"9905","author":"Leibe","year":"2016","journal-title":"Computer Vision\u2014ECCV 2016"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1562","DOI":"10.1109\/TPAMI.2019.2957464","article-title":"GOT-10k: A Large High-Diversity Benchmark for Generic Object Tracking in the Wild","volume":"43","author":"Huang","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"749","DOI":"10.1007\/978-3-319-46448-0_45","article-title":"Learning to Track at 100 FPS with Deep Regression Networks","volume":"9905","author":"Leibe","year":"2016","journal-title":"Computer Vision\u2014ECCV 2016"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Hong, Z., Chen, Z., Wang, C., Mei, X., Prokhorov, D.V., and Tao, D. (2015, January 7\u201312). MUlti-Store Tracker (MUSTer): A cognitive psychology inspired approach to object tracking. Proceedings of the 2015 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298675"},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"188","DOI":"10.1007\/978-3-319-10599-4_13","article-title":"MEEM: Robust Tracking via Multiple Experts Using Entropy Minimization","volume":"8694","author":"Fleet","year":"2014","journal-title":"Computer Vision\u2014ECCV 2014"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Hare, S., Saffari, A., and Torr, P.H.S. (2011, January 6\u201313). Struck: Structured output tracking with kernels. Proceedings of the IEEE International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126251"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"1409","DOI":"10.1109\/TPAMI.2011.239","article-title":"Tracking-Learning-Detection","volume":"34","author":"Kalal","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Galoogahi, H.K., Fagg, A., and Lucey, S. (2017, January 22\u201329). Learning Background-Aware Correlation Filters for Visual Tracking. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.129"},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"583","DOI":"10.1109\/TPAMI.2014.2345390","article-title":"High-Speed Tracking with Kernelized Correlation Filters","volume":"37","author":"Henriques","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_36","unstructured":"Cai, H., Gan, C., Wang, T., Zhang, Z., and Han, S. (2020, January 26\u201330). Once-for-All: Train One Network and Specialize it for Efficient Deployment. Proceedings of the 8th International Conference on Learning Representations, Addis Ababa, Ethiopia."},{"key":"ref_37","unstructured":"Canziani, A., Paszke, A., and Culurciello, E. (2016). An Analysis of Deep Neural Network Models for Practical Applications. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Wan, A., Dai, X., Zhang, P., He, Z., Tian, Y., Xie, S., Wu, B., Yu, M., Xu, T., and Chen, K. (2020, January 13\u201319). FBNetV2: Differentiable Neural Architecture Search for Spatial and Channel Dimensions. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01298"},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep Residual Learning for Image Recognition. Proceedings of the 2016 IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/4\/1585\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:22:06Z","timestamp":1760134926000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/4\/1585"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,2,18]]},"references-count":39,"journal-issue":{"issue":"4","published-online":{"date-parts":[[2022,2]]}},"alternative-id":["s22041585"],"URL":"https:\/\/doi.org\/10.3390\/s22041585","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2022,2,18]]}}}