{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,19]],"date-time":"2026-05-19T23:40:48Z","timestamp":1779234048709,"version":"3.51.4"},"reference-count":41,"publisher":"MDPI AG","issue":"17","license":[{"start":{"date-parts":[[2022,9,2]],"date-time":"2022-09-02T00:00:00Z","timestamp":1662076800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"China Scholarship Council (CSC)"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Finding a template in a search image is an important task underlying many computer vision applications. This is typically solved by calculating a similarity map using features extracted from the separate images. Recent approaches perform template matching in a deep feature space, produced by a convolutional neural network (CNN), which is found to provide more tolerance to changes in appearance. Inspired by these findings, in this article we investigate whether enhancing the CNN\u2019s encoding of shape information can produce more distinguishable features that improve the performance of template matching. By comparing features from the same CNN trained using different shape\u2013texture training methods, we determined a feature space which improves the performance of most template matching algorithms. When combining the proposed method with the Divisive Input Modulation (DIM) template matching algorithm, its performance is greatly improved, and the resulting method produces state-of-the-art results on a standard benchmark. To confirm these results, we create a new benchmark and show that the proposed method outperforms existing techniques on this new dataset.<\/jats:p>","DOI":"10.3390\/s22176658","type":"journal-article","created":{"date-parts":[[2022,9,8]],"date-time":"2022-09-08T04:18:32Z","timestamp":1662610712000},"page":"6658","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Shape\u2013Texture Debiased Training for Robust Template Matching"],"prefix":"10.3390","volume":"22","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3930-1815","authenticated-orcid":false,"given":"Bo","family":"Gao","sequence":"first","affiliation":[{"name":"Department of Informatics, King\u2019s College London, London WC2R 2LS, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9531-2813","authenticated-orcid":false,"given":"Michael W.","family":"Spratling","sequence":"additional","affiliation":[{"name":"Department of Informatics, King\u2019s College London, London WC2R 2LS, UK"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2022,9,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Gao, B., and Spratling, M.W. (2022). Explaining away results in more robust visual tracking. Vis. Comput., 1\u201315.","DOI":"10.1007\/s00371-022-02466-6"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"108628","DOI":"10.1016\/j.sigpro.2022.108628","article-title":"More Robust Object Tracking via Shape and Motion Cue Integration","volume":"22","author":"Gao","year":"2022","journal-title":"Signal Process."},{"key":"ref_3","first-page":"1368","article-title":"Object recognition by template matching using correlations and phase angle method","volume":"2","author":"Ahuja","year":"2013","journal-title":"Int. J. Adv. Res. Comput. Commun. Eng."},{"key":"ref_4","unstructured":"Dai, J., Li, Y., He, K., and Sun, J. (2016, January 5\u201310). R-fcn: Object detection via region-based fully convolutional networks. Proceedings of the Advances in Neural Information Processing Systems, Barcelona, Spain."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1023\/A:1014573219977","article-title":"A taxonomy and evaluation of dense two-frame stereo correspondence algorithms","volume":"47","author":"Scharstein","year":"2002","journal-title":"Int. J. Comput. Vis."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Chhatkuli, A., Pizarro, D., and Bartoli, A. (2014, January 23\u201328). Stable template-based isometric 3D reconstruction in all imaging conditions by linear least-squares. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Washington, DC, USA.","DOI":"10.1109\/CVPR.2014.96"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1799","DOI":"10.1109\/TPAMI.2017.2737424","article-title":"Best-buddies similarity\u2014Robust template matching using mutual nearest neighbors","volume":"40","author":"Oron","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"Distinctive image features from scale-invariant keypoints","volume":"60","author":"Lowe","year":"2004","journal-title":"Int. J. Comput. Vis."},{"key":"ref_9","unstructured":"Dalal, N., and Triggs, B. (2005, January 20\u201325). Histograms of oriented gradients for human detection. Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905), San Diego, CA, USA."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"850","DOI":"10.1016\/j.ijleo.2018.06.094","article-title":"Robust image matching based on the information of SIFT","volume":"171","author":"Dou","year":"2018","journal-title":"Optik"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Lee, H., Kwon, H., Robinson, R.M., and Nothwang, W.D. (2016, January 20\u201325). DTM: Deformable template matching. Proceedings of the 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Shanghai, China.","DOI":"10.1109\/ICASSP.2016.7472020"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Sibiryakov, A. (2011, January 20\u201325). Fast and high-performance template matching method. Proceedings of the CVPR 2011, Washington, DC, USA.","DOI":"10.1109\/CVPR.2011.5995391"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Arslan, O., Demirci, B., Altun, H., and Tunaboylu, N.S. (2013, January 9\u201311). A novel rotation-invariant template matching based on HOG and AMDF for industrial laser cutting applications. Proceedings of the 2013 9th International Symposium on Mechatronics and Its Applications (ISMA), Amman, Jordan.","DOI":"10.1109\/ISMA.2013.6547367"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Antipov, G., Berrani, S.A., Ruchaud, N., and Dugelay, J.L. (2015, January 26\u201330). Learned vs. hand-crafted features for pedestrian gender recognition. Proceedings of the 23rd ACM international Conference on Multimedia, Brisbane, Australia.","DOI":"10.1145\/2733373.2806332"},{"key":"ref_15","doi-asserted-by":"crossref","first-page":"5017","DOI":"10.1109\/TIP.2015.2475625","article-title":"PCANet: A simple deep learning baseline for image classification?","volume":"24","author":"Chan","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Wang, F., Jiang, M., Qian, C., Yang, S., Li, C., Zhang, H., Wang, X., and Tang, X. (2017, January 21\u201326). Residual attention network for image classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.683"},{"key":"ref_17","unstructured":"Liang, M., and Hu, X. (2015, January 7\u201312). Recurrent convolutional neural network for object recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Wohlhart, P., and Lepetit, V. (2015, January 7\u201312). Learning descriptors for object recognition and 3d pose estimation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298930"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Bertinetto, L., Valmadre, J., Henriques, J.F., Vedaldi, A., and Torr, P.H. (2016, January 8\u201316). Fully-convolutional siamese networks for object tracking. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-48881-3_56"},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"2709","DOI":"10.1109\/TPAMI.2018.2865311","article-title":"Robust visual tracking via hierarchical convolutional features","volume":"41","author":"Ma","year":"2018","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Cheng, J., Wu, Y., AbdAlmageed, W., and Natarajan, P. (2019, January 16\u201317). QATM: Quality-Aware Template Matching For Deep Learning. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01182"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Kat, R., Jevnisek, R., and Avidan, S. (2018, January 18\u201323). Matching pixels using co-occurrence statistics. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00188"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Kim, J., Kim, J., Choi, S., Hasan, M.A., and Kim, C. (2017, January 12\u201315). Robust template matching using scale-adaptive deep convolutional features. Proceedings of the 2017 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC), Kuala Lumpur, Malaysia.","DOI":"10.1109\/APSIPA.2017.8282124"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Talmi, I., Mechrez, R., and Zelnik-Manor, L. (2017, January 21\u201326). Template matching with deformable diversity similarity. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.144"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"6787","DOI":"10.1109\/TII.2020.2972290","article-title":"Weighted smallest deformation similarity for NN-based template matching","volume":"16","author":"Zhang","year":"2020","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"107029","DOI":"10.1016\/j.patcog.2019.107029","article-title":"Fast and robust template matching with majority neighbour similarity and annulus projection transformation","volume":"98","author":"Lai","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"417","DOI":"10.1146\/annurev-vision-082114-035447","article-title":"Deep neural networks: A new framework for modeling biological vision and brain information processing","volume":"1","author":"Kriegeskorte","year":"2015","journal-title":"Annu. Rev. Vis. Sci."},{"key":"ref_28","unstructured":"Geirhos, R., Rubisch, P., Michaelis, C., Bethge, M., Wichmann, F.A., and Brendel, W. (2018). ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv."},{"key":"ref_29","unstructured":"Li, Y., Yu, Q., Tan, M., Mei, J., Tang, P., Shen, W., Yuille, A., and Xie, C. (2020). Shape-Texture Debiased Neural Network Training. arXiv."},{"key":"ref_30","doi-asserted-by":"crossref","first-page":"107337","DOI":"10.1016\/j.patcog.2020.107337","article-title":"Explaining away results in accurate and tolerant template matching","volume":"104","author":"Spratling","year":"2020","journal-title":"Pattern Recognit."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Gao, B., and Spratling, M.W. (2021, January 15\u201317). Robust Template Matching via Hierarchical Convolutional Features from a Shape Biased CNN. Proceedings of the The International Conference on Image, Vision and Intelligent Systems (ICIVIS 2021), Changsha, China.","DOI":"10.1007\/978-981-16-6963-7_31"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Korman, S., Milam, M., and Soatto, S. (2018, January 18\u201323). OATM: Occlusion aware template matching by consensus set maximization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00283"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"271","DOI":"10.1146\/annurev.psych.55.090902.142005","article-title":"Object perception as Bayesian inference","volume":"55","author":"Kersten","year":"2004","journal-title":"Annu. Rev. Psychol."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1162\/NECO_a_00222","article-title":"Unsupervised learning of generative and discriminative weights encoding elementary image components in a predictive coding model of cortical function","volume":"24","author":"Spratling","year":"2012","journal-title":"Neural Comput."},{"key":"ref_35","unstructured":"Simonyan, K., and Zisserman, A. (2014). Very deep convolutional networks for large-scale image recognition. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Huang, X., and Belongie, S. (2017, January 21\u201326). Arbitrary style transfer in real-time with adaptive instance normalization. Proceedings of the IEEE International Conference on Computer Vision, Honolulu, HI, USA.","DOI":"10.1109\/ICCV.2017.167"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"774","DOI":"10.1016\/j.conb.2011.05.018","article-title":"Neural processing as causal inference","volume":"21","author":"Lochmann","year":"2011","journal-title":"Curr. Opin. Neurobiol."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"4179","DOI":"10.1523\/JNEUROSCI.0817-11.2012","article-title":"Perceptual inference predicts contextual modulations of sensory responses","volume":"32","author":"Lochmann","year":"2012","journal-title":"J. Neurosci."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"1834","DOI":"10.1109\/TPAMI.2014.2388226","article-title":"Object tracking benchmark","volume":"37","author":"Wu","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"5630","DOI":"10.1109\/TIP.2015.2482905","article-title":"Encoding color information for visual tracking: Algorithms and benchmark","volume":"24","author":"Liang","year":"2015","journal-title":"IEEE Trans. Image Process."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/17\/6658\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T00:22:43Z","timestamp":1760142163000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/22\/17\/6658"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,2]]},"references-count":41,"journal-issue":{"issue":"17","published-online":{"date-parts":[[2022,9]]}},"alternative-id":["s22176658"],"URL":"https:\/\/doi.org\/10.3390\/s22176658","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,2]]}}}