{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,21]],"date-time":"2025-02-21T03:21:08Z","timestamp":1740108068566,"version":"3.37.3"},"reference-count":45,"publisher":"Springer Science and Business Media LLC","issue":"18","license":[{"start":{"date-parts":[[2021,3,19]],"date-time":"2021-03-19T00:00:00Z","timestamp":1616112000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2021,3,19]],"date-time":"2021-03-19T00:00:00Z","timestamp":1616112000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100005642","name":"Universit\u00e0 degli Studi di Sassari","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100005642","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Neural Comput &amp; Applic"],"published-print":{"date-parts":[[2021,9]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Traditional local image descriptors such as SIFT and SURF are based on processings similar to those that take place in the early visual cortex. Nowadays, convolutional neural networks still draw inspiration from the human vision system, integrating computational elements typical of higher visual cortical areas. Deep CNN\u2019s architectures are intrinsically hard to interpret, so much effort has been made to dissect them in order to understand which type of features they learn. However, considering the resemblance to the human vision system, no enough attention has been devoted to understand if the image features learned by deep CNNs and used for classification correlate with features that humans select when viewing images, the so-called human fixations, nor if they correlate with earlier developed handcrafted features such as SIFT and SURF. Exploring these correlations is highly meaningful since what we require from CNNs, and features in general, is to recognize and correctly classify objects or subjects relevant to humans. In this paper, we establish the correlation between three families of image interest points: human fixations, handcrafted and CNN features. We extract features from the feature maps of selected layers of several deep CNN\u2019s architectures, from the shallowest to the deepest. All features and fixations are then compared with two types of measures, global and local, which unveil the degree of similarity of the areas of interest of the three families. From the experiments carried out on ETD human fixations database, it turns out that human fixations are positively correlated with handcrafted features and even more with deep layers of CNNs and that handcrafted features highly correlate between themselves as some CNNs do.<\/jats:p>","DOI":"10.1007\/s00521-021-05863-5","type":"journal-article","created":{"date-parts":[[2021,3,19]],"date-time":"2021-03-19T13:04:49Z","timestamp":1616159089000},"page":"11905-11922","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":11,"title":["On the correlation between human fixations, handcrafted and CNN features"],"prefix":"10.1007","volume":"33","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3970-7797","authenticated-orcid":false,"given":"Marinella","family":"Cadoni","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Andrea","family":"Lagorio","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Souad","family":"Khellat-Kihel","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Enrico","family":"Grosso","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2021,3,19]]},"reference":[{"issue":"3","key":"5863_CR1","doi-asserted-by":"publisher","first-page":"346","DOI":"10.1016\/j.cviu.2007.09.014","volume":"110","author":"H Bay","year":"2008","unstructured":"Bay H, Ess A, Tuytelaars T, Van Gool L (2008) Speeded-up robust features (surf). Comput Vis Image Underst 110(3):346\u2013359","journal-title":"Comput Vis Image Underst"},{"issue":"5","key":"5863_CR2","doi-asserted-by":"publisher","first-page":"2916","DOI":"10.1214\/10-AOS799","volume":"38","author":"ZI Botev","year":"2010","unstructured":"Botev ZI, Grotowski JF, Kroese DP et al (2010) Kernel density estimation via diffusion. Ann Stat 38(5):2916\u20132957","journal-title":"Ann Stat"},{"issue":"4","key":"5863_CR3","doi-asserted-by":"publisher","first-page":"325","DOI":"10.2307\/1942268","volume":"27","author":"JR Bray","year":"1957","unstructured":"Bray JR, Curtis JT (1957) An ordination of the upland forest communities of Southern Wisconsin. Ecol Monogr 27(4):325\u2013349. https:\/\/doi.org\/10.2307\/1942268","journal-title":"Ecol Monogr"},{"issue":"3","key":"5863_CR4","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1167\/9.3.5","volume":"9","author":"ND Bruce","year":"2009","unstructured":"Bruce ND, Tsotsos JK (2009) Saliency, attention, and visual search: an information theoretic approach. J Vis 9(3):5","journal-title":"J Vis"},{"key":"5863_CR5","doi-asserted-by":"crossref","unstructured":"Cadoni M, Lagorio A, Grosso E (2014) Iconic methods for multimodal face recognition: a comparative study. In: 2014 22nd international conference on pattern recognition, pp 4612\u20134617","DOI":"10.1109\/ICPR.2014.789"},{"issue":"C","key":"5863_CR6","doi-asserted-by":"publisher","first-page":"42","DOI":"10.1016\/j.imavis.2016.05.008","volume":"52","author":"M Cadoni","year":"2016","unstructured":"Cadoni M, Lagorio A, Grosso E (2016) Large scale face identification by combined iconic features and 3d joint invariant signatures. Image Vis Comput 52(C):42\u201355","journal-title":"Image Vis Comput"},{"key":"5863_CR7","doi-asserted-by":"publisher","first-page":"38","DOI":"10.1016\/j.patrec.2019.02.019","volume":"122","author":"M Cadoni","year":"2019","unstructured":"Cadoni M, Lagorio A, Grosso E (2019) Incremental models based on features persistence for object recognition. Pattern Recognit Lett 122:38\u201344. https:\/\/doi.org\/10.1016\/j.patrec.2019.02.019","journal-title":"Pattern Recognit Lett"},{"key":"5863_CR8","doi-asserted-by":"publisher","unstructured":"Cadoni M, Lagorio A, Grosso E (2020) Do cnn\u2019s features correlate with human fixations? In: Petkov N, Strisciuglio N, Travieso-Gonz\u00e1lez CM (eds) APPIS 2020: 3rd international conference on applications of intelligent systems, APPIS 2020, Las Palmas de Gran Canaria Spain, 7\u20139 January 2020, ACM, pp 13:1\u201313:6. https:\/\/doi.org\/10.1145\/3378184.3378197","DOI":"10.1145\/3378184.3378197"},{"key":"5863_CR9","doi-asserted-by":"crossref","unstructured":"Chatfield K, Simonyan K, Vedaldi A, Zisserman A (2014) Return of the devil in the details: delving deep into convolutional nets. In: Proceedings of the British machine vision conference (BMVC), pp 1\u201310","DOI":"10.5244\/C.28.6"},{"key":"5863_CR10","doi-asserted-by":"crossref","unstructured":"Chen S, Zhao Q (2018) Boosted attention: leveraging human attention for image captioning. arXiv:1904.00767","DOI":"10.1007\/978-3-030-01252-6_5"},{"key":"5863_CR11","unstructured":"Dave A, Dubey R, Ghanem B (2012) Do humans fixate on interest points? In: Proceedings of the 21st international conference on pattern recognition (ICPR2012). IEEE, pp 2784\u20132787"},{"key":"5863_CR12","unstructured":"Erhan D, Bengio Y, Courville A, Vincent P (2009) Visualizing higher-layer features of a deep network. Technical Report, University of Montreal 1341(3)"},{"issue":"5","key":"5863_CR13","doi-asserted-by":"publisher","first-page":"476","DOI":"10.1007\/s11263-017-1048-0","volume":"126","author":"A Gonzalez-Garcia","year":"2018","unstructured":"Gonzalez-Garcia A, Modolo D, Ferrari V (2018) Do semantic parts emerge in convolutional neural networks? Int J Comput Vis 126(5):476\u2013494","journal-title":"Int J Comput Vis"},{"key":"5863_CR14","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Identity mappings in deep residual networks. In: European conference on computer vision (ECCV), pp 630\u2013645","DOI":"10.1007\/978-3-319-46493-0_38"},{"key":"5863_CR15","doi-asserted-by":"crossref","unstructured":"Hou X, Zhang L (2007) Saliency detection: a spectral residual approach. In: 2007 IEEE conference on computer vision and pattern recognition. IEEE, pp 1\u20138","DOI":"10.1109\/CVPR.2007.383267"},{"key":"5863_CR16","doi-asserted-by":"crossref","unstructured":"Huang G, Liu Z, van\u00a0der Maaten L, Weinberger KQ (2017) Densely connected convolutional networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition (CVPR)","DOI":"10.1109\/CVPR.2017.243"},{"issue":"1","key":"5863_CR17","doi-asserted-by":"publisher","first-page":"106","DOI":"10.1113\/jphysiol.1962.sp006837","volume":"160","author":"D Hubel","year":"1962","unstructured":"Hubel D, Wiesel T (1962) Receptive fields, binocular interaction and functional architecture in the cat\u2019s visual cortex. J Physiol 160(1):106\u2013154. https:\/\/doi.org\/10.1113\/jphysiol.1962.sp006837","journal-title":"J Physiol"},{"issue":"10\u201312","key":"5863_CR18","doi-asserted-by":"publisher","first-page":"1489","DOI":"10.1016\/S0042-6989(99)00163-7","volume":"40","author":"L Itti","year":"2000","unstructured":"Itti L, Koch C (2000) A saliency-based search mechanism for overt and covert shifts of visual attention. Vis Res 40(10\u201312):1489\u20131506","journal-title":"Vis Res"},{"key":"5863_CR19","first-page":"337","volume":"11","author":"C Jones","year":"1996","unstructured":"Jones C, Marron JS, Sheather SJ (1996) Progress in data-based bandwidth selection for kernel density estimation. Comput Stat 11:337\u2013381","journal-title":"Comput Stat"},{"key":"5863_CR20","doi-asserted-by":"crossref","unstructured":"Judd T, Ehinger K, Durand F, Torralba A (2009) Learning to predict where humans look. In: 2009 IEEE 12th international conference on computer vision. IEEE, pp 2106\u20132113","DOI":"10.1109\/ICCV.2009.5459462"},{"key":"5863_CR21","unstructured":"Krizhevsky A, Sutskever I, Hinton GE (2012) Imagenet classification with deep convolutional neural networks. In: Advances in neural information processing systems, pp 1097\u20131105"},{"key":"5863_CR22","unstructured":"K\u00fcmmerer M, Theis L, Bethge M (2015) Deep gaze i: boosting saliency prediction with feature maps trained on imagenet. In: ICLR workshop, arXiv:1411.1045"},{"key":"5863_CR23","doi-asserted-by":"crossref","unstructured":"Li K, Wu Z, Peng K, Ernst J, Fu Y (2018) Tell me where to look: guided attention inference network. 2018 IEEE\/CVF conference on computer vision and pattern recognition, pp 9215\u20139223","DOI":"10.1109\/CVPR.2018.00960"},{"key":"5863_CR24","doi-asserted-by":"publisher","first-page":"145","DOI":"10.1109\/18.61115","volume":"37","author":"J Lin","year":"1991","unstructured":"Lin J (1991) Divergence measures based on the Shannon entropy. IEEE Trans Inf Theory 37:145\u2013151","journal-title":"IEEE Trans Inf Theory"},{"key":"5863_CR25","doi-asserted-by":"crossref","unstructured":"Mahendran A, Vedaldi A (2015) Understanding deep image representations by inverting them. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 5188\u20135196","DOI":"10.1109\/CVPR.2015.7299155"},{"issue":"5","key":"5863_CR26","doi-asserted-by":"publisher","first-page":"2116","DOI":"10.1109\/TIP.2018.2881920","volume":"28","author":"KR Mopuri","year":"2019","unstructured":"Mopuri KR, Garg U, Venkatesh Babu R (2019) Cnn fixations: an unraveling approach to visualize the discriminative image regions. IEEE Trans Image Process 28(5):2116\u20132125","journal-title":"IEEE Trans Image Process"},{"key":"5863_CR27","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1016\/j.visres.2011.12.006","volume":"57","author":"MS Mould","year":"2012","unstructured":"Mould MS, Foster DH, Amano K, Oakley JP (2012) A simple nonparametric method for classifying eye fixations. Vis Res 57:18\u201325. https:\/\/doi.org\/10.1016\/j.visres.2011.12.006","journal-title":"Vis Res"},{"key":"5863_CR28","doi-asserted-by":"publisher","first-page":"65","DOI":"10.1007\/978-3-319-01796-9_7","volume-title":"Genetic and evolutionary computing","author":"T Nguyen","year":"2014","unstructured":"Nguyen T, Park EA, Han J, Park DC, Min SY (2014) Object detection using scale invariant feature transform. In: Pan JS, Kr\u00f6mer P, Sn\u00e1\u0161el V (eds) Genetic and evolutionary computing. Springer, Cham, pp 65\u201372"},{"key":"5863_CR29","doi-asserted-by":"publisher","unstructured":"Rs R, Cogswell M, Das A, Vedantam R, Parikh D, Batra D (2017) Grad-cam: visual explanations from deep networks via gradient-based localization, pp 618\u2013626. https:\/\/doi.org\/10.1109\/ICCV.2017.74","DOI":"10.1109\/ICCV.2017.74"},{"key":"5863_CR30","doi-asserted-by":"publisher","first-page":"411","DOI":"10.1109\/TPAMI.2007.56","volume":"3","author":"T Serre","year":"2007","unstructured":"Serre T, Wolf L, Bileschi S, Riesenhuber M, Poggio T (2007) Robust object recognition with cortex-like mechanisms. IEEE Trans Pattern Anal Mach Intell 3:411\u2013426","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"5863_CR31","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. In: 3rd international conference on learning representations (ICLR), pp 1\u201310"},{"key":"5863_CR32","unstructured":"Simonyan K, Vedaldi A, Zisserman A (2013) Deep inside convolutional networks: visualising image classification models and saliency maps. arXiv preprint arXiv:13126034"},{"key":"5863_CR33","doi-asserted-by":"crossref","unstructured":"Stephens M, Harris C (1988) A combined corner and edge detector. In: Alvey vision conference, vol\u00a015","DOI":"10.5244\/C.2.23"},{"key":"5863_CR34","doi-asserted-by":"crossref","unstructured":"Szegedy C, Vanhoucke V, Ioffe S, Shlens J, Wojna Z (2016) Rethinking the inception architecture for computer vision. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 2818\u20132826","DOI":"10.1109\/CVPR.2016.308"},{"key":"5863_CR35","unstructured":"Tan M, Le Q (2019) Efficientnet: rethinking model scaling for convolutional neural networks. In: International conference on machine learning, pp 6105\u20136114"},{"key":"5863_CR36","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017) Attention is all you need. CoRR, arXiv:1706.03762"},{"key":"5863_CR37","unstructured":"Xu K, Ba J, Kiros R, Cho K, Courville A, Salakhudinov R, Zemel R, Bengio Y (2015) Show, attend and tell: Neural image caption generation with visual attention. In: Bach F, Blei D (eds) Proceedings of the 32nd international conference on machine learning, PMLR, Lille, France, Proceedings of Machine Learning Research, vol\u00a037, pp 2048\u20132057"},{"issue":"3","key":"5863_CR38","doi-asserted-by":"publisher","first-page":"506","DOI":"10.1109\/JSTSP.2020.2987729","volume":"14","author":"W Xu","year":"2020","unstructured":"Xu W, Wang J, Wang Y, Xu G, Lin D, Dai W, Wu Y (2020) Where is the model looking at? Concentrate and explain the network attention. IEEE J Sel Top Signal Process 14(3):506\u2013516. https:\/\/doi.org\/10.1109\/JSTSP.2020.2987729","journal-title":"IEEE J Sel Top Signal Process"},{"issue":"3","key":"5863_CR39","doi-asserted-by":"publisher","first-page":"356","DOI":"10.1038\/nn.4244","volume":"19","author":"DL Yamins","year":"2016","unstructured":"Yamins DL, DiCarlo JJ (2016) Using goal-driven deep learning models to understand sensory cortex. Nat Neurosci 19(3):356","journal-title":"Nat Neurosci"},{"issue":"23","key":"5863_CR40","doi-asserted-by":"publisher","first-page":"8619","DOI":"10.1073\/pnas.1403112111","volume":"111","author":"DL Yamins","year":"2014","unstructured":"Yamins DL, Hong H, Cadieu CF, Solomon EA, Seibert D, DiCarlo JJ (2014) Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proc Natl Acad Sci 111(23):8619\u20138624","journal-title":"Proc Natl Acad Sci"},{"key":"5863_CR41","unstructured":"Yosinski J, Clune J, Nguyen A, Fuchs T, Lipson H (2015) Understanding neural networks through deep visualization. arXiv preprint arXiv:150606579"},{"key":"5863_CR42","unstructured":"Zeiler MD, Fergus R (2013) Visualizing and understanding convolutional networks. CoRR, arXiv:1311.2901"},{"key":"5863_CR43","unstructured":"Zhou B, Khosla A, Lapedriza A, Oliva A, Torralba A (2015) Object detectors emerge in deep scene cnns. arXiv:1412.6856"},{"key":"5863_CR44","doi-asserted-by":"publisher","first-page":"2131","DOI":"10.1109\/TPAMI.2018.2858759","volume":"41","author":"B Zhou","year":"2019","unstructured":"Zhou B, Bau D, Oliva A, Torralba A (2019) Interpreting deep visual representations via network dissection. IEEE Trans Pattern Anal Mach Intell 41:2131\u20132145","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"5863_CR45","doi-asserted-by":"publisher","first-page":"345","DOI":"10.1016\/j.cviu.2008.08.006","volume":"113","author":"H Zhou","year":"2009","unstructured":"Zhou H, Yuan Y, Shi C (2009) Object tracking using sift features and mean shift. Comput Vis Image Underst 113:345\u2013352. https:\/\/doi.org\/10.1016\/j.cviu.2008.08.006","journal-title":"Comput Vis Image Underst"}],"container-title":["Neural Computing and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-021-05863-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s00521-021-05863-5\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s00521-021-05863-5.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2021,8,23]],"date-time":"2021-08-23T20:13:50Z","timestamp":1629749630000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s00521-021-05863-5"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,3,19]]},"references-count":45,"journal-issue":{"issue":"18","published-print":{"date-parts":[[2021,9]]}},"alternative-id":["5863"],"URL":"https:\/\/doi.org\/10.1007\/s00521-021-05863-5","relation":{},"ISSN":["0941-0643","1433-3058"],"issn-type":[{"type":"print","value":"0941-0643"},{"type":"electronic","value":"1433-3058"}],"subject":[],"published":{"date-parts":[[2021,3,19]]},"assertion":[{"value":"31 July 2020","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 February 2021","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 March 2021","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}