{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T20:39:43Z","timestamp":1779914383566,"version":"3.53.1"},"reference-count":61,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2019,10,29]],"date-time":"2019-10-29T00:00:00Z","timestamp":1572307200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2019,10,29]],"date-time":"2019-10-29T00:00:00Z","timestamp":1572307200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100008332","name":"Graz University of Technology","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100008332","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2020,2]]},"abstract":"<jats:title>Abstract<\/jats:title>\n<jats:p>As the success of deep models has led to their deployment in all areas of computer vision, it is increasingly important to understand how these representations work and what they are capturing. In this paper, we shed light on deep spatiotemporal representations by visualizing the internal representation of models that have been trained to recognize actions in video. We visualize multiple two-stream architectures to show that local detectors for appearance and motion objects arise to form distributed representations for recognizing human actions. Key observations include the following. First, cross-stream fusion enables the learning of true spatiotemporal features rather than simply separate appearance and motion features. Second, the networks can learn local representations that are highly class specific, but also generic representations that can serve a range of classes. Third, throughout the hierarchy of the network, features become more abstract and show increasing invariance to aspects of the data that are unimportant to desired distinctions (e.g. motion patterns across various speeds). Fourth, visualizations can be used not only to shed light on learned representations, but also to reveal idiosyncrasies of training data and to explain failure cases of the system.<\/jats:p>","DOI":"10.1007\/s11263-019-01225-w","type":"journal-article","created":{"date-parts":[[2019,10,30]],"date-time":"2019-10-30T20:31:57Z","timestamp":1572467517000},"page":"420-437","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":21,"title":["Deep Insights into Convolutional Networks for Video Recognition"],"prefix":"10.1007","volume":"128","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9756-7238","authenticated-orcid":false,"given":"Christoph","family":"Feichtenhofer","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Axel","family":"Pinz","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Richard P.","family":"Wildes","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Andrew","family":"Zisserman","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2019,10,29]]},"reference":[{"key":"1225_CR1","unstructured":"Animated Manuscript. \nhttp:\/\/feichtenhofer.github.io\/pubs\/Feichtenhofer_IJCV19.pdf\n\n."},{"key":"1225_CR2","doi-asserted-by":"crossref","unstructured":"Bau, D., Zhou, B., Khosla, A., Oliva, A., & Torralba, A. (2017). Network dissection: Quantifying interpretability of deep visual representations. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2017.354"},{"key":"1225_CR3","doi-asserted-by":"crossref","unstructured":"Carreira, J., & Zisserman, A. (2017). Quo vadis, action recognition? A new model and the kinetics dataset. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2017.502"},{"key":"1225_CR4","doi-asserted-by":"crossref","unstructured":"Dollar, P., Rabaud, V., Cottrell, G., & Belongie, S. (2005). Behavior recognition via sparse spatio-temporal features. In ICCV VS-PETS.","DOI":"10.1109\/VSPETS.2005.1570899"},{"key":"1225_CR5","unstructured":"Dosovitskiy, A., & Brox, T. (2016). Generating images with perceptual similarity metrics based on deep networks. In NIPS."},{"key":"1225_CR6","unstructured":"Erhan, D., Bengio, Y., Courville, A., & Vincent, P. (2009). it Visualizing higher-layer features of a deep network. Technical report 1341, University of Montreal."},{"key":"1225_CR7","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., & Wildes, R. (2015). Dynamically encoded actions based on spacetime saliency. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2015.7298892"},{"key":"1225_CR8","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., & Wildes, R. (2016a). Spatiotemporal residual networks for video action recognition. In NIPS.","DOI":"10.1109\/CVPR.2017.787"},{"key":"1225_CR9","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., & Wildes, R. P. (2017). Spatiotemporal multiplier networks for video action recognition. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2017.787"},{"key":"1225_CR10","unstructured":"Feichtenhofer, C., Pinz, A., Wildes, R. P., & Zisserman, A. (2018). What have we learned from deep representations for action recognition? In Proceedings of CVPR."},{"key":"1225_CR11","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., & Zisserman, A. (2016b). Convolutional two-stream network fusion for video action recognition. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2016.213"},{"issue":"1","key":"1225_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1093\/cercor\/1.1.1","volume":"1","author":"DJ Felleman","year":"1991","unstructured":"Felleman, D. J., & Van Essen, D. C. (1991). Distributed hierarchical processing in the primate cerebral cortex. Cerebral Cortex, 1(1), 1\u201347.","journal-title":"Cerebral Cortex"},{"issue":"2","key":"1225_CR13","doi-asserted-by":"publisher","first-page":"194","DOI":"10.1162\/neco.1991.3.2.194","volume":"3","author":"P F\u00f6ldi\u00e1k","year":"1991","unstructured":"F\u00f6ldi\u00e1k, P. (1991). Learning invariance from transformation sequences. Neural Computation, 3(2), 194\u2013200.","journal-title":"Neural Computation"},{"issue":"9","key":"1225_CR14","doi-asserted-by":"publisher","first-page":"891","DOI":"10.1109\/34.93808","volume":"13","author":"W Freeman","year":"1991","unstructured":"Freeman, W., & Adelson, E. (1991). The design and use of steerable filters. IEEE PAMI, 13(9), 891\u2013906.","journal-title":"IEEE PAMI"},{"key":"1225_CR15","unstructured":"Galloway, A., Tanay, T., & Taylor, G. W. (2018). Adversarial training versus weight decay. arXiv preprint \narXiv:1804.03308\n\n."},{"issue":"1","key":"1225_CR16","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1016\/0166-2236(92)90344-8","volume":"15","author":"MA Goodale","year":"1992","unstructured":"Goodale, M. A., & Milner, A. D. (1992). Separate visual pathways for perception and action. Trends in Neurosciences, 15(1), 20\u201325.","journal-title":"Trends in Neurosciences"},{"key":"1225_CR17","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., & Bengio, Y. (2014). Generative adversarial nets. In NIPS."},{"issue":"7","key":"1225_CR18","doi-asserted-by":"publisher","first-page":"1450","DOI":"10.1364\/JOSAA.10.001450","volume":"10","author":"A Gorea","year":"1993","unstructured":"Gorea, A., & Papathomas, T. V. (1993). Double opponency as a generalized concept in texture segregation illustrated with stimuli defined by color, luminance, and orientation. Journal of the Optical Society of America A, 10(7), 1450\u20131462.","journal-title":"Journal of the Optical Society of America A"},{"key":"1225_CR19","unstructured":"Goroshin, R., Bruna, J., Tompson, J., Eigen, D., & LeCun, Y. (2015). Unsupervised feature learning from temporal data. In Proceedings of ICCV."},{"issue":"3","key":"1225_CR20","doi-asserted-by":"publisher","first-page":"583","DOI":"10.1113\/jphysiol.1974.sp010545","volume":"238","author":"P Gouras","year":"1974","unstructured":"Gouras, P. (1974). Opponent-colour cells in different layers of foveal striate cortex. The Journal of Physiology, 238(3), 583\u2013602.","journal-title":"The Journal of Physiology"},{"key":"1225_CR21","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2016.90"},{"issue":"7","key":"1225_CR22","doi-asserted-by":"publisher","first-page":"1527","DOI":"10.1162\/neco.2006.18.7.1527","volume":"18","author":"GE Hinton","year":"2006","unstructured":"Hinton, G. E., Osindero, S., & Teh, Y. W. (2006). A fast learning algorithm for deep belief nets. Neural Computation, 18(7), 1527\u20131554.","journal-title":"Neural Computation"},{"key":"1225_CR23","unstructured":"Ioffe, S., & Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. In Proceedings of ICML."},{"key":"1225_CR24","doi-asserted-by":"crossref","unstructured":"Johnson, J., Alahi, A., & Fei-Fei, L. (2016). Perceptual losses for real-time style transfer and super-resolution. In Proceedings of ECCV.","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"1225_CR25","unstructured":"Kingma, D., & Ba, J. (2015). Adam: A method for stochastic optimization. In Proceedings of ICLR."},{"key":"1225_CR26","doi-asserted-by":"crossref","unstructured":"Kl\u00e4ser, A., Marsza\u0142ek, M., & Schmid, C. (2008). A spatio-temporal descriptor based on 3D-gradients. In Proceedings of BMVC.","DOI":"10.5244\/C.22.99"},{"issue":"1","key":"1225_CR27","doi-asserted-by":"publisher","first-page":"48","DOI":"10.1162\/08989290051137594","volume":"12","author":"Z Kourtzi","year":"2000","unstructured":"Kourtzi, Z., & Kanwisher, N. (2000). Activation in human MT\/MST by static images with implied motion. Journal of Cognitive Neuroscience, 12(1), 48\u201355.","journal-title":"Journal of Cognitive Neuroscience"},{"key":"1225_CR28","doi-asserted-by":"crossref","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., & Serre, T. (2011). HMDB: A large video database for human motion recognition. In Proceedings of ICCV.","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"1225_CR29","unstructured":"Le, Q., Ranzato, M., Monga, R., Devin, M., Chen, K., Corrado, G., Dean, J., & Ng, A. (2012). Building high-level features using large scale unsupervised learning. In Proceedings of ICML."},{"key":"1225_CR30","unstructured":"Liu, C., Zoph, B., Shlens, J., Hua, W., Li, L. J., Fei-Fei, L., Yuille, A., Huang, J., & Murphy, K. (2017). Progressive neural architecture search. arXiv preprint \narXiv:1712.00559\n\n."},{"issue":"1","key":"1225_CR31","doi-asserted-by":"publisher","first-page":"309","DOI":"10.1523\/JNEUROSCI.04-01-00309.1984","volume":"4","author":"MS Livingstone","year":"1984","unstructured":"Livingstone, M. S., & Hubel, D. H. (1984). Anatomy and physiology of a color system in the primate visual cortex. Journal of Neuroscience, 4(1), 309\u2013356.","journal-title":"Journal of Neuroscience"},{"key":"1225_CR32","doi-asserted-by":"crossref","unstructured":"Mahendran, A., & Vedaldi, A. (2016a). Salient deconvolutional networks. In Proceedings of ECCV.","DOI":"10.1007\/978-3-319-46466-4_8"},{"issue":"3","key":"1225_CR33","doi-asserted-by":"publisher","first-page":"233","DOI":"10.1007\/s11263-016-0911-8","volume":"120","author":"A Mahendran","year":"2016","unstructured":"Mahendran, A., & Vedaldi, A. (2016b). Visualizing deep convolutional neural networks using natural pre-images. IJCV, 120(3), 233\u2013255.","journal-title":"IJCV"},{"key":"1225_CR34","doi-asserted-by":"publisher","first-page":"414","DOI":"10.1016\/0166-2236(83)90190-X","volume":"6","author":"M Mishkin","year":"1983","unstructured":"Mishkin, M., Ungerleider, L. G., & Macko, K. A. (1983). Object vision and spatial vision: Two cortical pathways. Trends in Neurosciences, 6, 414\u2013417.","journal-title":"Trends in Neurosciences"},{"key":"1225_CR35","unstructured":"Mordvintsev, A., Olah, C., & Tyka., M. (2015). Inceptionism: Going deeper into neural networks. Google Research Blog. Retrieved from June 20, \nhttps:\/\/research.googleblog.com\/2015\/06\/inceptionism-going-deeper-into-neural.html\n\n."},{"key":"1225_CR36","unstructured":"Nguyen, A., Dosovitskiy, A., Yosinski, J., Brox, T., & Clune, J. (2016). Synthesizing the preferred inputs for neurons in neural networks via deep generator networks. In NIPS."},{"key":"1225_CR37","doi-asserted-by":"crossref","unstructured":"Nguyen, A., Yosinski, J., Bengio, Y., Dosovitskiy, A., & Clune, J. (2017). Plug & play generative networks: Conditional iterative generation of images in latent space. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2017.374"},{"key":"1225_CR38","doi-asserted-by":"crossref","unstructured":"Nguyen, A., Yosinski, J., & Clune, J. (2015). Deep neural networks are easily fooled: High confidence predictions for unrecognizable images. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2015.7298640"},{"issue":"1","key":"1225_CR39","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/BF00595226","volume":"46","author":"W Reichardt","year":"1983","unstructured":"Reichardt, W., Poggio, T., & Hausen, K. (1983). Figure-ground discrimination by relative movement in the visual system of the fly. Biological Cybernetics, 46(1), 1\u201330.","journal-title":"Biological Cybernetics"},{"issue":"13","key":"1225_CR40","doi-asserted-by":"publisher","first-page":"5083","DOI":"10.1523\/JNEUROSCI.20-13-05083.2000","volume":"20","author":"K Saleem","year":"2000","unstructured":"Saleem, K., Suzuki, W., Tanaka, K., & Hashikawa, T. (2000). Connections between anterior inferotemporal cortex and superior temporal sulcus regions in the macaque monkey. Journal of Neuroscience, 20(13), 5083\u20135101.","journal-title":"Journal of Neuroscience"},{"key":"1225_CR41","unstructured":"Selvaraju, R. R., Das, A., Vedantam, R., Cogswell, M., Parikh, D., & Batra, D. (2016). Grad-cam: Why did you say that? Visual explanations from deep networks via gradient-based localization. arXiv preprint \narXiv:1610.02391\n\n."},{"key":"1225_CR42","unstructured":"Simonyan, K., Vedaldi, A., & Zisserman, A. (2014). Deep inside convolutional networks: Visualising image classification models and saliency maps. In ICLR workshop."},{"key":"1225_CR43","unstructured":"Simonyan, K., & Zisserman, A. (2014). Two-stream convolutional networks for action recognition in videos. In NIPS."},{"key":"1225_CR44","unstructured":"Simonyan, K., & Zisserman, A. (2015). Very deep convolutional networks for large-scale image recognition. In International conference on learning representations."},{"key":"1225_CR45","unstructured":"Soomro, K., Zamir, A. R., & Shah, M. (2012). UCF101: A dataset of 101 human actions classes from videos in the wild. Technical report CRCV-TR-12-01."},{"key":"1225_CR46","unstructured":"Springenberg, J. T., Dosovitskiy, A., Brox, T., & Riedmiller, M. (2015). Striving for simplicity: The all convolutional net. In ICLR workshop."},{"issue":"8","key":"1225_CR47","doi-asserted-by":"publisher","first-page":"876","DOI":"10.1364\/JOSAA.1.000876","volume":"1","author":"C Stromeyer","year":"1984","unstructured":"Stromeyer, C., Kronauer, R., Madsen, J., & Klein, S. (1984). Opponent-movement mechanisms in human vision. Journal of the Optical Society of America A, 1(8), 876\u2013884.","journal-title":"Journal of the Optical Society of America A"},{"key":"1225_CR48","doi-asserted-by":"crossref","unstructured":"Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., & Rabinovich, A. (2015). Going deeper with convolutions. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2015.7298594"},{"key":"1225_CR49","unstructured":"Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., & Wojna, Z. (2015). Rethinking the inception architecture for computer vision. arXiv preprint \narXiv:1512.00567\n\n."},{"key":"1225_CR50","unstructured":"Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., & Fergus, R. (2014). Intriguing properties of neural networks. In Proceedings of ICLR."},{"key":"1225_CR51","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., & Paluri, M. (2015). Learning spatiotemporal features with 3D convolutional networks. In Proceedings of ICCV.","DOI":"10.1109\/ICCV.2015.510"},{"key":"1225_CR52","doi-asserted-by":"crossref","unstructured":"Wang, H., & Schmid, C. (2013). Action recognition with improved trajectories. In Proceedings of ICCV.","DOI":"10.1109\/ICCV.2013.441"},{"key":"1225_CR53","doi-asserted-by":"crossref","unstructured":"Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., & Van Gool, L. (2016). Temporal segment networks: Towards good practices for deep action recognition. In ECCV.","DOI":"10.1007\/978-3-319-46484-8_2"},{"issue":"4","key":"1225_CR54","doi-asserted-by":"publisher","first-page":"715","DOI":"10.1162\/089976602317318938","volume":"14","author":"L Wiskott","year":"2002","unstructured":"Wiskott, L., & Sejnowski, T. J. (2002). Slow feature analysis: Unsupervised learning of invariances. Neural computation, 14(4), 715\u2013770.","journal-title":"Neural computation"},{"key":"1225_CR55","unstructured":"Yosinski, J., Clune, J., Nguyen, A., Fuchs, T., & Lipson, H. (2015). Understanding neural networks through deep visualization. In ICML workshop."},{"key":"1225_CR56","unstructured":"Zach, C., Pock, T., & Bischof, H. (2007). A duality based approach for realtime TV-L1 optical flow. In Proceedings of DAGM."},{"key":"1225_CR57","unstructured":"Zeiler, M. D., & Fergus, R. (2013). Visualizing and understanding convolutional networks. CoRR\n\narXiv:1311.2901\n\n."},{"key":"1225_CR58","doi-asserted-by":"crossref","unstructured":"Zhang, H., Yang, J., Zhang, Y., & Huang, T. S. (2010). Non-local kernel regression for image and video restoration. In Proceedings of ECCV.","DOI":"10.1007\/978-3-642-15558-1_41"},{"key":"1225_CR59","doi-asserted-by":"crossref","unstructured":"Zhang, J., Lin, Z., Brandt, J., Shen, X., & Sclaroff, S. (2016). Top-down neural attention by excitation backprop. In Proceedings of ECCV.","DOI":"10.1007\/978-3-319-46493-0_33"},{"key":"1225_CR60","unstructured":"Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., & Torralba, A. (2014). Object detectors emerge in deep scene CNNs. In Proceedings of ICLR."},{"key":"1225_CR61","doi-asserted-by":"crossref","unstructured":"Zhou, B., Khosla, A., Lapedriza, A., Oliva, A., & Torralba, A. (2016). Learning deep features for discriminative localization. In Proceedings of CVPR.","DOI":"10.1109\/CVPR.2016.319"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-019-01225-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11263-019-01225-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-019-01225-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2020,10,28]],"date-time":"2020-10-28T00:39:30Z","timestamp":1603845570000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11263-019-01225-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,10,29]]},"references-count":61,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2020,2]]}},"alternative-id":["1225"],"URL":"https:\/\/doi.org\/10.1007\/s11263-019-01225-w","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,10,29]]},"assertion":[{"value":"26 September 2018","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 August 2019","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 October 2019","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}