{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,12]],"date-time":"2026-05-12T23:30:08Z","timestamp":1778628608051,"version":"3.51.4"},"reference-count":92,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T00:00:00Z","timestamp":1734652800000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T00:00:00Z","timestamp":1734652800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Nat Comput Sci"],"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Understanding how visual information is encoded in biological and artificial systems often requires the generation of appropriate stimuli to test specific hypotheses, but available methods for video generation are scarce. Here we introduce the spatiotemporal style transfer (STST) algorithm, a dynamic visual stimulus generation framework that allows the manipulation and synthesis of video stimuli for vision research. We show how stimuli can be generated that match the low-level spatiotemporal features of their natural counterparts, but lack their high-level semantic features, providing a useful tool to study object recognition. We used these stimuli to probe PredNet, a predictive coding deep network, and found that its next-frame predictions were not disrupted by the omission of high-level information, with human observers also confirming the preservation of low-level features and lack of high-level information in the generated stimuli. We also introduce a procedure for the independent spatiotemporal factorization of dynamic stimuli. Testing such factorized stimuli on humans and deep vision models suggests a spatial bias in how humans and deep vision models encode dynamic visual information. These results showcase potential applications of the STST algorithm as a versatile tool for dynamic stimulus generation in vision science.<\/jats:p>","DOI":"10.1038\/s43588-024-00746-w","type":"journal-article","created":{"date-parts":[[2024,12,20]],"date-time":"2024-12-20T10:01:20Z","timestamp":1734688880000},"page":"155-169","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":8,"title":["A spatiotemporal style transfer algorithm for dynamic visual stimulus generation"],"prefix":"10.1038","volume":"5","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9502-4667","authenticated-orcid":false,"given":"Antonino","family":"Greco","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5115-936X","authenticated-orcid":false,"given":"Markus","family":"Siegel","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2024,12,20]]},"reference":[{"key":"746_CR1","unstructured":"Marr, D. Vision: A Computational Approach (MIT Press, 1982)."},{"key":"746_CR2","doi-asserted-by":"publisher","first-page":"167","DOI":"10.1016\/j.neuroimage.2019.03.028","volume":"193","author":"D Proklova","year":"2019","unstructured":"Proklova, D., Kaiser, D. & Peelen, M. V. MEG sensor patterns reflect perceptual but not categorical similarity of animate and inanimate objects. Neuroimage 193, 167\u2013177 (2019).","journal-title":"Neuroimage"},{"key":"746_CR3","doi-asserted-by":"publisher","first-page":"578","DOI":"10.1038\/nn1669","volume":"9","author":"AA Stocker","year":"2006","unstructured":"Stocker, A. A. & Simoncelli, E. P. Noise characteristics and prior expectations in human visual speed perception. Nat. Neurosci. 9, 578\u2013585 (2006).","journal-title":"Nat. Neurosci."},{"key":"746_CR4","doi-asserted-by":"publisher","DOI":"10.1038\/srep19739","volume":"6","author":"AJ Davies","year":"2016","unstructured":"Davies, A. J., Chaplin, T. A., Rosa, M. G. P. & Yu, H.-H. Natural motion trajectory enhances the coding of speed in primate extrastriate cortex. Sci. Rep. 6, 19739 (2016).","journal-title":"Sci. Rep."},{"key":"746_CR5","doi-asserted-by":"publisher","first-page":"108309","DOI":"10.1016\/j.jneumeth.2019.06.001","volume":"324","author":"AP Murphy","year":"2019","unstructured":"Murphy, A. P. & Leopold, D. A. A parameterized digital 3D model of the Rhesus macaque face for investigating the visual processing of social cues. J. Neurosci. Methods 324, 108309 (2019).","journal-title":"J. Neurosci. Methods"},{"key":"746_CR6","unstructured":"Raistrick, A. et al. Proc. 2023 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2023)."},{"key":"746_CR7","unstructured":"Greff, K. et al. Proc. 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2022)."},{"key":"746_CR8","doi-asserted-by":"publisher","first-page":"391","DOI":"10.1088\/0954-898X_14_3_302","volume":"14","author":"A Torralba","year":"2003","unstructured":"Torralba, A. & Oliva, A. Statistics of natural image categories. Network 14, 391\u2013412 (2003).","journal-title":"Network"},{"key":"746_CR9","doi-asserted-by":"publisher","first-page":"436","DOI":"10.1038\/nature14539","volume":"521","author":"Y LeCun","year":"2015","unstructured":"LeCun, Y., Bengio, Y. & Hinton, G. Deep learning. Nature 521, 436\u2013444 (2015).","journal-title":"Nature"},{"key":"746_CR10","unstructured":"Krizhevsky, A., Sutskever, I. & Hinton, G. E. ImageNet classification with deep convolutional neural networks. In Proc. Advances in Neural Information Processing Systems (eds Pereira, F. et al.) 1097\u20131105 (Curran Associates, Inc., 2012)."},{"key":"746_CR11","doi-asserted-by":"publisher","first-page":"114","DOI":"10.1016\/j.conb.2016.02.001","volume":"37","author":"DLK Yamins","year":"2016","unstructured":"Yamins, D. L. K. & DiCarlo, J. J. Eight open questions in the computational modeling of higher sensory cortex. Curr. Opin. Neurobiol. 37, 114\u2013120 (2016).","journal-title":"Curr. Opin. Neurobiol."},{"key":"746_CR12","doi-asserted-by":"publisher","DOI":"10.1126\/science.aav9436","volume":"364","author":"P Bashivan","year":"2019","unstructured":"Bashivan, P., Kar, K. & DiCarlo, J. J. Neural population control via deep image synthesis. Science 364, eaav9436 (2019).","journal-title":"Science"},{"key":"746_CR13","unstructured":"Simonyan, K., Vedaldi, A. & Zisserman, A. Proc. 2nd International Conference on Learning Representations (ICLR, 2014)."},{"key":"746_CR14","unstructured":"Mordvintsev, A., Olah, C. & Tyka, M. Inceptionism: going deeper into neural networks. Google Research Blog http:\/\/googleresearch.blogspot.co.uk\/2015\/06\/inceptionism-going-deeper-into-neural.html (2015)."},{"key":"746_CR15","unstructured":"Szegedy, C. et al. Proc. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2015)."},{"key":"746_CR16","doi-asserted-by":"publisher","first-page":"15982","DOI":"10.1038\/s41598-017-16316-2","volume":"7","author":"K Suzuki","year":"2017","unstructured":"Suzuki, K., Roseboom, W., Schwartzman, D. J. & Seth, A. K. A deep-dream virtual reality platform for studying altered perceptual phenomenology. Sci. Rep. 7, 15982 (2017).","journal-title":"Sci. Rep."},{"key":"746_CR17","doi-asserted-by":"publisher","first-page":"839","DOI":"10.3390\/e23070839","volume":"23","author":"A Greco","year":"2021","unstructured":"Greco, A., Gallitto, G., D\u2019Alessandro, M. & Rastelli, C. Increased entropic brain dynamics during deepdream-induced altered perceptual phenomenology. Entropy 23, 839 (2021).","journal-title":"Entropy"},{"key":"746_CR18","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-022-08047-w","volume":"12","author":"C Rastelli","year":"2022","unstructured":"Rastelli, C., Greco, A., Kenett, Y. N., Finocchiaro, C. & De Pisapia, N. Simulated visual hallucinations in virtual reality enhance cognitive flexibility. Sci. Rep. 12, 4027 (2022).","journal-title":"Sci. Rep."},{"key":"746_CR19","doi-asserted-by":"publisher","first-page":"2060","DOI":"10.1038\/s41593-019-0517-x","volume":"22","author":"EY Walker","year":"2019","unstructured":"Walker, E. Y. et al. Inception loops discover what excites neurons most using deep predictive models. Nat. Neurosci. 22, 2060\u20132065 (2019).","journal-title":"Nat. Neurosci."},{"key":"746_CR20","doi-asserted-by":"publisher","first-page":"e1007973","DOI":"10.1371\/journal.pcbi.1007973","volume":"16","author":"W Xiao","year":"2020","unstructured":"Xiao, W. & Kreiman, G. XDream: finding preferred stimuli for visual neurons using generative networks and gradient-free optimization. PLoS Comput. Biol. 16, e1007973 (2020).","journal-title":"PLoS Comput. Biol."},{"key":"746_CR21","unstructured":"Gatys, L. A., Ecker, A. S. & Bethge, M. Proc. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2016)."},{"key":"746_CR22","unstructured":"Gatys, L., Ecker, A. S. & Bethge, M. Texture synthesis using convolutional neural networks. In Proc. Advances in Neural Information Processing Systems (eds Cortes, C. et al.) 262\u2013270 (Curran Associates, Inc., 2015)."},{"key":"746_CR23","doi-asserted-by":"crossref","unstructured":"Johnson, J., Alahi, A. & Fei-Fei, L. Perceptual losses for real-time style transfer and super-resolution. In Computer Vision \u2013 ECCV 2016 (eds Leibe, B. et al.) 694\u2013711 (Springer, 2016).","DOI":"10.1007\/978-3-319-46475-6_43"},{"key":"746_CR24","unstructured":"Gatys, L. A., Ecker, A. S., Bethge, M., Hertzmann, A. & Shechtman, E. Proc. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2017)."},{"key":"746_CR25","doi-asserted-by":"publisher","first-page":"5","DOI":"10.1167\/17.12.5","volume":"17","author":"TS Wallis","year":"2017","unstructured":"Wallis, T. S. et al. A parametric texture model based on deep convolutional features closely matches texture appearance for humans. J. Vis. 17, 5 (2017).","journal-title":"J. Vis."},{"key":"746_CR26","doi-asserted-by":"publisher","first-page":"105621","DOI":"10.1016\/j.cognition.2023.105621","volume":"241","author":"EO Nadler","year":"2023","unstructured":"Nadler, E. O. et al. Divergences in color perception between deep neural networks and humans. Cognition 241, 105621 (2023).","journal-title":"Cognition"},{"key":"746_CR27","doi-asserted-by":"publisher","first-page":"15","DOI":"10.1038\/s41593-018-0284-0","volume":"22","author":"MH Turner","year":"2019","unstructured":"Turner, M. H., Sanchez Giraldo, L. G., Schwartz, O. & Rieke, F. Stimulus- and goal-oriented frameworks for understanding natural vision. Nat. Neurosci. 22, 15\u201324 (2019).","journal-title":"Nat. Neurosci."},{"key":"746_CR28","doi-asserted-by":"publisher","first-page":"199","DOI":"10.1016\/j.conb.2019.09.009","volume":"58","author":"A Pasupathy","year":"2019","unstructured":"Pasupathy, A., Kim, T. & Popovkina, D. V. Object shape and surface properties are jointly encoded in mid-level ventral visual cortex. Curr. Opin. Neurobiol. 58, 199\u2013208 (2019).","journal-title":"Curr. Opin. Neurobiol."},{"key":"746_CR29","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2115302119","volume":"119","author":"AV Jagadeesh","year":"2022","unstructured":"Jagadeesh, A. V. & Gardner, J. L. Texture-like representation of objects in human visual cortex. Proc. Natl Acad. Sci. USA 119, e2115302119 (2022).","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"746_CR30","doi-asserted-by":"publisher","first-page":"10","DOI":"10.1167\/14.4.10","volume":"14","author":"EI Nitzany","year":"2014","unstructured":"Nitzany, E. I. & Victor, J. D. The statistics of local motion signals in naturalistic movies. J. Vis. 14, 10 (2014).","journal-title":"J. Vis."},{"key":"746_CR31","unstructured":"Sinno, Z. & Bovik, A. C. Proc. 2019 IEEE International Conference on Image Processing (ICIP) (IEEE, 2019)."},{"key":"746_CR32","unstructured":"Funke, C. M., Gatys, L. A., Ecker, A. S. & Bethge, M. Synthesising dynamic textures using convolutional neural networks. Preprint at https:\/\/arxiv.org\/abs\/1702.07006v1 (2017)."},{"key":"746_CR33","unstructured":"Tesfaldet, M., Brubaker, M. A. & Derpanis, K. G. Proc. 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (IEEE, 2018)."},{"key":"746_CR34","doi-asserted-by":"publisher","first-page":"2017","DOI":"10.1038\/s41593-023-01442-0","volume":"26","author":"J Feather","year":"2023","unstructured":"Feather, J., Leclerc, G., M\u0105dry, A. & McDermott, J. H. Model metamers reveal divergent invariances between biological and artificial neural networks. Nat. Neurosci. 26, 2017\u20132034 (2023).","journal-title":"Nat. Neurosci."},{"key":"746_CR35","unstructured":"Simonyan, K. & Zisserman, A. Two-stream convolutional networks for action recognition in videos. In Proc. Advances in Neural Information Processing Systems (eds Ghahramani, Z. et al.) 568\u2013576 (Curran Associates, Inc., 2014)."},{"key":"746_CR36","unstructured":"Feichtenhofer, C., Pinz, A. & Wildes, R. Proc. Advances in Neural Information Processing Systems (Curran Associates, Inc., 2016)."},{"key":"746_CR37","doi-asserted-by":"publisher","first-page":"154","DOI":"10.1038\/349154a0","volume":"349","author":"MA Goodale","year":"1991","unstructured":"Goodale, M. A., Milner, A. D., Jakobson, L. S. & Carey, D. P. A neurological dissociation between perceiving objects and grasping them. Nature 349, 154\u2013156 (1991).","journal-title":"Nature"},{"key":"746_CR38","doi-asserted-by":"publisher","first-page":"20","DOI":"10.1016\/0166-2236(92)90344-8","volume":"15","author":"MA Goodale","year":"1992","unstructured":"Goodale, M. A. & Milner, A. D. Separate visual pathways for perception and action. Trends Neurosci. 15, 20\u201325 (1992).","journal-title":"Trends Neurosci."},{"key":"746_CR39","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1016\/S0959-4388(98)80042-1","volume":"8","author":"VA Lamme","year":"1998","unstructured":"Lamme, V. A., Sup\u00e8r, H. & Spekreijse, H. Feedforward, horizontal and feedback processing in the visual cortex. Curr. Opin. Neurobiol. 8, 529\u2013535 (1998).","journal-title":"Curr. Opin. Neurobiol."},{"key":"746_CR40","doi-asserted-by":"publisher","first-page":"1019","DOI":"10.1038\/14819","volume":"2","author":"M Riesenhuber","year":"1999","unstructured":"Riesenhuber, M. & Poggio, T. Hierarchical models of object recognition in cortex. Nat. Neurosci. 2, 1019\u20131025 (1999).","journal-title":"Nat. Neurosci."},{"key":"746_CR41","doi-asserted-by":"publisher","first-page":"8619","DOI":"10.1073\/pnas.1403112111","volume":"111","author":"DLK Yamins","year":"2014","unstructured":"Yamins, D. L. K. et al. Performance-optimized hierarchical models predict neural responses in higher visual cortex. Proc. Natl Acad. Sci. USA 111, 8619\u20138624 (2014).","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"746_CR42","doi-asserted-by":"publisher","first-page":"356","DOI":"10.1038\/nn.4244","volume":"19","author":"DL Yamins","year":"2016","unstructured":"Yamins, D. L. & DiCarlo, J. J. Using goal-driven deep learning models to understand sensory cortex. Nat. Neurosci. 19, 356\u2013365 (2016).","journal-title":"Nat. Neurosci."},{"key":"746_CR43","unstructured":"Simonyan, K. & Zisserman, A. Proc. 3rd International Conference on Learning Representations (ICLR, 2015)."},{"key":"746_CR44","doi-asserted-by":"publisher","first-page":"1193","DOI":"10.1109\/TPAMI.2011.221","volume":"34","author":"KGP Derpanis","year":"2011","unstructured":"Derpanis, K. G. P. & Wildes, R. Spacetime texture representation and recognition based on a spatiotemporal orientation analysis. IEEE Trans. Pattern Anal. Mach. Intell. 34, 1193\u20131205 (2011).","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"746_CR45","doi-asserted-by":"publisher","first-page":"259","DOI":"10.1016\/0167-2789(92)90242-F","volume":"60","author":"LI Rudin","year":"1992","unstructured":"Rudin, L. I., Osher, S. & Fatemi, E. Nonlinear total variation based noise removal algorithms. Physica D 60, 259\u2013268 (1992).","journal-title":"Physica D"},{"key":"746_CR46","doi-asserted-by":"publisher","DOI":"10.23915\/distill.00007","volume":"2","author":"C Olah","year":"2017","unstructured":"Olah, C., Mordvintsev, A. & Schubert, L. Feature visualization. Distill 2, e7 (2017).","journal-title":"Distill"},{"key":"746_CR47","doi-asserted-by":"publisher","first-page":"34","DOI":"10.1109\/38.946629","volume":"21","author":"E Reinhard","year":"2001","unstructured":"Reinhard, E., Adhikhmin, M., Gooch, B. & Shirley, P. Color transfer between images. IEEE Comput. Graph. Appl. 21, 34\u201341 (2001).","journal-title":"IEEE Comput. Graph. Appl."},{"key":"746_CR48","doi-asserted-by":"publisher","first-page":"123","DOI":"10.1016\/j.cviu.2006.11.011","volume":"107","author":"F Piti\u00e9","year":"2007","unstructured":"Piti\u00e9, F., Kokaram, A. C. & Dahyot, R. Automated colour grading using colour distribution transfer. Comput. Vis. Image Underst. 107, 123\u2013137 (2007).","journal-title":"Comput. Vis. Image Underst."},{"key":"746_CR49","unstructured":"Abu-El-Haija, S. et al. YouTube-8M: a large-scale video classification benchmark. Preprint at https:\/\/arxiv.org\/abs\/1609.08675 (2016)."},{"key":"746_CR50","doi-asserted-by":"publisher","first-page":"10645","DOI":"10.1523\/JNEUROSCI.3663-13.2014","volume":"34","author":"K Vinken","year":"2014","unstructured":"Vinken, K., Vermaercke, B. & Op De Beeck, H. P. Visual categorization of natural movies by rats. J. Neurosci. 34, 10645\u201310658 (2014).","journal-title":"J. Neurosci."},{"key":"746_CR51","unstructured":"Kay, W. et al. The kinetics human action video dataset. Preprint at https:\/\/arxiv.org\/abs\/1705.06950 (2017)."},{"key":"746_CR52","doi-asserted-by":"crossref","unstructured":"Zeiler, M. D. & Fergus, R. Visualizing and understanding convolutional networks. In Proc. 13th European Conference on Computer Vision (eds Fleet, D. et al.) 818\u2013833 (Springer, 2014).","DOI":"10.1007\/978-3-319-10590-1_53"},{"key":"746_CR53","unstructured":"Kornblith, S., Norouzi, M., Lee, H. & Hinton, G. Proc. 36th International Conference on Machine Learning (PMLR, 2019)."},{"key":"746_CR54","unstructured":"He, K., Zhang, X., Ren, S. & Sun, J. Proc. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2016)."},{"key":"746_CR55","unstructured":"Liu, Z. et al. Proc. 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2022)."},{"key":"746_CR56","unstructured":"Tran, D. et al. Proc. 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (IEEE, 2018)."},{"key":"746_CR57","doi-asserted-by":"crossref","unstructured":"Xie, S., Sun, C., Huang, J., Tu, Z. & Murphy, K. Rethinking spatiotemporal feature learning: speed-accuracy trade-offs in video classification. In Proc. European Conference on Computer Vision (ECCV) (eds Ferrari, V. et al.) 318\u2013335 (Springer, 2018).","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"746_CR58","unstructured":"Millidge, B., Seth, A. & Buckley, C. L. Predictive coding: a theoretical and experimental review. Preprint at https:\/\/arxiv.org\/abs\/2107.12979 (2022)."},{"key":"746_CR59","unstructured":"Salvatori, T. et al. Brain-inspired computational intelligence via predictive coding. Preprint at https:\/\/arxiv.org\/abs\/2308.07870 (2023)."},{"key":"746_CR60","unstructured":"Lotter, W., Kreiman, G. & Cox, D. Proc. International Conference on Learning Representations (ICLR, 2017)."},{"key":"746_CR61","doi-asserted-by":"publisher","first-page":"69273","DOI":"10.1109\/ACCESS.2020.2987281","volume":"8","author":"Y Zhou","year":"2020","unstructured":"Zhou, Y., Dong, H. & El Saddik, A. Deep learning in next-frame prediction: a benchmark review. IEEE Access 8, 69273\u201369283 (2020).","journal-title":"IEEE Access"},{"key":"746_CR62","doi-asserted-by":"publisher","first-page":"210","DOI":"10.1038\/s42256-020-0170-9","volume":"2","author":"W Lotter","year":"2020","unstructured":"Lotter, W., Kreiman, G. & Cox, D. A neural network trained for prediction mimics diverse features of biological neurons and perception. Nat. Mach. Intell. 2, 210\u2013219 (2020).","journal-title":"Nat. Mach. Intell."},{"key":"746_CR63","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.2014196118","volume":"118","author":"C Zhuang","year":"2021","unstructured":"Zhuang, C. et al. Unsupervised neural network models of the ventral visual stream. Proc. Natl Acad. Sci. USA 118, e2014196118 (2021).","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"746_CR64","unstructured":"Rane, R. P. Sz\u00fcgyi, E., Saxena, V., Ofner, A. & Stober, S. Proc. 2020 International Conference on Multimedia Retrieval (ACM, 2020)."},{"key":"746_CR65","doi-asserted-by":"publisher","first-page":"600","DOI":"10.1109\/TIP.2003.819861","volume":"13","author":"Z Wang","year":"2004","unstructured":"Wang, Z., Bovik, A. C., Sheikh, H. R. & Simoncelli, E. P. Image quality assessment: from error visibility to structural similarity. IEEE Trans. Image Process. 13, 600\u2013612 (2004).","journal-title":"IEEE Trans. Image Process."},{"key":"746_CR66","doi-asserted-by":"publisher","first-page":"861","DOI":"10.21105\/joss.00861","volume":"3","author":"L McInnes","year":"2018","unstructured":"McInnes, L., Healy, J., Saul, N. & Gro\u00dfberger, L. UMAP: Uniform Manifold Approximation and Projection. J. Open Source Softw. 3, 861 (2018).","journal-title":"J. Open Source Softw."},{"key":"746_CR67","unstructured":"Huang, H. et al. Proc. 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2017)."},{"key":"746_CR68","doi-asserted-by":"publisher","first-page":"1199","DOI":"10.1007\/s11263-018-1089-z","volume":"126","author":"M Ruder","year":"2018","unstructured":"Ruder, M., Dosovitskiy, A. & Brox, T. Artistic style transfer for videos and spherical images. Int. J. Comput. Vis. 126, 1199\u20131219 (2018).","journal-title":"Int. J. Comput. Vis."},{"key":"746_CR69","unstructured":"Gao, W., Li, Y., Yin, Y. & Yang, M.-H. Proc. IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV) (IEEE, 2020)."},{"key":"746_CR70","doi-asserted-by":"publisher","first-page":"29330","DOI":"10.1073\/pnas.1912334117","volume":"117","author":"T Golan","year":"2020","unstructured":"Golan, T., Raju, P. C. & Kriegeskorte, N. Controversial stimuli: pitting neural networks against each other as models of human cognition. Proc. Natl Acad. Sci. USA 117, 29330\u201329337 (2020).","journal-title":"Proc. Natl Acad. Sci. USA"},{"key":"746_CR71","unstructured":"Golan, T., Guo, W., Sch\u00fctt, H. H. & Kriegeskorte, N. Proc. SVRHM 2022 Workshop at NeurIPS (International Conference on Neural Information Processing Systems, 2022); https:\/\/neurips.cc\/virtual\/2022\/65923"},{"key":"746_CR72","unstructured":"Gaziv, G., Lee, M. J. & DiCarlo, J. J. Proc. 37th International Conference on Neural Information Processing Systems (Curran Associates, Inc., 2024)."},{"key":"746_CR73","doi-asserted-by":"publisher","first-page":"e1003963","DOI":"10.1371\/journal.pcbi.1003963","volume":"10","author":"CF Cadieu","year":"2014","unstructured":"Cadieu, C. F. et al. Deep neural networks rival the representation of primate IT cortex for core visual object recognition. PLoS Comput. Biol. 10, e1003963 (2014).","journal-title":"PLoS Comput. Biol."},{"key":"746_CR74","unstructured":"Vaswani, A. et al. Attention is all you need. In Proc. Advances in Neural Information Processing Systems (eds Guyon, I. et al.) 6000\u20136010 (Curran Associates, Inc., 2017)."},{"key":"746_CR75","unstructured":"Dosovitskiy, A. et al. Proc. International Conference on Learning Representations (ICLR, 2021)."},{"key":"746_CR76","unstructured":"Raghu, M., Unterthiner, T., Kornblith, S., Zhang, C. & Dosovitskiy, A. Do vision transformers see like convolutional neural networks? In Proc. Advances in Neural Information Processing Systems (eds Ranzato, M. et al.) 12116\u201312128 (Curran Associates, Inc., 2021)."},{"key":"746_CR77","doi-asserted-by":"crossref","unstructured":"Bakhtiari, S., Mineault, P., Lillicrap, T., Pack, C. & Richards, B. The functional specialization of visual cortex emerges from training parallel pathways with self-supervised predictive learning. In Proc. Advances in Neural Information Processing Systems (eds Ranzato, M. et al.) 25164\u201325178 (Curran Associates, Inc., 2021).","DOI":"10.1101\/2021.06.18.448989"},{"key":"746_CR78","doi-asserted-by":"crossref","unstructured":"Mineault, P., Bakhtiari, S., Richards, B. & Pack, C. Your head is there to move you around: goal-driven models of the primate dorsal pathway. In Proc. Advances in Neural Information Processing Systems (eds Ranzato, M. e al.) 28757\u201328771 (Curran Associates, Inc., 2021).","DOI":"10.1101\/2021.07.09.451701"},{"key":"746_CR79","doi-asserted-by":"publisher","first-page":"429","DOI":"10.1098\/rstb.1992.0119","volume":"337","author":"A Verri","year":"1992","unstructured":"Verri, A., Straforini, M. & Torre, V. Computational aspects of motion perception in natural and artificial vision systems. Phil. Trans. R. Soc. B 337, 429\u2013443 (1992).","journal-title":"Phil. Trans. R. Soc. B"},{"key":"746_CR80","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1038\/nrn1057","volume":"4","author":"MA Giese","year":"2003","unstructured":"Giese, M. A. & Poggio, T. Neural mechanisms for the recognition of biological movements. Nat. Rev. Neurosci. 4, 179\u2013192 (2003).","journal-title":"Nat. Rev. Neurosci."},{"key":"746_CR81","unstructured":"Deng, J. et al. Proc. 2009 IEEE Conference on Computer Vision and Pattern Recognition (IEEE, 2009)."},{"key":"746_CR82","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky, O. et al. ImageNet large scale visual recognition challenge. Int. J. Comput. Vis. 115, 211\u2013252 (2015).","journal-title":"Int. J. Comput. Vis."},{"key":"746_CR83","doi-asserted-by":"publisher","DOI":"10.23915\/distill.00012","volume":"3","author":"A Mordvintsev","year":"2018","unstructured":"Mordvintsev, A., Pezzotti, N., Schubert, L. & Olah, C. Differentiable image parameterizations. Distill 3, e12 (2018).","journal-title":"Distill"},{"key":"746_CR84","doi-asserted-by":"publisher","DOI":"10.23915\/distill.00003","volume":"1","author":"A Odena","year":"2016","unstructured":"Odena, A., Dumoulin, V. & Olah, C. Deconvolution and checkerboard artifacts. Distill 1, e3 (2016).","journal-title":"Distill"},{"key":"746_CR85","unstructured":"Soomro, K., Zamir, A. R. & Shah, M. UCF101: a dataset of 101 human actions classes from videos in the wild. Preprint at https:\/\/arxiv.org\/abs\/1212.0402 (2012)."},{"key":"746_CR86","unstructured":"Mahendran, A. & Vedaldi, A. Proc. 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE, 2015)."},{"key":"746_CR87","doi-asserted-by":"publisher","first-page":"1143","DOI":"10.1109\/TIP.2005.864170","volume":"15","author":"D Coltuc","year":"2006","unstructured":"Coltuc, D., Bolon, P. & Chassery, J.-M. Exact histogram specification. IEEE Trans. Image Process. 15, 1143\u20131152 (2006).","journal-title":"IEEE Trans. Image Process."},{"key":"746_CR88","doi-asserted-by":"crossref","unstructured":"Farneb\u00e4ck, G. in Image Analysis Vol. 2749 (eds Bigun, J. & Gustavsson, T.) 363\u2013370 (Springer, 2003).","DOI":"10.1007\/3-540-45103-X_50"},{"key":"746_CR89","unstructured":"torchvision: PyTorch\u2019s Computer Vision library. GitHub https:\/\/github.com\/pytorch\/vision (2016)."},{"key":"746_CR90","doi-asserted-by":"publisher","first-page":"1231","DOI":"10.1177\/0278364913491297","volume":"32","author":"A Geiger","year":"2013","unstructured":"Geiger, A., Lenz, P., Stiller, C. & Urtasun, R. Vision meets robotics: the KITTI dataset. Int. J. Robot. Res. 32, 1231\u20131237 (2013).","journal-title":"Int. J. Robot. Res."},{"key":"746_CR91","doi-asserted-by":"crossref","unstructured":"Reimers, N. & Gurevych, I. Sentence-BERT: sentence embeddings using Siamese BERT-Networks. In Proc. 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (eds Inui, K. et al.) 3982\u20133992 (ACL, 2019).","DOI":"10.18653\/v1\/D19-1410"},{"key":"746_CR92","doi-asserted-by":"publisher","unstructured":"Greco, A. antoninogreco\/STST: STST v1.0. Zenodo https:\/\/doi.org\/10.5281\/zenodo.14168471 (2024).","DOI":"10.5281\/zenodo.14168471"}],"container-title":["Nature Computational Science"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.nature.com\/articles\/s43588-024-00746-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s43588-024-00746-w","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.nature.com\/articles\/s43588-024-00746-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,25]],"date-time":"2025-02-25T23:02:38Z","timestamp":1740524558000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.nature.com\/articles\/s43588-024-00746-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,12,20]]},"references-count":92,"journal-issue":{"issue":"2","published-online":{"date-parts":[[2025,2]]}},"alternative-id":["746"],"URL":"https:\/\/doi.org\/10.1038\/s43588-024-00746-w","relation":{},"ISSN":["2662-8457"],"issn-type":[{"value":"2662-8457","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,12,20]]},"assertion":[{"value":"24 April 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"21 November 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 December 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The authors declare no competing interests.","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Competing interests"}}]}}