{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,2,18]],"date-time":"2025-02-18T02:40:23Z","timestamp":1739846423334,"version":"3.37.3"},"reference-count":38,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T00:00:00Z","timestamp":1737936000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T00:00:00Z","timestamp":1737936000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/100018341","name":"Leuphana Universit\u00e4t L\u00fcneburg","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100018341","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[2025,2]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>Automatically labeling trajectories of multiple agents is key to behavioral analyses but usually requires a large amount of manual annotations. This also applies to the domain of team sport analyses. In this paper, we specifically show how pretraining transformer models improves the classification performance on tracking data from professional soccer. For this purpose, we propose a novel self-supervised masked autoencoder for multiagent trajectories to effectively learn from only a few labeled sequences. Our approach builds upon a factorized transformer architecture for multiagent trajectory data and employs a masking scheme on the level of individual agent trajectories. As a result, our model allows for a reconstruction of masked trajectory segments while being permutation equivariant with respect to the agent trajectories. In addition to experiments on soccer, we demonstrate the usefulness of the proposed pretraining approach on multiagent pose data from entomology. In contrast to related work, our approach is conceptually much simpler, does not require handcrafted features and naturally allows for permutation invariance in downstream tasks.<\/jats:p>","DOI":"10.1007\/s10994-024-06647-3","type":"journal-article","created":{"date-parts":[[2025,1,27]],"date-time":"2025-01-27T21:44:04Z","timestamp":1738014244000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Masked autoencoder for multiagent trajectories"],"prefix":"10.1007","volume":"114","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-5677-3318","authenticated-orcid":false,"given":"Yannick","family":"Rudolph","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ulf","family":"Brefeld","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,1,27]]},"reference":[{"key":"6647_CR1","doi-asserted-by":"crossref","unstructured":"Aksan, E., Kaufmann, M., Cao, P., & Hilliges, O. (2021). A Spatio-temporal Transformer for 3D Human Motion Prediction. International Conference on 3D Vision. arXiv:2004.08692 [cs.CV].","DOI":"10.1109\/3DV53792.2021.00066"},{"key":"6647_CR2","unstructured":"Anzer, G., Bauer, P., Brefeld, U., & Fassmeyer, D. (2022). Detection of tactical patterns using semi-supervised graph neural networks. In MIT Sloan Sports Analytics Conference."},{"key":"6647_CR3","doi-asserted-by":"publisher","first-page":"295","DOI":"10.1007\/s10618-021-00810-3","volume":"36","author":"G Anzer","year":"2022","unstructured":"Anzer, G., & Bauer, P. (2022). Expected passes: Determining the difficulty of a pass in football (soccer) using spatio-temporal data. Data Mining And Knowledge Discovery, 36, 295\u2013317.","journal-title":"Data Mining And Knowledge Discovery"},{"key":"6647_CR4","unstructured":"Ba, J. L., Kiros, J. R., & Hinton, G. E. (2016). Layer Normalization. arXiv preprint arXiv:1607.06450 ."},{"key":"6647_CR5","doi-asserted-by":"publisher","first-page":"2009","DOI":"10.1007\/s10618-021-00763-7","volume":"35","author":"P Bauer","year":"2021","unstructured":"Bauer, P., & Anzer, G. (2021). Data-driven detection of counterpressing in professional football: A supervised machine learning task based on synchronized positional and event data with expert-based feature extraction. Data Mining and Knowledge Discovery, 35, 2009\u20132049.","journal-title":"Data Mining and Knowledge Discovery"},{"issue":"1","key":"6647_CR6","doi-asserted-by":"publisher","first-page":"39","DOI":"10.3233\/JSA-220620","volume":"9","author":"P Bauer","year":"2023","unstructured":"Bauer, P., Anzer, G., & Shaw, L. (2023). Putting team formations in association football into context. Journal of Sports Analytics, 9(1), 39\u201359.","journal-title":"Journal of Sports Analytics"},{"key":"6647_CR7","first-page":"1877","volume":"33","author":"T Brown","year":"2020","unstructured":"Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., \u2026 Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877\u20131901.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"6647_CR8","doi-asserted-by":"crossref","unstructured":"Casas, S., Gulino, C., Suo, S., Luo, K., Liao, R., & Urtasun, R. (2020). Implicit Latent Variable Model for Scene-Consistent Motion Forecasting. In European Conference on Computer Vision.","DOI":"10.1007\/978-3-030-58592-1_37"},{"issue":"2","key":"6647_CR9","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3105576","volume":"3","author":"S Chawla","year":"2017","unstructured":"Chawla, S., Estephan, J., Gudmundsson, J., & Horton, M. (2017). Classification of passes in football matches using spatiotemporal data. ACM Transactions on Spatial Algorithms and Systems, 3(2), 1\u201330.","journal-title":"ACM Transactions on Spatial Algorithms and Systems"},{"key":"6647_CR10","unstructured":"Chen, T., Kornblith, S., Swersky, K., Norouzi, M., & Hinton, G. (2020). Big Self-Supervised Models are Strong Semi-Supervised Learners. In Advances in Neural Information Processing Systems."},{"key":"6647_CR11","doi-asserted-by":"crossref","unstructured":"Chen, H., Wang, J., Shao, K., Liu, F., Hao, J., Guan, C., Chen, G., & Heng, P. A. (2023). Traj-MAE: Masked Autoencoders for Trajectory Prediction. In International Conference on Computer Vision.","DOI":"10.1109\/ICCV51070.2023.00767"},{"key":"6647_CR12","unstructured":"Co-Reyes, J., Liu, Y., Gupta, A., Eysenbach, B., Abbeel, P., & Levine, S. (2018). Self-consistent trajectory autoencoder: Hierarchical reinforcement learning with trajectory embeddings. In International Conference on Machine Learning, pp. 1009\u20131018. PMLR."},{"key":"6647_CR13","unstructured":"Dai, A. M., & Le, Q. V. (2015). Semi-supervised sequence learning. In Advances in Neural Information Processing Systems, Volume\u00a028."},{"key":"6647_CR14","unstructured":"Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018), October. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805. [cs.CL]."},{"key":"6647_CR15","doi-asserted-by":"publisher","DOI":"10.3389\/fspor.2021.682986","volume":"3","author":"U Dick","year":"2021","unstructured":"Dick, U., Tavakol, M., & Brefeld, U. (2021). Rating player actions in soccer. Frontiers in Sports and Active Living, 3, 682986.","journal-title":"Frontiers in Sports and Active Living"},{"key":"6647_CR16","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., & Houlsby, N. (2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In International Conference on Learning Representations."},{"key":"6647_CR17","doi-asserted-by":"crossref","unstructured":"Eyjolfsdottir, E., Branson, S., Burgos-Artizzu, X. P., Hoopfer, E. D., Schor, J., Anderson, D. J., & Perona, P. (2014). Detecting social actions of fruit flies. In European Conference on Computer Vision.","DOI":"10.1007\/978-3-319-10605-2_50"},{"key":"6647_CR18","doi-asserted-by":"publisher","DOI":"10.3389\/fspor.2021.725431","volume":"3","author":"D Fassmeyer","year":"2021","unstructured":"Fassmeyer, D., Anzer, G., Bauer, P., & Brefeld, U. (2021). Toward Automatically Labeling Situations in Soccer. Frontiers in Sports and Active Living, 3, 725431.","journal-title":"Frontiers in Sports and Active Living"},{"key":"6647_CR19","unstructured":"Girgis, R., Golemo, F., Codevilla, F., Weiss, M., D\u2019Souza, J. A., Kahou, S. E., Heide, F., & Pal, C. (2022). Latent variable sequential set transformers for joint multi-agent motion prediction."},{"key":"6647_CR20","doi-asserted-by":"crossref","unstructured":"He, K., Chen, X., Xie, S., Li, Y., Doll\u00e1r, P., & Girshick, R. (2022). Masked Autoencoders Are Scalable Vision Learners. In IEEE Conference on Computer Vision and Pattern Recognition.","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"6647_CR21","doi-asserted-by":"crossref","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., & Girshick, R. (2020). Momentum contrast for unsupervised visual representation learning. In IEEE Conference on Computer Vision and Pattern Recognition, pp. 9729\u20139738.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"6647_CR22","unstructured":"Henaff, O. (2020). Data-efficient image recognition with contrastive predictive coding. In International Conference on Machine Learning, pp. 4182\u20134192. PMLR."},{"key":"6647_CR23","unstructured":"Kingma, D. P., & Ba, J. L. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 ."},{"key":"6647_CR24","unstructured":"Pascanu, R., Mikolov, T., & Bengio, Y. (2013). On the difficulty of training recurrent neural networks. In International Conference on Machine Learning, pp. 1310\u20131318."},{"key":"6647_CR25","first-page":"8024","volume":"32","author":"A Paszke","year":"2019","unstructured":"Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., K\u00f6pf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., \u2026 Chintala, S. (2019). PyTorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems, 32, 8024\u20138035.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"6647_CR26","doi-asserted-by":"crossref","unstructured":"Pettersen, S. A., Halvorsen, P., Johansen, D., Johansen, H., Berg-Johansen, V., Gaddam, V. R., Mortensen, A., Langseth, R., Griwodz, C., & Stensland, H. K. (2014). Soccer Video and Player Position Dataset. In ACM Multimedia Systems Conference, pp. 18\u201323.","DOI":"10.1145\/2557642.2563677"},{"key":"6647_CR27","doi-asserted-by":"crossref","unstructured":"Power, P., Ruiz, H., Wei, X., & Lucey, P. (2017). Not all passes are created equal: Objectively measuring the risk and reward of passes in soccer from tracking data. In SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 1605\u20131613.","DOI":"10.1145\/3097983.3098051"},{"key":"6647_CR28","doi-asserted-by":"crossref","unstructured":"Sanford, R., Gorji, S., Hafemann, L. G., Pourbabaee, B., & Javan, M. (2020). Group Activity Detection From Trajectory and Video Data in Soccer. In (Workshop) IEEE Conference on Computer Vision and Pattern Recognition, pp. 898\u2013899.","DOI":"10.1109\/CVPRW50498.2020.00457"},{"issue":"1","key":"6647_CR29","first-page":"1929","volume":"15","author":"N Srivastava","year":"2014","unstructured":"Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., & Salakhutdinov, R. (2014). Dropout: a simple way to prevent neural networks from overfitting. The Journal of Machine Learning Research, 15(1), 1929\u20131958.","journal-title":"The Journal of Machine Learning Research"},{"key":"6647_CR30","unstructured":"St\u00f6ckl, M., Seidl, T., Marley, D., & Power, P. (2021). Making Offensive Play Predictable \u2013 Using a Graph Convolutional Network to Understand Defensive Performance in Soccer. In MIT Sloan Sports Analytics Conference."},{"key":"6647_CR31","doi-asserted-by":"crossref","unstructured":"Sun, J. J., Kennedy, A., Zhan, E., Anderson, D. J., Yue, Y., & Perona, P. (2021). Task Programming: Learning Data Efficient Behavior Representations. In IEEE Conference on Computer Vision and Pattern Recognition.","DOI":"10.1109\/CVPR46437.2021.00290"},{"key":"6647_CR32","unstructured":"van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., & Kavukcuoglu, K. (2016). Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 ."},{"key":"6647_CR33","first-page":"5998","volume":"30","author":"A Vaswani","year":"2017","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, \u0141, & Polosukhin, I. (2017). Attention Is all you need. Advances in Neural Information Processing Systems, 30, 5998\u20136008.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"6647_CR34","doi-asserted-by":"crossref","unstructured":"Vincent, P., Larochelle, H., Bengio, Y., & Manzagol, P. A. (2008). Extracting and composing robust features with denoising autoencoders. In International Conference on Machine Learning, pp. 1096\u20131103. ACM.","DOI":"10.1145\/1390156.1390294"},{"key":"6647_CR35","first-page":"3371","volume":"11","author":"P Vincent","year":"2010","unstructured":"Vincent, P., Larochelle, H., Lajoie, I., Bengio, Y., & Manzagol, P. A. (2010). Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion. Journal of Machine Learning Research, 11, 3371\u20133408.","journal-title":"Journal of Machine Learning Research"},{"key":"6647_CR36","doi-asserted-by":"crossref","unstructured":"Yeh, R. A., Schwing, A. G., Huang, J., & Murphy, K. (2019). Diverse Generation for Multi-Agent Sports Games. In IEEE Conference on Computer Vision and Pattern Recognition.","DOI":"10.1109\/CVPR.2019.00474"},{"key":"6647_CR37","first-page":"3391","volume":"30","author":"M Zaheer","year":"2017","unstructured":"Zaheer, M., Kottur, S., Ravanbakhsh, S., Pocz\u00f3s, B., Salakhutdinov, R. R., & Smola, A. J. (2017). Deep sets. Advances in Neural Information Processing Systems, 30, 3391\u20133401.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"6647_CR38","unstructured":"Zhan, E., Tseng, A., Yue, Y., Swaminathan, A., & Hausknecht, M. (2020). Learning Calibratable Policies using Programmatic Style-Consistency."}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-024-06647-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10994-024-06647-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10994-024-06647-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,2,18]],"date-time":"2025-02-18T02:02:05Z","timestamp":1739844125000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10994-024-06647-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,1,27]]},"references-count":38,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2025,2]]}},"alternative-id":["6647"],"URL":"https:\/\/doi.org\/10.1007\/s10994-024-06647-3","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"type":"print","value":"0885-6125"},{"type":"electronic","value":"1573-0565"}],"subject":[],"published":{"date-parts":[[2025,1,27]]},"assertion":[{"value":"5 October 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"31 May 2024","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 October 2024","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"27 January 2025","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors have no conflict of interest to declare that are relevant to the content of this article.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethics approval"}},{"value":"Not applicable.","order":4,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent to participate"}},{"value":"Not applicable.","order":5,"name":"Ethics","group":{"name":"EthicsHeading","label":"Consent for publication"}},{"value":"Yannick Rudolph performed large part of his work on this article while employed at SAP SE, Berlin.","order":6,"name":"Ethics","group":{"name":"EthicsHeading","label":"Employment"}}],"article-number":"44"}}