{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,11]],"date-time":"2026-07-11T22:13:55Z","timestamp":1783808035289,"version":"3.55.0"},"reference-count":93,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2024,5,7]],"date-time":"2024-05-07T00:00:00Z","timestamp":1715040000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2024,5,7]],"date-time":"2024-05-07T00:00:00Z","timestamp":1715040000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001659","name":"Deutsche Forschungsgemeinschaft","doi-asserted-by":"crossref","award":["EXC 2117 - 422037984"],"award-info":[{"award-number":["EXC 2117 - 422037984"]}],"id":[{"id":"10.13039\/501100001659","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100002347","name":"Bundesministerium f\u00fcr Bildung und Forschung","doi-asserted-by":"publisher","award":["01IS23046B"],"award-info":[{"award-number":["01IS23046B"]}],"id":[{"id":"10.13039\/501100002347","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2024,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Markerless methods for animal posture tracking have been rapidly developing recently, but frameworks and benchmarks for tracking large animal groups in 3D are still lacking. To overcome this gap in the literature, we present\u00a03D-MuPPET, a framework to estimate and track 3D poses of up to 10 pigeons at interactive speed using multiple camera views. We train a pose estimator to infer 2D keypoints and bounding boxes of multiple pigeons, then triangulate the keypoints to 3D. For identity matching of individuals in all views, we first dynamically match 2D detections to global identities in the first frame, then use a 2D tracker to maintain IDs across views in subsequent frames. We achieve comparable accuracy to a state of the art 3D pose estimator in terms of median error and Percentage of Correct Keypoints. Additionally, we benchmark the inference speed of\u00a03D-MuPPET, with up to 9.45 fps in 2D and 1.89 fps in 3D, and perform quantitative tracking evaluation, which yields encouraging results. Finally, we showcase two novel applications for 3D-MuPPET. First, we train a model with data of single pigeons and achieve comparable results in 2D and 3D posture estimation for up to 5 pigeons. Second, we show that\u00a03D-MuPPET also works in outdoors without additional annotations from natural environments. Both use cases simplify the domain shift to new species and environments, largely reducing annotation effort needed for 3D posture tracking. To the best of our knowledge we are the first to present a framework for 2D\/3D animal posture and trajectory tracking that works in both indoor and outdoor environments for up to 10 individuals. We hope that the framework can open up new opportunities in studying animal collective behaviour and encourages further developments in 3D multi-animal posture tracking.<\/jats:p>","DOI":"10.1007\/s11263-024-02074-y","type":"journal-article","created":{"date-parts":[[2024,5,7]],"date-time":"2024-05-07T13:02:12Z","timestamp":1715086932000},"page":"4235-4252","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":18,"title":["3D-MuPPET: 3D Multi-Pigeon Pose Estimation and Tracking"],"prefix":"10.1007","volume":"132","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1626-9253","authenticated-orcid":false,"given":"Urs","family":"Waldmann","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5405-7155","authenticated-orcid":false,"given":"Alex Hoi Hang","family":"Chan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7627-1726","authenticated-orcid":false,"given":"Hemal","family":"Naik","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8817-087X","authenticated-orcid":false,"given":"M\u00e1t\u00e9","family":"Nagy","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8556-4558","authenticated-orcid":false,"given":"Iain D.","family":"Couzin","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5803-2185","authenticated-orcid":false,"given":"Oliver","family":"Deussen","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3427-4029","authenticated-orcid":false,"given":"Bastian","family":"Goldluecke","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4534-6630","authenticated-orcid":false,"given":"Fumihiro","family":"Kano","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2024,5,7]]},"reference":[{"issue":"3\u20134","key":"2074_CR1","doi-asserted-by":"publisher","first-page":"227","DOI":"10.1163\/156853974X00534","volume":"49","author":"J Altmann","year":"1974","unstructured":"Altmann, J. (1974). Observational study of behavior: Sampling methods. Behaviour, 49(3\u20134), 227\u2013266.","journal-title":"Behaviour"},{"issue":"1","key":"2074_CR2","doi-asserted-by":"publisher","first-page":"7727","DOI":"10.1038\/s41467-023-43483-w","volume":"14","author":"L An","year":"2023","unstructured":"An, L., Ren, J., Yu, T., Hai, T., Jia, Y., & Liu, Y. (2023). Three-dimensional surface motion capture of multiple freely moving pigs using mammal. Nature Communications, 14(1), 7727.","journal-title":"Nature Communications"},{"issue":"1","key":"2074_CR3","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1016\/j.neuron.2014.09.005","volume":"84","author":"D Anderson","year":"2014","unstructured":"Anderson, D., & Perona, P. (2014). Toward a science of computational ethology. Neuron, 84(1), 18\u201331.","journal-title":"Neuron"},{"key":"2074_CR4","doi-asserted-by":"crossref","unstructured":"Badger, M. , Wang, Y. , Modh, A. , Perkes, A. , Kolotouros, N. , Pfrommer, B.G. , & Daniilidis, K. (2020). 3d bird reconstruction: A dataset, model, and shape recovery from a single view. In European conference on computer vision (pp. 1\u201317).","DOI":"10.1007\/978-3-030-58523-5_1"},{"key":"2074_CR5","doi-asserted-by":"publisher","first-page":"4560","DOI":"10.1038\/s41467-020-18441-5","volume":"11","author":"PC Bala","year":"2020","unstructured":"Bala, P. C., Eisenreich, B. R., Yoo, S. B. M., Hayden, B. Y., Park, H. S., & Zimmermann, J. (2020). Automated markerless pose estimation in freely moving macaques with openmonkeystudio. Nature Communication, 11, 4560.","journal-title":"Nature Communication"},{"key":"2074_CR6","doi-asserted-by":"crossref","unstructured":"Bekuzarov, M. , Bermudez, A. , Lee, J.- Y. , & Li, H. (2023 October). Xmem++: Production-level video segmentation from few annotated frames. In Proceedings of the IEEE\/CVF international conference on computer vision (iccv) (pp.\u00a0635\u2013644).","DOI":"10.1109\/ICCV51070.2023.00065"},{"key":"2074_CR7","doi-asserted-by":"crossref","unstructured":"Berman, G. J. (2018). Measuring behavior across scales. BMC Biology\u00a016(23).","DOI":"10.1186\/s12915-018-0494-7"},{"key":"2074_CR8","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1155\/2008\/246309","volume":"2008","author":"K Bernardin","year":"2008","unstructured":"Bernardin, K., & Stiefelhagen, R. (2008). Evaluating multiple object tracking performance: the clear mot metrics. EURASIP Journal on Image and Video Processing, 2008, 1\u201310.","journal-title":"EURASIP Journal on Image and Video Processing"},{"key":"2074_CR9","volume-title":"The co-ordination and regulation of movements","author":"N Bernshtein","year":"1967","unstructured":"Bernshtein, N. (1967). The co-ordination and regulation of movements. Pergamon Press."},{"key":"2074_CR10","doi-asserted-by":"crossref","unstructured":"Bewley, A. , Ge, Z. , Ott, L. , Ramos, F. , & Upcroft, B. (2016). Simple online and realtime tracking. In IEEE international conference on image processing. (pp. 3464\u20133468).","DOI":"10.1109\/ICIP.2016.7533003"},{"key":"2074_CR11","doi-asserted-by":"crossref","unstructured":"Biggs, B. , Roddick, T. , Fitzgibbon, A. , & Cipolla, R. (2019). Creatures great and smal: Recovering the shape and motion of animals from video. In Proceedings of the Asian conference on computer vision (pp. 3\u201319).","DOI":"10.1007\/978-3-030-20873-8_1"},{"key":"2074_CR12","doi-asserted-by":"publisher","first-page":"378","DOI":"10.1038\/s41592-021-01103-9","volume":"18","author":"LA Bola\u00f1os","year":"2021","unstructured":"Bola\u00f1os, L. A., Xiao, D., Ford, N. L., LeDue, J. M., Gupta, P. K., Doebeli, C., & Murphy, T. H. (2021). A three-dimensional virtual mouse generates synthetic training data for behavioral analysis. Nature Methods, 18, 378\u2013381.","journal-title":"Nature Methods"},{"key":"2074_CR13","doi-asserted-by":"crossref","unstructured":"Bridgeman, L. , Volino, M. , Guillemaut, J.- Y. , & Hilton, A. (2019). Multi-person 3d pose estimation and tracking in sports. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition workshops.","DOI":"10.1109\/CVPRW.2019.00304"},{"issue":"2","key":"2074_CR14","doi-asserted-by":"publisher","first-page":"249","DOI":"10.1037\/h0061438","volume":"25","author":"RD Chard","year":"1938","unstructured":"Chard, R. D., & Gundlach, R. H. (1938). The structure of the eye of the homing pigeon. Journal of Comparative Psychology, 25(2), 249.","journal-title":"Journal of Comparative Psychology"},{"key":"2074_CR15","doi-asserted-by":"crossref","unstructured":"Chen, X. , Zhai, H. , Liu, D. , Li, W. , Ding, C. , Xie, Q. , & Han, H. (2020). Siambomb: A real-time ai-based system for home-cage animal tracking, segmentation and behavioral analysis. In Proceedings of the twenty-ninth international conference on international joint conferences on artificial intelligence (pp. 5300\u20135302).","DOI":"10.24963\/ijcai.2020\/776"},{"key":"2074_CR16","doi-asserted-by":"crossref","unstructured":"Couzin, I.D. , & Heins, C. (2023). Emerging technologies for behavioral research in changing environments. Trends in Ecology & Evolution","DOI":"10.1016\/j.tree.2022.11.008"},{"issue":"7","key":"2074_CR17","doi-asserted-by":"publisher","first-page":"417","DOI":"10.1016\/j.tree.2014.05.004","volume":"29","author":"AI Dell","year":"2014","unstructured":"Dell, A. I., Bender, J. A., Branson, K., Couzin, I. D., de Polavieja, G. G., Noldus, L. P., & Brose, U. (2014). Automated image-based tracking and its application in ecology. Trends in Ecology & Evolution, 29(7), 417\u2013428.","journal-title":"Trends in Ecology & Evolution"},{"key":"2074_CR18","unstructured":"Dendorfer, P. (2020). Motchallengeevalkit. https:\/\/github.com\/dendorferpatrick\/MOTChallengeEvalKit."},{"issue":"4","key":"2074_CR19","doi-asserted-by":"publisher","first-page":"845","DOI":"10.1007\/s11263-020-01393-0","volume":"129","author":"P Dendorfer","year":"2021","unstructured":"Dendorfer, P., Osep, A., Milan, A., Schindler, K., Cremers, D., Reid, I., & Leal-Taix\u00e9, L. (2021). Motchallenge: A benchmark for single-camera multiple target tracking. International Journal of Computer Vision, 129(4), 845\u2013881.","journal-title":"International Journal of Computer Vision"},{"key":"2074_CR20","doi-asserted-by":"crossref","unstructured":"Deng, J. , Dong, W. , Socher, R. , Li, L.- J. , Li, K. , & Fei-Fei, L. (2009). Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition (pp.\u00a0248-255).","DOI":"10.1109\/CVPR.2009.5206848"},{"issue":"5","key":"2074_CR21","doi-asserted-by":"publisher","first-page":"564","DOI":"10.1038\/s41592-021-01106-6","volume":"18","author":"TW Dunn","year":"2021","unstructured":"Dunn, T. W., Marshall, J. D., Severson, K. S., Aldarondo, D. E., Hildebrand, D. G., Chettih, S. N., et al. (2021). Geometric deep learning enables 3d kinematic profiling across species and environments. Nature Methods, 18(5), 564\u2013573.","journal-title":"Nature Methods"},{"issue":"3","key":"2074_CR22","doi-asserted-by":"publisher","first-page":"369","DOI":"10.1002\/rse2.195","volume":"7","author":"I Duporge","year":"2021","unstructured":"Duporge, I., Isupova, O., Reece, S., Macdonald, D. W., & Wang, T. (2021). Using very-high-resolution satellite imagery and deep learning to detect and count African elephants in heterogeneous landscapes. Remote Sensing in Ecology and Conservation, 7(3), 369\u2013381.","journal-title":"Remote Sensing in Ecology and Conservation"},{"issue":"1","key":"2074_CR23","doi-asserted-by":"publisher","first-page":"155","DOI":"10.1038\/s41598-022-25087-4","volume":"13","author":"AS Ebrahimi","year":"2023","unstructured":"Ebrahimi, A. S., Orlowska-Feuer, P., Huang, Q., Zippo, A. G., Martial, F. P., Petersen, R. S., & Storchi, R. (2023). Three-dimensional unsupervised probabilistic pose reconstruction (3d-upper) for freely moving animals. Scientific Reports, 13(1), 155.","journal-title":"Scientific Reports"},{"issue":"9","key":"2074_CR24","doi-asserted-by":"publisher","first-page":"1072","DOI":"10.1111\/2041-210X.13436","volume":"11","author":"AC Ferreira","year":"2020","unstructured":"Ferreira, A. C., Silva, L. R., Renna, F., Brandl, H. B., Renoult, J. P., Farine, D. R., & Doutrelant, C. (2020). Deep learning-based methods for individual recognition in small birds. Methods in Ecology and Evolution, 11(9), 1072\u20131085.","journal-title":"Methods in Ecology and Evolution"},{"key":"2074_CR25","unstructured":"Ferrero, F.R. , Bergomi, M.G. , Heras, F.J. , Hinz, R. , de Polavieja, G.G. , & the Champalimaud\u00a0Foundation. (2017). idtracker.ai. https:\/\/idtrackerai.readthedocs.io\/en\/latest"},{"key":"2074_CR26","doi-asserted-by":"crossref","unstructured":"Giebenhain, S. , Waldmann, U. , Johannsen, O. , & Goldluecke, B. (2022). Neural puppeteer: Keypoint-based neural rendering of dynamic shapes. In Proceedings of the Asian conference on computer vision (ACCV) (pp. 2830\u20132847).","DOI":"10.1007\/978-3-031-26316-3_15"},{"key":"2074_CR27","doi-asserted-by":"publisher","first-page":"1455","DOI":"10.1038\/nn.3812","volume":"17","author":"A Gomez-Marin","year":"2014","unstructured":"Gomez-Marin, A., Paton, J., Kampff, A. R., Costa, R. M., & Mainen, Z. F. (2014). Big behavioral data: Psychology, ethology and the foundations of neuroscience. Nature Neuroscience, 17, 1455\u20131462.","journal-title":"Nature Neuroscience"},{"key":"2074_CR28","doi-asserted-by":"publisher","first-page":"975","DOI":"10.1038\/s41592-021-01226-z","volume":"18","author":"A Gosztolai","year":"2021","unstructured":"Gosztolai, A., G\u00fcnel, S., Lobato-R\u00edos, V., Pietro Abrate, M., Morales, D., Rhodin, H., & Ramdya, P. (2021). Liftpose3d, a deep learning-based approach for transforming two-dimensional to three-dimensional poses in laboratory animals. Nature Methods, 18, 975\u2013981.","journal-title":"Nature Methods"},{"key":"2074_CR29","doi-asserted-by":"publisher","unstructured":"Graving, J.M. , Chae, D. , Naik, H. , Li, L. , Koger, B. , Costelloe, B.R. , & Couzin, I.D. (2019). Deepposekit, a software toolkit for fast and robust animal pose estimation using deep learning. eLife\u00a08, e47994 https:\/\/doi.org\/10.7554\/eLife.47994","DOI":"10.7554\/eLife.47994"},{"key":"2074_CR30","doi-asserted-by":"crossref","unstructured":"G\u00fcnel, S. , Rhodin, H. , Morales, D. , Campagnolo, J. , Ramdya, P. , & Fua, P. (2019). Deepfly3d, a deep learning-based approach for 3d limb and appendage tracking in tethered, adult Drosophila. eLife\u00a08, e48571.","DOI":"10.7554\/eLife.48571"},{"key":"2074_CR31","doi-asserted-by":"crossref","unstructured":"Han, Y. , Chen, K. , Wang, Y. , Liu, W. , Wang, X. , Liao, J. , & et\u00a0al. (2023). Social behavior atlas: A computational framework for tracking and mapping 3d close interactions of free-moving animals. bioRxiv\u00a02023\u201303","DOI":"10.1101\/2023.03.05.531235"},{"key":"2074_CR32","doi-asserted-by":"crossref","unstructured":"He, K. , Gkioxari, G. , Dollar, P. , & Girshick, R. (2017). Mask r-cnn. In Proceedings of the IEEE international conference on computer vision","DOI":"10.1109\/ICCV.2017.322"},{"key":"2074_CR33","doi-asserted-by":"crossref","unstructured":"He, K. , Zhang, X. , Ren, S. , & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2016.90"},{"issue":"9","key":"2074_CR34","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1371\/journal.pcbi.1007354","volume":"15","author":"FJH Heras","year":"2019","unstructured":"Heras, F. J. H., Romero-Ferrero, F., Hinz, R. C., & de Polavieja, G. G. (2019). Deep attention networks reveal the rules of collective motion in zebrafish. PLOS Computational Biology, 15(9), 1\u201323.","journal-title":"PLOS Computational Biology"},{"key":"2074_CR35","doi-asserted-by":"crossref","unstructured":"Huang, C. , Jiang, S. , Li, Y. , Zhang, Z. , Traish, J. , Deng, C. , & Da\u00a0Xu, R.Y. (2020). End-to-end dynamic matching network for multi-view multi-person 3d pose estimation. In Computer vision ECCV 2020: 16th European conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part XXVIII 16 (pp. 477\u2013493).","DOI":"10.1007\/978-3-030-58604-1_29"},{"key":"2074_CR36","doi-asserted-by":"crossref","unstructured":"Ionescu, C., Papava, D., Olaru, V., & Sminchisescu, C. (2014). Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7), 1325\u20131339.","DOI":"10.1109\/TPAMI.2013.248"},{"key":"2074_CR37","doi-asserted-by":"crossref","unstructured":"Iskakov, K. , Burkov, E. , Lempitsky, V. , & Malkov, Y. (2019). Learnable triangulation of human pose. In Proceedings of the IEEE\/CVF international conference on computer vision","DOI":"10.1109\/ICCV.2019.00781"},{"issue":"1","key":"2074_CR38","doi-asserted-by":"publisher","first-page":"1","DOI":"10.2502\/janip.72.1.1","volume":"72","author":"A Itahara","year":"2022","unstructured":"Itahara, A., & Kano, F. (2022). \u201ccorvid tracking studio\u2019\u2019: A custom-built motion capture system to track head movements of corvids. Japanese Journal of Animal Psychology, 72(1), 1\u201316.","journal-title":"Japanese Journal of Animal Psychology"},{"key":"2074_CR39","doi-asserted-by":"publisher","unstructured":"Itahara, A. , & Kano, F. (2023). Gaze tracking of large-billed crows (corvus macrorhynchos) in a motion-capture system. bioRxivhttps:\/\/doi.org\/10.1101\/2023.08.10.552747","DOI":"10.1101\/2023.08.10.552747"},{"key":"2074_CR40","unstructured":"Jocher, G. , Chaurasia, A. , & Qiu, J. (2023). Yolo by ultralytics. https:\/\/github.com\/ultralytics\/ultralytics"},{"key":"2074_CR41","doi-asserted-by":"crossref","unstructured":"Joska, D. , Clark, L. , Muramatsu, N. , Jericevich, R. , Nicolls, F. , Mathis, A. , & Patel, A. (2021). Acinoset: A 3d pose estimation dataset and baseline models for cheetahs in the wild. In 2021 ieee international conference on robotics and automation (icra) (pp. 13901\u201313908).","DOI":"10.1109\/ICRA48506.2021.9561338"},{"issue":"1","key":"2074_CR42","doi-asserted-by":"publisher","first-page":"35","DOI":"10.1115\/1.3662552","volume":"82","author":"RE Kalman","year":"1960","unstructured":"Kalman, R. E. (1960). A new approach to linear filtering and prediction problems. Journal of Basic Engineering, 82(1), 35\u201345.","journal-title":"Journal of Basic Engineering"},{"key":"2074_CR43","doi-asserted-by":"publisher","DOI":"10.7554\/eLife.61909","volume":"9","author":"GA Kane","year":"2020","unstructured":"Kane, G. A., Lopes, G., Saunders, J. L., Mathis, A., & Mathis, M. W. (2020). Real-time, low-latency closed-loop feedback using markerless posture tracking. Elife, 9, e61909.","journal-title":"Elife"},{"issue":"1","key":"2074_CR44","doi-asserted-by":"publisher","first-page":"19113","DOI":"10.1038\/s41598-022-21931-9","volume":"12","author":"F Kano","year":"2022","unstructured":"Kano, F., Naik, H., Keskin, G., Couzin, I. D., & Nagy, M. (2022). Head-tracking of freely-behaving pigeons in a motion-capture system reveals the selective use of visual field regions. Scientific Reports, 12(1), 19113.","journal-title":"Scientific Reports"},{"key":"2074_CR45","unstructured":"Karaev, N. , Rocco, I. , Graham, B. , Neverova, N. , Vedaldi, A. , & Rupprecht, C. (2023). Cotracker: It is better to track together. arXiv preprintarXiv:2307.07635"},{"issue":"13","key":"2074_CR46","doi-asserted-by":"publisher","DOI":"10.1016\/j.celrep.2021.109730","volume":"36","author":"P Karashchuk","year":"2021","unstructured":"Karashchuk, P., Rupp, K. L., Dickinson, E. S., Walling-Bell, S., Sanders, E., Azim, E., & Tuthill, J. C. (2021). Anipose: A toolkit for robust markerless 3d pose estimation. Cell Reports, 36(13), 109730.","journal-title":"Cell Reports"},{"key":"2074_CR47","doi-asserted-by":"crossref","unstructured":"Kays, R. , Crofoot, M.C. , Jetz, W. , & Wikelski, M. (2015). Terrestrial animal tracking as an eye on life and planet. Science\u00a0348(6240), aaa2478","DOI":"10.1126\/science.aaa2478"},{"key":"2074_CR48","doi-asserted-by":"crossref","unstructured":"Kirillov, A. , Mintun, E. , Ravi, N. , Mao, H. , Rolland, C. , Gustafson, L. , & others (2023). Segment anything. arXiv preprintarXiv:2304.02643","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"2074_CR49","doi-asserted-by":"crossref","unstructured":"Koger, B. , Deshpande, A. , Kerby, J.T. , Graving, J.M. , Costelloe, B.R. , & Couzin, I.D. (2023). Quantifying the movement, behaviour and environmental context of group-living animals using drones and computer vision. Journal of Animal Ecology","DOI":"10.1101\/2022.06.30.498251"},{"key":"2074_CR50","doi-asserted-by":"publisher","first-page":"268","DOI":"10.3389\/fnbeh.2020.581154","volume":"14","author":"R Labuguen","year":"2021","unstructured":"Labuguen, R., Matsumoto, J., Negrete, S. B., Nishimaru, H., Nishijo, H., Takada, M., & Shibata, T. (2021). Macaquepose: A novel \u201cin the wild\u2019\u2019 macaque monkey pose dataset for markerless motion capture. Frontiers in Behavioral Neuroscience, 14, 268.","journal-title":"Frontiers in Behavioral Neuroscience"},{"key":"2074_CR51","doi-asserted-by":"publisher","first-page":"496","DOI":"10.1038\/s41592-022-01443-0","volume":"19","author":"J Lauer","year":"2022","unstructured":"Lauer, J., Zhou, M., Ye, S., Menegas, W., Schneider, S., Nath, T., & Mathis, A. (2022). Multi-animal pose estimation, identification and tracking with deeplabcut. Nature Methods, 19, 496\u2013504.","journal-title":"Nature Methods"},{"key":"2074_CR52","doi-asserted-by":"crossref","unstructured":"Li, Y. , Huang, C. , & Nevatia, R. (2009). Learning to associate: Hybridboosted multi-target tracker for crowded scene. In 2009 IEEE conference on computer vision and pattern recognition (p.\u00a02953-2960).","DOI":"10.1109\/CVPR.2009.5206735"},{"key":"2074_CR53","doi-asserted-by":"crossref","unstructured":"Lin, T.- Y. , Dollar, P. , Girshick, R. , He, K. , Hariharan, B. , & Belongie, S. (2017). Feature pyramid networks for object detection. In 2009 IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2017.106"},{"key":"2074_CR54","unstructured":"Luiten, J. , & Hoffhues, A. (2020). Trackeval. https:\/\/github.com\/JonathonLuiten\/TrackEval."},{"issue":"2","key":"2074_CR55","doi-asserted-by":"publisher","first-page":"548","DOI":"10.1007\/s11263-020-01375-2","volume":"129","author":"J Luiten","year":"2021","unstructured":"Luiten, J., Osep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taix\u00e9, L., & Leibe, B. (2021). Hota: A higher order metric for evaluating multi-object tracking. International Journal of Computer Vision, 129(2), 548\u2013578.","journal-title":"International Journal of Computer Vision"},{"key":"2074_CR56","doi-asserted-by":"crossref","unstructured":"Marshall, J.D. , Klibaite, U. , Gellis, A. , Aldarondo, D.E. , \u00d6lveczky, B.P. , & Dunn, T.W. (2021). The pair-r24m dataset for multi-animal 3d pose estimation. bioRxiv\u00a02021\u201311","DOI":"10.1101\/2021.11.23.469743"},{"key":"2074_CR57","doi-asserted-by":"publisher","first-page":"1281","DOI":"10.1038\/s41593-018-0209-y","volume":"21","author":"A Mathis","year":"2018","unstructured":"Mathis, A., Mamidanna, P., Cury, K. M., Abe, T., Murthy, V. N., Mathis, M. W., & Bethge, M. (2018). Deeplabcut: Markerless pose estimation of user-defined body parts with deep learning. Nature Neuroscience, 21, 1281\u20131289.","journal-title":"Nature Neuroscience"},{"issue":"6","key":"2074_CR58","doi-asserted-by":"publisher","first-page":"1497","DOI":"10.1007\/s11263-022-01733-2","volume":"131","author":"S Mi\u00f1ano","year":"2023","unstructured":"Mi\u00f1ano, S., Golodetz, S., Cavallari, T., & Taylor, G. K. (2023). Through hawks\u2019 eyes: synthetically reconstructing the visual field of a bird in flight. International Journal of Computer Vision, 131(6), 1497\u20131531.","journal-title":"International Journal of Computer Vision"},{"issue":"7290","key":"2074_CR59","doi-asserted-by":"publisher","first-page":"890","DOI":"10.1038\/nature08891","volume":"464","author":"M Nagy","year":"2010","unstructured":"Nagy, M., \u00c1kos, Z., Biro, D., & Vicsek, T. (2010). Hierarchical group dynamics in pigeon flocks. Nature, 464(7290), 890\u2013893.","journal-title":"Nature"},{"key":"2074_CR60","doi-asserted-by":"crossref","unstructured":"Nagy, M. , Naik, H. , Fumihiro, K. , Nora, C.V. , Koblitz, J.C. , Wikelski, M. , & Couzin, I.D. (2023). Smart-barn: Scalable multimodal arena for real-time tracking behavior of animals in large numbers. Science Advances (in press)","DOI":"10.1126\/sciadv.adf8068"},{"issue":"32","key":"2074_CR61","doi-asserted-by":"publisher","first-page":"13049","DOI":"10.1073\/pnas.1305552110","volume":"110","author":"M Nagy","year":"2013","unstructured":"Nagy, M., V\u00e1s\u00e1rhelyi, G., Pettit, B., Roberts-Mariani, I., Vicsek, T., & Biro, D. (2013). Context-dependent hierarchies in pigeons. Proceedings of the National Academy of Sciences, 110(32), 13049\u201313054.","journal-title":"Proceedings of the National Academy of Sciences"},{"key":"2074_CR62","volume-title":"Xr for all: Closed-loop visual stimulation techniques for human and non-human animals (Dissertation)","author":"H Naik","year":"2021","unstructured":"Naik, H. (2021). Xr for all: Closed-loop visual stimulation techniques for human and non-human animals (Dissertation). M\u00fcnchen: Technische Universit\u00e4t M\u00fcnchen."},{"issue":"5","key":"2074_CR63","doi-asserted-by":"publisher","first-page":"2073","DOI":"10.1109\/TVCG.2020.2973063","volume":"26","author":"H Naik","year":"2020","unstructured":"Naik, H., Bastien, R., Navab, N., & Couzin, I. D. (2020). Animals in virtual environments. IEEE Transactions on Visualization and Computer Graphics, 26(5), 2073\u20132083.","journal-title":"IEEE Transactions on Visualization and Computer Graphics"},{"key":"2074_CR64","doi-asserted-by":"crossref","unstructured":"Naik, H. , Chan, A.H.H. , Yang, J. , Delacoux, M. , Couzin, I.D. , Kano, F. , & Nagy, M. (2023 June). 3d-pop - an automated annotation approach to facilitate markerless 2d-3d tracking of freely moving birds with marker-based motion capture. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp.\u00a021274-21284).","DOI":"10.1109\/CVPR52729.2023.02038"},{"key":"2074_CR65","doi-asserted-by":"publisher","first-page":"2152","DOI":"10.1038\/s41596-019-0176-0","volume":"14","author":"T Nath","year":"2019","unstructured":"Nath, T., Mathis, A., Chen, A. C., Patel, A., Bethge, M., & Mathis, M. W. (2019). Using deeplabcut for 3d markerless pose estimation across species and behaviors. Nature Protocol, 14, 2152\u20132176.","journal-title":"Nature Protocol"},{"key":"2074_CR66","doi-asserted-by":"crossref","unstructured":"Newell, A. , Yang, K. , & Deng, J. (2016). Stacked hourglass networks for human pose estimation. In Computer vision\u2013eccv 2016: 14th European conference, Amsterdam, The Netherlands, October 11\u201314, 2016, proceedings, part viii 14 (pp. 483\u2013499).","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"2074_CR67","doi-asserted-by":"publisher","first-page":"1052","DOI":"10.1038\/s41592-020-0961-2","volume":"17","author":"A Nourizonoz","year":"2020","unstructured":"Nourizonoz, A., Zimmermann, R., Ho, C. L. A., Pellat, S., Ormen, Y., Pr\u00e9vost-Soli\u00e9, C., & Huber, D. (2020). Etholoop: automated closed-loop neuroethology in naturalistic environments. Nature Methods, 17, 1052\u20131059.","journal-title":"Nature Methods"},{"issue":"1","key":"2074_CR68","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pcbi.1009772","volume":"18","author":"M Papadopoulou","year":"2022","unstructured":"Papadopoulou, M., Hildenbrandt, H., Sankey, D. W., Portugal, S. J., & Hemelrijk, C. K. (2022). Self-organization of collective escape in pigeon flocks. PLoS Computational Biology, 18(1), e1009772.","journal-title":"PLoS Computational Biology"},{"key":"2074_CR69","unstructured":"Paszke, A. , Gross, S. , Massa, F. , Lerer, A. , Bradbury, J. , Chanan, G. , & Chintala, S. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in Neural Information Processing Systems."},{"key":"2074_CR70","doi-asserted-by":"crossref","unstructured":"Pedersen, M. , Haurum, J.B. , Bengtson, S.H. , & Moeslund, T.B. (2020). 3d-zef: A 3d zebrafish tracking benchmark dataset. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR42600.2020.00250"},{"key":"2074_CR71","doi-asserted-by":"publisher","first-page":"117","DOI":"10.1038\/s41592-018-0234-5","volume":"16","author":"TD Pereira","year":"2019","unstructured":"Pereira, T. D., Aldarondo, D. E., Willmore, L., Kislin, M., Wang, S.S.-H., Murthy, M., & Shaevitz, J. W. (2019). Fast animal pose estimation using deep neural networks. Nature Methods, 16, 117\u2013125.","journal-title":"Nature Methods"},{"key":"2074_CR72","doi-asserted-by":"publisher","first-page":"486","DOI":"10.1038\/s41592-022-01426-1","volume":"19","author":"TD Pereira","year":"2022","unstructured":"Pereira, T. D., Tabris, N., Matsliah, A., Turner, D. M., Li, J., Ravindranath, S., & Murthy, M. (2022). Sleap: A deep learning system for multi-animal pose tracking. Nature Methods, 19, 486\u2013495.","journal-title":"Nature Methods"},{"key":"2074_CR73","doi-asserted-by":"crossref","unstructured":"Ristani, E. , Solera, F. , Zou, R. , Cucchiara, R. , & Tomasi, C. (2016). Performance measures and a data set for multi-target, multi-camera tracking. In European conference on computer vision (pp. 17\u201335).","DOI":"10.1007\/978-3-319-48881-3_2"},{"key":"2074_CR74","doi-asserted-by":"publisher","first-page":"179","DOI":"10.1038\/s41592-018-0295-5","volume":"16","author":"F Romero-Ferrero","year":"2019","unstructured":"Romero-Ferrero, F., Bergomi, M. G., Hinz, R. C., Heras, F. J. H., & de Polavieja, G. G. (2019). idtracker.ai: tracking all individuals in small or large collectives of unmarked animals. Nature Methods, 16, 179\u2013182.","journal-title":"Nature Methods"},{"key":"2074_CR75","doi-asserted-by":"crossref","unstructured":"Sanakoyeu, A. , Khalidov, V. , McCarthy, M.S. , Vedaldi, A. , & Neverova, N. (2020 June). Transferring dense pose to proximal animal classes. In Proceedings of the ieee\/cvf conference on computer vision and pattern recognition (cvpr).","DOI":"10.1109\/CVPR42600.2020.00528"},{"issue":"1","key":"2074_CR76","doi-asserted-by":"publisher","first-page":"15049","DOI":"10.1038\/ncomms15049","volume":"8","author":"T Sasaki","year":"2017","unstructured":"Sasaki, T., & Biro, D. (2017). Cumulative culture can emerge from collective intelligence in animal groups. Nature Communications, 8(1), 15049.","journal-title":"Nature Communications"},{"key":"2074_CR77","doi-asserted-by":"crossref","unstructured":"Sun, J.J. , Karashchuk, L. , Dravid, A. , Ryou, S. , Fereidooni, S. , Tuthill, J.C. , & others (2023). Bkind-3d: Self-supervised 3d keypoint discovery from multi-view videos. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 9001\u20139010).","DOI":"10.1109\/CVPR52729.2023.00869"},{"key":"2074_CR78","doi-asserted-by":"crossref","unstructured":"Van\u00a0Horn, G. , Branson, S. , Farrell, R. , Haber, S. , Barry, J. , Ipeirotis, P. , & Belongie, S. (2015). Building a bird recognition app and large scale dataset with citizen scientists: The fine print in fine-grained dataset collection. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2015.7298658"},{"key":"2074_CR79","doi-asserted-by":"crossref","unstructured":"Waldmann, U. , Bamberger, J. , Johannsen, O. , Deussen, O. , & Goldl\u00fccke, B. (2022). Improving unsupervised label propagation for pose tracking and video object segmentation. In Dagm German conference on pattern recognition (pp. 230\u2013245).","DOI":"10.1007\/978-3-031-16788-1_15"},{"key":"2074_CR80","doi-asserted-by":"crossref","unstructured":"Waldmann, U. , Johannsen, O. , & Goldluecke, B. (2023). Neural texture puppeteer: A framework for neural geometry and texture rendering of articulated shapes, enabling re-identification at interactive speed. arXiv preprint arXiv:2311.17109","DOI":"10.1109\/WACVW60836.2024.00016"},{"key":"2074_CR81","doi-asserted-by":"crossref","unstructured":"Waldmann, U. , Naik, H. , M\u00e1t\u00e9, N. , Kano, F. , Couzin, I.D. , Deussen, O. , & Goldl\u00fccke, B. (2022). I-muppet: Interactive multi-pigeon pose estimation and tracking. In Dagm German conference on pattern recognition (pp. 513\u2013528).","DOI":"10.1007\/978-3-031-16788-1_31"},{"key":"2074_CR82","doi-asserted-by":"crossref","unstructured":"Walter, T. , & Couzin, I.D. (2021). Trex, a fast multi-animal tracking system with markerless identification, and 2d estimation of posture and visual fields. eLife\u00a010, e64000","DOI":"10.7554\/eLife.64000"},{"key":"2074_CR83","doi-asserted-by":"crossref","unstructured":"Wang, J. , & Yuille, A.L. (2015). Semantic part segmentation using compositional model combining shape and appearance. In Proceedings of the IEEE conference on computer vision and pattern recognition.","DOI":"10.1109\/CVPR.2015.7298788"},{"key":"2074_CR84","doi-asserted-by":"crossref","unstructured":"Wang, P. , Shen, X. , Lin, Z. , Cohen, S. , Price, B. , & Yuille, A.L. (2015). Joint object and part segmentation using deep learned potentials. In Proceedings of the IEEE international conference on computer vision","DOI":"10.1109\/ICCV.2015.184"},{"key":"2074_CR85","unstructured":"Welinder, P. , Branson, S. , Mita, T. , Wah, C. , Schroff, F. , Belongie, S. , & Perona, P. (2010). Caltech-UCSD Birds 200 Tech. Rep. No. CNS-TR-2010-001. California Institute of Technology."},{"key":"2074_CR86","doi-asserted-by":"crossref","unstructured":"Wojke, N. , & Bewley, A. (2018). Deep cosine metric learning for person re-identification. In 2018 IEEE winter conference on applications of computer vision (wacv) (pp. 748\u2013756).","DOI":"10.1109\/WACV.2018.00087"},{"key":"2074_CR87","doi-asserted-by":"crossref","unstructured":"Xiao, B. , Wu, H. , & Wei, Y. (2018). Simple baselines for human pose estimation and tracking. In Proceedings of the European conference on computer vision (ECCV).","DOI":"10.1007\/978-3-030-01231-1_29"},{"key":"2074_CR88","unstructured":"Xu, Y. , Zhang, J. , Zhang, Q. , & Tao, D. (2022). ViTPose: Simple vision transformer baselines for human pose estimation. Advances in Neural Information Processing Systems."},{"key":"2074_CR89","unstructured":"Yang, J. , Gao, M. , Li, Z. , Gao, S. , Wang, F. , & Zheng, F. (2023). Track anything: Segment anything meets videos. arXiv preprint arXiv:2304.11968"},{"issue":"12","key":"2074_CR90","doi-asserted-by":"publisher","first-page":"2878","DOI":"10.1109\/TPAMI.2012.261","volume":"35","author":"Y Yang","year":"2013","unstructured":"Yang, Y., & Ramanan, D. (2013). Articulated human detection with flexible mixtures of parts. IEEE Transactions on Pattern Analysis and Machine Intelligence, 35(12), 2878\u20132890.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"10","key":"2074_CR91","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0140558","volume":"10","author":"M Yomosa","year":"2015","unstructured":"Yomosa, M., Mizuguchi, T., V\u00e1s\u00e1rhelyi, G., & Nagy, M. (2015). Coordinated behaviour in pigeon flocks. Plos One, 10(10), e0140558.","journal-title":"Plos One"},{"key":"2074_CR92","doi-asserted-by":"crossref","unstructured":"Zhang, L., Gao, J., Xiao, Z., & Fan, H. (2023). Animaltrack: A benchmark for multi-animal tracking in the wild. International Journal of Computer Vision, 131(2), 496\u2013513.","DOI":"10.1007\/s11263-022-01711-8"},{"key":"2074_CR93","unstructured":"Zuffi, S. , Rhodin, H. , Park, H.S. , Beery, S. , Kanazawa, A. , Nobuhara, S. , & Zamansky, A. (2023). Cv4animals: Computer vision for animal behavior tracking and modeling. https:\/\/www.cv4animals.com\/"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-024-02074-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-024-02074-y\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-024-02074-y.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,10,4]],"date-time":"2024-10-04T06:16:34Z","timestamp":1728022594000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-024-02074-y"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,5,7]]},"references-count":93,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2024,10]]}},"alternative-id":["2074"],"URL":"https:\/\/doi.org\/10.1007\/s11263-024-02074-y","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,5,7]]},"assertion":[{"value":"22 August 2023","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"28 March 2024","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"7 May 2024","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"All authors certify that they have no affiliations with or involvement in any organization or entity with any financial interest or non-financial interest in the subject matter or materials discussed in this manuscript.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}}]}}