{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,20]],"date-time":"2026-05-20T02:14:14Z","timestamp":1779243254586,"version":"3.51.4"},"reference-count":55,"publisher":"Springer Science and Business Media LLC","issue":"10","license":[{"start":{"date-parts":[[2025,7,22]],"date-time":"2025-07-22T00:00:00Z","timestamp":1753142400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,7,22]],"date-time":"2025-07-22T00:00:00Z","timestamp":1753142400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100020952","name":"AIP Network Laboratory","doi-asserted-by":"publisher","award":["JPMJCR20U1"],"award-info":[{"award-number":["JPMJCR20U1"]}],"id":[{"id":"10.13039\/501100020952","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100020952","name":"AIP Network Laboratory","doi-asserted-by":"publisher","award":["JPMJCR20U1"],"award-info":[{"award-number":["JPMJCR20U1"]}],"id":[{"id":"10.13039\/501100020952","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2025,10]]},"abstract":"<jats:title>Abstract<\/jats:title>\n          <jats:p>In the development of science, accurate and reproducible documentation of the experimental process is crucial. Automatic recognition of the actions in experiments from videos would help experimenters by complementing the recording of experiments. Towards this goal, we propose FineBio, a new fine-grained video dataset of people performing biological experiments. The dataset consists of multi-view videos of 32 participants performing mock biological experiments with a total duration of 14.5 hours. One experiment forms a hierarchical structure, where a protocol consists of several steps, each further decomposed into a set of atomic operations. The uniqueness of biological experiments is that while they require strict adherence to steps described in each protocol, there is freedom in the order of atomic operations. We provide hierarchical annotation on protocols, steps, atomic operations, object locations, and their manipulation states, providing new challenges for structured activity understanding and hand-object interaction recognition. To find out challenges on activity understanding in biological experiments, we introduce baseline models and results on four different tasks, including (i) step segmentation, (ii) atomic operation detection (iii) object detection, and (iv) manipulated\/affected object detection. Dataset and code are available from <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/github.com\/aistairc\/FineBio\" ext-link-type=\"uri\">https:\/\/github.com\/aistairc\/FineBio<\/jats:ext-link>.<\/jats:p>","DOI":"10.1007\/s11263-025-02523-2","type":"journal-article","created":{"date-parts":[[2025,7,22]],"date-time":"2025-07-22T11:20:36Z","timestamp":1753183236000},"page":"7352-7367","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation"],"prefix":"10.1007","volume":"133","author":[{"given":"Takuma","family":"Yagi","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Misaki","family":"Ohashi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yifei","family":"Huang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryosuke","family":"Furuta","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Shungo","family":"Adachi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Toutai","family":"Mitsuyama","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0097-4537","authenticated-orcid":false,"given":"Yoichi","family":"Sato","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,7,22]]},"reference":[{"issue":"4","key":"2523_CR1","doi-asserted-by":"publisher","first-page":"103","DOI":"10.1049\/htl.2018.5098","volume":"6","author":"MM Alam","year":"2019","unstructured":"Alam, M. M., & Islam, M. T. (2019). Machine learning approach of automatic identification and counting of blood cells. Healthcare technology letters, 6(4), 103\u2013108.","journal-title":"Healthcare technology letters"},{"key":"2523_CR2","doi-asserted-by":"crossref","unstructured":"Alayrac, J. -B., Bojanowski, P., Agrawal, N., Sivic, J., Laptev, I., & Lacoste-Julien, S. (2016). Unsupervised learning from narrated instruction videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, (pp. 4575\u20134583).","DOI":"10.1109\/CVPR.2016.495"},{"key":"2523_CR3","unstructured":"Ashutosh, K., Ramakrishnan, S. K., Afouras, T., & Grauman, K. (2023). Video-mined task graphs for keystep recognition in instructional videos. Proceedings of the Advances in neural information processing systems 36."},{"key":"2523_CR4","doi-asserted-by":"crossref","unstructured":"Bansal, S., Arora, C., & Jawahar, C. (2022). My view is the best view: Procedure learning from egocentric videos. In Proceedings of the European Conference on Computer Vision, (pp. 657\u2013675).","DOI":"10.1007\/978-3-031-19778-9_38"},{"key":"2523_CR5","doi-asserted-by":"crossref","unstructured":"Butler, D. J., Wulff, J., Stanley, G. B., & Black, M. J. (2012). A naturalistic open source movie for optical flow evaluation. In Proceedings of the European Conference on Computer Vision, (pp. 611\u2013625).","DOI":"10.1007\/978-3-642-33783-3_44"},{"key":"2523_CR6","doi-asserted-by":"crossref","unstructured":"Carreira, J., & Zisserman, A. (2017). Quo vadis, action recognition? a new model and the kinetics dataset. In Proceedings of the IEEE Computer Vision and Pattern Recognition, (pp. 4724\u20134733).","DOI":"10.1109\/CVPR.2017.502"},{"key":"2523_CR7","unstructured":"Chen, K., Wang, J., Pang, J., Cao, Y., Xiong, Y., Li, X., Sun, S., Feng, W., Liu, Z., Xu, J., Zhang, Z., Cheng, D., Zhu, C., Cheng, T., Zhao, Q., Li, B., Lu, X., Zhu, R., Wu, Y., Dai, J., Wang, J., Shi, J., Ouyang, W., Loy, C. C., & Lin, D. (2019). MMDetection: Open mmlab detection toolbox and benchmark. arXiv preprint arXiv:1906.07155"},{"key":"2523_CR8","unstructured":"Cui, J., Gong, Z., Jia, B., Huang, S., Zheng, Z., Ma, J., & Zhu, Y. (2023). Probio: A protocol-guided multimodal dataset for molecular biology lab. Advances in Neural Information Processing Systems DataBase and Benchmarks Track 36."},{"key":"2523_CR9","doi-asserted-by":"publisher","first-page":"33","DOI":"10.1007\/s11263-021-01531-2","volume":"130","author":"D Damen","year":"2022","unstructured":"Damen, D., Doughty, H., Farinella, G. M., Furnari, A., Ma, J., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., & Wray, M. (2022). Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal of Computer Vision, 130, 33\u201355.","journal-title":"International Journal of Computer Vision"},{"key":"2523_CR10","first-page":"13745","volume":"35","author":"A Darkhalil","year":"2022","unstructured":"Darkhalil, A., Shan, D., Zhu, B., Ma, J., Kar, A., Higgins, R., Fidler, S., Fouhey, D., & Damen, D. (2022). Epic-kitchens visor benchmark: Video segmentations and object relations. Advances in Neural Information Processing Systems, 35, 13745\u201313758.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2523_CR11","unstructured":"DINO. https:\/\/github.com\/IDEA-Research\/DINO"},{"issue":"9","key":"2523_CR12","doi-asserted-by":"publisher","first-page":"1038","DOI":"10.1038\/s41592-021-01249-6","volume":"18","author":"C Edlund","year":"2021","unstructured":"Edlund, C., Jackson, T. R., Khalid, N., Bevan, N., Dale, T., Dengel, A., Ahmed, S., Trygg, J., & Sj\u00f6gren, R. (2021). Livecell\u2014a large-scale dataset for label-free live cell segmentation. Nature methods, 18(9), 1038\u20131045.","journal-title":"Nature methods"},{"key":"2523_CR13","doi-asserted-by":"crossref","unstructured":"Fu, Q., Liu, X., & Kitani, K. (2022). Sequential voting with relational box fields for active object detection. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, (pp. 2374\u20132383).","DOI":"10.1109\/CVPR52688.2022.00241"},{"key":"2523_CR14","doi-asserted-by":"crossref","unstructured":"Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., Liu, X., et al. (2022). Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition, (pp. 18995\u201319012).","DOI":"10.1109\/CVPR52688.2022.01842"},{"key":"2523_CR15","doi-asserted-by":"crossref","unstructured":"Grauman, K., Westbury, A., Torresani, L., Kitani, K., Malik, J., Afouras, T., Ashutosh, K., Baiyya, V., Bansal, S., Boote, B., et al. (2024). Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition, (pp. 19383\u201319400).","DOI":"10.1109\/CVPR52733.2024.01834"},{"key":"2523_CR16","doi-asserted-by":"crossref","unstructured":"Heilbron, F. C., Escorcia, V., Ghanem, B., & Niebles, J. C. (2015). Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE Computer Vision and Pattern Recognition, (pp. 961\u2013970).","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"2523_CR17","doi-asserted-by":"publisher","DOI":"10.3389\/fbioe.2020.571777","volume":"8","author":"I Holland","year":"2020","unstructured":"Holland, I., & Davies, J. A. (2020). Automation in the life science research laboratory. Frontiers in Bioengineering and Biotechnology, 8, Article 571777.","journal-title":"Frontiers in Bioengineering and Biotechnology"},{"issue":"7209","key":"2523_CR18","doi-asserted-by":"publisher","first-page":"47","DOI":"10.1038\/455047a","volume":"455","author":"D Howe","year":"2008","unstructured":"Howe, D., Costanzo, M., Fey, P., Gojobori, T., Hannick, L., Hide, W., Hill, D. P., Kania, R., Schaeffer, M., St Pierre, S., et al. (2008). The future of biocuration. Nature, 455(7209), 47\u201350.","journal-title":"Nature"},{"key":"2523_CR19","first-page":"1671","volume":"35","author":"A Kaku","year":"2022","unstructured":"Kaku, A., Liu, K., Parnandi, A., Rajamohan, H. R., Venkataramanan, K., Venkatesan, A., Wirtanen, A., Pandit, N., Schambra, H., & Fernandez-Granda, C. (2022). Strokerehab: A benchmark dataset for sub-second action identification. Advances in Neural Information Processing Systems, 35, 1671\u20131684.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2523_CR20","unstructured":"Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980"},{"key":"2523_CR21","doi-asserted-by":"crossref","unstructured":"Kuehne, H., Gall, J., & Serre, T. (2016). An end-to-end generative framework for video segmentation and recognition. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision.","DOI":"10.1109\/WACV.2016.7477701"},{"key":"2523_CR22","doi-asserted-by":"crossref","unstructured":"Lee, S. -P., Lu, Z., Zhang, Z., Hoai, M., & Elhamifar, E. (2024). Error detection in egocentric procedural task videos. In: Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition, pp. 18655\u201318666.","DOI":"10.1109\/CVPR52733.2024.01765"},{"issue":"6","key":"2523_CR23","doi-asserted-by":"publisher","first-page":"6647","DOI":"10.1109\/TPAMI.2020.3021756","volume":"45","author":"S-J Li","year":"2023","unstructured":"Li, S.-J., AbuFarha, Y., Liu, Y., Cheng, M.-M., & Gall, J. (2023). Ms-tcn++: Multi-stage temporal convolutional network for action segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(6), 6647\u20136658.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2523_CR24","doi-asserted-by":"crossref","unstructured":"Lin, T. -Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., & Zitnick, C. L. (2014). Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision, (pp. 740\u2013755).","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"2523_CR25","doi-asserted-by":"crossref","unstructured":"Lin, Z., Wei, D., Petkova, M. D., Wu, Y., Ahmed, Z., Zou, S., Wendt, N., Boulanger-Weill, J., Wang, X., Dhanyasi, N., et al. (2021). Nucmm dataset: 3d neuronal nuclei instance segmentation at sub-cubic millimeter scale. In Proceedings of the International Conference on Medical Image Computing and Computer Assisted Intervention, (pp. 164\u2013174).","DOI":"10.1007\/978-3-030-87193-2_16"},{"key":"2523_CR26","doi-asserted-by":"crossref","unstructured":"Liu, X., Iwase, S., & Kitani, K. M. (2021). Stereobj-1m: Large-scale stereo image dataset for 6d object pose estimation. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 10870\u201310879).","DOI":"10.1109\/ICCV48922.2021.01069"},{"key":"2523_CR27","doi-asserted-by":"crossref","unstructured":"Miech, A., Zhukov, D., Alayrac, J. -B., Tapaswi, M., Laptev, I., & Sivic, J. (2019). Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 2630\u20132640).","DOI":"10.1109\/ICCV.2019.00272"},{"key":"2523_CR28","doi-asserted-by":"crossref","unstructured":"Moltisanti, D., Wray, M., Mayol-Cuevas, W., & Damen, D. (2017). Trespassing the boundaries: Labeling temporal bounds for object interactions in egocentric video. In Proceedings of the IEEE International Conference on Computer Vision, (pp. 2886\u20132894).","DOI":"10.1109\/ICCV.2017.314"},{"key":"2523_CR29","first-page":"7841","volume":"33","author":"S Narasimhaswamy","year":"2020","unstructured":"Narasimhaswamy, S., Nguyen, T., & Nguyen, M. H. (2020). Detecting hands and recognizing physical contact in the wild. Advances in neural information processing systems, 33, 7841\u20137851.","journal-title":"Advances in neural information processing systems"},{"key":"2523_CR30","doi-asserted-by":"crossref","unstructured":"Nishimura, T., Sakoda, K., Hashimoto, A., Ushiku, Y., Tanaka, N., Ono, F., Kameko, H., & Mori, S. (2021). Egocentric biochemical video-and-language dataset. In IEEE International Conference on Computer Vision Workshops, (pp. 3129\u20133133).","DOI":"10.1109\/ICCVW54120.2021.00348"},{"key":"2523_CR31","unstructured":"Nishimura, T., Yamamoto, K., Haneji, Y., Kajimura, K., Nishiwaki, C., Daikoku, E., Okuda, N., Ono, F., Kameko, H., & Mori, S. (2024). Biovl-qr: Egocentric biochemical video-and-language dataset using micro qr codes. arXiv preprint arXiv:2404.03161"},{"key":"2523_CR32","doi-asserted-by":"crossref","unstructured":"Papadopoulos, D. P., Mora, E., Chepurko, N., Huang, K. W., Ofli, F., & Torralba, A. (2022). Learning program representations for food images and cooking recipes. In: Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition, pp. 16559\u201316569.","DOI":"10.1109\/CVPR52688.2022.01606"},{"key":"2523_CR33","doi-asserted-by":"crossref","unstructured":"Ragusa, F., Furnari, A., Livatino, S., & Farinella, G. M. (2021). The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, (pp. 1569\u20131578).","DOI":"10.1109\/WACV48630.2021.00161"},{"key":"2523_CR34","doi-asserted-by":"crossref","unstructured":"Rai, N., Chen, H., Ji, J., Desai, R., Kozuka, K., Ishizaka, S., Adeli, E., & Niebles, J. C. (2021). Home action genome: Cooperative compositional action understanding. In: CVPR, pp. 11184\u201311193.","DOI":"10.1109\/CVPR46437.2021.01103"},{"key":"2523_CR35","unstructured":"Ren, S., He, K., Girshick, R., & Sun, J. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. In: Proceedings of the Advances in neural information processing systems, vol. 28."},{"key":"2523_CR36","doi-asserted-by":"crossref","unstructured":"Rodin, I., Furnari, A., Min, K., Tripathi, S., & Farinella, G. M. (2024). Action scene graphs for long-form understanding of egocentric videos. In Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition, (pp. 18622\u201318632).","DOI":"10.1109\/CVPR52733.2024.01762"},{"key":"2523_CR37","doi-asserted-by":"crossref","unstructured":"Sener, F., Chatterjee, D., Shelepov, D., He, K., Singhania, D., Wang, R., & Yao, A. (2022). Assembly101: A large-scale multi-view video dataset for understanding procedural activities. In: Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition, pp. 21096\u201321106.","DOI":"10.1109\/CVPR52688.2022.02042"},{"key":"2523_CR38","doi-asserted-by":"crossref","unstructured":"Shan, D., Geng, J., Shu, M., & Fouhey, D. F. (2020). Understanding human hands in contact at internet scale. In: Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition, pp. 9866\u20139875.","DOI":"10.1109\/CVPR42600.2020.00989"},{"key":"2523_CR39","first-page":"5898","volume":"34","author":"D Shan","year":"2021","unstructured":"Shan, D., Higgins, R., & Fouhey, D. (2021). Cohesiv: Contrastive object and hand embedding segmentation in video. Advances in Neural Information Processing Systems, 34, 5898\u20135909.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2523_CR40","doi-asserted-by":"crossref","unstructured":"Shao, D., Zhao, Y., Dai, B., & Lin, D. (2020). Finegym: A hierarchical video dataset for fine-grained action understanding. In: Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition, pp. 2616\u20132625.","DOI":"10.1109\/CVPR42600.2020.00269"},{"key":"2523_CR41","unstructured":"Song, Y., Byrne, E., Nagarajan, T., Wang, H., Martin, M., & Torresani, L. (2023). Ego4d goal-step: Toward hierarchical understanding of procedural activities. Proceedings of the Advances in neural information processing systems 36."},{"key":"2523_CR42","doi-asserted-by":"crossref","unstructured":"Tang, Y., Ding, D., Rao, Y., Zheng, Y., Zhang, D., Zhao, L., Lu, J., & Zhou, J. (2019). Coin: A large-scale dataset for comprehensive instructional video analysis. In: Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition.","DOI":"10.1109\/CVPR.2019.00130"},{"key":"2523_CR43","doi-asserted-by":"crossref","unstructured":"Teed, Z., & Deng, J. (2020). Raft: Recurrent all-pairs field transforms for optical flow. In: Proceedings of the European Conference on Computer Vision, pp. 402\u2013419.","DOI":"10.1007\/978-3-030-58536-5_24"},{"key":"2523_CR44","doi-asserted-by":"crossref","unstructured":"Wei, J., Suriawinata, A., Ren, B., Liu, X., Lisovsky, M., Vaickus, L., Brown, C., Baker, M., Tomita, N., Torresani, L., et al. (2021). A petri dish for histopathology image analysis. In: Artificial Intelligence in Medicine, pp. 11\u201324.","DOI":"10.1007\/978-3-030-77211-6_2"},{"key":"2523_CR45","doi-asserted-by":"crossref","unstructured":"Xu, J., Rao, Y., Yu, X., Chen, G., Zhou, J., & Lu, J. (2022). Finediving: A fine-grained dataset for procedure-aware action quality assessment. In: Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition, pp. 2949\u20132958.","DOI":"10.1109\/CVPR52688.2022.00296"},{"key":"2523_CR46","doi-asserted-by":"crossref","unstructured":"Yagi, T., Hasan, M. T., & Sato, Y. (2021). Hand-object contact prediction via motion-based pseudo-labeling and guided progressive label correction. In: Proceedings of the British Machine Vision Conference.","DOI":"10.5244\/C.35.25"},{"key":"2523_CR47","doi-asserted-by":"crossref","unstructured":"Yi, F., Wen, H., & Jiang, T. (2021). Asformer: Transformer for action segmentation. In: Proceedings of the British Machine Vision Conference.","DOI":"10.5244\/C.35.49"},{"key":"2523_CR48","doi-asserted-by":"crossref","unstructured":"Yu, Z., Huang, Y., Furuta, R., Yagi, T., Goutsu, Y., & Sato, Y. (2023). Fine-grained affordance annotation for egocentric hand-object interaction videos. In: Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, pp. 2154\u20132162.","DOI":"10.1109\/WACV56688.2023.00219"},{"key":"2523_CR49","unstructured":"Zhang, H., Li, F., Liu, S., Zhang, L., Su, H., Zhu, J., Ni, L., & Shum, H. -Y. (2023). Dino: Detr with improved denoising anchor boxes for end-to-end object detection. In: Proceedings of the International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=3mRwyG5one"},{"key":"2523_CR50","doi-asserted-by":"crossref","unstructured":"Zhang, C. -L., Wu, J., & Li, Y. (2022). Actionformer: Localizing moments of actions with transformers. In Proceedings of the European Conference on Computer Vision, vol. 13664, (pp. 492\u2013510).","DOI":"10.1007\/978-3-031-19772-7_29"},{"key":"2523_CR51","doi-asserted-by":"crossref","unstructured":"Zhang, L., Zhou, S., Stent, S., & Shi, J. (2022). Fine-grained egocentric hand-object segmentation: Dataset, model, and applications. In Proceedings of the European Conference on Computer Vision, (pp. 127\u2013145).","DOI":"10.1007\/978-3-031-19818-2_8"},{"issue":"11","key":"2523_CR52","doi-asserted-by":"publisher","first-page":"1330","DOI":"10.1109\/34.888718","volume":"22","author":"Z Zhang","year":"2000","unstructured":"Zhang, Z. (2000). A flexible new technique for camera calibration. IEEE Transactions on pattern analysis and machine intelligence, 22(11), 1330\u20131334.","journal-title":"IEEE Transactions on pattern analysis and machine intelligence"},{"key":"2523_CR53","doi-asserted-by":"crossref","unstructured":"Zhao, H., Torralba, A., Torresani, L., & Yan, Z. (2019). Hacs: Human action clips and segments dataset for recognition and temporal localization. In Proceedings of the IEEE\/CVF International Conference on Computer Vision, (pp. 8668\u2013.8678)","DOI":"10.1109\/ICCV.2019.00876"},{"key":"2523_CR54","unstructured":"Zhu, X., Su, W., Lu, L., Li, B., Wang, X., & Dai, J. (2021). Deformable detr: Deformable transformers for end-to-end object detection. In: Proceedings of the International Conference on Learning Representations."},{"key":"2523_CR55","doi-asserted-by":"crossref","unstructured":"Zhukov, D., Alayrac, J. -B., Cinbis, R. G., Fouhey, D., Laptev, I., & Sivic, J. (2019). Cross-task weakly supervised learning from instructional videos. In Proceedings of the IEEE\/CVF Computer Vision and Pattern Recognition.","DOI":"10.1109\/CVPR.2019.00365"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02523-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-025-02523-2\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-025-02523-2.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T08:48:29Z","timestamp":1760086109000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-025-02523-2"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,22]]},"references-count":55,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2025,10]]}},"alternative-id":["2523"],"URL":"https:\/\/doi.org\/10.1007\/s11263-025-02523-2","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,22]]},"assertion":[{"value":"30 May 2024","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 July 2025","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"22 July 2025","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}