{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,3]],"date-time":"2026-07-03T16:58:19Z","timestamp":1783097899487,"version":"3.54.6"},"reference-count":69,"publisher":"Springer Science and Business Media LLC","issue":"11","license":[{"start":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T00:00:00Z","timestamp":1688947200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"},{"start":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T00:00:00Z","timestamp":1688947200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.springernature.com\/gp\/researchers\/text-and-data-mining"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Comput Vis"],"published-print":{"date-parts":[[2023,11]]},"DOI":"10.1007\/s11263-023-01842-6","type":"journal-article","created":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T18:02:42Z","timestamp":1689012162000},"page":"2994-3018","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":37,"title":["HiEve: A Large-Scale Benchmark for Human-Centric Video Analysis in Complex Events"],"prefix":"10.1007","volume":"131","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-8307-7107","authenticated-orcid":false,"given":"Weiyao","family":"Lin","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huabin","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shizhan","family":"Liu","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yuxi","family":"Li","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Hongkai","family":"Xiong","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guojun","family":"Qi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nicu","family":"Sebe","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,7,10]]},"reference":[{"key":"1842_CR1","doi-asserted-by":"crossref","unstructured":"Andriluka, M., Iqbal, U., Insafutdinov, E., Pishchulin, L., Milan, A., Gall, J., & Schiele, B. (2018). Posetrack: A benchmark for human pose estimation and tracking. In CVPR (pp. 5167\u20135176).","DOI":"10.1109\/CVPR.2018.00542"},{"key":"1842_CR2","doi-asserted-by":"crossref","unstructured":"Andriluka, M., Pishchulin, L., Gehler, P., & Schiele, B. (2014). 2D human pose estimation: New benchmark and state of the art analysis. In CVPR.","DOI":"10.1109\/CVPR.2014.471"},{"key":"1842_CR3","unstructured":"Bertasius, G., Wang, H., & Torresani, L. (2021). Is space-time attention all you need for video understanding?. In ICML (Vol.\u00a02, p.\u00a04)."},{"key":"1842_CR4","doi-asserted-by":"crossref","unstructured":"Bewley, A., Ge, Z., Ott, L., Ramos, F., & Upcroft, B. (2016). Simple online and realtime tracking. In 2016 IEEE international conference on image processing (ICIP) (pp. 3464\u20133468). IEEE.","DOI":"10.1109\/ICIP.2016.7533003"},{"key":"1842_CR5","doi-asserted-by":"crossref","unstructured":"Bochinski, E., Eiselein, V., & Sikora, T. (2017). High-speed tracking-by-detection without using image information. In 2017 14th IEEE international conference on advanced video and signal based surveillance (AVSS). IEEE.","DOI":"10.1109\/AVSS.2017.8078516"},{"key":"1842_CR6","doi-asserted-by":"crossref","unstructured":"Cai, Y., Wang, Z., Luo, Z., Yin, B., Du, A., Wang, H., Zhang, X., Zhou, X., Zhou, E., & Sun, J. (2020). Learning delicate local representations for multi-person pose estimation. In European conference on computer vision (pp. 455\u2013472). Springer.","DOI":"10.1007\/978-3-030-58580-8_27"},{"key":"1842_CR7","doi-asserted-by":"crossref","unstructured":"Carreira, J., & Zisserman, A. (2017). Quo Vadis, action recognition? A new model and the kinetics dataset. In CVPR (pp. 6299\u20136308).","DOI":"10.1109\/CVPR.2017.502"},{"key":"1842_CR8","doi-asserted-by":"crossref","unstructured":"Chen, L., Ai, H., Zhuang, Z., & Shang, C. (2018). Real-time multiple people tracking with deeply learned candidate selection and person re-identification. In ICME.","DOI":"10.1109\/ICME.2018.8486597"},{"key":"1842_CR9","doi-asserted-by":"crossref","unstructured":"Chen, Y., Zhao, P., Qi, M., Zhao, Y., Jia, W., & Wang, R. (2022). Audio matters in video super-resolution by implicit semantic guidance. IEEE Transactions on Multimedia.","DOI":"10.1109\/TMM.2022.3152941"},{"key":"1842_CR10","doi-asserted-by":"crossref","unstructured":"Cheng, B., Xiao, B., Wang, J., Shi, H., Huang, T.S., & Zhang, L. (2020). Higherhrnet: Scale-aware representation learning for bottom-up human pose estimation. In CVPR","DOI":"10.1109\/CVPR42600.2020.00543"},{"issue":"11","key":"1842_CR11","doi-asserted-by":"publisher","first-page":"4125","DOI":"10.1109\/TPAMI.2020.2991965","volume":"43","author":"D Damen","year":"2021","unstructured":"Damen, D., Doughty, H., Farinella, G. M., Fidler, S., Furnari, A., Kazakos, E., Moltisanti, D., Munro, J., Perrett, T., Price, W., & Wray, M. (2021). The epic-kitchens dataset: Collection, challenges and baselines. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 43(11), 4125\u20134141. https:\/\/doi.org\/10.1109\/TPAMI.2020.2991965","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)"},{"key":"1842_CR12","unstructured":"Dendorfer, P., Rezatofighi, H., Milan, A., Shi, J., Cremers, D., Reid, I., Roth, S., Schindler, K., Leal-Taix\u00e9, L. (2020). Mot20: A benchmark for multi object tracking in crowded scenes. arXiv:2003.09003."},{"issue":"7","key":"1842_CR13","doi-asserted-by":"publisher","first-page":"3010","DOI":"10.1109\/TIP.2016.2552404","volume":"25","author":"Y Du","year":"2016","unstructured":"Du, Y., Fu, Y., & Wang, L. (2016). Representation learning of temporal dynamics for skeleton-based action recognition. IEEE Transactions on Image Processing, 25(7), 3010\u20133022.","journal-title":"IEEE Transactions on Image Processing"},{"key":"1842_CR14","doi-asserted-by":"crossref","unstructured":"Eichner, M., &Ferrari, V. (2010). We are family: Joint pose estimation of multiple persons. In ECCV.","DOI":"10.1007\/978-3-642-15549-9_17"},{"key":"1842_CR15","doi-asserted-by":"crossref","unstructured":"Fang, H.S., Xie, S., Tai, Y.W., & Lu, C. (2017). Rmpe: Regional multi-person pose estimation. In IEEE international conference on computer vision.","DOI":"10.1109\/ICCV.2017.256"},{"key":"1842_CR16","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., & He, K. (2019). Slowfast networks for video recognition. In IEEE international conference on computer vision.","DOI":"10.1109\/ICCV.2019.00630"},{"key":"1842_CR17","doi-asserted-by":"crossref","unstructured":"Ferryman, J., & Shahrokni, A. (2009). Pets2009: Dataset and challenge. In 2009 Twelfth IEEE international workshop on performance evaluation of tracking and surveillance (pp. 1\u20136. IEEE).","DOI":"10.1109\/PETS-WINTER.2009.5399556"},{"key":"1842_CR18","doi-asserted-by":"crossref","unstructured":"Geiger, A., Lenz, P., & Urtasun, R. (2012). Are we ready for autonomous driving? The Kitti vision benchmark suite. In 2012 CVPR. IEEE","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"1842_CR19","doi-asserted-by":"crossref","unstructured":"Geng, Z., Sun, K., Xiao, B., Zhang, Z., & Wang, J. (2021). Bottom-up human pose estimation via disentangled keypoint regression. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 14676\u201314686).","DOI":"10.1109\/CVPR46437.2021.01444"},{"key":"1842_CR20","unstructured":"Girdhar, R., Carreira, J., Doersch, C., & Zisserman, A. (2018). A better baseline for AVA. arXiv:1807.10066."},{"key":"1842_CR21","doi-asserted-by":"crossref","unstructured":"Girdhar, R., Carreira, J., Doersch, C., & Zisserman, A. (2019). Video action transformer network. In CVPR (pp. 244\u2013253).","DOI":"10.1109\/CVPR.2019.00033"},{"key":"1842_CR22","doi-asserted-by":"crossref","unstructured":"Goyal, R., Ebrahimi\u00a0Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., & Mueller-Freitag, M., et\u00a0al. (2017). The\" something something\" video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision (pp. 5842\u20135850).","DOI":"10.1109\/ICCV.2017.622"},{"key":"1842_CR23","doi-asserted-by":"crossref","unstructured":"Gu, C., Sun, C., Ross, D.A., Vondrick, C., Pantofaru, C., Li, Y., Vijayanarasimhan, S., Toderici, G., Ricco, S., & Sukthankar, R., et\u00a0al. (2018). Ava: A video dataset of spatio-temporally localized atomic visual actions. In CVPR.","DOI":"10.1109\/CVPR.2018.00633"},{"key":"1842_CR24","doi-asserted-by":"crossref","unstructured":"Gu, C., Sun, C., Vijayanarasimhan, S., Pantofaru, C., Ross, D.A., Toderici, G., Li, Y., Ricco, S., & Sukthankar, R. (2018). Ava: A video dataset of spatio-temporally localized atomic visual actions. In CVPR.","DOI":"10.1109\/CVPR.2018.00633"},{"key":"1842_CR25","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., & Sun, J. (2016) Deep residual learning for image recognition. In CVPR (pp. 770\u2013778).","DOI":"10.1109\/CVPR.2016.90"},{"key":"1842_CR26","doi-asserted-by":"crossref","unstructured":"Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In CVPR.","DOI":"10.1109\/CVPR.2018.00745"},{"key":"1842_CR27","doi-asserted-by":"crossref","unstructured":"Iqbal, U., Garbade, M., & Gall, J. (2017). Pose for action-action for pose. In 2017 12th IEEE international conference on automatic face and gesture recognition (FG 2017) (pp. 438\u2013445). IEEE.","DOI":"10.1109\/FG.2017.61"},{"key":"1842_CR28","doi-asserted-by":"crossref","unstructured":"Johnson, S., & Everingham, M. (2010). Clustered pose and nonlinear appearance models for human pose estimation. In: bmvc.","DOI":"10.5244\/C.24.12"},{"key":"1842_CR29","doi-asserted-by":"crossref","unstructured":"Kalogeiton, V., Weinzaepfel, P., Ferrari, V., & Schmid, C. (2017). Action tubelet detector for spatio-temporal action localization. In Proceedings of the IEEE international conference on computer vision (pp. 4405\u20134413).","DOI":"10.1109\/ICCV.2017.472"},{"key":"1842_CR30","doi-asserted-by":"crossref","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., & Serre, T. (2011). Hmdb: A large video database for human motion recognition. In 2011 international conference on computer vision. IEEE.","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"1842_CR31","doi-asserted-by":"crossref","unstructured":"Li, J., Wang, C., Zhu, H., Mao, Y., Fang, H.S., & Lu, C. (2019). Crowdpose: Efficient crowded scenes pose estimation and a new benchmark. In CVPR (pp. 10863\u201310872).","DOI":"10.1109\/CVPR.2019.01112"},{"key":"1842_CR32","doi-asserted-by":"crossref","unstructured":"Li, Y., Zhang, B., Li, J., Wang, Y., Lin, W., Wang, C., Li, J., & Huang, F. (2021). Lstc: Boosting atomic action detection with long-short-term context. In Proceedings of the 29th ACM international conference on multimedia (pp. 2158\u20132166).","DOI":"10.1145\/3474085.3475374"},{"key":"1842_CR33","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., & Zitnick, C.L. (2014). Microsoft coco: Common objects in context. In ECCV.","DOI":"10.1007\/978-3-319-10602-1_48"},{"issue":"4","key":"1842_CR34","doi-asserted-by":"publisher","first-page":"1586","DOI":"10.1109\/TIP.2017.2785279","volume":"27","author":"J Liu","year":"2017","unstructured":"Liu, J., Wang, G., Duan, L. Y., Abdiyeva, K., & Kot, A. C. (2017). Skeleton-based human action recognition with global context-aware attention LSTM networks. IEEE Transactions on Image Processing, 27(4), 1586\u20131599.","journal-title":"IEEE Transactions on Image Processing"},{"key":"1842_CR35","doi-asserted-by":"crossref","unstructured":"Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., & Hu, H. (2022). Video swin transformer. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 3202\u20133211).","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"1842_CR36","doi-asserted-by":"crossref","unstructured":"Lu, C., Shi, J., & Jia, J. (2013). Abnormal event detection at 150 fps in matlab. In IEEE international conference on computer vision (pp. 2720\u20132727)","DOI":"10.1109\/ICCV.2013.338"},{"key":"1842_CR37","doi-asserted-by":"publisher","first-page":"548","DOI":"10.1007\/s11263-020-01375-2","volume":"129","author":"J Luiten","year":"2021","unstructured":"Luiten, J., Osep, A., Dendorfer, P., Torr, P., Geiger, A., Leal-Taix\u00e9, L., & Leibe, B. (2021). Hota: A higher order metric for evaluating multi-object tracking. International Journal of Computer Vision, 129, 548\u2013578.","journal-title":"International Journal of Computer Vision"},{"key":"1842_CR38","doi-asserted-by":"crossref","unstructured":"Luvizon, D.C., Picard, D., &Tabia, H. (2018). 2d\/3d pose estimation and action recognition using multitask deep learning. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 5137\u20135146).","DOI":"10.1109\/CVPR.2018.00539"},{"issue":"3","key":"1842_CR39","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/2487268.2487269","volume":"9","author":"T Mei","year":"2013","unstructured":"Mei, T., Tang, L. X., Tang, J., & Hua, X. S. (2013). Near-lossless semantic video summarization and its applications to video analysis. ACM Transactions on Multimedia Computing, Communications, and Applications, 9(3), 1\u201323.","journal-title":"ACM Transactions on Multimedia Computing, Communications, and Applications"},{"key":"1842_CR40","unstructured":"Milan, A., Leal-Taix\u00e9, L., Reid, I., Roth, S., & Schindler, K. (2016). Mot16: A benchmark for multi-object tracking. arXiv:1603.00831."},{"key":"1842_CR41","doi-asserted-by":"crossref","unstructured":"Ning, G., Huang, & H. (2019). Lighttrack: A generic framework for online top-down human pose tracking. arXiv:1905.02822.","DOI":"10.1109\/CVPRW50498.2020.00525"},{"key":"1842_CR42","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107480","volume":"107","author":"J Peng","year":"2020","unstructured":"Peng, J., Wang, T., Lin, W., Wang, J., See, J., Wen, S., & Ding, E. (2020). Tpm: Multiple object tracking with tracklet-plane matching. Pattern Recognition, 107, 107480.","journal-title":"Pattern Recognition"},{"key":"1842_CR43","doi-asserted-by":"crossref","unstructured":"Pishchulin, L., Insafutdinov, E., Tang, S., Andres, B., Andriluka, M., Gehler, P.V., & Schiele, B. (2016). Deepcut: Joint subset partition and labeling for multi person pose estimation. In CVPR (pp. 4929\u20134937).","DOI":"10.1109\/CVPR.2016.533"},{"key":"1842_CR44","doi-asserted-by":"crossref","unstructured":"Ren, L., Lu, J., Wang, Z., Tian, Q., & Zhou, J. (2018). Collaborative deep reinforcement learning for multi-object tracking. In ECCV (pp. 586\u2013602).","DOI":"10.1007\/978-3-030-01219-9_36"},{"key":"1842_CR45","unstructured":"Ren, S., He, K., Girshick, R., & Sun, J. (2015). Faster r-cnn: Towards real-time object detection with region proposal networks. In Advances in Neural Information Processing Systems (pp. 91\u201399)."},{"key":"1842_CR46","doi-asserted-by":"crossref","unstructured":"Ristani, E., Solera, F., Zou, R., Cucchiara, R., & Tomasi, C. (2016). Performance measures and a data set for multi-target, multi-camera tracking. In ECCV (pp. 17\u201335).","DOI":"10.1007\/978-3-319-48881-3_2"},{"key":"1842_CR47","doi-asserted-by":"crossref","unstructured":"Sapp, B., & Taskar, B. (2013). Modec: Multimodal decomposable models for human pose estimation. In: CVPR (pp. 3674\u20133681).","DOI":"10.1109\/CVPR.2013.471"},{"key":"1842_CR48","doi-asserted-by":"crossref","unstructured":"Shahroudy, A., Liu, J., Ng, T.T., & Wang, G.: Ntu rgb+ d . (2016). A large scale dataset for 3d human activity analysis. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1010\u20131019).","DOI":"10.1109\/CVPR.2016.115"},{"key":"1842_CR49","unstructured":"Shu, X., Tang, J., Qi, G., Liu, W., & Yang, J. (2019). Hierarchical long short-term concurrent memory for human interaction recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence."},{"key":"1842_CR50","doi-asserted-by":"crossref","unstructured":"Singh, G., Saha, S., Sapienza, M., Torr, P.H., & Cuzzolin, F. (2017). Online real-time multiple spatiotemporal action localisation and prediction. In Proceedings of the IEEE international conference on computer vision (pp. 3637\u20133646).","DOI":"10.1109\/ICCV.2017.393"},{"key":"1842_CR51","unstructured":"Soomro, K., Zamir, A.R., & Shah, M. (2012). Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv:1212.0402."},{"key":"1842_CR52","doi-asserted-by":"crossref","unstructured":"Sultani, W., Chen, C., & Shah, M. (2018). Real-world anomaly detection in surveillance videos. In CVPR (pp. 6479\u20136488).","DOI":"10.1109\/CVPR.2018.00678"},{"key":"1842_CR53","doi-asserted-by":"crossref","unstructured":"Sun, K., Xiao, B., Liu, D., & Wang, J. (2019). Deep high-resolution representation learning for human pose estimation. In CVPR.","DOI":"10.1109\/CVPR.2019.00584"},{"key":"1842_CR54","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., & Polosukhin, I. (2017). Attention is all you need. In: Advances in neural information processing systems (pp. 5998\u20136008)."},{"key":"1842_CR55","doi-asserted-by":"crossref","unstructured":"Veeriah, V., Zhuang, N., & Qi, G.J. (2015). Differential recurrent neural networks for action recognition. In IEEE international conference on computer vision (pp. 4041\u20134049).","DOI":"10.1109\/ICCV.2015.460"},{"key":"1842_CR56","doi-asserted-by":"crossref","unstructured":"Wang, H., & Wang, L. (2018). Beyond joints: Learning representations from primitive geometries for skeleton-based action recognition and detection. IEEE Transactions on Image Processing (pp. 4382\u20134394).","DOI":"10.1109\/TIP.2018.2837386"},{"key":"1842_CR57","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., & He, K. (2018). Non-local neural networks. In CVPR.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"1842_CR58","doi-asserted-by":"crossref","unstructured":"Wang, Z., Zheng, L., Liu, Y., & Wang, S. (2019). Towards real-time multi-object tracking. arXiv:1909.12605.","DOI":"10.1007\/978-3-030-58621-8_7"},{"key":"1842_CR59","doi-asserted-by":"crossref","unstructured":"Wojke, N., Bewley, A., &Paulus, D. (2017). Simple online and realtime tracking with a deep association metric. In 2017 IEEE international conference on image processing (pp. 3645\u20133649). IEEE.","DOI":"10.1109\/ICIP.2017.8296962"},{"key":"1842_CR60","doi-asserted-by":"crossref","unstructured":"Wu, C.Y., Feichtenhofer, C., Fan, H., He, K., Krahenbuhl, P., & Girshick, R. (2019). Long-term feature banks for detailed video understanding. In Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (pp. 284\u2013293).","DOI":"10.1109\/CVPR.2019.00037"},{"key":"1842_CR61","doi-asserted-by":"crossref","unstructured":"Xiao, B., Wu, H., & Wei, Y. (2018). Simple baselines for human pose estimation and tracking. In ECCV.","DOI":"10.1007\/978-3-030-01231-1_29"},{"key":"1842_CR62","unstructured":"Xiaohan\u00a0Nie, B., Xiong, C., & Zhu, S.C. (2015). Joint action recognition and pose estimation from video. In CVPR (pp. 1293\u20131301)."},{"key":"1842_CR63","unstructured":"Xiu, Y., Li, J., Wang, H., Fang, Y., & Lu, C. (2018). Pose flow: Efficient online pose tracking. arXiv:1802.00977."},{"key":"1842_CR64","doi-asserted-by":"crossref","unstructured":"Xu, M., Liu, Y., Hu, R., & He, F. (2018). Find who to look at: Turning from action to saliency. IEEE Transactions on Image Processing.","DOI":"10.1109\/TIP.2018.2837106"},{"key":"1842_CR65","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., & Lin, D. (2018). Spatial temporal graph convolutional networks for skeleton-based action recognition. In Thirty-second AAAI conference on artificial intelligence.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"1842_CR66","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., & Lin, D. (2018). Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proceedings of the AAAI conference on artificial intelligence (vol.\u00a032).","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"1842_CR67","unstructured":"Yuan, Y., Fu, R., Huang, L., Lin, W., Zhang, C., Chen, X., & Wang, J. (2021). Hrformer: High-resolution vision transformer for dense predict. Advances in Neural Information Processing Systems."},{"key":"1842_CR68","unstructured":"Zhang, Y., Wang, C., Wang, X., Zeng, W., & Liu, W. (2020). Fairmot: On the fairness of detection and re-identification in multiple object tracking. arXiv:2004.01888."},{"key":"1842_CR69","doi-asserted-by":"crossref","unstructured":"Zhou, X., Koltun, V., & Kr\u00e4henb\u00fchl, P. (2020). Tracking objects as points. In European conference on computer vision (pp. 474\u2013490).","DOI":"10.1007\/978-3-030-58548-8_28"}],"container-title":["International Journal of Computer Vision"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-023-01842-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11263-023-01842-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11263-023-01842-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,25]],"date-time":"2023-09-25T04:10:29Z","timestamp":1695615029000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11263-023-01842-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,10]]},"references-count":69,"journal-issue":{"issue":"11","published-print":{"date-parts":[[2023,11]]}},"alternative-id":["1842"],"URL":"https:\/\/doi.org\/10.1007\/s11263-023-01842-6","relation":{},"ISSN":["0920-5691","1573-1405"],"issn-type":[{"value":"0920-5691","type":"print"},{"value":"1573-1405","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,10]]},"assertion":[{"value":"19 August 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 June 2023","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 July 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflicts of interest. All videos in this paper are either collected where the human participants were informed in advance and their consents for data publication were obtained, or obtained from online repositories where the publishing approvals from the video authors were obtained and the human identity information was guaranteed to be properly hidden or blurred.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflicts of interest"}},{"value":"The authors do not condone with AI systems developed for malicious\/unethical surveillance and tracking systems. Any use of the proposed video dataset must adhere to all relevant laws and regulations, including those related to data protection, privacy, and ethical considerations. The proposed video dataset is not to be used for any purpose that violates individual privacy or other legal or ethical standards. The authors are committed to ensuring that the proposed video dataset is used in ways that benefit society and do not cause harm.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Intended use of HiEve"}}]}}