{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T01:25:06Z","timestamp":1760145906584,"version":"build-2065373602"},"reference-count":46,"publisher":"MDPI AG","issue":"9","license":[{"start":{"date-parts":[[2024,8,31]],"date-time":"2024-08-31T00:00:00Z","timestamp":1725062400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Natural Science Foundation of China","award":["62362003","62371343","62301087","20232BAB202017","YC2023-S864"],"award-info":[{"award-number":["62362003","62371343","62301087","20232BAB202017","YC2023-S864"]}]},{"DOI":"10.13039\/501100004479","name":"Natural Science Foundation of Jiangxi Province","doi-asserted-by":"publisher","award":["62362003","62371343","62301087","20232BAB202017","YC2023-S864"],"award-info":[{"award-number":["62362003","62371343","62301087","20232BAB202017","YC2023-S864"]}],"id":[{"id":"10.13039\/501100004479","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Graduate Innovation Funding Program of Jiangxi Province","award":["62362003","62371343","62301087","20232BAB202017","YC2023-S864"],"award-info":[{"award-number":["62362003","62371343","62301087","20232BAB202017","YC2023-S864"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Imaging"],"abstract":"<jats:p>Methods based on deep learning have achieved great success in the field of video action recognition. When these methods are applied to real-world scenarios that require fine-grained analysis of actions, such as being tested on a tea ceremony, limitations may arise. To promote the development of fine-grained action recognition, a fine-grained video action dataset is constructed by collecting videos of tea ceremony actions. This dataset includes 2745 video clips. By using a hierarchical fine-grained action classification approach, these clips are divided into 9 basic action classes and 31 fine-grained action subclasses. To better establish a fine-grained temporal model for tea ceremony actions, a method named TSM-ConvNeXt is proposed that integrates a TSM into the high-performance convolutional neural network ConvNeXt. Compared to a baseline method using ResNet50, the experimental performance of TSM-ConvNeXt is improved by 7.31%. Furthermore, compared with the state-of-the-art methods for action recognition on the FineTea and Diving48 datasets, the proposed approach achieves the best experimental results. The FineTea dataset is publicly available.<\/jats:p>","DOI":"10.3390\/jimaging10090216","type":"journal-article","created":{"date-parts":[[2024,9,2]],"date-time":"2024-09-02T06:51:17Z","timestamp":1725259877000},"page":"216","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["FineTea: A Novel Fine-Grained Action Recognition Video Dataset for Tea Ceremony Actions"],"prefix":"10.3390","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0009-0005-3534-0384","authenticated-orcid":false,"given":"Changwei","family":"Ouyang","sequence":"first","affiliation":[{"name":"School of Mathematics and Computer Science, Gannan Normal University, Ganzhou 341000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5644-8002","authenticated-orcid":false,"given":"Yun","family":"Yi","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Gannan Normal University, Ganzhou 341000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-9999-4871","authenticated-orcid":false,"given":"Hanli","family":"Wang","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Technology, Tongji University, Shanghai 201804, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jin","family":"Zhou","sequence":"additional","affiliation":[{"name":"School of Mathematics and Computer Science, Gannan Normal University, Ganzhou 341000, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tao","family":"Tian","sequence":"additional","affiliation":[{"name":"School of Computer Science and Artificial Intelligence, Chaohu University, Hefei 238024, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2024,8,31]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Al-Faris, M., Chiverton, J., Ndzi, D., and Ahmed, A.I. (2020). A review on computer vision-based methods for human action recognition. J. Imaging, 6.","DOI":"10.3390\/jimaging6060046"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Rohrbach, M., Amin, S., Andriluka, M., and Schiele, B. (2012, January 16\u201321). A database for fine grained activity detection of cooking activities. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Providence, RI, USA.","DOI":"10.1109\/CVPR.2012.6247801"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Piergiovanni, A., and Ryoo, M.S. (2018, January 18\u201322). Fine-grained activity recognition in baseball videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00226"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Li, Y., Li, Y., and Vasconcelos, N. (2018, January 8\u201314). Resound: Towards action recognition without representation bias. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01231-1_32"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"20429","DOI":"10.1007\/s11042-020-08917-3","article-title":"Fine grained sport action recognition with Twin spatio-temporal convolutional neural networks: Application to table tennis","volume":"79","author":"Martin","year":"2020","journal-title":"Multimed. Tools Appl."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_7","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017, January 4\u20139). Attention is all you need. Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_8","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning transferable visual models from natural language supervision. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., and Serre, T. (2011, January 6\u201313). HMDB: A large video database for human motion recognition. Proceedings of the IEEE International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"ref_10","unstructured":"Soomro, K., Zamir, A.R., and Shah, M. (2012). UCF101: A dataset of 101 human actions classes from videos in the wild. arXiv."},{"key":"ref_11","unstructured":"Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, S., Viola, F., Green, T., Back, T., and Natsev, P. (2017). The kinetics human action video dataset. arXiv."},{"key":"ref_12","unstructured":"Bertasius, G., Wang, H., and Torresani, L. (2021, January 18\u201324). Is space-time attention all you need for video understanding?. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Liu, Z., Ning, J., Cao, Y., Wei, Y., Zhang, Z., Lin, S., and Hu, H. (2022, January 18\u201324). Video swin transformer. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00320"},{"key":"ref_14","unstructured":"Tong, Z., Song, Y., Wang, J., and Wang, L. (December, January 28). Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Proceedings of the Advances in Neural Information Processing Systems, New Orleans, LA, USA."},{"key":"ref_15","unstructured":"Yang, T., Zhu, Y., Xie, Y., Zhang, A., Chen, C., and Li, M. (2023, January 1\u20135). AIM: Adapting image models for efficient video action recognition. Proceedings of the International Conference on Learning Representations, Kigali, Rwanda."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Liu, Z., Mao, H., Wu, C.Y., Feichtenhofer, C., Darrell, T., and Xie, S. (2022, January 18\u201324). A convnet for the 2020s. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"ref_17","unstructured":"Lin, J., Gan, C., and Han, S. (November, January 27). TSM: Temporal shift module for efficient video understanding. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Schuldt, C., Laptev, I., and Caputo, B. (2004, January 23\u201326). Recognizing human actions: A local SVM approach. Proceedings of the International Conference on Pattern Recognition, Cambridge, UK.","DOI":"10.1109\/ICPR.2004.1334462"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"2247","DOI":"10.1109\/TPAMI.2007.70711","article-title":"Actions as space-time shapes","volume":"29","author":"Gorelick","year":"2007","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Caba Heilbron, F., Escorcia, V., Ghanem, B., and Carlos Niebles, J. (2015, January 7\u201312). Activitynet: A large-scale video benchmark for human activity understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1016\/j.cviu.2016.10.018","article-title":"The thumos challenge on action recognition for videos \u201cin the wild\u201d","volume":"155","author":"Idrees","year":"2017","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., and Gupta, A. (2016, January 11\u201314). Hollywood in homes: Crowdsourcing data collection for activity understanding. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46448-0_31"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Wang, L., Xiong, Y., Wang, Z., Qiao, Y., Lin, D., Tang, X., and Van Gool, L. (2016, January 11\u201314). Temporal segment networks: Towards good practices for deep action recognition. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_2"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Goyal, R., Ebrahimi Kahou, S., Michalski, V., Materzynska, J., Westphal, S., Kim, H., Haenel, V., Fruend, I., Yianilos, P., and Mueller-Freitag, M. (2017, January 22\u201329). The \u201csomething something\u201d video database for learning and evaluating visual common sense. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.622"},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Shao, D., Zhao, Y., Dai, B., and Lin, D. (2020, January 13\u201319). Finegym: A hierarchical video dataset for fine-grained action understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00269"},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Ullah, H., and Munir, A. (2023). Human activity recognition using cascaded dual attention cnn and bi-directional gru framework. J. Imaging, 9.","DOI":"10.3390\/jimaging9070130"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Host, K., Pobar, M., and Ivasic-Kos, M. (2023). Analysis of movement and activities of handball players using deep neural networks. J. Imaging, 9.","DOI":"10.3390\/jimaging9040080"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Karpathy, A., Toderici, G., Shetty, S., Leung, T., Sukthankar, R., and Fei-Fei, L. (2014, January 23\u201328). Large-scale video classification with convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.223"},{"key":"ref_29","unstructured":"Simonyan, K., and Zisserman, A. (2014, January 8\u201313). Two-stream convolutional networks for action recognition in videos. Proceedings of the Advances in Neural Information Processing Systems, Montreal, QC, Canada."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., and Zisserman, A. (2016, January 27\u201330). Convolutional two-stream network fusion for video action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.213"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo vadis, action recognition? A new model and the kinetics dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201322). Non-local neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_33","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (November, January 27). Slowfast networks for video recognition. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_34","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16 \u00d7 16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., Lin, S., and Guo, B. (2021, January 10\u201317). Swin transformer: Hierarchical vision transformer using shifted windows. Proceedings of the IEEE International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.00986"},{"key":"ref_36","unstructured":"Ioffe, S., and Szegedy, C. (2015, January 6\u201311). Batch normalization: Accelerating deep network training by reducing internal covariate shift. Proceedings of the International Conference on Machine Learning, Lille, France."},{"key":"ref_37","unstructured":"Ba, J.L., Kiros, J.R., and Hinton, G.E. (2016). Layer normalization. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Fei-Fei, L. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_39","unstructured":"MMAction2 Contributors (2024, August 27). OpenMMLab\u2019s Next Generation Video Understanding Toolbox and Benchmark. Available online: https:\/\/github.com\/open-mmlab\/mmaction2."},{"key":"ref_40","doi-asserted-by":"crossref","first-page":"330","DOI":"10.1007\/s42452-023-05568-5","article-title":"Towards efficient video-based action recognition: Context-aware memory attention network","volume":"5","author":"Koh","year":"2023","journal-title":"SN Appl. Sci."},{"key":"ref_41","doi-asserted-by":"crossref","first-page":"7279","DOI":"10.1109\/TIP.2022.3221292","article-title":"Spatio-temporal collaborative module for efficient action recognition","volume":"31","author":"Hao","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"7120","DOI":"10.1109\/TCSVT.2022.3169842","article-title":"Attention in attention: Modeling context correlation for efficient video classification","volume":"32","author":"Hao","year":"2022","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Zhang, C., Gupta, A., and Zisserman, A. (2021, January 20\u201325). Temporal query networks for fine-grained video understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00446"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Xie, S., Sun, C., Huang, J., Tu, Z., and Murphy, K. (2018, January 8\u201314). Rethinking spatiotemporal feature learning: Speed-accuracy trade-offs in video classification. Proceedings of the European Conference on Computer Vision, Munich, Germany.","DOI":"10.1007\/978-3-030-01267-0_19"},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"109905","DOI":"10.1016\/j.patcog.2023.109905","article-title":"Relative-position embedding based spatially and temporally decoupled Transformer for action recognition","volume":"145","author":"Ma","year":"2024","journal-title":"Pattern Recognit."},{"key":"ref_46","unstructured":"Kim, M., Kwon, H., Wang, C., Kwak, S., and Cho, M. (2021, January 6\u201316). Relational self-attention: What\u2019s missing in attention for video understanding. Proceedings of the Advances in Neural Information Processing Systems, Virtual."}],"container-title":["Journal of Imaging"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2313-433X\/10\/9\/216\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T15:46:29Z","timestamp":1760111189000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2313-433X\/10\/9\/216"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,31]]},"references-count":46,"journal-issue":{"issue":"9","published-online":{"date-parts":[[2024,9]]}},"alternative-id":["jimaging10090216"],"URL":"https:\/\/doi.org\/10.3390\/jimaging10090216","relation":{},"ISSN":["2313-433X"],"issn-type":[{"type":"electronic","value":"2313-433X"}],"subject":[],"published":{"date-parts":[[2024,8,31]]}}}