{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T01:27:33Z","timestamp":1760059653117,"version":"build-2065373602"},"reference-count":54,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2025,6,27]],"date-time":"2025-06-27T00:00:00Z","timestamp":1750982400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Natural Science Foundation of China","award":["62072362","12101479","2021JQ-660","2024JC-YBMS-531","2024ZC-KJXX-034","23CGZHCYH0008"],"award-info":[{"award-number":["62072362","12101479","2021JQ-660","2024JC-YBMS-531","2024ZC-KJXX-034","23CGZHCYH0008"]}]},{"name":"Natural Science Basis Research Plan in Shaanxi Province of China","award":["62072362","12101479","2021JQ-660","2024JC-YBMS-531","2024ZC-KJXX-034","23CGZHCYH0008"],"award-info":[{"award-number":["62072362","12101479","2021JQ-660","2024JC-YBMS-531","2024ZC-KJXX-034","23CGZHCYH0008"]}]},{"name":"Shaanxi Provincial Innovation Capacity Support Programme Project","award":["62072362","12101479","2021JQ-660","2024JC-YBMS-531","2024ZC-KJXX-034","23CGZHCYH0008"],"award-info":[{"award-number":["62072362","12101479","2021JQ-660","2024JC-YBMS-531","2024ZC-KJXX-034","23CGZHCYH0008"]}]},{"name":"Xi\u2019an Major Scientific and Technological Achievements Transformation Industrialization Project","award":["62072362","12101479","2021JQ-660","2024JC-YBMS-531","2024ZC-KJXX-034","23CGZHCYH0008"],"award-info":[{"award-number":["62072362","12101479","2021JQ-660","2024JC-YBMS-531","2024ZC-KJXX-034","23CGZHCYH0008"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Algorithms"],"abstract":"<jats:p>Recognizing hand actions and poses from first-person RGB videos is crucial for applications like human\u2013computer interaction. However, the recognition accuracy is often affected by factors such as occlusion and blurring. In this study, we propose a unified framework for action recognition and hand pose estimation in first-person RGB videos. The framework consists of two main modules: the Hand Pose Estimation Module and the Action Recognition Module. In the Hand Pose Estimation Module, each video frame is fed into a multi-layer transformer encoder after passing through a feature extractor. The hand pose results and object categories for each frame are obtained through multi-layer perceptron prediction using a dual residual network structure. The above prediction results are concatenated with the feature information corresponding to each frame for subsequent action recognition tasks. In the Action Recognition Module, the feature vectors from each frame are aggregated by a multi-layer transformer encoder to capture the temporal information of the hand between video frames and obtain the motion trajectory. The final output is the category of hand movements in consecutive video frames. We conducted experiments on two publicly available datasets, FPHA and H2O, and the results show that our method achieves significant improvements on both datasets, with action recognition accuracies of 94.82% and 87.92%, respectively.<\/jats:p>","DOI":"10.3390\/a18070393","type":"journal-article","created":{"date-parts":[[2025,6,30]],"date-time":"2025-06-30T12:10:31Z","timestamp":1751285431000},"page":"393","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["A Unified Framework for Recognizing Dynamic Hand Actions and Estimating Hand Pose from First-Person RGB Videos"],"prefix":"10.3390","volume":"18","author":[{"given":"Jiayi","family":"Yang","sequence":"first","affiliation":[{"name":"School of Computer Science, Xi\u2019an Polytechnic University, Xi\u2019an 710600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jiao","family":"Liang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Xi\u2019an Polytechnic University, Xi\u2019an 710600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-4596-9207","authenticated-orcid":false,"given":"Huimin","family":"Pan","sequence":"additional","affiliation":[{"name":"School of Computer Science, Xi\u2019an Polytechnic University, Xi\u2019an 710600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuting","family":"Cai","sequence":"additional","affiliation":[{"name":"School of Computer Science, Xi\u2019an Polytechnic University, Xi\u2019an 710600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7119-8144","authenticated-orcid":false,"given":"Quanli","family":"Gao","sequence":"additional","affiliation":[{"name":"School of Computer Science, Xi\u2019an Polytechnic University, Xi\u2019an 710600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xihan","family":"Wang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Xi\u2019an Polytechnic University, Xi\u2019an 710600, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,6,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Garcia-Hernando, G., Yuan, S., Baek, S., and Kim, T.-K. (2018, January 18\u201322). First-person hand action benchmark with rgb-d videos and 3d hand pose annotations. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00050"},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Kwon, T., Tekin, B., St\u00fchmer, J., Bogo, F., and Pollefeys, M. (2021, January 11\u201317). H2o: Two hands manipulating objects for first person interaction recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.00998"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Tekin, B., Bogo, F., and Pollefeys, M. (2019, January 15\u201320). H+ o: Unified egocentric recognition of 3d hand-object poses and interactions. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00464"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Wen, Y., Pan, H., Yang, L., Pan, J., Komura, T., and Wang, W. (2023, January 18\u201322). Hierarchical temporal transformer for 3D hand pose estimation and action recognition from egocentric RGB videos. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, Canada.","DOI":"10.1109\/CVPR52729.2023.02035"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Yang, S., Liu, J., Lu, S., Er, M.H., and Kot, A.C. (2020, January 23\u201328). Collaborative learning of gesture recognition and 3D hand pose estimation with multi-order feature analysis. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Part III 16, 2020.","DOI":"10.1007\/978-3-030-58580-8_45"},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Fan, Z., Liu, J., and Wang, Y. (2020, January 23\u201328). Adaptive computationally efficient network for monocular 3d hand pose estimation. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Part IV 16, 2020.","DOI":"10.1007\/978-3-030-58548-8_8"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Iqbal, U., Molchanov, P., Gall, T.B.J., and Kautz, J. (2018, January 18\u201322). Hand pose estimation via latent 2.5 d heatmap regression. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01252-6_8"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Kim, D.U., Kim, K.I., and Baek, S. (2021, January 11\u201317). End-to-end detection and pose estimation of two interacting hands. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01100"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Lin, K., Wang, L., and Liu, Z. (2021, January 11\u201317). Mesh graphormer. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, BC, Canada.","DOI":"10.1109\/ICCV48922.2021.01270"},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Meng, H., Jin, S., Liu, W., Qian, C., Lin, M., Ouyang, W., and Luo, P. (2022, January 23\u201327). 3d interacting hand pose estimation by hand de-occlusion and removal. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-20068-7_22"},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Moon, G., Yu, S.-I., Wen, H., Shiratori, T., and Lee, K.M. (2020, January 23\u201328). Interhand2. 6m: A dataset and baseline for 3d interacting hand pose estimation from a single rgb image. Proceedings of the Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK. Part XX 16. 2020.","DOI":"10.1007\/978-3-030-58565-5_33"},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Mueller, F., Bernard, F., Sotnychenko, O., Mehta, D., Sridhar, S., Casas, D., and Theobalt, C. (2018, January 18\u201322). Ganerated hands for real-time 3d hand tracking from monocular rgb. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00013"},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Pan, H., Cai, Y., Yang, J., Niu, S., Gao, Q., and Wang, X. (2024). HandFI: Multilevel Interacting Hand Reconstruction Based on Multilevel Feature Fusion in RGB Images. Sensors, 25.","DOI":"10.3390\/s25010088"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Spurr, A., Iqbal, U., Molchanov, P., Hilliges, O., and Kautz, J. (2020, January 23\u201328). Weakly supervised 3d hand pose estimation via biomechanical constraints. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58520-4_13"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zimmermann, C., and Brox, T. (2017, January 22\u201329). Learning to estimate 3d hand pose from single rgb images. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.525"},{"key":"ref_16","unstructured":"Cai, Y., Ge, L., Liu, J., Cai, J., Cham, T.-J., Yuan, J., and Thalmann, N.M. (November, January 27). Exploiting spatial-temporal relationships for 3d pose estimation via graph convolutional networks. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Chen, L., Lin, S.-Y., Xie, Y., Lin, Y.-Y., and Xie, X. (2021, January 5\u20139). Temporal-aware self-supervised learning for 3d hand pose and mesh estimation in videos. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Virtual.","DOI":"10.1109\/WACV48630.2021.00109"},{"key":"ref_18","first-page":"1","article-title":"Rgb2hands: Real-time tracking of 3d hand interactions from monocular rgb video","volume":"39","author":"Wang","year":"2020","journal-title":"ACM Trans. Graph. (ToG)"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Cosma, A., and Radoi, E. (2023). GaitFormer: Learning Gait Representations with Noisy Multi-Task Learning. arXiv.","DOI":"10.3390\/s22186803"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Hu, L., Gao, L., Liu, Z., and Feng, W. (2023, January 18\u201322). Continuous sign language recognition with correlation network. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00249"},{"key":"ref_21","unstructured":"Kim, J.-H., Kim, N., and Won, C.S. (2023). Multi Modal Facial Expression Recognition with Transformer-Based Fusion Networks and Dynamic Sampling. arXiv."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"8590","DOI":"10.1109\/TIP.2020.3018222","article-title":"Revealing the invisible with model and data shrinking for composite-database micro-expression recognition","volume":"29","author":"Xia","year":"2020","journal-title":"IEEE Trans. Image Process."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Zhu, X., Huang, P.-Y., Liang, J., De Melo, C.M., and Hauptmann, A.G. (2023, January 18\u201322). Stmt: A spatial-temporal mesh transformer for mocap-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00153"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Carreira, J., and Zisserman, A. (2017, January 21\u201326). Quo vadis, action recognition? a new model and the kinetics dataset. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.502"},{"key":"ref_25","unstructured":"Feichtenhofer, C., Fan, H., Malik, J., and He, K. (November, January 27). Slowfast networks for video recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Republic of Korea."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Feichtenhofer, C., Pinz, A., and Zisserman, A. (2016, January 27\u201330). Convolutional two-stream network fusion for video action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.213"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Liu, J., Shahroudy, A., Xu, D., and Wang, G. (2016, January 11\u201314). Spatio-temporal lstm with trust gates for 3d human action recognition. Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46487-9_50"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2019, January 15\u201320). Two-stream adaptive graph convolutional networks for skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01230"},{"key":"ref_29","first-page":"568","article-title":"Two-stream convolutional networks for action recognition in videos","volume":"27","author":"Simonyan","year":"2014","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Yan, S., Xiong, Y., and Lin, D. (2018, January 2\u20137). Spatial temporal graph convolutional networks for skeleton-based action recognition. Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA.","DOI":"10.1609\/aaai.v32i1.12328"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Tran, D., Bourdev, L., Fergus, R., Torresani, L., and Paluri, M. (2015, January 13\u201316). Learning spatiotemporal features with 3d convolutional networks. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.510"},{"key":"ref_32","first-page":"6000","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., and Zagoruyko, S. (2020, January 23\u201328). End-to-end object detection with transformers. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_13"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Lin, K., Wang, L., and Liu, Z. (2021, January 20\u201325). End-to-end human pose and mesh reconstruction with transformers. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA.","DOI":"10.1109\/CVPR46437.2021.00199"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Hampali, S., Sarkar, S.D., Rad, M., and Lepetit, V. (2022, January 18\u201324). Keypoint transformer: Solving joint identification in challenging hands and object interactions for accurate 3d pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01081"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Beyer, L., Izmailov, P., Kolesnikov, A., Caron, M., Kornblith, S., Zhai, X., Minderer, M., Tschannen, M., Alabdulmohsin, I., and Pavetic, F. (2023, January 18\u201322). Flexivit: One model for all patch sizes. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.01393"},{"key":"ref_37","first-page":"3200","article-title":"Human action recognition from various data modalities: A review","volume":"45","author":"Sun","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_38","unstructured":"Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2020). An image is worth 16x16 words: Transformers for image recognition at scale. arXiv."},{"key":"ref_39","unstructured":"Lee, J., and Toutanova, K. (2018). Pre-training of deep bidirectional transformers for language understanding. arXiv."},{"key":"ref_40","unstructured":"Kingma, D.P., and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Hasson, Y., Tekin, B., Bogo, F., Laptev, I., Pollefeys, M., and Schmid, C. (2020, January 13\u201319). Leveraging photometric consistency over time for sparsely supervised hand-object reconstruction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00065"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Aboukhadra, A.T., Malik, J., Elhayek, A., Robertini, N., and Stricker, D. (2023, January 3\u20137). Thor-net: End-to-end graformer-based realistic two hands and object reconstruction with self-supervision. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV56688.2023.00106"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Hu, J.-F., Zheng, W.-S., Lai, J., and Zhang, J. (2015, January 8\u201310). Jointly learning heterogeneous features for RGB-D activity recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299172"},{"key":"ref_44","doi-asserted-by":"crossref","first-page":"111343","DOI":"10.1016\/j.patcog.2025.111343","article-title":"HAN: An efficient hierarchical self-attention network for skeleton-based gesture recognition","volume":"162","author":"Liu","year":"2025","journal-title":"Pattern Recognit."},{"key":"ref_45","doi-asserted-by":"crossref","first-page":"2179","DOI":"10.1109\/TCDS.2023.3242988","article-title":"An efficient graph convolution network for skeleton-based dynamic hand gesture recognition","volume":"15","author":"Peng","year":"2023","journal-title":"IEEE Trans. Cogn. Dev. Syst."},{"key":"ref_46","doi-asserted-by":"crossref","first-page":"120735","DOI":"10.1016\/j.eswa.2023.120735","article-title":"SBI-DHGR: Skeleton-based intelligent dynamic hand gestures recognition","volume":"232","author":"Narayan","year":"2023","journal-title":"Expert Syst. Appl."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Prasse, K., Jung, S., Zhou, Y., and Keuper, M. (2023, January 19\u201322). Local spherical harmonics improve skeleton-based hand action recognition. Proceedings of the DAGM German Conference on Pattern Recognition, Heidelberg, Germany.","DOI":"10.1007\/978-3-031-54605-1_5"},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Chen, Y., Zhang, Z., Yuan, C., Li, B., Deng, Y., and Hu, W. (2021, January 10\u201317). Channel-wise topology refinement graph convolution for skeleton-based action recognition. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Virtual.","DOI":"10.1109\/ICCV48922.2021.01311"},{"key":"ref_49","doi-asserted-by":"crossref","first-page":"84","DOI":"10.1007\/s00138-022-01328-4","article-title":"Graph convolutional networks and LSTM for first-person multimodal hand action recognition","volume":"33","author":"Li","year":"2022","journal-title":"Mach. Vis. Appl."},{"key":"ref_50","doi-asserted-by":"crossref","unstructured":"Mucha, W., and Kampel, M. (2024, January 27\u201331). In my perspective, in my hands: Accurate egocentric 2d hand pose and action recognition. Proceedings of the 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG), Istanbul, Turkey.","DOI":"10.1109\/FG59268.2024.10582035"},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"246","DOI":"10.1109\/TCDS.2020.3048883","article-title":"Trear: Transformer-based rgb-d egocentric action recognition","volume":"14","author":"Li","year":"2021","journal-title":"IEEE Trans. Cogn. Dev. Syst."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Wang, X., Girshick, R., Gupta, A., and He, K. (2018, January 18\u201322). Non-local neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00813"},{"key":"ref_53","doi-asserted-by":"crossref","unstructured":"Duan, H., Zhao, Y., Chen, K., Lin, D., and Dai, B. (2022, January 18\u201324). Revisiting skeleton-based action recognition. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00298"},{"key":"ref_54","doi-asserted-by":"crossref","first-page":"2208","DOI":"10.1109\/TNNLS.2020.3044176","article-title":"SymNet: A simple symmetric positive definite manifold deep learning method for image set classification","volume":"33","author":"Wang","year":"2021","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."}],"container-title":["Algorithms"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/7\/393\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:00:12Z","timestamp":1760032812000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-4893\/18\/7\/393"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,27]]},"references-count":54,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2025,7]]}},"alternative-id":["a18070393"],"URL":"https:\/\/doi.org\/10.3390\/a18070393","relation":{},"ISSN":["1999-4893"],"issn-type":[{"type":"electronic","value":"1999-4893"}],"subject":[],"published":{"date-parts":[[2025,6,27]]}}}