{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T02:03:26Z","timestamp":1783130606322,"version":"3.54.6"},"reference-count":49,"publisher":"MDPI AG","issue":"12","license":[{"start":{"date-parts":[[2023,6,16]],"date-time":"2023-06-16T00:00:00Z","timestamp":1686873600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62233003"],"award-info":[{"award-number":["62233003"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62073072"],"award-info":[{"award-number":["62073072"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["BE2020006"],"award-info":[{"award-number":["BE2020006"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["BE2020006-1"],"award-info":[{"award-number":["BE2020006-1"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["JCYJ20210324132202005"],"award-info":[{"award-number":["JCYJ20210324132202005"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["JCYJ20220818101206014"],"award-info":[{"award-number":["JCYJ20220818101206014"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Key Projects of the Key R&amp;D Program of Jiangsu Province","award":["62233003"],"award-info":[{"award-number":["62233003"]}]},{"name":"Key Projects of the Key R&amp;D Program of Jiangsu Province","award":["62073072"],"award-info":[{"award-number":["62073072"]}]},{"name":"Key Projects of the Key R&amp;D Program of Jiangsu Province","award":["BE2020006"],"award-info":[{"award-number":["BE2020006"]}]},{"name":"Key Projects of the Key R&amp;D Program of Jiangsu Province","award":["BE2020006-1"],"award-info":[{"award-number":["BE2020006-1"]}]},{"name":"Key Projects of the Key R&amp;D Program of Jiangsu Province","award":["JCYJ20210324132202005"],"award-info":[{"award-number":["JCYJ20210324132202005"]}]},{"name":"Key Projects of the Key R&amp;D Program of Jiangsu Province","award":["JCYJ20220818101206014"],"award-info":[{"award-number":["JCYJ20220818101206014"]}]},{"name":"Shenzhen Natural Science Foundation","award":["62233003"],"award-info":[{"award-number":["62233003"]}]},{"name":"Shenzhen Natural Science Foundation","award":["62073072"],"award-info":[{"award-number":["62073072"]}]},{"name":"Shenzhen Natural Science Foundation","award":["BE2020006"],"award-info":[{"award-number":["BE2020006"]}]},{"name":"Shenzhen Natural Science Foundation","award":["BE2020006-1"],"award-info":[{"award-number":["BE2020006-1"]}]},{"name":"Shenzhen Natural Science Foundation","award":["JCYJ20210324132202005"],"award-info":[{"award-number":["JCYJ20210324132202005"]}]},{"name":"Shenzhen Natural Science Foundation","award":["JCYJ20220818101206014"],"award-info":[{"award-number":["JCYJ20220818101206014"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Intelligent devices, which significantly improve the quality of life and work efficiency, are now widely integrated into people\u2019s daily lives and work. A precise understanding and analysis of human motion is essential for achieving harmonious coexistence and efficient interaction between intelligent devices and humans. However, existing human motion prediction methods often fail to fully exploit the dynamic spatial correlations and temporal dependencies inherent in motion sequence data, which leads to unsatisfactory prediction results. To address this issue, we proposed a novel human motion prediction method that utilizes dual-attention and multi-granularity temporal convolutional networks (DA-MgTCNs). Firstly, we designed a unique dual-attention (DA) model that combines joint attention and channel attention to extract spatial features from both joint and 3D coordinate dimensions. Next, we designed a multi-granularity temporal convolutional networks (MgTCNs) model with varying receptive fields to flexibly capture complex temporal dependencies. Finally, the experimental results from two benchmark datasets, Human3.6M and CMU-Mocap, demonstrated that our proposed method significantly outperformed other methods in both short-term and long-term prediction, thereby verifying the effectiveness of our algorithm.<\/jats:p>","DOI":"10.3390\/s23125653","type":"journal-article","created":{"date-parts":[[2023,6,16]],"date-time":"2023-06-16T10:19:22Z","timestamp":1686910762000},"page":"5653","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Human Motion Prediction via Dual-Attention and Multi-Granularity Temporal Convolutional Networks"],"prefix":"10.3390","volume":"23","author":[{"given":"Biaozhang","family":"Huang","sequence":"first","affiliation":[{"name":"Key Laboratory Measurement and Control of CSE Ministry of Education, School of Automation, Southeast University, Nanjing 210002, China"},{"name":"Nanjing Center for Applied Mathematics, Nanjing 211135, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Xinde","family":"Li","sequence":"additional","affiliation":[{"name":"Key Laboratory Measurement and Control of CSE Ministry of Education, School of Automation, Southeast University, Nanjing 210002, China"},{"name":"Nanjing Center for Applied Mathematics, Nanjing 211135, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2023,6,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1109\/MSP.2020.2984780","article-title":"3d point cloud processing and learning for autonomous driving: Impacting map creation, localization, and perception","volume":"38","author":"Chen","year":"2020","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Gui, L.Y., Zhang, K., Wang, Y.X., Liang, X., Moura, J.M., and Veloso, M. (2018, January 1\u20135). Teaching robots to predict human motion. Proceedings of the 2018 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain.","DOI":"10.1109\/IROS.2018.8594452"},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1109\/TPAMI.2015.2430335","article-title":"Anticipating human activities using object affordances for reactive robotic response","volume":"38","author":"Koppula","year":"2015","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"107868","DOI":"10.1016\/j.patcog.2021.107868","article-title":"Multi-task learning for gait-based identity recognition and emotion recognition using attention enhanced temporal graph convolutional network","volume":"114","author":"Sheng","year":"2021","journal-title":"Pattern Recognit."},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"13643","DOI":"10.1007\/s11042-017-4979-0","article-title":"Automatic analysis of complex athlete techniques in broadcast taekwondo video","volume":"77","author":"Kong","year":"2018","journal-title":"Multimed. Tools Appl."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"3607","DOI":"10.1109\/TFUZZ.2021.3079495","article-title":"Evidential reasoning with hesitant fuzzy belief structures for human activity recognition","volume":"29","author":"Dong","year":"2021","journal-title":"IEEE Trans. Fuzzy Syst."},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"5530","DOI":"10.1109\/TII.2022.3182780","article-title":"Multi-Source Weighted Domain Adaptation With Evidential Reasoning for Activity Recognition","volume":"19","author":"Dong","year":"2022","journal-title":"IEEE Trans. Ind. Inform."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Lehrmann, A.M., Gehler, P.V., and Nowozin, S. (2014, January 23\u201328). Efficient nonlinear Markov models for human motion. Proceedings of the Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.171"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"283","DOI":"10.1109\/TPAMI.2007.1167","article-title":"Gaussian Process Dynamical Models for Human Motion","volume":"30","author":"Wang","year":"2007","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_10","first-page":"1345","article-title":"Modeling human motion using binary latent variables","volume":"19","author":"Taylor","year":"2006","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Li, C., Zhang, Z., Lee, W.S., and Lee, G.H. (2018, January 18\u201322). Convolutional sequence to sequence model for human dynamics. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA.","DOI":"10.1109\/CVPR.2018.00548"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"3316","DOI":"10.1109\/TPAMI.2021.3053765","article-title":"Symbiotic Graph Neural Networks for 3D Skeleton-Based Human Action Recognition and Motion Prediction","volume":"44","author":"Li","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Li, M., Chen, S., Zhao, Y., Zhang, Y., Wang, Y., and Tian, Q. (2020, January 13\u201319). Dynamic multiscale graph neural networks for 3d skeleton based human motion prediction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00029"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Zhong, C., Hu, L., Zhang, Z., Ye, Y., and Xia, S. (2022, January 18\u201324). Spatio-Temporal Gating-Adjacency GCN For Human Motion Prediction. Proceedings of the Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00634"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Fragkiadaki, K., Levine, S., Felsen, P., and Malik, J. (2015, January 7\u201313). Recurrent network models for human dynamics. Proceedings of the 2015 IEEE International Conference on Computer Vision (ICCV), IEEE Computer Society, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.494"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Jain, A., Zamir, A.R., Savarese, S., and Saxena, A. (2016, January 27\u201330). Structural-rnn: Deep learning on spatio-temporal graphs. Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.573"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Martinez, J., Black, M.J., and Romero, J. (2017, January 21\u201326). On human motion prediction using recurrent neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.497"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Liu, Z., Wu, S., Jin, S., Liu, Q., Lu, S., Zimmermann, R., and Cheng, L. (2019, January 15\u201320). Towards natural and accurate future motion prediction of humans and animals. Proceedings of the Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.01024"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"3300","DOI":"10.1109\/TPAMI.2021.3050918","article-title":"Spatiotemporal co-attention recurrent neural networks for human-skeleton motion prediction","volume":"44","author":"Shu","year":"2021","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"681","DOI":"10.1109\/TPAMI.2021.3139918","article-title":"Investigating pose representations and motion contexts modeling for 3D motion prediction","volume":"45","author":"Liu","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_21","unstructured":"Lebailly, T., Kiciroglu, S., Salzmann, M., Fua, P., and Wang, W. (December, January 30). Motion prediction using temporal inception module. Proceedings of the Asian Conference on Computer Vision, Kyoto, Japan."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"427","DOI":"10.1016\/j.ins.2020.08.123","article-title":"Efficient human motion prediction using temporal convolutional generative adversarial network","volume":"545","author":"Cui","year":"2021","journal-title":"Inf. Sci."},{"key":"ref_23","doi-asserted-by":"crossref","first-page":"2513","DOI":"10.1007\/s11263-021-01483-7","article-title":"Multi-level motion attention for human motion prediction","volume":"129","author":"Mao","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Medjaouri, O., and Desai, K. (2022, January 18\u201324). Hr-stan: High-resolution spatio-temporal attention network for 3d human motion prediction. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPRW56347.2022.00286"},{"key":"ref_25","unstructured":"Mao, W., Liu, M., Salzmann, M., and Li, H. (November, January 27). Learning trajectory dependencies for human motion prediction. Proceedings of the IEEE International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_26","unstructured":"Shi, L., Zhang, Y., Cheng, J., and Lu, H. (2020). Decoupled spatial-temporal attention network for skeleton-based action recognition. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Aksan, E., Kaufmann, M., Cao, P., and Hilliges, O. (2021, January 1\u20133). A spatio-temporal transformer for 3d human motion prediction. Proceedings of the 2021 International Conference on 3D Vision (3DV), London, UK.","DOI":"10.1109\/3DV53792.2021.00066"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"1143","DOI":"10.1007\/s00371-019-01692-9","article-title":"Efficient convolutional hierarchical autoencoder for human motion prediction","volume":"35","author":"Li","year":"2019","journal-title":"Vis. Comput."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Chiu, H.K., Adeli, E., Wang, B., Huang, D.A., and Niebles, J.C. (2018, January 7\u201311). Action-agnostic human pose forecasting. Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa Village, HI, USA.","DOI":"10.1109\/WACV.2019.00156"},{"key":"ref_30","unstructured":"Guo, X., and Choi, J. (February, January 27). Human motion prediction via learning local structure representations and temporal dependencies. Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA."},{"key":"ref_31","unstructured":"Bai, S., Kolter, J.Z., and Koltun, V. (2018). An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Lea, C., Flynn, M.D., Vidal, R., Reiter, A., and Hager, G.D. (2017, January 21\u201326). Temporal convolutional networks for action segmentation and detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.113"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Farha, Y.A., and Gall, J. (2019, January 15\u201320). Ms-tcn: Multi-stage temporal convolutional network for action segmentation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00369"},{"key":"ref_34","unstructured":"Dauphin, Y.N., Fan, A., Auli, M., and Grangier, D. (2017, January 6\u201311). Language modeling with gated convolutional networks. Proceedings of the International conference on machine learning. PMLR, Sydney, Australia."},{"key":"ref_35","unstructured":"van den Oord, A., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. (2016). Wavenet: A generative model for raw audio. arXiv."},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Pavllo, D., Feichtenhofer, C., Grangier, D., and Auli, M. (2019, January 15\u201320). 3d human pose estimation in video with temporal convolutions and semi-supervised training. Proceedings of the IEEE\/CVF conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00794"},{"key":"ref_37","unstructured":"Yu, F., and Koltun, V. (2015). Multi-scale context aggregation by dilated convolutions. arXiv."},{"key":"ref_38","doi-asserted-by":"crossref","unstructured":"Reis, M.S. (2019). Multiscale and multi-granularity process analytics: A review. Processes, 7.","DOI":"10.3390\/pr7020061"},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"851","DOI":"10.1007\/s40747-022-00834-2","article-title":"Multi-granularity scenarios understanding network for trajectory prediction","volume":"9","author":"Yang","year":"2023","journal-title":"Complex Intell. Syst."},{"key":"ref_40","unstructured":"Chorowski, J.K., Bahdanau, D., Serdyuk, D., Cho, K., and Bengio, Y. (2015). Attention-based models for speech recognition. arXiv."},{"key":"ref_41","unstructured":"Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, \u0141., and Polosukhin, I. (2017). Attention is all you need. arXiv."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1315","DOI":"10.1109\/TIP.2015.2397314","article-title":"Vector sparse representation of color image using quaternion matrix analysis","volume":"24","author":"Xu","year":"2015","journal-title":"IEEE Trans. Image Process."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Tang, Y., Ma, L., Liu, W., and Zheng, W. (2018). Long-term human motion prediction by modeling motion context and enhancing motion dynamic. arXiv.","DOI":"10.24963\/ijcai.2018\/130"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Cai, Y., Huang, L., Wang, Y., Cham, T.J., Cai, J., Yuan, J., Liu, J., Yang, X., Zhu, Y., and Shen, X. (2020, January 23\u201328). Learning progressive joint propagation for human motion prediction. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58571-6_14"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Ma, T., Nie, Y., Long, C., Zhang, Q., and Li, G. (2022, January 18\u201324). Progressively Generating Better Initial Guesses Towards Next Stages for High-Quality Human Motion Prediction. Proceedings of the Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00633"},{"key":"ref_46","unstructured":"Loshchilov, I., and Hutter, F. (2017). Decoupled weight decay regularization. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","first-page":"1325","DOI":"10.1109\/TPAMI.2013.248","article-title":"Human3. 6m: Large scale datasets and predictive methods for 3d human sensing in natural environments","volume":"36","author":"Ionescu","year":"2013","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_48","doi-asserted-by":"crossref","unstructured":"Mao, W., Liu, M., and Salzmann, M. (2020, January 23\u201328). History Repeats Itself: Human Motion Prediction via Motion Attention. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58568-6_28"},{"key":"ref_49","doi-asserted-by":"crossref","unstructured":"Dang, L., Nie, Y., Long, C., Zhang, Q., and Li, G. (2021, January 10\u201317). MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion Prediction. Proceedings of the International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01127"}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5653\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T19:56:43Z","timestamp":1760126203000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/12\/5653"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,6,16]]},"references-count":49,"journal-issue":{"issue":"12","published-online":{"date-parts":[[2023,6]]}},"alternative-id":["s23125653"],"URL":"https:\/\/doi.org\/10.3390\/s23125653","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,6,16]]}}}