{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,1]],"date-time":"2026-08-01T10:54:48Z","timestamp":1785581688198,"version":"3.56.0"},"reference-count":51,"publisher":"Association for Computing Machinery (ACM)","issue":"5","license":[{"start":{"date-parts":[[2024,2,7]],"date-time":"2024-02-07T00:00:00Z","timestamp":1707264000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Multimedia Comput. Commun. Appl."],"published-print":{"date-parts":[[2024,5,31]]},"abstract":"<jats:p>As a fundamental aspect of human life, two-person interactions contain meaningful information about people\u2019s activities, relationships, and social settings. Human action recognition serves as the foundation for many smart applications, with a strong focus on personal privacy. However, recognizing two-person interactions poses more challenges due to increased body occlusion and overlap compared to single-person actions. In this article, we propose a point cloud-based network named Two-stream Multi-level Dynamic Point Transformer for two-person interaction recognition. Our model addresses the challenge of recognizing two-person interactions by incorporating local-region spatial information, appearance information, and motion information. To achieve this, we introduce a designed frame selection method named Interval Frame Sampling (IFS), which efficiently samples frames from videos, capturing more discriminative information in a relatively short processing time. Subsequently, a frame features learning module and a two-stream multi-level feature aggregation module extract global and partial features from the sampled frames, effectively representing the local-region spatial information, appearance information, and motion information related to the interactions. Finally, we apply a transformer to perform self-attention on the learned features for the final classification. Extensive experiments are conducted on two large-scale datasets, the interaction subsets of NTU RGB+D 60 and NTU RGB+D 120. The results show that our network outperforms state-of-the-art approaches in most standard evaluation settings.<\/jats:p>","DOI":"10.1145\/3639470","type":"journal-article","created":{"date-parts":[[2024,1,5]],"date-time":"2024-01-05T20:13:20Z","timestamp":1704485600000},"page":"1-22","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":8,"title":["Two-stream Multi-level Dynamic Point Transformer for Two-person Interaction Recognition"],"prefix":"10.1145","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-5271-0536","authenticated-orcid":false,"given":"Yao","family":"Liu","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, University of New South Wales, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0000-8604-7667","authenticated-orcid":false,"given":"Gangfeng","family":"Cui","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, University of New South Wales, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-6220-3927","authenticated-orcid":false,"given":"Jiahui","family":"Luo","sequence":"additional","affiliation":[{"name":"School of Computing and Information Systems, University of Melbourne, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7778-8807","authenticated-orcid":false,"given":"Xiaojun","family":"Chang","sequence":"additional","affiliation":[{"name":"Faculty of Engineering and Information Technology, University of Technology Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4149-839X","authenticated-orcid":false,"given":"Lina","family":"Yao","sequence":"additional","affiliation":[{"name":"Data 61, CSIRO and School of Computer Science and Engineering, University of New South Wales, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,2,7]]},"reference":[{"key":"e_1_3_1_2_2","series-title":"Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event","first-page":"813","volume":"139","author":"Bertasius Gedas","year":"2021","unstructured":"Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding?. In Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event(Proceedings of Machine Learning Research, Vol. 139), Marina Meila and Tong Zhang (Eds.). PMLR, 813\u2013824. Retrieved from http:\/\/proceedings.mlr.press\/v139\/bertasius21a.html"},{"key":"e_1_3_1_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2015.12.001"},{"key":"e_1_3_1_4_2","first-page":"56","volume-title":"CAIP 2021: Proceedings of the 1st International Conference on AI for People: Towards Sustainable AI, CAIP 2021, 20-24 November 2021","author":"Chiu Shian-Yu","year":"2021","unstructured":"Shian-Yu Chiu, Kun-Ru Wu, and Yu-Chee Tseng. 2021. Two-person mutual action recognition using joint dynamics and coordinate transformation. In CAIP 2021: Proceedings of the 1st International Conference on AI for People: Towards Sustainable AI, CAIP 2021, 20-24 November 2021. European Alliance for Innovation, 56."},{"key":"e_1_3_1_5_2","doi-asserted-by":"publisher","DOI":"10.1145\/3441628"},{"key":"e_1_3_1_6_2","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly Jakob Uszkoreit and Neil Houlsby. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. In 9th International Conference on Learning Representations (ICLR\u201921 Virtual Event Austria May 3-7 2021) OpenReview.net."},{"key":"e_1_3_1_7_2","first-page":"6804","volume-title":"Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision, ICCV 2021, October 10-17, 2021","author":"Fan Haoqi","year":"2021","unstructured":"Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. 2021. Multiscale vision transformers. In Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision, ICCV 2021, October 10-17, 2021. IEEE, 6804\u20136815. DOI:10.1109\/ICCV48922.2021.00675"},{"key":"e_1_3_1_8_2","first-page":"14204","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021","author":"Fan Hehe","year":"2021","unstructured":"Hehe Fan, Yi Yang, and Mohan S. Kankanhalli. 2021. Point 4D transformer networks for spatio-temporal modeling in point cloud videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021. Computer Vision Foundation\/IEEE, 14204\u201314213. DOI:10.1109\/CVPR46437.2021.01398"},{"key":"e_1_3_1_9_2","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Multimedia and Expo, ICME 2022, July 18-22, 2022","author":"Gao Feng","year":"2022","unstructured":"Feng Gao, Hailun Xia, and Zhihao Tang. 2022. Attention interactive graph convolutional network for skeleton-based human interaction recognition. In Proceedings of the IEEE International Conference on Multimedia and Expo, ICME 2022, July 18-22, 2022. IEEE, 1\u20136. DOI:10.1109\/ICME52920.2022.9859618"},{"key":"e_1_3_1_10_2","doi-asserted-by":"crossref","first-page":"101","DOI":"10.1007\/978-3-319-11839-0_9","volume-title":"Proceedings of the 5th International Workshop Human Behavior Understanding.","author":"Gemeren Coert Van","year":"2014","unstructured":"Coert Van Gemeren, Robby T. Tan, Ronald Poppe, and Remco C. Veltkamp. 2014. Dyadic interaction detection from pose and flow. In Proceedings of the 5th International Workshop Human Behavior Understanding.Hyun Soo Park, Albert Ali Salah, Yong Jae Lee, Louis-Philippe Morency, Yaser Sheikh, and Rita Cucchiara (Eds.), Lecture Notes in Computer Science, Vol. 7948, Springer, 101\u2013115. DOI:10.1007\/978-3-319-11839-0_9"},{"key":"e_1_3_1_11_2","first-page":"1533","volume-title":"Proceedings of the 25th International Joint Conference on Artificial Intelligence","author":"Hammerla Nils Y.","year":"2016","unstructured":"Nils Y. Hammerla, Shane Halloran, and Thomas Pl\u00f6tz. 2016. Deep, convolutional, and recurrent models for human activity recognition using wearables. In Proceedings of the 25th International Joint Conference on Artificial Intelligence. IJCAI\/AAAI Press, 1533\u20131540. Retrieved from http:\/\/www.ijcai.org\/Abstract\/16\/220"},{"key":"e_1_3_1_12_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.sigpro.2015.08.006"},{"key":"e_1_3_1_13_2","doi-asserted-by":"publisher","DOI":"10.1145\/3177757"},{"key":"e_1_3_1_14_2","first-page":"2036","volume-title":"Proceedings of the 31st AAAI Conference on Artificial Intelligence","author":"Huang Zhiwu","year":"2017","unstructured":"Zhiwu Huang and Luc Van Gool. 2017. A riemannian network for SPD matrix learning. In Proceedings of the 31st AAAI Conference on Artificial Intelligence. Satinder Singh and Shaul Markovitch (Eds.). AAAI Press, 2036\u20132042. Retrieved from http:\/\/aaai.org\/ocs\/index.php\/AAAI\/AAAI17\/paper\/view\/14633"},{"key":"e_1_3_1_15_2","first-page":"231","volume-title":"Proceedings of the 2022 IEEE International Conference on Image Processing","author":"Ito Yoshiki","year":"2022","unstructured":"Yoshiki Ito, Quan Kong, Kenichi Morita, and Tomoaki Yoshinaga. 2022. Efficient and accurate skeleton-based two-person interaction recognition using inter-and intra-body graphs. In Proceedings of the 2022 IEEE International Conference on Image Processing. IEEE, 231\u2013235. DOI:10.1109\/ICIP46576.2022.9897250"},{"key":"e_1_3_1_16_2","doi-asserted-by":"crossref","unstructured":"Salman Khan Muzammal Naseer Munawar Hayat Syed Waqas Zamir Fahad Shahbaz Khan and Mubarak Shah. 2022. Transformers in vision: A survey. ACM Comput. Surv. 54 10s Article 200 (Sep. 2022) 41 pages.","DOI":"10.1145\/3505244"},{"key":"e_1_3_1_17_2","volume-title":"Proceedings of the 3rd International Conference on Learning Representations","author":"Kingma Diederik P.","year":"2015","unstructured":"Diederik P. Kingma and Jimmy Ba. 2015. Adam: A method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations. Yoshua Bengio and Yann LeCun (Eds.). Retrieved from http:\/\/arxiv.org\/abs\/1412.6980"},{"key":"e_1_3_1_18_2","first-page":"3595","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Li Maosen","year":"2019","unstructured":"Maosen Li, Siheng Chen, Xu Chen, Ya Zhang, Yanfeng Wang, and Qi Tian. 2019. Actional-structural graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Computer Vision Foundation \/ IEEE, 3595\u20133603. DOI:10.1109\/CVPR.2019.00371"},{"key":"e_1_3_1_19_2","unstructured":"Xing Li Qian Huang Zhijian Wang Zhenjie Hou and Tianjin Yang. 2021. SequentialPointNet: A strong parallelized point cloud sequence network for 3D action recognition. arXiv:2111.08492 Retrieved from https:\/\/arxiv.org\/abs\/2111.08492"},{"key":"e_1_3_1_20_2","first-page":"1797","volume-title":"Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Lin Zhi-Hao","year":"2020","unstructured":"Zhi-Hao Lin, Sheng-Yu Huang, and Yu-Chiang Frank Wang. 2020. Convolution in the cloud: Learning deformable kernels in 3D graph convolution networks for point cloud analysis. In Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Computer Vision Foundation \/ IEEE, 1797\u20131806. DOI:10.1109\/CVPR42600.2020.00187"},{"key":"e_1_3_1_21_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2019.05.020"},{"key":"e_1_3_1_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2916873"},{"key":"e_1_3_1_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2898954"},{"key":"e_1_3_1_24_2","series-title":"Proceedings of the 14th European Conference Computer Vision - ECCV 2016.","first-page":"816","volume":"9907","author":"Liu Jun","year":"2016","unstructured":"Jun Liu, Amir Shahroudy, Dong Xu, and Gang Wang. 2016. Spatio-temporal LSTM with trust gates for 3D human action recognition. In Proceedings of the 14th European Conference Computer Vision - ECCV 2016.Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling (Eds.), Lecture Notes in Computer Science, Vol. 9907, Springer, 816\u2013833. DOI:10.1007\/978-3-319-46487-9_50"},{"key":"e_1_3_1_25_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2785279"},{"key":"e_1_3_1_26_2","first-page":"3671","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition","author":"Liu Jun","year":"2017","unstructured":"Jun Liu, Gang Wang, Ping Hu, Ling-Yu Duan, and Alex C. Kot. 2017. Global context-aware attention LSTM networks for 3D action recognition. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. IEEE Computer Society, 3671\u20133680. DOI:10.1109\/CVPR.2017.391"},{"key":"e_1_3_1_27_2","doi-asserted-by":"publisher","DOI":"10.1002\/INT.23087"},{"key":"e_1_3_1_28_2","first-page":"13359","volume-title":"Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision","author":"Nguyen Xuan Son","year":"2021","unstructured":"Xuan Son Nguyen. 2021. GeomNet: A neural network based on riemannian geometries of SPD matrix space and cholesky space for 3D skeleton-based interaction recognition. In Proceedings of the 2021 IEEE\/CVF International Conference on Computer Vision. IEEE, 13359\u201313369. DOI:10.1109\/ICCV48922.2021.01313"},{"key":"e_1_3_1_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2021.3050642"},{"key":"e_1_3_1_30_2","first-page":"77","volume-title":"Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition","author":"Qi Charles Ruizhongtai","year":"2017","unstructured":"Charles Ruizhongtai Qi, Hao Su, Kaichun Mo, and Leonidas J. Guibas. 2017. PointNet: Deep learning on point sets for 3D classification and segmentation. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition. IEEE Computer Society, 77\u201385. DOI:10.1109\/CVPR.2017.16"},{"key":"e_1_3_1_31_2","first-page":"5099","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017","author":"Qi Charles Ruizhongtai","year":"2017","unstructured":"Charles Ruizhongtai Qi, Li Yi, Hao Su, and Leonidas J. Guibas. 2017. PointNet++: Deep hierarchical feature learning on point sets in a metric space. In Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017. Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.). 5099\u20135108. Retrieved from https:\/\/proceedings.neurips.cc\/paper\/2017\/hash\/d8bf84be3800d12f74d8b05e9b89836f-Abstract.html"},{"issue":"1","key":"e_1_3_1_32_2","doi-asserted-by":"crossref","first-page":"110","DOI":"10.3390\/app7010110","article-title":"A comprehensive review on handcrafted and learning-based action representation approaches for human activity recognition","volume":"7","author":"Sargano Allah Bux","year":"2017","unstructured":"Allah Bux Sargano, Plamen Angelov, and Zulfiqar Habib. 2017. A comprehensive review on handcrafted and learning-based action representation approaches for human activity recognition. Applied Sciences 7, 1 (2017), 110.","journal-title":"Applied Sciences"},{"key":"e_1_3_1_33_2","first-page":"1010","volume-title":"Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition","author":"Shahroudy Amir","year":"2016","unstructured":"Amir Shahroudy, Jun Liu, Tian-Tsong Ng, and Gang Wang. 2016. NTU RGB+D: A large scale dataset for 3D human activity analysis. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition. IEEE Computer Society, 1010\u20131019. DOI:10.1109\/CVPR.2016.115"},{"key":"e_1_3_1_34_2","first-page":"12026","volume-title":"Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition","author":"Shi Lei","year":"2019","unstructured":"Lei Shi, Yifan Zhang, Jian Cheng, and Hanqing Lu. 2019. Two-stream adaptive graph convolutional networks for skeleton-based action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. Computer Vision Foundation \/ IEEE, 12026\u201312035. DOI:10.1109\/CVPR.2019.01230"},{"key":"e_1_3_1_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2020.3028207"},{"key":"e_1_3_1_36_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485665"},{"key":"e_1_3_1_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2019.102799"},{"key":"e_1_3_1_38_2","doi-asserted-by":"publisher","DOI":"10.1145\/3472722"},{"key":"e_1_3_1_39_2","doi-asserted-by":"crossref","first-page":"47","DOI":"10.1145\/3132515.3132517","volume-title":"Proceedings of the Workshop on Multimodal Understanding of Social, Affective and Subjective Attributes","author":"Trabelsi Rim","year":"2017","unstructured":"Rim Trabelsi, Jagannadan Varadarajan, Yong Pei, Le Zhang, Issam Jabri, Ammar Bouallegue, and Pierre Moulin. 2017. Robust multi-modal cues for dyadic human interaction recognition. In Proceedings of the Workshop on Multimodal Understanding of Social, Affective and Subjective Attributes. Xavier Alameda-Pineda, Miriam Redi, Mohammad Soleymani, Nicu Sebe, Shih-Fu Chang, and Samuel D. Gosling (Eds.). ACM, 47\u201353. DOI:10.1145\/3132515.3132517"},{"key":"e_1_3_1_40_2","first-page":"5998","volume-title":"Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017Isabelle Guyon, Ulrike von Luxburg, Samy Bengio, Hanna M. Wallach, Rob Fergus, S. V. N. Vishwanathan, and Roman Garnett (Eds.). 5998\u20136008."},{"key":"e_1_3_1_41_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2018.04.007"},{"key":"e_1_3_1_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2021.3112214"},{"key":"e_1_3_1_43_2","first-page":"508","volume-title":"Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Wang Yancheng","year":"2020","unstructured":"Yancheng Wang, Yang Xiao, Fu Xiong, Wenxiang Jiang, Zhiguo Cao, Joey Tianyi Zhou, and Junsong Yuan. 2020. 3DV: 3D dynamic voxel for action recognition in depth video. In Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. Computer Vision Foundation \/ IEEE, 508\u2013517. DOI:10.1109\/CVPR42600.2020.00059"},{"key":"e_1_3_1_44_2","series-title":"Proceedings of the 15th European Conference Computer Vision - ECCV 2018.","first-page":"3","volume":"11211","author":"Woo Sanghyun","year":"2018","unstructured":"Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. 2018. CBAM: Convolutional block attention module. In Proceedings of the 15th European Conference Computer Vision - ECCV 2018.Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss (Eds.), Lecture Notes in Computer Science, Vol. 11211, Springer, 3\u201319. DOI:10.1007\/978-3-030-01234-2_1"},{"key":"e_1_3_1_45_2","doi-asserted-by":"publisher","DOI":"10.1109\/THMS.2017.2776211"},{"key":"e_1_3_1_46_2","doi-asserted-by":"crossref","first-page":"357","DOI":"10.1109\/WACV.2015.54","volume-title":"Proceedings of the 2015 IEEE Winter Conference on Applications of Computer Vision","author":"Xia Lu","year":"2015","unstructured":"Lu Xia, Ilaria Gori, Jake K. Aggarwal, and Michael S. Ryoo. 2015. Robot-centric activity recognition from first-person RGB-D videos. In Proceedings of the 2015 IEEE Winter Conference on Applications of Computer Vision. IEEE Computer Society, 357\u2013364. DOI:10.1109\/WACV.2015.54"},{"key":"e_1_3_1_47_2","doi-asserted-by":"publisher","DOI":"10.1145\/3450410"},{"key":"e_1_3_1_48_2","doi-asserted-by":"publisher","DOI":"10.1145\/3538749"},{"key":"e_1_3_1_49_2","first-page":"7444","volume-title":"Proceedings of the 32nd AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18)","author":"Yan Sijie","year":"2018","unstructured":"Sijie Yan, Yuanjun Xiong, and Dahua Lin. 2018. Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proceedings of the 32nd AAAI Conference on Artificial Intelligence, (AAAI-18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence (EAAI-18). Sheila A. McIlraith and Kilian Q. Weinberger (Eds.). AAAI Press, 7444\u20137452. Retrieved from https:\/\/www.aaai.org\/ocs\/index.php\/AAAI\/AAAI18\/paper\/view\/17135"},{"key":"e_1_3_1_50_2","first-page":"2166","volume-title":"Proceedings of the IEEE International Conference on Image Processing","author":"Yang Chao-Lung","year":"2020","unstructured":"Chao-Lung Yang, Aji Setyoko, Hendrik Tampubolon, and Kai-Lung Hua. 2020. Pairwise adjacency matrix on spatial temporal graph convolution network for skeleton-based two-person interaction recognition. In Proceedings of the IEEE International Conference on Image Processing. IEEE, 2166\u20132170. DOI:10.1109\/ICIP40778.2020.9190680"},{"key":"e_1_3_1_51_2","first-page":"28","volume-title":"Proceedings of the 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops","author":"Yun Kiwon","year":"2012","unstructured":"Kiwon Yun, Jean Honorio, Debaleena Chattopadhyay, Tamara L. Berg, and Dimitris Samaras. 2012. Two-person interaction detection using body-pose features and multiple instance learning. In Proceedings of the 2012 IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops. IEEE Computer Society, 28\u201335. DOI:10.1109\/CVPRW.2012.6239234"},{"key":"e_1_3_1_52_2","doi-asserted-by":"publisher","DOI":"10.3390\/A16040190"}],"container-title":["ACM Transactions on Multimedia Computing, Communications, and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639470","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3639470","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T22:53:37Z","timestamp":1750287217000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3639470"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,2,7]]},"references-count":51,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,5,31]]}},"alternative-id":["10.1145\/3639470"],"URL":"https:\/\/doi.org\/10.1145\/3639470","relation":{},"ISSN":["1551-6857","1551-6865"],"issn-type":[{"value":"1551-6857","type":"print"},{"value":"1551-6865","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,2,7]]},"assertion":[{"value":"2023-07-21","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-12-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-02-07","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}