{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,21]],"date-time":"2026-07-21T14:44:35Z","timestamp":1784645075995,"version":"3.55.0"},"reference-count":58,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2022,7,4]],"date-time":"2022-07-04T00:00:00Z","timestamp":1656892800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2022,7,4]]},"abstract":"<jats:p>Recent advances in sensor based human activity recognition (HAR) have exploited deep hybrid networks to improve the performance. These hybrid models combine Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs) to leverage their complementary advantages, and achieve impressive results. However, the roles and associations of different sensors in HAR are not fully considered by these models, leading to insufficient multi-modal fusion. Besides, the commonly used RNNs in HAR suffer from the 'forgetting' defect, which raises difficulties in capturing long-term information. To tackle these problems, an HAR framework composed of an Inertial Measurement Unit (IMU) fusion block and an applied ConvTransformer subnet is proposed in this paper. Inspired by the complementary filter, our IMU fusion block performs multi-modal fusion of commonly used sensors according to their physical relationships. Consequently, the features of different modalities can be aggregated more effectively. Then, the extracted features are fed into the applied ConvTransformer subnet for classification. Thanks to its convolutional subnet and self-attention layers, ConvTransformer can better capture local features and construct long-term dependencies. Extensive experiments on eight benchmark datasets demonstrate the superior performance of our framework. The source code will be published soon.<\/jats:p>","DOI":"10.1145\/3534584","type":"journal-article","created":{"date-parts":[[2022,7,7]],"date-time":"2022-07-07T18:50:18Z","timestamp":1657219818000},"page":"1-26","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":45,"title":["IF-ConvTransformer"],"prefix":"10.1145","volume":"6","author":[{"given":"Ye","family":"Zhang","sequence":"first","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Longguang","family":"Wang","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Huiling","family":"Chen","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Aosheng","family":"Tian","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shilin","family":"Zhou","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Yulan","family":"Guo","sequence":"additional","affiliation":[{"name":"College of Electronic Science and Technology, National University of Defense Technology, Changsha, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2022,7,7]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448083"},{"key":"e_1_2_1_2_1","first-page":"3","article-title":"A public domain dataset for human activity recognition using smartphones","volume":"3","author":"Anguita Davide","year":"2013","unstructured":"Davide Anguita, Alessandro Ghio, Luca Oneto, Xavier Parra, Jorge Luis Reyes-Ortiz, et al. 2013. A public domain dataset for human activity recognition using smartphones. In Esann, Vol. 3. 3.","journal-title":"Esann"},{"key":"e_1_2_1_3_1","volume-title":"An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271","author":"Bai Shaojie","year":"2018","unstructured":"Shaojie Bai, J Zico Kolter, and Vladlen Koltun. 2018. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271 (2018)."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1109\/SMC42975.2020.9283381"},{"key":"e_1_2_1_5_1","volume-title":"International Conference on Information and Communication Technologies for Ageing Well and E-Health. Springer, 100--118","author":"Chatzaki Charikleia","year":"2016","unstructured":"Charikleia Chatzaki, Matthew Pediaditis, George Vavoulas, and Manolis Tsiknakis. 2016. Human daily activity and fall recognition using a smartphone's acceleration sensor. In International Conference on Information and Communication Technologies for Ageing Well and E-Health. Springer, 100--118."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2012.12.014"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3447744"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.2991\/icaita-16.2016.13"},{"key":"e_1_2_1_9_1","volume-title":"Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805","author":"Devlin Jacob","year":"2018","unstructured":"Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805 (2018)."},{"key":"e_1_2_1_10_1","unstructured":"Alexey Dosovitskiy Lucas Beyer Alexander Kolesnikov Dirk Weissenborn Xiaohua Zhai Thomas Unterthiner Mostafa Dehghani Matthias Minderer Georg Heigold Sylvain Gelly et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020)."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2014.6854641"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00779-010-0293-9"},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2858933"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/3090076"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/IJCNN.2016.7727224"},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.5555\/3060832.3060835"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3410531.3414306"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.123"},{"key":"e_1_2_1_19_1","volume-title":"Jessica Sena, and William Robson Schwartz.","author":"Jordao Artur","year":"2018","unstructured":"Artur Jordao, Antonio C Nazare Jr, Jessica Sena, and William Robson Schwartz. 2018. Human activity recognition based on wearable sensor data: A standardization of the state-of-the-art. arXiv preprint arXiv:1806.05226 (2018)."},{"key":"e_1_2_1_20_1","doi-asserted-by":"crossref","unstructured":"Dongwon Jung and Panagiotis Tsiotras. 2007. Inertial attitude and position reference system development for a small UAV. In In AIAA Infotech@ Aerospace 2007 Conference and Exhibit. 2763--2778.","DOI":"10.2514\/6.2007-2763"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.3390\/s20020340"},{"key":"e_1_2_1_22_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267242.3267258"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3432208"},{"key":"e_1_2_1_25_1","volume-title":"ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis. arXiv preprint arXiv:2011.10185","author":"Liu Zhouyong","year":"2020","unstructured":"Zhouyong Liu, Shun Luo, Wubin Li, Jingben Lu, Yufan Wu, Chunguo Li, and Luxi Yang. 2020. ConvTransformer: A Convolutional Transformer Network for Video Frame Synthesis. arXiv preprint arXiv:2011.10185 (2020)."},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2019\/431"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.3233\/FAIA200236"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/CDC.2005.1582367"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAC.2008.923738"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3302505.3310068"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.pmcj.2020.101132"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3123021.3123046"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267242.3267287"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.3390\/s16010115"},{"key":"e_1_2_1_35_1","volume-title":"Twenty-Second International Joint Conference on Artificial Intelligence.","author":"Pl\u00f6tz Thomas","year":"2011","unstructured":"Thomas Pl\u00f6tz, Nils Y Hammerla, and Patrick L Olivier. 2011. Feature learning for activity recognition in ubiquitous computing. In Twenty-Second International Joint Conference on Artificial Intelligence."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISWC.2012.13"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2015.07.085"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.74"},{"key":"e_1_2_1_39_1","volume-title":"Striving for simplicity: The all convolutional net. arXiv preprint arXiv:1412.6806","author":"Springenberg Jost Tobias","year":"2014","unstructured":"Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller. 2014. Striving for simplicity: The all convolutional net. arXiv preprint arXiv:1412.6806 (2014)."},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.01625"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/2809695.2809718"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.308"},{"key":"e_1_2_1_43_1","volume-title":"Efficient transformers: A survey. arXiv preprint arXiv:2009.06732","author":"Tay Yi","year":"2020","unstructured":"Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler. 2020. Efficient transformers: A survey. arXiv preprint arXiv:2009.06732 (2020)."},{"key":"e_1_2_1_44_1","volume-title":"Smartphone based human activity monitoring and recognition using ML and DL: a comprehensive survey. Journal of Ambient Intelligence and Humanized Computing","author":"Thakur Dipanwita","year":"2020","unstructured":"Dipanwita Thakur and Suparna Biswas. 2020. Smartphone based human activity monitoring and recognition using ML and DL: a comprehensive survey. Journal of Ambient Intelligence and Humanized Computing (2020), 1--12."},{"key":"e_1_2_1_45_1","volume-title":"Hierarchical Self Attention Based Autoencoder for Open-Set Human Activity Recognition. In Pacific-Asia Conference on Knowledge Discovery and Data Mining. Springer, 351--363","author":"Tonmoy M","year":"2021","unstructured":"M Tonmoy, Saif Mahmud, AKM Mahbubur Rahman, M Ashraful Amin, and Amin Ahsan Ali. 2021. Hierarchical Self Attention Based Autoencoder for Open-Set Human Activity Recognition. In Pacific-Asia Conference on Knowledge Discovery and Data Mining. Springer, 351--363."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0061691"},{"key":"e_1_2_1_47_1","volume-title":"Attention is all you need. Advances in Neural Information Processing Systems 30","author":"Vaswani Ashish","year":"2017","unstructured":"Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, \u0141ukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in Neural Information Processing Systems 30 (2017)."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2018.02.010"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267305.3267531"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2890675"},{"key":"e_1_2_1_51_1","volume-title":"Twenty-Fourth International Joint Conference on Artificial Intelligence","volume":"15","author":"Yang Jianbo","year":"2015","unstructured":"Jianbo Yang, Minh Nhut Nguyen, Phyo Phyo San, Xiaoli Li, and Shonali Krishnaswamy. 2015. Deep convolutional neural networks on multichannel time series for human activity recognition.. In Twenty-Fourth International Joint Conference on Artificial Intelligence, Vol. 15. Buenos Aires, Argentina, 3995--4001."},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.12.024"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3038912.3052577"},{"key":"e_1_2_1_54_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFOCOM.2019.8737500"},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1145\/3267242.3267286"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTSP.2020.2987728"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1145\/3410530.3414355"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341162.3345571"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3534584","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3534584","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,14]],"date-time":"2025-07-14T04:32:21Z","timestamp":1752467541000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3534584"}},"subtitle":["A Framework for Human Activity Recognition Using IMU Fusion and ConvTransformer"],"short-title":[],"issued":{"date-parts":[[2022,7,4]]},"references-count":58,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2022,7,4]]}},"alternative-id":["10.1145\/3534584"],"URL":"https:\/\/doi.org\/10.1145\/3534584","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,7,4]]},"assertion":[{"value":"2022-07-07","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}