{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,2]],"date-time":"2026-08-02T19:37:38Z","timestamp":1785699458789,"version":"3.56.0"},"reference-count":63,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T00:00:00Z","timestamp":1781481600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/legalcode"}],"funder":[{"name":"ARC Centre of Excellence for Automated Decision-Making and Society","award":["CE200100005"],"award-info":[{"award-number":["CE200100005"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2026,6,15]]},"abstract":"<jats:p>The goal of creating intelligent, human-centered wearable systems for continuous activity understanding faces a fundamental trade-off: Egocentric video-based models capture rich semantic information and have demonstrated strong performance in human activity recognition (HAR), but their high power consumption, privacy concerns, and dependence on lighting limit their feasibility for continuous on-device recognition. In contrast, inertial measurement unit (IMU) sensors offer an energy-efficient, privacy-preserving alternative, yet lack large-scale annotated datasets, leading to weaker generalization. To bridge this gap, we propose COMODO, a cross-modal self-supervised distillation framework that transfers semantic knowledge from video to IMU without requiring labels. COMODO leverages a pretrained and frozen video encoder to construct a dynamic instance queue to align the feature distributions of video and IMU embeddings. This enables the IMU encoder to inherit rich semantic structure from video while maintaining its efficiency for real-world applications. Experiments on multiple egocentric HAR datasets show that COMODO consistently improves downstream performance, matching or surpassing fully supervised models, and demonstrating strong cross-dataset generalization. Benefiting from its simplicity and flexibility, COMODO is compatible with diverse pretrained video and time-series models, offering the potential to leverage more powerful teacher and student foundation models in future ubiquitous computing research. The code is available at this repository: https:\/\/github.com\/cruiseresearchgroup\/COMODO.<\/jats:p>","DOI":"10.1145\/3810218","type":"journal-article","created":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T17:06:41Z","timestamp":1781543201000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition"],"prefix":"10.1145","volume":"10","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-8617-9635","authenticated-orcid":false,"given":"Baiyu","family":"Chen","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, The University of New South Wales, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-0896-1941","authenticated-orcid":false,"given":"Wilson","family":"Wongso","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, The University of New South Wales, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-6900-833X","authenticated-orcid":false,"given":"Zechen","family":"Li","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, The University of New South Wales, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4297-6274","authenticated-orcid":false,"given":"Yonchanok","family":"Khaokaew","sequence":"additional","affiliation":[{"name":"King Mongkut's University of Technology North Bangkok, Bangkok, Thailand and School of Computer Science and Engineering, The University of New South Wales, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1700-9215","authenticated-orcid":false,"given":"Hao","family":"Xue","sequence":"additional","affiliation":[{"name":"The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China and School of Computer Science and Engineering, The University of New South Wales, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1237-1664","authenticated-orcid":false,"given":"Flora D.","family":"Salim","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, The University of New South Wales, Sydney, Australia"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2026,6,15]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448083"},{"key":"e_1_2_1_2_1","first-page":"3","article-title":"A public domain dataset for human activity recognition using smartphones","volume":"3","author":"Anguita Davide","year":"2013","unstructured":"Davide Anguita, Alessandro Ghio, Luca Oneto, Xavier Parra, Jorge Luis Reyes-Ortiz, et al. 2013. A public domain dataset for human activity recognition using smartphones. In Esann, Vol. 3. 3\u20134.","journal-title":"Esann"},{"key":"e_1_2_1_3_1","first-page":"4","article-title":"Is space-time attention all you need for video understanding?","volume":"2","author":"Bertasius Gedas","year":"2021","unstructured":"Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding?. In ICML, Vol. 2. 4.","journal-title":"ICML"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642242"},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00951"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.502"},{"key":"e_1_2_1_8_1","volume-title":"International journal of computer vision 73, 3","author":"Charpiat Guillaume","year":"2007","unstructured":"Guillaume Charpiat, Pierre Maurel, J-P Pons, Renaud Keriven, and Olivier Faugeras. 2007. Generalized gradients: Priors on minimization flows. International journal of computer vision 73, 3 (2007), 325\u2013344."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2012.12.014"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3616855.3635795"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550316"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3717608"},{"key":"e_1_2_1_13_1","volume-title":"SEED: Self-supervised Distillation For Visual Representation. International Conference on Learning Representations","author":"Fang Zhiyuan","year":"2021","unstructured":"Zhiyuan Fang, Jianfeng Wang, Lijuan Wang, Lei Zhang, Yezhou Yang, and Zicheng Liu. 2021. SEED: Self-supervised Distillation For Visual Representation. International Conference on Learning Representations (2021)."},{"key":"e_1_2_1_14_1","volume-title":"Mantis: Lightweight Calibrated Foundation Model for User-Friendly Time Series Classification. arXiv preprint arXiv:2502.15637","author":"Feofanov Vasilii","year":"2025","unstructured":"Vasilii Feofanov, Songkang Wen, Marius Alonso, Romain Ilbert, Hongbo Guo, Malik Tiomoko, Lujia Pan, Jianfeng Zhang, and levgen Redko. 2025. Mantis: Lightweight Calibrated Foundation Model for User-Friendly Time Series Classification. arXiv preprint arXiv:2502.15637 (2025)."},{"key":"e_1_2_1_15_1","volume-title":"Unsupervised scalable representation learning for multivariate time series. Advances in neural information processing systems 32","author":"Franceschi Jean-Yves","year":"2019","unstructured":"Jean-Yves Franceschi, Aymeric Dieuleveut, and Martin Jaggi. 2019. Unsupervised scalable representation learning for multivariate time series. Advances in neural information processing systems 32 (2019)."},{"key":"e_1_2_1_16_1","volume-title":"2025 IEEE International Conference on Pervasive Computing and Communications (PerCom). IEEE, 1\u201312","author":"Fritsch Stefan Gerd","year":"2025","unstructured":"Stefan Gerd Fritsch, Cennet Oguz, Vitor Fortes Rey, Lala Ray, Maximilian Kiefer-Emmanouilidis, and Paul Lukowicz. 2025. Mujo: Multimodal joint feature space learning for human activity recognition. In 2025 IEEE International Conference on Pervasive Computing and Communications (PerCom). IEEE, 1\u201312."},{"key":"e_1_2_1_17_1","first-page":"1","article-title":"Mmtsa: Multi-modal temporal segment attention network for efficient human activity recognition","volume":"7","author":"Gao Ziqi","year":"2023","unstructured":"Ziqi Gao, Yuntao Wang, Jianguo Chen, Junliang Xing, Shwetak Patel, Xin Liu, and Yuanchun Shi. 2023. Mmtsa: Multi-modal temporal segment attention network for efficient human activity recognition. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 7, 3 (2023), 1\u201326.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_2_1_18_1","volume-title":"Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 2755\u20132764","author":"Garcia Nuno Cruz","year":"2021","unstructured":"Nuno Cruz Garcia, Sarah Adel Bargal, Vitaly Ablavsky, Pietro Morerio, Vittorio Murino, and Stan Sclaroff. 2021. Distillation multiple choice learning for multimodal action recognition. In Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision. 2755\u20132764."},{"key":"e_1_2_1_19_1","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6481\u20136491","author":"Gong Xinyu","year":"2023","unstructured":"Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, Yilei Li, Zhangyang Wang, and Rakesh Ranjan. 2023. Mmg-ego4d: Multimodal generalization in egocentric action recognition. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 6481\u20136491."},{"key":"e_1_2_1_20_1","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Goswami Mononito","year":"2024","unstructured":"Mononito Goswami, Konrad Szafer, Arjun Choudhry, Yifu Cai, Shuo Li, and Artur Dubrawski. 2024. MOMENT: a family of open time-series foundation models. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML'24). JMLR.org, Article 642, 38 pages."},{"key":"e_1_2_1_21_1","volume-title":"Proceedings of the IEEE international conference on computer vision. 5842\u20135850","author":"Goyal Raghav","year":"2017","unstructured":"Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. 2017. The\u201c something something\u201d video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision. 5842\u20135850."},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01842"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01834"},{"key":"e_1_2_1_24_1","volume-title":"MiniLLM: Knowledge Distillation of Large Language Models. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=5h0qf7IBZZ","author":"Gu Yuxian","year":"2024","unstructured":"Yuxian Gu, Li Dong, Furu Wei, and Minlie Huang. 2024. MiniLLM: Knowledge Distillation of Large Language Models. In The Twelfth International Conference on Learning Representations. https:\/\/openreview.net\/forum?id=5h0qf7IBZZ"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3729467"},{"key":"e_1_2_1_26_1","volume-title":"European Conference on Computer Vision. Springer, 182\u2013199","author":"Hatano Masashi","year":"2024","unstructured":"Masashi Hatano, Ryo Hachiuma, Ryo Fujii, and Hideo Saito. 2024. Multimodal cross-domain few-shot learning for egocentric action recognition. In European Conference on Computer Vision. Springer, 182\u2013199."},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3659597"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3517246"},{"key":"e_1_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411841"},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/3530910"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3678545"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01826"},{"key":"e_1_2_1_34_1","volume-title":"Zara: Training-free motion time-series reasoning via evidence-grounded llm agents. arXiv preprint arXiv:2508.04038","author":"Li Zechen","year":"2025","unstructured":"Zechen Li, Baiyu Chen, Hao Xue, and Flora D Salim. 2025. Zara: Training-free motion time-series reasoning via evidence-grounded llm agents. arXiv preprint arXiv:2508.04038 (2025)."},{"key":"e_1_2_1_35_1","volume-title":"Sensorllm: Aligning large language models with motion sensors for human activity recognition. arXiv preprint arXiv:2410.10624","author":"Li Zechen","year":"2024","unstructured":"Zechen Li, Shohreh Deldari, Linyao Chen, Hao Xue, and Flora D Salim. 2024. Sensorllm: Aligning large language models with motion sensors for human activity recognition. arXiv preprint arXiv:2410.10624 (2024)."},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3706598.3713391"},{"key":"e_1_2_1_37_1","first-page":"6467","article-title":"ConGen: Unsupervised control and generalization distillation for sentence representation","volume":"2022","author":"Limkonchotiwat Peerat","year":"2022","unstructured":"Peerat Limkonchotiwat, Wuttikorn Ponwitayarat, Lalita Lowphansirikul, Can Udomcharoenchaikit, Ekapol Chuangsuwanich, and Sarana Nutanong. 2022. ConGen: Unsupervised control and generalization distillation for sentence representation. In Findings of the Association for Computational Linguistics: EMNLP 2022. 6467\u20136480.","journal-title":"Findings of the Association for Computational Linguistics: EMNLP"},{"key":"e_1_2_1_38_1","first-page":"2579","article-title":"Visualizing data using t-SNE","author":"van der Maaten Laurens","year":"2008","unstructured":"Laurens van der Maaten and Geoffrey Hinton. 2008. Visualizing data using t-SNE. Journal of machine learning research 9, Nov (2008), 2579\u20132605.","journal-title":"Journal of machine learning research 9"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3699736"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-emnlp.883"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642164"},{"key":"e_1_2_1_42_1","volume-title":"ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 4448\u20134452","author":"Ni Jianyuan","year":"2022","unstructured":"Jianyuan Ni, Raunak Sarbajna, Yang Liu, Anne HH Ngu, and Yan Yan. 2022. Cross-modal knowledge distillation for vision-to-sensor action recognition. In ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 4448\u20134452."},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3432700"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.3390\/s16010115"},{"key":"e_1_2_1_45_1","volume-title":"Christopher Kanan, and Stefan Wermter.","author":"Parisi German I","year":"2019","unstructured":"German I Parisi, Ronald Kemker, Jose L Part, Christopher Kanan, and Stefan Wermter. 2019. Continual lifelong learning with neural networks: A review. Neural networks 113 (2019), 54\u201371."},{"key":"e_1_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00481"},{"key":"e_1_2_1_47_1","volume-title":"International conference on machine learning. PmLR, 8748\u20138763","author":"Radford Alec","year":"2021","unstructured":"Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning. PmLR, 8748\u20138763."},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISWC.2012.13"},{"key":"e_1_2_1_49_1","volume-title":"Antoine Chassang, Carlo Gatta, and Yoshua Bengio.","author":"Romero Adriana","year":"2015","unstructured":"Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio. 2015. FitNets: Hints for Thin Deep Nets. In 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings, Yoshua Bengio and Yann LeCun (Eds.). http:\/\/arxiv.org\/abs\/1412.6550"},{"key":"e_1_2_1_50_1","first-page":"33485","article-title":"Egodistill: Egocentric head motion distillation for efficient video understanding","volume":"36","author":"Tan Shuhan","year":"2023","unstructured":"Shuhan Tan, Tushar Nagarajan, and Kristen Grauman. 2023. Egodistill: Egocentric head motion distillation for efficient video understanding. Advances in Neural Information Processing Systems 36 (2023), 33485\u201333498.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448112"},{"key":"e_1_2_1_52_1","volume-title":"Contrastive Representation Distillation. In International Conference on Learning Representations.","author":"Tian Yonglong","year":"2020","unstructured":"Yonglong Tian, Dilip Krishnan, and Phillip Isola. 2020. Contrastive Representation Distillation. In International Conference on Learning Representations."},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1145\/3494995"},{"key":"e_1_2_1_54_1","volume-title":"Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems 35","author":"Tong Zhan","year":"2022","unstructured":"Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems 35 (2022), 10078\u201310093."},{"key":"e_1_2_1_55_1","volume-title":"International Conference on Learning Representations.","author":"Wu Haixu","year":"2023","unstructured":"Haixu Wu, Tengge Hu, Yong Liu, Hang Zhou, Jianmin Wang, and Mingsheng Long. 2023. TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis. In International Conference on Learning Representations."},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2023.3295899"},{"key":"e_1_2_1_57_1","unstructured":"Zihui Xue Zhengqi Gao Sucheng Ren and Hang Zhao. 2023. The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge Distillation. In ICLR."},{"key":"e_1_2_1_58_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision. 854\u2013863","author":"Xue Zihui","year":"2021","unstructured":"Zihui Xue, Sucheng Ren, Zhengqi Gao, and Hang Zhao. 2021. Multimodal knowledge expansion. In Proceedings of the IEEE\/CVF International Conference on Computer Vision. 854\u2013863."},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i8.20881"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v37i9.26317"},{"key":"e_1_2_1_61_1","volume-title":"European Conference on Computer Vision. Springer, 312\u2013330","author":"Zhang Mingfang","year":"2024","unstructured":"Mingfang Zhang, Yifei Huang, Ruicong Liu, and Yoichi Sato. 2024. Masked video and body-worn IMU autoencoder for egocentric action recognition. In European Conference on Computer Vision. Springer, 312\u2013330."},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370216.2370438"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i12.17325"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3810218","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,6,15]],"date-time":"2026-06-15T17:07:44Z","timestamp":1781543264000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3810218"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,6,15]]},"references-count":63,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2026,6,15]]}},"alternative-id":["10.1145\/3810218"],"URL":"https:\/\/doi.org\/10.1145\/3810218","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,6,15]]},"assertion":[{"value":"2026-06-15","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}