{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,4]],"date-time":"2026-08-04T01:38:03Z","timestamp":1785807483602,"version":"3.56.0"},"reference-count":74,"publisher":"Association for Computing Machinery (ACM)","issue":"2","license":[{"start":{"date-parts":[[2019,6,21]],"date-time":"2019-06-21T00:00:00Z","timestamp":1561075200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100007601","name":"Horizon 2020 Framework Programme","doi-asserted-by":"publisher","award":["737422"],"award-info":[{"award-number":["737422"]}],"id":[{"id":"10.13039\/501100007601","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2019,6,21]]},"abstract":"<jats:p>Deep learning methods are successfully used in applications pertaining to ubiquitous computing, pervasive intelligence, health, and well-being. Specifically, the area of human activity recognition (HAR) is primarily transformed by the convolutional and recurrent neural networks, thanks to their ability to learn semantic representations directly from raw input. However, in order to extract generalizable features massive amounts of well-curated data are required, which is a notoriously challenging task; hindered by privacy issues and annotation costs. Therefore, unsupervised representation learning (i.e., learning without manually labeling the instances) is of prime importance to leverage the vast amount of unlabeled data produced by smart devices. In this work, we propose a novel self-supervised technique for feature learning from sensory data that does not require access to any form of semantic labels, i.e., activity classes. We learn a multi-task temporal convolutional network to recognize transformations applied on an input signal. By exploiting these transformations, we demonstrate that simple auxiliary tasks of the binary classification result in a strong supervisory signal for extracting useful features for the down-stream task. We extensively evaluate the proposed approach on several publicly available datasets for smartphone-based HAR in unsupervised, semi-supervised and transfer learning settings. Our method achieves performance levels superior to or comparable with fully-supervised networks trained directly with activity labels, and it performs significantly better than unsupervised learning through autoencoders. Notably, for the semi-supervised case, the self-supervised features substantially boost the detection rate by attaining a kappa score between 0.7 - 0.8 with only 10 labeled examples per class. We get similar impressive performance even if the features are transferred from a different data source. Self-supervision drastically reduces the requirement of labeled activity data, effectively narrowing the gap between supervised and unsupervised techniques for learning meaningful representations. While this paper focuses on HAR as the application domain, the proposed approach is general and could be applied to a wide variety of problems in other areas.<\/jats:p>","DOI":"10.1145\/3328932","type":"journal-article","created":{"date-parts":[[2019,6,24]],"date-time":"2019-06-24T13:45:01Z","timestamp":1561383901000},"page":"1-30","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":303,"title":["Multi-task Self-Supervised Learning for Human Activity Detection"],"prefix":"10.1145","volume":"3","author":[{"given":"Aaqib","family":"Saeed","sequence":"first","affiliation":[{"name":"Eindhoven University of Technology, Eindhoven, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Tanir","family":"Ozcelebi","sequence":"additional","affiliation":[{"name":"Eindhoven University of Technology, Eindhoven, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Johan","family":"Lukkien","sequence":"additional","affiliation":[{"name":"Eindhoven University of Technology, Eindhoven, The Netherlands"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,6,21]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.13"},{"key":"e_1_2_2_2_1","unstructured":"Davide Anguita Alessandro Ghio Luca Oneto Xavier Parra and Jorge Luis Reyes-Ortiz. 2013. A public domain dataset for human activity recognition using smartphones.. In ESANN.  Davide Anguita Alessandro Ghio Luca Oneto Xavier Parra and Jorge Luis Reyes-Ortiz. 2013. A public domain dataset for human activity recognition using smartphones.. In ESANN."},{"key":"e_1_2_2_3_1","unstructured":"Relja Arandjelovi\u0107 and Andrew Zisserman. 2017. Objects that sound. arXiv preprint arXiv:1712.06651 (2017).  Relja Arandjelovi\u0107 and Andrew Zisserman. 2017. Objects that sound. arXiv preprint arXiv:1712.06651 (2017)."},{"key":"e_1_2_2_4_1","volume-title":"Soundnet: Learning sound representations from unlabeled video. In Advances in Neural Information Processing Systems. 892--900.","author":"Aytar Yusuf","year":"2016"},{"key":"e_1_2_2_5_1","unstructured":"Shaojie Bai J Zico Kolter and Vladlen Koltun. 2018. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271 (2018).  Shaojie Bai J Zico Kolter and Vladlen Koltun. 2018. An empirical evaluation of generic convolutional and recurrent networks for sequence modeling. arXiv preprint arXiv:1803.01271 (2018)."},{"key":"e_1_2_2_6_1","volume-title":"Proceedings of ICML workshop on unsupervised and transfer learning. 37--49","author":"Baldi Pierre","year":"2012"},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1137\/1.9781611972818.60"},{"key":"e_1_2_2_8_1","doi-asserted-by":"crossref","unstructured":"Yoshua Bengio Aaron Courville and Pascal Vincent. 2013. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence 35 8 (2013) 1798--1828.  Yoshua Bengio Aaron Courville and Pascal Vincent. 2013. Representation learning: A review and new perspectives. IEEE transactions on pattern analysis and machine intelligence 35 8 (2013) 1798--1828.","DOI":"10.1109\/TPAMI.2013.50"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.pmcj.2014.05.006"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1007379606734"},{"key":"e_1_2_2_11_1","volume-title":"International Conference on Information and Communication Technologies for Ageing Well and e-Health. Springer, 100--118","author":"Chatzaki Charikleia","year":"2016"},{"key":"e_1_2_2_12_1","unstructured":"Zhicheng Cui Wenlin Chen and Yixin Chen. 2016. Multi-scale convolutional neural networks for time series classification. arXiv preprint arXiv:1603.06995 (2016).  Zhicheng Cui Wenlin Chen and Yixin Chen. 2016. Multi-scale convolutional neural networks for time series classification. arXiv preprint arXiv:1603.06995 (2016)."},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2015.167"},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.226"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.607"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00779-010-0293-9"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3131895"},{"key":"e_1_2_2_18_1","unstructured":"Spyros Gidaris Praveer Singh and Nikos Komodakis. 2018. Unsupervised Representation Learning by Predicting Image Rotations. arXiv preprint arXiv:1803.07728 (2018).  Spyros Gidaris Praveer Singh and Nikos Komodakis. 2018. Unsupervised Representation Learning by Predicting Image Rotations. arXiv preprint arXiv:1803.07728 (2018)."},{"key":"e_1_2_2_19_1","doi-asserted-by":"crossref","unstructured":"Lluis Gomez Yash Patel Mar\u00e7al Rusi\u00f1ol Dimosthenis Karatzas and CV Jawahar. 2017. Self-supervised learning of visual features through embedding images into text topic spaces. arXiv preprint arXiv:1705.08631 (2017).  Lluis Gomez Yash Patel Mar\u00e7al Rusi\u00f1ol Dimosthenis Karatzas and CV Jawahar. 2017. Self-supervised learning of visual features through embedding images into text topic spaces. arXiv preprint arXiv:1705.08631 (2017).","DOI":"10.1109\/CVPR.2017.218"},{"key":"e_1_2_2_20_1","unstructured":"Nils Y Hammerla Shane Halloran and Thomas Ploetz. 2016. Deep convolutional and recurrent models for human activity recognition using wearables. arXiv preprint arXiv:1604.08880 (2016).  Nils Y Hammerla Shane Halloran and Thomas Ploetz. 2016. Deep convolutional and recurrent models for human activity recognition using wearables. arXiv preprint arXiv:1604.08880 (2016)."},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41591-018-0268-3"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/D17-1206"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/P18-1031"},{"key":"e_1_2_2_24_1","doi-asserted-by":"crossref","unstructured":"Simon Jenni and Paolo Favaro. 2018. Self-Supervised Feature Learning by Learning to Spot Artifacts. arXiv preprint arXiv:1806.05024 (2018).  Simon Jenni and Paolo Favaro. 2018. Self-Supervised Feature Learning by Learning to Spot Artifacts. arXiv preprint arXiv:1806.05024 (2018).","DOI":"10.1109\/CVPR.2018.00289"},{"key":"e_1_2_2_25_1","unstructured":"Alex Kendall Yarin Gal and Roberto Cipolla. {n. d.}. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. ({n. d.}).  Alex Kendall Yarin Gal and Roberto Cipolla. {n. d.}. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. ({n. d.})."},{"key":"e_1_2_2_26_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014"},{"key":"e_1_2_2_27_1","unstructured":"Bruno Korbar Du Tran and Lorenzo Torresani. 2018. Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization. In Advances in Neural Information Processing Systems. 7774--7785.  Bruno Korbar Du Tran and Lorenzo Torresani. 2018. Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization. In Advances in Neural Information Processing Systems. 7774--7785."},{"key":"e_1_2_2_28_1","unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105.   Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105."},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/1964897.1964918"},{"key":"e_1_2_2_30_1","first-page":"7","article-title":"Colorization as a proxy task for visual understanding","volume":"2","author":"Larsson Gustav","year":"2017","journal-title":"CVPR"},{"key":"e_1_2_2_31_1","doi-asserted-by":"crossref","unstructured":"Yann LeCun Yoshua Bengio and Geoffrey Hinton. 2015. Deep learning. Nature 521 (27 May 2015) 436 EP --.  Yann LeCun Yoshua Bengio and Geoffrey Hinton. 2015. Deep learning. Nature 521 (27 May 2015) 436 EP --.","DOI":"10.1038\/nature14539"},{"key":"e_1_2_2_32_1","unstructured":"Yann LeCun John S Denker and Sara A Solla. 1990. Optimal brain damage. In Advances in neural information processing systems. 598--605.   Yann LeCun John S Denker and Sara A Solla. 1990. Optimal brain damage. In Advances in neural information processing systems. 598--605."},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553453"},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.79"},{"key":"e_1_2_2_35_1","unstructured":"Chunyuan Li Heerad Farkhoor Rosanne Liu and Jason Yosinski. 2018. Measuring the intrinsic dimension of objective landscapes. arXiv preprint arXiv:1804.08838 (2018).  Chunyuan Li Heerad Farkhoor Rosanne Liu and Jason Yosinski. 2018. Measuring the intrinsic dimension of objective landscapes. arXiv preprint arXiv:1804.08838 (2018)."},{"key":"e_1_2_2_36_1","doi-asserted-by":"crossref","unstructured":"Chi Li M Zeeshan Zia Quoc-Huy Tran Xiang Yu Gregory D Hager and Manmohan Chandraker. 2016. Deep supervision with shape concepts for occlusion-aware 3d object parsing. arXiv preprint arXiv:1612.02699 (2016).  Chi Li M Zeeshan Zia Quoc-Huy Tran Xiang Yu Gregory D Hager and Manmohan Chandraker. 2016. Deep supervision with shape concepts for occlusion-aware 3d object parsing. arXiv preprint arXiv:1612.02699 (2016).","DOI":"10.1109\/CVPR.2017.49"},{"key":"e_1_2_2_37_1","volume-title":"Mining Intelligence and Knowledge Exploration","author":"Li Yongmou"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-39601-9_4"},{"key":"e_1_2_2_39_1","first-page":"2579","article-title":"Visualizing data using t-SNE","author":"van der Maaten Laurens","year":"2008","journal-title":"Journal of machine learning research 9"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3195258.3195260"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.3390\/app7101101"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_32"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASL.2011.2109382"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/2971763.2971764"},{"key":"e_1_2_2_45_1","unstructured":"Ari S Morcos David GT Barrett Neil C Rabinowitz and Matthew Botvinick. 2018. On the importance of single directions for generalization. arXiv preprint arXiv:1803.06959 (2018).  Ari S Morcos David GT Barrett Neil C Rabinowitz and Matthew Botvinick. 2018. On the importance of single directions for generalization. arXiv preprint arXiv:1803.06959 (2018)."},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.5555\/3104322.3104425"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46466-4_5"},{"key":"e_1_2_2_48_1","unstructured":"Jeeheh Oh Jiaxuan Wang and Jenna Wiens. 2018. Learning to Exploit Invariances in Clinical Time-Series Data using Sequence Transformer Networks. arXiv preprint arXiv:1808.06725 (2018).  Jeeheh Oh Jiaxuan Wang and Jenna Wiens. 2018. Learning to Exploit Invariances in Clinical Time-Series Data using Sequence Transformer Networks. arXiv preprint arXiv:1808.06725 (2018)."},{"key":"e_1_2_2_49_1","doi-asserted-by":"crossref","unstructured":"Chris Olah Arvind Satyanarayan Ian Johnson Shan Carter Ludwig Schubert Katherine Ye and Alexander Mordvintsev. 2018. The Building Blocks of Interpretability. Distill (2018). https:\/\/doi.org\/undefined https:\/\/distill.pub\/2018\/building-blocks.  Chris Olah Arvind Satyanarayan Ian Johnson Shan Carter Ludwig Schubert Katherine Ye and Alexander Mordvintsev. 2018. The Building Blocks of Interpretability. Distill (2018). https:\/\/doi.org\/undefined https:\/\/distill.pub\/2018\/building-blocks.","DOI":"10.23915\/distill.00010"},{"key":"e_1_2_2_50_1","unstructured":"Avital Oliver Augustus Odena Colin Raffel Ekin D Cubuk and Ian J Goodfellow. 2018. Realistic Evaluation of Deep Semi-Supervised Learning Algorithms. (2018).   Avital Oliver Augustus Odena Colin Raffel Ekin D Cubuk and Ian J Goodfellow. 2018. Realistic Evaluation of Deep Semi-Supervised Learning Algorithms. (2018)."},{"key":"e_1_2_2_51_1","doi-asserted-by":"crossref","unstructured":"Andrew Owens and Alexei A Efros. 2018. Audio-visual scene analysis with self-supervised multisensory features. arXiv preprint arXiv:1804.03641 (2018).  Andrew Owens and Alexei A Efros. 2018. Audio-visual scene analysis with self-supervised multisensory features. arXiv preprint arXiv:1804.03641 (2018).","DOI":"10.1007\/978-3-030-01231-1_39"},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-46448-0_48"},{"key":"e_1_2_2_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2009.191"},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.5555\/3305890.3305968"},{"key":"e_1_2_2_55_1","volume-title":"IJCAI Proceedings-International Joint Conference on Artificial Intelligence","volume":"22","author":"Pl\u00f6tz Thomas","year":"2011"},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3161174"},{"key":"e_1_2_2_57_1","volume-title":"Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability. In Advances in Neural Information Processing Systems. 6076--6085.","author":"Raghu Maithra","year":"2017"},{"key":"e_1_2_2_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/1273496.1273592"},{"key":"e_1_2_2_59_1","volume-title":"Machine Learning for Healthcare Conference. 73--100","author":"Razavian Narges","year":"2016"},{"key":"e_1_2_2_60_1","doi-asserted-by":"publisher","DOI":"10.3390\/s18092967"},{"key":"e_1_2_2_61_1","unstructured":"Aaqib Saeed and Stojan Trajanovski. 2017. Personalized Driver Stress Detection with Multi-task Neural Networks using Physiological Signals. arXiv preprint arXiv:1711.06116 (2017).  Aaqib Saeed and Stojan Trajanovski. 2017. Personalized Driver Stress Detection with Multi-task Neural Networks using Physiological Signals. arXiv preprint arXiv:1711.06116 (2017)."},{"key":"e_1_2_2_62_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW.2014.131"},{"key":"e_1_2_2_63_1","unstructured":"Karen Simonyan Andrea Vedaldi and Andrew Zisserman. 2013. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034 (2013).  Karen Simonyan Andrea Vedaldi and Andrew Zisserman. 2013. Deep inside convolutional networks: Visualising image classification models and saliency maps. arXiv preprint arXiv:1312.6034 (2013)."},{"key":"e_1_2_2_64_1","doi-asserted-by":"publisher","DOI":"10.1145\/2809695.2809718"},{"key":"e_1_2_2_65_1","unstructured":"Ilya Sutskever Oriol Vinyals and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Advances in neural information processing systems. 3104--3112.   Ilya Sutskever Oriol Vinyals and Quoc V Le. 2014. Sequence to sequence learning with neural networks. In Advances in neural information processing systems. 3104--3112."},{"key":"e_1_2_2_66_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2014.220"},{"key":"e_1_2_2_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/3136755.3136817"},{"key":"e_1_2_2_68_1","doi-asserted-by":"crossref","unstructured":"Jindong Wang Yiqiang Chen Shuji Hao Xiaohui Peng and Lisha Hu. 2018. Deep learning for sensor-based activity recognition: A survey. Pattern Recognition Letters (2018).  Jindong Wang Yiqiang Chen Shuji Hao Xiaohui Peng and Lisha Hu. 2018. Deep learning for sensor-based activity recognition: A survey. Pattern Recognition Letters (2018).","DOI":"10.1016\/j.patrec.2018.02.010"},{"key":"e_1_2_2_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3265689.3265705"},{"key":"e_1_2_2_70_1","volume-title":"2015 Federated Conference on Computer Science and Information Systems (FedCSIS). 411--416","author":"Wawrzyniak S."},{"key":"e_1_2_2_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00840"},{"key":"e_1_2_2_72_1","first-page":"3995","article-title":"Deep Convolutional Neural Networks on Multichannel Time Series for Human Activity Recognition","volume":"15","author":"Yang Jianbo","year":"2015","journal-title":"Ijcai"},{"key":"e_1_2_2_73_1","doi-asserted-by":"publisher","DOI":"10.1145\/3264954"},{"key":"e_1_2_2_74_1","first-page":"5","article-title":"Split-brain autoencoders: Unsupervised learning by cross-channel prediction","volume":"1","author":"Zhang Richard","year":"2017","journal-title":"CVPR"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3328932","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3328932","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:54:41Z","timestamp":1750204481000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3328932"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,6,21]]},"references-count":74,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2019,6,21]]}},"alternative-id":["10.1145\/3328932"],"URL":"https:\/\/doi.org\/10.1145\/3328932","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,6,21]]},"assertion":[{"value":"2019-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-04-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-06-21","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}