{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,16]],"date-time":"2026-07-16T14:24:29Z","timestamp":1784211869214,"version":"3.55.0"},"reference-count":44,"publisher":"Association for Computing Machinery (ACM)","issue":"1","license":[{"start":{"date-parts":[[2024,3,6]],"date-time":"2024-03-06T00:00:00Z","timestamp":1709683200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100006374","name":"National Science Foundation","doi-asserted-by":"publisher","award":["SaTC-1801472"],"award-info":[{"award-number":["SaTC-1801472"]}],"id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100006374","name":"JPMorgan Chase and Company","doi-asserted-by":"publisher","id":[{"id":"10.13039\/501100006374","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2024,3,6]]},"abstract":"<jats:p>Audio-based human activity recognition (HAR) is very popular because many human activities have unique sound signatures that can be detected using machine learning (ML) approaches. These audio-based ML HAR pipelines often use common featurization techniques, such as extracting various statistical and spectral features by converting time domain signals to the frequency domain (using an FFT) and using them to train ML models. Some of these approaches also claim privacy benefits by preventing the identification of human speech. However, recent deep learning-based automatic speech recognition (ASR) models pose new privacy challenges to these featurization techniques. In this paper, we systematically evaluate various featurization approaches for audio data, assessing their privacy risks through metrics like speech intelligibility (PER and WER) while considering the utility tradeoff in terms of ML-based activity recognition accuracy. Our findings reveal the susceptibility of these approaches to speech content recovery when exposed to recent ASR models, especially under re-tuning or retraining conditions. Notably, fine-tuned ASR models achieved an average Phoneme Error Rate (PER) of 39.99% and Word Error Rate (WER) of 44.43% in speech recognition for these approaches. To overcome these privacy concerns, we propose Kirigami, a lightweight machine learning-based audio speech filter that removes human speech segments reducing the efficacy of ASR models (70.48% PER and 101.40% WER) while also maintaining HAR accuracy (76.0% accuracy). We show that Kirigami can be implemented on common edge microcontrollers with limited computational capabilities and memory, providing a path to deployment on a variety of IoT devices. Finally, we conducted a real-world user study and showed the robustness of Kirigami on a laptop and an ARM Cortex-M4F microcontroller under three different background noises.<\/jats:p>","DOI":"10.1145\/3643502","type":"journal-article","created":{"date-parts":[[2024,3,6]],"date-time":"2024-03-06T13:12:36Z","timestamp":1709730756000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":7,"title":["Kirigami"],"prefix":"10.1145","volume":"8","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3743-1384","authenticated-orcid":false,"given":"Sudershan","family":"Boovaraghavan","sequence":"first","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-7315-6033","authenticated-orcid":false,"given":"Haozhe","family":"Zhou","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1237-7545","authenticated-orcid":false,"given":"Mayank","family":"Goel","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9304-6080","authenticated-orcid":false,"given":"Yuvraj","family":"Agarwal","sequence":"additional","affiliation":[{"name":"Carnegie Mellon University, Pittsburgh, United States"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,3,6]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Alexei Baevski Yuhao Zhou Abdelrahman Mohamed and Michael Auli. 2020. Wav2vec 2.0: a framework for self-supervised learning of speech representations. Advances in neural information processing systems 33 12449--12460."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.comnet.2020.107447"},{"key":"e_1_2_1_3_1","first-page":"36","article-title":"How many phonemes does the english language have","volume":"5","author":"Bizzocchi Aldo Luiz","year":"2017","unstructured":"Aldo Luiz Bizzocchi. 2017. How many phonemes does the english language have? International Journal on Studies in English Language and Literature (IJSELL), 5, 10, 36--46.","journal-title":"International Journal on Studies in English Language and Literature (IJSELL)"},{"key":"e_1_2_1_4_1","first-page":"11","article-title":"The multiple dimensions of privacy: testing lay expectations of privacy","author":"Blumenthal Jeremy A","year":"2008","unstructured":"Jeremy A Blumenthal, Meera Adya, and Jacqueline Mogle. 2008. The multiple dimensions of privacy: testing lay expectations of privacy. U. Pa. J. Const. L., 11, 331.","journal-title":"U. Pa. J. Const."},{"key":"e_1_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3580865"},{"key":"e_1_2_1_6_1","volume-title":"Connectionist speech recognition: a hybrid approach","author":"Bourlard Herve A","unstructured":"Herve A Bourlard and Nelson Morgan. 1994. Connectionist speech recognition: a hybrid approach. Vol. 247. Springer Science & Business Media."},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1459359.1459472"},{"key":"e_1_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3517258"},{"key":"e_1_2_1_9_1","unstructured":"Jacob Devlin Ming-Wei Chang Kenton Lee and Kristina Toutanova. 2018. Bert: pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805."},{"key":"e_1_2_1_10_1","unstructured":"The CMU Pronouncing Dictionary. [n. d.] http:\/\/www.speech.cs.cmu.edu\/cgi-bin\/cmudict."},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIT.1983.1056650"},{"key":"e_1_2_1_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/3478091"},{"key":"e_1_2_1_13_1","unstructured":"Zhiyun Fan Meng Li Shiyu Zhou and Bo Xu. 2020. Exploring wav2vec 2.0 on speaker verification and language identification. arXiv preprint arXiv:2012.06185."},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11023-020-09548-1"},{"key":"e_1_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.6028\/NIST.IR.4930"},{"key":"e_1_2_1_16_1","doi-asserted-by":"crossref","unstructured":"Yuan Gong Yu-An Chung and James Glass. 2021. Ast: audio spectrogram transformer. arXiv preprint arXiv:2104.01778.","DOI":"10.21437\/Interspeech.2021-698"},{"key":"e_1_2_1_17_1","doi-asserted-by":"crossref","unstructured":"Alex Graves. 2012. Sequence transduction with recurrent neural networks. arXiv preprint arXiv:1211.3711.","DOI":"10.1007\/978-3-642-24797-2"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1143844.1143891"},{"key":"e_1_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP39728.2021.9413668"},{"key":"e_1_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1145\/3411764.3445169"},{"key":"e_1_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.comnet.2018.09.003"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISM.2012.12"},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2699343.2699366"},{"key":"e_1_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1145\/3242587.3242609"},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/3025453.3025773"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/2030112.2030163"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.2478\/popets-2019-0068"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3376897.3379158"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3550284"},{"key":"e_1_2_1_30_1","volume-title":"Better English Pronunciation","author":"O'Connor Joseph D","unstructured":"Joseph D O'Connor. 1980. Better English Pronunciation. Cambridge University Press."},{"key":"e_1_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11265-022-01748-5"},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2015.7178964"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/3173574.3173867"},{"key":"e_1_2_1_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2806390"},{"key":"e_1_2_1_35_1","unstructured":"Pincelate. [n. d.] https:\/\/github.com\/aparrish\/pincelate\/."},{"key":"e_1_2_1_36_1","volume-title":"Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever.","author":"Radford Alec","year":"2022","unstructured":"Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022. Robust speech recognition via large-scale weak supervision. arXiv preprint arXiv:2212.04356."},{"key":"e_1_2_1_37_1","unstructured":"Mirco Ravanelli et al. 2021. Speechbrain: a general-purpose speech toolkit. arXiv preprint arXiv:2106.04624."},{"key":"e_1_2_1_38_1","first-page":"1816","article-title":"A scalable noisy speech dataset and online subjective test framework","volume":"2019","author":"Reddy Chandan KA","year":"2019","unstructured":"Chandan KA Reddy, Ebrahim Beyrami, Jamie Pool, Ross Cutler, Sriram Srinivasan, and Johannes Gehrke. 2019. A scalable noisy speech dataset and online subjective test framework. Proc. Interspeech 2019, 1816--1820.","journal-title":"Proc. Interspeech"},{"key":"e_1_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3264945"},{"key":"e_1_2_1_40_1","unstructured":"The Verge. 2023. A new hack can turn an Echo into a live microphone. https:\/\/www.theverge.com\/2017\/8\/1\/16079044\/amazon-echo-hack-microphone-listen-in-mark-barnes. (2023)."},{"key":"e_1_2_1_41_1","unstructured":"Yingzhi Wang Abdelmoumene Boumadane and Abdelwahab Heba. 2021. A fine-tuned wav2vec 2.0\/hubert benchmark for speech emotion recognition speaker verification and spoken language understanding. arXiv preprint arXiv:2111.02735."},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448124"},{"key":"e_1_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448078"},{"key":"e_1_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ASRU51503.2021.9688232"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643502","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3643502","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,8,22]],"date-time":"2025-08-22T13:02:23Z","timestamp":1755867743000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3643502"}},"subtitle":["Lightweight Speech Filtering for Privacy-Preserving Activity Recognition using Audio"],"short-title":[],"issued":{"date-parts":[[2024,3,6]]},"references-count":44,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2024,3,6]]}},"alternative-id":["10.1145\/3643502"],"URL":"https:\/\/doi.org\/10.1145\/3643502","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,3,6]]},"assertion":[{"value":"2024-03-06","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}