{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,27]],"date-time":"2026-07-27T22:08:48Z","timestamp":1785190128892,"version":"3.55.0"},"reference-count":80,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2023,9,27]],"date-time":"2023-09-27T00:00:00Z","timestamp":1695772800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2023,9,27]]},"abstract":"<jats:p>Understanding social interactions is relevant across many domains and applications, including psychology, behavioral sciences, human computer interaction, and healthcare. In this paper, we present a practical approach for automatically detecting face-to-face conversations by leveraging the acoustic sensing capabilities of an off-the-shelf, unmodified smartwatch. Our proposed framework incorporates feature representations extracted from different neural network setups and shows the benefits of feature fusion. The framework does not require an acoustic model specifically trained to the speech of the individual wearing the watch or of those nearby. We evaluate our framework with 39 participants in 18 homes in a semi-naturalistic study and with four participants in free living, obtaining an F1 score of 83.2% and 83.3% respectively for detecting user's conversations with the watch. Additionally, we study the real-time capability of our framework by deploying a system on an actual smartwatch and discuss several strategies to improve its practicality in real life. To support further work in this area by the research community, we also release our annotated dataset of conversations.<\/jats:p>","DOI":"10.1145\/3610882","type":"journal-article","created":{"date-parts":[[2023,9,27]],"date-time":"2023-09-27T15:45:03Z","timestamp":1695829503000},"page":"1-29","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":11,"title":["Automated Face-To-Face Conversation Detection on a Commodity Smartwatch with Acoustic Sensing"],"prefix":"10.1145","volume":"7","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7641-4299","authenticated-orcid":false,"given":"Dawei","family":"Liang","sequence":"first","affiliation":[{"name":"The University of Texas at Austin, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-0148-095X","authenticated-orcid":false,"given":"Alice","family":"Zhang","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8188-0457","authenticated-orcid":false,"given":"Edison","family":"Thomaz","sequence":"additional","affiliation":[{"name":"The University of Texas at Austin, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,9,27]]},"reference":[{"key":"e_1_2_1_1_1","unstructured":"Mart\u00edn Abadi Ashish Agarwal Paul Barham Eugene Brevdo Zhifeng Chen Craig Citro Greg S. Corrado Andy Davis Jeffrey Dean Matthieu Devin Sanjay Ghemawat Ian Goodfellow Andrew Harp Geoffrey Irving Michael Isard Yangqing Jia Rafal Jozefowicz Lukasz Kaiser Manjunath Kudlur Josh Levenberg Dandelion Man\u00e9 Rajat Monga Sherry Moore Derek Murray Chris Olah Mike Schuster Jonathon Shlens Benoit Steiner Ilya Sutskever Kunal Talwar Paul Tucker Vincent Vanhoucke Vijay Vasudevan Fernanda Vi\u00e9gas Oriol Vinyals Pete Warden Martin Wattenberg Martin Wicke Yuan Yu and Xiaoqiang Zheng. 2015. TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems. https:\/\/www.tensorflow.org\/ Software available from tensorflow.org."},{"key":"e_1_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/3448090"},{"key":"e_1_2_1_3_1","volume-title":"Deep learning using rectified linear units (relu). arXiv preprint arXiv:1803.08375","author":"Agarap Abien Fred","year":"2018","unstructured":"Abien Fred Agarap. 2018. Deep learning using rectified linear units (relu). arXiv preprint arXiv:1803.08375 (2018)."},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-15582-1_5"},{"key":"e_1_2_1_5_1","unstructured":"Apple. 2023. AirPods redefine the personal audio experience. https:\/\/www.apple.com\/newsroom\/2023\/06\/airpods-redefine-the-personal-audio-experience\/. Accessed: 2023-07-13."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1145\/3191734"},{"key":"e_1_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/3432210"},{"key":"e_1_2_1_8_1","volume-title":"Conversational scene analysis. Ph. D. Dissertation","author":"Basu Sumit","unstructured":"Sumit Basu. 2002. Conversational scene analysis. Ph. D. Dissertation. Massachusetts Institute of Technology."},{"key":"e_1_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1086\/260265"},{"key":"e_1_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3341163.3347735"},{"key":"e_1_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMC.2015.2465376"},{"key":"e_1_2_1_12_1","volume-title":"Sensing and modeling human networks. Ph. D. Dissertation","author":"Choudhury Tanzeem Khalid","unstructured":"Tanzeem Khalid Choudhury. 2004. Sensing and modeling human networks. Ph. D. Dissertation. Massachusetts Institute of Technology."},{"key":"e_1_2_1_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/3191736"},{"key":"e_1_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0218429"},{"key":"e_1_2_1_15_1","volume-title":"One kind of speech act: How do we know when we're conversing?","author":"Donaldson Susan Kay","year":"1979","unstructured":"Susan Kay Donaldson. 1979. One kind of speech act: How do we know when we're conversing? (1979)."},{"key":"e_1_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.3390\/s18082474"},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/3211960.3211975"},{"key":"e_1_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2017.7952261"},{"key":"e_1_2_1_19_1","unstructured":"Google. 2020. Sound classification with YAMNet. https:\/\/www.tensorflow.org\/hub\/tutorials\/yamnet. Accessed: 2022-08-01."},{"key":"e_1_2_1_20_1","unstructured":"Hall and Watson. 1970. NASA Exercise: Survival on the Moon. https:\/\/www.humber.ca\/centreforteachingandlearning\/assets\/files\/pdfs\/MoonExercise.pdf. Accessed: 2020-10-01."},{"key":"e_1_2_1_21_1","doi-asserted-by":"crossref","unstructured":"Judith Holler Kobin H Kendrick Marisa Casillas and Stephen C Levinson. 2016. Turn-taking in human communicative interaction. Frontiers Media SA.","DOI":"10.3389\/978-2-88919-825-2"},{"key":"e_1_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2750858.2807526"},{"key":"e_1_2_1_23_1","volume-title":"Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861","author":"Howard Andrew G","year":"2017","unstructured":"Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017)."},{"key":"e_1_2_1_24_1","volume-title":"Workshops at the Twenty-Sixth AAAI Conference on Artificial Intelligence.","author":"Hsiao Joey","year":"2012","unstructured":"Joey Chiao-yin Hsiao, Wan-rong Jih, and Jane Yung-jen Hsu. 2012. Recognizing continuous social engagement level in dyadic conversation by using turn-taking and speech emotion patterns. In Workshops at the Twenty-Sixth AAAI Conference on Artificial Intelligence."},{"key":"e_1_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2668332.2668338"},{"key":"e_1_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3383652.3423907"},{"key":"e_1_2_1_27_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00286"},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/3362743.3362959"},{"key":"e_1_2_1_29_1","volume-title":"Multisensor data fusion: A review of the state-of-the-art. Information fusion 14, 1","author":"Khaleghi Bahador","year":"2013","unstructured":"Bahador Khaleghi, Alaa Khamis, Fakhreddine O Karray, and Saiedeh N Razavi. 2013. Multisensor data fusion: A review of the state-of-the-art. Information fusion 14, 1 (2013), 28--44."},{"key":"e_1_2_1_30_1","volume-title":"Meeting mediator: enhancing group collaborationusing sociometric feedback. In Proceedings of the 2008 ACM conference on Computer supported cooperative work. 457--466","author":"Kim Taemie","year":"2008","unstructured":"Taemie Kim, Agnes Chang, Lindsey Holland, and Alex Sandy Pentland. 2008. Meeting mediator: enhancing group collaborationusing sociometric feedback. In Proceedings of the 2008 ACM conference on Computer supported cooperative work. 457--466."},{"key":"e_1_2_1_31_1","volume-title":"Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980","author":"Kingma Diederik P","year":"2014","unstructured":"Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980 (2014)."},{"key":"e_1_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-01516-8_13"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASLP.2020.3030497"},{"key":"e_1_2_1_34_1","volume-title":"The measurement of observer agreement for categorical data. biometrics","author":"Richard Landis J","year":"1977","unstructured":"J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics (1977), 159--174."},{"key":"e_1_2_1_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2750858.2804262"},{"key":"e_1_2_1_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/3242587.3242609"},{"key":"e_1_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3290605.3300568"},{"key":"e_1_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2007.sigdial-1.33"},{"key":"e_1_2_1_39_1","doi-asserted-by":"crossref","unstructured":"Youngki Lee Chulhong Min Chanyou Hwang Jaeung Lee Inseok Hwang Younghyun Ju Chungkuk Yoo Miri Moon Uichin Lee and Junehwa Song. 2013. Sociophone: Everyday face-to-face interaction monitoring platform using multi-phone sensor fusion. In Proceeding of the 11th annual international conference on Mobile systems applications and services. 375--388.","DOI":"10.1145\/2462456.2465426"},{"key":"e_1_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1109\/THMS.2015.2401391"},{"key":"e_1_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/BSN.2013.6575509"},{"key":"e_1_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1145\/3544794.3558471"},{"key":"e_1_2_1_43_1","volume-title":"Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study. arXiv preprint arXiv:2110.03174","author":"Liang Dawei","year":"2021","unstructured":"Dawei Liang, Yangyang Shi, Yun Wang, Nayan Singhal, Alex Xiao, Jonathan Shaw, Edison Thomaz, Ozlem Kalinli, and Mike Seltzer. 2021. Transferring Voice Knowledge for Acoustic Event Detection: An Empirical Study. arXiv preprint arXiv:2110.03174 (2021)."},{"key":"e_1_2_1_44_1","volume-title":"22nd International Conference on Human-Computer Interaction with Mobile Devices and Services. 1--10.","author":"Liang Dawei","unstructured":"Dawei Liang, Wenting Song, and Edison Thomaz. 2020. Characterizing the Effect of Audio Degradation on Privacy Perception And Inference Performance in Audio-Based Human Activity Recognition. In 22nd International Conference on Human-Computer Interaction with Mobile Devices and Services. 1--10."},{"key":"e_1_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1145\/3314404"},{"key":"e_1_2_1_46_1","doi-asserted-by":"crossref","unstructured":"Bethany Little Ossama Alshabrawy Daniel Stow I Nicol Ferrier Roisin McNaney Daniel G Jackson Karim Ladha Cassim Ladha Thomas Ploetz Jaume Bacardit et al. 2020. Deep learning-based automated speech detection as a marker of social functioning in late-life depression. Psychological Medicine (2020) 1--10.","DOI":"10.1017\/S0033291719003994"},{"key":"e_1_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/ISPA\/IUCC.2017.00148"},{"key":"e_1_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-642-21726-5_12"},{"key":"e_1_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1145\/2370216.2370270"},{"key":"e_1_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/1555816.1555834"},{"key":"e_1_2_1_51_1","doi-asserted-by":"publisher","DOI":"10.1145\/1869983.1869992"},{"key":"e_1_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/2517351.2517353"},{"key":"e_1_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.4108\/icst.pervasivehealth.2012.248689"},{"key":"e_1_2_1_54_1","volume-title":"The Electronically Activated Recorder (EAR): A device for sampling naturalistic daily activities and conversations. Behavior research methods, instruments, & computers 33, 4","author":"Mehl Matthias R","year":"2001","unstructured":"Matthias R Mehl, James W Pennebaker, D Michael Crow, James Dabbs, and John H Price. 2001. The Electronically Activated Recorder (EAR): A device for sampling naturalistic daily activities and conversations. Behavior research methods, instruments, & computers 33, 4 (2001), 517--523."},{"key":"e_1_2_1_55_1","doi-asserted-by":"publisher","DOI":"10.1177\/0956797610362675"},{"key":"e_1_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2019.8683244"},{"key":"e_1_2_1_57_1","doi-asserted-by":"publisher","DOI":"10.1109\/BSN.2019.8771081"},{"key":"e_1_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICC.2015.7248384"},{"key":"e_1_2_1_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/2030112.2030164"},{"key":"e_1_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1145\/1864349.1864393"},{"key":"e_1_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1145\/2077546.2077557"},{"key":"e_1_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1145\/1978942.1978945"},{"key":"e_1_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1177\/0963721414560811"},{"key":"e_1_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2018.2800728"},{"key":"e_1_2_1_65_1","volume-title":"Interpersonal communication: Lifeblood of an organization. IUP Journal of Soft Skills 3","author":"Sethi Deepa","year":"2009","unstructured":"Deepa Sethi and Manisha Seth. 2009. Interpersonal communication: Lifeblood of an organization. IUP Journal of Soft Skills 3 (2009)."},{"key":"e_1_2_1_66_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0023176"},{"key":"e_1_2_1_67_1","volume-title":"2011 Proceedings IEEE INFOCOM. IEEE, 2291--2299","author":"Tang Shaojie","year":"2011","unstructured":"Shaojie Tang, Jing Yuan, Xufei Mao, Xiang-Yang Li, Wei Chen, and Guojun Dai. 2011. Relationship classification in large scale online social networks and its impact on information propagation. In 2011 Proceedings IEEE INFOCOM. IEEE, 2291--2299."},{"key":"e_1_2_1_68_1","doi-asserted-by":"publisher","DOI":"10.1145\/2750858.2807545"},{"key":"e_1_2_1_69_1","doi-asserted-by":"publisher","DOI":"10.1145\/3313831.3376855"},{"key":"e_1_2_1_70_1","volume-title":"Features of naturalness in conversation","author":"Warren Martin","unstructured":"Martin Warren. 2006. Features of naturalness in conversation. Vol. 152. John Benjamins Publishing."},{"key":"e_1_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1080\/08351818709389286"},{"key":"e_1_2_1_72_1","doi-asserted-by":"crossref","unstructured":"Danny Wyatt Tanzeem Choudhury and Jeff A Bilmes. 2007. Conversation detection and speaker segmentation in privacy-sensitive situated speech data.. In Interspeech. 586--589.","DOI":"10.21437\/Interspeech.2007-256"},{"key":"e_1_2_1_73_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2007.367201"},{"key":"e_1_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1145\/2493432.2493435"},{"key":"e_1_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1145\/2494091.2494118"},{"key":"e_1_2_1_76_1","volume-title":"Social interaction and information diffusion in Social Internet of Things: dynamics, cloud-edge, traceability","author":"Yi Yinxue","year":"2020","unstructured":"Yinxue Yi, Zufan Zhang, Laurence T Yang, Xianjun Deng, Lingzhi Yi, and Xiaokang Wang. 2020. Social interaction and information diffusion in Social Internet of Things: dynamics, cloud-edge, traceability. IEEE Internet of things Journal 8, 4 (2020), 2177--2192."},{"key":"e_1_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2017-65"},{"key":"e_1_2_1_78_1","volume-title":"Vishnu Vidyadhara Raju Vegesna, and Anil Kumar Vuppala","author":"Zarazaga Pablo P\u00e9rez","year":"2019","unstructured":"Pablo P\u00e9rez Zarazaga, Sneha Das, Tom B\u00e4ckstr\u00f6m, Vishnu Vidyadhara Raju Vegesna, and Anil Kumar Vuppala. 2019. Sound Privacy: A Conversational Speech Corpus for Quantifying the Experience of Privacy.. In Interspeech. 3720--3724."},{"key":"e_1_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/2973750.2973775"},{"key":"e_1_2_1_80_1","volume-title":"To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878","author":"Zhu Michael","year":"2017","unstructured":"Michael Zhu and Suyog Gupta. 2017. To prune, or not to prune: exploring the efficacy of pruning for model compression. arXiv preprint arXiv:1710.01878 (2017)."}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3610882","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3610882","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,7,28]],"date-time":"2025-07-28T16:25:33Z","timestamp":1753719933000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3610882"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,9,27]]},"references-count":80,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,9,27]]}},"alternative-id":["10.1145\/3610882"],"URL":"https:\/\/doi.org\/10.1145\/3610882","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,9,27]]},"assertion":[{"value":"2023-09-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}