{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,12]],"date-time":"2025-10-12T02:51:26Z","timestamp":1760237486797,"version":"build-2065373602"},"reference-count":38,"publisher":"MDPI AG","issue":"10","license":[{"start":{"date-parts":[[2020,5,22]],"date-time":"2020-05-22T00:00:00Z","timestamp":1590105600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100001691","name":"Japan Society for the Promotion of Science","doi-asserted-by":"publisher","award":["JP15H01771","JP17H01753"],"award-info":[{"award-number":["JP15H01771","JP17H01753"]}],"id":[{"id":"10.13039\/501100001691","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100002241","name":"Japan Science and Technology Agency","doi-asserted-by":"publisher","award":["JPMJCE1309"],"award-info":[{"award-number":["JPMJCE1309"]}],"id":[{"id":"10.13039\/501100002241","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Speech discrimination that determines whether a participant is speaking at a given moment is essential in investigating human verbal communication. Specifically, in dynamic real-world situations where multiple people participate in, and form, groups in the same space, simultaneous speakers render speech discrimination that is solely based on audio sensing difficult. In this study, we focused on physical activity during speech, and hypothesized that combining audio and physical motion data acquired by wearable sensors can improve speech discrimination. Thus, utterance and physical activity data of students in a university participatory class were recorded, using smartphones worn around their neck. First, we tested the temporal relationship between manually identified utterances and physical motions and confirmed that physical activities in wide-frequency ranges co-occurred with utterances. Second, we trained and tested classifiers for each participant and found a higher performance with the audio-motion classifier (average accuracy 92.2%) than both the audio-only (80.4%) and motion-only (87.8%) classifiers. Finally, we tested inter-individual classification and obtained a higher performance with the audio-motion combined classifier (83.2%) than the audio-only (67.7%) and motion-only (71.9%) classifiers. These results show that audio-motion multimodal sensing using widely available smartphones can provide effective utterance discrimination in dynamic group communications.<\/jats:p>","DOI":"10.3390\/s20102948","type":"journal-article","created":{"date-parts":[[2020,5,22]],"date-time":"2020-05-22T10:18:18Z","timestamp":1590142698000},"page":"2948","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["Speech Discrimination in Real-World Group Communication Using Audio-Motion Multimodal Sensing"],"prefix":"10.3390","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6300-4373","authenticated-orcid":false,"given":"Takayuki","family":"Nozawa","sequence":"first","affiliation":[{"name":"Research Institute for the Earth Inclusive Sensing, Tokyo Institute of Technology, Tokyo 152-8550, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mizuki","family":"Uchiyama","sequence":"additional","affiliation":[{"name":"School of Computing, Tokyo Institute of Technology, Yokohama 226-8502, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Keigo","family":"Honda","sequence":"additional","affiliation":[{"name":"School of Computing, Tokyo Institute of Technology, Yokohama 226-8502, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Tamio","family":"Nakano","sequence":"additional","affiliation":[{"name":"Institute for Liberal Arts, Tokyo Institute of Technology, Tokyo 152-8550, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yoshihiro","family":"Miyake","sequence":"additional","affiliation":[{"name":"Department of Computer Science, Tokyo Institute of Technology, Yokohama 226-8502, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2020,5,22]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Sato, N., Tsuji, S., Yano, K., Otsuka, R., Moriwaki, N., Ara, K., Wakisaka, Y., Ohkubo, N., Hayakawa, M., and Horry, Y. (2009, January 17\u201319). Knowledge-creating behavior index for improving knowledge workers\u2019 productivity. Proceedings of the 2009 Sixth International Conference on Networked Sensing Systems (INSS), Pittsburgh, PA, USA.","DOI":"10.1109\/INSS.2009.5409923"},{"key":"ref_2","unstructured":"Olguin, D.O., Paradiso, J.A., and Pentland, A. (2006, January 11\u201314). Wearable Communicator Badge: Designing a New Platform for Revealing Organizational Dynamics. Proceedings of the 10th International Symposium on Wearable Computers (student colloquium), Montreux, Switzerland."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Wu, L., Waber, B.N., Aral, S., Brynjolfsson, E., and Pentland, A. (2008). Mining face-to-face interaction networks using sociometric badges: Predicting productivity in an IT configuration task. Available SSRN.","DOI":"10.2139\/ssrn.1130251"},{"key":"ref_4","first-page":"54","article-title":"SPPAS\u2014Multi-lingual approaches to the automatic annotation of speech","volume":"111","author":"Bigi","year":"2015","journal-title":"Phonetician"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1109\/MSP.2015.2462851","article-title":"Speaker recognition by machines and humans: A tutorial review","volume":"32","author":"Hansen","year":"2015","journal-title":"IEEE Signal Process. Mag."},{"key":"ref_6","first-page":"1","article-title":"Voice activity detection. Fundamentals and speech recognition system robustness","volume":"6","author":"Grimm","year":"2007","journal-title":"Robust Speech: Recognition and Understanding"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1032","DOI":"10.1109\/TMM.2014.2305632","article-title":"Simultaneous-speaker voice activity detection and localization using mid-fusion of SVM and HMMs","volume":"16","author":"Minotto","year":"2014","journal-title":"IEEE Trans. Multimed."},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Bertrand, A., and Moonen, M. (2010, January 14\u201319). Energy-based multi-speaker voice activity detection with an ad hoc microphone array. Proceedings of the 2010 IEEE International Conference on Acoustics, Speech and Signal Processing, Dallas, TX, USA.","DOI":"10.1109\/ICASSP.2010.5496183"},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"356","DOI":"10.1109\/TASL.2011.2125954","article-title":"Speaker diarization: A review of recent research","volume":"20","author":"Bozonnet","year":"2012","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"1557","DOI":"10.1109\/TASL.2006.878256","article-title":"An overview of automatic speaker diarization systems","volume":"14","author":"Tranter","year":"2006","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Boakye, K., Trueba-Hornero, B., Vinyals, O., and Friedland, G. (April, January 31). Overlapped speech detection for improved speaker diarization in multiparty meetings. Proceedings of the 2008 IEEE International Conference on Acoustics, Speech and Signal Processing, Las Vegas, NV, USA.","DOI":"10.1109\/ICASSP.2008.4518619"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"371","DOI":"10.1109\/TASL.2011.2158419","article-title":"The ICSI RT-09 speaker diarization system","volume":"20","author":"Friedland","year":"2011","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_13","doi-asserted-by":"crossref","unstructured":"Hung, H., and Ba, S.O. (2009). Speech\/non-speech Detection in Meetings from Automatically Extracted Low Resolution Visual Features, Idiap.","DOI":"10.1109\/ICASSP.2010.5494913"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Friedland, G., Hung, H., and Yeo, C. (2009, January 19\u201324). Multi-modal Speaker Diarization of Real-World Meetings Using Compressed-Domain Video Features. Proceedings of the 2009 IEEE International Conference on Acoustics, Speech and Signal Processing, Taipei, Taiwan.","DOI":"10.1109\/ICASSP.2009.4960522"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Garau, G., Dielmann, A., and Bourlard, H. (2010, January 26\u201330). Audio-visual synchronisation for speaker diarisation. Proceedings of the Eleventh Annual Conference of the International Speech Communication Association, Chiba, Japan.","DOI":"10.21437\/Interspeech.2010-704"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"79","DOI":"10.1109\/TPAMI.2011.47","article-title":"Multimodal speaker diarization","volume":"34","author":"Noulas","year":"2011","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1146\/annurev.anthro.26.1.109","article-title":"Gesture","volume":"26","author":"Kendon","year":"1997","journal-title":"Annu. Rev. Anthropol."},{"key":"ref_18","doi-asserted-by":"crossref","first-page":"98","DOI":"10.1037\/h0027035","article-title":"Body movement and speech rhythm in social conversation","volume":"11","author":"Dittmann","year":"1969","journal-title":"J. Personal. Soc. Psychol."},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1177\/009365085012002004","article-title":"Listeners\u2019 body movements and speaking turns","volume":"12","author":"Harrigan","year":"1985","journal-title":"Commun. Res."},{"key":"ref_20","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1111\/1467-8721.ep13175642","article-title":"Why do we gesture when we speak?","volume":"7","author":"Krauss","year":"1998","journal-title":"Curr. Dir. Psychol. Sci."},{"key":"ref_21","unstructured":"Vossen, D.L. (2012). Detecting Speaker Status in a Social Setting with a Single Triaxial Accelerometer. Unpublished. [Master\u2019s Thesis, Department of Artificial Intelligence, University of Amsterdam]."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"36","DOI":"10.1109\/MPRV.2011.79","article-title":"Sensing the \u201chealth state\u201d of a community","volume":"11","author":"Madan","year":"2012","journal-title":"IEEE Pervasive Comput."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Lee, Y., Min, C., Hwang, C., Lee, J., Hwang, I., Ju, Y., Yoo, C., Moon, M., Lee, U., and Song, J. (2013, January 25\u201328). Sociophone: Everyday face-to-face interaction monitoring platform using multi-phone sensor fusion. Proceedings of the 11th Annual International Conference on Mobile Systems, Applications, and Services, Taipei, Taiwan.","DOI":"10.1145\/2462456.2465426"},{"key":"ref_24","doi-asserted-by":"crossref","first-page":"74","DOI":"10.1145\/1964897.1964918","article-title":"Activity recognition using cell phone accelerometers","volume":"12","author":"Kwapisz","year":"2011","journal-title":"ACM SigKDD Explor. Newsl."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Anjum, A., and Ilyas, M.U. (2013, January 11\u201314). Activity Recognition Using Smartphone Sensors. Proceedings of the 2013 IEEE 10th Consumer Communications and Networking Conference (CCNC), Las Vegas, NV, USA.","DOI":"10.1109\/CCNC.2013.6488584"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"252","DOI":"10.1109\/LSP.2015.2495219","article-title":"Voice activity detection: Merging source and filter-based information","volume":"23","author":"Drugman","year":"2015","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_27","doi-asserted-by":"crossref","first-page":"239","DOI":"10.1093\/biomet\/33.3.239","article-title":"The treatment of ties in ranking problems","volume":"33","author":"Kendall","year":"1945","journal-title":"Biometrika"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Agresti, A. (2010). Analysis of Ordinal Categorical Data, Wiley. [2nd ed.].","DOI":"10.1002\/9780470594001"},{"key":"ref_29","first-page":"18","article-title":"Classification and regression by randomForest","volume":"2","author":"Liaw","year":"2002","journal-title":"R News"},{"key":"ref_30","unstructured":"Team, R.C. (2018). R: A Language and Environment for Statistical Computing, R Foundation for Statistical Computing."},{"key":"ref_31","doi-asserted-by":"crossref","first-page":"321","DOI":"10.1613\/jair.953","article-title":"SMOTE: Synthetic minority over-sampling technique","volume":"16","author":"Chawla","year":"2002","journal-title":"J. Artif. Intell. Res."},{"key":"ref_32","doi-asserted-by":"crossref","first-page":"1086","DOI":"10.1109\/TPAMI.2017.2648793","article-title":"Audio-visual speaker diarization based on spatiotemporal bayesian fusion","volume":"40","author":"Gebru","year":"2017","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Hung, H., Englebienne, G., and Kools, J. (2013, January 8\u201312). Classifying social actions with a single accelerometer. Proceedings of the 2013 ACM International Joint Conference on Pervasive and Ubiquitous Computing, Zurich, Switzerland.","DOI":"10.1145\/2493432.2493513"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/97.736233","article-title":"A statistical model-based voice activity detection","volume":"6","author":"Sohn","year":"1999","journal-title":"IEEE Signal Process. Lett."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"697","DOI":"10.1109\/TASL.2012.2229986","article-title":"Deep belief networks based voice activity detection","volume":"21","author":"Zhang","year":"2012","journal-title":"IEEE Trans. Audio Speech Lang. Process."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"9017","DOI":"10.1109\/ACCESS.2018.2800728","article-title":"A convolutional neural network smartphone app for real-time voice activity detection","volume":"6","author":"Sehgal","year":"2018","journal-title":"IEEE Access"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Martin, A., Charlet, D., and Mauuary, L. (2001, January 7\u201311). Robust speech\/non-speech detection using LDA applied to MFCC. Proceedings of the 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No. 01CH37221), Salt Lake City, UT, USA.","DOI":"10.21437\/Eurospeech.2001-269"},{"key":"ref_38","doi-asserted-by":"crossref","first-page":"197","DOI":"10.1109\/LSP.2013.2237903","article-title":"Unsupervised speech activity detection using voicing measures and perceptual spectral flux","volume":"20","author":"Sadjadi","year":"2013","journal-title":"IEEE Signal Process. Lett."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/10\/2948\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T09:31:36Z","timestamp":1760175096000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/20\/10\/2948"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,5,22]]},"references-count":38,"journal-issue":{"issue":"10","published-online":{"date-parts":[[2020,5]]}},"alternative-id":["s20102948"],"URL":"https:\/\/doi.org\/10.3390\/s20102948","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2020,5,22]]}}}