{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,4]],"date-time":"2026-06-04T22:14:14Z","timestamp":1780611254795,"version":"3.54.1"},"reference-count":41,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2018,9,18]],"date-time":"2018-09-18T00:00:00Z","timestamp":1537228800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"College of Engineering, Nanyang Technological University","award":["Seed Grant"],"award-info":[{"award-number":["Seed Grant"]}]},{"DOI":"10.13039\/501100006512","name":"Nanyang Technological University","doi-asserted-by":"publisher","award":["SUG"],"award-info":[{"award-number":["SUG"]}],"id":[{"id":"10.13039\/501100006512","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["Proc. ACM Interact. Mob. Wearable Ubiquitous Technol."],"published-print":{"date-parts":[[2018,9,18]]},"abstract":"<jats:p>Recent years have seen the increasing need of location awareness by mobile applications. This paper presents a room-level indoor localization approach based on the measured room's echos in response to a two-millisecond single-tone inaudible chirp emitted by a smartphone's loudspeaker. Different from other acoustics-based room recognition systems that record full-spectrum audio for up to ten seconds, our approach records audio in a narrow inaudible band for 0.1 seconds only to preserve the user's privacy. However, the short-time and narrowband audio signal carries limited information about the room's characteristics, presenting challenges to accurate room recognition. This paper applies deep learning to effectively capture the subtle fingerprints in the rooms' acoustic responses. Our extensive experiments show that a two-layer convolutional neural network fed with the spectrogram of the inaudible echos achieve the best performance, compared with alternative designs using other raw data formats and deep models. Based on this result, we design a RoomRecognize cloud service and its mobile client library that enable the mobile application developers to readily implement the room recognition functionality without resorting to any existing infrastructures and add-on hardware. Extensive evaluation shows that RoomRecognize achieves 99.7%, 97.7%, 99%, and 89% accuracy in differentiating 22 and 50 residential\/office rooms, 19 spots in a quiet museum, and 15 spots in a crowded museum, respectively. Compared with the state-of-the-art approaches based on support vector machine, RoomRecognize significantly improves the Pareto frontier of recognition accuracy versus robustness against interfering sounds (e.g., ambient music).<\/jats:p>","DOI":"10.1145\/3264945","type":"journal-article","created":{"date-parts":[[2018,9,19]],"date-time":"2018-09-19T11:58:41Z","timestamp":1537358321000},"page":"1-28","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":28,"title":["Deep Room Recognition Using Inaudible Echos"],"prefix":"10.1145","volume":"2","author":[{"given":"Qun","family":"Song","sequence":"first","affiliation":[{"name":"School of Computer Science and Engineering, Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Chaojie","family":"Gu","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Rui","family":"Tan","sequence":"additional","affiliation":[{"name":"School of Computer Science and Engineering, Nanyang Technological University, Singapore"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2018,9,18]]},"reference":[{"key":"e_1_2_2_1_1","unstructured":"2018. AutoML. http:\/\/www.ml4aad.org\/automl.  2018. AutoML. http:\/\/www.ml4aad.org\/automl."},{"key":"e_1_2_2_2_1","unstructured":"2018. Flask. http:\/\/flask.pocoo.org.  2018. Flask. http:\/\/flask.pocoo.org."},{"key":"e_1_2_2_3_1","unstructured":"2018. Near Ultrasound Tests. https:\/\/source.android.com\/compatibility\/cts\/near-ultrasound.  2018. Near Ultrasound Tests. https:\/\/source.android.com\/compatibility\/cts\/near-ultrasound."},{"key":"e_1_2_2_4_1","unstructured":"2018. Python_speech_features. https:\/\/github.com\/jameslyons\/python_speech_features.  2018. Python_speech_features. https:\/\/github.com\/jameslyons\/python_speech_features."},{"key":"e_1_2_2_5_1","unstructured":"2018. scikit-learn: Machine Learning in Python. http:\/\/scikit-learn.org\/stable\/.  2018. scikit-learn: Machine Learning in Python. http:\/\/scikit-learn.org\/stable\/."},{"key":"e_1_2_2_6_1","volume-title":"Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467","author":"Abadi Mart\u00edn","year":"2016","unstructured":"Mart\u00edn Abadi , Ashish Agarwal , Paul Barham , Eugene Brevdo , Zhifeng Chen , Craig Citro , Greg S Corrado , Andy Davis , Jeffrey Dean , Matthieu Devin , 2016 . Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467 (2016). Mart\u00edn Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, et al. 2016. Tensorflow: Large-scale machine learning on heterogeneous distributed systems. arXiv preprint arXiv:1603.04467 (2016)."},{"key":"e_1_2_2_7_1","doi-asserted-by":"publisher","DOI":"10.1145\/1614320.1614350"},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/INFCOM.2000.832252"},{"key":"e_1_2_2_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/1067170.1067191"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/2307636.2307653"},{"key":"e_1_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1145\/1999995.2000010"},{"key":"e_1_2_2_12_1","first-page":"2493","article-title":"Natural language processing (almost) from scratch","author":"Collobert Ronan","year":"2011","unstructured":"Ronan Collobert , Jason Weston , L\u00e9on Bottou , Michael Karlen , Koray Kavukcuoglu , and Pavel Kuksa . 2011 . Natural language processing (almost) from scratch . Journal of Machine Learning Research 12 , Aug (2011), 2493 -- 2537 . Ronan Collobert, Jason Weston, L\u00e9on Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa. 2011. Natural language processing (almost) from scratch. Journal of Machine Learning Research 12, Aug (2011), 2493--2537.","journal-title":"Journal of Machine Learning Research 12"},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1145\/2809695.2809722"},{"key":"e_1_2_2_14_1","volume-title":"UniLoc: A Unified Mobile Localization Framework Exploiting Scheme Diversity. In The 38th IEEE International Conference on Distributed Computing Systems (ICDCS). IEEE.","author":"Du Wan","year":"2018","unstructured":"Wan Du , Panrong Tong , and Mo Li . 2018 . UniLoc: A Unified Mobile Localization Framework Exploiting Scheme Diversity. In The 38th IEEE International Conference on Distributed Computing Systems (ICDCS). IEEE. Wan Du, Panrong Tong, and Mo Li. 2018. UniLoc: A Unified Mobile Localization Framework Exploiting Scheme Diversity. In The 38th IEEE International Conference on Distributed Computing Systems (ICDCS). IEEE."},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3131672.3131698"},{"key":"e_1_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.1145\/2634317.2634320"},{"key":"e_1_2_2_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/2207676.2208331"},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1145\/1023720.1023728"},{"key":"e_1_2_2_19_1","volume-title":"Perceptual linear predictive (PLP) analysis of speech. the Journal of the Acoustical Society of America 87, 4","author":"Hermansky Hynek","year":"1990","unstructured":"Hynek Hermansky . 1990. Perceptual linear predictive (PLP) analysis of speech. the Journal of the Acoustical Society of America 87, 4 ( 1990 ), 1738--1752. Hynek Hermansky. 1990. Perceptual linear predictive (PLP) analysis of speech. the Journal of the Acoustical Society of America 87, 4 (1990), 1738--1752."},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/11551201_10"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/MSP.2012.2205597"},{"key":"e_1_2_2_22_1","volume-title":"Densely connected convolutional networks. arXiv preprint arXiv:1608.06993","author":"Huang Gao","year":"2016","unstructured":"Gao Huang , Zhuang Liu , Kilian Q Weinberger , and Laurens van der Maaten . 2016. Densely connected convolutional networks. arXiv preprint arXiv:1608.06993 ( 2016 ). Gao Huang, Zhuang Liu, Kilian Q Weinberger, and Laurens van der Maaten. 2016. Densely connected convolutional networks. arXiv preprint arXiv:1608.06993 (2016)."},{"key":"e_1_2_2_23_1","unstructured":"Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105.   Alex Krizhevsky Ilya Sutskever and Geoffrey E Hinton. 2012. Imagenet classification with deep convolutional neural networks. In Advances in neural information processing systems. 1097--1105."},{"key":"e_1_2_2_24_1","volume-title":"Symbolic object localization through active sampling of acceleration and sound signatures. UbiComp 2007: Ubiquitous Computing","author":"Kunze Kai","year":"2007","unstructured":"Kai Kunze and Paul Lukowicz . 2007. Symbolic object localization through active sampling of acceleration and sound signatures. UbiComp 2007: Ubiquitous Computing ( 2007 ), 163--180. Kai Kunze and Paul Lukowicz. 2007. Symbolic object localization through active sampling of acceleration and sound signatures. UbiComp 2007: Ubiquitous Computing (2007), 163--180."},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/2750858.2804262"},{"key":"e_1_2_2_26_1","volume-title":"Deep learning. Nature 521, 7553","author":"LeCun Yann","year":"2015","unstructured":"Yann LeCun , Yoshua Bengio , and Geoffrey Hinton . 2015. Deep learning. Nature 521, 7553 ( 2015 ), 436--444. Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015. Deep learning. Nature 521, 7553 (2015), 436--444."},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2994551.2994569"},{"key":"e_1_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/2742647.2742674"},{"key":"e_1_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/3131897"},{"key":"e_1_2_2_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/1322263.1322265"},{"key":"e_1_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1145\/2459236.2459252"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/11428572_1"},{"key":"e_1_2_2_33_1","unstructured":"K. Simonyan and A. Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. ArXiv e-prints (2014). https:\/\/arxiv.org\/abs\/1409.1556.  K. Simonyan and A. Zisserman. 2014. Very Deep Convolutional Networks for Large-Scale Image Recognition. ArXiv e-prints (2014). https:\/\/arxiv.org\/abs\/1409.1556."},{"key":"e_1_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.1145\/2971648.2971684"},{"key":"e_1_2_2_35_1","unstructured":"Stephen Tarzia. 2018. Batphone. https:\/\/itunes.apple.com\/us\/app\/batphone\/id405396715?mt=8.  Stephen Tarzia. 2018. Batphone. https:\/\/itunes.apple.com\/us\/app\/batphone\/id405396715?mt=8."},{"key":"e_1_2_2_36_1","doi-asserted-by":"publisher","DOI":"10.1145\/1999995.2000011"},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/2789168.2790102"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2973750.2973764"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1145\/3131672.3131675"},{"key":"e_1_2_2_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/2307636.2307638"},{"key":"e_1_2_2_41_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02943243"}],"container-title":["Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3264945","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3264945","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T02:08:00Z","timestamp":1750212480000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3264945"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,9,18]]},"references-count":41,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2018,9,18]]}},"alternative-id":["10.1145\/3264945"],"URL":"https:\/\/doi.org\/10.1145\/3264945","relation":{},"ISSN":["2474-9567"],"issn-type":[{"value":"2474-9567","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,9,18]]},"assertion":[{"value":"2018-02-01","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-09-01","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2018-09-18","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}