{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T04:25:18Z","timestamp":1750220718661,"version":"3.41.0"},"publisher-location":"New York, NY, USA","reference-count":24,"publisher":"ACM","license":[{"start":{"date-parts":[[2020,6,8]],"date-time":"2020-06-08T00:00:00Z","timestamp":1591574400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2020,6,8]]},"DOI":"10.1145\/3372278.3390730","type":"proceedings-article","created":{"date-parts":[[2020,6,2]],"date-time":"2020-06-02T04:35:27Z","timestamp":1591072527000},"page":"301-305","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":4,"title":["At the Speed of Sound: Efficient Audio Scene Classification"],"prefix":"10.1145","author":[{"given":"Bo","family":"Dong","sequence":"first","affiliation":[{"name":"NEC Laboratories America &amp; University of Texas at Dallas, Richardson, TX, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Cristian","family":"Lumezanu","sequence":"additional","affiliation":[{"name":"NEC Laboratories America, Princeton, NJ, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Yuncong","family":"Chen","sequence":"additional","affiliation":[{"name":"NEC Laboratories America, Princeton, NJ, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Dongjin","family":"Song","sequence":"additional","affiliation":[{"name":"NEC Laboratories America, Princeton, NJ, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Takehiko","family":"Mizoguchi","sequence":"additional","affiliation":[{"name":"NEC Laboratories America, Princeton, NJ, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Haifeng","family":"Chen","sequence":"additional","affiliation":[{"name":"NEC Laboratories America, Princeton, NJ, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Latifur","family":"Khan","sequence":"additional","affiliation":[{"name":"University of Texas at Dallas, Richardson, TX, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2020,6,8]]},"reference":[{"key":"e_1_3_2_1_1_1","volume-title":"Electronics","volume":"8","author":"Aziz Sumair","year":"2019","unstructured":"Sumair Aziz , Muhammad Awais , Talha Akram , Muhammad Umar , Khursheed Khursheed , and Musaed Alhussein . 2019 . Automatic Scene Recognition through Acoustic Classification for Behavioral Robotics . Electronics , Vol. 8 (04 2019). https:\/\/doi.org\/10.3390\/electronics8050483 10.3390\/electronics8050483 Sumair Aziz, Muhammad Awais, Talha Akram, Muhammad Umar, Khursheed Khursheed, and Musaed Alhussein. 2019. Automatic Scene Recognition through Acoustic Classification for Behavioral Robotics. Electronics, Vol. 8 (04 2019). https:\/\/doi.org\/10.3390\/electronics8050483"},{"key":"e_1_3_2_1_2_1","volume-title":"Neural Machine Translation by Jointly Learning to Align and Translate. CoRR","author":"Bahdanau Dzmitry","year":"2014","unstructured":"Dzmitry Bahdanau , Kyunghyun Cho , and Yoshua Bengio . 2014. Neural Machine Translation by Jointly Learning to Align and Translate. CoRR , Vol. abs\/ 1409 .0473 ( 2014 ). Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014. Neural Machine Translation by Jointly Learning to Align and Translate. CoRR, Vol. abs\/1409.0473 (2014)."},{"key":"e_1_3_2_1_3_1","first-page":"3","article-title":"A Review of Audio Fingerprinting","volume":"41","author":"Cano Pedro","year":"2005","unstructured":"Pedro Cano , Eloi Batlle , Ton Kalker , and Jaap Haitsma . 2005 . A Review of Audio Fingerprinting . J. VLSI Signal Process. Syst. , Vol. 41 , 3 (Nov. 2005), 271--284. Pedro Cano, Eloi Batlle, Ton Kalker, and Jaap Haitsma. 2005. A Review of Audio Fingerprinting. J. VLSI Signal Process. Syst., Vol. 41, 3 (Nov. 2005), 271--284.","journal-title":"J. VLSI Signal Process. Syst."},{"key":"e_1_3_2_1_4_1","volume-title":"Neural Network Distillation on IoT Platforms for Sound Event Detection. (08","author":"Cerutti Gianmarco","year":"2019","unstructured":"Gianmarco Cerutti , Rahul Prasad , Alessio Brutti , and Elisabetta Farella . 2019. Neural Network Distillation on IoT Platforms for Sound Event Detection. (08 2019 ). Gianmarco Cerutti, Rahul Prasad, Alessio Brutti, and Elisabetta Farella. 2019. Neural Network Distillation on IoT Platforms for Sound Event Detection. (08 2019)."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/3206025.3206067"},{"key":"e_1_3_2_1_6_1","volume-title":"Signal Processing: Image Communication","volume":"70","author":"Dai Xuerui","year":"2018","unstructured":"Xuerui Dai and Xueye Wei . 2018 . HybridNet: A fast vehicle detection system for autonomous driving . Signal Processing: Image Communication , Vol. 70 (09 2018). https:\/\/doi.org\/10.1016\/j.image.2018.09.002 10.1016\/j.image.2018.09.002 Xuerui Dai and Xueye Wei. 2018. HybridNet: A fast vehicle detection system for autonomous driving. Signal Processing: Image Communication, Vol. 70 (09 2018). https:\/\/doi.org\/10.1016\/j.image.2018.09.002"},{"key":"e_1_3_2_1_7_1","volume-title":"18th Annual Conference of the International Speech Communication Association","author":"Guo Jinxi","year":"2017","unstructured":"Jinxi Guo , Ning Xu , Li-Jia Li , and Abeer Alwan . 2017 . Attention Based CLDNNs for Short-Duration Acoustic Scene Classification. In Interspeech 2017 , 18th Annual Conference of the International Speech Communication Association , Stockholm, Sweden, August 20--24 , 2017. 469--473. http:\/\/www.isca-speech.org\/archive\/Interspeech_2017\/abstracts\/0440.html Jinxi Guo, Ning Xu, Li-Jia Li, and Abeer Alwan. 2017. Attention Based CLDNNs for Short-Duration Acoustic Scene Classification. In Interspeech 2017, 18th Annual Conference of the International Speech Communication Association, Stockholm, Sweden, August 20--24, 2017. 469--473. http:\/\/www.isca-speech.org\/archive\/Interspeech_2017\/abstracts\/0440.html"},{"key":"e_1_3_2_1_8_1","first-page":"2","article-title":"Information Extraction from Sound for Medical","volume":"10","author":"Istrate D.","year":"2006","unstructured":"D. Istrate , E. Castelli , M. Vacher , L. Besacier , and J. F. Serignat . 2006 . Information Extraction from Sound for Medical Telemonitoring. Trans. Info. Tech. Biomed. , Vol. 10 , 2 (April 2006). D. Istrate, E. Castelli, M. Vacher, L. Besacier, and J. F. Serignat. 2006. Information Extraction from Sound for Medical Telemonitoring. Trans. Info. Tech. Biomed., Vol. 10, 2 (April 2006).","journal-title":"Telemonitoring. Trans. Info. Tech. Biomed."},{"key":"e_1_3_2_1_9_1","volume-title":"Saurous","author":"Jansen Aren","year":"2017","unstructured":"Aren Jansen , Manoj Plakal , Ratheet Pandya , Daniel P. W. Ellis , Shawn Hershey , Jiayang Liu , R. Channing Moore , and Rif A . Saurous . 2017 . Unsupervised Learning of Semantic Audio Representations. CoRR , Vol. abs\/ 1711 .02209 (2017). arxiv: 1711.02209 http:\/\/arxiv.org\/abs\/1711.02209 Aren Jansen, Manoj Plakal, Ratheet Pandya, Daniel P. W. Ellis, Shawn Hershey, Jiayang Liu, R. Channing Moore, and Rif A. Saurous. 2017. Unsupervised Learning of Semantic Audio Representations. CoRR, Vol. abs\/1711.02209 (2017). arxiv: 1711.02209 http:\/\/arxiv.org\/abs\/1711.02209"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1002\/0470841834"},{"key":"e_1_3_2_1_11_1","volume-title":"TUT Database for Acoustic Scene Classification and Sound Event Detection. In 24th European Signal Processing Conference 2016 (EUSIPCO","author":"Mesaros Annamaria","year":"2016","unstructured":"Annamaria Mesaros , Toni Heittola , and Tuomas Virtanen . 2016 . TUT Database for Acoustic Scene Classification and Sound Event Detection. In 24th European Signal Processing Conference 2016 (EUSIPCO 2016). Budapest, Hungary. Annamaria Mesaros, Toni Heittola, and Tuomas Virtanen. 2016. TUT Database for Acoustic Scene Classification and Sound Event Detection. In 24th European Signal Processing Conference 2016 (EUSIPCO 2016). Budapest, Hungary."},{"key":"e_1_3_2_1_12_1","volume-title":"Jorgenson","author":"Moore Arthur William","year":"1993","unstructured":"Arthur William Moore and James W . Jorgenson . 1993 . Median filtering for removal of low-frequency background drift. Analytical chemistry, Vol. 65 2 (1993), 188--91. Arthur William Moore and James W. Jorgenson. 1993. Median filtering for removal of low-frequency background drift. Analytical chemistry, Vol. 65 2 (1993), 188--91."},{"key":"e_1_3_2_1_13_1","volume-title":"Hirokazu Kameoka, and Shigeki Sagayama.","author":"Ono Nobutaka","year":"2008","unstructured":"Nobutaka Ono , Kenichi Miyamoto , Jonathan Le Roux , Hirokazu Kameoka, and Shigeki Sagayama. 2008 . Separation of a monaural audio signal into harmonic\/percussive components by complementary diffusion on spectrogram. (01 2008). Nobutaka Ono, Kenichi Miyamoto, Jonathan Le Roux, Hirokazu Kameoka, and Shigeki Sagayama. 2008. Separation of a monaural audio signal into harmonic\/percussive components by complementary diffusion on spectrogram. (01 2008)."},{"key":"e_1_3_2_1_14_1","volume-title":"Lam Dang Pham, Philipp Koch, Maarten De Vos, Ian Vince McLoughlin, and Alfred Mertins.","author":"Phan Huy","year":"2019","unstructured":"Huy Phan , Oliver Y. Ch\u00e9 n , Lam Dang Pham, Philipp Koch, Maarten De Vos, Ian Vince McLoughlin, and Alfred Mertins. 2019 . Spatio-Temporal Attention Pooling for Audio Scene Classification. CoRR , Vol. abs\/ 1904 .03543 (2019). arxiv: 1904.03543 http:\/\/arxiv.org\/abs\/1904.03543 Huy Phan, Oliver Y. Ch\u00e9 n, Lam Dang Pham, Philipp Koch, Maarten De Vos, Ian Vince McLoughlin, and Alfred Mertins. 2019. Spatio-Temporal Attention Pooling for Audio Scene Classification. CoRR, Vol. abs\/1904.03543 (2019). arxiv: 1904.03543 http:\/\/arxiv.org\/abs\/1904.03543"},{"key":"e_1_3_2_1_15_1","volume-title":"18th Annual Conference of the International Speech Communication Association","author":"Phan Huy","year":"2017","unstructured":"Huy Phan , Philipp Koch , Fabrice Katzberg , Marco Maa\u00df , Radoslaw Mazur , and Alfred Mertins . 2017 . Audio Scene Classification with Deep Recurrent Neural Networks. In Interspeech 2017 , 18th Annual Conference of the International Speech Communication Association , Stockholm, Sweden, August 20--24 , 2017. 3043--3047. http:\/\/www.isca-speech.org\/archive\/Interspeech_2017\/abstracts\/0101.html Huy Phan, Philipp Koch, Fabrice Katzberg, Marco Maa\u00df, Radoslaw Mazur, and Alfred Mertins. 2017. Audio Scene Classification with Deep Recurrent Neural Networks. In Interspeech 2017, 18th Annual Conference of the International Speech Communication Association, Stockholm, Sweden, August 20--24, 2017. 3043--3047. http:\/\/www.isca-speech.org\/archive\/Interspeech_2017\/abstracts\/0101.html"},{"key":"e_1_3_2_1_16_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCKE.2014.6993383"},{"key":"e_1_3_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"e_1_3_2_1_18_1","volume-title":"Proceedings of the 27th International Conference on Neural Information Processing Systems -","volume":"2","author":"Sutskever Ilya","unstructured":"Ilya Sutskever , Oriol Vinyals , and Quoc V. Le . 2014. Sequence to Sequence Learning with Neural Networks . In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2 (NIPS'14). Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014. Sequence to Sequence Learning with Neural Networks. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 2 (NIPS'14)."},{"key":"#cr-split#-e_1_3_2_1_19_1.1","doi-asserted-by":"crossref","unstructured":"Nicolas Turpault Romain Serizel and Emmanuel Vincent. 2019. Semi-supervised Triplet Loss Based Learning of Ambient Audio Embeddings. 760--764. https:\/\/doi.org\/10.1109\/ICASSP.2019.8683774 10.1109\/ICASSP.2019.8683774","DOI":"10.1109\/ICASSP.2019.8683774"},{"key":"#cr-split#-e_1_3_2_1_19_1.2","doi-asserted-by":"crossref","unstructured":"Nicolas Turpault Romain Serizel and Emmanuel Vincent. 2019. Semi-supervised Triplet Loss Based Learning of Ambient Audio Embeddings. 760--764. https:\/\/doi.org\/10.1109\/ICASSP.2019.8683774","DOI":"10.1109\/ICASSP.2019.8683774"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.5555\/3295222.3295349"},{"key":"e_1_3_2_1_21_1","volume-title":"Sampling Matters in Deep Embedding Learning. CoRR","author":"Wu Chao-Yuan","year":"2017","unstructured":"Chao-Yuan Wu , R. Manmatha , Alexander J. Smola , and Philipp Kr\"a henb\u00fc hl. 2017. Sampling Matters in Deep Embedding Learning. CoRR , Vol. abs\/ 1706 .07567 ( 2017 ). arxiv: 1706.07567 http:\/\/arxiv.org\/abs\/1706.07567 Chao-Yuan Wu, R. Manmatha, Alexander J. Smola, and Philipp Kr\"a henb\u00fc hl. 2017. Sampling Matters in Deep Embedding Learning. CoRR, Vol. abs\/1706.07567 (2017). arxiv: 1706.07567 http:\/\/arxiv.org\/abs\/1706.07567"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"crossref","unstructured":"Hong-Bo Zhang Yi-Xiang Zhang Bineng Zhong Qing Lei Lijie Yang Ji-Xiang Du and Duan-Sheng Chen. 2019. A Comprehensive Survey of Vision-Based Human Action Recognition Methods. In Sensors .  Hong-Bo Zhang Yi-Xiang Zhang Bineng Zhong Qing Lei Lijie Yang Ji-Xiang Du and Duan-Sheng Chen. 2019. A Comprehensive Survey of Vision-Based Human Action Recognition Methods. In Sensors .","DOI":"10.3390\/s19051005"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1007\/BF02943243"}],"event":{"name":"ICMR '20: International Conference on Multimedia Retrieval","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Dublin Ireland","acronym":"ICMR '20"},"container-title":["Proceedings of the 2020 International Conference on Multimedia Retrieval"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3372278.3390730","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3372278.3390730","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T22:33:25Z","timestamp":1750199605000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3372278.3390730"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,6,8]]},"references-count":24,"alternative-id":["10.1145\/3372278.3390730","10.1145\/3372278"],"URL":"https:\/\/doi.org\/10.1145\/3372278.3390730","relation":{},"subject":[],"published":{"date-parts":[[2020,6,8]]},"assertion":[{"value":"2020-06-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}