{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T02:09:24Z","timestamp":1777342164222,"version":"3.51.4"},"publisher-location":"New York, NY, USA","reference-count":36,"publisher":"ACM","license":[{"start":{"date-parts":[[2025,10,27]],"date-time":"2025-10-27T00:00:00Z","timestamp":1761523200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/100000001","name":"NSF (National Science Foundation)","doi-asserted-by":"publisher","award":["2349713"],"award-info":[{"award-number":["2349713"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746027.3758298","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T07:26:55Z","timestamp":1761377215000},"page":"13369-13375","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":1,"title":["DogSpeak: A Canine Vocalization Classification Dataset"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-7256-7503","authenticated-orcid":false,"given":"Hridayesh","family":"Lekhak","sequence":"first","affiliation":[{"name":"The University of Texas at Arlington, Arlington, Texas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-5329-9620","authenticated-orcid":false,"given":"Theron S.","family":"Wang","sequence":"additional","affiliation":[{"name":"The University of Texas at Arlington, Arlington, Texas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-0686-5184","authenticated-orcid":false,"given":"Tuan M.","family":"Dang","sequence":"additional","affiliation":[{"name":"The University of Texas at Arlington, Arlington, Texas, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3782-3230","authenticated-orcid":false,"given":"Kenny Q.","family":"Zhu","sequence":"additional","affiliation":[{"name":"The University of Texas at Arlington, Arlington, TX, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"e_1_3_2_2_1_1","volume-title":"Humberto P\u00e9rez Espinosa, and Rada Mihalcea","author":"Abzaliev Artem","year":"2024","unstructured":"Artem Abzaliev, Humberto P\u00e9rez Espinosa, and Rada Mihalcea. 2024. Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification. arXiv preprint arXiv:2404.18739 (2024)."},{"key":"e_1_3_2_2_2_1","volume-title":"Raven food calls indicate sender's age and sex. Frontiers in zoology","author":"Boeckle Markus","year":"2018","unstructured":"Markus Boeckle, Georgine Szipl, and Thomas Bugnyar. 2018. Raven food calls indicate sender's age and sex. Frontiers in zoology, Vol. 15 (2018), 1-9."},{"key":"e_1_3_2_2_3_1","volume-title":"Random forests. Machine learning","author":"Breiman Leo","year":"2001","unstructured":"Leo Breiman. 2001. Random forests. Machine learning, Vol. 45 (2001), 5-32."},{"key":"e_1_3_2_2_4_1","volume-title":"Beats: Audio pre-training with acoustic tokenizers. arXiv preprint arXiv:2212.09058","author":"Chen Sanyuan","year":"2022","unstructured":"Sanyuan Chen, Yu Wu, Chengyi Wang, Shujie Liu, Daniel Tompkins, Zhuo Chen, and Furu Wei. 2022. Beats: Audio pre-training with acoustic tokenizers. arXiv preprint arXiv:2212.09058 (2022)."},{"key":"e_1_3_2_2_5_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939785"},{"key":"e_1_3_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/TASSP.1980.1163420"},{"key":"e_1_3_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Florian Eyben Klaus R Scherer Bj\u00f6rn W Schuller Johan Sundberg Elisabeth Andr\u00e9 Carlos Busso Laurence Y Devillers Julien Epps Petri Laukka Shrikanth S Narayanan et al. 2015. The Geneva minimalistic acoustic parameter set (GeMAPS) for voice research and affective computing. IEEE transactions on affective computing Vol. 7 2 (2015) 190-202.","DOI":"10.1109\/TAFFC.2015.2457417"},{"key":"e_1_3_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/1873951.1874246"},{"key":"e_1_3_2_2_9_1","volume-title":"Biocommunication of animals","author":"Farag\u00f3 Tam\u00e1s","unstructured":"Tam\u00e1s Farag\u00f3, Simon Townsend, and Friederike Range. 2013. The information content of wolf (and dog) social communication. In Biocommunication of animals. Springer, 41-62."},{"key":"e_1_3_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1098\/rsos.150372"},{"key":"e_1_3_2_2_11_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.beproc.2024.105028"},{"key":"e_1_3_2_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8461757"},{"key":"e_1_3_2_2_13_1","volume-title":"The elements of statistical learning: data mining, inference, and prediction","author":"Hastie Trevor","unstructured":"Trevor Hastie, Robert Tibshirani, Jerome H Friedman, and Jerome H Friedman. 2009. The elements of statistical learning: data mining, inference, and prediction. Vol. 2. Springer."},{"key":"e_1_3_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1093\/gerona\/glx061"},{"key":"e_1_3_2_2_15_1","volume-title":"Kushal Lakhotia","author":"Hsu Wei-Ning","year":"2021","unstructured":"Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed. 2021. Hubert: Self-supervised speech representation learning by masked prediction of hidden units. IEEE\/ACM transactions on audio, speech, and language processing, Vol. 29 (2021), 3451-3460."},{"key":"e_1_3_2_2_16_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.findings-acl.869"},{"key":"e_1_3_2_2_17_1","volume-title":"K Hammerschmidt, and Bernhard Englitz.","author":"Ivanenko Alexander","year":"2020","unstructured":"Alexander Ivanenko, Paul Watkins, Marcel AJ van Gerven, K Hammerschmidt, and Bernhard Englitz. 2020. Classifying sex and strain from mouse ultrasonic vocalizations using deep learning. PLoS computational biology, Vol. 16, 6 (2020), e1007918."},{"key":"e_1_3_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.388089"},{"key":"e_1_3_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/3478384.3478385"},{"key":"e_1_3_2_2_20_1","volume-title":"Comparing supervised learning methods for classifying sex, age, context and individual Mudi dogs from barking. Animal cognition","author":"Larranaga Ana","year":"2015","unstructured":"Ana Larranaga, Concha Bielza, P\u00e9ter Pongr\u00e1cz, Tam\u00e1s Farag\u00f3, Anna B\u00e1lint, and Pedro Larranaga. 2015. Comparing supervised learning methods for classifying sex, age, context and individual Mudi dogs from barking. Animal cognition, Vol. 18, 2 (2015), 405-421."},{"key":"e_1_3_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.21437\/Interspeech.2025-1287"},{"key":"e_1_3_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.3390\/ani12223106"},{"key":"e_1_3_2_2_23_1","volume-title":"Separate anything you describe. arXiv preprint arXiv:2308.05037","author":"Liu Xubo","year":"2023","unstructured":"Xubo Liu, Qiuqiang Kong, Yan Zhao, Haohe Liu, Yi Yuan, Yuzhuo Liu, Rui Xia, Yuxuan Wang, Mark D Plumbley, and Wenwu Wang. 2023. Separate anything you describe. arXiv preprint arXiv:2308.05037 (2023)."},{"key":"e_1_3_2_2_24_1","doi-asserted-by":"publisher","DOI":"10.25080\/Majora-7b98e3ed-003"},{"key":"e_1_3_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10071-007-0129-9"},{"key":"e_1_3_2_2_26_1","volume-title":"Can humans discriminate between dogs on the base of the acoustic parameters of barks? Behavioural processes","author":"Moln\u00e1r Csaba","year":"2006","unstructured":"Csaba Moln\u00e1r, P\u00e9ter Pongr\u00e1cz, Antal D\u00f3ka, and \u00c1d\u00e1m Mikl\u00f3si. 2006. Can humans discriminate between dogs on the base of the acoustic parameters of barks? Behavioural processes, Vol. 73, 1 (2006), 76-83."},{"key":"e_1_3_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1371\/journal.pone.0295840"},{"key":"e_1_3_2_2_28_1","doi-asserted-by":"publisher","DOI":"10.3390\/ani9080543"},{"key":"e_1_3_2_2_29_1","doi-asserted-by":"publisher","DOI":"10.3233\/JIFS-169509"},{"key":"e_1_3_2_2_30_1","volume-title":"Do children understand man's best friend? Classification of dog barks by pre-adolescents and adults. Applied animal behaviour science","author":"Pongr\u00e1cz P\u00e9ter","year":"2011","unstructured":"P\u00e9ter Pongr\u00e1cz, Csaba Moln\u00e1r, Antal D\u00f3ka, and \u00c1d\u00e1m Mikl\u00f3si. 2011. Do children understand man's best friend? Classification of dog barks by pre-adolescents and adults. Applied animal behaviour science, Vol. 135, 1-2 (2011), 95-102."},{"key":"e_1_3_2_2_31_1","doi-asserted-by":"publisher","DOI":"10.1037\/0735-7036.119.2.136"},{"key":"e_1_3_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.3390\/ani10122390"},{"key":"e_1_3_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2024.findings-emnlp.816"},{"key":"e_1_3_2_2_34_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2025.acl-long.451"},{"key":"e_1_3_2_2_35_1","volume-title":"Age influences domestic dog cognitive performance independent of average breed lifespan. Animal cognition","author":"Watowich Marina M","year":"2020","unstructured":"Marina M Watowich, Evan L MacLean, Brian Hare, Josep Call, Juliane Kaminski, \u00c1d\u00e1m Mikl\u00f3si, and Noah Snyder-Mackler. 2020. Age influences domestic dog cognitive performance independent of average breed lifespan. Animal cognition, Vol. 23 (2020), 795-805."},{"key":"e_1_3_2_2_36_1","volume-title":"Barking in domestic dogs: context specificity and individual identification. Animal behaviour","author":"Yin Sophia","year":"2004","unstructured":"Sophia Yin and Brenda McCowan. 2004. Barking in domestic dogs: context specificity and individual identification. Animal behaviour, Vol. 68, 2 (2004), 343-355."}],"event":{"name":"MM '25: The 33rd ACM International Conference on Multimedia","location":"Dublin Ireland","acronym":"MM '25","sponsor":["SIGMM ACM Special Interest Group on Multimedia"]},"container-title":["Proceedings of the 33rd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/abs\/10.1145\/3746027.3758298","content-type":"text\/html","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3758298","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3758298","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,10]],"date-time":"2025-12-10T05:07:01Z","timestamp":1765343221000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746027.3758298"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":36,"alternative-id":["10.1145\/3746027.3758298","10.1145\/3746027"],"URL":"https:\/\/doi.org\/10.1145\/3746027.3758298","relation":{},"subject":[],"published":{"date-parts":[[2025,10,27]]},"assertion":[{"value":"2025-10-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}