{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T20:00:26Z","timestamp":1765310426769,"version":"3.46.0"},"publisher-location":"New York, NY, USA","reference-count":90,"publisher":"ACM","funder":[{"name":"National Natural Science Foundation of China","award":["No. 62372364"],"award-info":[{"award-number":["No. 62372364"]}]},{"name":"Technical Innovation Guidance Plan of Shaanxi Province, China","award":["No. 2024QCY-KXJ-199"],"award-info":[{"award-number":["No. 2024QCY-KXJ-199"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,10,27]]},"DOI":"10.1145\/3746027.3755149","type":"proceedings-article","created":{"date-parts":[[2025,10,25]],"date-time":"2025-10-25T07:30:51Z","timestamp":1761377451000},"page":"1346-1355","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["Ear with Eye: Lightweight Multimodal Audio-Visual Network Inspired by Bionic Structures"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0009-0009-8452-4664","authenticated-orcid":false,"given":"Xuanming","family":"Jiang","sequence":"first","affiliation":[{"name":"Xi'an Jiaotong University, Xi'an, China and Xi'an Jiyun Technology Co., Ltd., Xi'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1499-5892","authenticated-orcid":false,"given":"Baoyi","family":"An","sequence":"additional","affiliation":[{"name":"Xi'an Jiyun Technology Co., Ltd., Xi'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-2718-304X","authenticated-orcid":false,"given":"Zhengwei","family":"Zou","sequence":"additional","affiliation":[{"name":"Xi'an Jiaotong University, Xi'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-2881-4075","authenticated-orcid":false,"given":"Dingyu","family":"Nie","sequence":"additional","affiliation":[{"name":"Xi'an Jiyun Technology Co., Ltd., Xi'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4560-8509","authenticated-orcid":false,"given":"Jialie","family":"Shen","sequence":"additional","affiliation":[{"name":"City St George's, University of London, London, United Kingdom"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3173-6307","authenticated-orcid":false,"given":"Xueming","family":"Qian","sequence":"additional","affiliation":[{"name":"Xi'an Jiaotong University, Xi'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4392-8450","authenticated-orcid":false,"given":"Guoshuai","family":"Zhao","sequence":"additional","affiliation":[{"name":"Xi'an Jiaotong University, Xi'an, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,10,27]]},"reference":[{"key":"e_1_3_2_1_1_1","doi-asserted-by":"publisher","DOI":"10.1083\/jcb.87.2.451"},{"key":"e_1_3_2_1_2_1","doi-asserted-by":"publisher","DOI":"10.1093\/gbe\/evz111"},{"key":"e_1_3_2_1_3_1","first-page":"177","article-title":"Structure and function of the bird fovea. Anatomia, Histologia","volume":"48","author":"Bringmann Andreas","year":"2019","unstructured":"Andreas Bringmann. 2019. Structure and function of the bird fovea. Anatomia, Histologia, Embryologia, Vol. 48, 3 (2019), 177-200.","journal-title":"Embryologia"},{"volume-title":"Artificial Intelligence and Computational Intelligence","author":"Chen Chen","key":"e_1_3_2_1_4_1","unstructured":"Chen Chen, Weijun Li, and Liang Chen. 2011. An improved biomimetic image processing method. In Artificial Intelligence and Computational Intelligence. Springer Berlin Heidelberg, Berlin, Heidelberg, 246-254."},{"key":"e_1_3_2_1_5_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP43922.2022.9746312"},{"key":"e_1_3_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2017.2675998"},{"key":"e_1_3_2_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3447085"},{"key":"e_1_3_2_1_8_1","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681679"},{"key":"e_1_3_2_1_9_1","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681595"},{"key":"e_1_3_2_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSEN.2021.3064588"},{"key":"e_1_3_2_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01181"},{"key":"e_1_3_2_1_12_1","first-page":"1","article-title":"Hearing in birds and reptiles: An overview","volume":"13","author":"Dooling Robert J.","year":"2012","unstructured":"Robert J. Dooling, Richard R. Fay, and Arthur N. Popper. 2012. Hearing in birds and reptiles: An overview. Comparative Hearing: Birds and Reptiles, Vol. 13 (2012), 1-361.","journal-title":"Comparative Hearing: Birds and Reptiles"},{"volume-title":"A History of Discoveries on Hearing","author":"Dooling Robert J","key":"e_1_3_2_1_13_1","unstructured":"Robert J Dooling and Georg M Klump. 2023. Birds as a model in hearing research. In A History of Discoveries on Hearing. Springer, New York, USA, 151-185."},{"key":"e_1_3_2_1_14_1","doi-asserted-by":"publisher","DOI":"10.1121\/1.1494447"},{"key":"e_1_3_2_1_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681572"},{"volume-title":"Neural Principles in Vision","author":"Gallego A","key":"e_1_3_2_1_16_1","unstructured":"A Gallego. 1976. Comparative study of the horizontal cells in the vertebrate retina: Mammals and birds. In Neural Principles in Vision. Springer Berlin Heidelberg, Berlin, Heidelberg, 26-62."},{"key":"e_1_3_2_1_17_1","first-page":"37","article-title":"Metric learning-based multimodal audio-visual emotion recognition","volume":"27","author":"Ghaleb Esam","year":"2020","unstructured":"Esam Ghaleb, Mirela Popa, and Stylianos Asteriadis. 2020. Metric learning-based multimodal audio-visual emotion recognition. IEEE MultiMedia, Vol. 27, 1 (2020), 37-48.","journal-title":"IEEE MultiMedia"},{"key":"e_1_3_2_1_18_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neuron.2009.12.009"},{"key":"e_1_3_2_1_19_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i10.21315"},{"key":"e_1_3_2_1_20_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.febslet.2015.08.030"},{"key":"e_1_3_2_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.01265"},{"key":"e_1_3_2_1_22_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.tics.2023.11.002"},{"key":"e_1_3_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1109\/IGARSS.2018.8519248"},{"key":"e_1_3_2_1_24_1","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2019.2918242"},{"key":"e_1_3_2_1_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/TETCI.2017.2784878"},{"key":"e_1_3_2_1_26_1","doi-asserted-by":"publisher","DOI":"10.1002\/cne.24896"},{"key":"e_1_3_2_1_27_1","volume-title":"Proceedings of the 38th International Conference on Machine Learning. PMLR, Virtual, 4651-4664","author":"Jaegle Andrew","year":"2021","unstructured":"Andrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals, Andrew Zisserman, and Joao Carreira. 2021. Perceiver: General perception with iterative attention. In Proceedings of the 38th International Conference on Machine Learning. PMLR, Virtual, 4651-4664."},{"key":"e_1_3_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.celrep.2021.108900"},{"key":"e_1_3_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1038\/nrn1606"},{"key":"e_1_3_2_1_30_1","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v39i17.33940"},{"key":"e_1_3_2_1_31_1","doi-asserted-by":"publisher","DOI":"10.32604\/cmc.2024.053204"},{"key":"e_1_3_2_1_32_1","doi-asserted-by":"publisher","DOI":"10.1053\/j.jepm.2007.03.012"},{"key":"e_1_3_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neuron.2018.03.044"},{"key":"e_1_3_2_1_34_1","volume-title":"Computer Vision - ECCV","author":"Kim Donghyun","year":"2024","unstructured":"Donghyun Kim, Byeongho Heo, and Dongyoon Han. 2024. DenseNets reloaded: Paradigm shift beyond ResNets and ViTs. In Computer Vision - ECCV 2024. Springer International Publishing, Cham, 395-415."},{"key":"e_1_3_2_1_35_1","volume-title":"Kingma and Jimmy Ba","author":"Diederik","year":"2014","unstructured":"Diederik P. Kingma and Jimmy Ba. 2014. Adam: A Method for Stochastic Optimization. arXiv:1412.6980 [cs.LG]"},{"key":"e_1_3_2_1_36_1","volume-title":"Advances in Neural Information Processing Systems","volume":"25","author":"Krizhevsky Alex","year":"2012","unstructured":"Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. ImageNet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, Vol. 25. Curran Associates, Inc., USA."},{"key":"e_1_3_2_1_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3688986"},{"key":"e_1_3_2_1_38_1","doi-asserted-by":"publisher","DOI":"10.1364\/JOSA.61.000001"},{"key":"e_1_3_2_1_39_1","doi-asserted-by":"publisher","DOI":"10.3390\/biomimetics9080453"},{"key":"e_1_3_2_1_40_1","doi-asserted-by":"publisher","DOI":"10.1145\/3664647.3681626"},{"key":"e_1_3_2_1_41_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.00228"},{"key":"e_1_3_2_1_42_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.apacoust.2021.108213"},{"key":"e_1_3_2_1_43_1","doi-asserted-by":"publisher","DOI":"10.1038\/nn1367"},{"key":"e_1_3_2_1_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01167"},{"key":"e_1_3_2_1_45_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2018.10.010"},{"key":"e_1_3_2_1_46_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41578-024-00750-6"},{"key":"e_1_3_2_1_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2019.2963387"},{"key":"e_1_3_2_1_48_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.image.2021.116466"},{"key":"e_1_3_2_1_49_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10336-007-0213-6"},{"key":"e_1_3_2_1_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-19836-6_18"},{"key":"e_1_3_2_1_51_1","first-page":"14200","article-title":"Attention bottlenecks for multimodal fusion","volume":"34","author":"Nagrani Arsha","year":"2021","unstructured":"Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, and Chen Sun. 2021. Attention bottlenecks for multimodal fusion. Advances in Neural Information Processing Systems, Vol. 34 (2021), 14200-14213.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_1_52_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10097236"},{"key":"e_1_3_2_1_53_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.02627"},{"key":"e_1_3_2_1_54_1","first-page":"325","article-title":"Skull asymmetry, ear structure and function, and auditory localization in Tengmalm's owl, Aegolius funereus (Linn\u00e9). Philosophical Transactions of the Royal Society of London. B","volume":"282","author":"Norberg R \u00c5ke","year":"1978","unstructured":"R \u00c5ke Norberg. 1978. Skull asymmetry, ear structure and function, and auditory localization in Tengmalm's owl, Aegolius funereus (Linn\u00e9). Philosophical Transactions of the Royal Society of London. B, Biological Sciences, Vol. 282, 991 (1978), 325-410.","journal-title":"Biological Sciences"},{"key":"e_1_3_2_1_55_1","doi-asserted-by":"crossref","unstructured":"James D Paruk. 2018. The cornell lab of ornithology handbook of bird biology.","DOI":"10.1642\/AUK-18-104.1"},{"key":"e_1_3_2_1_56_1","doi-asserted-by":"publisher","DOI":"10.1145\/3583133.3596301"},{"key":"e_1_3_2_1_57_1","first-page":"973","volume-title":"Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics","volume":"1","author":"P\u00e9rez-Rosas Ver\u00f3nica","year":"2013","unstructured":"Ver\u00f3nica P\u00e9rez-Rosas, Rada Mihalcea, and Louis-Philippe Morency. 2013. Utterance-level multimodal sentiment analysis. In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics, Vol. 1. Association for Computational Linguistics, Sofia, Bulgaria, 973-982."},{"key":"e_1_3_2_1_58_1","doi-asserted-by":"publisher","DOI":"10.1145\/2733373.2806390"},{"volume-title":"Seminars in Cell & Developmental Biology","author":"Potier Simon","key":"e_1_3_2_1_59_1","unstructured":"Simon Potier, Mindaugas Mitkus, and Almut Kelber. 2020. Visual adaptations of diurnal and nocturnal raptors. In Seminars in Cell & Developmental Biology, Vol. 106. Elsevier, Amsterdam, Netherlands, 116-126."},{"key":"e_1_3_2_1_60_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW56347.2022.00278"},{"key":"e_1_3_2_1_61_1","doi-asserted-by":"publisher","DOI":"10.1242\/jeb.25.3.299"},{"key":"e_1_3_2_1_62_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.neubiorev.2022.104942"},{"key":"e_1_3_2_1_63_1","doi-asserted-by":"publisher","DOI":"10.1145\/2647868.2655045"},{"key":"e_1_3_2_1_64_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11042-025-20709-1"},{"key":"e_1_3_2_1_65_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP49357.2023.10096110"},{"volume-title":"MultiMedia Modeling","author":"Shen Jialie","key":"e_1_3_2_1_66_1","unstructured":"Jialie Shen, Liqiang Nie, and Tat-Seng Chua. 2016. Smart ambient sound analysis via structured statistical Modeling. In MultiMedia Modeling. Springer International Publishing, Cham, 231-243."},{"key":"e_1_3_2_1_67_1","doi-asserted-by":"publisher","DOI":"10.1111\/jmi.13192"},{"key":"e_1_3_2_1_68_1","volume-title":"Science","volume":"388","author":"Sun Yuwei","year":"2025","unstructured":"Yuwei Sun, Minhui Ren, Yu Zhang, Shuting Li, Zhengnan Luo, Suhong Sun, Shunji He, Guangqin Wang, Di Zhang, Suzanne L. Mansour, Lei Song, and Zhiyong Liu. 2025. Casz1 is required for both inner hair cell fate stabilization and outer hair cell survival. Science, Vol. 388 (2025), eado4930."},{"key":"e_1_3_2_1_69_1","first-page":"6105","volume-title":"Proceedings of the 36th International Conference on Machine Learning","volume":"97","author":"Tan Mingxing","year":"2019","unstructured":"Mingxing Tan and Quoc Le. 2019. EfficientNet: Rethinking model scaling for convolutional neural networks. In Proceedings of the 36th International Conference on Machine Learning, Vol. 97. PMLR, Long Beach, CA, USA, 6105-6114."},{"key":"e_1_3_2_1_70_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01216-8_16"},{"key":"e_1_3_2_1_71_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00532"},{"key":"e_1_3_2_1_72_1","first-page":"155","article-title":"Multimodal sentiment sensing and emotion recognition based on cognitive computing using hidden Markov model with extreme learning machine","volume":"14","author":"Verma Diksha","year":"2022","unstructured":"Diksha Verma, Sweta Kumari Barnwal, Amit Barve, MK Jayanthi Kannan, Rajesh Gupta, and R Swaminathan. 2022. Multimodal sentiment sensing and emotion recognition based on cognitive computing using hidden Markov model with extreme learning machine. International Journal of Communication Networks and Information Security, Vol. 14, 2 (2022), 155-167.","journal-title":"International Journal of Communication Networks and Information Security"},{"key":"e_1_3_2_1_73_1","first-page":"342","article-title":"The bird's eye view","volume":"78","author":"Waldvogel Jerry A","year":"1990","unstructured":"Jerry A Waldvogel. 1990. The bird's eye view. American Scientist, Vol. 78, 4 (1990), 342-353.","journal-title":"American Scientist"},{"key":"e_1_3_2_1_74_1","doi-asserted-by":"publisher","DOI":"10.1088\/1748-3190\/ad2085"},{"key":"e_1_3_2_1_75_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW63382.2024.00190"},{"key":"e_1_3_2_1_76_1","unstructured":"Pete Warden. 2018. Speech commands: A dataset for limited-vocabulary speech recognition. arXiv:1804.03209 [cs.CL]"},{"key":"e_1_3_2_1_77_1","doi-asserted-by":"publisher","DOI":"10.1109\/CCGrid59990.2024.00053"},{"key":"e_1_3_2_1_78_1","doi-asserted-by":"publisher","DOI":"10.1002\/cne.25524"},{"key":"e_1_3_2_1_79_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00581"},{"key":"e_1_3_2_1_80_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01936"},{"key":"e_1_3_2_1_81_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.eswa.2024.123768"},{"key":"e_1_3_2_1_82_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.00529"},{"key":"e_1_3_2_1_83_1","first-page":"1","article-title":"Incremental audio-visual fusion for person recognition in earthquake scene","volume":"20","author":"You Sisi","year":"2023","unstructured":"Sisi You, Yukun Zuo, Hantao Yao, and Changsheng Xu. 2023. Incremental audio-visual fusion for person recognition in earthquake scene. ACM Transactions on Multimedia Computing, Communications and Applications, Vol. 20, 2 (2023), 1-19.","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"e_1_3_2_1_84_1","doi-asserted-by":"publisher","DOI":"10.1016\/0042-6989(84)90098-1"},{"key":"e_1_3_2_1_85_1","doi-asserted-by":"publisher","DOI":"10.1145\/3503161.3547869"},{"key":"e_1_3_2_1_86_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3318015"},{"key":"e_1_3_2_1_87_1","doi-asserted-by":"publisher","DOI":"10.1109\/WACV61041.2025.00218"},{"key":"e_1_3_2_1_88_1","doi-asserted-by":"publisher","DOI":"10.1145\/3649447"},{"key":"e_1_3_2_1_89_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2024.3420239"},{"key":"e_1_3_2_1_90_1","doi-asserted-by":"publisher","DOI":"10.1109\/JPROC.2020.3004555"}],"event":{"name":"MM '25: The 33rd ACM International Conference on Multimedia","sponsor":["SIGMM ACM Special Interest Group on Multimedia"],"location":"Dublin Ireland","acronym":"MM '25"},"container-title":["Proceedings of the 33rd ACM International Conference on Multimedia"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3746027.3755149","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T19:56:29Z","timestamp":1765310189000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3746027.3755149"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,10,27]]},"references-count":90,"alternative-id":["10.1145\/3746027.3755149","10.1145\/3746027"],"URL":"https:\/\/doi.org\/10.1145\/3746027.3755149","relation":{},"subject":[],"published":{"date-parts":[[2025,10,27]]},"assertion":[{"value":"2025-10-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}