{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,6]],"date-time":"2026-02-06T09:39:56Z","timestamp":1770370796410,"version":"3.49.0"},"reference-count":35,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T00:00:00Z","timestamp":1770249600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T00:00:00Z","timestamp":1770249600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004837","name":"Ministerio de Ciencia e Innovaci\u00f3n","doi-asserted-by":"publisher","award":["PID2022-137048OA-C43"],"award-info":[{"award-number":["PID2022-137048OA-C43"]}],"id":[{"id":"10.13039\/501100004837","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100012818","name":"Comunidad de Madrid","doi-asserted-by":"publisher","award":["TEC-2024\/COM-360"],"award-info":[{"award-number":["TEC-2024\/COM-360"]}],"id":[{"id":"10.13039\/100012818","id-type":"DOI","asserted-by":"publisher"}]},{"name":"Universidad Carlos III"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Advances in object tracking and acoustic beamforming are driving new capabilities in surveillance, human-computer interaction, and robotics. This work presents an embedded system that integrates deep learning\u2013based tracking with beamforming to achieve precise sound source localization and directional audio capture in dynamic environments. The approach combines single-camera depth estimation and stereo vision to enable accurate 3D localization of moving objects. A planar concentric circular microphone array constructed with MEMS microphones provides a compact, energy-efficient platform supporting 2D beam steering across azimuth and elevation. Real-time tracking outputs continuously adapt the array\u2019s focus, synchronizing the acoustic response with the target\u2019s position. By uniting learned spatial awareness with dynamic steering, the system maintains robust performance in the presence of multiple or moving sources. Experimental evaluation demonstrates significant gains in signal-to-interference ratio, making the design well-suited for teleconferencing, smart home devices, and assistive technologies.<\/jats:p>","DOI":"10.1007\/s11227-026-08261-7","type":"journal-article","created":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T18:26:00Z","timestamp":1770315960000},"update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Real-time object tracking with on-device deep learning for adaptive beamforming in dynamic acoustic environments"],"prefix":"10.1007","volume":"82","author":[{"given":"Jorge","family":"Ortigoso-Narro","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jose A.","family":"Belloch","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Adrian","family":"Amor-Martin","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Sandra","family":"Roger","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Maximo","family":"Cobos","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2026,2,5]]},"reference":[{"key":"8261_CR1","doi-asserted-by":"publisher","DOI":"10.3390\/s22228791","author":"S-C Hsia","year":"2022","unstructured":"Hsia S-C, Wang S-H, Wei C-M, Chang C-Y (2022) Intelligent object tracking with an automatic image zoom algorithm for a camera sensing surveillance system. Sensors. https:\/\/doi.org\/10.3390\/s22228791","journal-title":"Sensors"},{"key":"8261_CR2","doi-asserted-by":"publisher","unstructured":"Price TPW, Howard DM, Lewis AV, Tyrrell AM (1999) Adaptive microphone array beamforming for teleconferencing using vhdl and parallel architectures. In: Proceedings of the Seventh Euromicro Workshop on Parallel and Distributed Processing. PDP\u201999. https:\/\/doi.org\/10.1109\/EMPDP.1999.746639","DOI":"10.1109\/EMPDP.1999.746639"},{"key":"8261_CR3","doi-asserted-by":"publisher","DOI":"10.3390\/s25010214","author":"M Trigka","year":"2025","unstructured":"Trigka M, Dritsas E (2025) A comprehensive survey of machine learning techniques and models for object detection. Sensors. https:\/\/doi.org\/10.3390\/s25010214","journal-title":"Sensors"},{"key":"8261_CR4","doi-asserted-by":"publisher","unstructured":"Jacob B, Kligys S, Chen B, Zhu M, Tang M, Howard A, Adam H, Kalenichenko D (2018) Quantization and training of neural networks for efficient integer-arithmetic-only inference. In: 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition. https:\/\/doi.org\/10.1109\/CVPR.2018.00286","DOI":"10.1109\/CVPR.2018.00286"},{"key":"8261_CR5","unstructured":"Han S, Pool J, Tran J, Dally WJ (2015) Learning both weights and connections for efficient neural networks. In: Proceedings of the 29th International Conference on Neural Information Processing Systems - Volume 1. NIPS\u201915, pp. 1135\u20131143. MIT Press, Cambridge, MA, USA"},{"key":"8261_CR6","unstructured":"Howard AG, Zhu M, Chen B, Kalenichenko D, Wang W, Weyand T, Andreetto M, Adam H (2017) MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. https:\/\/arxiv.org\/abs\/1704.04861"},{"key":"8261_CR7","unstructured":"Ortigoso-Narro, J., Moreno, R., De\u00a0La\u00a0Prida, D., Raiola, M., Azpicueta-Ruiz, L.A.: 64-Microphone Module for a Massive Acoustic Camera, Faro (2024)"},{"key":"8261_CR8","doi-asserted-by":"publisher","unstructured":"Mizoguchi H, Tamai Y, Shinoda K, Kagami S, Nagasghima K (2004) Visually steerable sound beam forming system based on face tracking and speaker array. In: Proceedings of the 17th International Conference on Pattern Recognition, 2004. ICPR 2004. 3, 977\u20139803 . https:\/\/doi.org\/10.1109\/ICPR.2004.1334692","DOI":"10.1109\/ICPR.2004.1334692"},{"key":"8261_CR9","doi-asserted-by":"publisher","DOI":"10.3390\/app13106056","author":"Z Shi","year":"2023","unstructured":"Shi Z, Zhang L, Wang D (2023) Audio-visual sound source localization and tracking based on mobile robot for the cocktail party problem. Appl Sci. https:\/\/doi.org\/10.3390\/app13106056","journal-title":"Appl Sci"},{"issue":"2","key":"8261_CR10","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1023\/b:visi.0000029664.99615.94","volume":"60","author":"DG Lowe","year":"2004","unstructured":"Lowe DG (2004) Distinctive image features from scale-invariant keypoints. Int J Comput Vision 60(2):91\u2013110. https:\/\/doi.org\/10.1023\/b:visi.0000029664.99615.94","journal-title":"Int J Comput Vision"},{"key":"8261_CR11","doi-asserted-by":"publisher","unstructured":"Dalal N, Triggs B (2005) Histograms of oriented gradients for human detection. In: 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR\u201905) 1:886\u20138931. https:\/\/doi.org\/10.1109\/CVPR.2005.177","DOI":"10.1109\/CVPR.2005.177"},{"key":"8261_CR12","unstructured":"Csurka G, Dance CR, Fan L, Willamowski JK, Bray C (2002) Visual categorization with bags of keypoints. In: European Conference on Computer Vision"},{"key":"8261_CR13","doi-asserted-by":"crossref","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 770\u2013778","DOI":"10.1109\/CVPR.2016.90"},{"key":"8261_CR14","unstructured":"Simonyan K, Zisserman A (2015) Very deep convolutional networks for large-scale image recognition. https:\/\/arxiv.org\/abs\/1409.1556"},{"key":"8261_CR15","unstructured":"Khanam R, Hussain M (2024) YOLOv11: an overview of the key architectural enhancements. https:\/\/arxiv.org\/abs\/2410.17725"},{"key":"8261_CR16","unstructured":"Yang L, Kang B, Huang Z, Zhao Z, Xu X, Feng J, Zhao H (2024) Depth anything v2. arXiv:2406.09414"},{"key":"8261_CR17","doi-asserted-by":"crossref","unstructured":"Godard C, Mac Aodha O, Firman M, Brostow GJ (2019) Digging into self-supervised monocular depth prediction. In: The International Conference on Computer Vision (ICCV)","DOI":"10.1109\/ICCV.2019.00393"},{"key":"8261_CR18","doi-asserted-by":"publisher","unstructured":"Feng,C, Chen Z, Zhang C, Hu W, Li B, Ge L (2024) Real-time monocular depth estimation on embedded systems. In: Proc. IEEE Int. Conf. Image Process. (ICIP), Abu Dhabi, United Arab Emirates, pp. 3464\u20133470. https:\/\/doi.org\/10.1109\/ICIP51287.2024.10648152","DOI":"10.1109\/ICIP51287.2024.10648152"},{"key":"8261_CR19","doi-asserted-by":"crossref","unstructured":"Lipson L, Teed Z, Deng J (2021) Raft-stereo: Multilevel recurrent field transforms for stereo matching. In: International Conference on 3D Vision (3DV)","DOI":"10.1109\/3DV53792.2021.00032"},{"key":"8261_CR20","doi-asserted-by":"crossref","unstructured":"Li J, Wang P, Xiong P, Cai T, Yan Z, Yang L, Liu J, Fan H, Liu S (2022) Practical stereo matching via cascaded recurrent network with adaptive correlation. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, pp. 16263\u201316272","DOI":"10.1109\/CVPR52688.2022.01578"},{"key":"8261_CR21","unstructured":"Contributors X-S (2021) X-StereoLab stereo matching and stereo 3D object detection toolbox. https:\/\/github.com\/meteorshowers\/X-StereoLab"},{"key":"8261_CR22","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511811685","volume-title":"Multiple View Geometry in Computer Vision","author":"RI Hartley","year":"2004","unstructured":"Hartley RI, Zisserman A (2004) Multiple View Geometry in Computer Vision, 2nd edn. Cambridge University Press, Cambridge, U.K","edition":"2"},{"key":"8261_CR23","unstructured":"Quigley M (2009) Ros: an open-source robot operating system. In: IEEE International Conference on Robotics and Automation"},{"issue":"19","key":"8261_CR24","doi-asserted-by":"publisher","first-page":"6033","DOI":"10.3390\/s25196033","volume":"25","author":"H Qian","year":"2025","unstructured":"Qian H, Wang M, Zhu M, Wang H (2025) A review of multi-sensor fusion in autonomous driving. Sensors 25(19):6033. https:\/\/doi.org\/10.3390\/s25196033","journal-title":"Sensors"},{"issue":"2","key":"8261_CR25","doi-asserted-by":"publisher","first-page":"216","DOI":"10.1109\/JPROC.2004.840301","volume":"93","author":"M Frigo","year":"2005","unstructured":"Frigo M, Johnson SG (2005) The design and implementation of fftw3. Proc IEEE 93(2):216\u2013231. https:\/\/doi.org\/10.1109\/JPROC.2004.840301","journal-title":"Proc IEEE"},{"key":"8261_CR26","unstructured":"Bochkovskii A, Delaunoy A, Germain H, Santos M, Zhou Y, Richter SR, Koltun V (2024) Depth pro: Sharp monocular metric depth in less than a second. arXiv"},{"issue":"2","key":"8261_CR27","doi-asserted-by":"publisher","first-page":"328","DOI":"10.1109\/TPAMI.2007.1166","volume":"30","author":"H Hirschmuller","year":"2008","unstructured":"Hirschmuller H (2008) Stereo processing by semiglobal matching and mutual information. IEEE Trans Pattern Anal Mach Intell 30(2):328\u2013341. https:\/\/doi.org\/10.1109\/TPAMI.2007.1166","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"8261_CR28","doi-asserted-by":"publisher","unstructured":"Hamid U, Qamar RA, Waqas K (2014) Performance comparison of time-domain and frequency-domain beamforming techniques for sensor array processing. In: Proc. 11th Int. Bhurban Conf. Applied Sciences and Technology (IBCAST), Islamabad, Pakistan. https:\/\/doi.org\/10.1109\/IBCAST.2014.6778172","DOI":"10.1109\/IBCAST.2014.6778172"},{"issue":"3","key":"8261_CR29","doi-asserted-by":"publisher","first-page":"1778","DOI":"10.1109\/LRA.2017.2657002","volume":"2","author":"M Mancini","year":"2017","unstructured":"Mancini M, Costante G, Valigi P, Ciarfuglia TA, Delmerico J, Scaramuzza D (2017) Toward domain independence for learning-based monocular depth estimation. IEEE Robot Autom Lett 2(3):1778\u20131785. https:\/\/doi.org\/10.1109\/LRA.2017.2657002","journal-title":"IEEE Robot Autom Lett"},{"key":"8261_CR30","doi-asserted-by":"crossref","unstructured":"Geiger A, Lenz P, Urtasun R (2012) Are we ready for autonomous driving? the kitti vision benchmark suite. In: Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2012.6248074"},{"key":"8261_CR31","doi-asserted-by":"crossref","unstructured":"Menze M, Heipke C, Geiger A (2018) Object scene flow. ISPRS J Photogramm Remote Sens (JPRS)","DOI":"10.1016\/j.isprsjprs.2017.09.013"},{"key":"8261_CR32","doi-asserted-by":"publisher","unstructured":"Rudolph M, Dawoud Y, G\u00fcldenring R, Nalpantidis L, Belagiannis V (2022) Lightweight monocular depth estimation through guided decoding. In: 2022 International Conference on Robotics and Automation (ICRA), pp. 2344\u20132350. https:\/\/doi.org\/10.1109\/ICRA46639.2022.9812220","DOI":"10.1109\/ICRA46639.2022.9812220"},{"key":"8261_CR33","doi-asserted-by":"publisher","unstructured":"Wang Q, Zheng S, Yan Q, Deng F, Zhao K, Chu X (2021) Irs: A large naturalistic indoor robotics stereo dataset to train deep models for disparity and surface normal estimation. In: 2021 IEEE International Conference on Multimedia and Expo (ICME), pp. 1\u20136. https:\/\/doi.org\/10.1109\/ICME51207.2021.9428423","DOI":"10.1109\/ICME51207.2021.9428423"},{"key":"8261_CR34","doi-asserted-by":"crossref","unstructured":"Garofolo JS, Lamel LF, Fisher WM, Fiscus JG, Pallett DS, Dahlgren NL (1993) DARPA TIMIT Acoustic Phonetic Continuous Speech Corpus CDROM. NIST","DOI":"10.6028\/NIST.IR.4930"},{"issue":"10","key":"8261_CR35","doi-asserted-by":"publisher","DOI":"10.1063\/1.3021098","volume":"104","author":"YC Shiah","year":"2008","unstructured":"Shiah YC, Her H-C, Huang JH, Huang B (2008) Parametric analysis for a miniature loudspeaker used in cellular phones. J Appl Phys 104(10):104905. https:\/\/doi.org\/10.1063\/1.3021098","journal-title":"J Appl Phys"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-026-08261-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11227-026-08261-7","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-026-08261-7.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,2,5]],"date-time":"2026-02-05T18:26:05Z","timestamp":1770315965000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11227-026-08261-7"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2,5]]},"references-count":35,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2026,2]]}},"alternative-id":["8261"],"URL":"https:\/\/doi.org\/10.1007\/s11227-026-08261-7","relation":{},"ISSN":["1573-0484"],"issn-type":[{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2,5]]},"assertion":[{"value":"15 October 2025","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"19 January 2026","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"5 February 2026","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Conflict of interest"}},{"value":"Not applicable.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Ethical approval"}}],"article-number":"131"}}