{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,24]],"date-time":"2026-03-24T19:49:19Z","timestamp":1774381759444,"version":"3.50.1"},"reference-count":25,"publisher":"MDPI AG","issue":"14","license":[{"start":{"date-parts":[[2023,7,8]],"date-time":"2023-07-08T00:00:00Z","timestamp":1688774400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"Arcadyan Holding (BVI) Corp.","award":["E109-I04-023"],"award-info":[{"award-number":["E109-I04-023"]}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>In recent years, many things have been held via video conferences due to the impact of the COVID-19 epidemic around the world. A webcam will be used in conjunction with a computer and the Internet. However, the network camera cannot automatically turn and cannot lock the screen to the speaker. Therefore, this study uses the objection detector YOLO to capture the upper body of all people on the screen and judge whether each person opens or closes their mouth. At the same time, the Time Difference of Arrival (TDOA) is used to detect the angle of the sound source. Finally, the person\u2019s position obtained by YOLO is reversed to the person\u2019s position in the spatial coordinates through the distance between the person and the camera. Then, the spatial coordinates are used to calculate the angle between the person and the camera through inverse trigonometric functions. Finally, the angle obtained by the camera, and the angle of the sound source obtained by the microphone array, are matched for positioning. The experimental results show that the recall rate of positioning through YOLOX-Tiny reached 85.2%, and the recall rate of TDOA alone reached 88%. Integrating YOLOX-Tiny and TDOA for positioning, the recall rate reached 86.7%, the precision rate reached 100%, and the accuracy reached 94.5%. Therefore, the method proposed in this study can locate the speaker, and it has a better effect than using only one source.<\/jats:p>","DOI":"10.3390\/s23146250","type":"journal-article","created":{"date-parts":[[2023,7,10]],"date-time":"2023-07-10T01:02:50Z","timestamp":1688950970000},"page":"6250","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Automatic Speaker Positioning in Meetings Based on YOLO and TDOA"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-7716-7306","authenticated-orcid":false,"given":"Chen-Chiung","family":"Hsieh","sequence":"first","affiliation":[{"name":"Department of Computer Science and Engineering, Tatung University, Taipei City 104, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Men-Ru","family":"Lu","sequence":"additional","affiliation":[{"name":"Department of Computer Science and Engineering, Tatung University, Taipei City 104, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8289-5236","authenticated-orcid":false,"given":"Hsiao-Ting","family":"Tseng","sequence":"additional","affiliation":[{"name":"Department of Information Management, National Central University, Taoyuan City 320, Taiwan"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,7,8]]},"reference":[{"key":"ref_1","unstructured":"Li, J., Wang, J., and He, L. (2010, January 25\u201327). Design and implementation of a distributed video conference system based on SIP protocol. Proceedings of the International Conference on Computer Design and Applications, Qinhuangdao, China."},{"key":"ref_2","unstructured":"Sumra, H. (2019, August 02). How to Position Your Smart Camera for the Best Field of View?. Available online: https:\/\/www.ooma.com\/blog\/home-security\/how-to-position-your-smart-camera-for-the-best-field-of-view\/."},{"key":"ref_3","unstructured":"(2023, July 03). AXIS Communications. Positioning Cameras. Available online: https:\/\/www.axis.com\/products\/positioning-cameras."},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Mathur, G., Somwanshi, D., and Bundele, M.M. (2018, January 22\u201325). Intelligent video surveillance based on object tracking. In Proceeding of the 3rd International Conference and Workshops on Recent Advances and Innovations in Engineering (ICRAIE), Jaipur, India.","DOI":"10.1109\/ICRAIE.2018.8710421"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"3981","DOI":"10.1007\/s11042-020-09749-x","article-title":"Real time object detection and tracking system for video surveillance system","volume":"80","author":"Jha","year":"2021","journal-title":"Multimed. Tools Appl."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Ingle, P.Y., and Kim, Y.-G. (2022). Real-time abnormal object detection for video surveillance in smart cities. Sensors, 22.","DOI":"10.3390\/s22103862"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zheng, L., Zhou, T., Jiang, R., and Peng, Y. (2021, January 24\u201326). Survey of video object detection algorithms based on deep learning. Proceedings of the 2021 4th International Conference on Algorithms, Computing and Artificial Intelligence (ACAI \u201821), Sanya, China.","DOI":"10.1145\/3508546.3508622"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Girshick, R., Donahue, J., Darrell, T., and Malik, J. (2014, January 23\u201324). Rich feature hierarchies for accurate object detection and semantic segmentation. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.81"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Girshick, R. (2015, January 7\u201313). Fast R-CNN. Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile.","DOI":"10.1109\/ICCV.2015.169"},{"key":"ref_10","first-page":"1137","article-title":"Faster R-CNN: Towards real-time object detection with region proposal networks","volume":"39","author":"Ren","year":"2017","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_11","unstructured":"Howard, A.G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H. (2017). MobileNets: Efficient convolutional neural networks for mobile vision applications. arXiv, Available online: https:\/\/arxiv.org\/abs\/1704.04861."},{"key":"ref_12","unstructured":"Redmon, J., Divvala, S., Girshick, R., and Farhadi, A. (July, January 26). You Only Look Once: Unified, real-time object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA."},{"key":"ref_13","unstructured":"Aver Inc. (2022, July 13). AVer CAM550. Available online: https:\/\/reurl.cc\/41Z8DD."},{"key":"ref_14","unstructured":"B & H Foto & Electronics Corp (2023, July 02). OBSBOT Tiny 4K Overview. Available online: https:\/\/www.bhphotovideo.com\/c\/product\/1678153-REG\/obsbot_owb_2105_ce_tiny_4k_ai_powered_ptz.html\/overview."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Redmon, J., and Farhadi, A. (2017, January 21\u201326). YOLO9000: Better, faster, stronger. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.690"},{"key":"ref_16","unstructured":"Redmon, J., and Farhadi, A. (2018). YOLOv3: An incremental improvement. Computer Science. arXiv, Available online: http:\/\/arxiv.org\/abs\/1804.02767."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., and Sun, J. (2016, January 27\u201330). Deep residual learning for image recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.","DOI":"10.1109\/CVPR.2016.90"},{"key":"ref_18","unstructured":"Bochkovskiy, A., Wang, C., and Liao, H.M. (2020). YOLOv4: Optimal speed and accuracy of object detection. arXiv, Available online: https:\/\/arxiv.org\/abs\/2004.10934."},{"key":"ref_19","unstructured":"Ge, Z., Liu, S., Wang, F., Li, Z., and Sun, J. (2021). YOLOX: Exceeding YOLO series in 2021. arXiv."},{"key":"ref_20","unstructured":"Wang, C., Yeh, I., and Liao, H. (2021). You only learn one representation: Unified network for multiple tasks. arXiv, Available online: https:\/\/arxiv.org\/pdf\/2105.04206.pdf."},{"key":"ref_21","unstructured":"Van Veen, B.D., and Buckley, K.M. (2023, July 05). Beamforming: A versatile Approach to Spatial Filtering. IEEE, (April 1998). Available online: https:\/\/ieeexplore.ieee.org\/document\/665."},{"key":"ref_22","unstructured":"Uchiyama, T., Kawamura, A., Fujisaka, Y.I., Hiruma, N., and Iiguni, Y. (2017, January 6\u20138). Blind source separation with distance estimation based on variance of phase difference. Proceedings of the International Workshop on Smart Info-Media Systems in Asia (SISA 2017), Fukuoka, Japan."},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Ma, Q., Li, X., Li, G., Ning, B., Bai, M., and Wang, X. (2020). MRLIHT: Mobile RFID-based localization for indoor human tracking. Sensors, 20.","DOI":"10.3390\/s20061711"},{"key":"ref_24","unstructured":"Antsfeld, L., Chidlovskii, B., and Sansano-Sansano, E. (2020). Deep smartphone sensors-WiFi fusion for indoor positioning and tracking. arXiv, Available online: https:\/\/arxiv.org\/pdf\/2011.10799.pdf."},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"101360","DOI":"10.1016\/j.csl.2022.101360","article-title":"Deep learning based multi-source localization with source splitting and its effectiveness in multi-talker speech recognition","volume":"75","author":"Subramanian","year":"2022","journal-title":"Comput. Speech Lang."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/14\/6250\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T20:09:01Z","timestamp":1760126941000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/14\/6250"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,7,8]]},"references-count":25,"journal-issue":{"issue":"14","published-online":{"date-parts":[[2023,7]]}},"alternative-id":["s23146250"],"URL":"https:\/\/doi.org\/10.3390\/s23146250","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,7,8]]}}}