{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,14]],"date-time":"2025-10-14T00:37:23Z","timestamp":1760402243152,"version":"build-2065373602"},"reference-count":34,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2021,4,2]],"date-time":"2021-04-02T00:00:00Z","timestamp":1617321600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100013058","name":"Jiangsu Provincial Key Research and Development Program","doi-asserted-by":"publisher","award":["BE2019311"],"award-info":[{"award-number":["BE2019311"]}],"id":[{"id":"10.13039\/501100013058","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100012431","name":"Jiangsu Agricultural Science and Technology Independent Innovation Fund","doi-asserted-by":"publisher","award":["CX(20)2013"],"award-info":[{"award-number":["CX(20)2013"]}],"id":[{"id":"10.13039\/501100012431","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Multiple-camera systems can expand coverage and mitigate occlusion problems. However, temporal synchronization remains a problem for budget cameras and capture devices. We propose an out-of-the-box framework to temporally synchronize multiple cameras using semantic human pose estimation from the videos. Human pose predictions are obtained with an out-of-the-shelf pose estimator for each camera. Our method firstly calibrates each pair of cameras by minimizing an energy function related to epipolar distances. We also propose a simple yet effective multiple-person association algorithm across cameras and a score-regularized energy function for improved performance. Secondly, we integrate the synchronized camera pairs into a graph and derive the optimal temporal displacement configuration for the multiple-camera system. We evaluate our method on four public benchmark datasets and demonstrate robust sub-frame synchronization accuracy on all of them.<\/jats:p>","DOI":"10.3390\/s21072464","type":"journal-article","created":{"date-parts":[[2021,4,2]],"date-time":"2021-04-02T10:34:09Z","timestamp":1617359649000},"page":"2464","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":5,"title":["Semantically Synchronizing Multiple-Camera Systems with Human Pose Estimation"],"prefix":"10.3390","volume":"21","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3632-988X","authenticated-orcid":false,"given":"Zhe","family":"Zhang","sequence":"first","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chunyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Microsoft Research Asia, Beijing 100089, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Wenhu","family":"Qin","sequence":"additional","affiliation":[{"name":"School of Instrument Science and Engineering, Southeast University, Nanjing 210096, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2021,4,2]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","unstructured":"Hou, Y., Zheng, L., and Gould, S. (2020, January 23\u201328). Multiview detection with feature perspective transformation. Proceedings of the 16th European Conference, Glasgow, UK.","DOI":"10.1007\/978-3-030-58571-6_1"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1023\/A:1021849801764","article-title":"M2Tracker: A multi-view approach to segmenting and tracking people in a cluttered scene","volume":"51","author":"Mittal","year":"2003","journal-title":"Int. J. Comput. Vis."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Fang, Z., V\u00e1zquez, D., and L\u00f3pez, A.M. (2017). On-board detection of pedestrian intentions. Sensors, 17.","DOI":"10.3390\/s17102193"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Wang, C., Qiu, W., Qin, W., and Zeng, W. (2020). AdaFuse: Adaptive Multiview Fusion for Accurate Human Pose Estimation in the Wild. Int. J. Comput. Vis., 1\u201316.","DOI":"10.1007\/s11263-020-01398-9"},{"key":"ref_5","doi-asserted-by":"crossref","unstructured":"Zhang, Z., Wang, C., Qin, W., and Zeng, W. (2020, January 14\u201319). Fusing Wearable IMUs with Multi-View Images for Human Pose Estimation: A Geometric Approach. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00227"},{"key":"ref_6","unstructured":"Qiu, H., Wang, C., Wang, J., Wang, N., and Zeng, W. Cross View Fusion for 3D Human Pose Estimation. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea."},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Tu, H., Wang, C., and Zeng, W. (2020, January 23\u201328). VoxelPose: Towards Multi-Camera 3D Human Pose Estimation in Wild Environment. Proceedings of the 16th European Conference, Glasgow, UK.","DOI":"10.1007\/978-3-030-58452-8_12"},{"key":"ref_8","doi-asserted-by":"crossref","unstructured":"Xie, R., Wang, C., and Wang, Y. (2020, January 13\u201319). Metafuse: A pre-trained fusion model for human pose estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.01370"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Liu, P., Zhang, Z., Meng, Z., and Gao, N. (2021). Monocular Depth Estimation with Joint Attention Feature Distillation and Wavelet-Based Loss Function. Sensors, 21.","DOI":"10.3390\/s21010054"},{"key":"ref_10","unstructured":"Saito, S., Huang, Z., Natsume, R., Morishima, S., Kanazawa, A., and Li, H. (27\u20132, January 27). Pifu: Pixel-aligned implicit function for high-resolution clothed human digitization. Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV), Seoul, Korea."},{"key":"ref_11","doi-asserted-by":"crossref","unstructured":"Chen, L., Ai, H., Chen, R., Zhuang, Z., and Liu, S. (2020, January 13\u201319). Cross-View Tracking for Multi-Human 3D Pose Estimation at over 100 FPS. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00334"},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"190","DOI":"10.1109\/TPAMI.2017.2782743","article-title":"Panoptic studio: A massively multiview system for social interaction capture","volume":"41","author":"Joo","year":"2019","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_13","unstructured":"Zhang, Z. (1999, January 20\u201327). Flexible camera calibration by viewing a plane from unknown orientations. Proceedings of the Seventh IEEE International Conference on Computer Vision, Kerkyra, Greece."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Shrestha, P., Weda, H., Barbieri, M., and Sekulovski, D. (2006, January 23\u201327). Synchronization of multiple video recordings based on still camera flashes. Proceedings of the 14th ACM International Conference on Multimedia, Santa Barbara, CA, USA.","DOI":"10.1145\/1180639.1180679"},{"key":"ref_15","unstructured":"Sinha, S.N., Pollefeys, M., and McMillan, L. (July, January 27). Camera network calibration from dynamic silhouettes. Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Washington, DC, USA."},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Takahashi, K., Mikami, D., Isogawa, M., and Kimata, H. (2018, January 18\u201322). Human pose as calibration pattern; 3D human pose estimation with multiple unsynchronized and uncalibrated cameras. Proceedings of the 2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Salt Lake City, UT, USA.","DOI":"10.1109\/CVPRW.2018.00230"},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Cheng, B., Xiao, B., Wang, J., Shi, H., Huang, T.S., and Zhang, L. (2020, January 13\u201319). HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose Estimation. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00543"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Doll\u00e1r, P., and Zitnick, C.L. (2014, January 6\u201312). Microsoft coco: Common objects in context. Proceedings of the 13th European Conference, Zurich, Switzerland.","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"1325","DOI":"10.1109\/TPAMI.2013.248","article-title":"Human3.6m: Large scale datasets and predictive methods for 3D human sensing in natural environments","volume":"36","author":"Ionescu","year":"2014","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Belagiannis, V., Amin, S., Andriluka, M., Schiele, B., Navab, N., and Ilic, S. (2014, January 23\u201328). 3D pictorial structures for multiple human pose estimation. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.216"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Shrstha, P., Barbieri, M., and Weda, H. (2007, January 24\u201329). Synchronization of multi-camera video recordings based on audio. Proceedings of the 15th ACM International Conference on Multimedia, Augsburg, Germany.","DOI":"10.1145\/1291233.1291367"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Hasler, N., Rosenhahn, B., Thormahlen, T., Wand, M., Gall, J., and Seidel, H.P. (2009, January 20\u201325). Markerless motion capture with unsynchronized moving cameras. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPRW.2009.5206859"},{"key":"ref_23","first-page":"51","article-title":"Reconstructing the 3D Trajectory of a Ball with Unsynchronized Cameras","volume":"14","author":"Tamaki","year":"2015","journal-title":"Int. J. Comput. Sci. Sport"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Andriluka, M., Pishchulin, L., Gehler, P., and Schiele, B. (2014, January 23\u201328). 2D Human Pose Estimation: New Benchmark and State of the Art Analysis. Proceedings of the 2014 IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA.","DOI":"10.1109\/CVPR.2014.471"},{"key":"ref_25","unstructured":"Wu, J., Zheng, H., Zhao, B., Li, Y., Yan, B., Liang, R., Wang, W., Zhou, S., Lin, G., and Fu, Y. (2017). Ai challenger: A large-scale dataset for going deeper in image understanding. arXiv."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Xiao, B., Wu, H., and Wei, Y. (2018, January 8\u201314). Simple baselines for human pose estimation and tracking. Proceedings of the 15th European Conference, Munich, Germany.","DOI":"10.1007\/978-3-030-01231-1_29"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Sun, K., Xiao, B., Liu, D., and Wang, J. (2019, January 15\u201320). Deep High-Resolution Representation Learning for Human Pose Estimation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00584"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Newell, A., Yang, K., and Deng, J. (2016, January 11\u201314). Stacked hourglass networks for human pose estimation. Proceedings of the14th European Conference, Amsterdam, The Netherlands.","DOI":"10.1007\/978-3-319-46484-8_29"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Sun, X., Xiao, B., Wei, F., Liang, S., and Wei, Y. (2018, January 8\u201314). Integral human pose regression. Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany.","DOI":"10.1007\/978-3-030-01231-1_33"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Cao, Z., Simon, T., Wei, S.E., and Sheikh, Y. (2017, January 21\u201326). Realtime multi-person 2d pose estimation using part affinity fields. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.143"},{"key":"ref_31","unstructured":"Newell, A., Huang, Z., and Deng, J. (2017, January 4\u20139). Associative embedding: End-to-end learning for joint detection and grouping. Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Hartley, R., and Zisserman, A. (2003). Multiple View Geometry in Computer Vision, Cambridge University Press.","DOI":"10.1017\/CBO9780511811685"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1090\/S0002-9939-1956-0078686-7","article-title":"On the shortest spanning subtree of a graph and the traveling salesman problem","volume":"7","author":"Kruskal","year":"1956","journal-title":"Proc. Am. Math. Soc."},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"35","DOI":"10.1115\/1.3662552","article-title":"A New Approach to Linear Filtering And Prediction Problems","volume":"82","author":"Kalman","year":"1960","journal-title":"ASME J. Basic Eng."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/7\/2464\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,13]],"date-time":"2025-10-13T13:33:37Z","timestamp":1760362417000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/21\/7\/2464"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,4,2]]},"references-count":34,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2021,4]]}},"alternative-id":["s21072464"],"URL":"https:\/\/doi.org\/10.3390\/s21072464","relation":{},"ISSN":["1424-8220"],"issn-type":[{"type":"electronic","value":"1424-8220"}],"subject":[],"published":{"date-parts":[[2021,4,2]]}}}