{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,5,13]],"date-time":"2025-05-13T16:31:04Z","timestamp":1747153864414,"version":"3.40.5"},"reference-count":36,"publisher":"Springer Science and Business Media LLC","issue":"2","license":[{"start":{"date-parts":[[2023,3,8]],"date-time":"2023-03-08T00:00:00Z","timestamp":1678233600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,3,8]],"date-time":"2023-03-08T00:00:00Z","timestamp":1678233600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Mach. Intell. Res."],"published-print":{"date-parts":[[2023,4]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>Association, aiming to link bounding boxes of the same identity in a video sequence, is a central component in multi-object tracking (MOT). To train association modules, e.g., parametric networks, real video data are usually used. However, annotating person tracks in consecutive video frames is expensive, and such real data, due to its inflexibility, offer us limited opportunities to evaluate the system performance w.r.t. changing tracking scenarios. In this paper, we study whether 3D synthetic data can replace real-world videos for association training. Specifically, we introduce a large-scale synthetic data engine named MOTX, where the motion characteristics of cameras and objects are manually configured to be similar to those of real-world datasets. We show that, compared with real data, association knowledge obtained from synthetic data can achieve very similar performance on real-world test sets without domain adaption techniques. Our intriguing observation is credited to two factors. First and foremost, 3D engines can well simulate motion factors such as camera movement, camera view, and object movement so that the simulated videos can provide association modules with effective motion features. Second, the experimental results show that the appearance domain gap hardly harms the learning of association knowledge. In addition, the strong customization ability of MOTX allows us to quantitatively assess the impact of motion factors on MOT, which brings new insights to the community.<\/jats:p>","DOI":"10.1007\/s11633-022-1380-x","type":"journal-article","created":{"date-parts":[[2023,4,4]],"date-time":"2023-04-04T12:23:43Z","timestamp":1680611023000},"page":"194-206","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":1,"title":["A Study of Using Synthetic Data for Effective Association Knowledge Learning"],"prefix":"10.1007","volume":"20","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-9061-6180","authenticated-orcid":false,"given":"Yuchi","family":"Liu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4483-8783","authenticated-orcid":false,"given":"Zhongdao","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1526-0548","authenticated-orcid":false,"given":"Xiangxin","family":"Zhou","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-1464-9500","authenticated-orcid":false,"given":"Liang","family":"Zheng","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2023,3,8]]},"reference":[{"key":"1380_CR1","doi-asserted-by":"publisher","first-page":"6246","DOI":"10.1109\/CVPR42600.2020.00628","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"G Bras\u00f3","year":"2020","unstructured":"G. Bras\u00f3, L. Leal-Taix\u00e9. Learning a neural solver for multiple object tracking. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Seattle, USA, pp. 6246\u20136256, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00628."},{"key":"1380_CR2","doi-asserted-by":"publisher","first-page":"6786","DOI":"10.1109\/CVPR42600.2020.00682","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"Y H Xu","year":"2020","unstructured":"Y. H. Xu, A. \u015cep, Y. T. Ban, R. Horaud, L. Leal-Taix\u00e9, X. Alameda-Pineda. How to train your deep multi-object tracker. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Seattle, USA, pp. 6786\u20136795, 2020. DOI: https:\/\/doi.org\/10.1109\/CVPR42600.2020.00682."},{"key":"1380_CR3","unstructured":"L. Leal-Taix\u00e9, A. Milan, I. Reid, S. Roth, K. Schindler. MOTChallenge 2015: Towards a benchmark for multi-target tracking. [Online], Available: https:\/\/arxiv.org\/abs\/1504.01942, 2015."},{"key":"1380_CR4","unstructured":"A. Milan, L. Leal-Taix\u00e9, I. Reid, S. Roth, K. Schindler. MOT16: A benchmark for multi-object tracking. [Online], Available: https:\/\/arxiv.org\/abs\/1603.00831, 2016."},{"key":"1380_CR5","doi-asserted-by":"publisher","first-page":"193","DOI":"10.1007\/978-3-030-01261-8_12","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"S B\u0105k","year":"2018","unstructured":"S. B\u0105k, P. Carr, J. F. Lalonde. Domain adaptation through synthesis for unsupervised person re-identification. In Proceedings of the 15th European Conference on Computer Vision, Springer, Munich, Germany, pp. 193\u2013209, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01261-8_12."},{"key":"1380_CR6","unstructured":"H. Z. Dou, W. H. Zhang, P. Z. Zhang, Y. H. Zhao, S. Y. Li, Z. Q. Qin, F. Wu, L. Dong, X. Li. VersatileGait: A large-scale synthetic gait dataset with fine-grained attributes and complicated scenarios. [Online], Available: https:\/\/arxiv.org\/abs\/2101.01394, 2021."},{"key":"1380_CR7","unstructured":"Z. F. Xue, W. J. Mao, L. Zheng. Learning to simulate complex scenes. [Online], Available: https:\/\/arxiv.org\/abs\/2006.14611, 2020."},{"key":"1380_CR8","doi-asserted-by":"publisher","first-page":"775","DOI":"10.1007\/978-3-030-58539-6_46","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"Y Yao","year":"2020","unstructured":"Y. Yao, L. Zheng, X. D. Yang, M. Naphade, T. Gedeon. Simulating content consistent vehicle datasets with attribute descent. In Proceedings of the 16th European Conference on Computer Vision, Springer, Glasgow, UK, pp. 775\u2013791, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-030-58539-6_46."},{"key":"1380_CR9","doi-asserted-by":"publisher","unstructured":"J. H. Li, X. Gao, T. T. Jiang. Graph networks for multiple object tracking. In Proceedings of IEEE Winter Conference on Applications of Computer Vision, Snowmass, USA, pp. 708\u2013717, 2020. DOI: https:\/\/doi.org\/10.1109\/WACV45572.2020.9093347.","DOI":"10.1109\/WACV45572.2020.9093347"},{"key":"1380_CR10","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1007\/978-3-030-58621-8_7","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"Z D Wang","year":"2020","unstructured":"Z. D. Wang, L. Zheng, Y. X. Liu, Y. L. Li, S. J. Wang. Towards real-time multi-object tracking. In Proceedings of the 16th European Conference on Computer Vision, Springer, Glasgow, UK, pp. 107\u2013122, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-030-58621-8_7."},{"key":"1380_CR11","doi-asserted-by":"publisher","unstructured":"N. Wojke, A. Bewley, D. Paulus. Simple online and real-time tracking with a deep association metric. In Proceedings of IEEE International Conference on Image Processing, Beijing, China, pp. 3645\u20133649, 2017. DOI: https:\/\/doi.org\/10.1109\/ICIP.2017.8296962.","DOI":"10.1109\/ICIP.2017.8296962"},{"key":"1380_CR12","unstructured":"Y. F. Zhan, C. Y. Wang, X. G. Wang, W. J. Zeng, W. Y. Liu. A simple baseline for multi-object tracking. [Online], Available: https:\/\/arxiv.org\/abs\/2004.01888v1, 2020."},{"key":"1380_CR13","doi-asserted-by":"publisher","first-page":"1809","DOI":"10.1109\/ICPR.2018.8545450","volume-title":"Proceedings of the 24th International Conference on Pattern Recognition","author":"Z W Zhou","year":"2018","unstructured":"Z. W. Zhou, J. L. Xing, M. D. Zhang, W. M. Hu. Online multi-target tracking with tensor-based high-order graph matching. In Proceedings of the 24th International Conference on Pattern Recognition, IEEE, Beijing, China, pp. 1809\u20131814, 2018. DOI: https:\/\/doi.org\/10.1109\/ICPR.2018.8545450."},{"key":"1380_CR14","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1007\/978-3-030-01228-1_23","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"J Zhu","year":"2018","unstructured":"J. Zhu, H. Yang, N. Liu, M. Kim, W. J. Zhang, M. H. Yang. Online multi-object tracking with dual matching attention networks. In Proceedings of the 15th European Conference on Computer Vision, Springer, Munich, Germany, pp. 379\u2013396, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01228-1_23."},{"issue":"1","key":"1380_CR15","doi-asserted-by":"publisher","first-page":"31","DOI":"10.1007\/s11633-010-0031-9","volume":"7","author":"Q C Wang","year":"2010","unstructured":"Q. C. Wang, Y. H. Gong, C. H. Yang, C. H. Li. Robust object tracking under appearance change conditions. International Journal of Automation and Computing, vol. 7, no. 1, pp. 31\u201338, 2010. DOI: https:\/\/doi.org\/10.1007\/s11633-010-0031-9.","journal-title":"International Journal of Automation and Computing"},{"issue":"1\u20132","key":"1380_CR16","doi-asserted-by":"publisher","first-page":"83","DOI":"10.1002\/nav.3800020109","volume":"2","author":"H W Kuhn","year":"1955","unstructured":"H. W. Kuhn. The Hungarian method for the assignment problem. Naval Research Logistics Quarterly, vol. 2, no. 1\u20132, pp. 83\u201397, 1955. DOI: https:\/\/doi.org\/10.1002\/nav.3800020109.","journal-title":"Naval Research Logistics Quarterly"},{"key":"1380_CR17","volume-title":"An Introduction to the Kalman Filter","author":"G Welch","year":"1995","unstructured":"G. Welch, G. Bishop. An Introduction to the Kalman Filter. University of North Carolina at Chapel Hill, Chapel Hill, USA, 1995."},{"key":"1380_CR18","unstructured":"I. Papakis, A. Sarkar, A. Karpatne. GCNNMatch: Graph convolutional neural networks for multi-object tracking via Sinkhorn normalization. [Online], Available: https:\/\/arxiv.org\/abs\/2010.00067, 2020."},{"key":"1380_CR19","unstructured":"X. C. Peng, B. Usman, N. Kaushik, J. Hoffman, D. Q. Wang, K. Saenko. VisDA: The visual domain adaptation challenge. [Online], Available: https:\/\/arxiv.org\/abs\/1710.06924, 2017."},{"key":"1380_CR20","doi-asserted-by":"publisher","first-page":"2021","DOI":"10.1109\/CVPRW.2018.00271","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"X C Peng","year":"2018","unstructured":"X. C. Peng, B. Usman, N. Kaushik, D. Q. Wang, J. Hoffman, K. Saenko. VisDA: A synthetic-to-real benchmark for visual domain adaptation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops, IEEE, Salt Lake City, USA, pp. 2021\u20132026, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPRW.2018.00271."},{"key":"1380_CR21","unstructured":"Y. Cabon, N. Murray, M. Humenberger. Virtual KITTI 2. [Online], Available: https:\/\/arxiv.org\/abs\/2001.10773, 2020."},{"key":"1380_CR22","doi-asserted-by":"publisher","unstructured":"A. Gaidon, Q. Wang, Y. Cabon, E. Vig. Virtual Worlds as proxy for multi-object tracking analysis. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, USA, pp. 4340\u20134349, 2016. DOI: https:\/\/doi.org\/10.1109\/CVPR.2016.470.","DOI":"10.1109\/CVPR.2016.470"},{"key":"1380_CR23","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1007\/978-3-030-58571-6_1","volume-title":"Proceedings of the 16th European Conference on Computer Vision","author":"Y Z Hou","year":"2020","unstructured":"Y. Z. Hou, L. Zheng, S. Gould. Multiview detection with feature perspective transformation. In Proceedings of the 16th European Conference on Computer Vision, Springer, Glasgow, UK, pp. 1\u201318, 2020. DOI: https:\/\/doi.org\/10.1007\/978-3-030-58571-6_1."},{"key":"1380_CR24","doi-asserted-by":"publisher","first-page":"450","DOI":"10.1007\/978-3-030-01225-0_27","volume-title":"Proceedings of the 15th European Conference on Computer Vision","author":"M Fabbri","year":"2018","unstructured":"M. Fabbri, F. Lanzi, S. Calderara, A. Palazzi, R. Vezzani, R. Cucchiara. Learning to detect and track visible and occluded body joints in a virtual world. In Proceedings of the 15th European Conference on Computer Vision, Springer, Munich, Germany, pp. 450\u2013456, 2018. DOI: https:\/\/doi.org\/10.1007\/978-3-030-01225-0_27."},{"key":"1380_CR25","doi-asserted-by":"publisher","first-page":"3752","DOI":"10.1109\/CVPR.2018.00395","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"S Sankaranarayanan","year":"2018","unstructured":"S. Sankaranarayanan, Y. Balaji, A. Jain, S. Nam Lim, R. Chellappa. Learning from synthetic data: Addressing domain shift for semantic segmentation. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, USA, pp. 3752\u20133761, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00395."},{"key":"1380_CR26","unstructured":"C. Doersch, A. Zisserman. Sim2real transfer learning for 3D human pose estimation: Motion to the rescue. In Proceedings of the 33rd Conference on Neural Information Processing Systems, Vancouver, Canada, 2019."},{"key":"1380_CR27","unstructured":"E. Kolve, R. Mottaghi, W. Han, E. VanderBilt, L. Weihs, A. Herrasti, M. Deitke, K. Ehsani, D. Gordon, Y. K. Zhu, A. Kembhavi, A. Gupta, A. Farhadi. AI2-THOR: An interactive 3D environment for visual AI. [Online], Available: https:\/\/arxiv.org\/abs\/1712.05474, 2017."},{"key":"1380_CR28","doi-asserted-by":"publisher","first-page":"4550","DOI":"10.1109\/ICCV.2019.00465","volume-title":"Proceedings of IEEE\/CVF International Conference on Computer Vision","author":"A Kar","year":"2019","unstructured":"A. Kar, A. Prakash, M. Y. Liu, E. Cameracci, J. Yuan, M. Rusiniak, D. Acuna, A. Torralba, S. Fidler. Meta-Sim: Learning to generate synthetic datasets. In Proceedings of IEEE\/CVF International Conference on Computer Vision, IEEE, Seoul, Republic of Korea, pp. 4550\u20134559, 2019. DOI: https:\/\/doi.org\/10.1109\/ICCV.2019.00465."},{"key":"1380_CR29","unstructured":"A. Juliani, V. P. Berges, E. Teng, A. Cohen, J. Harper, C. Elion, C. Goy, Y. Gao, H. Henry, M. Mattar, D. Lange. Unity: A general platform for intelligent agents. [Online], Available: https:\/\/arxiv.org\/abs\/1809.02627, 2018."},{"key":"1380_CR30","doi-asserted-by":"publisher","first-page":"608","DOI":"10.1109\/CVPR.2019.00070","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"X X Sun","year":"2019","unstructured":"X. X. Sun, L. Zheng. Dissecting person re-identification from the viewpoint of viewpoint. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Long Beach, USA, pp. 608\u2013617, 2019. DOI: https:\/\/doi.org\/10.1109\/CVPR.2019.00070."},{"issue":"1","key":"1380_CR31","doi-asserted-by":"publisher","first-page":"61","DOI":"10.1109\/TNN.2008.2005605","volume":"20","author":"F Scarselli","year":"2009","unstructured":"F. Scarselli, M. Gori, A. C. Tsoi, M. Hagenbuchner, G. Monfardini. The graph neural network model. IEEE Transactions on Neural Networks, vol. 20, no. 1, pp. 61\u201380, 2009. DOI: https:\/\/doi.org\/10.1109\/TNN.2008.2005605.","journal-title":"IEEE Transactions on Neural Networks"},{"key":"1380_CR32","doi-asserted-by":"publisher","unstructured":"A. Bewley, Z. Y. Ge, L. Ott, F. Ramos, B. Upcroft. Simple online and realtime tracking. In Proceedings of IEEE International Conference on Image Processing, Phoenix, USA, pp. 3464\u20133468, 2016. DOI: https:\/\/doi.org\/10.1109\/ICIP.2016.7533003.","DOI":"10.1109\/ICIP.2016.7533003"},{"key":"1380_CR33","unstructured":"Y. H. Du, Y. Song, B. Yang, Y. Y. Zhao. StrongSORT: Make deepSORT great again. [Online], Available: https:\/\/arxiv.org\/abs\/2202.13514, 2022."},{"key":"1380_CR34","doi-asserted-by":"crossref","unstructured":"K. Bernardin, R. Stiefelhagen. Evaluating multiple object tracking performance: The clear mot metrics. EURASIP Journal on Image and Video Processing, vol. 2008, Article number 246309, 2008.","DOI":"10.1155\/2008\/246309"},{"key":"1380_CR35","unstructured":"P. Dendorfer, H. Rezatofighi, A. Milan, J. Shi, D. Cremers, I. Reid, S. Roth, K. Schindler, L. Leal-Taix\u00e9. MOT20: A benchmark for multi object tracking in crowded scenes. [Online], Available: https:\/\/arxiv.org\/abs\/2003.09003. 2020."},{"key":"1380_CR36","doi-asserted-by":"publisher","first-page":"994","DOI":"10.1109\/CVPR.2018.00110","volume-title":"Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition","author":"W J Deng","year":"2018","unstructured":"W. J. Deng, L. Zheng, Q. X. Ye, G. L. Kang, Y. Yang, J. B. Jiao. Image-image domain adaptation with preserved self-similarity and domain-dissimilarity for person re-identification. In Proceedings of IEEE\/CVF Conference on Computer Vision and Pattern Recognition, IEEE, Salt Lake City, USA, pp. 994\u20131003, 2018. DOI: https:\/\/doi.org\/10.1109\/CVPR.2018.00110."}],"container-title":["Machine Intelligence Research"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-022-1380-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11633-022-1380-x\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11633-022-1380-x.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,7,11]],"date-time":"2024-07-11T06:19:58Z","timestamp":1720678798000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11633-022-1380-x"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,3,8]]},"references-count":36,"journal-issue":{"issue":"2","published-print":{"date-parts":[[2023,4]]}},"alternative-id":["1380"],"URL":"https:\/\/doi.org\/10.1007\/s11633-022-1380-x","relation":{},"ISSN":["2731-538X","2731-5398"],"issn-type":[{"type":"print","value":"2731-538X"},{"type":"electronic","value":"2731-5398"}],"subject":[],"published":{"date-parts":[[2023,3,8]]},"assertion":[{"value":"19 July 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"13 October 2022","order":2,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"8 March 2023","order":3,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}