{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,28]],"date-time":"2026-08-28T01:08:21Z","timestamp":1787879301602,"version":"build-2784847793"},"reference-count":66,"publisher":"MDPI AG","issue":"3","license":[{"start":{"date-parts":[[2022,3,16]],"date-time":"2022-03-16T00:00:00Z","timestamp":1647388800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100000780","name":"European Union","doi-asserted-by":"publisher","award":["101031646"],"award-info":[{"award-number":["101031646"]}],"id":[{"id":"10.13039\/501100000780","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Future Internet"],"abstract":"<jats:p>In most Computer Vision applications, Deep Learning models achieve state-of-the-art performances. One drawback of Deep Learning is the large amount of data needed to train the models. Unfortunately, in many applications, data are difficult or expensive to collect. Data augmentation can alleviate the problem, generating new data from a smaller initial dataset. Geometric and color space image augmentation methods can increase accuracy of Deep Learning models but are often not enough. More advanced solutions are Domain Randomization methods or the use of simulation to artificially generate the missing data. Data augmentation algorithms are usually specifically designed for single images. Most recently, Deep Learning models have been applied to the analysis of video sequences. The aim of this paper is to perform an exhaustive study of the novel techniques of video data augmentation for Deep Learning models and to point out the future directions of the research on this topic.<\/jats:p>","DOI":"10.3390\/fi14030093","type":"journal-article","created":{"date-parts":[[2022,3,16]],"date-time":"2022-03-16T03:34:13Z","timestamp":1647401653000},"page":"93","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":49,"title":["Survey on Videos Data Augmentation for Deep Learning Models"],"prefix":"10.3390","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-9611-0655","authenticated-orcid":false,"given":"Nino","family":"Cauli","sequence":"first","affiliation":[{"name":"Department of Mathematics and Computer Science, University of Cagliari, Via Ospedale 72, 09124 Cagliari, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8646-6183","authenticated-orcid":false,"given":"Diego","family":"Reforgiato Recupero","sequence":"additional","affiliation":[{"name":"Department of Mathematics and Computer Science, University of Cagliari, Via Ospedale 72, 09124 Cagliari, Italy"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"1968","published-online":{"date-parts":[[2022,3,16]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"172231","DOI":"10.1109\/ACCESS.2019.2956508","article-title":"A survey on the new generation of deep learning in image processing","volume":"7","author":"Jiao","year":"2019","journal-title":"IEEE Access"},{"key":"ref_2","doi-asserted-by":"crossref","first-page":"3212","DOI":"10.1109\/TNNLS.2018.2876865","article-title":"Object detection with deep learning: A review","volume":"30","author":"Zhao","year":"2019","journal-title":"IEEE Trans. Neural Netw. Learn. Syst."},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., and Li, F.-F. (2009, January 20\u201325). Imagenet: A large-scale hierarchical image database. Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA.","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Liu, Z., Luo, P., Wang, X., and Tang, X. (2015, January 7\u201313). Deep Learning Face Attributes in the Wild. Proceedings of the International Conference on Computer Vision (ICCV), Santiago, Chile.","DOI":"10.1109\/ICCV.2015.425"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"1231","DOI":"10.1177\/0278364913491297","article-title":"Vision meets robotics: The kitti dataset","volume":"32","author":"Geiger","year":"2013","journal-title":"Int. J. Robot. Res."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Yu, F., Chen, H., Wang, X., Xian, W., Chen, Y., Liu, F., Madhavan, V., and Darrell, T. (2020, January 13\u201319). Bdd100k: A diverse driving dataset for heterogeneous multitask learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00271"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"1173","DOI":"10.1109\/TBME.2021.3117407","article-title":"Domain adaptation for medical image analysis: A survey","volume":"69","author":"Guan","year":"2021","journal-title":"IEEE Trans. Biomed. Eng."},{"key":"ref_8","unstructured":"Pereira, F., Burges, C.J.C., Bottou, L., and Weinberger, K.Q. (2012). ImageNet Classification with Deep Convolutional Neural Networks. Advances in Neural Information Processing Systems, Curran Associates, Inc.. Available online: https:\/\/proceedings.neurips.cc\/paper\/2012\/file\/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf."},{"key":"ref_9","unstructured":"Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Generative adversarial nets. Adv. Neural Inf. Process. Syst., 27, Available online: https:\/\/proceedings.neurips.cc\/paper\/2014\/file\/5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Karras, T., Laine, S., and Aila, T. (2019, January 15\u201320). A style-based generator architecture for generative adversarial networks. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA.","DOI":"10.1109\/CVPR.2019.00453"},{"key":"ref_11","unstructured":"Tremblay, J., To, T., Sundaralingam, B., Xiang, Y., Fox, D., and Birchfield, S. (2018). Deep object pose estimation for semantic robotic grasping of household objects. arXiv."},{"key":"ref_12","unstructured":"Technologies, U. (2022, February 14). Unity Homepage. Available online: https:\/\/unity.com\/."},{"key":"ref_13","unstructured":"Games, E. (2022, February 14). Unreal Engine Homepage. Available online: https:\/\/www.unrealengine.com\/en-US\/."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Tobin, J., Fong, R., Ray, A., Schneider, J., Zaremba, W., and Abbeel, P. (2017, January 24\u201328). Domain randomization for transferring deep neural networks from simulation to the real world. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8202133"},{"key":"ref_15","unstructured":"Simonyan, K., and Zisserman, A. (2014). Two-stream convolutional networks for action recognition in videos. arXiv."},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"221","DOI":"10.1109\/TPAMI.2012.59","article-title":"3D convolutional neural networks for human action recognition","volume":"35","author":"Ji","year":"2012","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_17","doi-asserted-by":"crossref","unstructured":"Yue-Hei Ng, J., Hausknecht, M., Vijayanarasimhan, S., Vinyals, O., Monga, R., and Toderici, G. (2015, January 7\u201312). Beyond short snippets: Deep networks for video classification. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7299101"},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Lee, N., Choi, W., Vernaza, P., Choy, C.B., Torr, P.H., and Chandraker, M. (2017, January 21\u201326). Desire: Distant future prediction in dynamic scenes with interacting agents. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.233"},{"key":"ref_19","doi-asserted-by":"crossref","first-page":"60","DOI":"10.1186\/s40537-019-0197-0","article-title":"A survey on image data augmentation for deep learning","volume":"6","author":"Shorten","year":"2019","journal-title":"J. Big Data"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Khalifa, N.E., Loey, M., and Mirjalili, S. (2021). A comprehensive survey of recent trends in deep learning for digital images augmentation. Artif. Intell. Rev., 1\u201327. Available online: https:\/\/link.springer.com\/article\/10.1007\/s10462-021-10066-4.","DOI":"10.1007\/s10462-021-10066-4"},{"key":"ref_21","doi-asserted-by":"crossref","first-page":"15503","DOI":"10.1007\/s00521-020-04748-3","article-title":"A survey on face data augmentation for the training of deep neural networks","volume":"32","author":"Wang","year":"2020","journal-title":"Neural Comput. Appl."},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"545","DOI":"10.1111\/1754-9485.13261","article-title":"A review of medical image data augmentation techniques for deep learning applications","volume":"65","author":"Chlap","year":"2021","journal-title":"J. Med. Imaging Radiat. Oncol."},{"key":"ref_23","unstructured":"Naveed, H. (2021). Survey: Image mixing and deleting for data augmentation. arXiv."},{"key":"ref_24","unstructured":"Scopus (2022, February 14). Scopus Homepage. Available online: https:\/\/www.scopus.com\/."},{"key":"ref_25","doi-asserted-by":"crossref","unstructured":"Charalambous, C., and Bharath, A. (2016, January 19\u201322). A data augmentation methodology for training machine\/deep learning gait recognition algorithms. Proceedings of the British Machine Vision Conference 2016, BMVC 2016, York, UK.","DOI":"10.5244\/C.30.110"},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1016\/j.patrec.2017.04.004","article-title":"Three-stream CNNs for action recognition","volume":"92","author":"Wang","year":"2017","journal-title":"Pattern Recognit. Lett."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"De Souza, C., Gaidon, A., Cabon, Y., and L\u00f3pez, A. (2017, January 21\u201326). Procedural generation of videos to train deep action recognition networks. Proceedings of the 30th IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2017, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.278"},{"key":"ref_28","doi-asserted-by":"crossref","first-page":"38","DOI":"10.1109\/TIP.2017.2754941","article-title":"Video Salient Object Detection via Fully Convolutional Networks","volume":"27","author":"Wang","year":"2018","journal-title":"IEEE Trans. Image Process."},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"286","DOI":"10.1016\/j.simpat.2018.03.003","article-title":"A system for the generation of synthetic Wide Area Aerial surveillance imagery","volume":"84","author":"Griffith","year":"2018","journal-title":"Simul. Model. Pract. Theory"},{"key":"ref_30","doi-asserted-by":"crossref","unstructured":"Lu, S.P., You, J., Ceulemans, B., Wang, M., and Munteanu, A. (2018, January 7\u201310). Synthesis of Shaking Video Using Motion Capture Data and Dynamic 3D Scene Modeling. Proceedings of the International Conference on Image Processing, ICIP, Athens, Greece.","DOI":"10.1109\/ICIP.2018.8451475"},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Dong, J., Li, X., Xu, C., Yang, G., and Wang, X. (2018, January 22\u201326). Feature re-learning with data augmentation for content-based video recommendation. Proceedings of the MM 2018\u20142018 ACM Multimedia Conference, Seoul, Korea.","DOI":"10.1145\/3240508.3266441"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Angus, M., Elbalkini, M., Khan, S., Harakeh, A., Andrienko, O., Reading, C., Waslander, S., and Czarnecki, K. (2018, January 4\u20137). Unlimited Road-scene Synthetic Annotation (URSA) Dataset. Proceedings of the IEEE Conference on Intelligent Transportation Systems, ITSC, Maui, HI, USA.","DOI":"10.1109\/ITSC.2018.8569519"},{"key":"ref_33","doi-asserted-by":"crossref","first-page":"219","DOI":"10.1111\/cgf.13632","article-title":"Deep Video-Based Performance Cloning","volume":"38","author":"Aberman","year":"2019","journal-title":"Comput. Graph. Forum"},{"key":"ref_34","doi-asserted-by":"crossref","unstructured":"Rimboux, A., Dupre, R., Daci, E., Lagkas, T., Sarigiannidis, P., Remagnino, P., and Argyriou, V. (2019, January 29\u201331). Smart IoT cameras for crowd analysis based on augmentation for automatic pedestrian detection, simulation and annotation. Proceedings of the 15th Annual International Conference on Distributed Computing in Sensor Systems, DCOSS 2019, Santorini Island, Greece.","DOI":"10.1109\/DCOSS.2019.00070"},{"key":"ref_35","doi-asserted-by":"crossref","unstructured":"Fonder, M., and Van Droogenbroeck, M. (2019, January 16\u201317). Mid-air: A multi-modal dataset for extremely low altitude drone flights. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, Long Beach, CA, USA.","DOI":"10.1109\/CVPRW.2019.00081"},{"key":"ref_36","doi-asserted-by":"crossref","unstructured":"Wu, D., Chen, J., Sharma, N., Pan, S., Long, G., and Blumenstein, M. (2019, January 14\u201319). Adversarial Action Data Augmentation for Similar Gesture Action Recognition. Proceedings of the International Joint Conference on Neural Networks, Budapest, Hungary.","DOI":"10.1109\/IJCNN.2019.8851993"},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Sakkos, D., Shum, H., and Ho, E. (2019, January 26\u201328). Illumination-based data augmentation for robust background subtraction. Proceedings of the 2019 13th International Conference on Software, Knowledge, Information Management and Applications, SKIMA 2019, Island of Ulkulhas, Maldives.","DOI":"10.1109\/SKIMA47702.2019.8982527"},{"key":"ref_38","first-page":"490","article-title":"Dynamic hand gesture recognition using multi-direction 3D convolutional neural networks","volume":"27","author":"Li","year":"2019","journal-title":"Eng. Lett."},{"key":"ref_39","doi-asserted-by":"crossref","unstructured":"Sakkos, D., Ho, E., Shum, H., and Elvin, G. (2020). Image editing-based data augmentation for illumination-insensitive background subtraction. J. Enterp. Inf. Manag., Available online: https:\/\/www.emerald.com\/insight\/content\/doi\/10.1108\/JEIM-02-2020-0042\/full\/html.","DOI":"10.1108\/JEIM-02-2020-0042"},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Kwon, Y., Petrangeli, S., Kim, D., Wang, H., Park, E., Swaminathan, V., and Fuchs, H. (2020, January 23\u201328). Rotationally-Temporally Consistent Novel View Synthesis of Human Performance Video. Proceedings of the European Conference on Computer Vision, Glasgow, UK.","DOI":"10.1007\/978-3-030-58548-8_23"},{"key":"ref_41","unstructured":"Chai, L., Liu, Y., Liu, W., Han, G., and He, S. (2020). CrowdGAN: Identity-free Interactive Crowd Video Generation and Beyond. IEEE Trans. Pattern Anal. Mach. Intell., Available online: https:\/\/www.computer.org\/csdl\/journal\/tp\/5555\/01\/09286483\/1por0TYwZvG."},{"key":"ref_42","doi-asserted-by":"crossref","first-page":"1505","DOI":"10.1007\/s11263-019-01222-z","article-title":"Generating Human Action Videos by Coupling 3D Game Engines and Probabilistic Graphical Models","volume":"128","author":"Gaidon","year":"2020","journal-title":"Int. J. Comput. Vis."},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Namitha, K., Narayanan, A., and Geetha, M. (2020, January 10\u201312). A Synthetic Video Dataset Generation Toolbox for Surveillance Video Synopsis Applications. Proceedings of the 2020 IEEE International Conference on Communication and Signal Processing, ICCSP 2020, Nanjing, China.","DOI":"10.1109\/ICCSP48568.2020.9182084"},{"key":"ref_44","doi-asserted-by":"crossref","unstructured":"Isobe, T., Han, J., Zhuz, F., Liy, Y., and Wang, S. (2020, January 25\u201328). Intra-Clip Aggregation for Video Person Re-Identification. Proceedings of the International Conference on Image Processing, ICIP, Abu Dhabi, United Arab Emirates.","DOI":"10.1109\/ICIP40778.2020.9190839"},{"key":"ref_45","doi-asserted-by":"crossref","unstructured":"Zhang, Y., Jia, G., Chen, L., Zhang, M., and Yong, J. (2020, January 12\u201316). Self-Paced Video Data Augmentation by Generative Adversarial Networks with Insufficient Samples. Proceedings of the MM 2020\u201428th ACM International Conference on Multimedia, Seattle, WA, USA.","DOI":"10.1145\/3394171.3414003"},{"key":"ref_46","unstructured":"Yun, S., Oh, S.J., Heo, B., Han, D., and Kim, J. (2020). Videomix: Rethinking data augmentation for video classification. arXiv."},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Ye, Y., Yang, K., Xiang, K., Wang, J., and Wang, K. (2020, January 11\u201314). Universal semantic segmentation for fisheye urban driving images. Proceedings of the 2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Toronto, ON, Canada.","DOI":"10.1109\/SMC42975.2020.9283099"},{"key":"ref_48","doi-asserted-by":"crossref","first-page":"225","DOI":"10.1007\/s11263-020-01365-4","article-title":"Pixel-Wise Crowd Understanding via Synthetic Data","volume":"129","author":"Wang","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_49","unstructured":"Hwang, H., Jang, C., Park, G., Cho, J., and Kim, I. (2020). ElderSim: A Synthetic Data Generation Platform for Human Action Recognition in Eldercare Applications. arXiv."},{"key":"ref_50","unstructured":"Tsou, Y.Y., Lee, Y.A., and Hsu, C.T. (December, January 30). Multi-task Learning for Simultaneous Video Generation and Remote Photoplethysmography Estimation. Proceedings of the Asian Conference on Computer Vision, Kyoto, Japan."},{"key":"ref_51","doi-asserted-by":"crossref","first-page":"2457","DOI":"10.1109\/TMM.2020.3011290","article-title":"GAC-GAN: A General Method for Appearance-Controllable Human Video Motion Transfer","volume":"23","author":"Wei","year":"2021","journal-title":"IEEE Trans. Multimed."},{"key":"ref_52","doi-asserted-by":"crossref","unstructured":"Chen, Y., Rong, F., Duggal, S., Wang, S., Yan, X., Manivasagam, S., Xue, S., Yumer, E., and Urtasun, R. (2021, January 13\u201319). GeoSim: Realistic Video Simulation via Geometry-Aware Composition for Self-Driving. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR46437.2021.00715"},{"key":"ref_53","first-page":"1946","article-title":"Feature Re-Learning with Data Augmentation for Video Relevance Prediction","volume":"33","author":"Dong","year":"2021","journal-title":"IEEE Trans. Knowl. Data Eng."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Hu, L., Huang, S., Wang, S., Liu, W., and Ning, J. (2021, January 20\u201324). Do We Really Need Frame-by-Frame Annotation Datasets for Object Tracking?. Proceedings of the MM 2021\u201429th ACM International Conference on Multimedia, Chengdu, China.","DOI":"10.1145\/3474085.3475365"},{"key":"ref_55","doi-asserted-by":"crossref","first-page":"104187","DOI":"10.1016\/j.imavis.2021.104187","article-title":"Using synthetic data for person tracking under adverse weather conditions","volume":"111","author":"Kerim","year":"2021","journal-title":"Image Vis. Comput."},{"key":"ref_56","doi-asserted-by":"crossref","first-page":"2264","DOI":"10.1007\/s11263-021-01467-7","article-title":"Synthetic Humans for Action Recognition from Unseen Viewpoints","volume":"129","author":"Varol","year":"2021","journal-title":"Int. J. Comput. Vis."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Hu, Y.T., Wang, J., Yeh, R., and Schwing, A. (2021, January 19\u201325). SAIL-VOS 3D: A synthetic dataset and baselines for object detection and 3d mesh reconstruction from video data. Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition Workshops, Nashville, TN, USA.","DOI":"10.1109\/CVPRW53098.2021.00375"},{"key":"ref_58","doi-asserted-by":"crossref","unstructured":"Bongini, F., Berlincioni, L., Bertini, M., and Del Bimbo, A. (2021, January 20\u201324). Partially Fake it Till you Make It: Mixing Real and Fake Thermal Images for Improved Object Detection. Proceedings of the MM 2021\u201429th ACM International Conference on Multimedia, Chengdu, China.","DOI":"10.1145\/3474085.3475679"},{"key":"ref_59","doi-asserted-by":"crossref","first-page":"848","DOI":"10.1109\/TPAMI.2020.3002500","article-title":"Dynamic Facial Expression Generation on Hilbert Hypersphere with Conditional Wasserstein Generative Adversarial Nets","volume":"44","author":"Otberdout","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Zhang, H., Cisse, M., Dauphin, Y.N., and Lopez-Paz, D. (2017). mixup: Beyond empirical risk minimization. arXiv.","DOI":"10.1007\/978-1-4899-7687-1_79"},{"key":"ref_61","unstructured":"Yun, S., Han, D., Oh, S.J., Chun, S., Choe, J., and Yoo, Y. (November, January 27). Cutmix: Regularization strategy to train strong classifiers with localizable features. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Seoul, Korea."},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"Sadeghi, F., and Levine, S. (2016). Cad2rl: Real single-image flight without a single real image. arXiv.","DOI":"10.15607\/RSS.2017.XIII.034"},{"key":"ref_63","unstructured":"Blender (2022, February 14). Blender Homepage. Available online: https:\/\/www.blender.org\/."},{"key":"ref_64","unstructured":"Shi, X., Chen, Z., Wang, H., Yeung, D.Y., Wong, W.K., and Woo, W.C. (2015). Convolutional LSTM network: A machine learning approach for precipitation nowcasting. Adv. Neural Inf. Process. Syst., 28, Available online: https:\/\/proceedings.neurips.cc\/paper\/2015\/file\/07563a3fe3bbe7e3ba84431ad9d055af-Paper.pdf."},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Siam, M., Valipour, S., Jagersand, M., and Ray, N. (2017, January 17\u201320). Convolutional gated recurrent networks for video segmentation. Proceedings of the 2017 IEEE International Conference on Image Processing (ICIP), Beijing, China.","DOI":"10.1109\/ITSC.2017.8317600"},{"key":"ref_66","unstructured":"To, T., Tremblay, J., McKay, D., Yamaguchi, Y., Leung, K., Balanon, A., Cheng, J., Hodge, W., and Birchfield, S. (2022, February 14). NDDS: NVIDIA Deep Learning Dataset Synthesizer. Available online: https:\/\/github.com\/NVIDIA\/Dataset_Synthesizer."}],"container-title":["Future Internet"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/3\/93\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T22:37:15Z","timestamp":1760135835000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1999-5903\/14\/3\/93"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,3,16]]},"references-count":66,"journal-issue":{"issue":"3","published-online":{"date-parts":[[2022,3]]}},"alternative-id":["fi14030093"],"URL":"https:\/\/doi.org\/10.3390\/fi14030093","relation":{},"ISSN":["1999-5903"],"issn-type":[{"value":"1999-5903","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,3,16]]}}}