{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,11,26]],"date-time":"2025-11-26T16:41:26Z","timestamp":1764175286323,"version":"3.37.3"},"reference-count":40,"publisher":"Springer Science and Business Media LLC","issue":"7","license":[{"start":{"date-parts":[[2022,9,15]],"date-time":"2022-09-15T00:00:00Z","timestamp":1663200000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,9,15]],"date-time":"2022-09-15T00:00:00Z","timestamp":1663200000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100004593","name":"Universidad Aut\u00f3noma de Madrid","doi-asserted-by":"crossref","id":[{"id":"10.13039\/501100004593","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Ministerio de Econom\u00eda, Industria y Competividad, Gobierno de Espa\u00f1a"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"published-print":{"date-parts":[[2023,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In this work we explore enhancing performance of reinforcement learning algorithms in video game environments by feeding it better, more relevant data. For this purpose, we use semantic segmentation to transform the images that would be used as input for the reinforcement learning algorithm from their original domain to a simplified semantic domain with just silhouettes and class labels instead of textures and colors, and then we train the reinforcement learning algorithm with these simplified images. We have conducted different experiments to study multiple aspects: feasibility of our proposal, and potential benefits to model generalization and transfer learning. Experiments have been performed with the Super Mario Bros video game as the testing environment. Our results show multiple advantages for this method. First, it proves that using semantic segmentation enables reaching higher performance than the baseline reinforcement learning algorithm without modifying the actual algorithm, and in fewer episodes; second, it shows noticeable performance improvements when training on multiple levels at the same time; and finally, it allows to apply transfer learning for models trained on visually different environments. We conclude that using semantic segmentation can certainly help reinforcement learning algorithms that work with visual data, by refining it. Our results also suggest that other computer vision techniques may also be beneficial for data prepossessing. <jats:italic>Models and code will be available on github upon acceptance.<\/jats:italic><\/jats:p>","DOI":"10.1007\/s11042-022-13695-1","type":"journal-article","created":{"date-parts":[[2022,9,15]],"date-time":"2022-09-15T17:02:56Z","timestamp":1663261376000},"page":"10961-10979","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":3,"title":["Exploiting semantic segmentation to boost reinforcement learning in video game environments"],"prefix":"10.1007","volume":"82","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-6610-1566","authenticated-orcid":false,"given":"Javier","family":"Montalvo","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"\u00c1lvaro","family":"Garc\u00eda-Mart\u00edn","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jes\u00fas","family":"Besc\u00f3s","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2022,9,15]]},"reference":[{"issue":"11","key":"13695_CR1","doi-asserted-by":"publisher","first-page":"3119","DOI":"10.1007\/s11263-021-01511-6","volume":"129","author":"H Blum","year":"2021","unstructured":"Blum H, Sarlin P-E, Nieto J, Siegwart R, Cadena C (2021) The fishyscapes benchmark: measuring blind spots in semantic segmentation. Int J Comput Vis 129(11):3119\u20133135","journal-title":"Int J Comput Vis"},{"key":"13695_CR2","unstructured":"Brockman G, Cheung V, Pettersson L, Schneider J, Schulman J, Tang J, Zaremba W (2016) OpenAI Gym"},{"key":"13695_CR3","doi-asserted-by":"crossref","unstructured":"Chen Y, Li W, Chen X, Gool LV (2019) Learning semantic segmentation from synthetic data: a geometrically guided input-output adaptation approach. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","DOI":"10.1109\/CVPR.2019.00194"},{"issue":"4","key":"13695_CR4","doi-asserted-by":"publisher","first-page":"834","DOI":"10.1109\/TPAMI.2017.2699184","volume":"40","author":"L-C Chen","year":"2017","unstructured":"Chen L-C, Papandreou G, Kokkinos I, Murphy K, Yuille AL (2017) Deeplab: semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs. IEEE Trans Pattern Anal Mach Intell 40(4):834\u2013848","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"13695_CR5","unstructured":"Chen L-C, Papandreou G, Schroff F, Adam H (2017) Rethinking Atrous convolution for semantic image segmentation"},{"key":"13695_CR6","unstructured":"Cobbe K, Klimov O, Hesse C, Kim T, Schulman J (2019) Quantifying generalization in reinforcement learning. In: International Conference on Machine Learning, pp 1282\u20131289. PMLR"},{"key":"13695_CR7","doi-asserted-by":"crossref","unstructured":"Cordts M, Omran M, Ramos S, Rehfeld T, Enzweiler M, Benenson R, Franke U, Roth S, Schiele B (2016) The cityscapes dataset for semantic urban scene understanding","DOI":"10.1109\/CVPR.2016.350"},{"key":"13695_CR8","doi-asserted-by":"crossref","unstructured":"Dabney W, Ostrovski G, Silver D, Munos R (2018) Implicit quantile networks for distributional reinforcement learning. In: International Conference on Machine Learning, pp 1096\u20131105. PMLR","DOI":"10.1609\/aaai.v32i1.11791"},{"key":"13695_CR9","doi-asserted-by":"publisher","unstructured":"Deschaud J-E, Duque D, Richa JP, Velasco-Forero S, Marcotegui B, Goulette F (2021) Paris-carla-3d: a real and synthetic outdoor point cloud dataset for challenging tasks in 3d mapping. Remote Sensing 13(22). https:\/\/doi.org\/10.3390\/rs13224713","DOI":"10.3390\/rs13224713"},{"key":"13695_CR10","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S et al (2020) An image is worth 16x16 words: Transformers for image recognition at scale. arXiv:http:\/\/arxiv.org\/abs\/2010.11929"},{"key":"13695_CR11","unstructured":"Dosovitskiy A, Ros G, Codevilla F, Lopez A, Koltun V (2017) CARLA: an open urban driving simulator. In: Proceedings of the 1st annual conference on robot learning, pp 1\u201316"},{"key":"13695_CR12","doi-asserted-by":"crossref","unstructured":"Fu J, Liu J, Tian H, Li Y, Bao Y, Fang Z, Lu H (2019) Dual attention network for scene segmentation. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 3146\u20133154","DOI":"10.1109\/CVPR.2019.00326"},{"key":"13695_CR13","first-page":"2613","volume":"23","author":"H Hasselt","year":"2010","unstructured":"Hasselt H (2010) Double q-learning. Advan Neural Inform Process Syst 23:2613\u20132621","journal-title":"Advan Neural Inform Process Syst"},{"key":"13695_CR14","doi-asserted-by":"crossref","unstructured":"Hu X, Yang K, Fei L, Wang K (2019) Acnet: attention based network to exploit complementary features for rgbd semantic segmentation. In: 2019 IEEE International Conference on Image Processing (ICIP), pp 1440\u20131444. IEEE","DOI":"10.1109\/ICIP.2019.8803025"},{"key":"13695_CR15","doi-asserted-by":"crossref","unstructured":"Huang Z, Wang X, Huang L, Huang C, Wei Y, Liu W (2019) Ccnet: criss-cross attention for semantic segmentation. In: Proceedings of the IEEE\/CVF international conference on computer vision (ICCV)","DOI":"10.1109\/ICCV.2019.00069"},{"key":"13695_CR16","doi-asserted-by":"crossref","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft coco: Common objects in context. In: European Conference on Computer Vision, pp 740\u2013755 . Springer","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"13695_CR17","doi-asserted-by":"crossref","unstructured":"Long J, Shelhamer E, Darrell T (2015) Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3431\u20133440","DOI":"10.1109\/CVPR.2015.7298965"},{"key":"13695_CR18","doi-asserted-by":"crossref","unstructured":"Lu X, Wang W, Danelljan M, Zhou T, Shen J, Gool LV (2020) Video object segmentation with episodic graph memory networks. In: European conference on computer vision, pp 661\u2013679. Springer","DOI":"10.1007\/978-3-030-58580-8_39"},{"key":"13695_CR19","doi-asserted-by":"crossref","unstructured":"Lu X, Wang W, Ma C, Shen J, Shao L, Porikli F (2019) See more, know more: unsupervised video object segmentation with co-attention siamese networks. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition (CVPR)","DOI":"10.1109\/CVPR.2019.00374"},{"key":"13695_CR20","doi-asserted-by":"crossref","unstructured":"Lu X, Wang W, Shen J, Crandall D, Luo J (2020) Zero-shot video object segmentation with co-attention siamese networks. IEEE transactions on pattern analysis and machine intelligence","DOI":"10.1109\/TPAMI.2020.3040258"},{"key":"13695_CR21","doi-asserted-by":"crossref","unstructured":"Lu X, Wang W, Shen J, Crandall D, Van Gool L (2021) Segmenting objects from relational visual data. IEEE transactions on pattern analysis and machine intelligence","DOI":"10.1109\/TPAMI.2021.3115815"},{"issue":"3","key":"13695_CR22","doi-asserted-by":"publisher","first-page":"473","DOI":"10.5194\/isprs-annals-III-3-473-2016","volume":"2016","author":"D Marmanis","year":"2016","unstructured":"Marmanis D, Wegner JD, Galliani S, Schindler K, Datcu M, Stilla U (2016) Semantic segmentation of aerial images with an ensemble of cnss. ISPRS Annal Photogrammetry, Remote Sens Spatial Inform Sci 2016(3):473\u2013480","journal-title":"ISPRS Annal Photogrammetry, Remote Sens Spatial Inform Sci"},{"key":"13695_CR23","unstructured":"Mnih V, Kavukcuoglu K, Silver D, Graves A, Antonoglou I, Wierstra D, Riedmiller M (2013) Playing atari with deep reinforcement learning"},{"issue":"7","key":"13695_CR24","doi-asserted-by":"publisher","first-page":"11201","DOI":"10.1007\/s11042-020-10248-2","volume":"80","author":"R Muthalagu","year":"2021","unstructured":"Muthalagu R, Bolimera A, Kalaichelvi V (2021) Vehicle lane markings segmentation and keypoint determination using deep convolutional neural networks. Multimed Tools Appl 80(7):11201\u201311215","journal-title":"Multimed Tools Appl"},{"key":"13695_CR25","unstructured":"Ng MH, Radia K, Chen J, Wang D, Gog I, Gonzalez JE (2020) Bev-seg: bird\u2019s eye view semantic segmentation using geometry and semantic point cloud. arXiv:http:\/\/arxiv.org\/abs\/2006.11436"},{"key":"13695_CR26","doi-asserted-by":"publisher","unstructured":"Niranjan DR, VinayKarthik BC (2021) Mohana: deep learning based object detection model for autonomous driving research using carla simulator. In: 2021 2nd international conference on smart electronics and communication (ICOSEC), pp 1251\u20131258. https:\/\/doi.org\/10.1109\/ICOSEC51865.2021.9591747https:\/\/doi.org\/10.1109\/ICOSEC51865.2021.9591747","DOI":"10.1109\/ICOSEC51865.2021.9591747 10.1109\/ICOSEC51865.2021.9591747"},{"key":"13695_CR27","unstructured":"Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, Killeen T, Lin Z, Gimelshein N, Antiga L, Desmaison A, Kopf A, Yang E, DeVito Z, Raison M, Tejani A, Chilamkurthy S, Steiner B, Fang L, Bai J, Chintala S (2019) Pytorch: an imperative style, high-performance deep learning library. In: Wallach H, Larochelle H, Beygelzimer A, d\u2019Alch\u00e9-buc F, Fox E, Garnett R (eds) Advances in neural information processing systems 32, pp 8024\u20138035"},{"key":"13695_CR28","doi-asserted-by":"crossref","unstructured":"Richter SR, Vineet V, Roth S, Koltun V (2016) Playing for data: ground truth from computer games. In: European Conference on Computer Vision, pp 102\u2013118. Springer","DOI":"10.1007\/978-3-319-46475-6_7"},{"key":"13695_CR29","doi-asserted-by":"crossref","unstructured":"Ros G, Sellart L, Materzynska J, Vazquez D, Lopez AM (2016) The synthia dataset: a large collection of synthetic images for semantic segmentation of urban scenes. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 3234\u20133243","DOI":"10.1109\/CVPR.2016.352"},{"issue":"7839","key":"13695_CR30","doi-asserted-by":"publisher","first-page":"604","DOI":"10.1038\/s41586-020-03051-4","volume":"588","author":"J Schrittwieser","year":"2020","unstructured":"Schrittwieser J, Antonoglou I, Hubert T, Simonyan K, Sifre L, Schmitt S, Guez A, Lockhart E, Hassabis D, Graepel T et al (2020) Mastering atari, go, chess and shogi by planning with a learned model. Nature 588 (7839):604\u2013609","journal-title":"Nature"},{"key":"13695_CR31","unstructured":"Schulman J, Wolski F, Dhariwal P, Radford A, Klimov O (2017) Proximal policy optimization algorithms. arXiv:http:\/\/arxiv.org\/abs\/1707.06347"},{"key":"13695_CR32","doi-asserted-by":"crossref","unstructured":"Sekkat AR, Dupuis Y, Vasseur P, Honeine P (2020) The omniscape dataset. In: 2020 IEEE International conference on robotics and automation (ICRA), pp 1603\u20131608. IEEE","DOI":"10.1109\/ICRA40945.2020.9197144"},{"issue":"26","key":"13695_CR33","doi-asserted-by":"publisher","first-page":"34203","DOI":"10.1007\/s11042-020-09840-3","volume":"80","author":"J Sun","year":"2021","unstructured":"Sun J, Li J, Liu L (2021) Semantic segmentation of brain tumor with nested residual attention networks. Multimed Tools Appl 80(26):34203\u201334220","journal-title":"Multimed Tools Appl"},{"key":"13695_CR34","doi-asserted-by":"crossref","unstructured":"van Hasselt H, Guez A, Silver D (2016) Deep reinforcement learning with double q-learning. In: Proceedings of the AAAI conference on artificial intelligence 30(1)","DOI":"10.1609\/aaai.v30i1.10295"},{"issue":"3","key":"13695_CR35","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1007\/BF00992698","volume":"8","author":"CJ Watkins","year":"1992","unstructured":"Watkins CJ, Dayan P (1992) Q-learning. Mach Learn 8(3):279\u2013292","journal-title":"Mach Learn"},{"key":"13695_CR36","unstructured":"Xie E, Wang W, Yu Z, Anandkumar A, Alvarez JM, Luo P (2021) Segformer: simple and efficient design for semantic segmentation with transformers. Adv Neural Inf Process Syst 34"},{"key":"13695_CR37","doi-asserted-by":"crossref","unstructured":"Yuan Y, Chen X, Wang J (2020) In: Vedaldi, A, Bischof, H, Brox, T, Frahm, J-M (eds.) Object-Contextual Representations for Semantic Segmentation, pp 173\u2013190. Springer, Cham","DOI":"10.1007\/978-3-030-58539-6_11"},{"key":"13695_CR38","unstructured":"Zhang H, Wu C, Zhang Z, Zhu Y, Lin H, Zhang Z, Sun Y, He T, Mueller J, Manmatha R et al (2020) Resnest: Split-attention networks. arXiv:http:\/\/arxiv.org\/abs\/2004.08955"},{"key":"13695_CR39","doi-asserted-by":"crossref","unstructured":"Zhang J, Yang K, Constantinescu A, Peng K, M\u00fcller K, Stiefelhagen R (2021) Trans4trans: efficient transformer for transparent object segmentation to help visually impaired people navigate in the real world. In: Proceedings of the IEEE\/CVF international conference on computer vision (ICCV) workshops, pp 1760\u20131770","DOI":"10.1109\/ICCVW54120.2021.00202"},{"key":"13695_CR40","doi-asserted-by":"crossref","unstructured":"Zheng S, Lu J, Zhao H, Zhu X, Luo Z, Wang Y, Fu Y, Feng J, Xiang T, Torr PH et al (2021) Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers. In: Proceedings of the IEEE\/CVF conference on computer vision and pattern recognition, pp 6881\u20136890","DOI":"10.1109\/CVPR46437.2021.00681"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-022-13695-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-022-13695-1\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-022-13695-1.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,3,2]],"date-time":"2023-03-02T16:35:13Z","timestamp":1677774913000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-022-13695-1"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,15]]},"references-count":40,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2023,3]]}},"alternative-id":["13695"],"URL":"https:\/\/doi.org\/10.1007\/s11042-022-13695-1","relation":{},"ISSN":["1380-7501","1573-7721"],"issn-type":[{"type":"print","value":"1380-7501"},{"type":"electronic","value":"1573-7721"}],"subject":[],"published":{"date-parts":[[2022,9,15]]},"assertion":[{"value":"29 March 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"29 May 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 August 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"15 September 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that no conflict of interest exists in this manuscript.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Competing interests"}}]}}