{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T14:30:36Z","timestamp":1779114636209,"version":"3.51.4"},"reference-count":0,"publisher":"Slovenian Association Informatika","issue":"13","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJCAI"],"abstract":"<jats:p>Artificial Intelligence (AI) algorithms such as Proximal Policy Optimization (PPO) help train agents for sequential decision-making tasks. Existing surveys already provide good coverage of the original PPO and its early developments but fell short of capturing the rapid evolution of PPO-based methods over the recent few years. Since 2023, a wave of algorithmic variants has emerged. These variants address different objectives, regularization approaches, exploration methods, training pipelines, and hybrid architectures across diverse applications. However, there has been no systematic effort to organize, compare, or critically assess these advances. This review addresses that research gap. This review analyzes 32 peer- reviewed studies (2023\u20132025) to evaluate 15+ PPO variants across six innovation categories and ten application domains. The study revealed that PPO delivers a typical performance improvement of 15\u2013 44% over baselines across metrics, including improved safety constraint satisfaction (+15%), computational efficiency (+18% SLA compliance), and sim-to-real transfer (+23% task success). The study analyzed advancements and developments by proposing a unified taxonomy focused on algorithmic advances and their performance in real-world scenarios. Three critical dimensions considered for evaluation are: generalization across tasks and environments, robustness and safety in deployment, and computational efficiency in training and inference. The review also identifies recurring limitations, inconsistent evaluation practices, and underexplored directions. It exposes gaps between simulation benchmarks and real-world deployment conditions, including operational constraints and challenges. By connecting theoretical improvements to empirical outcomes, this work serves as both a practical reference for engineers and researchers applying PPO today. The synthesized taxonomy provides a structured reference for analyzing recent PPO variants and their empirical trade-offs.<\/jats:p>","DOI":"10.31449\/inf.v50i13.12663","type":"journal-article","created":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T13:54:43Z","timestamp":1779112483000},"source":"Crossref","is-referenced-by-count":0,"title":["A Unified Taxonomy and Empirical Review of Recent Proximal Policy Optimization Variants and Their Real-World Applications"],"prefix":"10.31449","volume":"50","author":[{"given":"Vijaya Kittu","family":"Manda","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"16141","published-online":{"date-parts":[[2026,5,18]]},"container-title":["Informatica"],"original-title":[],"link":[{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/12663\/6710","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/12663\/6710","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T13:54:43Z","timestamp":1779112483000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/view\/12663"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,18]]},"references-count":0,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2026,5,18]]}},"URL":"https:\/\/doi.org\/10.31449\/inf.v50i13.12663","relation":{},"ISSN":["1854-3871","0350-5596"],"issn-type":[{"value":"1854-3871","type":"electronic"},{"value":"0350-5596","type":"print"}],"subject":[],"published":{"date-parts":[[2026,5,18]]}}}