{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T14:30:58Z","timestamp":1779114658208,"version":"3.51.4"},"reference-count":0,"publisher":"Slovenian Association Informatika","issue":"13","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJCAI"],"abstract":"<jats:p>Under the dynamic condition of short video platforms, the shortfall of conventional recommendation algorithms that pay too much attention to short-term indicators at the cost of long-term user behavior is increasingly obvious. To compensate for it, we utilized a Deep Reinforcement Learning (DRL) approach to develop an intelligent recommendation system framework supported by deep feature engineering, policy updating, and online interaction. We effectively cast the difficult recommendation process into a Markov Decision Process (MDP) in order to improve the user experience by maximizing long-term user value. Experimental findings illustrate that, relative to baseline models like collaborative filtering (MF) and deep neural networks (DNN), our DRL agent possesses a remarkable lead over key long-term engagement indicators, specifically gaining an improvement of more than 22% in average session time. Besides, an ablation study of the reward function confirmed that both immediate and delayed signals are necessary for a composite reward architecture in order to learn a good policy. The findings of this work have repercussions for how short video recommendation intelligence can be boosted and even indicate a new research path for the recommender systems community, shifting away from using short-term metrics towards maximizing long-term user value.<\/jats:p>","DOI":"10.31449\/inf.v50i13.13064","type":"journal-article","created":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T13:54:38Z","timestamp":1779112478000},"source":"Crossref","is-referenced-by-count":0,"title":["Optimizing Long-Term User Engagement in Short-Video Recommendation via Reinforcement Learning: A Markov Decision Process Framework with Composite Rewards"],"prefix":"10.31449","volume":"50","author":[{"given":"Juan","family":"Di","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"16141","published-online":{"date-parts":[[2026,5,18]]},"container-title":["Informatica"],"original-title":[],"link":[{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/13064\/6715","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/13064\/6715","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,18]],"date-time":"2026-05-18T13:54:38Z","timestamp":1779112478000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/view\/13064"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,5,18]]},"references-count":0,"journal-issue":{"issue":"13","published-online":{"date-parts":[[2026,5,18]]}},"URL":"https:\/\/doi.org\/10.31449\/inf.v50i13.13064","relation":{},"ISSN":["1854-3871","0350-5596"],"issn-type":[{"value":"1854-3871","type":"electronic"},{"value":"0350-5596","type":"print"}],"subject":[],"published":{"date-parts":[[2026,5,18]]}}}