{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,9,8]],"date-time":"2025-09-08T05:36:09Z","timestamp":1757309769251},"publisher-location":"California","reference-count":0,"publisher":"International Joint Conferences on Artificial Intelligence Organization","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2022,7]]},"abstract":"<jats:p>Reward  is  the  driving  force  for  reinforcement-learning agents.  We here set out to understand the expressivity of Markov reward as a way to capture tasks that we would want an agent to perform.  We frame this study around three new abstract notions of \"task\":  (1) a set of acceptable behaviors, (2) a partial ordering over behaviors, or (3) a partial ordering over trajectories. Our main results prove that while reward can express many of these tasks, there exist instances  of  each  task  type  that  no  Markov reward  function  can  capture.   We  then  provide  a set of polynomial-time algorithms that construct a Markov reward function that allows an agent to perform each task type, and correctly determine when no such reward function exists.<\/jats:p>","DOI":"10.24963\/ijcai.2022\/730","type":"proceedings-article","created":{"date-parts":[[2022,7,16]],"date-time":"2022-07-16T02:55:56Z","timestamp":1657940156000},"page":"5254-5258","source":"Crossref","is-referenced-by-count":1,"title":["On the Expressivity of Markov Reward (Extended Abstract)"],"prefix":"10.24963","author":[{"given":"David","family":"Abel","sequence":"first","affiliation":[{"name":"DeepMind"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Will","family":"Dabney","sequence":"additional","affiliation":[{"name":"DeepMind"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Anna","family":"Harutyunyan","sequence":"additional","affiliation":[{"name":"DeepMind"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mark K.","family":"Ho","sequence":"additional","affiliation":[{"name":"Princeton University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Michael L.","family":"Littman","sequence":"additional","affiliation":[{"name":"Brown University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Doina","family":"Precup","sequence":"additional","affiliation":[{"name":"DeepMind"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Satinder","family":"Singh","sequence":"additional","affiliation":[{"name":"DeepMind"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"10584","event":{"number":"31","sponsor":["International Joint Conferences on Artificial Intelligence Organization (IJCAI)"],"acronym":"IJCAI-2022","name":"Thirty-First International Joint Conference on Artificial Intelligence {IJCAI-22}","start":{"date-parts":[[2022,7,23]]},"theme":"Artificial Intelligence","location":"Vienna, Austria","end":{"date-parts":[[2022,7,29]]}},"container-title":["Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence"],"original-title":[],"deposited":{"date-parts":[[2022,7,18]],"date-time":"2022-07-18T11:11:19Z","timestamp":1658142679000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.ijcai.org\/proceedings\/2022\/730"}},"subtitle":[],"proceedings-subject":"Artificial Intelligence Research Articles","short-title":[],"issued":{"date-parts":[[2022,7]]},"references-count":0,"URL":"https:\/\/doi.org\/10.24963\/ijcai.2022\/730","relation":{},"subject":[],"published":{"date-parts":[[2022,7]]}}}