{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,22]],"date-time":"2026-05-22T15:05:34Z","timestamp":1779462334369,"version":"3.53.1"},"reference-count":41,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2026,2,1]],"date-time":"2026-02-01T00:00:00Z","timestamp":1769904000000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T00:00:00Z","timestamp":1774396800000},"content-version":"vor","delay-in-days":52,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Int J Softw Tools Technol Transfer"],"published-print":{"date-parts":[[2026,2]]},"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Agent behavior is shaped by latent decision parameters that govern how rewards are interpreted and traded off over time. A key example is the\n                    <jats:italic>discount factor<\/jats:italic>\n                    , which encodes time preference. Mis-specifying the discount factor can confound reward-centric behavioral models (e.g., inverse RL), motivating the need to infer time preference directly from behavior. This paper presents methods for\n                    <jats:italic>discount factor elicitation<\/jats:italic>\n                    in finite-state Markov Decision Processes via policy observations and controlled reward modifications. First, we introduce an algorithm that bounds the set of discount factors consistent with an agent\u2019s observed (near-)optimal policy, and show how observations across heterogeneous reward settings progressively tighten these bounds. Building on this result, we propose an\n                    <jats:italic>active elicitation framework<\/jats:italic>\n                    in which an ego agent strategically adjusts rewards (with fixed dynamics) to refine its estimate of another agent\u2019s discount factor. Through case studies, we demonstrate that active elicitation accelerates interval refinement relative to passive observation and enables\n                    <jats:italic>targeted exploration<\/jats:italic>\n                    in strategic multi-agent settings. Overall, our results establish reward modification as a principled mechanism for eliciting discount factors and improving behavioral modeling, prediction, and control.\n                  <\/jats:p>","DOI":"10.1007\/s10009-026-00851-3","type":"journal-article","created":{"date-parts":[[2026,3,25]],"date-time":"2026-03-25T13:19:58Z","timestamp":1774444798000},"page":"55-70","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Active discount factor elicitation via reward modification"],"prefix":"10.1007","volume":"28","author":[{"given":"Shadi","family":"Tasdighi\u00a0Kalat","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sriram","family":"Sankaranarayanan","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ashutosh","family":"Trivedi","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2026,3,25]]},"reference":[{"key":"851_CR1","doi-asserted-by":"crossref","unstructured":"Agranov, M., Kim, J., Yariv, L.: Coordination with differential time preferences: Experimental evidence. Tech. Rep., National Bureau of Economic Research (2023)","DOI":"10.3386\/w31288"},{"key":"851_CR2","first-page":"269","volume-title":"International Conference on Machine Learning","author":"R. Amit","year":"2020","unstructured":"Amit, R., Meir, R., Ciosek, K.: Discount factor as a regularizer in reinforcement learning. In: International Conference on Machine Learning, pp.\u00a0269\u2013278. PMLR (2020)"},{"issue":"3","key":"851_CR3","doi-asserted-by":"publisher","first-page":"583","DOI":"10.1111\/j.1468-0262.2008.00848.x","volume":"76","author":"S. Andersen","year":"2008","unstructured":"Andersen, S., Harrison, G.W., Lau, M.I., Rutstr\u00f6m, E.E.: Eliciting risk and time preferences. Econometrica 76(3), 583\u2013618 (2008)","journal-title":"Econometrica"},{"key":"851_CR4","doi-asserted-by":"crossref","unstructured":"Blackwell, D.: Discrete dynamic programming. Ann. Math. Stat., 719\u2013726 (1962)","DOI":"10.1214\/aoms\/1177704593"},{"key":"851_CR5","first-page":"239","volume-title":"AAAI\/IAAI","author":"C. Boutilier","year":"2002","unstructured":"Boutilier, C.: A pomdp formulation of preference elicitation problems. In: AAAI\/IAAI, Edmonton, AB, pp.\u00a0239\u2013246 (2002)"},{"issue":"6","key":"851_CR6","doi-asserted-by":"publisher","first-page":"1066","DOI":"10.1111\/j.1539-6924.2012.01907.x","volume":"33","author":"Y. Cao","year":"2013","unstructured":"Cao, Y., McGill, W.L.: Linkit: a ludic elicitation game for eliciting risk perceptions. Risk Anal. 33(6), 1066\u20131082 (2013)","journal-title":"Risk Anal."},{"issue":"2","key":"851_CR7","doi-asserted-by":"publisher","first-page":"571","DOI":"10.1016\/j.geb.2012.07.011","volume":"76","author":"B. Chen","year":"2012","unstructured":"Chen, B., Takahashi, S.: A folk theorem for repeated games with unequal discounting. Games Econ. Behav. 76(2), 571\u2013581 (2012)","journal-title":"Games Econ. Behav."},{"issue":"2","key":"851_CR8","first-page":"1","volume":"2","author":"Y. Chen","year":"2014","unstructured":"Chen, Y., Kash, I.A., Ruberry, M., Shnayder, V.: Eliciting predictions and recommendations for decision making. ACM Trans. Econ. Comput. 2(2), 1\u201327 (2014)","journal-title":"ACM Trans. Econ. Comput."},{"key":"851_CR9","doi-asserted-by":"publisher","first-page":"107","DOI":"10.1023\/A:1009986005690","volume":"2","author":"M. Coller","year":"1999","unstructured":"Coller, M., Williams, M.B.: Eliciting individual discount rates. Exp. Econ. 2, 107\u2013127 (1999)","journal-title":"Exp. Econ."},{"key":"851_CR10","doi-asserted-by":"publisher","first-page":"272","DOI":"10.1145\/800205.806346","volume-title":"Proceedings of the Third ACM Symposium on Symbolic and Algebraic Computation, SYMSAC\u201976","author":"G.E. Collins","year":"1976","unstructured":"Collins, G.E., Akritas, A.G.: Polynomial real root isolation using descarte\u2019s rule of signs. In: Proceedings of the Third ACM Symposium on Symbolic and Algebraic Computation, SYMSAC\u201976, pp.\u00a0272\u2013275. Association for Computing Machinery, New York (1976). https:\/\/doi.org\/10.1145\/800205.806346"},{"key":"851_CR11","volume-title":"Computational Methods of Linear Algebra","author":"D.K. Faddeev","year":"1963","unstructured":"Faddeev, D.K., Faddeeva, V.N., Williams, R.C.: Computational Methods of Linear Algebra (1963)"},{"key":"851_CR12","volume-title":"Game Theory for Data Science: Eliciting Truthful Information","author":"B. Faltings","year":"2022","unstructured":"Faltings, B., Radanovic, G.: Game Theory for Data Science: Eliciting Truthful Information. Springer, Berlin (2022)"},{"key":"851_CR13","volume-title":"Competitive Markov Decision Processes","author":"J. Filar","year":"2012","unstructured":"Filar, J., Vrieze, K.: Competitive Markov Decision Processes. Springer, Berlin (2012)"},{"key":"851_CR14","first-page":"1","volume":"43","author":"I. Fisher","year":"1930","unstructured":"Fisher, I.: The theory of interest. New York 43, 1\u201319 (1930)","journal-title":"New York"},{"key":"851_CR15","unstructured":"Fran\u00e7ois-Lavet, V., Fonteneau, R., Ernst, D.: How to discount deep reinforcement learning: Towards new dynamic strategies (2015). arXiv:1512.02011"},{"key":"851_CR16","doi-asserted-by":"publisher","first-page":"7786","DOI":"10.1109\/IROS51168.2021.9636479","volume-title":"2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"B.H. Giwa","year":"2021","unstructured":"Giwa, B.H., Lee, C.G.: A marginal log-likelihood approach for the estimation of discount factors of multiple experts in inverse reinforcement learning. In: 2021 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), pp.\u00a07786\u20137791 (2021). https:\/\/doi.org\/10.1109\/IROS51168.2021.9636479"},{"key":"851_CR17","unstructured":"Gurvich, V., Miltersen, P.B.: On the computational complexity of solving stochastic mean-payoff games. CoRR (2008). arXiv:0812.0486"},{"key":"851_CR18","series-title":"Proceedings of Machine Learning Research","first-page":"9072","volume-title":"Proceedings of the 39th International Conference on Machine Learning","author":"H. Hu","year":"2022","unstructured":"Hu, H., Yang, Y., Zhao, Q., Zhang, C.: On the role of discount factor in offline reinforcement learning. In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S. (eds.) Proceedings of the 39th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol.\u00a0162, pp.\u00a09072\u20139098 (2022). PMLR (17\u201323 Jul 2022)"},{"issue":"4","key":"851_CR19","doi-asserted-by":"publisher","first-page":"150","DOI":"10.1257\/mic.20140161","volume":"7","author":"M.O. Jackson","year":"2015","unstructured":"Jackson, M.O., Yariv, L.: Collective dynamic choice: the necessity of time inconsistency. Am. Econ. J. Microecon. 7(4), 150\u2013178 (2015)","journal-title":"Am. Econ. J. Microecon."},{"issue":"2","key":"851_CR20","doi-asserted-by":"publisher","first-page":"393","DOI":"10.1111\/1468-0262.00024","volume":"67","author":"E. Lehrer","year":"1999","unstructured":"Lehrer, E., Pauzner, A.: Repeated games with differential time preferences. Econometrica 67(2), 393\u2013412 (1999)","journal-title":"Econometrica"},{"issue":"5","key":"851_CR21","doi-asserted-by":"publisher","first-page":"587","DOI":"10.1016\/j.orl.2016.06.007","volume":"44","author":"E. Lehrer","year":"2016","unstructured":"Lehrer, E., Solan, E., Solan, O.N.: The value functions of Markov decision processes. Oper. Res. Lett. 44(5), 587\u2013591 (2016)","journal-title":"Oper. Res. Lett."},{"key":"851_CR22","unstructured":"Littman, M.L., Topcu, U., Fu, J., Isbell, C., Wen, M., MacGlashan, J.: Environment-independent task specifications via gltl (2017). arXiv:1704.04341"},{"issue":"9","key":"851_CR23","doi-asserted-by":"publisher","first-page":"1359","DOI":"10.1287\/mnsc.1050.0379","volume":"51","author":"N. Miller","year":"2005","unstructured":"Miller, N., Resnick, P., Zeckhauser, R.: Eliciting informative feedback: the peer-prediction method. Manag. Sci. 51(9), 1359\u20131373 (2005)","journal-title":"Manag. Sci."},{"key":"851_CR24","unstructured":"Mischel, W.: to willpower. The psychology of action: Linking cognition and motivation to behavior, p.\u00a0197 (1996)"},{"issue":"2","key":"851_CR25","doi-asserted-by":"publisher","DOI":"10.1037\/h0032198","volume":"21","author":"W. Mischel","year":"1972","unstructured":"Mischel, W., Ebbesen, E.B., Raskoff Zeiss, A.: Cognitive and attentional mechanisms in delay of gratification. J. Pers. Soc. Psychol. 21(2), 204 (1972)","journal-title":"J. Pers. Soc. Psychol."},{"key":"851_CR26","unstructured":"Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., Riedmiller, M.: Playing Atari with deep reinforcement learning (2013). arXiv:1312.5602"},{"key":"851_CR27","first-page":"2","volume-title":"Icml","author":"A.Y. Ng","year":"2000","unstructured":"Ng, A.Y., Russell, S., et al.: Algorithms for inverse reinforcement learning. In: Icml, vol.\u00a01, p.\u00a02 (2000)"},{"issue":"3","key":"851_CR28","doi-asserted-by":"publisher","first-page":"804","DOI":"10.1137\/S0363012996299557","volume":"37","author":"S.D. Patek","year":"1999","unstructured":"Patek, S.D., Bertsekas, D.P.: Stochastic shortest path games. SIAM J. Control Optim. 37(3), 804\u2013824 (1999)","journal-title":"SIAM J. Control Optim."},{"key":"851_CR29","volume-title":"Markov Decision Processes: Discrete Stochastic Dynamic Programming","author":"M.L. Puterman","year":"2014","unstructured":"Puterman, M.L.: Markov Decision Processes: Discrete Stochastic Dynamic Programming. Wiley, New York (2014)"},{"key":"851_CR30","series-title":"Algorithms and Computation in Mathematics","volume-title":"Algorithms in Real Algebraic Geometry","author":"M.F. Roy","year":"2006","unstructured":"Roy, M.F., Basu, S., Pollack, R.: Algorithms in Real Algebraic Geometry. Algorithms and Computation in Mathematics, vol.\u00a010 (2006)"},{"key":"851_CR31","doi-asserted-by":"publisher","first-page":"46","DOI":"10.1016\/j.jsc.2015.03.004","volume":"73","author":"M. Sagraloff","year":"2016","unstructured":"Sagraloff, M., Mehlhorn, K.: Computing real roots of real polynomials. J. Symb. Comput. 73, 46\u201386 (2016). https:\/\/www.sciencedirect.com\/science\/article\/pii\/S0747717115000292","journal-title":"J. Symb. Comput."},{"issue":"4","key":"851_CR32","doi-asserted-by":"publisher","first-page":"658","DOI":"10.1287\/opre.14.4.658","volume":"14","author":"R.D. Smallwood","year":"1966","unstructured":"Smallwood, R.D.: Optimum policy regions for Markov processes with discounting. Oper. Res. 14(4), 658\u2013669 (1966)","journal-title":"Oper. Res."},{"key":"851_CR33","volume-title":"Reinforcement Learning: An Introduction","author":"R.S. Sutton","year":"2018","unstructured":"Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. MIT Press, Cambridge (2018)"},{"key":"851_CR34","first-page":"322","volume-title":"International Conference on Quantitative Evaluation of Systems and Formal Modeling and Analysis of Timed Systems","author":"S. Tasdighi Kalat","year":"2024","unstructured":"Tasdighi Kalat, S., Sankaranarayanan, S., Trivedi, A.: What is your discount factor? In: International Conference on Quantitative Evaluation of Systems and Formal Modeling and Analysis of Timed Systems, pp.\u00a0322\u2013336. Springer, Berlin (2024)"},{"key":"851_CR35","unstructured":"Tessler, C.: Deep reinforcement learning works - now what? (2020). https:\/\/tesslerc.github.io\/posts\/drl_works_now_what\/"},{"key":"851_CR36","doi-asserted-by":"publisher","first-page":"2779","DOI":"10.1145\/3548606.3560554","volume-title":"Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security","author":"F. Tram\u00e8r","year":"2022","unstructured":"Tram\u00e8r, F., Shokri, R., San Joaquin, A., Le, H., Jagielski, M., Hong, S., Carlini, N.: Truth serum: poisoning machine learning models to reveal their secrets. In: Proceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security, pp.\u00a02779\u20132792 (2022)"},{"issue":"589","key":"851_CR37","doi-asserted-by":"publisher","first-page":"2116","DOI":"10.1111\/ecoj.12160","volume":"125","author":"S.T. Trautmann","year":"2015","unstructured":"Trautmann, S.T., van de Kuilen, G.: Belief elicitation: a horse race among truth serums. Econ. J. 125(589), 2116\u20132135 (2015)","journal-title":"Econ. J."},{"key":"851_CR38","doi-asserted-by":"publisher","first-page":"150","DOI":"10.1016\/j.jdeveco.2015.07.007","volume":"118","author":"D. Ubfal","year":"2016","unstructured":"Ubfal, D.: How general are time preferences? Eliciting good-specific discount rates. J. Dev. Econ. 118, 150\u2013170 (2016)","journal-title":"J. Dev. Econ."},{"issue":"7782","key":"851_CR39","doi-asserted-by":"publisher","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","volume":"575","author":"O. Vinyals","year":"2019","unstructured":"Vinyals, O., Babuschkin, I., Czarnecki, W.M., Mathieu, M., Dudzik, A., Chung, J., Choi, D.H., Powell, R., Ewalds, T., Georgiev, P., et al.: Grandmaster level in starcraft ii using multi-agent reinforcement learning. Nature 575(7782), 350\u2013354 (2019)","journal-title":"Nature"},{"key":"851_CR40","volume-title":"Mathematics for the Physical Sciences","author":"H.S. Wilf","year":"2013","unstructured":"Wilf, H.S.: Mathematics for the Physical Sciences. Courier Corporation (2013)"},{"issue":"16\u201317","key":"851_CR41","doi-asserted-by":"publisher","first-page":"1917","DOI":"10.1016\/j.artint.2008.08.005","volume":"172","author":"A. Zohar","year":"2008","unstructured":"Zohar, A., Rosenschein, J.S.: Mechanisms for information elicitation. Artif. Intell. 172(16\u201317), 1917\u20131939 (2008)","journal-title":"Artif. Intell."}],"container-title":["International Journal on Software Tools for Technology Transfer"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10009-026-00851-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s10009-026-00851-3","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s10009-026-00851-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,5,22]],"date-time":"2026-05-22T14:13:50Z","timestamp":1779459230000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s10009-026-00851-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,2]]},"references-count":41,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2026,2]]}},"alternative-id":["851"],"URL":"https:\/\/doi.org\/10.1007\/s10009-026-00851-3","relation":{},"ISSN":["1433-2779","1433-2787"],"issn-type":[{"value":"1433-2779","type":"print"},{"value":"1433-2787","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,2]]},"assertion":[{"value":"27 February 2026","order":1,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"25 March 2026","order":2,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}}]}}