{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,8]],"date-time":"2026-06-08T12:02:53Z","timestamp":1780920173030,"version":"3.54.1"},"reference-count":49,"publisher":"Emerald","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2024,8,15]]},"abstract":"<jats:p>This monograph introduces various value-based approaches for solving the policy evaluation problem in the online reinforcement learning (RL) scenario, which aims to learn the value function associated with a specific policy under a single Markov decision process (MDP). Approaches vary depending on whether they are implemented in an on-policy or off-policy manner: In on-policy settings, where the evaluation of the policy is conducted using data generated from the same policy that is being assessed, classical techniques such as TD(0), TD(\u03bb), and their extensions with function approximation or variance reduction are employed in this setting. For off-policy evaluation, where samples are collected under a different behavior policy, this monograph introduces gradient-based two-timescale algorithms like GTD2, TDC, and variance-reduced TDC. These algorithms are designed to minimize the mean-squared projected Bellman error (MSPBE) as the objective function. This monograph also discusses their finite-sample convergence upper bounds and sample complexity.<\/jats:p>","DOI":"10.1561\/2400000045","type":"journal-article","created":{"date-parts":[[2024,7,11]],"date-time":"2024-07-11T04:53:59Z","timestamp":1720673639000},"page":"145-192","source":"Crossref","is-referenced-by-count":3,"title":["Stochastic Optimization Methods for Policy Evaluation in Reinforcement Learning"],"prefix":"10.1108","volume":"6","author":[{"given":"Yi","family":"Zhou","sequence":"first","affiliation":[{"name":"University of Utah"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shacong","family":"Ma","sequence":"additional","affiliation":[{"name":"University of Utah"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2024,8,15]]},"reference":[{"key":"2026033014424364000_ref001","article-title":"An actor-critic algorithm for sequence prediction","author":"Bahdanau"},{"key":"2026033014424364000_ref002","first-page":"30","article-title":"Residual algorithms: Reinforcement learning with function approximation","author":"Baird"},{"key":"2026033014424364000_ref003","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-75894-2","volume-title":"Adaptive Algorithms and Stochastic Approximations","author":"Benveniste","year":"1990"},{"key":"2026033014424364000_ref004","article-title":"A finite time analysis of temporal difference learning with linear function approximation","volume-title":"arXiv preprint","author":"Bhandari","year":"2018"},{"key":"2026033014424364000_ref005","volume-title":"Stochastic approximation: a dynamical systems viewpoint","author":"Borkar","year":"2009"},{"key":"2026033014424364000_ref006","first-page":"177","article-title":"Large-scale machine learning with stochastic gradient descent","author":"Bottou"},{"key":"2026033014424364000_ref007","volume-title":"Openai gym","author":"Brockman","year":"2016"},{"key":"2026033014424364000_ref008","doi-asserted-by":"crossref","article-title":"Teaching a machine to read maps with deep reinforcement learning","author":"Brunner","DOI":"10.1609\/aaai.v32i1.11645"},{"issue":"3-4","key":"2026033014424364000_ref009","doi-asserted-by":"crossref","first-page":"231","DOI":"10.1561\/2200000050","article-title":"Convex optimization: Algorithms and complexity","volume":"8","author":"Bubeck","year":"2015","journal-title":"Foundations and Trends\u00ae in Machine Learning"},{"key":"2026033014424364000_ref010","first-page":"1052","article-title":"Generative adversarial user model for reinforcement learning based recommendation system","author":"Chen"},{"key":"2026033014424364000_ref011","article-title":"Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning","author":"Dalal"},{"key":"2026033014424364000_ref012","first-page":"1","article-title":"Finite sample analysis of two-timescale stochastic approximation with applications to reinforcement learning","volume":"75","author":"Dalal","year":"2018","journal-title":"Proceedings of Machine Learning Research"},{"key":"2026033014424364000_ref013","first-page":"1646","article-title":"SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives","volume-title":"Proc. Advances in Neural Information Processing Systems (NeurIPS)","author":"Defazio","year":"2014"},{"key":"2026033014424364000_ref014","volume-title":"A Survey on Policy Search for Robotics","author":"Deisenroth","year":"2013"},{"key":"2026033014424364000_ref015","first-page":"5200","article-title":"Sgd: General analysis and improved rates","author":"Gower"},{"key":"2026033014424364000_ref016","first-page":"315","article-title":"Accelerating stochastic gradient descent using predictive variance reduction","volume-title":"Advances in neural information processing systems","author":"Johnson","year":"2013"},{"key":"2026033014424364000_ref017","volume-title":"Finite time analysis of linear two-timescale stochastic approximation with markovian noise","author":"Kaledin","year":"2020"},{"key":"2026033014424364000_ref018","first-page":"9","volume-title":"Reinforcement learning in robotics: A survey","author":"Kober Jens","year":"2014"},{"key":"2026033014424364000_ref019","first-page":"626","article-title":"On TD (0) with function approximation: Concentration bounds and a centered variant with exponential convergence","author":"Korda"},{"key":"2026033014424364000_ref020","article-title":"Imagenet classification with deep convolutional neural networks","volume":"25","author":"Krizhevsky","year":"2012","journal-title":"Advances in neural information processing systems"},{"issue":"1","key":"2026033014424364000_ref021","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1002\/wics.57","article-title":"Stochastic approximation: A survey","volume":"2","author":"Kushner","year":"2010","journal-title":"Wiley Interdisciplinary Reviews: Computational Statistics"},{"key":"2026033014424364000_ref022","article-title":"A simpler approach to obtaining an O(1\/t) convergence rate for the projected stochastic subgradient method","volume-title":"arXiv preprint","author":"Lacoste-Julien","year":"2012"},{"key":"2026033014424364000_ref023","first-page":"1192","article-title":"Deep reinforcement learning for dialogue generation","author":"Li"},{"key":"2026033014424364000_ref024","first-page":"5679","article-title":"3DCNN-DQN-RNN: A deep reinforcement learning framework for semantic parsing of large-scale 3D point clouds","author":"Liu"},{"key":"2026033014424364000_ref025","first-page":"1296","article-title":"Data sampling affects the complexity of online sgd over dependent data","author":"Ma"},{"key":"2026033014424364000_ref026","article-title":"Greedy-gq with variance reduction: Finite-time analysis and improved complexity","author":"Ma"},{"key":"2026033014424364000_ref027","first-page":"14796","article-title":"Variance-reduced off-policy tdc learning: Non-asymptotic convergence analysis","volume":"33","author":"Ma","year":"2020","journal-title":"Advances in neural information processing systems"},{"key":"2026033014424364000_ref028","article-title":"Convergent temporal-difference learning with arbitrary smooth function approximation","volume":"22","author":"Maei","year":"2009","journal-title":"Advances in neural information processing systems"},{"key":"2026033014424364000_ref029","article-title":"Closing the gap between svrg and td-svrg with gradient splitting","volume-title":"arXiv preprint","author":"Mustafin","year":"2022"},{"key":"2026033014424364000_ref030","first-page":"2","volume-title":"Dynamic programming and optimal control","author":"P","year":"1995"},{"key":"2026033014424364000_ref031","doi-asserted-by":"crossref","DOI":"10.1002\/9780470316887","volume-title":"Markov Decision Processes","author":"Puterman","year":"1994"},{"key":"2026033014424364000_ref032","first-page":"1","article-title":"Deep reinforcement learning for delay-sensitive lte downlink scheduling","author":"Sharma"},{"key":"2026033014424364000_ref033","first-page":"2803","article-title":"Finite-time error bounds for linear stochastic approximation andtd learning","author":"Srikant"},{"key":"2026033014424364000_ref034","first-page":"2803","article-title":"Finite-time error bounds for linear stochastic approximation andtd learning","author":"Srikant"},{"key":"2026033014424364000_ref035","first-page":"19592","article-title":"Finite-time analysis of adaptive temporal difference learning with deep neural networks","volume":"35","author":"Sun","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"issue":"1","key":"2026033014424364000_ref036","doi-asserted-by":"crossref","first-page":"9","DOI":"10.1023\/A:1022633531479","article-title":"Learning to predict by the methods of temporal differences","volume":"3","author":"Sutton","year":"1988","journal-title":"Machine Learning"},{"key":"2026033014424364000_ref037","first-page":"993","article-title":"Fast gradient-descent methods for temporal-difference learning with linear function approximation","author":"Sutton"},{"issue":"21","key":"2026033014424364000_ref038","first-page":"1609","article-title":"A convergent o(n) algorithm for off-policy temporal-difference learning with linear function approximation","volume":"21","author":"Sutton","year":"2008","journal-title":"Proc. Advances in Neural Information Processing Systems (NIPS)"},{"key":"2026033014424364000_ref039","volume-title":"Reinforcement Learning: An Introduction","author":"Sutton","year":"2018"},{"issue":"3","key":"2026033014424364000_ref040","doi-asserted-by":"crossref","first-page":"241","DOI":"10.1023\/A:1007609817671","article-title":"On the convergence of temporal-difference learning with linear function approximation","volume":"42","author":"Tadic","year":"2001","journal-title":"Machine learning"},{"key":"2026033014424364000_ref041","article-title":"On the performance of temporal difference learning with neural networks","volume-title":"arXiv preprint","author":"Tian","year":"2023"},{"issue":"5","key":"2026033014424364000_ref042","doi-asserted-by":"crossref","first-page":"674","DOI":"10.1109\/9.580874","article-title":"An analysis of temporal-difference learning with function approximation","volume":"42","author":"Tsitsiklis","year":"1997","journal-title":"IEEE transactions on automatic control"},{"key":"2026033014424364000_ref043","article-title":"Attention is all you need","volume":"30","author":"Vaswani","year":"2017","journal-title":"Advances in neural information processing systems"},{"key":"2026033014424364000_ref044","volume-title":"Variance-reduced q-learning is minimax optimal","author":"Wainwright","year":"2019"},{"key":"2026033014424364000_ref045","first-page":"9747","article-title":"Non-asymptotic analysis for two time-scale tdc with general smooth function approximation","volume":"34","author":"Wang","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"2026033014424364000_ref046","first-page":"811","article-title":"Sample complexity bounds for two timescale value-based reinforcement learning algorithms","author":"Xu"},{"key":"2026033014424364000_ref047","article-title":"Reanalysis of variance reduced temporal difference learning","author":"Xu"},{"key":"2026033014424364000_ref048","first-page":"10633","article-title":"Two time-scale off-policy TD learning: Non-asymptotic analysis over Markovian samples","volume-title":"Proc. Advances in Neural Information Processing Systems (NeurIPS)","author":"Xu","year":"2019"},{"key":"2026033014424364000_ref049","first-page":"8665","article-title":"Finite-sample analysis for SARSA with linear function approximation","volume-title":"Advances in Neural Information Processing Systems","author":"Zou","year":"2019"}],"container-title":["Foundations and Trends\u00ae in Optimization"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftopt\/article-pdf\/6\/3\/145\/10976061\/2400000045en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftopt\/article-pdf\/6\/3\/145\/10976061\/2400000045en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T18:14:42Z","timestamp":1777486482000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftopt\/article\/6\/3\/145\/1324798\/Stochastic-Optimization-Methods-for-Policy"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,8,15]]},"references-count":49,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2024,8,15]]}},"URL":"https:\/\/doi.org\/10.1561\/2400000045","relation":{},"ISSN":["2167-3888","2167-3918"],"issn-type":[{"value":"2167-3888","type":"print"},{"value":"2167-3918","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,8,15]]}}}