{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,27]],"date-time":"2026-05-27T15:35:50Z","timestamp":1779896150189,"version":"3.53.1"},"reference-count":106,"publisher":"Emerald","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2013,12,19]]},"abstract":"<jats:p>A Markov Decision Process (MDP) is a natural framework for formulating sequential decision-making problems under uncertainty. In recent years, researchers have greatly advanced algorithms for learning and acting in MDPs. This article reviews such algorithms, beginning with well-known dynamic programming methods for solving MDPs such as policy iteration and value iteration, then describes approximate dynamic programming methods such as trajectory based value iteration, and finally moves to reinforcement learning methods such as Q-Learning, SARSA, and least-squares policy iteration. We describe algorithms in a unified framework, giving pseudocode together with memory and iteration complexity analysis for each. Empirical evaluations of these techniques with four representations across four domains, provide insight into how these algorithms perform with various feature sets in terms of running time and performance.<\/jats:p>","DOI":"10.1561\/2200000042","type":"journal-article","created":{"date-parts":[[2013,12,19]],"date-time":"2013-12-19T06:30:34Z","timestamp":1387434634000},"page":"375-451","source":"Crossref","is-referenced-by-count":74,"title":["A Tutorial on Linear Function Approximators for Dynamic Programming and Reinforcement Learning"],"prefix":"10.1108","volume":"6","author":[{"given":"Alborz","family":"Geramifard","sequence":"first","affiliation":[{"name":"MIT LIDS"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Thomas J.","family":"Walsh","sequence":"additional","affiliation":[{"name":"MIT LIDS"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Stefanie","family":"Tellex","sequence":"additional","affiliation":[{"name":"MIT CSAIL"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Girish","family":"Chowdhary","sequence":"additional","affiliation":[{"name":"MIT LIDS"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Nicholas","family":"Roy","sequence":"additional","affiliation":[{"name":"MIT CSAIL"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jonathan P.","family":"How","sequence":"additional","affiliation":[{"name":"MIT LIDS"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"140","published-online":{"date-parts":[[2013,12,19]]},"reference":[{"key":"2026033014111481400_ref001","unstructured":"RL competition\n          \n          http:\/\/www.rl-competition.org\/, 2012. Accessed: 20\/08\/2012."},{"key":"2026033014111481400_ref002","first-page":"1","volume-title":"International Joint Conference on Autonomous Agents and Multiagent Systems (AAMAS)","author":"Ahmadi","year":"2007"},{"key":"2026033014111481400_ref003","article-title":"Fitted Q-iteration in continuous action-space MDPs","volume-title":"Proceedings of Neural Information Processing Systems Conference (NIPS)","author":"Antos","year":"2007"},{"key":"2026033014111481400_ref004","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1007\/s10994-007-5038-2","article-title":"Learning near-optimal policies with bellman-residual minimization based fitted policy iteration and a single sample path","volume":"71","author":"Antos","year":"2008","journal-title":"Machine Learning"},{"key":"2026033014111481400_ref005","first-page":"19","volume-title":"International Conference on Uncertainty in Artificial Intelligence (UAI)","author":"Asmuth","year":"2009"},{"key":"2026033014111481400_ref006","first-page":"30","article-title":"Residual algorithms: Reinforcement learning with function approximation","volume-title":"ICML","author":"Baird","year":"1995"},{"key":"2026033014111481400_ref007","doi-asserted-by":"crossref","first-page":"454","DOI":"10.1016\/j.artint.2007.08.001","article-title":"Restricted gradient-descent algorithm for value-function approximation in reinforcement learning","volume":"172","author":"Barreto","year":"2008","journal-title":"Artificial Intelligence"},{"key":"2026033014111481400_ref008","first-page":"687","volume-title":"Neural Information Processing Systems (NIPS)","author":"Barto","year":"1994"},{"key":"2026033014111481400_ref009","doi-asserted-by":"crossref","first-page":"81","DOI":"10.1016\/0004-3702(94)00011-O","article-title":"Learning to act using real-time dynamic programming","volume":"72","author":"Barto","year":"1995","journal-title":"Artificial Intelligence"},{"key":"2026033014111481400_ref010","first-page":"271","volume-title":"Circuits and Systems, 2000. Proceedings. ISCAS 2000 Geneva. The 2000 IEEE International Symposium on","author":"Baxter","year":"2000"},{"key":"2026033014111481400_ref011","author":"Bellman","year":"1957"},{"key":"2026033014111481400_ref012","author":"Bertsekas","year":"1996"},{"key":"2026033014111481400_ref013","author":"Bertsekas","year":"1976"},{"key":"2026033014111481400_ref014","author":"Bertsekas","year":"1996"},{"key":"2026033014111481400_ref015","volume-title":"American Control Conference (ACC)","author":"Bethke","year":"2009"},{"key":"2026033014111481400_ref016","first-page":"105","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Bhatnagar","year":"2007"},{"key":"2026033014111481400_ref017","first-page":"379","volume-title":"International Joint Conference on Autonomous Agents and Multiagent Systems (AAMAS)","author":"Bowling","year":"2008"},{"key":"2026033014111481400_ref018","first-page":"369","volume-title":"Neural Information Processing Systems (NIPS)","author":"Boyan","year":"1995"},{"key":"2026033014111481400_ref019","first-page":"49","volume-title":"International Conference on Machine Learning (ICML)","author":"Boyan","year":"1999"},{"key":"2026033014111481400_ref020","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1023\/A:1018056104778","article-title":"Linear least-squares algorithms for temporal difference learning","volume":"22","author":"Bradtke","year":"1996","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"2026033014111481400_ref021","first-page":"213","article-title":"R-Max - A General Polynomial Time Algorithm for Near-Optimal Reinforcement Learning","volume":"3","author":"Brafman","year":"2002","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"2026033014111481400_ref022","author":"Bu\u015foniu","year":"2010"},{"key":"2026033014111481400_ref023","doi-asserted-by":"crossref","first-page":"227","DOI":"10.1613\/jair.639","article-title":"Hierarchical reinforcement learning with the MAXQ value function decomposition","volume":"13","author":"Dietterich","year":"2000","journal-title":"Journal of Artificial Intelligence and Research (JAIR)"},{"key":"2026033014111481400_ref024","article-title":"Reinforcement learning benchmarks and bake-offs II","volume-title":"Advances in Neural Information Processing Systems (NIPS) 17 Workshop","author":"Dutech","year":"2005"},{"key":"2026033014111481400_ref025","first-page":"154","article-title":"Bayes meets bellman: The gaussian process approach to temporal difference learning","volume-title":"International Conference on Machine Learning (ICML)","author":"Engel","year":"2003"},{"key":"2026033014111481400_ref026","author":"Farahmand","year":"2009"},{"key":"2026033014111481400_ref027","first-page":"441","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Farahmand","year":"2008"},{"key":"2026033014111481400_ref028","article-title":"Risk-sensitive reinforcement learning applied to chance constrained control","volume":"24","author":"Geibel","year":"2005","journal-title":"Journal of Artificial Intelligence and Research (JAIR)"},{"key":"2026033014111481400_ref029","first-page":"881","volume-title":"International Conference on Machine Learning (ICML)","author":"Geramifard","year":"2011"},{"key":"2026033014111481400_ref030","doi-asserted-by":"crossref","DOI":"10.1109\/ACC.2012.6314997","article-title":"Model estimation within planning and learning","volume-title":"American Control Conference (ACC)","author":"Geramifard","year":"2012"},{"key":"2026033014111481400_ref031","article-title":"RLPy: The Reinforcement Learning Library for Education and Research","author":"Geramifard","year":"2013"},{"key":"2026033014111481400_ref032","volume-title":"Proceedings of 29th Annual Conference on Uncertainty in Artificial Intelligence (UAI)","author":"Geramifard","year":"2013"},{"key":"2026033014111481400_ref033","article-title":"Feature Discovery in Reinforcement Learning using Genetic Programming","author":"Girgin","year":"2007"},{"key":"2026033014111481400_ref034","volume-title":"Matrix Computations","author":"Golub","year":"1996"},{"key":"2026033014111481400_ref035","first-page":"261","volume-title":"International Conference on Machine Learning (ICML)","author":"Gordon","year":"1995"},{"key":"2026033014111481400_ref036","doi-asserted-by":"crossref","first-page":"178","DOI":"10.1287\/ijoc.1080.0305","article-title":"Reinforcement learning: A tutorial survey and recent advances","volume":"21","author":"Gosavi","year":"2009","journal-title":"INFORMS J. on Computing"},{"key":"2026033014111481400_ref037","first-page":"1351","article-title":"Adaptive importance sampling with automatic model selection in value function approximation","volume-title":"Association for the Advancement of Artificial Intelligence (AAAI)","author":"Hachiya","year":"2008"},{"key":"2026033014111481400_ref038","author":"Haykin","year":"1994"},{"key":"2026033014111481400_ref039","volume-title":"Dynamic Programming and Markov Processes","author":"Howard","year":"1960"},{"key":"2026033014111481400_ref040","author":"Jaakkola","year":"1993"},{"key":"2026033014111481400_ref041","first-page":"1563","article-title":"Near-optimal regret bounds for reinforcement learning","volume":"11","author":"Jaksch","year":"2010","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"2026033014111481400_ref042","volume-title":"Generalization in Reinforcement Learning","author":"Josemans","year":"2009"},{"key":"2026033014111481400_ref043","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-15880-3_44","article-title":"Gaussian processes for sample efficient reinforcement learning with RMAX-like exploration","volume-title":"European Conference on Machine Learning (ECML)","author":"Jung","year":"2010"},{"key":"2026033014111481400_ref044","doi-asserted-by":"crossref","first-page":"237","DOI":"10.1613\/jair.301","article-title":"Reinforcement learning: A survey","volume":"4","author":"Kaelbling","year":"1996","journal-title":"Journal of Artificial Intelligence and Research (JAIR)"},{"key":"2026033014111481400_ref045","article-title":"Characterizing reinforcement learning methods through parameterized learning problems","volume-title":"Machine Learning","author":"Kalyanakrishnan","year":"2011"},{"key":"2026033014111481400_ref046","first-page":"521","volume-title":"International Conference on Machine Learning (ICML)","author":"Kolter","year":"2009"},{"key":"2026033014111481400_ref047","first-page":"834","article-title":"Comparison of cmacs and radial basis functions for local function approximators in reinforcement learning","volume":"2","author":"Kretchmar","year":"1997","journal-title":"International Conference on Neural Networks"},{"key":"2026033014111481400_ref048","first-page":"1719","article-title":"A non-parametric approach to dynamic programming","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Kroemer","year":"2011"},{"key":"2026033014111481400_ref049","first-page":"101","article-title":"Model-based reinforcement learning with an approximate, learned model","volume-title":"Proceeding of the ninth Yale workshop on adaptive and learning systems","author":"Kuvayev","year":"1996"},{"key":"2026033014111481400_ref050","first-page":"1107","article-title":"Least-squares policy iteration","volume":"4","author":"Lagoudakis","year":"2003","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"2026033014111481400_ref051","volume-title":"Reinforcement Learning: State of the Art","author":"Li","year":"2012"},{"key":"2026033014111481400_ref052","first-page":"733","volume-title":"International Joint Conference on Autonomous Agents and Multiagent Systems (AAMAS)","author":"Li","year":"2009"},{"key":"2026033014111481400_ref053","doi-asserted-by":"crossref","DOI":"10.21437\/Interspeech.2009-659","article-title":"Reinforcement learning for dialog management using least-squares policy iteration and fast feature selection","volume-title":"New York Academy of Sciences Symposium on Machine Learning","author":"Li","year":"2009"},{"key":"2026033014111481400_ref054","doi-asserted-by":"crossref","first-page":"543","DOI":"10.1109\/TSP.2007.907881","article-title":"The kernel least-mean-square algorithm","volume":"56","author":"Liu","year":"2008","journal-title":"IEEE Transactions on Signal Processing"},{"key":"2026033014111481400_ref055","author":"Liu","year":"2010"},{"key":"2026033014111481400_ref056","volume-title":"Proceedings of the Third Conference on Artificial General Intelligence (AGI)","author":"Maei","year":"2010"},{"key":"2026033014111481400_ref057","first-page":"719","volume-title":"International Conference on Machine Learning (ICML)","author":"Maei","year":"2010"},{"key":"2026033014111481400_ref058","article-title":"Representation policy iteration","volume-title":"International Conference on Uncertainty in Artificial Intelligence (UAI)","author":"Mahadevan","year":"2005"},{"key":"2026033014111481400_ref059","author":"Mausam","year":"2012"},{"key":"2026033014111481400_ref060","doi-asserted-by":"crossref","first-page":"664","DOI":"10.1145\/1390156.1390240","article-title":"An analysis of reinforcement learning with function approximation","volume-title":"International Conference on Machine Learning (ICML)","author":"Melo","year":"2008"},{"key":"2026033014111481400_ref061","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1023\/A:1017940631555","article-title":"Risk-sensitive reinforcement learning","volume":"49","author":"Mihatsch","year":"2002","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"2026033014111481400_ref062","doi-asserted-by":"crossref","first-page":"281","DOI":"10.1162\/neco.1989.1.2.281","article-title":"Fast learning in networks of locally-tuned processing units","volume":"1","author":"Moody","year":"1989","journal-title":"Neural Computation"},{"key":"2026033014111481400_ref063","first-page":"103","article-title":"Prioritized sweeping: Reinforcement learning with less data and less time","volume-title":"Machine Learning","author":"Moore","year":"1993"},{"key":"2026033014111481400_ref064","first-page":"1209","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Nouri","year":"2009"},{"key":"2026033014111481400_ref065","doi-asserted-by":"crossref","first-page":"737","DOI":"10.1145\/1273496.1273589","volume-title":"International Conference on Machine Learning (ICML)","author":"Parr","year":"2007"},{"key":"2026033014111481400_ref066","doi-asserted-by":"crossref","first-page":"752","DOI":"10.1145\/1390156.1390251","volume-title":"International Conference on Machine Learning (ICML)","author":"Parr","year":"2008"},{"key":"2026033014111481400_ref067","doi-asserted-by":"crossref","first-page":"2219","DOI":"10.1109\/IROS.2006.282564","volume-title":"2006 IEEE\/RSJ International Conference on Intelligent Robots and Systems","author":"Peters","year":"2006"},{"key":"2026033014111481400_ref068","doi-asserted-by":"crossref","first-page":"1180","DOI":"10.1016\/j.neucom.2007.11.026","article-title":"Natural actor-critic","volume":"71","author":"Peters","year":"2008","journal-title":"Neurocomputing"},{"key":"2026033014111481400_ref069","article-title":"Feature selection using regularization in approximate linear programs for Markov decision processes","volume-title":"International Conference on Machine Learning (ICML)","author":"Petrik","year":"2010"},{"key":"2026033014111481400_ref070","doi-asserted-by":"crossref","DOI":"10.1002\/9780470316887","volume-title":"Markov Decision Processes: Discrete Stochastic Dynamic Programming","author":"Puterman","year":"1994"},{"key":"2026033014111481400_ref071","volume-title":"Gaussian Processes for Machine Learning(ECML)","author":"Rasmussen","year":"2006"},{"key":"2026033014111481400_ref072","first-page":"347","article-title":"Sparse distributed memories for on-line value-based reinforcement learning","volume-title":"European Conference on Machine Learning (ECML)","author":"Ratitch","year":"2004"},{"key":"2026033014111481400_ref073","first-page":"254","article-title":"Evaluation of policy gradient methods and variants on the Cart-Pole benchmark","volume-title":"IEEE International Symposium on Approximate Dynamic Programming and Reinforcement Learning (ADPRL)","author":"Riedmiller","year":"2007"},{"key":"2026033014111481400_ref074","article-title":"Online Q-learning using connectionist systems (tech. rep. no. cued\/f-infeng\/tr 166)","volume-title":"Cambridge University Engineering Department","author":"Rummery","year":"1994"},{"key":"2026033014111481400_ref075","article-title":"International Probabilistic Planning Competition (IPPC) at International Joint Conference on Artificial Intelligence (IJCAI)","author":"Sanner","year":"2011"},{"key":"2026033014111481400_ref076","article-title":"Should one compute the temporal difference fix point or minimize the bellman residual? the unified oblique projection view","volume-title":"International Conference on Machine Learning (ICML)","author":"Scherrer","year":"2010"},{"key":"2026033014111481400_ref077","doi-asserted-by":"crossref","first-page":"1299","DOI":"10.1162\/089976698300017467","article-title":"Nonlinear component analysis as a kernel eigenvalue problem","volume":"10","author":"Sch\u00f6lkopf","year":"1998","journal-title":"Neural Computations"},{"key":"2026033014111481400_ref078","author":"Sch\u00f6lkopf","year":"2002"},{"key":"2026033014111481400_ref079","doi-asserted-by":"crossref","first-page":"568","DOI":"10.1016\/0022-247X(85)90317-8","article-title":"Generalized polynomial approximation in Markovian decision processes","volume":"110","author":"Schweitzer","year":"1985","journal-title":"Journal of mathematical analysis and applications"},{"key":"2026033014111481400_ref080","doi-asserted-by":"crossref","first-page":"968","DOI":"10.1145\/1390156.1390278","volume-title":"International Conference on Machine Learning (ICML)","author":"Silver","year":"2008"},{"key":"2026033014111481400_ref081","doi-asserted-by":"crossref","first-page":"183","DOI":"10.1007\/s10994-012-5280-0","article-title":"Temporal-difference search in computer go","volume":"87","author":"Silver","year":"2012","journal-title":"Machine Learning"},{"key":"2026033014111481400_ref082","first-page":"202","volume-title":"Proceeding of the Tenth National Conference on Artificial Intelligence","author":"Singh","year":"1992"},{"key":"2026033014111481400_ref083","doi-asserted-by":"crossref","first-page":"287","DOI":"10.1023\/A:1007678930559","article-title":"Convergence results for single-step on-policy reinforcement-learning algorithms","volume":"38","author":"Singh","year":"2000","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"2026033014111481400_ref084","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1177\/105971230501300301","article-title":"Reinforcement learning for RoboCup-soccer keepaway","volume":"13","author":"Stone","year":"2005","journal-title":"International Society for Adaptive Behavior"},{"issue":"3","key":"2026033014111481400_ref106","doi-asserted-by":"crossref","first-page":"165","DOI":"10.1177\/105971230501300301","article-title":"Reinforcement learning for RoboCup soccer keepaway","volume":"13","author":"Stone","year":"2005","journal-title":"Adaptive Behavior"},{"key":"2026033014111481400_ref085","first-page":"2413","article-title":"Reinforcement learning in finite mdps: Pac analysis","volume":"10","author":"Strehi","year":"2009","journal-title":"Journal of Machine Learning Research (JMLR)"},{"key":"2026033014111481400_ref086","first-page":"1038","volume-title":"Neural Information Processing Systems (NIPS)","author":"Sutton","year":"1996"},{"key":"2026033014111481400_ref087","author":"Sutton","year":"1998"},{"key":"2026033014111481400_ref088","first-page":"1057","article-title":"Policy gradient methods for reinforcement learning with function approximation","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Sutton","year":"2000"},{"key":"2026033014111481400_ref089","first-page":"993","volume-title":"International Conference on Machine Learning (ICML)","author":"Sutton","year":"2009"},{"key":"2026033014111481400_ref090","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-031-01551-9","volume-title":"Algorithms for Reinforcement Learning","author":"Szepesv\u00e1ri","year":"2010"},{"key":"2026033014111481400_ref091","first-page":"1031","article-title":"Model-based reinforcement learning with nearly tight exploration complexity bounds","volume-title":"International Conference on Machine Learning (ICML)","author":"Szita","year":"2010"},{"key":"2026033014111481400_ref092","first-page":"1017","volume-title":"International Conference on Machine Learning (ICML)","author":"Taylor","year":"2009"},{"key":"2026033014111481400_ref093","doi-asserted-by":"crossref","first-page":"674","DOI":"10.1109\/9.580874","article-title":"An analysis of temporal difference learning with function approximation","volume":"42","author":"Tsitsiklis","year":"1997","journal-title":"IEEE Transactions on Automatic Control"},{"key":"2026033014111481400_ref094","doi-asserted-by":"crossref","first-page":"1799","DOI":"10.1016\/S0005-1098(99)00099-0","article-title":"Average cost temporal-difference learning","volume":"35","author":"Tsitsiklis","year":"1999","journal-title":"Automatica"},{"key":"2026033014111481400_ref095","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-642-33486-3_7","article-title":"Adaptive Planning for Markov Decision Processes with Uncertain Transition Models via Incremental Feature Dependency Discovery","volume-title":"European Conference on Machine Learning (ECML)","author":"Ure","year":"2012"},{"key":"2026033014111481400_ref096","volume-title":"Models of Delayed Reinforcement Learning","author":"Watkins","year":"1989"},{"key":"2026033014111481400_ref097","first-page":"279","article-title":"Q-learning","volume":"8","author":"Watkins","year":"1992","journal-title":"Machine Learning"},{"key":"2026033014111481400_ref098","first-page":"279","article-title":"Q-learning","volume":"8","author":"Watkins","year":"1992","journal-title":"Machine Learning"},{"key":"2026033014111481400_ref099","first-page":"607","article-title":"A complexity analysis of cooperative mechanisms in reinforcement learning","volume-title":"Association for the Advancement of Artificial Intelligence (AAAI)","author":"Whitehead","year":"1991"},{"key":"2026033014111481400_ref100","first-page":"1","article-title":"Introduction to the special issue on empirical evaluations in reinforcement learning","volume-title":"Machine Learning","author":"Whiteson","year":"2011"},{"key":"2026033014111481400_ref101","first-page":"288","article-title":"Pattern-recognizing control systems","volume-title":"Computer and Information Sciences: Collected Papers on Learning, Adaptation and Control in Information Systems, COINS symposium proceedings","author":"Widrow","year":"1964"},{"key":"2026033014111481400_ref102","first-page":"229","article-title":"Simple statistical gradient-following algorithms for connectionist reinforcement learning","volume-title":"Machine Learning","author":"Williams","year":"1992"},{"key":"2026033014111481400_ref103","author":"Winograd","year":"1971"},{"key":"2026033014111481400_ref104","doi-asserted-by":"crossref","first-page":"593","DOI":"10.1287\/moor.1110.0516","article-title":"The simplex and policy-iteration methods are strongly polynomial for the markov decision problem with a fixed discount rate","volume":"36","author":"Ye","year":"2011","journal-title":"Math. Oper. Res."},{"key":"2026033014111481400_ref105","doi-asserted-by":"crossref","first-page":"306","DOI":"10.1287\/moor.1100.0441","article-title":"Error bounds for approximations from projected linear equations","volume":"35","author":"Yu","year":"2010","journal-title":"Math. Oper. Res."}],"container-title":["Foundations and Trends\u00ae in Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftmal\/article-pdf\/6\/4\/375\/11147187\/2200000042en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftmal\/article-pdf\/6\/4\/375\/11147187\/2200000042en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T18:10:42Z","timestamp":1777486242000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftmal\/article\/6\/4\/375\/1332155\/A-Tutorial-on-Linear-Function-Approximators-for"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2013,12,19]]},"references-count":106,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2013,12,19]]}},"URL":"https:\/\/doi.org\/10.1561\/2200000042","relation":{},"ISSN":["1935-8237","1935-8245"],"issn-type":[{"value":"1935-8237","type":"print"},{"value":"1935-8245","type":"electronic"}],"subject":[],"published":{"date-parts":[[2013,12,19]]}}}