{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,9,7]],"date-time":"2023-09-07T20:12:12Z","timestamp":1694117532227},"reference-count":27,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"9","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2018,9,1]]},"DOI":"10.1587\/transinf.2017edp7363","type":"journal-article","created":{"date-parts":[[2018,8,31]],"date-time":"2018-08-31T22:42:10Z","timestamp":1535755330000},"page":"2346-2355","source":"Crossref","is-referenced-by-count":0,"title":["Incremental Estimation of Natural Policy Gradient with Relative Importance Weighting"],"prefix":"10.1587","volume":"E101.D","author":[{"given":"Ryo","family":"IWAKI","sequence":"first","affiliation":[{"name":"Osaka University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hiroki","family":"YOKOYAMA","sequence":"additional","affiliation":[{"name":"Tamagawa University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Minoru","family":"ASADA","sequence":"additional","affiliation":[{"name":"Osaka University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] S. Kakade, \u201cA natural policy gradient,\u201d Advances in Neural Information Processing Systems 14, pp.227-242, 2001."},{"key":"2","doi-asserted-by":"crossref","unstructured":"[2] S.-I. Amari, \u201cNatural gradient works efficiently in learning,\u201d Neural Computation, vol.10, no.2, pp.251-276. 10.1162\/089976698300017746","DOI":"10.1162\/089976698300017746"},{"key":"3","unstructured":"[3] T. Morimura, E. Uchibe, and K. Doya, \u201cUtilizing natural gradient in temporal difference reinforcement learning with eligibility traces,\u201d Proc. 2nd International Symposium on Information Geometry and Its Applications, pp.256-263, 2005."},{"key":"4","unstructured":"[4] T. Morimura, \u201cEfficient task-independent reinforcement learning methods based on policy gradient,\u201d Ph.D. thesis, Nara Institute of Science and Technology, 2008."},{"key":"5","doi-asserted-by":"publisher","unstructured":"[5] S. Bhatnagar, R.S. Sutton, M. Ghavamzadeh, and M. Lee, \u201cNatural actor-critic algorithms,\u201d Automatica, vol.45, no.11, pp.2471-2482, 2009. 10.1016\/j.automatica.2009.07.008","DOI":"10.1016\/j.automatica.2009.07.008"},{"key":"6","doi-asserted-by":"crossref","unstructured":"[6] T. Degris, P.M. Pilarski, and R.S. Sutton, \u201cModel-free reinforcement learning with continuous action in practice,\u201d Proc. 2012 American Control Conference, pp.2177-2182, 2012. 10.1109\/acc.2012.6315022","DOI":"10.1109\/ACC.2012.6315022"},{"key":"7","unstructured":"[7] P.S. Thomas, \u201cBias in natural actor-critic algorithms,\u201d Proc. 31st International Conference on Machine Learning, pp.441-448, 2014."},{"key":"8","doi-asserted-by":"publisher","unstructured":"[8] V. Mnih, K. Kavukcuoglu, D. Silver, A.A. Rusu, J. Veness, M.G. Bellemare, A. Graves, M. Riedmiller, A.K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, \u201cHuman-level control through deep reinforcement learning,\u201d Nature, vol.518, pp.529-533, 2015. 10.1038\/nature14236","DOI":"10.1038\/nature14236"},{"key":"9","doi-asserted-by":"publisher","unstructured":"[9] D. Silver, A. Huang, C.J. Maddison, A. Guez, L. Sifre, G. van den Driessche, J. Schrittwieser, I. Antonoglou, V. Panneershelvam, M. Lanctot, S. Dieleman, D. Grewe, J. Nham, N. Kalchbrenner, I. Sutskever, T. Lillicrap, M. Leach, K. Kavukcuoglu, T. Graepel, and D. Hassabis, \u201cMastering the game of Go with deep neural networks and tree search,\u201d Nature, vol.529, pp.484-489, 2016. 10.1038\/nature16961","DOI":"10.1038\/nature16961"},{"key":"10","unstructured":"[10] J. Duchi, E. Hazan, and Y. Singer, \u201cAdaptive subgradient methods for online learning and stochastic optimization,\u201d The Journal of Machine Learning Research, vol.12, pp.2121-2159, 2011."},{"key":"11","unstructured":"[11] D.P. Kingma and J.L. Ba, \u201cAdam: A method for stochastic optimization,\u201d International Conference for Learning Representations, 2015."},{"key":"12","unstructured":"[12] T. Matsubara, T. Morimura, and J. Morimoto, \u201cAdaptive step-size policy gradients with average reward metric,\u201d Proc. 2nd Asian Conference on Machine Learning, pp.285-298, 2010."},{"key":"13","unstructured":"[13] M. Pirotta, M. Restelli, and L. Bascetta, \u201cAdaptive step-size for policy gradient methods,\u201d Advances in Neural Information Processing Systems 26, pp.1394-1402, 2013."},{"key":"14","unstructured":"[14] M. Hutter and S. Legg, \u201cTemporal difference updating without a learning rate,\u201d Advances in Neural Information Processing Systems 20, pp.705-712, 2007."},{"key":"15","doi-asserted-by":"crossref","unstructured":"[15] W. Dabney and A.G. Barto, \u201cAdaptive step-size for online temporal difference learning,\u201d Proc. 26th Conference on Artificial Intelligence (AAAI-12), pp.872-878, 2012.","DOI":"10.1609\/aaai.v26i1.8313"},{"key":"16","unstructured":"[16] N. Karampatziakis and J. Langford, \u201cOnline importance weight aware updates,\u201d Proc. 27th Conference on Uncertainty in Artificial Intelligence, pp.392-399, 2011."},{"key":"17","unstructured":"[17] R.S. Sutton, D. McAllester, S. Singh, and Y. Mansour, \u201cPolicy gradient methods for reinforcement learning with function approximation,\u201d Advances in Neural Information Processing Systems 12, pp.1057-1063, 1999."},{"key":"18","unstructured":"[18] L.C. Baird III, \u201cAdvantage updating,\u201d Tech. Rep. WL-TR-93-1146, Wright-Patterson Air Force Base Ohio: Wright Laboratory, 1993. 10.21236\/ada280862"},{"key":"19","unstructured":"[19] J. Peters, S. Vijayakumar, and S. Schaal, \u201cReinforcement learning for humanoid robotics,\u201d Proc. Third IEEE-RAS International Conference on Humanoid Robots, pp.103-123, 2003."},{"key":"20","doi-asserted-by":"publisher","unstructured":"[20] J. Peters and S. Schaal, \u201cNatural actor-critic,\u201d Neurocomputing, vol.71, no.7-9, pp.1180-1190, 2008. 10.1016\/j.neucom.2007.11.026","DOI":"10.1016\/j.neucom.2007.11.026"},{"key":"21","unstructured":"[21] H. van Seijen and R.S. Sutton, \u201cTrue online TD(\u03bb),\u201d Proc. 31st International Conference on Machine Learning, pp.692-700, 2014."},{"key":"22","doi-asserted-by":"crossref","unstructured":"[22] R.S. Sutton and A.G. Barto, Reinforcement Learning: An introduction, MIT Press, 1998.","DOI":"10.1109\/TNN.1998.712192"},{"key":"23","doi-asserted-by":"publisher","unstructured":"[23] D.P. Bertsekas and J.N. Tsitsiklis, \u201cGradient convergence in gradient methods with errors,\u201d SIAM Journal on Optimization, vol.10, no.3, pp.627-642, 2000. 10.1137\/s1052623497331063","DOI":"10.1137\/S1052623497331063"},{"key":"24","unstructured":"[24] J.A. Bagnell and J. Schneider, \u201cCovariant policy search,\u201d Proc. 18th International Joint Conference on Artificial Intelligence, pp.1019-1024, 2003."},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] K. Doya, \u201cReinforcement learning in continuous time and space,\u201d Neural Computation, vol.12, no.1, pp.219-245, 2000.","DOI":"10.1162\/089976600300015961"},{"key":"26","doi-asserted-by":"crossref","unstructured":"[26] G. Konidaris, S. Osentoski, and P. Thomas, \u201cValue function approximation in reinforcement learning using the Fourier basis,\u201d Proc. 25th Conference on Artificial Intelligence (AAAI-11), pp.380-385, 2011.","DOI":"10.1609\/aaai.v25i1.7903"},{"key":"27","doi-asserted-by":"publisher","unstructured":"[27] T. Morimura, E. Uchibe, J. Yoshimoto, J. Peters, and K. Doya, \u201cDerivatives of logarithmic stationary distributions for policy gradient reinforcement learning,\u201d Neural Computation, vol.22, no.2, pp.342-376, 2010. 10.1162\/neco.2009.12-08-922","DOI":"10.1162\/neco.2009.12-08-922"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E101.D\/9\/E101.D_2017EDP7363\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,4]],"date-time":"2023-09-04T19:29:15Z","timestamp":1693855755000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E101.D\/9\/E101.D_2017EDP7363\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2018,9,1]]},"references-count":27,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2018]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2017edp7363","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"value":"0916-8532","type":"print"},{"value":"1745-1361","type":"electronic"}],"subject":[],"published":{"date-parts":[[2018,9,1]]}}}