{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,30]],"date-time":"2025-06-30T00:00:46Z","timestamp":1751241646337},"reference-count":38,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"9","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2020,9,1]]},"DOI":"10.1587\/transinf.2019edp7298","type":"journal-article","created":{"date-parts":[[2020,8,31]],"date-time":"2020-08-31T22:14:46Z","timestamp":1598912086000},"page":"1960-1970","source":"Crossref","is-referenced-by-count":4,"title":["Hybrid of Reinforcement and Imitation Learning for Human-Like Agents"],"prefix":"10.1587","volume":"E103.D","author":[{"given":"Rousslan F. J.","family":"DOSSA","sequence":"first","affiliation":[{"name":"Graduate School of System Informatics, Kobe University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Xinyu","family":"LIAN","sequence":"additional","affiliation":[{"name":"Graduate School of System Informatics, Kobe University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hirokazu","family":"NOMOTO","sequence":"additional","affiliation":[{"name":"EQUOS RESEARCH Co., Ltd."}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Takashi","family":"MATSUBARA","sequence":"additional","affiliation":[{"name":"Graduate School of Engineering Science, Osaka University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kuniaki","family":"UEHARA","sequence":"additional","affiliation":[{"name":"Faculty of Business Administration, Osaka Gakuin University"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","doi-asserted-by":"publisher","unstructured":"[1] D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis, \u201cMastering the game of go without human knowledge,\u201d Nature, vol.550, no.7676, pp.354-359, 2017. 10.1038\/nature24270","DOI":"10.1038\/nature24270"},{"key":"2","doi-asserted-by":"publisher","unstructured":"[2] A.E. Sallab, M. Abdou, E. Perot, and S. Yogamani, \u201cDeep reinforcement learning framework for autonomous driving,\u201d Electronic Imaging, vol.2017, no.19, pp.70-76, 2017. 10.2352\/issn.2470-1173.2017.19.avm-023","DOI":"10.2352\/ISSN.2470-1173.2017.19.AVM-023"},{"key":"3","unstructured":"[3] S. Shalev-Shwartz, S. Shammah, and A. Shashua, \u201cSafe, multi-agent, reinforcement learning for autonomous driving,\u201d CoRR, abs\/1610.03295, 2016."},{"key":"4","doi-asserted-by":"crossref","unstructured":"[4] D. Isele, R. Rahimi, A. Cosgun, K. Subramanian, and K. Fujimura, \u201cNavigating occluded intersections with autonomous vehicles using deep reinforcement learning,\u201d Proceedings of the International Conference on Robotics and Automation (IEEE ICRA), pp.2034-2039, 2018. 10.1109\/icra.2018.8461233","DOI":"10.1109\/ICRA.2018.8461233"},{"key":"5","unstructured":"[5] B. Vikas, \u201cDeep reinforcement learning approach to autonomous navigation,\u201d https:\/\/github.com\/bhanuvikasr\/Deep-RL-TORCS, 2017."},{"key":"6","doi-asserted-by":"publisher","unstructured":"[6] V. Mnih, K. Kavukcuoglu, D. Silver, A.A. Rusu, J. Veness, M.G. Bellemare, A. Graves, M. Riedmiller, A.K. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, \u201cHuman-level control through deep reinforcement learning,\u201d Nature, vol.518, no.7540, pp.529-533, 2015. 10.1038\/nature14236","DOI":"10.1038\/nature14236"},{"key":"7","unstructured":"[7] V. Mnih et al., \u201cAsynchronous methods for deep reinforcement learning,\u201d Proc. Int. Conf. Mach. Learn. (ICML), vol.48, pp.1928-1937, 2016."},{"key":"8","unstructured":"[8] S. Ross, G. Gordon, and D. Bagnell, \u201cA reduction of imitation learning and structured prediction to no-regret online learning,\u201d Proc. Int. Conf. Artif. Intell. Stat. (AISTATS), vol.15, pp.627-635, 2011."},{"key":"9","doi-asserted-by":"publisher","unstructured":"[9] J. Ortega, N. Shaker, J. Togelius, and G.N. Yannakakis, \u201cImitating human playing styles in super mario bros,\u201d Entertainment Computing, vol.4, no.2, pp.93-104, 2013. 10.1016\/j.entcom.2012.10.001","DOI":"10.1016\/j.entcom.2012.10.001"},{"key":"10","unstructured":"[10] K.-W. Chang et al., \u201cLearning to search better than your teacher,\u201d Proc. Int. Conf. Mach. Learn. (ICML), vol.37, pp.2058-2066, 2015."},{"key":"11","doi-asserted-by":"crossref","unstructured":"[11] A. Nair, B. McGrew, M. Andrychowicz, W. Zaremba, and P. Abbeel, \u201cOvercoming exploration in reinforcement learning with demonstrations,\u201d Proceedings of International Conference on Robotics and Automation (IEEE ICRA), pp.6292-6299, 2018. 10.1109\/icra.2018.8463162","DOI":"10.1109\/ICRA.2018.8463162"},{"key":"12","unstructured":"[12] T. Hester et al., \u201cLearning from demonstrations for real world reinforcement learning,\u201d CoRR, abs\/1704.03732, 2017."},{"key":"13","unstructured":"[13] H.M. Le et al., \u201cHierarchical imitation and reinforcement learning,\u201d CoRR, abs\/1803.00590, 2018."},{"key":"14","doi-asserted-by":"publisher","unstructured":"[14] M.G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, \u201cThe arcade learning environment: An evaluation platform for general agents,\u201d Journal of Artificial Intelligence Research, vol.47, pp.253-279, 2013. 10.1613\/jair.3912","DOI":"10.1613\/jair.3912"},{"key":"15","unstructured":"[15] B. Wymann, C. Dimitrakakis, A.D. Sumner, E. Espi\u00e9, C. Guionneau, and R. Coulom, \u201cTorcs: The open racing car simulator,\u201d 2013."},{"key":"16","unstructured":"[16] R.F.J. Dossa, X. Lian, H. Nomoto, T. Matsubara, and K. Uehara, \u201cA human-like agent based on a hybrid of reinforcement and imitation learning,\u201d Proceedings of the International Joint Conference on Neural Networks (IJCNN), 2019. 10.1109\/ijcnn.2019.8852026"},{"key":"17","unstructured":"[17] H. Daum\u00e9 III, J. Langford, and S. Ross, \u201cEfficient programmable learning to search,\u201d CoRR, abs\/1406.1837, 2014."},{"key":"18","doi-asserted-by":"publisher","unstructured":"[18] J. Rao Doppa, A. Fern, and P. Tadepalli, \u201cHc-search: A learning framework for search-based structured prediction,\u201d The Journal of Artificial Intelligence Research (JAIR), vol.50, pp.369-407, 2014. 10.1613\/jair.4212","DOI":"10.1613\/jair.4212"},{"key":"19","unstructured":"[19] S. Ross and J.A. Bagnell, \u201cReinforcement and imitation learning via interactive no-regret learning,\u201d CoRR, abs\/1406.5979, 2014."},{"key":"20","unstructured":"[20] Z. Wang et al., \u201cDueling network architectures for deep reinforcement learning,\u201d Proc. Int. Conf. Mach. Learn. (ICML), vol.48, pp.1995-2003, 2016."},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] B. Balaguer and S. Carpin, \u201cCombining imitation and reinforcement learning to fold deformable planar objects,\u201d Proceedings of the International Conference on Intelligent Robots and Systems (IEEE\/RSJ IROS), pp.1405-1412, 2011. 10.1109\/iros.2011.6094992","DOI":"10.1109\/IROS.2011.6094992"},{"key":"22","unstructured":"[22] W. Sun, J.A. Bagnell, and B. Boots, \u201cTruncated horizon policy search: Combining reinforcement learning &amp; imitation learning,\u201d International Conference on Learning Representations (ICLR), 2018."},{"key":"23","unstructured":"[23] A.Y. Ng, D. Harada, and S. Russell, \u201cPolicy invariance under reward transformations: Theory and application to reward shaping,\u201d Proc. Int. Conf. Mach. Learn. (ICML), pp.278-287, 1999."},{"key":"24","unstructured":"[24] S. Hecker, D. Dai, and L. van Gool, \u201cLearning accurate, comfortable and human-like driving,\u201d abs\/1903.10995, 2019."},{"key":"25","doi-asserted-by":"crossref","unstructured":"[25] R.S. Sutton et al., Reinforcement Learning: An Introduction. MIT Press, 1998.","DOI":"10.1109\/TNN.1998.712192"},{"key":"26","unstructured":"[26] T.P. Lillicrap et al., \u201cContinuous control with deep reinforcement learning,\u201d CoRR, abs\/1509.02971, 2015."},{"key":"27","unstructured":"[27] D. Silver et al., \u201cDeterministic policy gradient algorithms,\u201d Proc. Int. Conf. Mach. Learn. (ICML), vol.32, pp.I-387-I-395, 2014."},{"key":"28","unstructured":"[28] M. Plappert et al., \u201cParameter space noise for exploration,\u201d CoRR, abs\/1706.01905, 2017."},{"key":"29","unstructured":"[29] G. Hinton, O. Vinyals, and J. Dean, \u201cDistilling the knowledge in a neural network,\u201d arXiv preprint, arXiv:1503.02531, 2015."},{"key":"30","unstructured":"[30] A.A. Rusu et al., \u201cPolicy distillation,\u201d CoRR, abs\/1511.06295, 2015."},{"key":"31","unstructured":"[31] J. Ho and S. Ermon, \u201cGenerative adversarial imitation learning,\u201d CoRR, abs\/1606.03476, 2016."},{"key":"32","unstructured":"[32] J. Schulman et al., \u201cTrust region policy optimization,\u201d arXiv e-prints, arXiv:1502.05477, 2015."},{"key":"33","unstructured":"[33] P. Dhariwal et al., \u201cOpenai baselines,\u201d https:\/\/github.com\/openai\/baselines, 2017."},{"key":"34","unstructured":"[34] D.P. Kingma and J. Ba, \u201cAdam: A method for stochastic optimization,\u201d CoRR, 2014."},{"key":"35","unstructured":"[35] N. Srivastava et al., \u201cDropout: A simple way to prevent neural networks from overfitting,\u201d Journal of Machine Learning Research, pp.1929-1958, 2014."},{"key":"36","unstructured":"[36] B. Lau, \u201cUsing keras and deep deterministic policy gradient to play torcs,\u201d https:\/\/yanpanlau.github.io\/2016\/10\/11\/Torcs-Keras.html, 2016."},{"key":"37","unstructured":"[37] Y. You, \u201cTorcs for reinforcement learning,\u201d https:\/\/github.com\/YurongYou\/rlTORCS"},{"key":"38","unstructured":"[38] N. Yoshida, \u201cGym torcs,\u201d https:\/\/github.com\/ugo-nama-kun\/gym-torcs, 2016."}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E103.D\/9\/E103.D_2019EDP7298\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2020,9,5]],"date-time":"2020-09-05T03:26:50Z","timestamp":1599276410000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E103.D\/9\/E103.D_2019EDP7298\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,9,1]]},"references-count":38,"journal-issue":{"issue":"9","published-print":{"date-parts":[[2020]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2019edp7298","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"value":"0916-8532","type":"print"},{"value":"1745-1361","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,9,1]]}}}