{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,25]],"date-time":"2026-04-25T09:50:50Z","timestamp":1777110650303,"version":"3.51.4"},"reference-count":41,"publisher":"Institute of Electronics, Information and Communications Engineers (IEICE)","issue":"10","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IEICE Trans. Inf. &amp; Syst."],"published-print":{"date-parts":[[2020,10,1]]},"DOI":"10.1587\/transinf.2019edp7170","type":"journal-article","created":{"date-parts":[[2020,9,30]],"date-time":"2020-09-30T22:31:47Z","timestamp":1601505107000},"page":"2143-2153","source":"Crossref","is-referenced-by-count":8,"title":["Towards Interpretable Reinforcement Learning with State Abstraction Driven by External Knowledge"],"prefix":"10.1587","volume":"E103.D","author":[{"given":"Nicolas","family":"BOUGIE","sequence":"first","affiliation":[{"name":"Sokendai, The Graduate University for Advanced Studies"},{"name":"National Institute of Informatics"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Ryutaro","family":"ICHISE","sequence":"additional","affiliation":[{"name":"Sokendai, The Graduate University for Advanced Studies"},{"name":"National Institute of Informatics"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"532","reference":[{"key":"1","unstructured":"[1] R.S. Sutton, \u201cGeneralization in reinforcement learning: Successful examples using sparse coarse coding,\u201d Proc. Advances in Neural Information Processing Systems, pp.1038-1044, 1996."},{"key":"2","doi-asserted-by":"publisher","unstructured":"[2] C.J. Watkins and P. Dayan, \u201cQ-learning,\u201d Machine learning, vol.8, no.3-4, pp.279-292, 1992. 10.1023\/a:1022676722315","DOI":"10.1023\/A:1022676722315"},{"key":"3","doi-asserted-by":"crossref","unstructured":"[3] P. Abbeel, A. Coates, M. Quigley, and A.Y. Ng, \u201cAn application of reinforcement learning to aerobatic helicopter flight,\u201d Proc. Advances in Neural Information Processing Systems, pp.1-8, 2007.","DOI":"10.7551\/mitpress\/7503.003.0006"},{"key":"4","unstructured":"[4] S. Levine, C. Finn, T. Darrell, and P. Abbeel, \u201cEnd-to-end training of deep visuomotor policies,\u201d The Journal of Machine Learning Research, vol.17, no.1, pp.1334-1373, 2016."},{"key":"5","unstructured":"[5] V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller, \u201cPlaying atari with deep reinforcement learning,\u201d arXiv preprint arXiv:1312.5602, 2013."},{"key":"6","unstructured":"[6] M.G. Bellemare, Y. Naddaf, J. Veness, and M. Bowling, \u201cThe arcade learning environment: an evaluation platform for general agents,\u201d Proc. International Conference on Artificial Intelligence, pp.4148-4152, 2015."},{"key":"7","doi-asserted-by":"publisher","unstructured":"[7] T. Hester and P. Stone, \u201cTexplore: real-time sample-efficient reinforcement learning for robots,\u201d Machine learning, vol.90, no.3, pp.385-429, 2013. 10.1007\/s10994-012-5322-7","DOI":"10.1007\/s10994-012-5322-7"},{"key":"8","doi-asserted-by":"publisher","unstructured":"[8] G. Adomavicius and A. Tuzhilin, \u201cToward the next generation of recommender systems: A survey of the state-of-the-art and possible extensions,\u201d IEEE Trans. Knowl. Data Eng., vol.17, no.6, pp.734-749, 2005. 10.1109\/tkde.2005.99","DOI":"10.1109\/TKDE.2005.99"},{"key":"9","unstructured":"[9] M. Garnelo, K. Arulkumaran, and M. Shanahan, \u201cTowards deep symbolic reinforcement learning,\u201d arXiv preprint arXiv:1609.05518, 2016."},{"key":"10","unstructured":"[10] C. Zhang, O. Vinyals, R. Munos, and S. Bengio, \u201cA study on overfitting in deep reinforcement learning,\u201d arXiv preprint arXiv:1804.06893, 2018."},{"key":"11","doi-asserted-by":"crossref","unstructured":"[11] N. Bougie and R. Ichise, \u201cDeep reinforcement learning boosted by external knowledge,\u201d Proc. ACM Symposium on Applied Computing, pp.331-338, 2018. 10.1145\/3167132.3167165","DOI":"10.1145\/3167132.3167165"},{"key":"12","unstructured":"[12] T. Hester, M. Vecerik, O. Pietquin, M. Lanctot, T. Schaul, B. Piot, D. Horgan, J. Quan, A. Sendonaris, I. Osband, et al., \u201cDeep q-learning from demonstrations,\u201d Thirty-Second AAAI Conference on Artificial Intelligence, pp.3223-3230, 2018."},{"key":"13","doi-asserted-by":"publisher","unstructured":"[13] S.P. Singh and R.S. Sutton, \u201cReinforcement learning with replacing eligibility traces,\u201d Machine Learning, vol.22, no.1-3, pp.123-158, 1996. 10.1007\/bf00114726","DOI":"10.1007\/BF00114726"},{"key":"14","doi-asserted-by":"crossref","unstructured":"[14] J. Kim and J. Canny, \u201cInterpretable learning for self-driving cars by visualizing causal attention,\u201d Proc. International Conference on Computer Vision, pp.2961-2969, IEEE, 2017. 10.1109\/iccv.2017.320","DOI":"10.1109\/ICCV.2017.320"},{"key":"15","unstructured":"[15] A. d&apos;Avila Garcez, A. Resende Riquetti Dutra, and E. Alonso, \u201cTowards symbolic reinforcement learning with common sense,\u201d arXiv preprint arXiv:1804.08597, 2018."},{"key":"16","unstructured":"[16] A. Verma, V. Murali, R. Singh, P. Kohli, and S. Chaudhuri, \u201cProgrammatically interpretable reinforcement learning,\u201d arXiv preprint arXiv:1804.02477, 2018."},{"key":"17","doi-asserted-by":"crossref","unstructured":"[17] S. Lange and M. Riedmiller, \u201cDeep auto-encoder neural networks in reinforcement learning,\u201d The 2010 International Joint Conference on Neural Networks (IJCNN), pp.1-8, IEEE, 2010. 10.1109\/ijcnn.2010.5596468","DOI":"10.1109\/IJCNN.2010.5596468"},{"key":"18","doi-asserted-by":"crossref","unstructured":"[18] M. Rosencrantz, G. Gordon, and S. Thrun, \u201cLearning low dimensional predictive representations,\u201d Proc. International Conference on Machine learning, p.88, 2004. 10.1145\/1015330.1015441","DOI":"10.1145\/1015330.1015441"},{"key":"19","doi-asserted-by":"publisher","unstructured":"[19] S. D\u017eeroski, L. De Raedt, and K. Driessens, \u201cRelational reinforcement learning,\u201d Machine learning, vol.43, no.1-2, pp.7-52, 2001. 10.1023\/a:1007694015589","DOI":"10.1023\/A:1007694015589"},{"key":"20","unstructured":"[20] D. Andre and S.J. Russell, \u201cState abstraction for programmable reinforcement learning agents,\u201d Proc. National Conference on Artificial Intelligence, pp.119-125, 2002."},{"key":"21","doi-asserted-by":"crossref","unstructured":"[21] M.K. Gunady and W. Gomaa, \u201cReinforcement learning generalization using state aggregation with a maze-solving problem,\u201d Proc. Conference on Electronics, Communications and Computers, pp.157-162, 2012. 10.1109\/jec-ecc.2012.6186975","DOI":"10.1109\/JEC-ECC.2012.6186975"},{"key":"22","doi-asserted-by":"crossref","unstructured":"[22] N. Bougie and R. Ichise, \u201cAbstracting reinforcement learning agents with prior knowledge,\u201d Proc. International Conference on Principles and Practice of Multi-Agent Systems, pp.431-439, Springer, 2018. 10.1007\/978-3-030-03098-8_27","DOI":"10.1007\/978-3-030-03098-8_27"},{"key":"23","doi-asserted-by":"publisher","unstructured":"[23] R.S. Sutton, \u201cLearning to predict by the methods of temporal differences,\u201d Machine learning, vol.3, no.1, pp.9-44, 1988. 10.1007\/bf00115009","DOI":"10.1007\/BF00115009"},{"key":"24","doi-asserted-by":"crossref","unstructured":"[24] R.S. Sutton and A.G. Barto, Reinforcement learning: an introduction, MIT press, Cambridge, 1998.","DOI":"10.1109\/TNN.1998.712192"},{"key":"25","unstructured":"[25] G.A. Rummery and M. Niranjan, \u201cOn-line Q-learning using connectionist systems,\u201d tech. rep., University of Cambridge, Oct. 04 1994."},{"key":"26","unstructured":"[26] J. Randl\u00f8v and P. Alstr\u00f8m, \u201cLearning to drive a bicycle using reinforcement learning and shaping,\u201d Proc. International Conference on Machine Learning, pp.463-471, 1998."},{"key":"27","doi-asserted-by":"publisher","unstructured":"[27] Y. LeCun, Y. Bengio, and G. Hinton, \u201cDeep learning,\u201d Nature, vol.521, no.7553, pp.436-444, 2015. 10.1038\/nature14539","DOI":"10.1038\/nature14539"},{"key":"28","unstructured":"[28] S. Nison, Japanese candlestick charting techniques: a contemporary guide to the ancient investment techniques of the Far East, Penguin, 2001."},{"key":"29","doi-asserted-by":"crossref","unstructured":"[29] M. Mashayekhi and R. Gras, \u201cRule extraction from random forest: the rf+hc methods,\u201d Proc. Canadian Conference on Artificial Intelligence, pp.223-237, Springer, 2015. 10.1007\/978-3-319-18356-5_20","DOI":"10.1007\/978-3-319-18356-5_20"},{"key":"30","doi-asserted-by":"publisher","unstructured":"[30] M. Pal, \u201cRandom forest classifier for remote sensing classification,\u201d Proc. International Journal of Remote Sensing, vol.26, no.1, pp.217-222, 2005. 10.1080\/01431160412331269698","DOI":"10.1080\/01431160412331269698"},{"key":"31","doi-asserted-by":"publisher","unstructured":"[31] S.R. Safavian and D. Landgrebe, \u201cA survey of decision tree classifier methodology,\u201d IEEE Trans. Syst., Man, Cybern., vol.21, no.3, pp.660-674, 1991. 10.1109\/21.97458","DOI":"10.1109\/21.97458"},{"key":"32","doi-asserted-by":"publisher","unstructured":"[32] Y. Bengio, O. Delalleau, and C. Simard, \u201cDecision trees do not generalize to new variations,\u201d Computational Intelligence, vol.26, no.4, pp.449-467, 2010. 10.1111\/j.1467-8640.2010.00366.x","DOI":"10.1111\/j.1467-8640.2010.00366.x"},{"key":"33","unstructured":"[33] I. Gulrajani, K. Kumar, F. Ahmed, A.A. Taiga, F. Visin, D. Vazquez, and A. Courville, \u201cPixelvae: A latent variable model for natural images,\u201d arXiv preprint arXiv:1611.05013, 2016."},{"key":"34","doi-asserted-by":"publisher","unstructured":"[34] S. Lloyd, \u201cLeast squares quantization in pcm,\u201d IEEE Trans. Inf. Theory, vol.28, no.2, pp.129-137, 1982. 10.1109\/tit.1982.1056489","DOI":"10.1109\/TIT.1982.1056489"},{"key":"35","unstructured":"[35] G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, \u201cOpenai gym.\u201d https:\/\/github.com\/openai\/gym, 2016."},{"key":"36","unstructured":"[36] M. Hausknecht and P. Stone, \u201cDeep recurrent q-learning for partially observable mdps,\u201d arXiv preprint arXiv:1507.06527, 2015."},{"key":"37","unstructured":"[37] T.P. Lillicrap, J.J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, \u201cContinuous control with deep reinforcement learning,\u201d arXiv preprint arXiv:1509.02971, 2015."},{"key":"38","unstructured":"[38] Z. Xiong, X.Y. Liu, S. Zhong, A. Walid, et al., \u201cPractical deep reinforcement learning approach for stock trading,\u201d arXiv preprint arXiv:1811.07522, 2018."},{"key":"39","doi-asserted-by":"crossref","unstructured":"[39] A.R. Azhikodan, A.G. Bhat, and M.V. Jadhav, \u201cStock trading bot using deep reinforcement learning,\u201d in Innovations in Computer Science and Engineering, pp.41-49, Springer, 2019. 10.1007\/978-981-10-8201-6_5","DOI":"10.1007\/978-981-10-8201-6_5"},{"key":"40","unstructured":"[40] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, \u201cProximal policy optimization algorithms,\u201d arXiv preprint arXiv:1707.06347, 2017."},{"key":"41","doi-asserted-by":"publisher","unstructured":"[41] V. Mnih, K. Kavukcuoglu, D. Silver, A.A. Rusu, J. Veness, M.G. Bellemare, A. Graves, M. Riedmiller, A.K. Fidjeland, G. Ostrovski,S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D.Kumaran, D. Wierstra, S. Legg, and D. Hassabis, \u201cHuman-level control through deep reinforcement learning,\u201d Nature, vol.518, no.7540, pp.529-533, 2015. 10.1038\/nature14236","DOI":"10.1038\/nature14236"}],"container-title":["IEICE Transactions on Information and Systems"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E103.D\/10\/E103.D_2019EDP7170\/_pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,8]],"date-time":"2023-10-08T19:41:42Z","timestamp":1696794102000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.jstage.jst.go.jp\/article\/transinf\/E103.D\/10\/E103.D_2019EDP7170\/_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2020,10,1]]},"references-count":41,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2020]]}},"URL":"https:\/\/doi.org\/10.1587\/transinf.2019edp7170","relation":{},"ISSN":["0916-8532","1745-1361"],"issn-type":[{"value":"0916-8532","type":"print"},{"value":"1745-1361","type":"electronic"}],"subject":[],"published":{"date-parts":[[2020,10,1]]}}}