{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2022,10,14]],"date-time":"2022-10-14T04:26:48Z","timestamp":1665721608707},"reference-count":37,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[2012,1,17]],"date-time":"2012-01-17T00:00:00Z","timestamp":1326758400000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J Supercomput"],"published-print":{"date-parts":[[2012,10]]},"DOI":"10.1007\/s11227-011-0738-6","type":"journal-article","created":{"date-parts":[[2012,1,16]],"date-time":"2012-01-16T14:18:28Z","timestamp":1326723508000},"page":"588-615","source":"Crossref","is-referenced-by-count":5,"title":["Towards a Multiple-Lookahead-Levels agent reinforcement-learning technique and its implementation in integrated circuits"],"prefix":"10.1007","volume":"62","author":[{"given":"H. S.","family":"Al-Dayaa","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"D. B.","family":"Megherbi","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2012,1,17]]},"reference":[{"key":"738_CR1","unstructured":"Angelo A, Florence D et al (1999) Efficient learning of variable-resolution cognitive maps for autonomous indoor navigation. IEEE Trans Robot Autom"},{"key":"738_CR2","unstructured":"Barto AG, Sutton RS, Watkins CJCH (1990) Learning and sequential decision making. Learning Comput Neurosci, 539\u2013602"},{"issue":"2","key":"738_CR3","doi-asserted-by":"crossref","first-page":"398","DOI":"10.1109\/TSMCB.2006.883264","volume":"37","author":"BN Araabi","year":"2007","unstructured":"Araabi BN, Mastoureshgh S, Ahmadabadi MN (2007) A study on expertise of agents and its effects on cooperative Q-Learning. IEEE Trans Syst Man Cybern, Part B, Cybern 37(2):398\u2013409","journal-title":"IEEE Trans Syst Man Cybern, Part B, Cybern"},{"key":"738_CR4","unstructured":"Watkins CJCH (1989) Learning with delayed rewards. PhD Thesis Cambridge University Psychology Department"},{"key":"738_CR5","first-page":"279","volume":"8","author":"C Watkins","year":"1992","unstructured":"Watkins C, Dayan P (1992) Q-Learning. Mach Learn 8:279\u2013292","journal-title":"Mach Learn"},{"issue":"2","key":"738_CR6","doi-asserted-by":"crossref","first-page":"279","DOI":"10.1109\/72.839000","volume":"11","author":"C Clausen","year":"2000","unstructured":"Clausen C, Wechsler H (2000) Quad-Q-Learning. IEEE Trans Neural Netw 11(2):279\u2013294","journal-title":"IEEE Trans Neural Netw"},{"key":"738_CR7","volume-title":"Proceedings of the 2006 international conference on machine learning; models, technologies & applications","author":"DB Megherbi","year":"2007","unstructured":"Megherbi DB, Al-Dayaa HS (2007) A Lyapunov-stability-based system hardware architecture for a real-time multiple-look-ahead-levels reinforcement learning. In: Proceedings of the 2006 international conference on machine learning; models, technologies & applications, Nevada, USA"},{"key":"738_CR8","series-title":"Unmanned Ground vehicle Technology","first-page":"419","volume-title":"Proceedings of the SPIE international conference on defense sensing","author":"DB Megherbi","year":"2001","unstructured":"Megherbi DB, Teirelbar A, Boulenouar AJ (2001) A time-varying-environment machine learning technique for autonomous agent shortest path planning. In: Proceedings of the SPIE international conference on defense sensing. Unmanned Ground vehicle Technology, Orlando, Florida, April 2001, pp 419\u2013428"},{"key":"738_CR9","volume-title":"Computer organization & design","author":"DA Patterson","year":"2004","unstructured":"Patterson DA, Hennessy JL (2004) Computer organization & design. Morgan Kaufmann, San Mateo"},{"key":"738_CR10","unstructured":"Ernst D, Geurts E, Wehenkel L (2005) Tree-based batch mode reinforcement learning. J Mach Learn Res"},{"key":"738_CR11","volume-title":"Advanced engineering mathematics","author":"E Kreyszig","year":"1993","unstructured":"Kreyszig E (1993) Advanced engineering mathematics, 7th edn. Wiley, New York","edition":"7"},{"key":"738_CR12","volume-title":"Proceedings of the 2006 international conference on machine learning; models, technologies & applications","author":"HS Al-Dayaa","year":"2006","unstructured":"Al-Dayaa HS, Megherbi DB (2006) Fast reinforcement learning technique via Multiple Lookahead Levels. In: Proceedings of the 2006 international conference on machine learning; models, technologies & applications, Nevada, USA"},{"key":"738_CR13","volume-title":"Proceedings of the 2006 international conference on machine learning; models, technologies & applications","author":"HS Al-Dayaa","year":"2006","unstructured":"Al-Dayaa HS, Megherbi DB (2006) Fast reinforcement learning techniques using the Euclidean distance and the agent state occurrence frequency. In: Proceedings of the 2006 international conference on machine learning; models, technologies & applications, Nevada, USA"},{"key":"738_CR14","unstructured":"IEEE (1985) IEEE Standard for binary floating point arithmetic. Institute of Electrical & Electronics Engineers, March 1985"},{"issue":"4","key":"738_CR15","doi-asserted-by":"crossref","first-page":"1014","DOI":"10.1109\/TSMCB.2008.922018","volume":"38","author":"J Valasek","year":"2008","unstructured":"Valasek J, Doebbler J, Tandale MD, Meade AJ (2008) Improved adaptive\u2013reinforcement learning control for morphing unmanned air vehicles. IEEE Trans Syst Man Cybern, Part B, Cybern 38(4):1014\u20131020","journal-title":"IEEE Trans Syst Man Cybern, Part B, Cybern"},{"key":"738_CR16","volume-title":"Proceedings of 2009 IEEE international conference on systems, man, and cybernetics","author":"K-S Hwang","year":"2009","unstructured":"Hwang K-S, Lo C-Y, Chen K-J (2009) Real-valued Q-Learning in multi-agent cooperation. In: Proceedings of 2009 IEEE international conference on systems, man, and cybernetics, Texas, USA"},{"key":"738_CR17","series-title":"Mathematics and its applications","doi-asserted-by":"crossref","DOI":"10.1007\/978-94-015-7939-1","volume-title":"Vector Lyapunov functions and stability analysis of nonlinear systems","author":"V Lakshmikantham","year":"1991","unstructured":"Lakshmikantham V et al (1991) Vector Lyapunov functions and stability analysis of nonlinear systems. Mathematics and its applications. Springer, Berlin"},{"issue":"3","key":"738_CR18","doi-asserted-by":"crossref","first-page":"1444","DOI":"10.1109\/TIE.2007.908526","volume":"55","author":"L Hu","year":"2008","unstructured":"Hu L, Zhou C, Sun Z (2008) Estimating biped gait using spline-based probability distribution function with Q-Learning. IEEE Trans Ind Electron 55(3):1444\u20131452","journal-title":"IEEE Trans Ind Electron"},{"issue":"5","key":"738_CR19","doi-asserted-by":"crossref","first-page":"2140","DOI":"10.1109\/TSMCB.2004.832154","volume":"34","author":"M Guo","year":"2004","unstructured":"Guo M, Liu Y, Malec J (2004) A new Q-Learning algorithm based on the metropolis criterion. IEEE Trans Syst Man Cybern, Part B, Cybern 34(5):2140\u20132143","journal-title":"IEEE Trans Syst Man Cybern, Part B, Cybern"},{"issue":"4","key":"738_CR20","doi-asserted-by":"crossref","first-page":"930","DOI":"10.1109\/TSMCB.2008.920231","volume":"38","author":"MA Wiering","year":"2008","unstructured":"Wiering MA, van Hasselt H (2008) Ensemble algorithms in reinforcement learning. IEEE Trans Syst Man Cybern, Part B, Cybern 38(4):930\u2013935","journal-title":"IEEE Trans Syst Man Cybern, Part B, Cybern"},{"key":"738_CR21","volume-title":"Complete digital design: a comprehensive guide to digital electronics and computer system architecture","author":"M Balch","year":"2003","unstructured":"Balch M (2003) Complete digital design: a comprehensive guide to digital electronics and computer system architecture. McGraw-Hill Professional, New York"},{"key":"738_CR22","unstructured":"Murphy SA (2005) A generalization error for Q-Learning. J Mach Learn Res, July"},{"key":"738_CR23","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4899-0013-5","volume-title":"Stabilization of control systems","author":"O Hijab","year":"1987","unstructured":"Hijab O (1987) Stabilization of control systems. Springer, New York"},{"key":"738_CR24","first-page":"341","volume":"8","author":"P Dayan","year":"1992","unstructured":"Dayan P (1992) The convergence of TD(\u03bb) for general \u03bb. Mach Learn 8:341\u2013362","journal-title":"Mach Learn"},{"key":"738_CR25","volume-title":"IEEE Press series on microelectronic systems","author":"R Jacob Baker","year":"2002","unstructured":"Jacob Baker R (2002) Mixed-signal circuit design. In: IEEE Press series on microelectronic systems"},{"key":"738_CR26","volume-title":"A mathematical introduction to robotic manipulation","author":"RM Murray","year":"1994","unstructured":"Murray RM, Li Z, Sastry SS (1994) A mathematical introduction to robotic manipulation. CRC Press LLC, Boca Raton"},{"key":"738_CR27","first-page":"251","volume":"22","author":"R Maclin","year":"1996","unstructured":"Maclin R, Shavlik JW (1996) Creating advice-taking reinforcement learners. Mach Learn 22:251\u2013281","journal-title":"Mach Learn"},{"key":"738_CR28","volume-title":"Proceedings of the int. conf. on the simulation of adaptive behavior","author":"R Riolo","year":"1991","unstructured":"Riolo R (1991) Lookahead planning and latent learning in a classifier system. In: Proceedings of the int. conf. on the simulation of adaptive behavior"},{"key":"738_CR29","volume-title":"Reinforcement learning: an introduction","author":"RS Sutton","year":"1998","unstructured":"Sutton RS, Barto AG (1998) Reinforcement learning: an introduction. MIT Press, Cambridge"},{"key":"738_CR30","first-page":"151","volume-title":"Working notes of 1991 AAAI spring symposium","author":"RS Sutton","year":"1991","unstructured":"Sutton RS (1991) Dyna, an integrated architecture for learning, planning, and reacting. In: Working notes of 1991 AAAI spring symposium, pp 151\u2013155"},{"key":"738_CR31","first-page":"216","volume-title":"Proceedings of the seventh international conference on machine learning","author":"RS Sutton","year":"1990","unstructured":"Sutton RS (1990) Integrated architectures for learning, planning, and reaction based on approximating dynamic programming. In: Proceedings of the seventh international conference on machine learning, pp 216\u2013224"},{"key":"738_CR32","doi-asserted-by":"crossref","unstructured":"Sutton RS, Barto AG, Williams RJ (1992) Reinforcement learning is direct adaptive optimal control. IEEE Control Syst Mag, April","DOI":"10.23919\/ACC.1991.4791776"},{"key":"738_CR33","volume-title":"IEEE Canadian conference on electrical and computer engineering","author":"R Hadidi","year":"2009","unstructured":"Hadidi R, Jeyasurya B (2009) Selective initial state criteria to enhance convergence rate of Q-Learning algorithm in power system stability application. In: IEEE Canadian conference on electrical and computer engineering, NL, Canada, May 2009"},{"key":"738_CR34","volume-title":"Design of feedback control systems","author":"RT Stefani","year":"2001","unstructured":"Stefani RT, Savant S, Hostetter et al (2001) Design of feedback control systems, 4th edn. Oxford University Press, London","edition":"4"},{"key":"738_CR35","volume-title":"Machine learning","author":"TM Mitchell","year":"1997","unstructured":"Mitchell TM (1997) Machine learning. McGraw-Hill, New York"},{"issue":"3","key":"738_CR36","doi-asserted-by":"crossref","first-page":"285","DOI":"10.1109\/TITS.2005.853698","volume":"6","author":"X Dai","year":"2005","unstructured":"Dai X, Li C-K, Rad AB (2005) An approach to tune fuzzy controllers based on reinforcement learning for autonomous vehicle control. IEEE Trans Intell Transp Syst 6(3):285\u2013293","journal-title":"IEEE Trans Intell Transp Syst"},{"issue":"1","key":"738_CR37","doi-asserted-by":"crossref","first-page":"526","DOI":"10.1007\/s11227-010-0451-x","volume":"59","author":"HS Al-Dayaa","year":"2012","unstructured":"Al-Dayaa HS, Megherbi DB (2012) Reinforcement learning technique using agent state occurrence frequency with analysis of knowledge sharing on the agent\u2019s learning process in multi-agent environments. J Supercomput 59(1), 526\u2013547. doi: 10.1007\/s11227-010-0451-x","journal-title":"J Supercomput"}],"container-title":["The Journal of Supercomputing"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-011-0738-6.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s11227-011-0738-6\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s11227-011-0738-6","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,6,22]],"date-time":"2019-06-22T14:56:49Z","timestamp":1561215409000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s11227-011-0738-6"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2012,1,17]]},"references-count":37,"journal-issue":{"issue":"1","published-print":{"date-parts":[[2012,10]]}},"alternative-id":["738"],"URL":"https:\/\/doi.org\/10.1007\/s11227-011-0738-6","relation":{},"ISSN":["0920-8542","1573-0484"],"issn-type":[{"value":"0920-8542","type":"print"},{"value":"1573-0484","type":"electronic"}],"subject":[],"published":{"date-parts":[[2012,1,17]]}}}