{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2023,5,3]],"date-time":"2023-05-03T22:14:25Z","timestamp":1683152065566},"reference-count":20,"publisher":"Springer Science and Business Media LLC","issue":"3","license":[{"start":{"date-parts":[[2007,7,24]],"date-time":"2007-07-24T00:00:00Z","timestamp":1185235200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Discrete Event Dyn Syst"],"published-print":{"date-parts":[[2007,8,27]]},"DOI":"10.1007\/s10626-007-0014-3","type":"journal-article","created":{"date-parts":[[2007,7,23]],"date-time":"2007-07-23T17:42:32Z","timestamp":1185212552000},"page":"307-327","source":"Crossref","is-referenced-by-count":6,"title":["Efficient PAC Learning for Episodic Tasks with Acyclic State Spaces"],"prefix":"10.1007","volume":"17","author":[{"given":"Spyros","family":"Reveliotis","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Theologos","family":"Bountourelis","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2007,7,24]]},"reference":[{"key":"14_CR1","doi-asserted-by":"crossref","first-page":"16","DOI":"10.1214\/aoms\/1177728845","volume":"25","author":"RE Bechhofer","year":"1954","unstructured":"Bechhofer RE (1954) A single-sample multiple decision procedure for ranking means of normal populations with known variances. Ann Math Stat 25:16\u201339","journal-title":"Ann Math Stat"},{"key":"14_CR2","volume-title":"Neuro-dynamic programming","author":"DP Bertsekas","year":"1996","unstructured":"Bertsekas DP, Tsitsiklis JN (1996) Neuro-dynamic programming. Athena Scientific, Belmont"},{"key":"14_CR3","first-page":"255","volume-title":"Proceedings of COLT\u201902","author":"E Even-Dar","year":"2002","unstructured":"Even-Dar E, Mannor S, Mansour Y (2002) PAC bounds for multi-armed bandit and Markov decision processes. In: Proceedings of COLT\u201902. ACM, New York, pp 255\u2013270"},{"key":"14_CR4","volume-title":"An introduction to probability theory and its applications, vol. II","author":"W Feller","year":"1971","unstructured":"Feller W (1971) An introduction to probability theory and its applications, vol. II, 2nd edn. Wiley, New York","edition":"2"},{"key":"14_CR5","first-page":"88","volume-title":"Proceedings of COLT\u201994","author":"CN Fiechter","year":"1994","unstructured":"Fiechter CN (1994) Efficient reinforcement learning. In: Proceedings of COLT\u201994. ACM, New York, pp 88\u201397"},{"key":"14_CR6","first-page":"116","volume-title":"Proceedings of ICML\u201997","author":"CN Fiechter","year":"1997","unstructured":"Fiechter CN (1997) Expected mistake bound model for on-line reinforcement learning. In: Proceedings of ICML\u201997. AAAI, Menlo Park, pp 116\u2013124"},{"key":"14_CR7","volume-title":"Operations management","author":"J Heizer","year":"2004","unstructured":"Heizer J, Render B (2004) Operations management, 7th edn. Pearson\/Prentice Hall, Upper Saddle River","edition":"7"},{"key":"14_CR8","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1080\/01621459.1963.10500830","volume":"58","author":"W Hoeffding","year":"1963","unstructured":"Hoeffding W (1963) Probability inequalities for sum of bounded random variables. J Am Stat Assoc 58:13\u201330","journal-title":"J Am Stat Assoc"},{"key":"14_CR9","first-page":"996","volume":"11","author":"M Kearns","year":"1999","unstructured":"Kearns M, Singh S (1999) Finite-sample convergence rates for Q-learning and indirect algorithms. Neural Inf Process Syst 11:996\u20131002","journal-title":"Neural Inf Process Syst"},{"key":"14_CR10","doi-asserted-by":"crossref","first-page":"209","DOI":"10.1023\/A:1017984413808","volume":"49","author":"M Kearns","year":"2002","unstructured":"Kearns M, Singh S (2002) Near-optimal reinforcement learning in polynomial time. Mach Learn 49:209\u2013232","journal-title":"Mach Learn"},{"key":"14_CR11","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/3897.001.0001","volume-title":"An introduction to computational learning theory","author":"MJ Kearns","year":"1994","unstructured":"Kearns MJ, Vazirani UV (1994) An introduction to computational learning theory. MIT Press, Cambridge"},{"key":"14_CR12","unstructured":"Kim S-H, Nelson BL (2004) Selecting the best system. Technical report, School of Industrial & Systems Eng., Georgia Tech"},{"key":"14_CR13","volume-title":"Machine learning","author":"TM Mitchell","year":"1997","unstructured":"Mitchell TM (1997) Machine learning. McGraw Hill, London"},{"key":"14_CR14","first-page":"135","volume-title":"Proceedings of the NSF\u2013IEEE\u2013ORSI international workshop on IT-enabled manufacturing, logistics and supply chain management","author":"SA Reveliotis","year":"2003","unstructured":"Reveliotis SA (2003) Uncertainty management in optimal disassembly planning through learning-based strategies. In: Proceedings of the NSF\u2013IEEE\u2013ORSI international workshop on IT-enabled manufacturing, logistics and supply chain management. NSF\/IEEE\/ORSI, Piscataway, pp 135\u2013141"},{"key":"14_CR15","first-page":"2625","volume-title":"IEEE international conference on robotics & automation","author":"SA Reveliotis","year":"2004","unstructured":"Reveliotis SA (2004) Modelling and controlling uncertainty in optimal disassembly planning through reinforcement learning. In: IEEE international conference on robotics & automation. IEEE, Piscataway, pp 2625\u20132632"},{"key":"14_CR16","doi-asserted-by":"crossref","first-page":"645","DOI":"10.1080\/07408170600897536","volume":"39","author":"SA Reveliotis","year":"2007","unstructured":"Reveliotis SA (2007) Uncertainty management in optimal disassembly planning through learning-based strategies. IIE Trans 39:645\u2013658","journal-title":"IIE Trans"},{"key":"14_CR17","first-page":"421","volume-title":"Proceedings of the 2006 IEEE international conference on automation science and engineering","author":"SA Reveliotis","year":"2006","unstructured":"Reveliotis SA, Bountourelis T (2006) Efficient learning algorithms for episodic tasks with acyclic state spaces. In: Proceedings of the 2006 IEEE international conference on automation science and engineering. IEEE, Piscataway, pp 421\u2013428"},{"key":"14_CR18","volume-title":"Reinforcement learning","author":"RS Sutton","year":"2000","unstructured":"Sutton RS, Barto AG (2000) Reinforcement learning. MIT Press, Cambridge"},{"key":"14_CR19","volume-title":"Probabilistic robotics","author":"S Thrun","year":"2005","unstructured":"Thrun S, Burgard W, Fox D (2005) Probabilistic robotics. MIT Press, Cambridge"},{"key":"14_CR20","unstructured":"Watkins CJCH (1989) Learning from delayed rewards. Ph.D. thesis, Cambridge University"}],"container-title":["Discrete Event Dynamic Systems"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s10626-007-0014-3.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/s10626-007-0014-3\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/s10626-007-0014-3","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,5,30]],"date-time":"2019-05-30T19:58:51Z","timestamp":1559246331000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/s10626-007-0014-3"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2007,7,24]]},"references-count":20,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2007,8,27]]}},"alternative-id":["14"],"URL":"https:\/\/doi.org\/10.1007\/s10626-007-0014-3","relation":{},"ISSN":["0924-6703","1573-7594"],"issn-type":[{"value":"0924-6703","type":"print"},{"value":"1573-7594","type":"electronic"}],"subject":[],"published":{"date-parts":[[2007,7,24]]}}}