{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,23]],"date-time":"2026-08-23T15:27:43Z","timestamp":1787498863514,"version":"build-2736575974"},"publisher-location":"Berlin, Heidelberg","reference-count":32,"publisher":"Springer Berlin Heidelberg","isbn-type":[{"value":"9783540423430","type":"print"},{"value":"9783540445814","type":"electronic"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2001]]},"DOI":"10.1007\/3-540-44581-1_41","type":"book-chapter","created":{"date-parts":[[2007,8,10]],"date-time":"2007-08-10T06:13:49Z","timestamp":1186726429000},"page":"616-629","source":"Crossref","is-referenced-by-count":6,"title":["Bounds on Sample Size for Policy Evaluation in Markov Environments"],"prefix":"10.1007","author":[{"given":"Leonid","family":"Peshkin","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Sayan","family":"Mukherjee","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2001,9,13]]},"reference":[{"key":"41_CR1","volume-title":"Proceedings of the Thirteenth Annual Conference on Computational Learning Theory","author":"P. Bartlett","year":"2000","unstructured":"P. Bartlett, S. Boucheron, and G. Lugosi. Model selection and error estimation. In Proceedings of the Thirteenth Annual Conference on Computational Learning Theory. ACM Press, New York, NY, 2000."},{"key":"41_CR2","volume-title":"Dynamic Programming","author":"R. Bellman","year":"1957","unstructured":"Richard Bellman. Dynamic Programming. Princeton University Press, Princeton, New Jersey, 1957."},{"key":"41_CR3","unstructured":"S.N. Bernstein. The Theory of Probability. Gostehizdat, Moscow, 1946."},{"issue":"3","key":"41_CR4","doi-asserted-by":"publisher","first-page":"247","DOI":"10.1023\/A:1010848128995","volume":"43","author":"N. Cesa-Bianchi","year":"2001","unstructured":"Nicol\u00f2 Cesa-Bianchi and G\u00e1bor Lugosi. Worst-case bounds for the logarithmic loss of predictors. Machine Learning, 43(3):247\u2013264, 2001.","journal-title":"Machine Learning"},{"issue":"11","key":"41_CR5","doi-asserted-by":"crossref","first-page":"1367","DOI":"10.1287\/mnsc.35.11.1367","volume":"35","author":"P. Glynn","year":"1989","unstructured":"Peter Glynn. Importance sampling for stochastic simulations. Management Science, 35(11):1367\u20131392, 1989.","journal-title":"Management Science"},{"key":"41_CR6","doi-asserted-by":"publisher","first-page":"78","DOI":"10.1016\/0890-5401(92)90010-D","volume":"100","author":"D. Haussler","year":"1992","unstructured":"D. Haussler. Decision theoretic generalizations of the pac model. Inf. and Comp., 100:78\u2013150, 1992.","journal-title":"Inf. and Comp."},{"key":"41_CR7","doi-asserted-by":"crossref","unstructured":"Leslie Pack Kaelbling, Michael L. Littman, and Andrew W. Moore. Reinforcement learning: A survey. Journal of Artificial Intelligence Research, 4, 1996.","DOI":"10.1613\/jair.301"},{"key":"41_CR8","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1287\/opre.1.5.263","volume":"1","author":"H. Kahn","year":"1953","unstructured":"H. Kahn and A. Marshall. Methods of reducing sample size in Monte Carlo computations. Journal of the Operations Research Society of America, 1:263\u2013278, 1953.","journal-title":"Journal of the Operations Research Society of America"},{"key":"41_CR9","unstructured":"Michael Kearns, Yishay Mansour, and Andrew Y. Ng. Approximate planning in large POMDPs via reusable trajectories. In Advances in Neural Information Processing Systems, 1999."},{"issue":"4","key":"41_CR10","first-page":"245","volume":"2","author":"N. Littlestone","year":"1988","unstructured":"Nick Littlestone. Learning quickly when irrelevant attributes abound: A new linear threshold algorithm. Machine Learning, 2(4):245\u2013318, 1988.","journal-title":"Machine Learning"},{"key":"41_CR11","first-page":"183","volume-title":"Proc. 12th Annu. Conf. on Comput. Learning Theory","author":"Y. Mansour","year":"1999","unstructured":"Yishay Mansour. Reinforcement learning and mistake bounded algorithms. In Proc. 12th Annu. Conf. on Comput. Learning Theory, pages 183\u2013192. ACM Press, New York, NY, 1999."},{"key":"41_CR12","doi-asserted-by":"crossref","unstructured":"C. McDiarmid. Surveys in Combinatorics, chapter On the method of bounded differences, pages 148\u2013188. Cambridge University Press, 1989.","DOI":"10.1017\/CBO9781107359949.008"},{"key":"41_CR13","unstructured":"Nicolas Meuleau, Leonid Peshkin, and Kee-Eung Kim. Exploration in gradientbased reinforcement learning. Technical Report 1713, MIT, 2000."},{"key":"41_CR14","unstructured":"Nicolas Meuleau, Leonid Peshkin, Kee-Eung Kim, and Leslie P. Kaelbling. Learning finite-state controllers for partially observable environments. In Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence, pages 427\u2013436. Morgan Kaufmann, 1999."},{"key":"41_CR15","doi-asserted-by":"crossref","unstructured":"M. Oh and J. Berger. Adaptive importance sampling in Monte Carlo integration, 1992.","DOI":"10.1080\/00949659208810398"},{"key":"41_CR16","unstructured":"M. Opper and D. Haussler. Worst case prediction over sequences in under log loss. In The Mathematics of Information coding, Extraction, and Distribution. Springer Verlag, 1997."},{"key":"41_CR17","unstructured":"Luis Ortiz and Leslie P. Kaelbling. Adaptive importance sampling for estimation in structured domains. In Proceedings of the Sixteenth Conference on Uncertainty in Artificial Intelligence, pages 446\u2013454, San Francisco, CA, 2000. Morgan Kaufmann Publishers."},{"issue":"3","key":"41_CR18","doi-asserted-by":"publisher","first-page":"441","DOI":"10.1287\/moor.12.3.441","volume":"12","author":"C. H. Papadimitriou","year":"1987","unstructured":"Christos H. Papadimitriou and John N. Tsitsiklis. The complexity of Markov decision processes. Mathematics of Operations Research, 12(3):441\u2013450, 1987.","journal-title":"Mathematics of Operations Research"},{"key":"41_CR19","unstructured":"Leonid Peshkin. Architectures for policy search. PhD thesis, Brown University, Providence, RI 02912, 2001. in preparation."},{"key":"41_CR20","unstructured":"Leonid Peshkin, Kee-Eung Kim, Nicolas Meuleau, and Leslie P. Kaelbling. Learning to cooperate via policy search. In Sixteenth Conference on Uncertainty in Artificial Intelligence, pages 307\u2013314, San Francisco, CA, 2000. Morgan Kaufmann."},{"key":"41_CR21","unstructured":"Leonid Peshkin, Nicolas Meuleau, and Leslie P. Kaelbling. Learning policies with external memory. In I. Bratko and S. Dzeroski, editors, Proceedings of the Sixteenth International Conference on Machine Learning, pages 307\u2013314, San Francisco, CA, 1999. Morgan Kaufmann."},{"key":"41_CR22","doi-asserted-by":"crossref","unstructured":"D. Pollard. Convergence of Stochastic Processes. Springer, 1984.","DOI":"10.1007\/978-1-4612-5254-2"},{"key":"41_CR23","doi-asserted-by":"publisher","first-page":"228","DOI":"10.1093\/imamat\/2.3.228","volume":"2","author":"M. Powell","year":"1966","unstructured":"M. Powell and J. Swann. Weighted uniform sampling-a Monte Carlo technique for reducing variance. Journal of the Institute of Mathematics and Applications, 2:228\u2013236, 1966.","journal-title":"Journal of the Institute of Mathematics and Applications"},{"key":"41_CR24","unstructured":"D. Precup, R.S. Sutton, and S. Singh. Eligibility traces for off-policy policy evaluation. In Proceedings of the Seventeenth International Conference on Machine Learning, 2000."},{"key":"41_CR25","doi-asserted-by":"crossref","DOI":"10.1002\/9780470316887","volume-title":"Markov Decision Processes","author":"M.L. Puterman","year":"1994","unstructured":"M.L. Puterman. Markov Decision Processes. John Wiley & Sons, New York, 1994."},{"key":"41_CR26","doi-asserted-by":"crossref","DOI":"10.1002\/9780470316511","volume-title":"Simulation and the Monte Carlo Method","author":"R.Y. Rubinstein","year":"1981","unstructured":"R.Y. Rubinstein. Simulation and the Monte Carlo Method. Wiley, New York, NY, 1981."},{"key":"41_CR27","unstructured":"J. Shtarkov. Universal sequential coding of single measures. Problems of Information Transmission, pages 175\u2013185, 1987."},{"key":"41_CR28","unstructured":"V. Vapnik. Statistical Learning Theory. Wiley, 1998."},{"key":"41_CR29","doi-asserted-by":"crossref","unstructured":"V. Vapnik, E. Levin, and Y. Le Cun. Measuring the VC-dimension of a learning machine. Neural Computation, 1994.","DOI":"10.1162\/neco.1994.6.5.851"},{"issue":"3","key":"41_CR30","first-page":"229","volume":"8","author":"R.J. Williams","year":"1992","unstructured":"R.J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3):229\u2013256, 1992.","journal-title":"Machine Learning"},{"key":"41_CR31","unstructured":"Ronald J. Williams. Reinforcement learning in connectionist networks: A mathematical analysis. Technical Report ICS-8605, Institute for Cognitive Science, University of California, San Diego, La Jolla, California, 1986."},{"key":"41_CR32","unstructured":"Y. Zhou. Adaptive Importance Sampling for Integration. PhD thesis, Stanford University, Palo-Alto, CA, 1998."}],"container-title":["Lecture Notes in Computer Science","Computational Learning Theory"],"original-title":[],"link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/3-540-44581-1_41","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,5,1]],"date-time":"2019-05-01T18:11:40Z","timestamp":1556734300000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/3-540-44581-1_41"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2001]]},"ISBN":["9783540423430","9783540445814"],"references-count":32,"URL":"https:\/\/doi.org\/10.1007\/3-540-44581-1_41","relation":{},"ISSN":["0302-9743"],"issn-type":[{"value":"0302-9743","type":"print"}],"subject":[],"published":{"date-parts":[[2001]]}}}