{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,13]],"date-time":"2026-05-13T13:45:09Z","timestamp":1778679909918,"version":"3.51.4"},"reference-count":31,"publisher":"Springer Science and Business Media LLC","issue":"1","license":[{"start":{"date-parts":[[1993,10,1]],"date-time":"1993-10-01T00:00:00Z","timestamp":749433600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/www.springer.com\/tdm"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Mach Learn"],"published-print":{"date-parts":[[1993,10]]},"DOI":"10.1007\/bf00993104","type":"journal-article","created":{"date-parts":[[2005,1,9]],"date-time":"2005-01-09T16:36:00Z","timestamp":1105288560000},"page":"103-130","source":"Crossref","is-referenced-by-count":271,"title":["Prioritized sweeping: Reinforcement learning with less data and less time"],"prefix":"10.1007","volume":"13","author":[{"given":"Andrew W.","family":"Moore","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Christopher G.","family":"Atkeson","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","reference":[{"key":"CR1","first-page":"35","volume-title":"Connectionist Models: Proceedings of the 1990 Summer School","author":"A.G. Barto","year":"1990","unstructured":"Barto, A.G., & Singh, S.P. (1990). On the computational economics of reinforcement learning. In D.S. Touretzky, J.L. Elman, T.J. Sejnowski, and G.E. Huiton (Eds.),Connectionist Models: Proceedings of the 1990 Summer School. San Mateo, CA: Morgan Kaufmann (pp. 35?44)."},{"key":"CR2","series-title":"COINS Technical Report","volume-title":"Learning and sequential decision making","author":"A.G. Barto","year":"1989","unstructured":"Barto, A.G., Sutton, R.S., & Watkins, C.J.C.H. (1989).Learning and sequential decision making (COINS Technical Report 89?95). Amherst, MA: University of Massachusetts."},{"key":"CR3","series-title":"COINS Technical Report","volume-title":"Real-time learning and control using asynchronous dynamic programming","author":"A.G. Barto","year":"1991","unstructured":"Barto, A.G., Bradtke, S.J., & Singh, S.P. (1991). Real-time learning and control using asynchronous dynamic programming (COINS Technical Report 91-57). Amherst, MA: University of Massachusetts."},{"key":"CR4","volume-title":"Dynamic programming","author":"R.E. Bellman","year":"1957","unstructured":"Bellman, R.E. (1957).Dynamic programming. Princeton, NJ: Princeton University Press."},{"key":"CR5","doi-asserted-by":"crossref","DOI":"10.1007\/978-94-015-3711-7","volume-title":"Bandit problems: Sequential allocation of experiments","author":"D.A. Berry","year":"1985","unstructured":"Berry, D.A., & Fristedt, B. (1985).Bandit problems: Sequential allocation of experiments. New York, NY: Chapman and Hall."},{"key":"CR6","volume-title":"Parallel and distributed computation","author":"D.P. Bertsekas","year":"1989","unstructured":"Bertsekas, D.P., & Tsitsiklis, J.N. (1989).Parallel and distributed computation. Englewood Cliffs, NJ: Prentice Hall."},{"key":"CR7","series-title":"Technical Report","volume-title":"Learning from delayed reinforcement in a complex domain","author":"D. Chapman","year":"1990","unstructured":"Chapman, D., & Kaelbling, L.P. (1990).Learning from delayed reinforcement in a complex domain (Technical Report No. TR-90-11). Teleos Research, Palo Alto, CA."},{"key":"CR8","doi-asserted-by":"crossref","first-page":"1224","DOI":"10.1109\/ROBOT.1990.126165","volume-title":"IEEE Conference on Robotics and Automation","author":"A.D. Christiansen","year":"1990","unstructured":"Christiansen, A.D., Mason, M.T., & Mitchell, T.M. (1990). Learning reliable manipulation strategies without initial physical models. InIEEE Conference on Robotics and Automation (pp. 1224?1230). IEEE Computer Society Press, Washington, DC."},{"issue":"3","key":"CR9","first-page":"341","volume":"8","author":"P. Dayan","year":"1992","unstructured":"Dayan, P. (1992). The convergence of TD(?) for general ?.Machine Learning, 8(3), 341?362.","journal-title":"Machine Learning"},{"key":"CR10","series-title":"Technical Report","volume-title":"Learning in embedded systems","author":"L.P. Kaelbling","year":"1990","unstructured":"Kaelbling, L.P. (1990).Learning in embedded systems. PhD. thesis, Department of Computer Science, Stanford University, Stanford CA. (Technical Report No. TR-90-04.)"},{"key":"CR11","volume-title":"Sorting and searching","author":"D.E. Knuth","year":"1973","unstructured":"Knuth, D.E. (1973).Sorting and searching. Reading, MA: Addison Wesley."},{"key":"CR12","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1016\/0004-3702(90)90054-4","volume":"42","author":"R.E. Korf","year":"1990","unstructured":"Korf, R.E. (1990). Real-time heuristic search.Artificial Intelligence, 42, 189?211.","journal-title":"Artificial Intelligence"},{"key":"CR13","volume-title":"Proceedings of the Ninth International Conference on Artificial Intelligence (AAAI-91)","author":"L.J. Lin","year":"1991","unstructured":"Lin, L.J. (1991). Programming robots using reinforcement learning and teaching. InProceedings of the Ninth International Conference on Artificial Intelligence (AAAI-91). Cambridge, MA: MIT Press."},{"key":"CR14","series-title":"Technical Report","volume-title":"Automatic programming of behavior-based robots using reinforcement learning","author":"S. Mahadevan","year":"1990","unstructured":"Mahadevan, S., & Connell, J. (1990). Automatic programming of behavior-based robots using reinforcement learning (Technical Report). IBM T.J. Watson Research Center. Yorktown Heights, NY."},{"key":"CR15","first-page":"137","volume-title":"Machine intelligence 2","author":"D. Michie","year":"1968","unstructured":"Michie, D., & Chambers, R.A. (1968). BOXES: An experiment in adaptive control. In E. Dale and D. Michie (Eds.),Machine intelligence 2. London: Oliver and Boyd, pp. 137?152."},{"key":"CR16","unstructured":"Moore, A.W., & Atkeson, C.G. (1992). Memory-based function approximators for learning control. In preparation."},{"key":"CR17","first-page":"333","volume-title":"Machine learning: Proceedings of the eighth international workshop","author":"A.W. Moore","year":"1991","unstructured":"Moore, A.W. (1991). Variable resolution dynamic programming: efficiently learning action maps in multivariate real-valued state-spaces. In L. Birnbaum & G. Collins (Eds.),Machine learning: Proceedings of the eighth international workshop. San Mateo, CA: Morgan Kaufman, pp. 333?337."},{"key":"CR18","volume-title":"Problem solving methods in artificial intelligence","author":"N.J. Nilsson","year":"1971","unstructured":"Nilsson, N.J. (1971).Problem solving methods in artificial intelligence. New York: McGraw Hill."},{"key":"CR19","volume-title":"Efficient search control in Dyna","author":"J. Peng","year":"1992","unstructured":"Peng, J. & Williams, R.J. (1992).Efficient search control in Dyna. College of Computer Science, Northeastern University, Boston, MA. (A revised version will appear as ?Efficient learning and planning within the dyna framework.?Proceedings of the Second International Conference on Simulation of Adaptive Behavior. Cambridge, MA: MIT Press, 1993.)"},{"key":"CR20","volume-title":"Optimum systems control","author":"A.P. Sage","year":"1977","unstructured":"Sage, A.P., & White, C.C. (1977).Optimum systems control. Englewood Cliffs, NJ: Prentice Hall."},{"issue":"3","key":"CR21","doi-asserted-by":"crossref","first-page":"210","DOI":"10.1147\/rd.33.0210","volume":"3","author":"A.L. Samuel","year":"1959","unstructured":"Samuel, A.L. (1959). Some studies in machine learning using the game of checkers.IBM Journal on Research and Development, 3, (3)210?229. Reprinted in E.A. Feigenbaum & J. Feldman (Eds.). (1963).Computers and thought. New York: McGraw-Hill, pp. 71?105.","journal-title":"IBM Journal on Research and Development"},{"issue":"5","key":"CR22","doi-asserted-by":"crossref","first-page":"667","DOI":"10.1109\/21.21595","volume":"18","author":"M. Sato","year":"1988","unstructured":"Sato, M., Abe, K., & Takeda, H. (1988). Learning control of finite Markov chains with an explicit trade-off between estimation and control.IEEE Transactions on Systems, Man, and Cybernetics, 18(5), 667?684.","journal-title":"IEEE Transactions on Systems, Man, and Cybernetics"},{"key":"CR23","doi-asserted-by":"crossref","unstructured":"Singh, S.P. (1991). Transfer of learning across compositions of sequential tasks. In L. Birnbaum & G. Collins (EDs.).Machine learning: Proceedings of the eighth international workshop. Morgan Kaufman, pp. 348?352.","DOI":"10.1016\/B978-1-55860-200-7.50072-6"},{"issue":"12","key":"CR24","doi-asserted-by":"crossref","first-page":"1213","DOI":"10.1145\/7902.7906","volume":"29","author":"C. Stanfill","year":"1986","unstructured":"Stanfill, C. & Waltz, D. (1986). Towards memory-based reasoning.Communications of the ACM, 29(12), 1213?1228.","journal-title":"Communications of the ACM"},{"key":"CR25","first-page":"497","volume-title":"Learning and computational neuroscience: Foundations of adaptive networks","author":"R.S. Sutton","year":"1990","unstructured":"Sutton, R.S., & Barto, A.G. (1990). Time-derivative models of Pavlovian reinforcement. In M. Gabriel & J. Moore (Eds.),Learning and computational neuroscience: Foundations of adaptive networks (pp. 497?537). Cambridge, MA: MIT Press."},{"key":"CR26","volume-title":"Temporal credit assignment in reinforcement learning","author":"R.S. Sutton","year":"1984","unstructured":"Sutton, R.S. (1984).Temporal credit assignment in reinforcement learning. Ph.D. thesis, Department of Computer and Information Sciences, University of Massachusetts, Amherst."},{"key":"CR27","first-page":"9","volume":"3","author":"R.S. Sutton","year":"1988","unstructured":"Sutton, R.S. (1988). Learning to predict by the methods of temporal differences.Machine Learning, 3, 9?44.","journal-title":"Machine Learning"},{"key":"CR28","volume-title":"Proceedings of the 7th International Conference on Machine Learning","author":"R.S. Sutton","year":"1990","unstructured":"Sutton, R.S. (1990). Integrated architecture for learning, planning, and reacting based on approximating dynamic programming. InProceedings of the 7th International Conference on Machine Learning. San Mateo, CA: Morgan Kaufman."},{"key":"CR29","volume-title":"Practical issues in temporal difference learning. Report RC 17223 (76307)","author":"G.J. Tesauro","year":"1991","unstructured":"Tesauro, G.J. (1991). Practical issues in temporal difference learning. Report RC 17223 (76307). IBM T.J. Watson Research Center, Yorktown Heights, NY."},{"key":"CR30","first-page":"531","volume-title":"Advances in neural information processing systems 4","author":"S.B. Thrun","year":"1992","unstructured":"Thrun, S.B., & M\u00f6ller, K. (1992). Active exploration in dynamic environments. In J.E. Moody, S.J. Hanson, & R.P. Lippman (Eds.),Advances in neural information processing systems 4. San Mateo, CA: Morgan Kaufmann, pp. 531?538."},{"key":"CR31","volume-title":"Learning from delayed rewards","author":"C.J.C.H. Watkins","year":"1989","unstructured":"Watkins, C.J.C.H. (1989).Learning from delayed rewards. Ph.D. thesis, King's College, University of Cambridge, United Kingdom."}],"container-title":["Machine Learning"],"original-title":[],"language":"en","link":[{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/BF00993104.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/article\/10.1007\/BF00993104\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"http:\/\/link.springer.com\/content\/pdf\/10.1007\/BF00993104","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2019,4,29]],"date-time":"2019-04-29T22:58:40Z","timestamp":1556578720000},"score":1,"resource":{"primary":{"URL":"http:\/\/link.springer.com\/10.1007\/BF00993104"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[1993,10]]},"references-count":31,"journal-issue":{"issue":"1","published-print":{"date-parts":[[1993,10]]}},"alternative-id":["BF00993104"],"URL":"https:\/\/doi.org\/10.1007\/bf00993104","relation":{},"ISSN":["0885-6125","1573-0565"],"issn-type":[{"value":"0885-6125","type":"print"},{"value":"1573-0565","type":"electronic"}],"subject":[],"published":{"date-parts":[[1993,10]]}}}