{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,23]],"date-time":"2026-07-23T05:23:21Z","timestamp":1784784201759,"version":"3.55.0"},"reference-count":69,"publisher":"American Association for the Advancement of Science (AAAS)","issue":"116","content-domain":{"domain":["www.science.org"],"crossmark-restriction":true},"short-container-title":["Sci. Robot."],"published-print":{"date-parts":[[2026,7,22]]},"abstract":"<jats:p>Real-world robotic manipulation in homes and factories demands reliability, efficiency, and robustness that approach or surpass skilled human operators. We present a real-world reinforcement learning (RL) framework, RL-100, for achieving complete task success under a predefined evaluation protocol built on diffusion visuomotor policies. RL-100 unifies imitation and RL under a single clipped proximal policy optimization surrogate objective applied in the denoising process, yielding conservative, stable improvements across offline and online stages. To meet deployment latency, a lightweight consistency distillation compresses multistep diffusion into a one-step controller for high-frequency control. The framework is task, embodiment, and representation agnostic and supports both single-action and action-chunking control. We evaluated RL-100 on eight diverse real-robot tasks, from pushing and bowling to pouring, cloth folding, unscrewing, multistage juicing, and long-horizon box folding. Under our predefined protocol, RL-100 achieved 100% success in the evaluated trials (1000 of 1000 episodes), including up to 250 of 250 consecutive trials on one task. It matched or surpassed expert teleoperators in time to completion. Without retraining, a single policy attained \u223c90% zero-shot success under environmental and dynamics shifts, adapted in a few-shot regime to substantial task variations (86.7%), and remained robust to human perturbations (about 96%). Our juicing robot served customers continuously for about 7 hours without failure when deployed zero-shot in a shopping mall. These results suggest a potential path to deployable robot learning by starting from human priors, aligning training objectives with human-grounded metrics, and reliably extending performance beyond human demonstrations.<\/jats:p>","DOI":"10.1126\/scirobotics.aed6267","type":"journal-article","created":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T17:58:13Z","timestamp":1784743093000},"update-policy":"https:\/\/doi.org\/10.34133\/aaas_crossmark","source":"Crossref","is-referenced-by-count":1,"title":["Performant robotic manipulation with real-world reinforcement learning"],"prefix":"10.1126","volume":"11","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-0758-1488","authenticated-orcid":true,"given":"Kun","family":"Lei","sequence":"first","affiliation":[{"name":"Shanghai Qi Zhi Institute, Shanghai, China."},{"name":"School of Computer Science, Shanghai Jiao Tong University, Shanghai, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-3812-8950","authenticated-orcid":true,"given":"Huanyu","family":"Li","sequence":"additional","affiliation":[{"name":"Shanghai Qi Zhi Institute, Shanghai, China."},{"name":"School of Computer Science, Shanghai Jiao Tong University, Shanghai, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3616-5400","authenticated-orcid":true,"given":"Dongjie","family":"Yu","sequence":"additional","affiliation":[{"name":"Shanghai Qi Zhi Institute, Shanghai, China."},{"name":"School of Computing and Data Science, University of Hong Kong, Hong Kong SAR, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-6218-0242","authenticated-orcid":true,"given":"Zhenyu","family":"Wei","sequence":"additional","affiliation":[{"name":"Department of Computer Science, University of North Carolina at Chapel Hill, Chapel Hill, NC, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-6888-2019","authenticated-orcid":true,"given":"Lingxiao","family":"Guo","sequence":"additional","affiliation":[{"name":"School of Computer Science, Carnegie Mellon University, Pittsburgh, PA, USA."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0005-0515-4126","authenticated-orcid":true,"given":"Zhennan","family":"Jiang","sequence":"additional","affiliation":[{"name":"Institute of Automation, Chinese Academy of Sciences, Beijing, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0003-4200-4763","authenticated-orcid":true,"given":"Ziyu","family":"Wang","sequence":"additional","affiliation":[{"name":"Institute for Interdisciplinary Information Sciences, Tsinghua University, Beijing, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shiyu","family":"Liang","sequence":"additional","affiliation":[{"name":"School of Computer Science, Shanghai Jiao Tong University, Shanghai, China."}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0001-8642-3519","authenticated-orcid":true,"given":"Huazhe","family":"Xu","sequence":"additional","affiliation":[{"name":"Shanghai Qi Zhi Institute, Shanghai, China."},{"name":"Institute for Interdisciplinary Information Sciences, Tsinghua University, Beijing, China."}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"221","reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.abd9461"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1126\/scirobotics.ads5033"},{"key":"e_1_3_2_4_2","doi-asserted-by":"publisher","DOI":"10.1177\/02783649241273668"},{"key":"e_1_3_2_5_2","doi-asserted-by":"crossref","unstructured":"Y. Ze G. Zhang K. Zhang C. Hu M. Wang H. Xu \u201c3D diffusion policy: Generalizable visuomotor policy learning via simple 3D representations\u201d in Proceedings of Robotics: Science and Systems D. Kulic G. Venture K. Bekris E Coronado Eds. (RSS Foundation 2024) 10.15607\/RSS.2024.XX.067.","DOI":"10.15607\/RSS.2024.XX.067"},{"key":"e_1_3_2_6_2","doi-asserted-by":"crossref","unstructured":"K. Black N. Brown D. Driess A. Esmail M. Equi C. Finn N. Fusai L. Groom K. Hausman B. Ichter S. Jakubczak T. Jones L. Ke S. Levine A. Li-Bell M. Mothukuri S. Nair K. Pertsch L. X. Shi J. Tanner Q. Vuong A. Walling H. Wang U. Zhilinsky \u03c00: A vision-language-action flow model for general robot control. arXiv:2410.24164 [cs.LG] (2024).","DOI":"10.15607\/RSS.2025.XXI.010"},{"key":"e_1_3_2_7_2","doi-asserted-by":"crossref","unstructured":"K. Black N. Brown J. Darpinian K. Dhabalia D. Driess A. Esmail M. Equi C. Finn N. Fusai M. Y. Galliker D. Ghosh L. Groom K. Hausman B. Ichter S. Jakubczak T. Jones L. Ke D. LeBlanc S. Levine A. Li-Bell M. Mothukuri S. Nair K. Pertsch A. Z. Ren L. X. Shi L. Smith J. T. Springenberg K. Stachowicz J. Tanner Q. Vuong H. Walke A. Walling H. Wang L. Yu U. Zhilinsky \u201c\u03c00.5: A vision-language-action model with open-world generalization\u201d in Proceedings of the 9th Conference on Robot Learning J. Lim S. Song H.-W. Park Eds. vol. 305 of Proceedings of Machine Learning Research (PMLR 2025) pp. 17\u201340.","DOI":"10.15607\/RSS.2025.XXI.010"},{"key":"e_1_3_2_8_2","unstructured":"Z. Yuan T. Wei L. Gu P. Hua T. Liang Y. Chen H. Xu HERMES: Human-to-robot embodied learning from multi-source motion data for mobile dexterous manipulation. arXiv:2508.20085 [cs.RO] (2025)."},{"key":"e_1_3_2_9_2","unstructured":"T. Lin K. Sachdev L. Fan J. Malik Y. Zhu \u201cSim-to-real reinforcement learning for vision-based dexterous manipulation on humanoids\u201d in Proceedings of the 9th Conference on Robot Learning J. Lim S. Song H.-W. Park Eds. vol. 305 of Proceedings of Machine Learning Research (PMLR 2025) pp. 4926\u20134940."},{"key":"e_1_3_2_10_2","doi-asserted-by":"crossref","unstructured":"J. Luo Z. Hu C. Xu Y. L. Tan J. Berg A. Sharma S. Schaal C. Finn A. Gupta S. Levine \u201cSERL: A software suite for sample-efficient robotic reinforcement learning\u201d in 2024 IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2024) pp. 16961\u201316969.","DOI":"10.1109\/ICRA57147.2024.10610040"},{"key":"e_1_3_2_11_2","unstructured":"L. Guo Z. Xue Z. Xu H. Xu \u201cDemoSpeedup: Accelerating visuomotor policies via entropy-guided demonstration acceleration\u201d in Proceedings of the 9th Conference on Robot Learning J. Lim S. Song H.-W. Park Eds. vol. 305 of Proceedings of Machine Learning Research (PMLR 2025) pp. 599\u2013609."},{"key":"e_1_3_2_12_2","unstructured":"J. Song C. Meng S. Ermon \u201cDenoising diffusion implicit models\u201d in International Conference on Learning Representations (ICLR 2021); https:\/\/openreview.net\/forum?id=St1giarCHLP."},{"key":"e_1_3_2_13_2","unstructured":"Y. Song P. Dhariwal M. Chen I. Sutskever \u201cConsistency models\u201d in Proceedings of the 40th International Conference on Machine Learning (ICML) A. Krause E. Brunskill K. Cho B. Engelhardt S. Sabato J. Scarlett Eds. vol. 202 of Proceedings of Machine Learning Research (PMLR 2023) pp. 32211\u201332252."},{"key":"e_1_3_2_14_2","doi-asserted-by":"crossref","unstructured":"T. Z. Zhao V. Kumar S. Levine C. Finn \u201cLearning fine-grained bimanual manipulation with low-cost hardware\u201d in Proceedings of Robotics: Science and Systems K. Bekris K. Hauser S. Herbert J. Yu (RSS Foundation 2023) 10.15607\/RSS.2023.XIX.016.","DOI":"10.15607\/RSS.2023.XIX.016"},{"key":"e_1_3_2_15_2","unstructured":"S. Mani S. Venkataraman A. Chandra A. Rizvi Y. Sirvi S. Bhattacharya A. Hazra DiffClone: Enhanced behaviour cloning in robotics with diffusion-driven policy learning. arXiv:2401.09243 [cs.RO] (2024)."},{"key":"e_1_3_2_16_2","unstructured":"S. F. Chen H. C. Wang M. H. Hsu C. M. Lai S. H. Sun \u201cDiffusion model-augmented behavioral cloning\u201d in Proceedings of the 41st International Conference on Machine Learning (ICML) R. Salakhutdinov Z. Kolter K. Heller A. Weller N. Oliver J. Scarlett F. Berkenkamp Eds. vol. 235 of Proceedings of Machine Learning Research (PMLR 2024) pp. 7486\u20137510."},{"key":"e_1_3_2_17_2","unstructured":"S. Nair A. Rajeswaran V. Kumar C. Finn A. Gupta \u201cR3M: A universal visual representation for robot manipulation\u201d in Proceedings of the 6th Conference on Robot Learning K. Liu D. Kulic J. Ichnowski Eds. vol. 205 of Proceedings of Machine Learning Research (PMLR 2023) pp. 892\u2013909."},{"key":"e_1_3_2_18_2","doi-asserted-by":"crossref","unstructured":"A. Majumdar K. Yadav S. Arnaud Y. Jason Ma C. Chen S. Silwal A. Jain V. P. Berges T. Wu J. Vakil P. Abbeel J. Malik D. Batra Y. Lin O. Maksymets A. Rajeswaran F. Meier \u201cWhere are we in the search for an artificial visual cortex for embodied intelligence?\u201d in Advances in Neural Information Processing Systems 36 (NeurIPS) A. Oh T. Naumann A. Globerson K. Saenko M. Hardt S. Levine Eds. (Curran Associates 2023) pp. 655\u2013677.","DOI":"10.52202\/075280-0031"},{"key":"e_1_3_2_19_2","doi-asserted-by":"crossref","unstructured":"K. He H. Fan Y. Wu S. Xie R. Girshick \u201cMomentum contrast for unsupervised visual representation learning\u201d in 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (IEEE 2020) pp. 9726\u20139735.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_20_2","doi-asserted-by":"crossref","unstructured":"Q. Zhang Z. Liu H. Fan G. Liu B. Zeng S. Liu \u201cFlowPolicy: Enabling fast and robust 3D flow-based policy via consistency flow matching for robot manipulation\u201d in Proceedings of the 39-th AAAI Conference on Artificial Intelligence (AAAI) (AAAI 2025) pp. 14754\u201314762.","DOI":"10.1609\/aaai.v39i14.33617"},{"key":"e_1_3_2_21_2","unstructured":"Y. Lu Y. Tian Z. Yuan X. Wang P. Hua Z. Xue H. Xu H3DP: Triply-hierarchical diffusion policy for visuomotor learning. arXiv:2505.07819 [cs.RO] (2025)."},{"key":"e_1_3_2_22_2","unstructured":"M. Janner Y. Du J. B. Tenenbaum S. Levine \u201cPlanning with diffusion for flexible behavior synthesis\u201d in Proceedings of the 39th International Conference on Machine Learning (ICML) K. Chaudhuri S. Jegelka L. Song C. Szepesvari G. Niu S. Sabato Eds. vol. 162 of Proceedings of Machine Learning Research (PMLR 2022) pp. 9902\u20139915."},{"key":"e_1_3_2_23_2","unstructured":"A. Ajay Y. Du A. Gupta J. Tenenbaum T. Jaakkola P. Agrawal \u201cIs conditional generative modeling all you need for decision making?\u201d in 11th International Conference on Learning Representations (ICLR) (ICLR 2023); https:\/\/openreview.net\/forum?id=sP1fo2K9DFG."},{"key":"e_1_3_2_24_2","unstructured":"S. Hegde S. Das G. Salhotra G. S. Sukhatme Latent weight diffusion: Generating reactive policies instead of trajectories. arXiv:2410.14040 [cs.LG] (2024)."},{"key":"e_1_3_2_25_2","unstructured":"Z. Wang M. Li A. Mandlekar Z. Xu J. Fan Y. Narang L. Fan Y. Zhu Y. Balaji M. Zhou M. Y. Liu Y. Zeng \u201cOne-step diffusion policy: Fast visuomotor policies via diffusion distillation\u201d in Proceedings of the 42nd International Conference on Machine Learning (ICML) A. Singh M. Fazel D. Hsu S. Lacoste-Julien F. Berkenkamp T. Maharaj K. Wagstaff J. Zhu Eds. vol. 267 of Proceedings of Machine Learning Research (PMLR 2025) pp. 63399\u201363416."},{"key":"e_1_3_2_26_2","unstructured":"A. Bartsch A. Car A. B. Farimani PinchBot: Long-horizon deformable manipulation with guided diffusion policy. arXiv:2507.17846 [cs.RO] (2025)."},{"key":"e_1_3_2_27_2","unstructured":"H. R. Walke K. Black T. Z. Zhao Q. Vuong C. Zheng P. Hansen-Estruch A. W. He V. Myers M. J. Kim M. Du A. Lee K. Fang C. Finn S. Levine \u201cBridgeData V2: A dataset for robot learning at scale\u201d in Proceedings of the 7th Conference on Robot Learning J. Tan M. Toussaint K. Darvish Eds. vol. 229 of Proceedings of Machine Learning Research (PMLR 2023) pp. 1723\u20131736."},{"key":"e_1_3_2_28_2","first-page":"1334","article-title":"End-to-end training of deep visuomotor policies","volume":"17","author":"Levine S.","year":"2016","unstructured":"S. Levine, C. Finn, T. Darrell, P. Abbeel, End-to-end training of deep visuomotor policies. J. Mach. Learn. Res. 17, 1334\u20131373 (2016).","journal-title":"J. Mach. Learn. Res."},{"key":"e_1_3_2_29_2","unstructured":"D. Kalashnikov A. Irpan P. Pastor J. Ibarz A. Herzog E. Jang D. Quillen E. Holly M. Kalakrishnan V. Vanhoucke S. Levine \u201cScalable deep reinforcement learning for vision-based robotic manipulation\u201d in Proceedings of the 2nd Conference on Robot Learning A. Billard A. Dragan J. Peters J. Morimoto Eds. vol. 87 of Proceedings of Machine Learning Research (PMLR 2018) pp.651\u2013673."},{"key":"e_1_3_2_30_2","unstructured":"T. Haarnoja A. Zhou P. Abbeel S. Levine \u201cSoft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor\u201d in Proceedings of the 35th International Conference on Machine Learning (ICML) J. Dy A. Krause Eds. vol. 80 of Proceedings of Machine Learning Research (PMLR 2018) pp. 1861\u20131870."},{"key":"e_1_3_2_31_2","unstructured":"S. Fujimoto H. Hoof D. Meger \u201cAddressing function approximation error in actor-critic methods\u201d in Proceedings of the 35th International Conference on Machine Learning (ICML) J. Dy A. Krause Eds. vol. 80 of Proceedings of Machine Learning Research (PMLR 2018) pp. 1587\u20131596."},{"key":"e_1_3_2_32_2","unstructured":"K. Chua R. Calandra R. McAllister S. Levine \u201cDeep reinforcement learning in a handful of trials using probabilistic dynamics models\u201d in Advances in Neural Information Processing Systems 31 (NeurIPS) S. Bengio H. Wallach H. Larochelle K. Grauman N. Cesa-Bianchi R. Garnett Eds. (Curran Associates 2018) pp. 4754\u20134765."},{"key":"e_1_3_2_33_2","unstructured":"M. Janner J. Fu M. Zhang S. Levine \u201cWhen to trust your model: Model-based policy optimization\u201d in Advances in Neural Information Processing Systems 32 (NeurIPS) H. Wallach H. Larochelle A. Beygelzimer F. d\u2019Alch\u00e9-Buc E. Fox R. Garnett Eds. (Curran Associates 2019) pp. 12519\u201312530."},{"key":"e_1_3_2_34_2","unstructured":"B. Eysenbach S. Gu J. Ibarz S. Levine \u201cLeave no trace: Learning to reset for safe and autonomous reinforcement learning\u201d in 6th International Conference on Learning Representations (ICLR) (ICLR 2018); https:\/\/openreview.net\/forum?id=S1vuO-bCW."},{"key":"e_1_3_2_35_2","doi-asserted-by":"crossref","unstructured":"A. Gupta J. Yu T. Z. Zhao V. Kumar A. Rovinsky K. Xu T. Devlin S. Levine \u201cReset-free reinforcement learning via multi-task learning: Learning dexterous manipulation behaviors without human intervention\u201d in IEEE International Conference on Robotics and Automation (ICRA) (IEEE 2021) pp. 6664\u20136671.","DOI":"10.1109\/ICRA48506.2021.9561384"},{"key":"e_1_3_2_36_2","doi-asserted-by":"crossref","unstructured":"A. Singh L. Yang K. Hartikainen C. Finn S. Levine \u201cEnd-to-end robotic reinforcement learning without reward engineering\u201d in Robotics: Science and Systems A. Bicchi H. Kress-Gazit S. Hutchinson Eds. (RSS Foundation 2019) 10.15607\/RSS.2019.XV.073.","DOI":"10.15607\/RSS.2019.XV.073"},{"key":"e_1_3_2_37_2","unstructured":"P. F. Christiano J. Leike T. B. Brown M. Martic S. Legg D. Amodei \u201cDeep reinforcement learning from human preferences\u201d in Advances in Neural Information Processing Systems 30 (NeurIPS) I. Guyon U. Von Luxburg S. Bengio H. Wallach R. Fergus S. Vishwanathan R. Garnett Eds. (Curran Associates 2017) pp. 4299\u20134307."},{"key":"e_1_3_2_38_2","unstructured":"K. Lei Z. He C. Lu K. Hu Y. Gao H. Xu \u201cUni-O4: Unifying online and offline deep reinforcement learning with multi-step on-policy optimization\u201d in 12th International Conference on Learning Representations (ICLR) B. Kim Y. Yue S. Chaudhuri K. Fragkiadaki M. Khan Y. Sun Eds. (ICLR 2024) pp. 32264\u201332297."},{"key":"e_1_3_2_39_2","unstructured":"Z. Zhuang K. Lei J. Liu D. Wang Y. Guo \u201cBehavior proximal policy optimization\u201d in 11th International Conference on Learning Representations (ICLR) (ICLR 2023); https:\/\/openreview.net\/forum?id=3c13LptpIph."},{"key":"e_1_3_2_40_2","unstructured":"I. Kostrikov A. Nair S. Levine \u201cOffline reinforcement learning with implicit Q-learning\u201d in 10th International Conference on Learning Representations (ICLR) (ICLR 2022); https:\/\/openreview.net\/forum?id=68n2s9ZJWF8."},{"key":"e_1_3_2_41_2","unstructured":"J. Ho A. Jain P. Abbeel \u201cDenoising diffusion probabilistic models\u201d in Advances in Neural Information Processing Systems 33 (NeurIPS) H. Larochelle M. Ranzato R. Hadsell M. F. Balcan H. Lin Eds. (Curran Associates 2020) pp. 6840\u20136851."},{"key":"e_1_3_2_42_2","unstructured":"Y. Song J. Sohl-Dickstein D. P. Kingma A. Kumar S. Ermon B. Poole \u201cScore-based generative modeling through stochastic differential equations\u201d in International Conference on Learning Representations (ICLR 2021); https:\/\/openreview.net\/forum?id=PxTIG12RRHS."},{"key":"e_1_3_2_43_2","unstructured":"Y. Lipman R. T. Chen H. Ben-Hamu M. Nickel M. Le \u201cFlow matching for generative modeling\u201d in 11th International Conference on Learning Representations (ICLR 2023); https:\/\/openreview.net\/forum?id=PqvMRDCJT9t."},{"key":"e_1_3_2_44_2","unstructured":"Z. Wang J. J. Hunt M. Zhou \u201cDiffusion policies as an expressive policy class for offline reinforcement learning\u201d in The Eleventh International Conference on Learning Representations (ICLR 2023); https:\/\/openreview.net\/forum?id=AHvFDPi-FA."},{"key":"e_1_3_2_45_2","doi-asserted-by":"crossref","unstructured":"B. Kang X. Ma C. Du T. Pang S. Yan \u201cEfficient diffusion policies for offline reinforcement learning\u201d in Advances in Neural Information Processing Systems 36 (NeurIPS) A. Oh T. Naumann A. Globerson K. Saenko M. Hardt S. Levine Eds. (Curran Associates 2023) pp. 67195\u201367212.","DOI":"10.52202\/075280-2937"},{"key":"e_1_3_2_46_2","unstructured":"C. Lu H. Chen J. Chen H. Su C. Li J. Zhu \u201cContrastive energy prediction for exact energy-guided diffusion sampling in offline reinforcement learning\u201d in Proceedings of the 40th International Conference on Machine Learning (ICML) A. Krause E. Brunskill K. Cho B. Engelhardt S. Sabato J. Scarlett Eds. vol. 202 of Proceedings of Machine Learning Research (PMLR 2023) pp. 22825\u201322855."},{"key":"e_1_3_2_47_2","doi-asserted-by":"crossref","unstructured":"S. Ding K. Hu Z. Zhang K. Ren W. Zhang J. Yu J. Wang Y. Shi \u201cDiffusion-based reinforcement learning via q-weighted variational policy optimization\u201d in Advances in Neural Information Processing Systems 37 (NeurIPS) A. Globerson L. Mackey D. Belgrave A. Fan U. Paquet J. Tomczak C. Zhang Eds. (Curran Associates 2024) pp. 53945\u201353968.","DOI":"10.52202\/079017-1708"},{"key":"e_1_3_2_48_2","unstructured":"M. Psenka A. Escontrela P. Abbeel Y. Ma \u201cLearning a diffusion model policy from rewards via Q-score matching\u201d in Proceedings of the 41st International Conference on Machine Learning (ICML) R. Salakhutdinov Z. Kolter K. Heller A. Weller N. Oliver J. Scarlett F. Berkenkamp Eds. vol. 235 of Proceedings of Machine Learning Research (PMLR 2024) pp. 41163\u201341182."},{"key":"e_1_3_2_49_2","doi-asserted-by":"crossref","unstructured":"H. He C. Bai K. Xu Z. Yang W. Zhang D. Wang B. Zhao X. Li \u201cDiffusion model is an effective planner and data synthesizer for multi-task reinforcement learning\u201d in Advances in Neural Information Processing Systems 36 (NeurIPS) A. Oh T. Naumann A. Globerson K. Saenko M. Hardt S. Levine Eds. (Curran Associates 2023) pp. 64896\u201364917.","DOI":"10.52202\/075280-2832"},{"key":"e_1_3_2_50_2","unstructured":"Z. Ding C. Jin \u201cConsistency models as a rich and efficient policy class for reinforcement learning\u201d in 12th International Conference on Learning Representations (ICLR) B. Kim Y. Yue S. Chaudhuri K. Fragkiadaki M. Khan Y. Sun Eds. (ICLR 2024) pp. 53047\u201353066."},{"key":"e_1_3_2_51_2","unstructured":"H. Chen C. Lu Z. Wang H. Su J. Zhu \u201cScore regularized policy optimization through diffusion behavior\u201d in 12th International Conference on Learning Representations (ICLR) B. Kim Y. Yue S. Chaudhuri K. Fragkiadaki M. Khan Y. Sun Eds. (ICLR 2024) pp. 10211\u201310230."},{"key":"e_1_3_2_52_2","doi-asserted-by":"crossref","unstructured":"H. Li Z. Jiang Y. CHEN D. Zhao \u201cGeneralizing consistency policy to visual RL with prioritized proximal experience regularization\u201d in Advances in Neural Information Processing Systems 37 (NeurIPS) A. Globerson L. Mackey D. Belgrave A. Fan U. Paquet J. Tomczak C. Zhang Eds. (Curran Associates 2024) pp. 109672\u2013109700.","DOI":"10.52202\/079017-3480"},{"key":"e_1_3_2_53_2","unstructured":"K. Black M. Janner Y. Du I. Kostrikov S. Levine \u201cTraining diffusion models with reinforcement learning\u201d in 12th International Conference on Learning Representations (ICLR) B. Kim Y. Yue S. Chaudhuri K. Fragkiadaki M. Khan Y. Sun Eds. (ICLR 2024) pp. 4965\u20134987."},{"key":"e_1_3_2_54_2","doi-asserted-by":"crossref","unstructured":"Y. Fan O. Watkins Y. Du H. Liu M. Ryu C. Boutilier P. Abbeel M. Ghavamzadeh K. Lee K. Lee \u201cDPOK: Reinforcement learning for fine-tuning text-to-image diffusion models\u201d in Advances in Neural Information Processing Systems 36 (NeurIPS) A. Oh T. Naumann A. Globerson K. Saenko M. Hardt S. Levine Eds. (Curran Associates 2023) pp. 79858\u201379885.","DOI":"10.52202\/075280-3497"},{"key":"e_1_3_2_55_2","unstructured":"A. Z. Ren J. Lidard L. L. Ankile A. Simeonov P. Agrawal A. Majumdar B. Burchfiel H. Dai M. Simchowitz \u201cDiffusion policy policy optimization\u201d in 13th International Conference on Learning Representations (ICLR) Y. Yue A. Garg N. Peng F. Sha R. Yu Eds. (ICLR 2025) pp. 77288\u201377329."},{"key":"e_1_3_2_56_2","unstructured":"S. Park Q. Li S. Levine \u201cFlow Q-learning\u201d in Proceedings of the 42nd International Conference on Machine Learning (ICML) A. Singh M. Fazel D. Hsu S. Lacoste-Julien F. Berkenkamp T. Maharaj K. Wagstaff J. Zhu Eds. vol. 267 of Proceedings of Machine Learning Research (PMLR 2025) pp. 48104\u201348127."},{"key":"e_1_3_2_57_2","unstructured":"T. Nguyen C. D. Yoo Revisiting diffusion Q-learning: From iterative denoising to one-step action generation. arXiv:2508.13904 [cs.LG] (2025)."},{"key":"e_1_3_2_58_2","unstructured":"A. Kumar A. Zhou G. Tucker S. Levine \u201cConservative q-learning for offline reinforcement learning\u201d in Advances in Neural Information Processing Systems 33 (NeurIPS) H. Larochelle M. Ranzato R. Hadsell M. F. Balcan H. Lin Eds. (Curran Associates 2020) pp. 1179\u20131191."},{"key":"e_1_3_2_59_2","unstructured":"X. B. Peng A. Kumar G. Zhang S. Levine Advantage-weighted regression: Simple and scalable off-policy reinforcement learning. arXiv:1910.00177 [cs.LG] (2019)."},{"key":"e_1_3_2_60_2","doi-asserted-by":"crossref","unstructured":"M. Nakamoto Y. Zhai A. Singh M. S. Mark Y. Ma C. Finn A. Kumar S. Levine \u201cCal-QL: Calibrated offline RL pre-training for efficient online fine-tuning\u201d in Advances in Neural Information Processing Systems 36 (NeurIPS) A. Oh T. Naumann A. Globerson K. Saenko M. Hardt S. Levine Eds. (Curran Associates 2023) pp. 62244\u201362269.","DOI":"10.52202\/075280-2719"},{"key":"e_1_3_2_61_2","unstructured":"J. Schulman F. Wolski P. Dhariwal A. Radford O. Klimov Proximal policy optimization algorithms. arXiv:1707.06347 [cs.LG] (2017)."},{"key":"e_1_3_2_62_2","unstructured":"P. J. Ball L. Smith I. Kostrikov S. Levine \u201cEfficient online reinforcement learning with offline data\u201d in Proceedings of the 40th International Conference on Machine Learning (ICML) A. Krause E. Brunskill K. Cho B. Engelhardt S. Sabato J. Scarlett Eds. vol. 202 of Proceedings of Machine Learning Research (PMLR 2023) pp. 1577\u20131594."},{"key":"e_1_3_2_63_2","unstructured":"A. Nair A. Gupta M. Dalal S. Levine AWAC: Accelerating online reinforcement learning with offline datasets. arXiv:2006.09359 [cs.LG] (2020)."},{"key":"e_1_3_2_64_2","unstructured":"A. Wagenmaker Y. Zhang M. Nakamoto S. Park W. Yagoub A. Nagabandi A. Gupta S. Levine \u201cSteering your diffusion policy with latent space reinforcement learning\u201d in Proceedings of the 9th Conference on Robot Learning J. Lim S. Song H.-W. Park Eds. vol. 305 of Proceedings of Machine Learning Research (PMLR 2025) pp. 258\u2013282."},{"key":"e_1_3_2_65_2","unstructured":"T. Zhang C. Yu S. Su Y. Wang \u201cReinFlow: Fine-tuning flow matching policy with online reinforcement learning\u201d in Advances in Neural Information Processing Systems 38 (NeurIPS) D. Belgrave C. Zhang H. Lin R. Pascanu P. Koniusz M. Ghassemi N. Chen Eds. (Curran Associates 2025) pp. 1\u201338."},{"key":"e_1_3_2_66_2","doi-asserted-by":"crossref","unstructured":"A. Rajeswaran V. Kumar A. Gupta G. Vezzani J. Schulman E. Todorov S. Levine \u201cLearning complex dexterous manipulation with deep reinforcement learning and demonstrations\u201d in Proceedings of Robotics: Science and Systems H. Kress-Gazit S. Srinivasa T. Howard N. Atanasov Eds. (RSS Foundation 2018) 10.15607\/RSS.2018.XIV.049.","DOI":"10.15607\/RSS.2018.XIV.049"},{"key":"e_1_3_2_67_2","unstructured":"T. Yu D. Quillen Z. He R. Julian K. Hausman C. Finn S. Levine \u201cMeta-world: A benchmark and evaluation for multi-task and meta reinforcement learning\u201d in Proceedings of the 2020 Conference on Robot Learning J. Kober F. Ramos C. Tomlin Eds. vol. 80 of Proceedings of Machine Learning Research (PMLR 2020) pp. 1094\u20131100."},{"key":"e_1_3_2_68_2","unstructured":"J. Fu A. Kumar O. Nachum G. Tucker S. Levine D4RL: Datasets for deep data-driven reinforcement learning. arXiv:2004.07219 [cs.LG] (2020)."},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","unstructured":"E. Todorov T. Erez Y. Tassa \u201cMuJoCo: A physics engine for model-based control\u201d in 2012 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IEEE 2012) pp. 5026\u20135033.","DOI":"10.1109\/IROS.2012.6386109"},{"key":"e_1_3_2_70_2","unstructured":"J. Schulman P. Moritz S. Levine M. I. Jordan P. Abbeel High-dimensional continuous control using generalized advantage estimation. arXiv:1506.02438 [cs.LG] (2015)."}],"container-title":["Science Robotics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.science.org\/doi\/pdf\/10.1126\/scirobotics.aed6267","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,7,22]],"date-time":"2026-07-22T17:58:23Z","timestamp":1784743103000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.science.org\/doi\/10.1126\/scirobotics.aed6267"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,7,22]]},"references-count":69,"journal-issue":{"issue":"116","published-print":{"date-parts":[[2026,7,22]]}},"alternative-id":["10.1126\/scirobotics.aed6267"],"URL":"https:\/\/doi.org\/10.1126\/scirobotics.aed6267","relation":{},"ISSN":["2470-9476"],"issn-type":[{"value":"2470-9476","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,7,22]]},"assertion":[{"value":"2025-11-05","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-06-23","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2026-07-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"eaed6267"}}