{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,28]],"date-time":"2026-04-28T20:19:16Z","timestamp":1777407556993,"version":"3.51.4"},"reference-count":22,"publisher":"Cambridge University Press (CUP)","issue":"3","license":[{"start":{"date-parts":[[2022,9,7]],"date-time":"2022-09-07T00:00:00Z","timestamp":1662508800000},"content-version":"unspecified","delay-in-days":0,"URL":"https:\/\/www.cambridge.org\/core\/terms"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Robotica"],"published-print":{"date-parts":[[2023,3]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>In the field of robot reinforcement learning (RL), the reality gap has always been a problem that restricts the robustness and generalization of algorithms. We propose Simulation Twin (SimTwin) : a deep RL framework that can help directly transfer the model from simulation to reality without any real-world training. SimTwin consists of a RL module and an adaptive correct module. We train the policy using the soft actor-critic algorithm only in a simulator with demonstration and domain randomization. In the adaptive correct module, we design and train a neural network to simulate the human error correction process using force feedback. Subsequently, we combine the above two modules through digital twin to control real-world robots, correct simulator parameters by comparing the difference between simulator and reality automatically, and then generalize the correct action through the trained policy network without additional training. We demonstrate the proposed method in an open cabinet task; the experiments show that our framework can reduce the reality gap without any real-world training.<\/jats:p>","DOI":"10.1017\/s0263574722001230","type":"journal-article","created":{"date-parts":[[2022,9,7]],"date-time":"2022-09-07T08:51:08Z","timestamp":1662540668000},"page":"1015-1024","source":"Crossref","is-referenced-by-count":10,"title":["Zero-shot sim-to-real transfer of reinforcement learning framework for robotics manipulation with demonstration and force feedback"],"prefix":"10.1017","volume":"41","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-0033-492X","authenticated-orcid":false,"given":"Yuanpei","family":"Chen","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chao","family":"Zeng","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhiping","family":"Wang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Peng","family":"Lu","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Chenguang","family":"Yang","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"56","published-online":{"date-parts":[[2022,9,7]]},"reference":[{"key":"S0263574722001230_ref19","doi-asserted-by":"publisher","DOI":"10.1016\/j.jmsy.2020.06.012"},{"key":"S0263574722001230_ref21","doi-asserted-by":"publisher","DOI":"10.1177\/0278364919887447"},{"key":"S0263574722001230_ref13","unstructured":"[13] Guha, A. and Annaswamy, A. , \u201cMrac-rl: a framework for on-line policy adaptation under parametric model uncertainty,\u201d arXiv preprint arXiv:2011.10562 (2020)."},{"key":"S0263574722001230_ref3","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3324926","article-title":"A survey of zero-shot learning: Settings, methods, and applications","volume":"10","author":"Wang","year":"2019","journal-title":"ACM Trans. Intell. Syst. Technol. (TIST)"},{"key":"S0263574722001230_ref10","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11757"},{"key":"S0263574722001230_ref7","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2019.8793789"},{"key":"S0263574722001230_ref12","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2021.3084880"},{"key":"S0263574722001230_ref20","doi-asserted-by":"publisher","DOI":"10.1109\/IROS51168.2021.9636292"},{"key":"S0263574722001230_ref15","article-title":"A unified parametric representation for robotic compliant skills","volume":"27","author":"Zeng","year":"2021","journal-title":"IEEE\/ASME Trans. Mechatron"},{"key":"S0263574722001230_ref6","doi-asserted-by":"publisher","DOI":"10.1109\/IRC.2019.00120"},{"key":"S0263574722001230_ref4","doi-asserted-by":"publisher","DOI":"10.1109\/IROS.2017.8202133"},{"key":"S0263574722001230_ref9","unstructured":"[9] Haarnoja, T. , Zhou, A. , Hartikainen, K. , Tucker, G. , Ha, S. , Tan, J. , Kumar, V. , Zhu, H. , Gupta, A. , P. Abbeel and S. Levine, \u201cSoft actor-critic algorithms and applications,\u201d arXiv preprint arXiv: 1812.05905 (2018)."},{"key":"S0263574722001230_ref17","doi-asserted-by":"publisher","DOI":"10.1109\/LRA.2020.3010739"},{"key":"S0263574722001230_ref16","doi-asserted-by":"publisher","DOI":"10.1109\/IROS40897.2019.8968201"},{"key":"S0263574722001230_ref22","unstructured":"[22] Makoviychuk, V. , Wawrzyniak, L. , Guo, Y. , Lu, M. , Storey, K. , Macklin, M. , Hoeller, D. , Rudin, N. , Allshire, A. , A. Handa and G. State, \u201cIsaac gym: High performance GPU-based physics simulation for robot learning,\u201d arXiv preprint arXiv: 2108.10470 (2021)."},{"key":"S0263574722001230_ref5","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8460528"},{"key":"S0263574722001230_ref8","unstructured":"[8] Lillicrap, T. P. , Hunt, J. J. , Pritzel, A. , Heess, N. , Erez, T. , Tassa, Y. , Silver, D. and Wierstra, D. , \u201cContinuous control with deep reinforcement learning,\u201d arXiv preprint arXiv: 1509.02971 (2015)."},{"key":"S0263574722001230_ref1","doi-asserted-by":"publisher","DOI":"10.1109\/SSCI47803.2020.9308468"},{"key":"S0263574722001230_ref2","unstructured":"[2] Gupta, A. , Devin, C. , Liu, Y. , Abbeel, P. and Levine, S. , \u201cLearning invariant feature spaces to transfer skills with reinforcement learning,\u201d arXiv preprint arXiv:1703.02949 (2017)."},{"key":"S0263574722001230_ref18","doi-asserted-by":"publisher","DOI":"10.1109\/CAC51589.2020.9327756"},{"key":"S0263574722001230_ref11","unstructured":"[11] Christiano, P. , Shah, Z. , Mordatch, I. , Schneider, J. , Blackwell, T. , Tobin, J. , Abbeel, P. and Zaremba, W. , \u201cTransfer from simulation to real world through learning deep inverse dynamics model,\u201d arXiv preprint arXiv: 1610.03518 (2016)."},{"key":"S0263574722001230_ref14","doi-asserted-by":"publisher","DOI":"10.1109\/TII.2021.3087337"}],"container-title":["Robotica"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.cambridge.org\/core\/services\/aop-cambridge-core\/content\/view\/S0263574722001230","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,2,8]],"date-time":"2023-02-08T05:28:31Z","timestamp":1675834111000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.cambridge.org\/core\/product\/identifier\/S0263574722001230\/type\/journal_article"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,9,7]]},"references-count":22,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2023,3]]}},"alternative-id":["S0263574722001230"],"URL":"https:\/\/doi.org\/10.1017\/s0263574722001230","relation":{},"ISSN":["0263-5747","1469-8668"],"issn-type":[{"value":"0263-5747","type":"print"},{"value":"1469-8668","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,9,7]]}}}