{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,11]],"date-time":"2025-10-11T02:46:04Z","timestamp":1760150764998,"version":"build-2065373602"},"reference-count":40,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2023,12,27]],"date-time":"2023-12-27T00:00:00Z","timestamp":1703635200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Robotics"],"abstract":"<jats:p>Robot manipulation in a physically constrained environment requires compliant manipulation. Compliant manipulation is a manipulation skill to adjust hand motion based on the force imposed by the environment. Recently, reinforcement learning (RL) has been applied to solve household operations involving compliant manipulation. However, previous RL methods have primarily focused on designing a policy for a specific operation that limits their applicability and requires separate training for every new operation. We propose a constraint-aware policy that is applicable to various unseen manipulations by grouping several manipulations together based on the type of physical constraint involved. The type of physical constraint determines the characteristic of the imposed force direction; thus, a generalized policy is trained in the environment and reward designed on the basis of this characteristic. This paper focuses on two types of physical constraints: prismatic and revolute joints. Experiments demonstrated that the same policy could successfully execute various compliant manipulation operations, both in the simulation and reality. We believe this study is the first step toward realizing a generalized household robot.<\/jats:p>","DOI":"10.3390\/robotics13010008","type":"journal-article","created":{"date-parts":[[2023,12,27]],"date-time":"2023-12-27T07:45:32Z","timestamp":1703663132000},"page":"8","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Constraint-Aware Policy for Compliant Manipulation"],"prefix":"10.3390","volume":"13","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-3843-7542","authenticated-orcid":false,"given":"Daichi","family":"Saito","sequence":"first","affiliation":[{"name":"School of Computing, Tokyo Institute of Technology, Tokyo 152-8550, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Kazuhiro","family":"Sasabuchi","sequence":"additional","affiliation":[{"name":"Applied Robotics Research, Microsoft, Redmond, WA 98052, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8278-2373","authenticated-orcid":false,"given":"Naoki","family":"Wake","sequence":"additional","affiliation":[{"name":"Applied Robotics Research, Microsoft, Redmond, WA 98052, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Atsushi","family":"Kanehira","sequence":"additional","affiliation":[{"name":"Applied Robotics Research, Microsoft, Redmond, WA 98052, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Jun","family":"Takamatsu","sequence":"additional","affiliation":[{"name":"Applied Robotics Research, Microsoft, Redmond, WA 98052, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hideki","family":"Koike","sequence":"additional","affiliation":[{"name":"School of Computing, Tokyo Institute of Technology, Tokyo 152-8550, Japan"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Katsushi","family":"Ikeuchi","sequence":"additional","affiliation":[{"name":"Applied Robotics Research, Microsoft, Redmond, WA 98052, USA"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,12,27]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"418","DOI":"10.1109\/TSMC.1981.4308708","article-title":"Compliance and force control for computer controlled manipulators","volume":"11","author":"Mason","year":"1981","journal-title":"IEEE Trans. Syst. Man Cybern."},{"key":"ref_2","doi-asserted-by":"crossref","unstructured":"Yahya, A., Li, A., Kalakrishnan, M., Chebotar, Y., and Levine, S. (2017, January 24\u201328). Collective robot reinforcement learning with distributed asynchronous guided policy search. Proceedings of the 2017 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS), Vancouver, BC, Canada.","DOI":"10.1109\/IROS.2017.8202141"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Gu, S., Holly, E., Lillicrap, T., and Levine, S. (June, January 29). Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore.","DOI":"10.1109\/ICRA.2017.7989385"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Rajeswaran, A., Kumar, V., Gupta, A., Vezzani, G., Schulman, J., Todorov, E., and Levine, S. (2018, January 26\u201330). Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations. Proceedings of the Robotics: Science and Systems (RSS), Pittsburgh, PA, USA.","DOI":"10.15607\/RSS.2018.XIV.049"},{"key":"ref_5","unstructured":"Urakami, Y., Hodgkinson, A., Carlin, C., Leu, R., Rigazio, L., and Abbeel, P. (2019). Doorgym: A scalable door opening environment and baseline agent. arXiv."},{"key":"ref_6","doi-asserted-by":"crossref","first-page":"137188","DOI":"10.1109\/ACCESS.2021.3118594","article-title":"Force-Vision Sensor Fusion Improves Learning-Based Approach for Self-Closing Door Pulling","volume":"9","author":"Sun","year":"2021","journal-title":"IEEE Access"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"161","DOI":"10.1109\/TRO.2015.2506154","article-title":"An adaptive control approach for opening doors and drawers under uncertainties","volume":"32","author":"Karayiannidis","year":"2016","journal-title":"IEEE Trans. Robot."},{"key":"ref_8","unstructured":"Ikeuchi, K., Wake, N., Arakawa, R., Sasabuchi, K., and Takamatsu, J. (2021). Semantic constraints to represent common sense required in household actions for multi-modal learning-from-observation robot. arXiv."},{"key":"ref_9","doi-asserted-by":"crossref","first-page":"368","DOI":"10.1109\/70.294211","article-title":"Toward an assembly plan from observation. I. Task recognition with polyhedral objects","volume":"10","author":"Ikeuchi","year":"1994","journal-title":"IEEE Trans. Robot. Autom."},{"key":"ref_10","unstructured":"Wake, N., Kanehira, A., Sasabuchi, K., Takamatsu, J., and Ikeuchi, K. (2022). Interactive Learning-from-Observation through multimodal human demonstration. arXiv."},{"key":"ref_11","unstructured":"Nagatani, K., and Yuta, S. (1995, January 5\u20139). An experiment on opening-door-behavior by an autonomous mobile robot with a manipulator. Proceedings of the 1995 IEEE\/RSJ International Conference on Intelligent Robots and Systems. Human Robot Interaction and Cooperative Robots, Pittsburgh, PA, USA."},{"key":"ref_12","doi-asserted-by":"crossref","unstructured":"Klingbeil, E., Saxena, A., and Ng, A.Y. (2010, January 18\u201322). Learning to open new doors. Proceedings of the 2010 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Taipei, Taiwan.","DOI":"10.1109\/IROS.2010.5649847"},{"key":"ref_13","first-page":"1289","article-title":"Learning to Generalize Kinematic Models to Novel Objects","volume":"Volume 100","author":"Kaelbling","year":"2020","journal-title":"Proceedings of the the Conference on Robot Learning, PMLR"},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"R\u00fchr, T., Sturm, J., Pangercic, D., Beetz, M., and Cremers, D. (2012, January 14\u201318). A generalized framework for opening doors and drawers in kitchen environments. Proceedings of the 2012 IEEE International Conference on Robotics and Automation, Saint Paul, MN, USA.","DOI":"10.1109\/ICRA.2012.6224929"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Li, X., Wang, H., Yi, L., Guibas, L.J., Abbott, A.L., and Song, S. (2020, January 13\u201319). Category-Level Articulated Object Pose Estimation. Proceedings of the 2020 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00376"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"409","DOI":"10.1007\/s11370-021-00366-7","article-title":"Robust and adaptive door operation with a mobile robot","volume":"14","author":"Arduengo","year":"2021","journal-title":"Intell. Serv. Robot."},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"1072","DOI":"10.1109\/TIP.2021.3138644","article-title":"Toward Real-World Category-Level Articulation Pose Estimation","volume":"31","author":"Liu","year":"2022","journal-title":"IEEE Trans. Image Process."},{"key":"ref_18","doi-asserted-by":"crossref","unstructured":"Jain, A., Lioutikov, R., Chuck, C., and Niekum, S. (June, January 30). ScrewNet: Category-Independent Articulation Model Estimation From Depth Images Using Screw Theory. Proceedings of the 2021 IEEE International Conference on Robotics and Automation (ICRA), Xi\u2019an, China.","DOI":"10.1109\/ICRA48506.2021.9561132"},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Eisner, B., Zhang, H., and Held, D. (2022). Flowbot3d: Learning 3d articulation flow to manipulate articulated objects. arXiv.","DOI":"10.15607\/RSS.2022.XVIII.018"},{"key":"ref_20","doi-asserted-by":"crossref","unstructured":"Wei, F., Chabra, R., Ma, L., Lassner, C., Zollh\u00f6fer, M., Rusinkiewicz, S., Sweeney, C., Newcombe, R., and Slavcheva, M. (2022, January 18\u201324). Self-supervised neural articulated shape and appearance models. Proceedings of the the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01536"},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Kessens, C.C., Rice, J.B., Smith, D.C., Biggs, S.J., and Garcia, R. (2010, January 18\u201322). Utilizing compliance to manipulate doors with unmodeled constraints. Proceedings of the 2010 IEEE\/RSJ International Conference on Intelligent Robots and Systems, Taipei, Taiwan.","DOI":"10.1109\/IROS.2010.5650927"},{"key":"ref_22","doi-asserted-by":"crossref","unstructured":"Jain, A., and Kemp, C.C. (2010, January 3\u20137). Pulling open doors and drawers: Coordinating an omni-directional base and a compliant arm with equilibrium point control. Proceedings of the 2010 IEEE International Conference on Robotics and Automation, Anchorage, AK, USA.","DOI":"10.1109\/ROBOT.2010.5509445"},{"key":"ref_23","unstructured":"Niemeyer, G., and Slotine, J.J. (1997, January 25). A simple strategy for opening an unknown door. Proceedings of the International Conference on Robotics and Automation, Albuquerque, NM, USA."},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Schmid, A.J., Gorges, N., Goger, D., and Worn, H. (2008, January 19\u201323). Opening a door with a humanoid robot using multi-sensory tactile feedback. Proceedings of the 2008 IEEE International Conference on Robotics and Automation, Pasadena, CA, USA.","DOI":"10.1109\/ROBOT.2008.4543222"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"3975","DOI":"10.1109\/TIE.2009.2025296","article-title":"Door-Opening Control of a Service Robot Using the Multifingered Robot Hand","volume":"56","author":"Chung","year":"2009","journal-title":"IEEE Trans. Ind. Electron."},{"key":"ref_26","doi-asserted-by":"crossref","first-page":"753","DOI":"10.3182\/20120905-3-HR-2030.00047","article-title":"Adaptive Force\/Velocity control for opening unknown doors1","volume":"45","author":"Karayiannidis","year":"2012","journal-title":"IFAC Proc. Vol."},{"key":"ref_27","first-page":"123","article-title":"A nonlinear adaptive compliance controller for rehabilitation","volume":"5","author":"Pilastro","year":"2016","journal-title":"IEEJ J. Ind. Appl."},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"Corr\u00e1, L., Oboe, R., and Shimono, T. (November, January 29). Adaptive optimal control for rehabilitation systems. Proceedings of the IECON 2017-43rd Annual Conference of the IEEE Industrial Electronics Society, Beijing, China.","DOI":"10.1109\/IECON.2017.8216899"},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Pareek, S., NIsar, H., and Kesavadas, T. (2023). AR3n: A Reinforcement Learning-based Assist-As-Needed Controller for Robotic Rehabilitation. arXiv.","DOI":"10.1109\/MRA.2023.3282434"},{"key":"ref_30","first-page":"9209","article-title":"Visual reinforcement learning with imagined goals","volume":"31","author":"Nair","year":"2018","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_31","unstructured":"Yu, T., Quillen, D., He, Z., Julian, R., Hausman, K., Finn, C., and Levine, S. (November, January 30). Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning. Proceedings of the Conference on Robot Learning, PMLR, Virtual Event."},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Dabis, J., Finn, C., Gopalakrishnan, K., Hausman, K., Herzog, A., and Hsu, J. (2022). Rt-1: Robotics transformer for real-world control at scale. arXiv.","DOI":"10.15607\/RSS.2023.XIX.025"},{"key":"ref_33","unstructured":"Reed, S., Zolna, K., Parisotto, E., Colmenarejo, S.G., Novikov, A., Barth-Maron, G., Gimenez, M., Sulsky, Y., Kay, J., and Springenberg, J.T. (2022). A generalist agent. arXiv."},{"key":"ref_34","unstructured":"Brohan, A., Brown, N., Carbajal, J., Chebotar, Y., Chen, X., Choromanski, K., Ding, T., Driess, D., Dubey, A., and Finn, C. (2023). Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv."},{"key":"ref_35","unstructured":"Shridhar, M., Manuelli, L., and Fox, D. (2023, January 6\u20139). Perceiver-actor: A multi-task transformer for robotic manipulation. Proceedings of the Conference on Robot Learning, PMLR, Atlanta, GA, USA."},{"key":"ref_36","unstructured":"Coumans, E., and Bai, Y. (2023, December 26). PyBullet, a Python Module for Physics Simulation for Games, Robotics and Machine Learning. 2016\u20132021. Available online: http:\/\/pybullet.org."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Saito, D., Sasabuchi, K., Wake, N., Takamatsu, J., Koike, H., and Ikeuchi, K. (2022, January 28\u201330). Task-grasping from a demonstrated human strategy. Proceedings of the 2022 IEEE-RAS 21st International Conference on Humanoid Robots (Humanoids), Ginowan, Japan.","DOI":"10.1109\/Humanoids53995.2022.10000167"},{"key":"ref_38","unstructured":"Wake, N., Sasabuchi, K., and Ikeuchi, K. (2020). Grasp-type recognition leveraging object affordance. arXiv."},{"key":"ref_39","unstructured":"Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"Quigley, M., Conley, K., Gerkey, B., Faust, J., Foote, T., Leibs, J., Wheeler, R., and Ng, A.Y. (2009, January 12\u201317). ROS: An open-source Robot Operating System. Proceedings of the ICRA Workshop on Open Source Software, Kobe, Japan.","DOI":"10.1109\/MRA.2010.936956"}],"container-title":["Robotics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2218-6581\/13\/1\/8\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T21:42:42Z","timestamp":1760132562000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2218-6581\/13\/1\/8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,12,27]]},"references-count":40,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2024,1]]}},"alternative-id":["robotics13010008"],"URL":"https:\/\/doi.org\/10.3390\/robotics13010008","relation":{},"ISSN":["2218-6581"],"issn-type":[{"type":"electronic","value":"2218-6581"}],"subject":[],"published":{"date-parts":[[2023,12,27]]}}}