{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,26]],"date-time":"2026-03-26T04:09:19Z","timestamp":1774498159169,"version":"3.50.1"},"reference-count":46,"publisher":"MDPI AG","issue":"1","license":[{"start":{"date-parts":[[2023,1,3]],"date-time":"2023-01-03T00:00:00Z","timestamp":1672704000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/501100012166","name":"National Key Research and Development Program of China","doi-asserted-by":"publisher","award":["2018YFC0806802"],"award-info":[{"award-number":["2018YFC0806802"]}],"id":[{"id":"10.13039\/501100012166","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Sensors"],"abstract":"<jats:p>Learning from visual observation for efficient robotic manipulation is a hitherto significant challenge in Reinforcement Learning (RL). Although the collocation of RL policies and convolution neural network (CNN) visual encoder achieves high efficiency and success rate, the method general performance for multi-tasks is still limited to the efficacy of the encoder. Meanwhile, the increasing cost of the encoder optimization for general performance could debilitate the efficiency advantage of the original policy. Building on the attention mechanism, we design a robotic manipulation method that significantly improves the policy general performance among multitasks with the lite Transformer based visual encoder, unsupervised learning, and data augmentation. The encoder of our method could achieve the performance of the original Transformer with much less data, ensuring efficiency in the training process and intensifying the general multi-task performances. Furthermore, we experimentally demonstrate that the master view outperforms the other alternative third-person views in the general robotic manipulation tasks when combining the third-person and egocentric views to assimilate global and local visual information. After extensively experimenting with the tasks from the OpenAI Gym Fetch environment, especially in the Push task, our method succeeds in 92% versus baselines that of 65%, 78% for the CNN encoder, 81% for the ViT encoder, and with fewer training steps.<\/jats:p>","DOI":"10.3390\/s23010515","type":"journal-article","created":{"date-parts":[[2023,1,4]],"date-time":"2023-01-04T02:54:55Z","timestamp":1672800895000},"page":"515","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Vision-Based Efficient Robotic Manipulation with a Dual-Streaming Compact Convolutional Transformer"],"prefix":"10.3390","volume":"23","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-1582-0007","authenticated-orcid":false,"given":"Hao","family":"Guo","sequence":"first","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Meichao","family":"Song","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Zhen","family":"Ding","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4180-1109","authenticated-orcid":false,"given":"Chunzhi","family":"Yi","sequence":"additional","affiliation":[{"name":"School of Medicine and Health, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Feng","family":"Jiang","sequence":"additional","affiliation":[{"name":"School of Computer Science and Technology, Harbin Institute of Technology, Harbin 150001, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2023,1,3]]},"reference":[{"key":"ref_1","unstructured":"Badia, A.P., Piot, B., Kapturowski, S., Sprechmann, P., Vitvitskyi, A., Guo, Z.D., and Blundell, C. (2020, January 13\u201318). Agent57: Outperforming the atari human benchmark. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_2","unstructured":"Berner, C., Brockman, G., Chan, B., Cheung, V., D\u0119biak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., and Hesse, C. (2019). Dota 2 with large scale deep reinforcement learning. arXiv."},{"key":"ref_3","doi-asserted-by":"crossref","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"ref_4","doi-asserted-by":"crossref","first-page":"350","DOI":"10.1038\/s41586-019-1724-z","article-title":"Grandmaster level in StarCraft II using multi-agent reinforcement learning","volume":"575","author":"Vinyals","year":"2019","journal-title":"Nature"},{"key":"ref_5","unstructured":"Yang, Y., Caluwaerts, K., Iscen, A., Zhang, T., Tan, J., and Sindhwani, V. (2020, January 16\u201318). Data efficient reinforcement learning for legged robots. Proceedings of the Conference on Robot Learning, Virtual."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Haarnoja, T., Ha, S., Zhou, A., Tan, J., Tucker, G., and Levine, S. (2018). Learning to walk via deep reinforcement learning. arXiv.","DOI":"10.15607\/RSS.2019.XV.011"},{"key":"ref_7","doi-asserted-by":"crossref","first-page":"3046","DOI":"10.1109\/LRA.2022.3144512","article-title":"Look Closer: Bridging Egocentric and Third-Person Views with Transformers for Robotic Manipulation","volume":"7","author":"Jangir","year":"2022","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_8","first-page":"3680","article-title":"Stabilizing deep q-learning with convnets and vision transformers under data augmentation","volume":"34","author":"Hansen","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_9","unstructured":"Chen, T., Xu, J., and Agrawal, P. (2022, January 8\u201311). A system for general in-hand object re-orientation. Proceedings of the Conference on Robot Learning, Auckland, New Zealand."},{"key":"ref_10","doi-asserted-by":"crossref","unstructured":"Johannink, T., Bahl, S., Nair, A., Luo, J., Kumar, A., Loskyll, M., Ojea, J.A., Solowjow, E., and Levine, S. (2019, January 20\u201324). Residual Reinforcement Learning for Robot Control. Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada.","DOI":"10.1109\/ICRA.2019.8794127"},{"key":"ref_11","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1177\/0278364919887447","article-title":"Learning dexterous in-hand manipulation","volume":"39","author":"Andrychowicz","year":"2020","journal-title":"Int. J. Robot. Res."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"421","DOI":"10.1177\/0278364917710318","article-title":"Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection","volume":"37","author":"Levine","year":"2018","journal-title":"Int. J. Robot. Res."},{"key":"ref_13","unstructured":"OpenAI, O., Plappert, M., Sampedro, R., Xu, T., Akkaya, I., Kosaraju, V., Welinder, P., D\u2019Sa, R., Petron, A., and Pinto, H.P.d.O. (2021). Asymmetric self-play for automatic goal discovery in robotic manipulation. arXiv."},{"key":"ref_14","doi-asserted-by":"crossref","unstructured":"Gu, S., Holly, E., Lillicrap, T., and Levine, S. (June, January 29). Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore.","DOI":"10.1109\/ICRA.2017.7989385"},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Peng, X.B., Andrychowicz, M., Zaremba, W., and Abbeel, P. (2018, January 21\u201325). Sim-to-real transfer of robotic control with dynamics randomization. Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, QLD, Australia.","DOI":"10.1109\/ICRA.2018.8460528"},{"key":"ref_16","doi-asserted-by":"crossref","first-page":"2950","DOI":"10.1109\/LRA.2020.2974685","article-title":"Learning fast adaptation with meta strategy optimization","volume":"5","author":"Yu","year":"2020","journal-title":"IEEE Robot. Autom. Lett."},{"key":"ref_17","unstructured":"James, S., Bloesch, M., and Davison, A.J. (2018, January 29\u201331). Task-embedded control networks for few-shot imitation learning. Proceedings of the Conference on Robot Learning, Z\u00fcrich, Switzerland."},{"key":"ref_18","unstructured":"Finn, C., Yu, T., Zhang, T., Abbeel, P., and Levine, S. (2017, January 13\u201315). One-shot visual imitation learning via meta-learning. Proceedings of the Conference on Robot Learning, Mountain View, CA, USA."},{"key":"ref_19","unstructured":"Duan, Y., Andrychowicz, M., Stadie, B., Jonathan Ho, O., Schneider, J., Sutskever, I., Abbeel, P., and Zaremba, W. (2017). One-shot imitation learning. arXiv."},{"key":"ref_20","unstructured":"Wu, Y.H., Charoenphakdee, N., Bao, H., Tangkaratt, V., and Sugiyama, M. (2019, January 9\u201315). Imitation learning from imperfect demonstration. Proceedings of the International Conference on Machine Learning, Long Beach, CA, USA."},{"key":"ref_21","unstructured":"Laskin, M., Srinivas, A., and Abbeel, P. (2020, January 13\u201318). Curl: Contrastive unsupervised representations for reinforcement learning. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_22","first-page":"19884","article-title":"Reinforcement learning with augmented data","volume":"33","author":"Laskin","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_23","unstructured":"Yarats, D., Zhang, A., Kostrikov, I., Amos, B., Pineau, J., and Fergus, R. (2019). Improving sample efficiency in model-free reinforcement learning from images. arXiv."},{"key":"ref_24","unstructured":"Zhan, A., Zhao, P., Pinto, L., Abbeel, P., and Laskin, M. (2020). A framework for efficient robotic manipulation. arXiv."},{"key":"ref_25","unstructured":"Hassani, A., Walton, S., Shah, N., Abuduweili, A., Li, J., and Shi, H. (2021). Escaping the big data paradigm with compact transformers. arXiv."},{"key":"ref_26","unstructured":"Wu, Z., Liu, Z., Lin, J., Lin, Y., and Han, S. (2020). Lite transformer with long-short range attention. arXiv."},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Wang, H., Wu, Z., Liu, Z., Cai, H., Zhu, L., Gan, C., and Han, S. (2020). Hat: Hardware-aware transformers for efficient natural language processing. arXiv.","DOI":"10.18653\/v1\/2020.acl-main.686"},{"key":"ref_28","unstructured":"Wu, B., Xu, C., Dai, X., Wan, A., Zhang, P., Yan, Z., Tomizuka, M., Gonzalez, J., Keutzer, K., and Vajda, P. (2020). Visual transformers: Token-based image representation and processing for computer vision. arXiv."},{"key":"ref_29","doi-asserted-by":"crossref","unstructured":"Graham, B., El-Nouby, A., Touvron, H., Stock, P., Joulin, A., J\u00e9gou, H., and Douze, M. (2021, January 10\u201317). LeViT: A Vision Transformer in ConvNet\u2019s Clothing for Faster Inference. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Montreal, QC, Canada.","DOI":"10.1109\/ICCV48922.2021.01204"},{"key":"ref_30","unstructured":"Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and J\u00e9gou, H. (2021, January 18\u201324). Training data-efficient image transformers & distillation through attention. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_31","unstructured":"Schwarzer, M., Anand, A., Goel, R., Hjelm, R.D., Courville, A., and Bachman, P. (2020). Data-efficient reinforcement learning with self-predictive representations. arXiv."},{"key":"ref_32","first-page":"4772","article-title":"Adaptive auxiliary task weighting for reinforcement learning","volume":"32","author":"Lin","year":"2019","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_33","unstructured":"Jaderberg, M., Mnih, V., Czarnecki, W.M., Schaul, T., Leibo, J.Z., Silver, D., and Kavukcuoglu, K. (2016). Reinforcement learning with unsupervised auxiliary tasks. arXiv."},{"key":"ref_34","unstructured":"Li, J., Zhan, X., Xiao, Z., and Zhou, G. (2021). Efficient Robotic Manipulation Through Offline-to-Online Reinforcement Learning and Goal-Aware State Information. arXiv."},{"key":"ref_35","first-page":"15084","article-title":"Decision transformer: Reinforcement learning via sequence modeling","volume":"34","author":"Chen","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_36","unstructured":"Chen, C., Wu, Y.F., Yoon, J., and Ahn, S. (2022). Transdreamer: Reinforcement learning with transformer world models. arXiv."},{"key":"ref_37","unstructured":"Tao, T., Reda, D., and van de Panne, M. (2022). Evaluating Vision Transformer Methods for Deep Reinforcement Learning from Pixels. arXiv."},{"key":"ref_38","unstructured":"Meng, L., Goodwin, M., Yazidi, A., and Engelstad, P. (2022). Deep Reinforcement Learning with Swin Transformer. arXiv."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"4687","DOI":"10.1109\/JSEN.2022.3146307","article-title":"Learning Automated Driving in Complex Intersection Scenarios Based on Camera Sensors: A Deep Reinforcement Learning Approach","volume":"22","author":"Li","year":"2022","journal-title":"IEEE Sens. J."},{"key":"ref_40","doi-asserted-by":"crossref","unstructured":"He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. (2020, January 13\u201319). Momentum contrast for unsupervised visual representation learning. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"ref_41","unstructured":"Henaff, O. (2020, January 13\u201318). Data-efficient image recognition with contrastive predictive coding. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_42","unstructured":"Chen, T., Kornblith, S., Norouzi, M., and Hinton, G. (2020, January 13\u201318). A simple framework for contrastive learning of visual representations. Proceedings of the International Conference on Machine Learning, Virtual."},{"key":"ref_43","first-page":"6256","article-title":"Unsupervised data augmentation for consistency training","volume":"33","author":"Xie","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_44","unstructured":"Berthelot, D., Carlini, N., Goodfellow, I., Papernot, N., Oliver, A., and Raffel, C.A. (2019). Mixmatch: A holistic approach to semi-supervised learning. arXiv."},{"key":"ref_45","first-page":"596","article-title":"FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence","volume":"33","author":"Sohn","year":"2020","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_46","unstructured":"Van den Oord, A., Li, Y., and Vinyals, O. (2018). Representation learning with contrastive predictive coding. arXiv."}],"container-title":["Sensors"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/1\/515\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T17:57:02Z","timestamp":1760119022000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/1424-8220\/23\/1\/515"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,1,3]]},"references-count":46,"journal-issue":{"issue":"1","published-online":{"date-parts":[[2023,1]]}},"alternative-id":["s23010515"],"URL":"https:\/\/doi.org\/10.3390\/s23010515","relation":{},"ISSN":["1424-8220"],"issn-type":[{"value":"1424-8220","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,1,3]]}}}