{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,16]],"date-time":"2026-06-16T05:13:51Z","timestamp":1781586831321,"version":"3.54.5"},"publisher-location":"New York, NY, USA","reference-count":49,"publisher":"ACM","funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["624B2025,62473050,92370203"],"award-info":[{"award-number":["624B2025,62473050,92370203"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":[],"published-print":{"date-parts":[[2025,12,15]]},"DOI":"10.1145\/3757377.3763966","type":"proceedings-article","created":{"date-parts":[[2025,12,8]],"date-time":"2025-12-08T16:30:41Z","timestamp":1765211441000},"page":"1-11","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":2,"title":["Unifying Latent Action and Latent State Pre-training for Policy Learning from Videos"],"prefix":"10.1145","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-2903-1957","authenticated-orcid":false,"given":"Guangyan","family":"Chen","sequence":"first","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3618-7423","authenticated-orcid":false,"given":"Meiling","family":"Wang","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0002-3265-3148","authenticated-orcid":false,"given":"Te","family":"Cui","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0006-5429-2207","authenticated-orcid":false,"given":"Luojie","family":"Yang","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0009-5971-4773","authenticated-orcid":false,"given":"Qi","family":"Shao","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8381-3397","authenticated-orcid":false,"given":"Lin","family":"Zhao","sequence":"additional","affiliation":[{"name":"JD, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0779-5905","authenticated-orcid":false,"given":"Tianle","family":"Zhang","sequence":"additional","affiliation":[{"name":"JD, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-8548-2499","authenticated-orcid":false,"given":"Yihang","family":"Li","sequence":"additional","affiliation":[{"name":"JD, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3964-2433","authenticated-orcid":false,"given":"Yi","family":"Yang","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6628-7946","authenticated-orcid":false,"given":"Yufeng","family":"Yue","sequence":"additional","affiliation":[{"name":"Beijing Institute of Technology, Beijing, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,12,14]]},"reference":[{"key":"e_1_3_3_2_2_1","unstructured":"Jean-Baptiste Alayrac Jeff Donahue Pauline Luc Antoine Miech Iain Barr Yana Hasson Karel Lenc Arthur Mensch Katherine Millican Malcolm Reynolds et\u00a0al. 2022. Flamingo: a visual language model for few-shot learning. NeurIPS 35 (2022) 23716\u201323736."},{"key":"e_1_3_3_2_3_1","doi-asserted-by":"crossref","unstructured":"Shikhar Bahl Abhinav Gupta and Deepak Pathak. 2022. Human-to-robot imitation in the wild. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2207.09450 (2022).","DOI":"10.15607\/RSS.2022.XVIII.026"},{"key":"e_1_3_3_2_4_1","unstructured":"Bowen Baker Ilge Akkaya Peter Zhokov Joost Huizinga Jie Tang Adrien Ecoffet Brandon Houghton Raul Sampedro and Jeff Clune. 2022. Video pretraining (vpt): Learning to act by watching unlabeled online videos. NeurIPS 35 (2022) 24639\u201324654."},{"key":"e_1_3_3_2_5_1","volume-title":"ICML","author":"Bruce Jake","year":"2024","unstructured":"Jake Bruce, Michael\u00a0D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, et\u00a0al. 2024. Genie: Generative interactive environments. In ICML."},{"key":"e_1_3_3_2_6_1","unstructured":"Qingwen Bu Jia Zeng Li Chen Yanchao Yang Guyue Zhou Junchi Yan Ping Luo Heming Cui Yi Ma and Hongyang Li. 2024. Closed-loop visuomotor control with generative expectation for robotic manipulation. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2409.09016 (2024)."},{"key":"e_1_3_3_2_7_1","doi-asserted-by":"crossref","unstructured":"Annie\u00a0S Chen Suraj Nair and Chelsea Finn. 2021a. Learning generalizable robotic reward functions from\" in-the-wild\" human videos. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2103.16817 (2021).","DOI":"10.15607\/RSS.2021.XVII.012"},{"key":"e_1_3_3_2_8_1","doi-asserted-by":"publisher","DOI":"10.15607\/RSS.2021.XVII.012"},{"key":"e_1_3_3_2_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52734.2025.00171"},{"key":"e_1_3_3_2_10_1","doi-asserted-by":"crossref","unstructured":"Guangyan Chen Meiling Wang Te Cui Yao Mu Haoyang Lu Zicai Peng Mengxiao Hu Tianxing Zhou Mengyin Fu Yi Yang et\u00a0al. 2025b. FMimic: Foundation Models are Fine-grained Action Learners from Human Videos. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2507.20622 (2025).","DOI":"10.1177\/02783649251377335"},{"key":"e_1_3_3_2_11_1","doi-asserted-by":"crossref","unstructured":"Guangyan Chen Meiling Wang Yao Mu\u00a0Te Cui Haoyang Lu Tianxing Zhou Zicai Peng Mengxiao Hu Haizhou Li Yuan Li Yi Yang et\u00a0al. 2024d. VLMimic: Vision Language Models are Visual Imitation Learner for Fine-grained Actions. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.20927 (2024).","DOI":"10.52202\/079017-2475"},{"key":"e_1_3_3_2_12_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.01709"},{"key":"e_1_3_3_2_13_1","unstructured":"Xiaoyu Chen Junliang Guo Tianyu He Chuheng Zhang Pushi Zhang Derek\u00a0Cathera Yang Li Zhao and Jiang Bian. 2024c. Igor: Image-goal representations are the atomic control units for foundation models in embodied ai. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2411.00785 (2024)."},{"key":"e_1_3_3_2_14_1","unstructured":"Yi Chen Yuying Ge Yizhuo Li Yixiao Ge Mingyu Ding Ying Shan and Xihui Liu. 2024b. Moto: Latent motion token as the bridging language for robot manipulation. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2412.04445 (2024)."},{"key":"e_1_3_3_2_15_1","first-page":"1930","volume-title":"Conference on Robot Learning","author":"Das Neha","year":"2021","unstructured":"Neha Das, Sarah Bechtle, Todor Davchev, Dinesh Jayaraman, Akshara Rai, and Franziska Meier. 2021. Model-based inverse reinforcement learning from visual demonstrations. In Conference on Robot Learning. PMLR, 1930\u20131942."},{"key":"e_1_3_3_2_16_1","unstructured":"Yilun Du Sherry Yang Bo Dai Hanjun Dai Ofir Nachum Josh Tenenbaum Dale Schuurmans and Pieter Abbeel. 2024. Learning universal policies via text-guided video generation. NeurIPS 36 (2024)."},{"key":"e_1_3_3_2_17_1","unstructured":"Alejandro Escontrela Ademi Adeniji Wilson Yan Ajay Jain Xue\u00a0Bin Peng Ken Goldberg Youngwoon Lee Danijar Hafner and Pieter Abbeel. 2024. Video prediction models as rewards for reinforcement learning. NeurIPS 36 (2024)."},{"key":"e_1_3_3_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01842"},{"key":"e_1_3_3_2_19_1","doi-asserted-by":"crossref","unstructured":"Sami Haddadin Sven Parusel Lars Johannsmeier Saskia Golz Simon Gabl Florian Walch Mohamadreza Sabaghian Christoph J\u00e4hne Lukas Hausperger and Simon Haddadin. 2022. The franka emika robot: A reference platform for robotics research and education. IEEE Robotics & Automation Magazine 29 2 (2022) 46\u201364.","DOI":"10.1109\/MRA.2021.3138382"},{"key":"e_1_3_3_2_20_1","unstructured":"Danijar Hafner Timothy Lillicrap Jimmy Ba and Mohammad Norouzi. 2019. Dream to control: Learning behaviors by latent imagination. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1912.01603 (2019)."},{"key":"e_1_3_3_2_21_1","doi-asserted-by":"crossref","unstructured":"Haoran He Chenjia Bai Ling Pan Weinan Zhang Bin Zhao and Xuelong Li. 2024. Learning an actionable discrete diffusion policy via large-scale actionless video pre-training. NeurIPS 37 (2024) 31124\u201331153.","DOI":"10.52202\/079017-0981"},{"key":"e_1_3_3_2_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01553"},{"key":"e_1_3_3_2_23_1","unstructured":"Zhi Hou Tianyi Zhang Yuwen Xiong Hengjun Pu Chengyang Zhao Ronglei Tong Yu Qiao Jifeng Dai and Yuntao Chen. 2024. Diffusion transformer policy. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.15959 (2024)."},{"key":"e_1_3_3_2_24_1","unstructured":"Aadhithya Iyer Zhuoran Peng Yinlong Dai Irmak Guzey Siddhant Haldar Soumith Chintala and Lerrel Pinto. 2024. Open teach: A versatile teleoperation system for robotic manipulation. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2403.07870 (2024)."},{"key":"e_1_3_3_2_25_1","volume-title":"ICML","author":"Karamcheti Siddharth","year":"2024","unstructured":"Siddharth Karamcheti, Suraj Nair, Ashwin Balakrishna, Percy Liang, Thomas Kollar, and Dorsa Sadigh. 2024. Prismatic vlms: Investigating the design space of visually-conditioned language models. In ICML."},{"key":"e_1_3_3_2_26_1","unstructured":"Moo\u00a0Jin Kim Karl Pertsch Siddharth Karamcheti Ted Xiao Ashwin Balakrishna Suraj Nair Rafael Rafailov Ethan Foster Grace Lam Pannag Sanketi et\u00a0al. 2024. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2406.09246 (2024)."},{"key":"e_1_3_3_2_27_1","unstructured":"Peiyan Li Hongtao Wu Yan Huang Chilam Cheang Liang Wang and Tao Kong. 2025. GR-MG: Leveraging Partially-Annotated Data Via Multi-Modal Goal-Conditioned Policy. RA-L (2025)."},{"key":"e_1_3_3_2_28_1","unstructured":"Bo Liu Yifeng Zhu Chongkai Gao Yihao Feng Qiang Liu Yuke Zhu and Peter Stone. 2024. Libero: Benchmarking knowledge transfer for lifelong robot learning. NeurIPS 36 (2024)."},{"key":"e_1_3_3_2_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA.2018.8462901"},{"key":"e_1_3_3_2_30_1","first-page":"1113","volume-title":"Conference on robot learning","author":"Lynch Corey","year":"2020","unstructured":"Corey Lynch, Mohi Khansari, Ted Xiao, Vikash Kumar, Jonathan Tompson, Sergey Levine, and Pierre Sermanet. 2020. Learning latent plans from play. In Conference on robot learning. Pmlr, 1113\u20131132."},{"key":"e_1_3_3_2_31_1","volume-title":"The Eleventh International Conference on Learning Representations","author":"Ma Yecheng\u00a0Jason","year":"2022","unstructured":"Yecheng\u00a0Jason Ma, Shagun Sodhani, Dinesh Jayaraman, Osbert Bastani, Vikash Kumar, and Amy Zhang. 2022. VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training. In The Eleventh International Conference on Learning Representations."},{"key":"e_1_3_3_2_32_1","doi-asserted-by":"crossref","unstructured":"Oier Mees Lukas Hermann Erick Rosete-Beas and Wolfram Burgard. 2022. Calvin: A benchmark for language-conditioned policy learning for long-horizon robot manipulation tasks. RA-L 7 3 (2022) 7327\u20137334.","DOI":"10.1109\/LRA.2022.3180108"},{"key":"e_1_3_3_2_33_1","unstructured":"Suraj Nair Aravind Rajeswaran Vikash Kumar Chelsea Finn and Abhinav Gupta. 2022. R3m: A universal visual representation for robot manipulation. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2203.12601 (2022)."},{"key":"e_1_3_3_2_34_1","first-page":"892","volume-title":"Conference on Robot Learning","author":"Nair Suraj","year":"2023","unstructured":"Suraj Nair, Aravind Rajeswaran, Vikash Kumar, Chelsea Finn, and Abhinav Gupta. 2023. R3M: A Universal Visual Representation for Robot Manipulation. In Conference on Robot Learning. PMLR, 892\u2013909."},{"key":"e_1_3_3_2_35_1","unstructured":"Jianmo Ni Gustavo\u00a0Hernandez Abrego Noah Constant Ji Ma Keith\u00a0B Hall Daniel Cer and Yinfei Yang. 2021. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2108.08877 (2021)."},{"key":"e_1_3_3_2_36_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICRA57147.2024.10611477"},{"key":"e_1_3_3_2_37_1","unstructured":"Krishan Rana Andrew Melnik and Niko S\u00fcnderhauf. 2023. Contrastive language action and state pre-training for robot learning. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2304.10782 (2023)."},{"key":"e_1_3_3_2_38_1","doi-asserted-by":"crossref","unstructured":"Lin Shao Toki Migimatsu Qiang Zhang Karen Yang and Jeannette Bohg. 2021. Concept2robot: Learning manipulation concepts from instructions and human demonstrations. IJRR 40 12-14 (2021) 1419\u20131434.","DOI":"10.1177\/02783649211046285"},{"key":"e_1_3_3_2_39_1","unstructured":"Pratyusha Sharma Deepak Pathak and Abhinav Gupta. 2019. Third-person visual imitation learning via decoupled hierarchical controller. NeurIPS 32 (2019)."},{"key":"e_1_3_3_2_40_1","first-page":"654","volume-title":"Conference on Robot Learning","author":"Shaw Kenneth","year":"2023","unstructured":"Kenneth Shaw, Shikhar Bahl, and Deepak Pathak. 2023. Videodex: Learning dexterity from internet videos. In Conference on Robot Learning. PMLR, 654\u2013665."},{"key":"e_1_3_3_2_41_1","doi-asserted-by":"crossref","unstructured":"Laura Smith Nikita Dhawan Marvin Zhang Pieter Abbeel and Sergey Levine. 2019. Avid: Learning multi-stage tasks via pixel-level translation of human videos. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/1912.04443 (2019).","DOI":"10.15607\/RSS.2020.XVI.024"},{"key":"e_1_3_3_2_42_1","unstructured":"Octo\u00a0Model Team Dibya Ghosh Homer Walke Karl Pertsch Kevin Black Oier Mees Sudeep Dasari Joey Hejna Tobias Kreiman Charles Xu et\u00a0al. 2024. Octo: An open-source generalist robot policy. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2405.12213 (2024)."},{"key":"e_1_3_3_2_43_1","doi-asserted-by":"crossref","unstructured":"Mohammad\u00a0Hassan Vali and Tom B\u00e4ckstr\u00f6m. 2022. NSVQ: Noise substitution in vector quantization for machine learning. IEEE Access 10 (2022) 13598\u201313610.","DOI":"10.1109\/ACCESS.2022.3147670"},{"key":"e_1_3_3_2_44_1","doi-asserted-by":"crossref","unstructured":"Jialong Wu Shaofeng Yin Ningya Feng Xu He Dong Li Jianye Hao and Mingsheng Long. 2024. ivideogpt: Interactive videogpts are scalable world models. NeurIPS 37 (2024) 68082\u201368119.","DOI":"10.52202\/079017-2173"},{"key":"e_1_3_3_2_45_1","unstructured":"Tete Xiao Ilija Radosavovic Trevor Darrell and Jitendra Malik. 2022. Masked visual pre-training for motor control. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2203.06173 (2022)."},{"key":"e_1_3_3_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/IROS51168.2021.9636080"},{"key":"e_1_3_3_2_47_1","unstructured":"Mengjiao Yang Yilun Du Kamyar Ghasemipour Jonathan Tompson Dale Schuurmans and Pieter Abbeel. 2023. Learning interactive real-world simulators. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2310.06114 (2023)."},{"key":"e_1_3_3_2_48_1","unstructured":"Seonghyeon Ye Joel Jang Byeongguk Jeon Sejune Joo Jianwei Yang Baolin Peng Ajay Mandlekar Reuben Tan Yu-Wei Chao Bill\u00a0Yuchen Lin et\u00a0al. 2024. Latent action pretraining from videos. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2410.11758 (2024)."},{"key":"e_1_3_3_2_49_1","first-page":"537","volume-title":"Conference on Robot Learning","author":"Zakka Kevin","year":"2022","unstructured":"Kevin Zakka, Andy Zeng, Pete Florence, Jonathan Tompson, Jeannette Bohg, and Debidatta Dwibedi. 2022. Xirl: Cross-embodiment inverse reinforcement learning. In Conference on Robot Learning. PMLR, 537\u2013546."},{"key":"e_1_3_3_2_50_1","unstructured":"Chuning Zhu Raymond Yu Siyuan Feng Benjamin Burchfiel Paarth Shah and Abhishek Gupta. 2025. Unified world models: Coupling video and action diffusion for pretraining on large robotic datasets. arXiv preprint arXiv:https:\/\/arXiv.org\/abs\/2504.02792 (2025)."}],"event":{"name":"SA Conference Papers '25: SIGGRAPH Asia 2025 Conference Papers","location":"Hong Kong Hong Kong","acronym":"SA Conference Papers '25","sponsor":["SIGGRAPH ACM Special Interest Group on Computer Graphics and Interactive Techniques"]},"container-title":["Proceedings of the SIGGRAPH Asia 2025 Conference Papers"],"original-title":[],"link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3757377.3763966","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,12,9]],"date-time":"2025-12-09T03:28:10Z","timestamp":1765250890000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3757377.3763966"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,12,14]]},"references-count":49,"alternative-id":["10.1145\/3757377.3763966","10.1145\/3757377"],"URL":"https:\/\/doi.org\/10.1145\/3757377.3763966","relation":{},"subject":[],"published":{"date-parts":[[2025,12,14]]},"assertion":[{"value":"2025-12-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}