{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,4]],"date-time":"2026-07-04T01:18:07Z","timestamp":1783127887390,"version":"3.54.6"},"reference-count":56,"publisher":"Association for Computing Machinery (ACM)","issue":"4","funder":[{"DOI":"10.13039\/100000001","name":"National Science Foundation","doi-asserted-by":"publisher","award":["2219864"],"award-info":[{"award-number":["2219864"]}],"id":[{"id":"10.13039\/100000001","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2025,8,1]]},"abstract":"<jats:p>While large vision-language models can generate motion graphics animations from text prompts, they regularly fail to include all spatio-temporal properties described in the prompt. We introduce MoVer, a motion verification DSL based on first-order logic that can check spatio-temporal properties of a motion graphics animation. We identify a general set of such properties that people commonly use to describe animations (e.g., the direction and timing of motions, the relative positioning of objects, etc.). We implement these properties as predicates in MoVer and provide an execution engine that can apply a MoVer program to any input SVG-based motion graphics animation. We then demonstrate how MoVer can be used in an LLM-based synthesis and verification pipeline for iteratively refining motion graphics animations. Given a text prompt, our pipeline synthesizes a motion graphics animation and a corresponding MoVer program. Executing the verification program on the animation yields a report of the predicates that failed and the report can be automatically fed back to LLM to iteratively correct the animation. To evaluate our pipeline, we build a synthetic dataset of 5600 text prompts paired with ground truth MoVer verification programs. We find that while our LLM-based pipeline is able to automatically generate a correct motion graphics animation for 58.8% of the test prompts without any iteration, this number raises to 93.6% with up to 50 correction iterations. Our code and dataset are at https:\/\/mover-dsl.github.io.<\/jats:p>","DOI":"10.1145\/3731209","type":"journal-article","created":{"date-parts":[[2025,7,27]],"date-time":"2025-07-27T04:02:22Z","timestamp":1753588942000},"page":"1-17","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":10,"title":["MoVer: Motion Verification for Motion Graphics Animations"],"prefix":"10.1145","volume":"44","author":[{"ORCID":"https:\/\/orcid.org\/0000-0003-2880-8506","authenticated-orcid":false,"given":"Jiaju","family":"Ma","sequence":"first","affiliation":[{"name":"Stanford University, Stanford, California, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8996-7327","authenticated-orcid":false,"given":"Maneesh","family":"Agrawala","sequence":"additional","affiliation":[{"name":"Stanford University, Stanford, California, USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2025,7,27]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"crossref","unstructured":"Ola \u00c5kerberg Hans Svensson Bastian Schulz and Pierre Nugues. 2003. CarSim: an automatic 3D text-to-scene conversion system applied to road accident reports. In Demonstrations.","DOI":"10.3115\/1067737.1067782"},{"key":"e_1_2_2_2_1","doi-asserted-by":"publisher","DOI":"10.1145\/182.358434"},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1162\/tacl_a_00209"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/310930.310975"},{"key":"e_1_2_2_5_1","unstructured":"Philippe Balbiani Jean-Fran\u00e7ois Condotta and L Farinas Del Cerro. 1998. A model for reasoning about bidimensional temporal relations. In PRINCIPLES OF KNOWLEDGE REPRESENTATION AND REASONING-INTERNATIONAL CONFERENCE-. MORGAN KAUFMANN PUBLISHERS 124\u2013130."},{"key":"e_1_2_2_6_1","doi-asserted-by":"publisher","DOI":"10.1109\/ACCESS.2023.3304903"},{"key":"e_1_2_2_7_1","doi-asserted-by":"crossref","unstructured":"Haoxin Chen Yong Zhang Xiaodong Cun Menghan Xia Xintao Wang Chao Weng and Ying Shan. 2024. VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion Models. arXiv:2401.09047 [cs.CV]","DOI":"10.1109\/CVPR52733.2024.00698"},{"key":"e_1_2_2_8_1","unstructured":"Jaemin Cho Yushi Hu Roopal Garg Peter Anderson Ranjay Krishna Jason Baldridge Mohit Bansal Jordi Pont-Tuset and Su Wang. 2024. Davidsonian Scene Graph: Improving Reliability in Fine-Grained Evaluation for Text-to-Image Generation. In ICLR."},{"key":"e_1_2_2_9_1","volume-title":"Introduction to algorithms","author":"Cormen Thomas H","unstructured":"Thomas H Cormen, Charles E Leiserson, Ronald L Rivest, and Clifford Stein. 2022. Introduction to algorithms. MIT press."},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3613904.3642579"},{"key":"e_1_2_2_11_1","unstructured":"Jack Doyle. 2025. Greensock Animation Platform v3. https:\/\/gsap.com\/docs\/v3\/"},{"key":"e_1_2_2_12_1","volume-title":"International conference on theory and applications of satisfiability testing. Springer, 502\u2013518","author":"E\u00e9n Niklas","year":"2003","unstructured":"Niklas E\u00e9n and Niklas S\u00f6rensson. 2003. An extensible SAT-solver. In International conference on theory and applications of satisfiability testing. Springer, 502\u2013518."},{"key":"e_1_2_2_13_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00414"},{"key":"e_1_2_2_14_1","unstructured":"Julien Garnier. 2025. anime.js \u00b7 JavaScript animation engine. https:\/\/animejs.com\/"},{"key":"e_1_2_2_15_1","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657447"},{"key":"e_1_2_2_16_1","doi-asserted-by":"crossref","unstructured":"Yuwei Guo Ceyuan Yang Anyi Rao Maneesh Agrawala Dahua Lin and Bo Dai. 2023a. SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models. arXiv:2311.16933 [cs.CV]","DOI":"10.1007\/978-3-031-72946-1_19"},{"key":"e_1_2_2_17_1","unstructured":"Yuwei Guo Ceyuan Yang Anyi Rao Zhengyang Liang Yaohui Wang Yu Qiao Maneesh Agrawala Dahua Lin and Bo Dai. 2023b. AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning."},{"key":"e_1_2_2_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52729.2023.01436"},{"key":"e_1_2_2_19_1","volume-title":"Ronan Le Bras, and Yejin Choi","author":"Hessel Jack","year":"2021","unstructured":"Jack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras, and Yejin Choi. 2021. CLIPScore: A Reference-free Evaluation Metric for Image Captioning. ArXiv abs\/2104.08718 (2021). https:\/\/api.semanticscholar.org\/CorpusID:233296711"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.5281\/zenodo.1212303"},{"key":"e_1_2_2_21_1","volume-title":"Proceedings of the 37th International Conference on Neural Information Processing Systems","author":"Hsu Joy","year":"2024","unstructured":"Joy Hsu, Jiayuan Mao, Joshua B. Tenenbaum, and Jiajun Wu. 2024. What's Left? concept grounding with logic-enhanced foundation models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS '23). Curran Associates Inc., Red Hook, NY, USA, Article 1684, 17 pages."},{"key":"e_1_2_2_22_1","volume-title":"Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 20406\u201320417","author":"Hu Yushi","unstructured":"Yushi Hu, Benlin Liu, Jungo Kasai, Yizhong Wang, Mari Ostendorf, Ranjay Krishna, and Noah A. Smith. 2023. TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering. In Proceedings of the IEEE\/CVF International Conference on Computer Vision (ICCV). 20406\u201320417."},{"key":"e_1_2_2_23_1","volume-title":"Proceedings of the 41st International Conference on Machine Learning","author":"Hu Ziniu","year":"2025","unstructured":"Ziniu Hu, Ahmet Iscen, Aashi Jain, Thomas Kipf, Yisong Yue, David A Ross, Cordelia Schmid, and Alireza Fathi. 2025. SceneCraft: an LLM agent for synthesizing 3D scenes as blender code. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML'24). JMLR.org, Article 776, 31 pages."},{"key":"e_1_2_2_24_1","unstructured":"Aaron Hurst Adam Lerer Adam P Goucher Adam Perelman Aditya Ramesh Aidan Clark AJ Ostrow Akila Welihinda Alan Hayes Alec Radford et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024)."},{"key":"e_1_2_2_25_1","doi-asserted-by":"publisher","DOI":"10.1145\/502348.502379"},{"key":"e_1_2_2_26_1","volume-title":"Proceedings of the IEEE International Conference on Computer Vision (ICCV).","author":"Johnson Justin","unstructured":"Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick. 2017. Inferring and Executing Programs for Visual Reasoning. In Proceedings of the IEEE International Conference on Computer Vision (ICCV)."},{"key":"e_1_2_2_27_1","volume-title":"A Solver-Aided Hierarchical Language for LLM-Driven CAD Design. arXiv preprint arXiv:2502.09819","author":"Jones Benjamin T","year":"2025","unstructured":"Benjamin T Jones, Felix H\u00e4hnlein, Zihan Zhang, Maaz Ahmad, Vladimir Kim, and Adriana Schulz. 2025. A Solver-Aided Hierarchical Language for LLM-Driven CAD Design. arXiv preprint arXiv:2502.09819 (2025)."},{"key":"e_1_2_2_28_1","volume-title":"A survey on semantic parsing. arXiv preprint arXiv:1812.00978","author":"Kamath Aishwarya","year":"2018","unstructured":"Aishwarya Kamath and Rajarshi Das. 2018. A survey on semantic parsing. arXiv preprint arXiv:1812.00978 (2018)."},{"key":"e_1_2_2_29_1","volume-title":"Didier Stricker, Sk Aziz Ali, and Muhammad Zeshan Afzal.","author":"Khan Mohammad Sadil","year":"2024","unstructured":"Mohammad Sadil Khan, Sankalp Sinha, Talha Uddin Sheikh, Didier Stricker, Sk Aziz Ali, and Muhammad Zeshan Afzal. 2024. Text2CAD: Generating Sequential CAD Models from Beginner-to-Expert Level Text Prompts. arXiv:2409.17106 [cs.CV] https:\/\/arxiv.org\/abs\/2409.17106"},{"key":"e_1_2_2_30_1","volume-title":"Proceedings of the 17th International Natural Language Generation Conference: Generation Challenges, Simon Mille and Miruna-Adriana Clinciu (Eds.). Association for Computational Linguistics","author":"Lapalme Guy","year":"2024","unstructured":"Guy Lapalme. 2024. pyrealb at the GEM'24 Data-to-text Task: Symbolic English Text Generation from RDF Triples. In Proceedings of the 17th International Natural Language Generation Conference: Generation Challenges, Simon Mille and Miruna-Adriana Clinciu (Eds.). Association for Computational Linguistics, Tokyo, Japan, 54\u201358. https:\/\/aclanthology.org\/2024.inlg-genchal.5\/"},{"key":"e_1_2_2_31_1","volume-title":"Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Benita Teufel, Marco Bellagente, Minguk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Li Fei-Fei, Jiajun Wu, Stefano Ermon, and Percy Liang.","author":"Lee Tony","year":"2023","unstructured":"Tony Lee, Michihiro Yasunaga, Chenlin Meng, Yifan Mai, Joon Sung Park, Agrim Gupta, Yunzhi Zhang, Deepak Narayanan, Hannah Benita Teufel, Marco Bellagente, Minguk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Li Fei-Fei, Jiajun Wu, Stefano Ermon, and Percy Liang. 2023. Holistic Evaluation of Text-To-Image Models. arXiv:2311.04287 [cs.CV] https:\/\/arxiv.org\/abs\/2311.04287"},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1145\/3706598.3714155"},{"key":"e_1_2_2_33_1","volume-title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint arXiv:1907.11692","author":"Liu Yinhan","year":"2019","unstructured":"Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. arXiv preprint arXiv:1907.11692 (2019)."},{"key":"e_1_2_2_34_1","unstructured":"Minhua Ma. 2006. Automatic conversion of natural language to 3D animation. Ph.D. Dissertation. University of Ulster."},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/3172944.3172972"},{"key":"e_1_2_2_36_1","volume-title":"WordNet: A Lexical Database for English. In Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8\u201311","author":"Miller George A.","year":"1994","unstructured":"George A. Miller. 1994. WordNet: A Lexical Database for English. In Human Language Technology: Proceedings of a Workshop held at Plainsboro, New Jersey, March 8\u201311, 1994. https:\/\/aclanthology.org\/H94-1111\/"},{"key":"e_1_2_2_37_1","unstructured":"MozDevNet. 2025. Using CSS animations - CSS: Cascading style sheets: MDN. https:\/\/developer.mozilla.org\/en-US\/docs\/Web\/CSS\/CSS_animations\/Using_CSS_animations"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10472-012-9327-5"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.aacl-demo.6"},{"key":"e_1_2_2_40_1","unstructured":"Sagi Polaczek Yuval Alaluf Elad Richardson Yael Vinker and Daniel Cohen-Or. 2025. NeuralSVG: An Implicit Representation for Text-to-Vector Generation. arXiv:2501.03992 [cs.CV] https:\/\/arxiv.org\/abs\/2501.03992"},{"key":"e_1_2_2_41_1","volume-title":"Design2code: How far are we from automating front-end engineering? arXiv e-prints","author":"Si Chenglei","year":"2024","unstructured":"Chenglei Si, Yanzhe Zhang, Zhengyuan Yang, Ruibo Liu, and Diyi Yang. 2024. Design2code: How far are we from automating front-end engineering? arXiv e-prints (2024), arXiv-2403."},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52733.2024.00826"},{"key":"e_1_2_2_43_1","volume-title":"DreamSync: Aligning Text-to-Image Generation with Image Understanding Models. In Synthetic Data for Computer Vision Workshop @ CVPR","author":"Sun Jiao","year":"2024","unstructured":"Jiao Sun, Yushi Hu, Deqing Fu, Royi Rassin, Sjoerd van Steenkiste, Dana Alon, Su Wang, Charles Herrmann, Ranjay Krishna, Da-Cheng Juan, and Cyrus Rashtchian. 2024. DreamSync: Aligning Text-to-Image Generation with Image Understanding Models. In Synthetic Data for Computer Vision Workshop @ CVPR 2024. https:\/\/arxiv.org\/abs\/2311.17946"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV51070.2023.01092"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1163\/9789004368828_008"},{"key":"e_1_2_2_46_1","volume-title":"Spatial orientation: Theory, research, and application","author":"Talmy Leonard","unstructured":"Leonard Talmy. 1983. How language structures space. In Spatial orientation: Theory, research, and application. Springer, 225\u2013282."},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-08-050754-5.50069-4"},{"key":"e_1_2_2_48_1","volume-title":"Keyframer: Empowering Animation Design using Large Language Models. arXiv:2402.06071 [cs.HC] https:\/\/arxiv.org\/abs\/2402.06071","author":"Tseng Tiffany","year":"2024","unstructured":"Tiffany Tseng, Ruijia Cheng, and Jeffrey Nichols. 2024. Keyframer: Empowering Animation Design using Large Language Models. arXiv:2402.06071 [cs.HC] https:\/\/arxiv.org\/abs\/2402.06071"},{"key":"e_1_2_2_49_1","doi-asserted-by":"publisher","DOI":"10.48550\/ARXIV.2209.11055"},{"key":"e_1_2_2_50_1","doi-asserted-by":"publisher","DOI":"10.1007\/3-540-69342-4_8"},{"key":"e_1_2_2_51_1","volume-title":"Compositionally Generalizable Robotic Manipulation. In The Eleventh International Conference on Learning Representations.","author":"Wang Renhao","year":"2023","unstructured":"Renhao Wang, Jiayuan Mao, Joy Hsu, Hang Zhao, Jiajun Wu, and Yang Gao. 2023. Programmatically Grounded, Compositionally Generalizable Robotic Manipulation. In The Eleventh International Conference on Learning Representations."},{"key":"e_1_2_2_52_1","doi-asserted-by":"publisher","DOI":"10.1145\/3687761"},{"key":"e_1_2_2_53_1","unstructured":"Ximing Xing Juncheng Hu Guotao Liang Jing Zhang Dong Xu and Qian Yu. 2024a. Empowering LLMs to Understand and Generate Complex Vector Graphics. (2024)."},{"key":"e_1_2_2_54_1","doi-asserted-by":"publisher","DOI":"10.1145\/3618316"},{"key":"e_1_2_2_55_1","volume-title":"The Scene Language: Representing Scenes with Programs, Words, and Embeddings. arXiv preprint arXiv:2410.16770","author":"Zhang Yunzhi","year":"2024","unstructured":"Yunzhi Zhang, Zizhang Li, Matt Zhou, Shangzhe Wu, and Jiajun Wu. 2024. The Scene Language: Representing Scenes with Programs, Words, and Embeddings. arXiv preprint arXiv:2410.16770 (2024)."},{"key":"e_1_2_2_56_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/S19-1032"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3731209","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T17:54:41Z","timestamp":1774634081000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3731209"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,7,27]]},"references-count":56,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2025,8,1]]}},"alternative-id":["10.1145\/3731209"],"URL":"https:\/\/doi.org\/10.1145\/3731209","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2025,7,27]]},"assertion":[{"value":"2025-07-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}