{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,3]],"date-time":"2026-05-03T00:23:05Z","timestamp":1777767785432,"version":"3.51.4"},"reference-count":77,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T00:00:00Z","timestamp":1774569600000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/onlinelibrary.wiley.com\/termsAndConditions#vor"},{"start":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T00:00:00Z","timestamp":1774569600000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"publisher","award":["62172367"],"award-info":[{"award-number":["62172367"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Sketching is a direct and inexpensive means of visual expression. Though image\u2010based sketching has been well studied, video\u2010based sketch animation generation is still very challenging due to the temporal coherence requirement. In this paper, we propose a novel end\u2010to\u2010end automatic generation approach for vector sketch animation. To solve the flickering issue, we introduce a Differentiable Motion Trajectory (DMT) representation that describes the frame\u2010wise movement of stroke control points using differentiable polynomial\u2010based trajectories. DMT enables global semantic gradient propagation across multiple frames, significantly improving the semantic consistency and temporal coherence, and producing high\u2010framerate output. DMT employs a Bernstein basis to balance the sensitivity of polynomial parameters, thus achieving more stable optimization. Instead of implicit fields, we introduce sparse track points for explicit spatial modeling, which improves efficiency and supports long\u2010duration video processing. Evaluations on DAVIS and LVOS datasets demonstrate the superiority of our approach over SOTA methods. Cross\u2010domain validation on 3D models and text\u2010to\u2010video data confirms the robustness and compatibility of our approach.<\/jats:p>","DOI":"10.1111\/cgf.70335","type":"journal-article","created":{"date-parts":[[2026,3,27]],"date-time":"2026-03-27T13:27:23Z","timestamp":1774618043000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Vector sketch animation generation with differentiable motion trajectories"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-3100-8481","authenticated-orcid":false,"given":"X.","family":"Zhu","sequence":"first","affiliation":[{"name":"Zhejiang University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0004-0865-3783","authenticated-orcid":false,"given":"X.","family":"Yang","sequence":"additional","affiliation":[{"name":"Zhejiang University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"S.","family":"Zheng","sequence":"additional","affiliation":[{"name":"Zhejiang University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0009-0007-5597-504X","authenticated-orcid":false,"given":"Z.","family":"Zhang","sequence":"additional","affiliation":[{"name":"Hangzhou Dianzi University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"F.","family":"Gao","sequence":"additional","affiliation":[{"name":"Zhejiang University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"J.","family":"Huang","sequence":"additional","affiliation":[{"name":"Zhejiang Gongshang University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-2780-6146","authenticated-orcid":false,"given":"J.","family":"Chen","sequence":"additional","affiliation":[{"name":"Zhejiang University of Technology"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,3,27]]},"reference":[{"key":"e_1_2_10_2_2","doi-asserted-by":"publisher","DOI":"10.1038\/nature13422"},{"key":"e_1_2_10_3_2","doi-asserted-by":"crossref","unstructured":"ArarE. FrenkelY. Cohen-OrD. ShamirA. VinkerY.: Swiftsketch: A diffusion model for image-to-vector sketch generation. InProceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers(2025) pp.1\u201312. 2 10","DOI":"10.1145\/3721238.3730612"},{"key":"e_1_2_10_4_2","doi-asserted-by":"publisher","DOI":"10.1145\/1015706.1015764"},{"key":"e_1_2_10_5_2","doi-asserted-by":"crossref","unstructured":"BellJ. B.:Solutions of ill-posed problems. 1978. 11","DOI":"10.2307\/2006360"},{"key":"e_1_2_10_6_2","doi-asserted-by":"publisher","DOI":"10.1145\/566654.566595"},{"key":"e_1_2_10_7_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2011.185"},{"key":"e_1_2_10_8_2","doi-asserted-by":"publisher","DOI":"10.1145\/2461912.2461964"},{"key":"e_1_2_10_9_2","first-page":"679","article-title":"A computational approach to edge detection","volume":"6","author":"Canny J.","year":"2009","journal-title":"IEEE Transactions on pattern analysis and machine intelligence"},{"key":"e_1_2_10_10_2","first-page":"16351","article-title":"Deepsvg: A hierarchical generative network for vector graphics animation","volume":"33","author":"Carlier A.","year":"2020","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_10_11_2","unstructured":"CuiQ. WuZ. XingC. ZhouZ. WuW.: Target temperature driven dynamic flame animation. InEurographics (Short Papers)(2015) pp.45\u201348. 3"},{"key":"e_1_2_10_12_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.14946"},{"key":"e_1_2_10_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4612-6333-3"},{"key":"e_1_2_10_14_2","unstructured":"DoubleL:RPG animations pack FREE. Unity Asset Store 2025. Accessed: 2025-09-01. URL:https:\/\/assetstore.unity.com\/packages\/3d\/animations\/rpg-animations-pack-free-288783. 8"},{"key":"e_1_2_10_15_2","doi-asserted-by":"publisher","DOI":"10.1145\/2766913"},{"key":"e_1_2_10_16_2","first-page":"632","volume-title":"European conference on computer vision","author":"Das A.","year":"2020"},{"key":"e_1_2_10_17_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cagd.2012.03.001"},{"key":"e_1_2_10_18_2","doi-asserted-by":"publisher","DOI":"10.1038\/s44159-023-00212-w"},{"key":"e_1_2_10_19_2","unstructured":"FangX. ChangM.: Video sketching using multi-domain guidance and implicit encoding: X. fang m. chang.The Visual Computer(2025) 1\u201312. 2 3 6 7 10 11"},{"key":"e_1_2_10_20_2","doi-asserted-by":"publisher","DOI":"10.1098\/rsta.1922.0009"},{"key":"e_1_2_10_21_2","doi-asserted-by":"publisher","DOI":"10.52202\/068431-0376"},{"key":"e_1_2_10_22_2","volume-title":"Deep learning","author":"Goodfellow I.","year":"2016"},{"key":"e_1_2_10_23_2","doi-asserted-by":"publisher","DOI":"10.1145\/2766893"},{"key":"e_1_2_10_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3355089.3356533"},{"key":"e_1_2_10_25_2","doi-asserted-by":"crossref","unstructured":"GalR. VinkerY. AlalufY. BermanoA. Cohen-OrD. ShamirA. ChechikG.: Breathing life into sketches using text-to-video priors. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024) pp.4325\u20134336. 3 7 9","DOI":"10.1109\/CVPR52733.2024.00414"},{"key":"e_1_2_10_26_2","unstructured":"HaD. EckD.: A neural representation of sketch drawings.arXiv preprint arXiv:1704.03477(2017). 2"},{"key":"e_1_2_10_27_2","doi-asserted-by":"crossref","unstructured":"HinzT. FisherM. WangO. ShechtmanE. WermterS.: Charactergan: Few-shot keypoint character animation and reposing. InProceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision(2022) pp.1988\u20131997. 3","DOI":"10.1109\/WACV51458.2022.00324"},{"key":"e_1_2_10_28_2","doi-asserted-by":"publisher","DOI":"10.1080\/00401706.1970.10488634"},{"key":"e_1_2_10_29_2","unstructured":"HongL. LiuZ. ChenW. TanC. FengY. ZhouX. GuoP. LiJ. ChenZ. GaoS. et al.: Lvos: A benchmark for large-scale long-term video object segmentation.IEEE Transactions on Pattern Analysis and Machine Intelligence(2025). 2 8"},{"key":"e_1_2_10_30_2","unstructured":"HastieT. TibshiraniR. FriedmanJ. et al.:The elements of statistical learning 2009. 4"},{"key":"e_1_2_10_31_2","unstructured":"IglesiasK.:Human basic motions FREE. Unity Asset Store 2025. Accessed: 2025-09-01. URL:https:\/\/assetstore.unity.com\/packages\/3d\/animations\/human-basic-motions-free-154271. 8"},{"key":"e_1_2_10_32_2","unstructured":"IsolaP. ZhuJ.-Y. ZhouT. EfrosA. A.: Image-to-image translation with conditional adversarial networks. InProceedings of the IEEE conference on computer vision and pattern recognition(2017) pp.1125\u20131134. 2"},{"key":"e_1_2_10_33_2","unstructured":"JeruzalskiT. LevinD. I. JacobsonA. LalondeP. NorouziM. TagliasacchiA.: Nilbs: Neural inverse linear blend skinning.arXiv preprint arXiv:2004.05980(2020). 3"},{"key":"e_1_2_10_34_2","doi-asserted-by":"crossref","unstructured":"JainA. XieA. AbbeelP.: Vectorfusion: Text-to-svg by abstracting pixel-based diffusion models. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2023) pp.1911\u20131920. 2","DOI":"10.1109\/CVPR52729.2023.00190"},{"key":"e_1_2_10_35_2","unstructured":"KhandelwalA.: Flexiclip: Locality-preserving free-form character animation.arXiv preprint arXiv:2501.08676(2025). 3"},{"key":"e_1_2_10_36_2","unstructured":"KaraevN. MakarovI. WangJ. NeverovaN. VedaldiA. RupprechtC.: Cotracker3: Simpler and better point tracking by pseudo-labelling real videos.arXiv preprint arXiv:2410.11831(2024). 6 11"},{"key":"e_1_2_10_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3478513.3480546"},{"key":"e_1_2_10_38_2","doi-asserted-by":"crossref","unstructured":"LeeJ. ChoiC. KimY. M. ParkJ.: Recovering dynamic 3d sketches from videos. InProceedings of the Computer Vision and Pattern Recognition Conference(2025) pp.12423\u201312432. 3","DOI":"10.1109\/CVPR52734.2025.01159"},{"key":"e_1_2_10_39_2","doi-asserted-by":"crossref","unstructured":"LiangG. HuJ. XingX. ZhangJ. YuQ.: Multi-object sketch animation with grouping and motion trajectory priors.arXiv preprint arXiv:2508.15535(2025). 3 9","DOI":"10.1145\/3746027.3754502"},{"key":"e_1_2_10_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3414685.3417763"},{"key":"e_1_2_10_41_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00154"},{"key":"e_1_2_10_42_2","unstructured":"LiuZ. NingJ. CaoY. WeiY. ZhangZ. LinS. HuH.: Video swin transformer. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2022) pp.3202\u20133211. 6"},{"key":"e_1_2_10_43_2","unstructured":"LiuJ. XinZ. FuY. ZhaoR. LanB. LiX.: Multi-object sketch animation by scene decomposition and motion planning.arXiv preprint arXiv:2503.19351(2025). 3 9"},{"key":"e_1_2_10_44_2","doi-asserted-by":"publisher","DOI":"10.1145\/3306346.3323045"},{"key":"e_1_2_10_45_2","doi-asserted-by":"crossref","unstructured":"OuyangH. WangQ. XiaoY. BaiQ. ZhangJ. ZhengK. ZhouX. ChenQ. ShenY.: Codef: Content deformation fields for temporally consistent video processing. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024) pp.8089\u20138099. 2 3 6","DOI":"10.1109\/CVPR52733.2024.00773"},{"key":"e_1_2_10_46_2","unstructured":"PooleB. JainA. BarronJ. T. MildenhallB.: Dreamfusion: Text-to-3d using 2d diffusion.arXiv preprint arXiv:2209.14988(2022). 2 3"},{"key":"e_1_2_10_47_2","unstructured":"Pont-TusetJ. PerazziF. CaellesS. Arbel\u00e1ezP. Sorkine-HornungA. Van GoolL.: The 2017 davis challenge on video object segmentation.arXiv preprint arXiv:1704.00675(2017). 2"},{"key":"e_1_2_10_48_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2020.107404"},{"key":"e_1_2_10_49_2","unstructured":"RombachR. BlattmannA. LorenzD. EsserP. OmmerB.: High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2022) pp.10684\u201310695. 2"},{"key":"e_1_2_10_50_2","doi-asserted-by":"crossref","unstructured":"ReddyP. GharbiM. LukacM. MitraN. J.: Im2vec: Synthesizing vector graphics without vector supervision. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2021) pp.7342\u20137351. 2","DOI":"10.1109\/CVPR46437.2021.00726"},{"key":"e_1_2_10_51_2","unstructured":"RadfordA. KimJ. W. HallacyC. RameshA. GohG. AgarwalS. SastryG. AskellA. MishkinP. ClarkJ. et al.: Learning transferable visual models from natural language supervision. InInternational conference on machine learning(2021) PmLR pp.8748\u20138763. 2 4 7"},{"key":"e_1_2_10_52_2","doi-asserted-by":"crossref","unstructured":"RaiG. SharmaO.: Enhancing sketch animation: Text-to-video diffusion models with temporal consistency and rigidity constraints.arXiv preprint arXiv:2411.19381(2024). 3","DOI":"10.5220\/0013304800003912"},{"key":"e_1_2_10_53_2","doi-asserted-by":"crossref","unstructured":"SuQ. BaiX. FuH. TaiC.-L. WangJ.: Live sketch: Video-driven dynamic deformation of static drawings. InProceedings of the 2018 chi conference on human factors in computing systems(2018) pp.1\u201312. 3","DOI":"10.1145\/3173574.3174236"},{"key":"e_1_2_10_54_2","doi-asserted-by":"crossref","unstructured":"SantosaS. ChevalierF. BalakrishnanR. SinghK.: Direct space-time trajectory control for visual media editing. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems(2013) pp.1149\u20131158. 3","DOI":"10.1145\/2470654.2466148"},{"key":"e_1_2_10_55_2","doi-asserted-by":"publisher","DOI":"10.1142\/S0129065704001899"},{"key":"e_1_2_10_56_2","doi-asserted-by":"publisher","DOI":"10.1111\/1467-8659.00547"},{"key":"e_1_2_10_57_2","article-title":"First order motion model for image animation","volume":"32","author":"Siarohin A.","year":"2019","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_10_58_2","unstructured":"SongJ. MengC. ErmonS.: Denoising diffusion implicit models.arXiv preprint arXiv:2010.02502(2020). 3"},{"key":"e_1_2_10_59_2","doi-asserted-by":"publisher","DOI":"10.2307\/3029337"},{"key":"e_1_2_10_60_2","doi-asserted-by":"publisher","DOI":"10.1145\/3592788"},{"key":"e_1_2_10_61_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2022.3220575"},{"key":"e_1_2_10_62_2","unstructured":"TanveerM. WangY. WangR. ZhaoN. Mahdavi-AmiriA. ZhangH.: Anamodiff: 2d analogical motion diffusion via disentangled denoising.arXiv preprint arXiv:2402.03549(2024). 3"},{"key":"e_1_2_10_63_2","unstructured":"VinkerY. AlalufY. Cohen-OrD. ShamirA.: Clipascene: Scene sketching with different types and levels of abstraction. InProceedings of the IEEE\/CVF International Conference on Computer Vision(2023) pp.4146\u20134156. 2 10"},{"key":"e_1_2_10_64_2","doi-asserted-by":"publisher","DOI":"10.1145\/3528223.3530068"},{"key":"e_1_2_10_65_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cag.2012.03.004"},{"key":"e_1_2_10_66_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-024-02306-1"},{"key":"e_1_2_10_67_2","unstructured":"WangJ. YuanH. ChenD. ZhangY. WangX. ZhangS.: Modelscope text-to-video technical report.arXiv preprint arXiv:2308.06571(2023). 3"},{"key":"e_1_2_10_68_2","doi-asserted-by":"crossref","unstructured":"XieS. TuZ.: Holistically-nested edge detection. InProceedings of the IEEE international conference on computer vision(2015) pp.1395\u20131403. 6 7","DOI":"10.1109\/ICCV.2015.164"},{"key":"e_1_2_10_69_2","first-page":"15869","article-title":"Diffsketcher: Text guided vector sketch synthesis through latent diffusion models","volume":"36","author":"Xing X.","year":"2023","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_10_70_2","doi-asserted-by":"crossref","unstructured":"XingX. ZhouH. WangC. ZhangJ. XuD. YuQ.: Svgdreamer: Text guided svg generation with diffusion model. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024) pp.4546\u20134555. 2","DOI":"10.1109\/CVPR52733.2024.00435"},{"key":"e_1_2_10_71_2","unstructured":"YingliuZhizhu1:Brothers i leave it to you do as you see fit #yingliu zhizhu challenge. Bilibili 2025. [Video; in Chinese]. URL:https:\/\/www.bilibili.com\/video\/BV1pmKEzRE5H. 8 12"},{"key":"e_1_2_10_72_2","unstructured":"YangZ. TengJ. ZhengW. DingM. HuangS. XuJ. YangY. HongW. ZhangX. FengG. et al.: Cogvideox: Text-to-video diffusion models with an expert transformer.arXiv preprint arXiv:2408.06072(2024). 9"},{"key":"e_1_2_10_73_2","unstructured":"YeX. YaoJ.-H. FengJ. MeiS. LanX. ChenS.: Vidanimator: User-guided stylized 3d character animation from human videos.arXiv preprint arXiv:2508.01878(2025). 3"},{"key":"e_1_2_10_74_2","unstructured":"YanW. ZhangY. AbbeelP. SrinivasA.: Videogpt: Video generation using vq-vae and transformers.arXiv preprint arXiv:2104.10157(2021). 6"},{"key":"e_1_2_10_75_2","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.15044"},{"key":"e_1_2_10_76_2","doi-asserted-by":"publisher","DOI":"10.1109\/TVCG.2009.9"},{"key":"e_1_2_10_77_2","doi-asserted-by":"crossref","unstructured":"ZhangR. IsolaP. EfrosA. A. ShechtmanE. WangO.: The unreasonable effectiveness of deep features as a perceptual metric. InProceedings of the IEEE conference on computer vision and pattern recognition(2018) pp.586\u2013595. 7","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_10_78_2","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2015.2409119"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70335","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70335","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70335","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T15:45:00Z","timestamp":1777477500000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70335"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,27]]},"references-count":77,"alternative-id":["10.1111\/cgf.70335"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70335","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,3,27]]},"assertion":[{"value":"2026-03-27","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70335"}}