{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,4]],"date-time":"2026-05-04T04:46:46Z","timestamp":1777870006480,"version":"3.51.4"},"reference-count":25,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T00:00:00Z","timestamp":1777507200000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T00:00:00Z","timestamp":1777507200000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"funder":[{"DOI":"10.13039\/100026024","name":"Adobe Research","doi-asserted-by":"publisher","id":[{"id":"10.13039\/100026024","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    Video transitions aim to synthesize intermediate frames between two clips, but na\u00efve approaches such as linear blending introduce artifacts that limit professional use or break temporal coherence. Traditional techniques (cross\u2010fades, morphing, frame interpolation) and recent generative inbetweening methods can produce high\u2010quality plausible intermediates, but they struggle with bridging diverse clips involving large temporal gaps or significant semantic differences, leaving a gap for content\u2010aware and visually coherent transitions. We address this challenge by drawing on artistic workflows, distilling strategies such as aligning silhouettes and interpolating salient features to preserve structure and perceptual continuity. Building on these strategies, we propose SAGE (Structure\u2010Aware Generative vidEo transitions) as a simple yet effective zeroshot approach that combines structural guidance, provided via line maps and motion flow, with generative synthesis, enabling smooth, motion\u2010consistent transitions without fine\u2010tuning. Extensive experiments and comparison with current alternatives, namely [RKT*22, ZCL*24, ZZX*24, JHM*25, ZRW*25], demonstrate that SAGE outperforms both classical and the latest generative baselines on quantitative metrics and user studies for producing transitions between diverse clips. The simple method effectively bypasses the need to acquire suitable training data, which is particularly difficult in our creative setting involving diverse clips. Code is available via the project page at\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/kan32501.github.io\/sage.github.io\/\">https:\/\/kan32501.github.io\/sage.github.io\/<\/jats:ext-link>\n                    .\n                  <\/jats:p>","DOI":"10.1111\/cgf.70420","type":"journal-article","created":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T15:09:08Z","timestamp":1777561748000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["SAGE: Structure\u2010Aware Generative Video Transitions between Diverse Clips"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0007-5590-6706","authenticated-orcid":false,"given":"Mia","family":"Kan","sequence":"first","affiliation":[{"name":"University College London"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-7336-1956","authenticated-orcid":false,"given":"Yilin","family":"Liu","sequence":"additional","affiliation":[{"name":"University College London"},{"name":"Autodesk Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2597-0914","authenticated-orcid":false,"given":"Niloy J.","family":"Mitra","sequence":"additional","affiliation":[{"name":"University College London"},{"name":"Adobe Research"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,4,30]]},"reference":[{"key":"e_1_2_9_2_2","doi-asserted-by":"crossref","unstructured":"BaoW. LaiW.-S. MaC. ZhangX. GaoZ. YangM.-H.: Depth-aware video frame interpolation. InProc. Computer Vision and Pattern Recognition (CVPR)(2019). 2 3","DOI":"10.1109\/CVPR.2019.00382"},{"key":"e_1_2_9_3_2","doi-asserted-by":"crossref","unstructured":"BeierT. NeelyS.: Feature-based image metamorphosis.Proc. SIGGRAPH(1992) 35\u201342. 2 3","DOI":"10.1145\/142920.134003"},{"key":"e_1_2_9_4_2","unstructured":"ChenX. WangY. ZhangL. ZhuangS. MaX. YuJ. WangY. LinD. QiaoY. LiuZ.: SEINE: Short-to-long video diffusion model for generative transition and prediction. InInternation Conference on Learning Representations(2023). 2 3 5"},{"key":"e_1_2_9_5_2","doi-asserted-by":"crossref","unstructured":"ChengM.-M. ZhengS. LinW.-Y. VineetV. SturgessP. CrookN. MitraN. TorrP.: Imagespirit: Verbal guided image parsing.ACM Transactions on Graphics(2014). 9","DOI":"10.1145\/2682628"},{"key":"e_1_2_9_6_2","doi-asserted-by":"crossref","unstructured":"DuttN. S. MuralikrishnanS. MitraN. J.: Diffusion 3d features (diff3f): Decorating untextured shapes with distilled semantic features. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(June2024) pp.4494\u20134504. 9","DOI":"10.1109\/CVPR52733.2024.00430"},{"key":"e_1_2_9_7_2","doi-asserted-by":"crossref","unstructured":"HuangZ. ZhangT. HengW. ShiB. LiuS.: RIFE: Real-time intermediate flow estimation for video frame interpolation. InEuropean Conference on Computer Vision (ECCV)(2022). 2 3","DOI":"10.1007\/978-3-031-19781-9_36"},{"key":"e_1_2_9_8_2","unstructured":"JainA. GharbiM. ZhangR. LiuC. FreemanW. T. DurandF. DekelT.: Video interpolation with diffusion models.Proc. Computer Vision and Pattern Recognition (CVPR)(2024). 2 3"},{"key":"e_1_2_9_9_2","unstructured":"JiangZ. HanZ. MaoC. ZhangJ. PanY. LiuY.: VACE: All-in-one video creation and editing.Arxiv(2025). arXiv:2503.07598. 1 2 3 5"},{"key":"e_1_2_9_10_2","doi-asserted-by":"crossref","unstructured":"KirillovA. MintunE. RaviN. MaoH. RollandC. GustafsonL. XiaoT. WhiteheadS. BergA. C. LoW. Doll\u00e1rP. GirshickR. B.: Segment anything. InInternational Conference on Computer Vision(2023) pp.3992\u20134003. 4","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"e_1_2_9_11_2","doi-asserted-by":"crossref","unstructured":"KyeD. RohC. KoS. EomC. OhJ.: AceVFI: A comprehensive survey of advances in video frame interpolation.Arxiv(2025). arXiv:2506.01061. 3","DOI":"10.1109\/TCSVT.2026.3672288"},{"key":"e_1_2_9_12_2","unstructured":"MaryWoodcock:Mastering the art of transitions for youtube videos.https:\/\/vimeo.com\/blog\/post\/adding-video-transitions 2025. Accessed: 2025-09-20. 3"},{"key":"e_1_2_9_13_2","volume-title":"The Conversations: Walter Murch and the Art of Editing Film","author":"Ondaatje M.","year":"1900"},{"key":"e_1_2_9_14_2","volume-title":"Cutting Rhythms: Intuitive Film Editing","author":"Pearlman K.","year":"2016"},{"key":"e_1_2_9_15_2","unstructured":"PardoA. PizzatiF. ZhangT. PondavenA. TorrP. PerezJ. C. GhanemB.: Matchdiffusion: Training-free generation of match-cuts.arXiv preprint arXiv:2411.18677(2024). 3"},{"key":"e_1_2_9_16_2","doi-asserted-by":"crossref","unstructured":"PautratR. Su\u00e1rezI. YuY. PollefeysM. LarssonV.: Gluestick: Robust image matching by sticking points and lines together. InInternational Conference on Computer Vision(2023) pp.9672\u20139682. 4","DOI":"10.1109\/ICCV51070.2023.00890"},{"key":"e_1_2_9_17_2","unstructured":"RedaF. KontkanenJ. TabellionE. SunD. PantofaruC. CurlessB.: FILM: Frame interpolation for large motion. InEuropean Conference on Computer Vision (ECCV)(2022). 1 2 3 5"},{"key":"e_1_2_9_18_2","unstructured":"VivianTejeda:Video transitions: The ultimate guide in 2025.https:\/\/www.descript.com\/blog\/article\/video-transitions 2025. Accessed: 2025-09-20. 3"},{"key":"e_1_2_9_19_2","first-page":"36","article-title":"SEA-RAFT: simple, efficient, accurate RAFT for optical flow","volume":"15065","author":"Wang Y.","year":"2024","journal-title":"European Conference on Computer Vision (ECCV)"},{"key":"e_1_2_9_20_2","doi-asserted-by":"crossref","first-page":"360","DOI":"10.1007\/s003710050148","article-title":"Image morphing: a survey","volume":"14","author":"Wolberg G.","year":"1998","journal-title":"The Visual Computer"},{"key":"e_1_2_9_21_2","unstructured":"WangX. ZhouB. CurlessB. Kemelmacher-ShlizermanI. HolynskiA. SeitzS. M.: Generative inbetweening: Adapting image-to-video models for keyframe interpolation.Arxiv(2025). arXiv:2408.15239. 3 5 9"},{"key":"e_1_2_9_22_2","unstructured":"YangZ. ZhangJ. YuY. LuS. BaiS.: Versatile transition generation with image-to-video diffusion.Arxiv(2025). 3"},{"key":"e_1_2_9_23_2","unstructured":"ZhangR. ChenY. LiuY. WangW. WenX. WangH.: Tvg: A training-free transition video generation method with diffusion models.Arxiv(2024). 1 2 3 5"},{"key":"e_1_2_9_24_2","doi-asserted-by":"crossref","unstructured":"ZhangZ. ChenH. ZhaoH. LuG. FuY. XuH. WuZ.: EDEN: Enhanced diffusion for high-quality large-motion video frame interpolation. InProc. Computer Vision and Pattern Recognition (CVPR)(2025). 2 3","DOI":"10.1109\/CVPR52734.2025.00202"},{"key":"e_1_2_9_25_2","doi-asserted-by":"crossref","unstructured":"ZhuT. RenD. WangQ. WuX. ZuoW.: Generative inbetweening through frame-wise conditions-driven video generation.Proc. Computer Vision and Pattern Recognition (CVPR)(2025). 1 2 5","DOI":"10.1109\/CVPR52734.2025.02604"},{"key":"e_1_2_9_26_2","doi-asserted-by":"crossref","unstructured":"ZhangK. ZhouY. XuX. DaiB. PanX.: Diffmorpher: Unleashing the capability of diffusion models for image morphing.Proc. Computer Vision and Pattern Recognition (CVPR)(2024) 7912\u20137921. 1 2 3 5","DOI":"10.1109\/CVPR52733.2024.00756"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70420","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70420","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70420","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,30]],"date-time":"2026-04-30T15:09:14Z","timestamp":1777561754000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70420"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,30]]},"references-count":25,"alternative-id":["10.1111\/cgf.70420"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70420","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,30]]},"assertion":[{"value":"2026-04-30","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70420"}}