{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T11:27:39Z","timestamp":1776166059673,"version":"3.50.1"},"reference-count":60,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T00:00:00Z","timestamp":1776124800000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T00:00:00Z","timestamp":1776124800000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>Video inbetweening creates smooth transitions between two frames making it an indispensable tool for video editing and longform video synthesis. Existing methods struggle with large or complex motion and offer limited control over intermediate frames, often misaligning with user intent. We introduce MultiCOIN, a video inbetweening framework supporting multi\u2010modal controls, including depth transitions and layering, motion trajectories, text prompts, and target regions for movement localization. It balances flexibility, usability, and fine\u2010grained precision. Built on a Diffusion Transformer (DiT), due to its proven capability to generate high\u2010quality long video, our model maps all motion controls into a unified sparse point\u2010based representation compatible with the denoising process. Further, to respect the variety of controls which operate at varying levels of granularity and influence, we separate content and motion into two branches, enabling dedicated generators for each. A stage\u2010wise training strategy ensures stable learning of multi\u2010modal controls. Extensive experiments show improved motion complexity, controllability, and narrative consistency. Project Page: MultiCOIN.<\/jats:p>","DOI":"10.1111\/cgf.70362","type":"journal-article","created":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T10:06:26Z","timestamp":1776161186000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["MultiCOIN: Multi\u2010Modal COntrollable INbetweening"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4310-8850","authenticated-orcid":false,"given":"M.","family":"Tanveer","sequence":"first","affiliation":[{"name":"Simon Fraser University"},{"name":"Adobe Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5070-6330","authenticated-orcid":false,"given":"Y.","family":"Zhou","sequence":"additional","affiliation":[{"name":"Adobe Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"S.","family":"Niklaus","sequence":"additional","affiliation":[{"name":"Adobe Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-4693-3565","authenticated-orcid":false,"given":"A. Mahdavi","family":"Amiri","sequence":"additional","affiliation":[{"name":"Simon Fraser University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-1991-119X","authenticated-orcid":false,"given":"H.","family":"Zhang","sequence":"additional","affiliation":[{"name":"Simon Fraser University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8066-6835","authenticated-orcid":false,"given":"K. K.","family":"Singh","sequence":"additional","affiliation":[{"name":"Adobe Research"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"N.","family":"Zhao","sequence":"additional","affiliation":[{"name":"Adobe Research"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2026,4,14]]},"reference":[{"key":"e_1_2_7_2_2","doi-asserted-by":"crossref","unstructured":"AhmadyanA. ZhangL. AblavatskiA. WeiJ. GrundmannM.: Objectron: A large scale dataset of object-centric videos in the wild with pose annotations. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2021) pp.7822\u20137831.","DOI":"10.1109\/CVPR46437.2021.00773"},{"key":"e_1_2_7_3_2","unstructured":"BlattmannA. DockhornT. KulalS. MendelevitchD. KilianM. LorenzD. LeviY. EnglishZ. VoletiV. LettsA. et al.: Stable video diffusion: Scaling latent video diffusion models to large datasets.arXiv preprint arXiv:2311.15127(2023)."},{"key":"e_1_2_7_4_2","doi-asserted-by":"crossref","unstructured":"BriedisK. M. DjelouahA. OrtizR. GrossM. SchroersC.: Controllable tracking-based video frame interpolation. InProceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers(2025) pp.1\u201311.","DOI":"10.1145\/3721238.3730598"},{"key":"e_1_2_7_5_2","doi-asserted-by":"crossref","unstructured":"BlattmannA. RombachR. LingH. DockhornT. KimS. W. FidlerS. KreisK.: Align your latents: High-resolution video synthesis with latent diffusion models. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2023) pp.22563\u201322575.","DOI":"10.1109\/CVPR52729.2023.02161"},{"issue":"3","key":"e_1_2_7_6_2","doi-asserted-by":"crossref","first-page":"603","DOI":"10.1109\/30.883418","article-title":"New frame rate up-conversion using bi-directional motion estimation","volume":"46","author":"Choi B.","year":"2000","journal-title":"IEEE Trans. Consumer Electron."},{"key":"e_1_2_7_7_2","unstructured":"CheferH. SingerU. ZoharA. KirstainY. PolyakA. TaigmanY. WolfL. SheyninS.: Videojam: Joint appearance-motion representations for enhanced motion generation in video models.arXiv preprint arXiv:2502.02492(2025)."},{"key":"e_1_2_7_8_2","unstructured":"ChenX. WangY. ZhangL. ZhuangS. MaX. YuJ. WangY. LinD. QiaoY. LiuZ.: Seine: Short-to-long video diffusion model for generative transition and prediction. InThe Twelfth International Conference on Learning Representations(2023)."},{"key":"e_1_2_7_9_2","doi-asserted-by":"crossref","first-page":"1472","DOI":"10.1609\/aaai.v38i2.27912","volume":"38","author":"Danier D.","year":"2024","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_2_7_10_2","unstructured":"FengH. DingZ. XiaZ. NiklausS. AbrevayaV. BlackM. J. ZhangX.: Explorative inbetweening of time and space.arXiv preprint arXiv:2403.14611(2024)."},{"key":"e_1_2_7_11_2","doi-asserted-by":"crossref","unstructured":"GengD. HerrmannC. HurJ. ColeF. ZhangS. PfaffT. Lopez-GuevaraT. AytarY. RubinsteinM. SunC. et al.: Motion prompting: Controlling video generation with motion trajectories. InProceedings of the Computer Vision and Pattern Recognition Conference(2025) pp.1\u201312.","DOI":"10.1109\/CVPR52734.2025.00010"},{"key":"e_1_2_7_12_2","doi-asserted-by":"crossref","unstructured":"GirdharR. SinghM. BrownA. DuvalQ. AzadiS. RambhatlaS. S. ShahA. YinX. ParikhD. MisraI.: Emu video: Factorizing text-to-video generation by explicit image conditioning.arXiv preprint arXiv:2311.10709(2023).","DOI":"10.1007\/978-3-031-73033-7_12"},{"key":"e_1_2_7_13_2","unstructured":"HoJ. ChanW. SahariaC. WhangJ. GaoR. GritsenkoA. KingmaD. P. PooleB. NorouziM. FleetD. J. et al.: Imagen video: High definition video generation with diffusion models.arXiv preprint arXiv:2210.02303(2022)."},{"key":"e_1_2_7_14_2","unstructured":"HongW. DingM. ZhengW. LiuX. TangJ.:Cogvideo: Large-scale pretraining for text-to-video generation via transformers 2022. URL:https:\/\/arxiv.org\/abs\/2205.15868 arXiv:2205.15868."},{"issue":"2","key":"e_1_2_7_15_2","doi-asserted-by":"crossref","first-page":"752","DOI":"10.1109\/TCE.2004.1309458","article-title":"Motion compensated frame interpolation by new block-based motion estimation algorithm","volume":"50","author":"Ha T.","year":"2004","journal-title":"IEEE Transactions on Consumer Electronics"},{"issue":"2","key":"e_1_2_7_16_2","doi-asserted-by":"crossref","first-page":"752","DOI":"10.1109\/TCE.2004.1309458","article-title":"Motion compensated frame interpolation by new block-based motion estimation algorithm","volume":"50","author":"Ha T.","year":"2004","journal-title":"IEEE Trans. Consumer Electron."},{"key":"e_1_2_7_17_2","first-page":"8633","article-title":"Video diffusion models","volume":"35","author":"Ho J.","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_7_18_2","unstructured":"HongS. SeoJ. ShinH. HongS. KimS.: Large language models are frame-level directors for zero-shot text-to-video generation. InFirst Workshop on Controllable Video Generation@ ICML24(2023)."},{"key":"e_1_2_7_19_2","unstructured":"HuangZ. ZhangT. HengW. ShiB. ZhouS.: Rife: real-time intermediate flow estimation for video frame interpolation. arxiv preprint arxiv.2011:06294. InRife: Real-time intermediate flow estimation for video frame interpolation. arXiv preprint arXiv: 2011.06294.2020."},{"key":"e_1_2_7_20_2","unstructured":"JangS. KiT. JoJ. YoonJ. KimS. Y. LinZ. HwangS. J.: Frame guidance: Training-free guidance for frame-level control in video diffusion models.arXiv preprint arXiv:2506.07177(2025)."},{"key":"e_1_2_7_21_2","doi-asserted-by":"crossref","unstructured":"JiangH. SunD. JampaniV. YangM.-H. Learned-MillerE. KautzJ.: Super slomo: High quality estimation of multiple intermediate frames for video interpolation. InProceedings of the IEEE conference on computer vision and pattern recognition(2018) pp.9000\u20139008.","DOI":"10.1109\/CVPR.2018.00938"},{"key":"e_1_2_7_22_2","unstructured":"KwonM. OhS. W. ZhouY. LiuD. LeeJ.-Y. CaiH. LiuB. LiuF. UhY.: Harivo: Harnessing text-to-image models for video generation.arXiv preprint arXiv:2410.07763(2024)."},{"key":"e_1_2_7_23_2","unstructured":"LiX. ChuW. WuY. YuanW. LiuF. ZhangQ. LiF. FengH. DingE. WangJ.: Videogen: A reference-guided latent diffusion approach for high definition text-to-video generation.arXiv preprint arXiv:2309.00398(2023)."},{"key":"e_1_2_7_24_2","doi-asserted-by":"crossref","first-page":"5031","DOI":"10.1609\/aaai.v39i5.32533","volume":"39","author":"Li Y.","year":"2025","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_2_7_25_2","unstructured":"MeyerS. Cornill\u00e8reV. DjelouahA. SchroersC. GrossM. H.: Deep video color propagation. InBMVC(2018) BMVA Press p. 128."},{"key":"e_1_2_7_26_2","first-page":"1410","volume-title":"CVPR","author":"Meyer S.","year":"2015"},{"key":"e_1_2_7_27_2","first-page":"713","volume-title":"WACV","author":"Niklaus S.","year":"2023"},{"key":"e_1_2_7_28_2","doi-asserted-by":"crossref","unstructured":"NiklausS. LiuF.: Context-aware synthesis for video frame interpolation. InCVPR(2018) Computer Vision Foundation \/ IEEE Computer Society pp.1701\u20131710.","DOI":"10.1109\/CVPR.2018.00183"},{"key":"e_1_2_7_29_2","doi-asserted-by":"crossref","unstructured":"NiklausS. LiuF.: Softmax splatting for video frame interpolation. InCVPR(2020) Computer Vision Foundation \/ IEEE pp.5436\u20135445.","DOI":"10.1109\/CVPR42600.2020.00548"},{"key":"e_1_2_7_30_2","doi-asserted-by":"crossref","unstructured":"NiklausS. MaiL. LiuF.: Video frame interpolation via adaptive convolution. InProceedings of the IEEE conference on computer vision and pattern recognition(2017) pp.670\u2013679.","DOI":"10.1109\/CVPR.2017.244"},{"key":"e_1_2_7_31_2","first-page":"261","volume-title":"ICCV","author":"Niklaus S.","year":"2017"},{"key":"e_1_2_7_32_2","doi-asserted-by":"crossref","first-page":"109","DOI":"10.1007\/978-3-030-58568-6_7","volume-title":"Computer Vision\u2013ECCV 2020: 16th European Conference, Glasgow, UK, August 23\u201328, 2020, Proceedings, Part XIV","author":"Park J.","year":"2020"},{"key":"e_1_2_7_33_2","unstructured":"Pont-TusetJ. PerazziF. CaellesS. Arbel\u00e1ezP. Sorkine-HornungA. Van GoolL.: The 2017 davis challenge on video object segmentation.arXiv preprint arXiv:1704.00675(2017)."},{"key":"e_1_2_7_34_2","doi-asserted-by":"crossref","unstructured":"PeeblesW. XieS.: Scalable diffusion models with transformers. InProceedings of the IEEE\/CVF International Conference on Computer Vision(2023) pp.4195\u20134205.","DOI":"10.1109\/ICCV51070.2023.00387"},{"key":"e_1_2_7_35_2","first-page":"1","volume-title":"2008 IEEE conference on computer vision and pattern recognition","author":"Rodriguez M. D.","year":"2008"},{"key":"e_1_2_7_36_2","doi-asserted-by":"crossref","unstructured":"RombachR. BlattmannA. LorenzD. EsserP. OmmerB.: High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2022) pp.10684\u201310695.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_7_37_2","first-page":"250","volume-title":"European Conference on Computer Vision","author":"Reda F.","year":"2022"},{"key":"e_1_2_7_38_2","unstructured":"RanzatoM. SzlamA. BrunaJ. MathieuM. CollobertR. ChopraS.:Video (language) modeling: a baseline for generative models of natural videos 2016. URL:https:\/\/arxiv.org\/abs\/1412.6604 arXiv:1412.6604."},{"key":"e_1_2_7_39_2","unstructured":"ShenX. LiX. ElhoseinyM.:Mostgan-v: Video generation with temporal motion styles 2023. URL:https:\/\/arxiv.org\/abs\/2304.02777 arXiv:2304.02777."},{"key":"e_1_2_7_40_2","unstructured":"SrivastavaN. MansimovE. SalakhutdinovR.:Unsupervised learning of video representations using lstms 2016. URL:https:\/\/arxiv.org\/abs\/1502.04681 arXiv:1502.04681."},{"key":"e_1_2_7_41_2","doi-asserted-by":"crossref","unstructured":"SaitoM. MatsumotoE. SaitoS.:Temporal generative adversarial nets with singular value clipping 2017. URL:https:\/\/arxiv.org\/abs\/1611.06624 arXiv:1611.06624.","DOI":"10.1109\/ICCV.2017.308"},{"key":"e_1_2_7_42_2","unstructured":"SingerU. PolyakA. HayesT. YinX. AnJ. ZhangS. HuQ. YangH. AshualO. GafniO. et al.: Make-avideo: Text-to-video generation without text-video data.arXiv preprint arXiv:2209.14792(2022)."},{"key":"e_1_2_7_43_2","unstructured":"SiyaoL. ZhaoS. YuW. SunW. MetaxasD. N. LoyC. C. LiuZ.: Deep animation video interpolation in the wild. InCVPR(2021) Computer Vision Foundation \/ IEEE pp.6587\u20136595."},{"key":"e_1_2_7_44_2","unstructured":"TulyakovS. LiuM.-Y. YangX. KautzJ.:Mocogan: Decomposing motion and content for video generation 2017. URL:https:\/\/arxiv.org\/abs\/1707.04993 arXiv:1707.04993."},{"key":"e_1_2_7_45_2","unstructured":"UnterthinerT. vanSteenkisteS. KurachK. MarinierR. MichalskiM. GellyS.: Fvd: A new metric for video generation. InICLR 2019 Workshop DeepGenStruct(2019)."},{"key":"e_1_2_7_46_2","doi-asserted-by":"crossref","unstructured":"WangY. BaoJ. WengW. FengR. YinD. YangT. ZhangJ. DaiQ. ZhaoZ. WangC. et al.: Microcinema: A divide-and-conquer approach for text-to-video generation. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024) pp.8414\u20138424.","DOI":"10.1109\/CVPR52733.2024.00804"},{"key":"e_1_2_7_47_2","doi-asserted-by":"crossref","unstructured":"WuJ. Z. GeY. WangX. LeiS. W. GuY. ShiY. HsuW. ShanY. QieX. ShouM. Z.: Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation. InProceedings of the IEEE\/CVF International Conference on Computer Vision(2023) pp.7623\u20137633.","DOI":"10.1109\/ICCV51070.2023.00701"},{"key":"e_1_2_7_48_2","unstructured":"WangW. WangQ. ZhengK. OuyangH. ChenZ. GongB. ChenH. ShenY. ShenC.: Framer: Interactive frame interpolation.arXiv preprint arXiv:2410.18978(2024)."},{"key":"e_1_2_7_49_2","doi-asserted-by":"crossref","unstructured":"WangZ. YuanZ. WangX. LiY. ChenT. XiaM. LuoP. ShanY.: Motionctrl: A unified and flexible motion controller for video generation. InACM SIGGRAPH 2024 Conference Papers(2024) pp.1\u201311.","DOI":"10.1145\/3641519.3657518"},{"issue":"10","key":"e_1_2_7_50_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/s11432-024-4592-3","article-title":"Unianimate: Taming unified video diffusion models for consistent human image animation","volume":"68","author":"Wang X.","year":"2025","journal-title":"Science China Information Sciences"},{"key":"e_1_2_7_51_2","unstructured":"XingJ. LiuH. XiaM. ZhangY. WangX. ShanY. WongT.-T.: Tooncrafter: Generative cartoon interpolation.arXiv preprint arXiv:2405.17933(2024)."},{"key":"e_1_2_7_52_2","first-page":"399","volume-title":"European Conference on Computer Vision","author":"Xing J.","year":"2025"},{"key":"e_1_2_7_53_2","unstructured":"YuL. ChengY. SohnK. LezamaJ. ZhangH. ChangH. HauptmannA. G. YangM.-H. HaoY. EssaI. JiangL.:Magvit: Masked generative video transformer 2023. URL:https:\/\/arxiv.org\/abs\/2212.05199 arXiv:2212.05199."},{"key":"e_1_2_7_54_2","unstructured":"YinS. WuC. LiangJ. ShiJ. LiH. MingG. DuanN.: Dragnuwa: Fine-grained control in video generation by integrating text image and trajectory.arXiv preprint arXiv:2308.08089(2023)."},{"key":"e_1_2_7_55_2","unstructured":"YanW. ZhangY. AbbeelP. SrinivasA.:Videogpt: Video generation using vq-vae and transformers 2021. URL:https:\/\/arxiv.org\/abs\/2104.10157 arXiv:2104.10157."},{"key":"e_1_2_7_56_2","doi-asserted-by":"crossref","unstructured":"ZhangR. IsolaP. EfrosA. A. ShechtmanE. WangO.: The unreasonable effectiveness of deep features as a perceptual metric. InProceedings of the IEEE conference on computer vision and pattern recognition(2018) pp.586\u2013595.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_7_57_2","unstructured":"ZhangZ. LiaoJ. LiM. QinL. WangW.: Tora: Trajectory-oriented diffusion transformer for video generation.arXiv preprint arXiv:2407.21705(2024)."},{"key":"e_1_2_7_58_2","doi-asserted-by":"crossref","first-page":"10743","DOI":"10.1609\/aaai.v39i10.33167","volume":"39","author":"Zhou H.","year":"2025","journal-title":"Proceedings of the AAAI Conference on Artificial Intelligence"},{"key":"e_1_2_7_59_2","doi-asserted-by":"crossref","unstructured":"ZengY. WeiG. ZhengJ. ZouJ. WeiY. ZhangY. LiH.: Make pixels dance: High-dynamic video generation. InProceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2024) pp.8850\u20138860.","DOI":"10.1109\/CVPR52733.2024.00845"},{"key":"e_1_2_7_60_2","doi-asserted-by":"crossref","unstructured":"ZhouY. YangJ. LiD. SaitoJ. AnejaD. KalogerakisE.: Audio-driven neural gesture reenactment with video motion graphs. InProceedings of the IEEE\/CVF conference on computer vision and pattern recognition(2022) pp.3418\u20133428.","DOI":"10.1109\/CVPR52688.2022.00341"},{"key":"e_1_2_7_61_2","unstructured":"ZhangY. YuanY. SongY. WangH. LiuJ.: Easy-control: Adding efficient and flexible control for diffusion transformer. InProceedings of the IEEE\/CVF International Conference on Computer Vision(2025) pp.19513\u201319524."}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70362","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70362","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70362","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,14]],"date-time":"2026-04-14T10:06:57Z","timestamp":1776161217000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70362"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,14]]},"references-count":60,"alternative-id":["10.1111\/cgf.70362"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70362","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,14]]},"assertion":[{"value":"2026-04-14","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70362"}}