{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,8,18]],"date-time":"2026-08-18T00:46:57Z","timestamp":1787014017075,"version":"3.56.0"},"reference-count":46,"publisher":"Wiley","license":[{"start":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T00:00:00Z","timestamp":1775606400000},"content-version":"vor","delay-in-days":0,"URL":"http:\/\/creativecommons.org\/licenses\/by\/4.0\/"},{"start":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T00:00:00Z","timestamp":1775606400000},"content-version":"tdm","delay-in-days":0,"URL":"http:\/\/doi.wiley.com\/10.1002\/tdm_license_1.1"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Computer Graphics Forum"],"abstract":"<jats:title>Abstract<\/jats:title>\n                  <jats:p>\n                    We present Story2Board, a training\u2010free framework for expressive storyboard generation from natural language. Existing methods narrowly focus on subject identity, overlooking key aspects of visual storytelling such as spatial composition, background evolution, and narrative pacing. To address this, we introduce a lightweight consistency framework composed of two components: Latent Panel Anchoring, which preserves a shared character reference across panels, and Reciprocal Attention Value Mixing, which softly blends visual features between token pairs with strong reciprocal attention. Together, these mechanisms enhance coherence without architectural changes or fine\u2010tuning, enabling state\u2010of\u2010the\u2010art diffusion models to generate visually diverse yet consistent storyboards. To structure generation, we use an off\u2010the\u2010shelf language model to convert free\u2010form stories into grounded panel\u2010level prompts. To evaluate, we propose the\n                    <jats:italic>Rich Storyboard Benchmark<\/jats:italic>\n                    , a suite of open\u2010domain narratives designed to assess layout diversity and background\u2010grounded storytelling, in addition to consistency. We also introduce a new Scene Diversity metric that quantifies spatial and pose variation across storyboards. Our qualitative and quantitative results, as well as a user study, show that Story2Board produces more dynamic, coherent, and narratively engaging storyboards than existing baselines.\n                    <jats:italic>Project page<\/jats:italic>\n                    :\n                    <jats:ext-link xmlns:xlink=\"http:\/\/www.w3.org\/1999\/xlink\" xlink:href=\"https:\/\/daviddinkevich.github.io\/Story2Board\/\">https:\/\/daviddinkevich.github.io\/Story2Board\/<\/jats:ext-link>\n                  <\/jats:p>","DOI":"10.1111\/cgf.70319","type":"journal-article","created":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T10:56:36Z","timestamp":1775645796000},"update-policy":"https:\/\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["Story2Board: A Training\u2010Free Approach for Expressive Visual Storytelling"],"prefix":"10.1111","author":[{"ORCID":"https:\/\/orcid.org\/0009-0000-0543-1035","authenticated-orcid":false,"given":"D.","family":"Dinkevich","sequence":"first","affiliation":[{"name":"Hebrew University of Jerusalem  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0009-0008-2915-1618","authenticated-orcid":false,"given":"M.","family":"Levy","sequence":"additional","affiliation":[{"name":"Hebrew University of Jerusalem  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-7628-7525","authenticated-orcid":false,"given":"O.","family":"Avrahami","sequence":"additional","affiliation":[{"name":"Hebrew University of Jerusalem  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-3573-2220","authenticated-orcid":false,"given":"D.","family":"Samuel","sequence":"additional","affiliation":[{"name":"OriginAI  Israel"},{"name":"Bar\u2010Ilan University  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6191-0361","authenticated-orcid":false,"given":"D.","family":"Lischinski","sequence":"additional","affiliation":[{"name":"Hebrew University of Jerusalem  Israel"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2026,4,8]]},"reference":[{"key":"e_1_2_6_2_2","doi-asserted-by":"publisher","DOI":"10.1145\/3680528.3687590"},{"key":"e_1_2_6_2_3","unstructured":"url:https:\/\/doi.org\/10.1145\/3680528.36875904."},{"key":"e_1_2_6_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3641519.3657430"},{"key":"e_1_2_6_3_3","unstructured":"url:https:\/\/doi.org\/10.1145\/3641519.36574303."},{"key":"e_1_2_6_4_2","unstructured":"Amazon Mechanical Turk.Amazon Mechanical Turk.https:\/\/www.mturk.com\/. Accessed: 2025-05-20.20258 20."},{"key":"e_1_2_6_5_2","unstructured":"Animator Island.Composition: What is Breathing Room?https:\/\/www.animatorisland.com\/composition-what-is-breathing-room\/. Accessed: 2025-05-12.20142."},{"key":"e_1_2_6_6_2","unstructured":"Black Forest Labs.FLUX.https:\/\/github.com\/black-forest-labs\/flux.20242 3 7 17."},{"key":"e_1_2_6_7_2","doi-asserted-by":"crossref","DOI":"10.4324\/9781315794839","volume-title":"The Visual Story: Creating the Visual Structure of Film, TV, and Digital Media","author":"Block Bruce","year":"2020"},{"key":"e_1_2_6_8_2","first-page":"432","volume-title":"European Conference on Computer Vision","author":"Dahary Omer","year":"2024"},{"key":"e_1_2_6_9_2","article-title":"Scaling Rectified Flow Transformers for High-Resolution Image Synthesis","volume":"2403","author":"Esser Patrick","year":"2024","journal-title":"ArXiv"},{"key":"e_1_2_6_10_2","doi-asserted-by":"crossref","unstructured":"Feng Haoran Huang Zehuan Li Lin et al. \u201cPersonalize anything for free with diffusion transformer\u201d.arXiv preprint arXiv:2503.12590(2025) 3 5.","DOI":"10.1609\/aaai.v40i5.37394"},{"key":"e_1_2_6_11_2","unstructured":"Filmmakers Academy.Negative Space: Film Composition Guide.https:\/\/www.filmmakersacademy.com\/blog-negative-space-film\/. Accessed: 2025-05-12.20252."},{"key":"e_1_2_6_12_2","article-title":"DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic Data","volume":"2306","author":"Fu Stephanie","year":"2023","journal-title":"ArXiv"},{"key":"e_1_2_6_13_2","unstructured":"Geyer Michal Bar-Tal Omer Bagon Shai andDekel Tali. \u201cTokenflow: Consistent diffusion features for consistent video editing\u201d.arXiv preprint arXiv:2307.10373(2023) 4."},{"key":"e_1_2_6_14_2","unstructured":"Ho Jonathan Jain Ajay andAbbeel Pieter. \u201cDenoising Diffusion Probabilistic Models\u201d.Proc. NeurIPS.20202 3."},{"key":"e_1_2_6_15_2","unstructured":"He Junjie Tuo Yuxiang Chen Binghui et al. \u201cAnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation\u201d.arXiv preprint arXiv:2501.09503(2025) 2 3 7."},{"key":"e_1_2_6_16_2","article-title":"In-Context LoRA for Diffusion Transformers","volume":"2410","author":"Huang Lianghua","year":"2024","journal-title":"ArXiv"},{"key":"e_1_2_6_17_2","unstructured":"He Huiguo Yang Huan Tuo Zixi et al. \u201cDreamstory: Open-domain story visualization by llm-guided multi-subject consistent diffusion\u201d.arXiv preprint arXiv:2407.12899(2024) 2 3 6 7 9 18."},{"key":"e_1_2_6_18_2","unstructured":"Kirillov Alexander Mintun Eric Ravi Nikhila et al.Segment Anything.2023. arXiv: 2304.02643 [cs.CV] 3."},{"key":"e_1_2_6_19_2","doi-asserted-by":"crossref","unstructured":"Kumari Nupur Zhang Bingliang Zhang Richard et al. \u201cMulti-concept customization of text-to-image diffusion\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2023 1931\u201319414.","DOI":"10.1109\/CVPR52729.2023.00192"},{"key":"e_1_2_6_20_2","unstructured":"Labs Black Forest Batifol Stephen Blattmann Andreas et al.FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.2025. arXiv: 2506.15742 [cs.GR]. url:https:\/\/arxiv.org\/abs\/2506.157427."},{"key":"e_1_2_6_21_2","unstructured":"Lin Zhiqiu Pathak Deepak Li Baiqi et al. \u201cEvaluating Text-to-Visual Generation with Image-to-Text Generation\u201d.European Conference on Computer Vision.20247 21."},{"key":"e_1_2_6_22_2","unstructured":"Liu Chang Wu Haoning Zhong Yujie et al.Intelligent Grimm \u2013 Open-ended Visual Storytelling via Latent Diffusion Models.2024. arXiv: 2306.00973 [cs.CV]. url:https:\/\/arxiv.org\/abs\/2306.009732 3 7."},{"key":"e_1_2_6_23_2","doi-asserted-by":"crossref","unstructured":"Liu Shaoteng Zhang Yuechen Li Wenbo et al. \u201cVideo-p2p: Video editing with cross-attention control\u201d.arXiv preprint arXiv:2303.04761(2023) 2.","DOI":"10.1109\/CVPR52733.2024.00821"},{"key":"e_1_2_6_24_2","first-page":"38","volume-title":"European conference on computer vision","author":"Liu Shilong","year":"2024"},{"key":"e_1_2_6_25_2","unstructured":"OpenAI Achiam Josh Adler Steven et al.GPT-4 Technical Report.2024. arXiv: 2303.08774 [cs.CL]. url:https:\/\/arxiv.org\/abs\/2303.087744 14."},{"key":"e_1_2_6_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TSMC.1979.4310076"},{"key":"e_1_2_6_27_2","article-title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","volume":"2307","author":"Podell Dustin","year":"2023","journal-title":"ArXiv"},{"key":"e_1_2_6_28_2","doi-asserted-by":"crossref","unstructured":"Rombach Robin Blattmann A. Lorenz Dominik et al. \u201cHigh-Resolution Image Synthesis with Latent Diffusion Models\u201d.2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)(2021) 10674\u2013106852 3.","DOI":"10.1109\/CVPR52688.2022.01042"},{"key":"e_1_2_6_29_2","unstructured":"Ramesh Aditya Dhariwal Prafulla Nichol Alex et al. \u201cHierarchical text-conditional image generation with CLIP latents\u201d.arXiv preprint arXiv:2204.06125(2022) 2 3."},{"key":"e_1_2_6_30_2","first-page":"8748","volume-title":"International conference on machine learning","author":"Radford Alec","year":"2021"},{"key":"e_1_2_6_31_2","first-page":"36479","article-title":"Photorealistic text-to-image diffusion models with deep language understanding","volume":"35","author":"Saharia Chitwan","year":"2022","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_6_32_2","unstructured":"Stability AI.Stable Diffusion 3: Next-Generation Text-to-Image Generation. urlhttps:\/\/stability.ai\/news\/stable-diffusion-3.20242."},{"key":"e_1_2_6_33_2","doi-asserted-by":"crossref","unstructured":"Tumanyan Narek Geyer Michal Bagon Shai andDekel Tali. \u201cPlug-and-play diffusion features for text-driven image-to-image translation\u201d.Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition.2023 1921\u201319304.","DOI":"10.1109\/CVPR52729.2023.00191"},{"key":"e_1_2_6_34_2","doi-asserted-by":"crossref","unstructured":"Tewel Yoad Gal Rinon Chechik Gal andAtzmon Yuval. \u201cKey-Locked Rank One Editing for Text-to-Image Personalization\u201d.ACM SIGGRAPH 2023 Conference Proceedings. SIGGRAPH '23. Los Angeles CA USA 20234.","DOI":"10.1145\/3588432.3591506"},{"key":"e_1_2_6_35_2","article-title":"Training-Free Consistent Text-to-Image Generation","volume":"2402","author":"Tewel Yoad","year":"2024","journal-title":"ArXiv"},{"issue":"4","key":"e_1_2_6_36_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3658157","article-title":"Training-free consistent text-to-image generation","volume":"43","author":"Tewel Yoad","year":"2024","journal-title":"ACM Transactions on Graphics (TOG)"},{"key":"e_1_2_6_37_2","unstructured":"Tan Zhenxiong Liu Songhua Yang Xingyi et al. \u201cOminiControl: Minimal and Universal Control for Diffusion Transformer\u201d.arXiv preprint arXiv:2411.15098(2024) 2 3 7."},{"key":"e_1_2_6_38_2","article-title":"Attention is all you need","volume":"30","author":"Vaswani Ashish","year":"2017","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_6_39_2","article-title":"ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation","volume":"2302","author":"Wei Yuxiang","year":"2023","journal-title":"ArXiv"},{"key":"e_1_2_6_40_2","first-page":"38571","article-title":"Vitpose: Simple vision transformer baselines for human pose estimation","volume":"35","author":"Xu Yufei","year":"2022","journal-title":"Advances in neural information processing systems"},{"key":"e_1_2_6_41_2","doi-asserted-by":"crossref","unstructured":"Yang Shuai Ge Yuying Li Yang et al. \u201cSeed-story: Multimodal long story generation with large language model\u201d.arXiv preprint arXiv:2407.08683(2024) 2.","DOI":"10.1109\/ICCVW69036.2025.00197"},{"key":"e_1_2_6_42_2","article-title":"IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models","volume":"2308","author":"Ye Hu","year":"2023","journal-title":"arXiv"},{"key":"e_1_2_6_43_2","doi-asserted-by":"crossref","unstructured":"Zhang Richard Isola Phillip Efros Alexei A. et al. \u201cThe Unreasonable Effectiveness of Deep Features as a Perceptual Metric\u201d.2018 IEEE\/CVF Conference on Computer Vision and Pattern Recognition(2018) 586\u201359521.","DOI":"10.1109\/CVPR.2018.00068"},{"key":"e_1_2_6_44_2","unstructured":"Zhang Kai Mo Lingbo Chen Wenhu et al. \u201cMagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing\u201d.Advances in Neural Information Processing Systems.20232."},{"key":"e_1_2_6_45_2","article-title":"StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation","volume":"2405","author":"Zhou Yupeng","year":"2024","journal-title":"ArXiv"}],"container-title":["Computer Graphics Forum"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70319","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1111\/cgf.70319","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/cgf.70319","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,8]],"date-time":"2026-04-08T10:56:48Z","timestamp":1775645808000},"score":1,"resource":{"primary":{"URL":"https:\/\/onlinelibrary.wiley.com\/doi\/10.1111\/cgf.70319"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,4,8]]},"references-count":46,"alternative-id":["10.1111\/cgf.70319"],"URL":"https:\/\/doi.org\/10.1111\/cgf.70319","archive":["Portico"],"relation":{},"ISSN":["0167-7055","1467-8659"],"issn-type":[{"value":"0167-7055","type":"print"},{"value":"1467-8659","type":"electronic"}],"subject":[],"published":{"date-parts":[[2026,4,8]]},"assertion":[{"value":"2026-04-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e70319"}}