{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,28]],"date-time":"2026-06-28T08:00:39Z","timestamp":1782633639204,"version":"3.54.5"},"reference-count":48,"publisher":"Association for Computing Machinery (ACM)","issue":"6","license":[{"start":{"date-parts":[[2019,11,8]],"date-time":"2019-11-08T00:00:00Z","timestamp":1573171200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001742","name":"Israel Science Foundation","doi-asserted-by":"publisher","award":["2216\/15"],"award-info":[{"award-number":["2216\/15"]}],"id":[{"id":"10.13039\/501100001742","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["61832016, 61902012, 61521002"],"award-info":[{"award-number":["61832016, 61902012, 61521002"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Graph."],"published-print":{"date-parts":[[2019,12,31]]},"abstract":"<jats:p>\n            We present\n            <jats:italic>Write-A-Video<\/jats:italic>\n            , a tool for the creation of video montage using mostly text-editing. Given an input themed text and a related video repository either from online websites or personal albums, the tool allows novice users to generate a video montage much more easily than current video editing tools. The resulting video illustrates the given narrative, provides diverse visual content, and follows cinematographic guidelines. The process involves three simple steps: (1) the user provides input, mostly in the form of editing the text, (2) the tool automatically searches for semantically matching candidate shots from the video repository, and (3) an optimization method assembles the video montage. Visual-semantic matching between segmented text and shots is performed by cascaded keyword matching and visual-semantic embedding, that have better accuracy than alternative solutions. The video assembly is formulated as a hybrid optimization problem over a graph of shots, considering temporal constraints, cinematography metrics such as camera movement and tone, and user-specified cinematography idioms. Using our system, users without video editing experience are able to generate appealing videos.\n          <\/jats:p>","DOI":"10.1145\/3355089.3356520","type":"journal-article","created":{"date-parts":[[2019,11,8]],"date-time":"2019-11-08T20:27:58Z","timestamp":1573244878000},"page":"1-13","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":60,"title":["Write-a-video"],"prefix":"10.1145","volume":"38","author":[{"given":"Miao","family":"Wang","sequence":"first","affiliation":[{"name":"State Key Lab of Virtual Reality Technology and Systems, Beihang University; Tsinghua University, Beijing"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Guo-Wei","family":"Yang","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shi-Min","family":"Hu","sequence":"additional","affiliation":[{"name":"Tsinghua University, Beijing"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Shing-Tung","family":"Yau","sequence":"additional","affiliation":[{"name":"Harvard University"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ariel","family":"Shamir","sequence":"additional","affiliation":[{"name":"The Interdisciplinary Center, Herzliya"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2019,11,8]]},"reference":[{"key":"e_1_2_2_1_1","doi-asserted-by":"publisher","DOI":"10.1145\/2601097.2601198"},{"key":"e_1_2_2_2_1","volume-title":"SURF: Speeded Up Robust Features. 404--417.","author":"Bay Herbert","year":"2006","unstructured":"Herbert Bay, Tinne Tuytelaars, and Luc Van Gool. 2006. SURF: Speeded Up Robust Features. 404--417."},{"key":"e_1_2_2_3_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-018-0115-y"},{"key":"e_1_2_2_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/2185520.2185563"},{"key":"e_1_2_2_5_1","volume-title":"ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding","author":"Heilbron Fabian Caba","unstructured":"Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. 2015. ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding. In IEEE CVPR."},{"key":"e_1_2_2_6_1","doi-asserted-by":"crossref","unstructured":"Minsuk Chang Anh Truong Oliver Wang Maneesh Agrawala and Juho Kim. 2019. How to Design Voice Based Navigation for How-To Videos. In ACM CHI. Article 701 701:1--701:11 pages.","DOI":"10.1145\/3290605.3300931"},{"key":"e_1_2_2_7_1","unstructured":"Pei-Yu Chi Joyce Liu Jason Linder Mira Dontcheva Wilmot Li and Bjoern Hartmann. 2013. Democut: generating concise instructional videos for physical demonstrations. In ACM UIST. 141--150."},{"key":"e_1_2_2_8_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICASSP.2018.8462105"},{"key":"e_1_2_2_9_1","doi-asserted-by":"crossref","unstructured":"W. S. Chu Yale Song and A. Jaimes. 2015. Video co-summarization: Video summarization by visual co-occurrence. In IEEE CVPR. 3584--3592.","DOI":"10.1109\/CVPR.2015.7298981"},{"key":"e_1_2_2_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201371"},{"key":"e_1_2_2_11_1","volume-title":"On Film Editing: An Introduction to the Art of Film Construction","author":"Dmytryk Edward","unstructured":"Edward Dmytryk. 1984. On Film Editing: An Introduction to the Art of Film Construction. Focal Press."},{"key":"e_1_2_2_12_1","doi-asserted-by":"crossref","unstructured":"David K Elson and Mark O Riedl. 2007. A Lightweight Intelligent Virtual Cinematography System for Machinima Production. (2007).","DOI":"10.21236\/ADA464770"},{"key":"e_1_2_2_13_1","volume-title":"Jamie Ryan Kiros, and Sanja Fidler","author":"Faghri Fartash","year":"2018","unstructured":"Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2018. VSE++: Improving Visual-Semantic Embeddings with Hard Negatives. In BMVC."},{"key":"e_1_2_2_14_1","doi-asserted-by":"publisher","DOI":"10.1145\/358669.358692"},{"key":"e_1_2_2_15_1","doi-asserted-by":"crossref","unstructured":"Quentin Galvane R\u00e9mi Ronfard Christophe Lino and Marc Christie. 2015. Continuity Editing for 3D Animation. In AAAI. 753--762.","DOI":"10.1609\/aaai.v29i1.9288"},{"key":"e_1_2_2_16_1","doi-asserted-by":"crossref","unstructured":"Andreas Girgensohn John Boreczky Patrick Chiu John Doherty Jonathan Foote Gene Golovchinsky Shingo Uchihashi and Lynn Wilcox. 2000. A Semi-automatic Approach to Home Video Editing. In ACM UIST. 81--89.","DOI":"10.1145\/354401.354415"},{"key":"e_1_2_2_17_1","volume-title":"Mask r-cnn","author":"He Kaiming","unstructured":"Kaiming He, Georgia Gkioxari, Piotr Doll\u00e1r, and Ross Girshick. 2017. Mask r-cnn. In IEEE ICCV. 2980--2988."},{"key":"e_1_2_2_18_1","volume-title":"Deep residual learning for image recognition","author":"He Kaiming","unstructured":"Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In IEEE CVPR. 770--778."},{"key":"e_1_2_2_19_1","doi-asserted-by":"publisher","DOI":"10.1145\/1198302.1198306"},{"key":"e_1_2_2_20_1","doi-asserted-by":"publisher","DOI":"10.1109\/TSMCC.2011.2109710"},{"key":"e_1_2_2_21_1","doi-asserted-by":"publisher","DOI":"10.1007\/s41095-016-0074-0"},{"key":"e_1_2_2_22_1","doi-asserted-by":"publisher","DOI":"10.1145\/2699644"},{"key":"e_1_2_2_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766954"},{"key":"e_1_2_2_24_1","volume-title":"Deep visual-semantic alignments for generating image descriptions","author":"Karpathy Andrej","unstructured":"Andrej Karpathy and Li Fei-Fei. 2015. Deep visual-semantic alignments for generating image descriptions. In IEEE CVPR. 3128--3137."},{"key":"e_1_2_2_25_1","doi-asserted-by":"crossref","unstructured":"G. Kim L. Sigal and E. P. Xing. 2014. Joint Summarization of Large-Scale Collections of Web Images and Videos for Storyline Reconstruction. In IEEE CVPR. 4225--4232.","DOI":"10.1109\/CVPR.2014.538"},{"key":"e_1_2_2_26_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073653"},{"key":"e_1_2_2_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/2766966"},{"key":"e_1_2_2_28_1","volume-title":"Microsoft coco: Common objects in context","author":"Lin Tsung-Yi","unstructured":"Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll\u00e1r, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In IEEE ECCV. 740--755."},{"key":"e_1_2_2_29_1","volume-title":"Learning Activity Progression in LSTMs for Activity Detection and Early Detection","author":"Ma Shugao","unstructured":"Shugao Ma, Leonid Sigal, and Stan Sclaroff. 2016. Learning Activity Progression in LSTMs for Activity Detection and Early Detection. In IEEE CVPR."},{"key":"e_1_2_2_30_1","volume-title":"Multimedia Computing and Networking","volume":"4673","author":"Machnicki Erik","year":"2001","unstructured":"Erik Machnicki and Lawrence A Rowe. 2001. Virtual director: Automating a webcast. In Multimedia Computing and Networking 2002, Vol. 4673. 208--226."},{"key":"e_1_2_2_31_1","volume-title":"Learning a text-video embedding from incomplete and heterogeneous data. arXiv preprint arXiv:1804.02516","author":"Miech Antoine","year":"2018","unstructured":"Antoine Miech, Ivan Laptev, and Josef Sivic. 2018. Learning a text-video embedding from incomplete and heterogeneous data. arXiv preprint arXiv:1804.02516 (2018)."},{"key":"e_1_2_2_32_1","doi-asserted-by":"publisher","DOI":"10.1109\/76.981844"},{"key":"e_1_2_2_33_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCSVT.2012.2189689"},{"key":"e_1_2_2_34_1","volume-title":"Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499","author":"van den Oord Aaron","year":"2016","unstructured":"Aaron van den Oord, Sander Dieleman, Heiga Zen, Karen Simonyan, Oriol Vinyals, Alex Graves, Nal Kalchbrenner, Andrew Senior, and Koray Kavukcuoglu. 2016. Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499 (2016)."},{"key":"e_1_2_2_35_1","doi-asserted-by":"publisher","DOI":"10.1145\/2807442.2807502"},{"key":"e_1_2_2_36_1","volume-title":"Skimmable Format for Informational Lecture Videos. In ACM UIST (UIST '14)","author":"Pavel Amy","year":"2014","unstructured":"Amy Pavel, Colorado Reed, Bj\u00f6rn Hartmann, and Maneesh Agrawala. 2014. Video Digests: A Browsable, Skimmable Format for Informational Lecture Videos. In ACM UIST (UIST '14). 573--582."},{"key":"e_1_2_2_37_1","doi-asserted-by":"publisher","DOI":"10.1145\/3072959.3073668"},{"key":"e_1_2_2_38_1","doi-asserted-by":"publisher","DOI":"10.1145\/2816795.2818123"},{"key":"e_1_2_2_39_1","doi-asserted-by":"publisher","DOI":"10.1109\/93.311653"},{"key":"e_1_2_2_40_1","article-title":"Videoscapes: Exploring Sparse","volume":"31","author":"Tompkin James","year":"2012","unstructured":"James Tompkin, Kwang In Kim, Jan Kautz, and Christian Theobalt. 2012. Videoscapes: Exploring Sparse, Unstructured Video Collections. ACM Trans. Graph. 31, 4, Article 68 (2012), 12 pages.","journal-title":"Unstructured Video Collections. ACM Trans. Graph."},{"key":"e_1_2_2_41_1","doi-asserted-by":"crossref","unstructured":"Anh Truong Floraine Berthouzoz Wilmot Li and Maneesh Agrawala. 2016. QuickCut: An Interactive Tool for Editing Narrated Video. In ACM UIST. 497--507.","DOI":"10.1145\/2984511.2984569"},{"key":"e_1_2_2_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2017.2749143"},{"key":"e_1_2_2_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2018.2884280"},{"key":"e_1_2_2_44_1","doi-asserted-by":"publisher","DOI":"10.1145\/3197517.3201284"},{"key":"e_1_2_2_45_1","doi-asserted-by":"publisher","DOI":"10.1111\/cgf.13559"},{"key":"e_1_2_2_46_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2015.2493959"},{"key":"e_1_2_2_47_1","doi-asserted-by":"publisher","DOI":"10.1109\/TMM.2018.2790163"},{"key":"e_1_2_2_48_1","doi-asserted-by":"publisher","DOI":"10.1007\/s11432-015-5494-4"}],"container-title":["ACM Transactions on Graphics"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3355089.3356520","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3355089.3356520","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T23:44:41Z","timestamp":1750203881000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3355089.3356520"}},"subtitle":["computational video montage from themed text"],"short-title":[],"issued":{"date-parts":[[2019,11,8]]},"references-count":48,"journal-issue":{"issue":"6","published-print":{"date-parts":[[2019,12,31]]}},"alternative-id":["10.1145\/3355089.3356520"],"URL":"https:\/\/doi.org\/10.1145\/3355089.3356520","relation":{},"ISSN":["0730-0301","1557-7368"],"issn-type":[{"value":"0730-0301","type":"print"},{"value":"1557-7368","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,11,8]]},"assertion":[{"value":"2019-11-08","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}