{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,23]],"date-time":"2026-01-23T08:18:48Z","timestamp":1769156328710,"version":"3.49.0"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,11,16]],"date-time":"2023-11-16T00:00:00Z","timestamp":1700092800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["J. Comput. Cult. Herit."],"published-print":{"date-parts":[[2023,12,31]]},"abstract":"<jats:p>Audiovisual (AV) archives, as an essential reservoir of our cultural assets, are suffering from the issue of accessibility. The complex nature of the medium itself made processing and interaction an open challenge still in the field of computer vision, multimodal learning, and human-computer interaction, as well as in culture and heritage. In recent years, with the raising of video retrieval tasks, methods in retrieving video content with natural language (text-to-video retrieval) gained quite some attention and have reached a performance level where real-world application is on the horizon. Appealing as it may sound, such methods focus on retrieving videos using plain visual-focused descriptions of what has happened in the video and finding videos such as instructions. It is too early to say such methods would be the new paradigms for accessing and encoding complex video content into high-dimensional data, but they are indeed innovative attempts and foundations to build future exploratory interfaces for AV archives (e.g., allow users to write stories and retrieve related snippets in the archive, or encoding video content at high-level for visualisation). This work filled the application gap by examining such text-to-video retrieval methods from an implementation point of view and proposed and verified a classifier-enhanced workflow to allow better results when dealing with in-situ queries that might have been different from the training dataset. Such a workflow is then applied to the real-world archive from T\u00e9l\u00e9vision Suisse Romande (RTS) to create a demo. At last, a human-centred evaluation is conducted to understand whether the text-to-video retrieval methods improve the overall experience of accessing AV archives.<\/jats:p>","DOI":"10.1145\/3627167","type":"journal-article","created":{"date-parts":[[2023,10,10]],"date-time":"2023-10-10T11:26:28Z","timestamp":1696937188000},"page":"1-16","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":6,"title":["Write What You Want: Applying Text-to-Video Retrieval to Audiovisual Archives"],"prefix":"10.1145","volume":"16","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-6866-1409","authenticated-orcid":false,"given":"Yuchen","family":"Yang","sequence":"first","affiliation":[{"name":"\u00c9cole Polytechnique F\u00e9d\u00e9rale de Lausanne - Laboratory for Experimental Museology+, Switzerland"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2023,11,16]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.618"},{"key":"e_1_3_2_3_2","article-title":"A context-based approach for dialogue act recognition using simple recurrent neural networks","author":"Bothe Chandrakant","year":"2018","unstructured":"Chandrakant Bothe, Cornelius Weber, Sven Magg, and Stefan Wermter. 2018. A context-based approach for dialogue act recognition using simple recurrent neural networks. arXiv preprint arXiv:1805.06280 (2018).","journal-title":"arXiv preprint arXiv:1805.06280"},{"issue":"1","key":"e_1_3_2_4_2","first-page":"145","article-title":"From grain to pixel? Notes on the technical dialectics in the small gauge film archive","volume":"7","author":"Cavallotti Diego","year":"2018","unstructured":"Diego Cavallotti. 2018. From grain to pixel? Notes on the technical dialectics in the small gauge film archive. NECSUS. European Journal of Media Studies 7, 1 (2018), 145\u2013164.","journal-title":"NECSUS. European Journal of Media Studies"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.5555\/2002472.2002497"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2022.103581"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1145\/997817.997857"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPRW53098.2021.00374"},{"key":"e_1_3_2_9_2","volume-title":"Audiovisual Archiving: Philosophy and Principles","author":"Edmondson. Ray","year":"2004","unstructured":"Ray Edmondson.2004. Audiovisual Archiving: Philosophy and Principles. Unesco Paris."},{"key":"e_1_3_2_10_2","article-title":"CLIP2Video: Mastering video-text retrieval via image clip","author":"Fang Han","year":"2021","unstructured":"Han Fang, Pengfei Xiong, Luhui Xu, and Yu Chen. 2021. CLIP2Video: Mastering video-text retrieval via image clip. arXiv preprint arXiv:2106.11097 (2021).","journal-title":"arXiv preprint arXiv:2106.11097"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00524"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1089\/cpb.2009.0138"},{"key":"e_1_3_2_13_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58548-8_13"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2013.379"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2010.57"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477495.3531960"},{"key":"e_1_3_2_17_2","article-title":"Deep fragment embeddings for bidirectional image sentence mapping","volume":"27","author":"Karpathy Andrej","year":"2014","unstructured":"Andrej Karpathy, Armand Joulin, and Li F. Fei-Fei. 2014. Deep fragment embeddings for bidirectional image sentence mapping. Advances in Neural Information Processing Systems 27 (2014).","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.4324\/9780367808433-1-3"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-83647-4_1"},{"key":"e_1_3_2_20_2","first-page":"2012","volume-title":"Proceedings of Coling 2016, the 26th International Conference on Computational Linguistics: Technical Papers","author":"Khanpour Hamed","year":"2016","unstructured":"Hamed Khanpour, Nishitha Guntakandla, and Rodney Nielsen. 2016. Dialogue act classification in domain-independent conversations using a deep recurrent neural network. In Proceedings of Coling 2016, the 26th International Conference on Computational Linguistics: Technical Papers. 2012\u20132021."},{"key":"e_1_3_2_21_2","article-title":"MDMMT-2: Multidomain multimodal transformer for video retrieval, one more step towards generalization","author":"Kunitsyn Alexander","year":"2022","unstructured":"Alexander Kunitsyn, Maksim Kalashnikov, Maksim Dzabraev, and Andrei Ivaniuta. 2022. MDMMT-2: Multidomain multimodal transformer for video retrieval, one more step towards generalization. arXiv preprint arXiv:2203.07086 (2022).","journal-title":"arXiv preprint arXiv:2203.07086"},{"key":"e_1_3_2_22_2","volume-title":"On Diary","author":"Lejeune Philippe","year":"2009","unstructured":"Philippe Lejeune. 2009. On Diary. University of Hawaii Press."},{"key":"e_1_3_2_23_2","article-title":"ACUTE-EVAL: Improved dialogue evaluation with optimized questions and multi-turn comparisons","author":"Li Margaret","year":"2019","unstructured":"Margaret Li, Jason Weston, and Stephen Roller. 2019. ACUTE-EVAL: Improved dialogue evaluation with optimized questions and multi-turn comparisons. arXiv preprint arXiv:1909.03087 (2019).","journal-title":"arXiv preprint arXiv:1909.03087"},{"key":"e_1_3_2_24_2","article-title":"A dual-attention hierarchical recurrent neural network for dialogue act classification","author":"Li Ruizhe","year":"2018","unstructured":"Ruizhe Li, Chenghua Lin, Matthew Collinson, Xiao Li, and Guanyi Chen. 2018. A dual-attention hierarchical recurrent neural network for dialogue act classification. arXiv preprint arXiv:1810.09154 (2018).","journal-title":"arXiv preprint arXiv:1810.09154"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.7551\/mitpress\/11214.001.0001"},{"issue":"4","key":"e_1_3_2_26_2","article-title":"Exploring digitised moving image collections: The SEMIA project, visual analysis and the turn to abstraction.","author":"Masson Eef","year":"2020","unstructured":"Eef Masson, Christian Gosvig Olesen, Nanne van Noord, and Giovanna Fossati. 2020. Exploring digitised moving image collections: The SEMIA project, visual analysis and the turn to abstraction. DHQ: Digital Humanities Quarterly4 (2020).","journal-title":"DHQ: Digital Humanities Quarterly"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00990"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00272"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.5860\/crln.79.6.296"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-77004-4_1"},{"key":"e_1_3_2_31_2","article-title":"Robust speech recognition via large-scale weak supervision","author":"Radford Alec","year":"2022","unstructured":"Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. 2022. Robust speech recognition via large-scale weak supervision. arXiv preprint arXiv:2212.04356 (2022).","journal-title":"arXiv preprint arXiv:2212.04356"},{"key":"e_1_3_2_32_2","article-title":"Dialogue act classification with context-aware self-attention","author":"Raheja Vipul","year":"2019","unstructured":"Vipul Raheja and Joel Tetreault. 2019. Dialogue act classification with context-aware self-attention. arXiv preprint arXiv:1904.02594 (2019).","journal-title":"arXiv preprint arXiv:1904.02594"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1007\/s11263-016-0987-1"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2015.7298682"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/MIS.2019.2954966"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/INCAE.2018.8579372"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01939"},{"key":"e_1_3_2_38_2","article-title":"Zero-shot video captioning with evolving pseudo-tokens","author":"Tewel Yoad","year":"2022","unstructured":"Yoad Tewel, Yoav Shalev, Roy Nadler, Idan Schwartz, and Lior Wolf. 2022. Zero-shot video captioning with evolving pseudo-tokens. arXiv preprint arXiv:2207.11100 (2022).","journal-title":"arXiv preprint arXiv:2207.11100"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1080\/01596306.2016.1148833"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3485447.3512022"},{"key":"e_1_3_2_41_2","first-page":"90","volume-title":"Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)","author":"Wang Sida I.","year":"2012","unstructured":"Sida I. Wang and Christopher D. Manning. 2012. Baselines and bigrams: Simple, good sentiment and topic classification. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers). 90\u201394."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00677"},{"key":"e_1_3_2_43_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00041"},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.571"},{"key":"e_1_3_2_45_2","article-title":"CLIP-ViP: Adapting pre-trained image-text model to video-language representation alignment","author":"Xue Hongwei","year":"2022","unstructured":"Hongwei Xue, Yuchong Sun, Bei Liu, Jianlong Fu, Ruihua Song, Houqiang Li, and Jiebo Luo. 2022. CLIP-ViP: Adapting pre-trained image-text model to video-language representation alignment. arXiv preprint arXiv:2209.06430 (2022).","journal-title":"arXiv preprint arXiv:2209.06430"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01136"},{"key":"e_1_3_2_47_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01234-2_29"}],"container-title":["Journal on Computing and Cultural Heritage"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3627167","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3627167","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,17]],"date-time":"2025-06-17T16:45:39Z","timestamp":1750178739000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3627167"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,11,16]]},"references-count":46,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,12,31]]}},"alternative-id":["10.1145\/3627167"],"URL":"https:\/\/doi.org\/10.1145\/3627167","relation":{},"ISSN":["1556-4673","1556-4711"],"issn-type":[{"value":"1556-4673","type":"print"},{"value":"1556-4711","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,11,16]]},"assertion":[{"value":"2023-02-28","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-07-25","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-11-16","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}