{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,28]],"date-time":"2026-07-28T14:15:40Z","timestamp":1785248140251,"version":"3.55.0"},"reference-count":46,"publisher":"Springer Science and Business Media LLC","issue":"24","license":[{"start":{"date-parts":[[2023,4,6]],"date-time":"2023-04-06T00:00:00Z","timestamp":1680739200000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2023,4,6]],"date-time":"2023-04-06T00:00:00Z","timestamp":1680739200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100001711","name":"Schweizerischer Nationalfonds zur F\u00f6rderung der Wissenschaftlichen Forschung","doi-asserted-by":"publisher","award":["CRSII5\u02d9193788"],"award-info":[{"award-number":["CRSII5\u02d9193788"]}],"id":[{"id":"10.13039\/501100001711","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100008375","name":"University of Basel","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100008375","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"published-print":{"date-parts":[[2023,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The past decades have seen an exponential growth in the amount of data which is produced by individuals. Smartphones which capture images, videos and sensor data have become commonplace, and wearables for fitness and health are growing in popularity. Lifelog retrieval systems aim to aid users in finding and exploring their personal history. We present two systems for lifelog retrieval: vitrivr and vitrivr-VR, which share a common retrieval model and backend for multi-modal multimedia retrieval. They differ in the user interface component, where vitrivr relies on a traditional desktop-based user interface and vitrivr-VR has a Virtual Reality user interface. Their effectiveness is evaluated at the Lifelog Search Challenge 2021, which offers an opportunity for interactive retrieval systems to compete with a focus on textual descriptions of past events. Our results show that the conventional user interface outperformed the VR user interface. However, the format of the evaluation campaign does not provide enough data for a thorough assessment and thus to make robust statements about the difference between the systems. Thus, we conclude by making suggestions for future interactive evaluation campaigns which would enable further insights.<\/jats:p>","DOI":"10.1007\/s11042-023-15082-w","type":"journal-article","created":{"date-parts":[[2023,4,6]],"date-time":"2023-04-06T09:02:56Z","timestamp":1680771776000},"page":"37829-37853","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":2,"title":["A tale of two interfaces: vitrivr at the lifelog search challenge"],"prefix":"10.1007","volume":"82","author":[{"ORCID":"https:\/\/orcid.org\/0000-0001-5386-330X","authenticated-orcid":false,"given":"Silvan","family":"Heller","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-3396-1516","authenticated-orcid":false,"given":"Florian","family":"Spiess","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-9865-6371","authenticated-orcid":false,"given":"Heiko","family":"Schuldt","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2023,4,6]]},"reference":[{"key":"15082_CR1","doi-asserted-by":"publisher","unstructured":"Ang W-H, Yen A-Z, Chu T-T, Huang H-H, Chen H-H (2021) LifeConcept: an interactive approach for multimodal lifelog retrieval through concept recommendation. In: 4th annual on lifelog search challenge. Association for Computing Machinery, New York, pp 47\u201351. https:\/\/doi.org\/10.1145\/3463948.3469070","DOI":"10.1145\/3463948.3469070"},{"issue":"4","key":"15082_CR2","doi-asserted-by":"publisher","first-page":"89","DOI":"10.1145\/1778765.1778826","volume":"29","author":"C Barnes","year":"2010","unstructured":"Barnes C, Goldman DB, Shechtman E, Finkelstein A. (2010) Video tapestries with continuous temporal zoom. ACM Trans Graph 29(4):89\u20131899. https:\/\/doi.org\/10.1145\/1778765.1778826","journal-title":"ACM Trans Graph"},{"key":"15082_CR3","unstructured":"Cer D, Yang Y, Kong S-Y, Hua N, Limtiaco N, John RS, Constant N, Guajardo-Cespedes M, Yuan S, Tar C, Sung Y-H, Strope B, Kurzweil R (2018) Universal sentence encoder. arXiv:1803.11175"},{"key":"15082_CR4","doi-asserted-by":"publisher","unstructured":"Gasser R, Rossetto L, Heller S, Schuldt H (2020) Cottontail DB: an open source database system for multimedia retrieval and analysis. In: Chen CW, Cucchiara R, Hua X-S, Qi G-J, Ricci E, Zhang Z, Zimmermann R (eds) International conference on multimedia (MM). Association for Computing Machinery, New York, pp 4465\u20134468. https:\/\/doi.org\/10.1145\/3394171.3414538","DOI":"10.1145\/3394171.3414538"},{"key":"15082_CR5","doi-asserted-by":"publisher","unstructured":"Gasser R, Rossetto L, Schuldt H (2019) Multimodal multimedia retrieval with vitrivr. In: International conference on multimedia retrieval (ICMR). Association for Computing Machinery, New York, pp 391\u2013394. https:\/\/doi.org\/10.1145\/3323873.3326921","DOI":"10.1145\/3323873.3326921"},{"key":"15082_CR6","unstructured":"Gasser R, Rossetto L, Schuldt H (2019) Towards an all-purpose content-based multimedia information retrieval system. arXiv:1902.03878"},{"key":"15082_CR7","doi-asserted-by":"publisher","unstructured":"Giangreco I (2018) Database support for large-scale multimedia retrieval. Thesis, University of Basel. https:\/\/doi.org\/10.5451\/unibas-006827345","DOI":"10.5451\/unibas-006827345"},{"key":"15082_CR8","doi-asserted-by":"publisher","DOI":"10.1145\/3210539","volume-title":"Proceedings of the 2018 ACM workshop, on the lifelog search challenge","author":"C Gurrin","year":"2018","unstructured":"Gurrin C, Schoeffmann K, Joho H, Dang-Nguyen D-T, Riegler M, Piras L (2018) Proceedings of the 2018 ACM workshop, on the lifelog search challenge. Association for Computing Machinery, New York"},{"issue":"2","key":"15082_CR9","doi-asserted-by":"publisher","first-page":"46","DOI":"10.3169\/mta.7.46","volume":"7","author":"C Gurrin","year":"2019","unstructured":"Gurrin C, Schoeffmann K, Joho H, Leibetseder A, Zhou L, Duane A, Dang-Nguyen D-T, Riegler M, Piras L, Tran M-T, Loko\u010d J, H\u00fcrst W (2019) [Invited papers] comparing approaches to interactive lifelog search at the lifelog search challenge (LSC2018). ITE Transactions on Media Technology and Applications 7(2):46\u201359. https:\/\/doi.org\/10.3169\/mta.7.46","journal-title":"ITE Transactions on Media Technology and Applications"},{"key":"15082_CR10","doi-asserted-by":"publisher","unstructured":"Heller S, Arnold R, Gasser R, Gsteiger V, Parian-Scherb M, Rossetto L, Sauter L, Spiess F, Schuldt H (2022) Multi-modal interactive video retrieval with temporal queries. In: J\u00f3nsson B@@, Gurrin C, Tran M-T, Dang-Nguyen D-T, Hu AM-C, Huynh Thi Thanh B, Huet B (eds) MultiMedia modeling. Springer International Publishing, Cham, pp 493\u2013498. https:\/\/doi.org\/10.1007\/978-3-030-98355-0_44","DOI":"10.1007\/978-3-030-98355-0_44"},{"key":"15082_CR11","doi-asserted-by":"publisher","unstructured":"Heller S, Gasser R, Illi C, Pasquinelli M, Sauter L, Spiess F, Schuldt H (2021) Towards explainable interactive multi-modal video retrieval with vitrivr. In: Loko\u010d J, Skopal T, Schoeffmann K, Mezaris V, Li X, Vrochidis S, Patras I (eds) MultiMedia modeling. Springer International Publishing, Cham, pp 435\u2013440. https:\/\/doi.org\/10.1007\/978-3-030-67835-7_41","DOI":"10.1007\/978-3-030-67835-7_41"},{"key":"15082_CR12","doi-asserted-by":"publisher","unstructured":"Heller S, Gasser R, Parian-Scherb M, Popovic S, Rossetto L, Sauter L, Spiess F, Schuldt H (2021) Interactive multimodal lifelog retrieval with vitrivr at LSC 2021. In: Gurrin C, Schoeffmann K, J\u00f3nsson B@@, Dang-Nguyen D-T, Lokoc J, Tran M-T, H\u00fcrst W, Rossetto L, Healy G (eds) Workshop on lifelog search challenge. Association for Computing Machinery, New York, pp 35\u201339. https:\/\/doi.org\/10.1145\/3463948.3469062","DOI":"10.1145\/3463948.3469062"},{"key":"15082_CR13","doi-asserted-by":"publisher","unstructured":"Heller S, Gsteiger V, Bailer W, Gurrin C, J\u00f3nsson B@@, Loko\u010d J, Leibetseder A, Mejzl\u00ed k F, Pe\u0161ka L, Rossetto L, Schall K, Schoeffmann K, Schuldt H, Spiess F, Tran L-D, Vadicamo L, Vesel\u00fd P, Vrochidis S, Wu J (2022) Interactive video retrieval evaluation at a distance: Comparing sixteen interactive video search systems in a remote setting at the 10th Video Browser Showdown. Int J Multimed Inf Retri 11(1):1\u201318. https:\/\/doi.org\/10.1007\/s13735-021-00225-2","DOI":"10.1007\/s13735-021-00225-2"},{"key":"15082_CR14","doi-asserted-by":"publisher","unstructured":"Heller S, Parian MA, Gasser R, Sauter L, Schuldt H (2020) Interactive lifelog retrieval with vitrivr. In: Gurrin C, Sch\u00f6ffmann K, J\u00f3nsson B@@, Dang-Nguyen D-T, Lokoc J, Tran M-T, H\u00fcrst W (eds) Third annual workshop on lifelog search challenge. Association for Computing Machinery, New York, pp 1\u20136. https:\/\/doi.org\/10.1145\/3379172.3391715","DOI":"10.1145\/3379172.3391715"},{"key":"15082_CR15","doi-asserted-by":"publisher","unstructured":"Heller S, Sauter L, Schuldt H, Rossetto L (2020) Multi-stage queries and temporal scoring in vitrivr. In: IEEE international conference on multimedia expo workshops (ICMEW). IEEE, New Jersey, pp 1\u20135. https:\/\/doi.org\/10.1109\/ICMEW46912.2020.9105954","DOI":"10.1109\/ICMEW46912.2020.9105954"},{"key":"15082_CR16","doi-asserted-by":"publisher","unstructured":"Li Y, Song Y, Cao L, Tetreault J, Goldberg L, Jaimes A, Luo J (2016) TGIF: a new dataset and benchmark on animated gif description. In: IEEE conference on computer vision and pattern recognition (CVPR). IEEE, pp 4641\u20134650. https:\/\/doi.org\/10.1109\/CVPR.2016.502","DOI":"10.1109\/CVPR.2016.502"},{"key":"15082_CR17","doi-asserted-by":"publisher","unstructured":"Lin T-Y, Maire M, Belongie S, Hays J, Perona P, Ramanan D, Doll\u00e1r P, Zitnick CL (2014) Microsoft COCO: common objects in context. In: Fleet D, Pajdla T, Schiele B, Tuytelaars T (eds) Computer Vision \u2013 ECCV 2014, vol 8693. Springer International Publishing, Cham, pp 740\u2013755. https:\/\/doi.org\/10.1007\/978-3-319-10602-1_48","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"15082_CR18","doi-asserted-by":"publisher","unstructured":"Loko\u010d J, Bailer W, Barthel KU, Gurrin C, Heller S, J\u00f3nsson B@@, Pe\u0161ka L, Rossetto L, Schoeffmann K, Vadicamo L, Vrochidis S, Wu J (2022) A task category space for user-centric comparative multimedia search evaluations. In: J\u00f3nsson B@@, Gurrin C, Tran M-T, Dang-Nguyen D-T, Hu AM-C, Huynh Thi Thanh B, Huet B (eds) MultiMedia modeling. Springer International Publishing, Cham, pp 193\u2013204. https:\/\/doi.org\/10.1007\/978-3-030-98358-1_16","DOI":"10.1007\/978-3-030-98358-1_16"},{"issue":"3","key":"15082_CR19","doi-asserted-by":"publisher","first-page":"91","DOI":"10.1145\/3445031","volume":"17","author":"J Loko\u010d","year":"2021","unstructured":"Loko\u010d J, Vesel\u00fd P, Mejzl\u00edk F, Koval\u010d\u00edk G, Sou\u010dek T, Rossetto L, Schoeffmann K, Bailer W, Gurrin C, Sauter L, Song J, Vrochidis S, Wu J, J\u00f3nsson B@@ (2021) Is the reign of interactive search eternal? Findings from the video browser showdown 2020. ACM Trans Multimed Comput Commun Appl 17(3):91\u201319126. https:\/\/doi.org\/10.1145\/3445031","journal-title":"ACM Trans Multimed Comput Commun Appl"},{"key":"15082_CR20","doi-asserted-by":"publisher","unstructured":"Peterhans S, Sauter L, Spiess F, Schuldt H (2022) Automatic generation of coherent image galleries in virtual reality. In: Linking theory and practice of digital libraries. Springer International Publishing, Cham, pp 282\u2013288. https:\/\/doi.org\/10.1007\/978-3-031-16802-4_23","DOI":"10.1007\/978-3-031-16802-4_23"},{"key":"15082_CR21","unstructured":"Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I (2021) Learning transferable visual models from natural language supervision. arXiv:2103.00020"},{"key":"15082_CR22","doi-asserted-by":"publisher","unstructured":"Rettig L, Shabani S, Sauter L, Cudr\u00e9-Mauroux P, Sokhn M, Schuldt H (2021) City-stories: combining entity linking multimedia retrieval, and crowdsourcing to make historical data accessible. In: Brambilla M, Chbeir R, Frasincar F, Manolescu I (eds) Web engineering. Springer International Publishing, Cham, pp 521\u2013524. https:\/\/doi.org\/10.1007\/978-3-030-74296-6_43","DOI":"10.1007\/978-3-030-74296-6_43"},{"key":"15082_CR23","doi-asserted-by":"publisher","unstructured":"Rossetto L (2018) Multi-modal video retrieval. Thesis, University of Basel. https:\/\/doi.org\/10.5451\/unibas-006859522","DOI":"10.5451\/unibas-006859522"},{"key":"15082_CR24","doi-asserted-by":"crossref","unstructured":"Rossetto L, Baumgartner M, Ashena N, Ruosch F, Pernischov\u00e1 R, Bernstein A (2020) LifeGraph: a knowledge graph for lifelogs. In: Proceedings of the third annual workshop on lifelog search challenge. Association for Computing Machinery, New York, pp 13\u201317","DOI":"10.1145\/3379172.3391717"},{"key":"15082_CR25","doi-asserted-by":"publisher","unstructured":"Rossetto L, Baumgartner M, Gasser R, Heitz L, Wang R, Bernstein A (2021) Exploring graph-querying approaches in lifegraph. In: Workshop on lifelog search challenge. Association for Computing Machinery, New York, pp 7\u201310. https:\/\/doi.org\/10.1145\/3463948.3469068","DOI":"10.1145\/3463948.3469068"},{"key":"15082_CR26","doi-asserted-by":"publisher","unstructured":"Rossetto L, Gasser R, Heller S, Parian MA, Schuldt H (2019) Retrieval of structured and unstructured data with vitrivr. In: Gurrin C, Sch\u00f6ffmann K, Joho H, Dang-Nguyen D-T, Riegler M, Piras L (eds) Workshop on lifelog search challenge. Association for Computing Machinery, New York, pp 27\u201331. https:\/\/doi.org\/10.1145\/3326460.3329160","DOI":"10.1145\/3326460.3329160"},{"issue":"4","key":"15082_CR27","doi-asserted-by":"publisher","first-page":"18","DOI":"10.1109\/MMUL.2021.3066779","volume":"28","author":"L Rossetto","year":"2021","unstructured":"Rossetto L, Gasser R, Heller S, Parian-Scherb M, Sauter L, Spiess F, Schuldt H, Pe\u0161ka L, Sou\u010dek T, Kratochv\u00edl M, Mejzl\u00edk F, Vesel\u00fd P, Loko\u010d J (2021) On the user-centric comparative remote evaluation of interactive video search systems. IEEE MultiMedia 28(4):18\u201328. https:\/\/doi.org\/10.1109\/MMUL.2021.3066779","journal-title":"IEEE MultiMedia"},{"key":"15082_CR28","doi-asserted-by":"publisher","unstructured":"Rossetto L, Gasser R, Sauter L, Bernstein A, Schuldt H (2021) A system for interactive multimedia retrieval evaluations. In: Loko\u010d J, Skopal T, Schoeffmann K, Mezaris V, Li X, Vrochidis S, Patras I (eds) MultiMedia modeling. Springer International Publishing, Cham, pp 385\u2013390. https:\/\/doi.org\/10.1007\/978-3-030-67835-7_33","DOI":"10.1007\/978-3-030-67835-7_33"},{"key":"15082_CR29","doi-asserted-by":"publisher","unstructured":"Rossetto L, Giangreco I, Heller S, Tanase C, Schuldt H (2016) Searching in video collections using sketches and sample images - the cineast system. In: Tian Q, Sebe N, Qi G-J, Huet B, Hong R, Liu X (eds) MultiMedia modeling, vol 9517. Springer International Publishing, Cham, pp 336\u2013341. https:\/\/doi.org\/10.1007\/978-3-319-27674-8_30","DOI":"10.1007\/978-3-319-27674-8_30"},{"key":"15082_CR30","doi-asserted-by":"publisher","unstructured":"Rossetto L, Giangreco I, Schuldt H (2014) Cineast: a multi-feature sketch-based video retrieval engine. In: IEEE international symposium on multimedia (ISM). IEEE, New Jersey, pp 18\u201323. https:\/\/doi.org\/10.1109\/ISM.2014.38","DOI":"10.1109\/ISM.2014.38"},{"key":"15082_CR31","doi-asserted-by":"publisher","unstructured":"Rossetto L, Giangreco I, Tanase C, Schuldt H (2016) Vitrivr: a flexible retrieval stack supporting multiple query modes for searching in multimedia collections. In: International conference on multimedia (MM). Association for Computing Machinery, New York, pp 1183\u20131186. https:\/\/doi.org\/10.1145\/2964284.2973797","DOI":"10.1145\/2964284.2973797"},{"key":"15082_CR32","doi-asserted-by":"publisher","unstructured":"Rossetto L, Parian MA, Gasser R, Giangreco I, Heller S, Schuldt H (2019) Deep learning-based concept detection in vitrivr. In: Kompatsiaris I, Huet B, Mezaris V, Gurrin C, Cheng W-H, Vrochidis S (eds) MultiMedia modeling, vol 11296. Springer International Publishing, Cham, pp 616\u2013621. https:\/\/doi.org\/10.1007\/978-3-030-05716-9_55","DOI":"10.1007\/978-3-030-05716-9_55"},{"key":"15082_CR33","doi-asserted-by":"publisher","unstructured":"Sauter L, Gasser R, Bernstein A, Schuldt H, Rossetto L (2022) An asynchronous scheme for the distributed evaluation of interactive multimedia retrieval. In: International workshop on interactive multimedia retrieval. Association for Computing Machinery, New York, pp 33\u201339. https:\/\/doi.org\/10.1145\/3552467.3554797","DOI":"10.1145\/3552467.3554797"},{"key":"15082_CR34","doi-asserted-by":"publisher","unstructured":"Sauter L, Rossetto L, Schuldt H (2018) Exploring cultural heritage in augmented reality with GoFind!. In: IEEE international conference on artificial intelligence and virtual reality (AIVR). IEEE, New Jersey, pp 187\u2013188. https:\/\/doi.org\/10.1109\/AIVR.2018.00041","DOI":"10.1109\/AIVR.2018.00041"},{"issue":"11","key":"15082_CR35","doi-asserted-by":"publisher","first-page":"2298","DOI":"10.1109\/TPAMI.2016.2646371","volume":"39","author":"B Shi","year":"2017","unstructured":"Shi B, Bai X, Yao C. (2017) An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE Trans Pattern Anal Mach Intell 39(11):2298\u20132304. https:\/\/doi.org\/10.1109\/TPAMI.2016.2646371","journal-title":"IEEE Trans Pattern Anal Mach Intell"},{"key":"15082_CR36","doi-asserted-by":"publisher","unstructured":"Sidorov O, Hu R, Rohrbach M, Singh A (2020) TextCaps: a dataset for image captioning with reading comprehension. In: Vedaldi A, Bischof H, Brox T, Frahm J-M (eds) Computer vision \u2013 ECCV 2020, vol 12347. Springer International Publishing, Cham, pp 742\u2013758. https:\/\/doi.org\/10.1007\/978-3-030-58536-5_44","DOI":"10.1007\/978-3-030-58536-5_44"},{"key":"15082_CR37","doi-asserted-by":"publisher","unstructured":"Spiess F, Gasser R, Heller S, Parian-Scherb M, Rossetto L, Sauter L, Schuldt H (2022) Multi-modal video retrieval in virtual reality with vitrivr-VR. In: J\u00f3nsson B@@, Gurrin C, Tran M-T, Dang-Nguyen D-T, Hu AM-C, Huynh Thi Thanh B, Huet B (eds) MultiMedia modeling. Springer International Publishing, Cham, pp 499\u2013504. https:\/\/doi.org\/10.1007\/978-3-030-98355-0_45","DOI":"10.1007\/978-3-030-98355-0_45"},{"key":"15082_CR38","doi-asserted-by":"publisher","unstructured":"Spiess F, Gasser R, Heller S, Rossetto L, Sauter L, van Zanten M, Schuldt H (2021) Exploring intuitive lifelog retrieval and interaction modes in virtual reality with vitrivr-vr. In: Gurrin C, Schoeffmann K, J\u00f3nsson B@@, Dang-Nguyen D-T, Lokoc J, Tran M-T, H\u00fcrst W, Rossetto L, Healy G (eds) Workshop on lifelog search challenge. Association for Computing Machinery, New York, pp 17\u201322. https:\/\/doi.org\/10.1145\/3463948.3469061","DOI":"10.1145\/3463948.3469061"},{"issue":"25","key":"15082_CR39","doi-asserted-by":"publisher","first-page":"33971","DOI":"10.1007\/s11042-021-11401-1","volume":"80","author":"N Spola\u00f4r","year":"2021","unstructured":"Spola\u00f4r N, Lee HD, Takaki WSR, Ensina LA, Parmezan ARS, Oliva JT, Coy CSR, Wu FC (2021) A video indexing and retrieval computational prototype based on transcribed speech. Multimed Tools Appl 80(25):33971\u201334017. https:\/\/doi.org\/10.1007\/s11042-021-11401-1","journal-title":"Multimed Tools Appl"},{"key":"15082_CR40","doi-asserted-by":"crossref","unstructured":"Szegedy C, Ioffe S, Vanhoucke V, Alemi AA (2017) Inception-v4, inception-resnet and the impact of residual connections on learning. In: Thirty-first AAAI conference, on artificial intelligence","DOI":"10.1609\/aaai.v31i1.11231"},{"key":"15082_CR41","doi-asserted-by":"publisher","unstructured":"Theus A, Rossetto L, Bernstein A (2022) HyText \u2013 a scene-text extraction method for video retrieval. In: J\u00f3nsson B@@, Gurrin C, Tran M-T, Dang-Nguyen D-T, Hu AM-C, Huynh Thi Thanh B, Huet B (eds) MultiMedia modeling. Springer International Publishing, Cham, pp 182\u2013193. https:\/\/doi.org\/10.1007\/978-3-030-98355-0_16","DOI":"10.1007\/978-3-030-98355-0_16"},{"key":"15082_CR42","doi-asserted-by":"crossref","unstructured":"Tran D, Bourdev L, Fergus R, Torresani L, Paluri M (2015) Learning spatiotemporal features with 3d convolutional networks. In: IEEE international conference on computer vision, pp 4489\u20134497","DOI":"10.1109\/ICCV.2015.510"},{"key":"15082_CR43","doi-asserted-by":"publisher","unstructured":"Wang X, Wu J, Chen J, Li L, Wang Y-F, Wang WY (2019) VaTeX: a large-scale high-quality multilingual dataset for video-and-language research. In: IEEE\/CVF international conference on computer vision (ICCV). IEEE, pp 4580\u20134590. https:\/\/doi.org\/10.1109\/ICCV.2019.00468","DOI":"10.1109\/ICCV.2019.00468"},{"key":"15082_CR44","doi-asserted-by":"publisher","unstructured":"Xu J, Mei T, Yao T, Rui Y (2016) MSR-VTT: a large video description dataset for bridging video and language. In: IEEE conference on computer vision and pattern recognition (CVPR). IEEE, pp 5288\u20135296. https:\/\/doi.org\/10.1109\/CVPR.2016.571","DOI":"10.1109\/CVPR.2016.571"},{"key":"15082_CR45","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1162\/tacl_a_00166","volume":"2","author":"P Young","year":"2014","unstructured":"Young P, Lai A, Hodosh M, Hockenmaier J (2014) From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Trans Assoc Comput Linguist 2:67\u201378. https:\/\/doi.org\/10.1162\/tacl_a_00166","journal-title":"Trans Assoc Comput Linguist"},{"key":"15082_CR46","doi-asserted-by":"crossref","unstructured":"Zhou X, Yao C, Wen H, Wang Y, Zhou S, He W, Liang J (2017) EAST: an efficient and accurate scene text detector. In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp 5551\u20135560","DOI":"10.1109\/CVPR.2017.283"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-023-15082-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-023-15082-w\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-023-15082-w.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,3]],"date-time":"2023-10-03T09:33:47Z","timestamp":1696325627000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-023-15082-w"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,4,6]]},"references-count":46,"journal-issue":{"issue":"24","published-print":{"date-parts":[[2023,10]]}},"alternative-id":["15082"],"URL":"https:\/\/doi.org\/10.1007\/s11042-023-15082-w","relation":{},"ISSN":["1380-7501","1573-7721"],"issn-type":[{"value":"1380-7501","type":"print"},{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,4,6]]},"assertion":[{"value":"29 April 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"1 February 2023","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"2 March 2023","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"6 April 2023","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"The code for the systems described in this paper is available open source at .","order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Statements and Declarations"}},{"value":"The logs analysed in this paper are available from corresponding authors or LSC organizers on reasonable request.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"Statements and Declarations"}},{"value":"The authors have no competing interests to declare that are relevant to the content of this article.","order":3,"name":"Ethics","group":{"name":"EthicsHeading","label":"Statements and Declarations"}}]}}