{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,5,2]],"date-time":"2026-05-02T22:58:35Z","timestamp":1777762715071,"version":"3.51.4"},"reference-count":303,"publisher":"Emerald","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2009,4,28]]},"abstract":"<jats:p>In this paper, we review 300 references on video retrieval, indicating when text-only solutions are unsatisfactory and showing the promising alternatives which are in majority concept-based. Therefore, central to our discussion is the notion of a semantic concept: an objective linguistic description of an observable entity. Specifically, we present our view on how its automated detection, selection under uncertainty, and interactive usage might solve the major scientific problem for video retrieval: the semantic gap. To bridge the gap, we lay down the anatomy of a concept-based video search engine. We present a component-wise decomposition of such an interdisciplinary multimedia system, covering influences from information retrieval, computer vision, machine learning, and human\u2013computer interaction. For each of the components we review state-of-the-art solutions in the literature, each having different characteristics and merits. Because of these differences, we cannot understand the progress in video retrieval without serious evaluation efforts such as carried out in the NIST TRECVID benchmark. We discuss its data, tasks, results, and the many derived community initiatives in creating annotations and baselines for repeatable experiments. We conclude with our perspective on future challenges and opportunities.<\/jats:p>","DOI":"10.1561\/1500000014","type":"journal-article","created":{"date-parts":[[2009,6,5]],"date-time":"2009-06-05T03:48:43Z","timestamp":1244173723000},"page":"215-322","source":"Crossref","is-referenced-by-count":238,"title":["Concept-Based Video Retrieval"],"prefix":"10.1108","volume":"2","author":[{"given":"Cees G. M.","family":"Snoek","sequence":"first","affiliation":[{"name":"University of Amsterdam, Science Park 107, 1098 XG Amsterdam ,","place":["The Netherlands"]}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Marcel","family":"Worring","sequence":"additional","affiliation":[{"name":"University of Amsterdam, Science Park 107, 1098 XG Amsterdam ,","place":["The Netherlands"]}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"140","published-online":{"date-parts":[[2009,4,28]]},"reference":[{"key":"2026040314322783500_ref001","volume-title":"Proceedings of the 11th Text Retrieval Conference","author":"Adams","year":"2002"},{"key":"2026040314322783500_ref002","first-page":"170","article-title":"\u201cSemantic indexing of multimedia content using visual, audio, and text cues\u201d","volume":"2003","author":"Adams","year":"2003","journal-title":"EURASIP Journal on Applied Signal Processing"},{"key":"2026040314322783500_ref003","doi-asserted-by":"crossref","first-page":"644","DOI":"10.1145\/1282280.1282372","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Adcock","year":"2007"},{"key":"2026040314322783500_ref004","first-page":"465","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Adcock","year":"2008"},{"key":"2026040314322783500_ref005","doi-asserted-by":"crossref","first-page":"832","DOI":"10.1145\/182.358434","article-title":"\u201cMaintaining knowledge about temporal intervals\u201d","volume":"26","author":"Allen","year":"1983","journal-title":"Communications of the ACM"},{"key":"2026040314322783500_ref006","volume-title":"Proceedings of the TRECVID Workshop","author":"Amir","year":"2003"},{"issue":"7-8","key":"2026040314322783500_ref007","first-page":"692","article-title":"\u201cEvaluation of active learning strategies for video indexing\u201d","volume":"22","author":"Ayache","year":"2007","journal-title":"Image Communication"},{"key":"2026040314322783500_ref008","first-page":"187","volume-title":"European Conference on Information Retrieval","author":"Ayache","year":"2008"},{"key":"2026040314322783500_ref009","first-page":"494","volume-title":"European Conference on Information Retrieval","author":"Ayache","year":"2007"},{"key":"2026040314322783500_ref010","doi-asserted-by":"crossref","first-page":"68","DOI":"10.1109\/6046.985555","article-title":"\u201cEvent based indexing of broadcasted sports video by intermodal collaboration\u201d","volume":"4","author":"Babaguchi","year":"2002","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref011","doi-asserted-by":"crossref","first-page":"575","DOI":"10.1109\/TMM.2004.830811","article-title":"\u201cPersonalized abstraction of broadcasted American football video by highlight selection\u201d","volume":"6","author":"Babaguchi","year":"2004","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref012","first-page":"805","volume-title":"International Joint Conference on Artificial Intelligence","author":"Banerjee","year":"2003"},{"key":"2026040314322783500_ref013","doi-asserted-by":"crossref","first-page":"346","DOI":"10.1016\/j.cviu.2007.09.014","article-title":"\u201cSpeeded-Up Robust Features (SURF)\u201d","volume":"110","author":"Bay","year":"2008","journal-title":"Computer Vision and Image Understanding"},{"key":"2026040314322783500_ref014","volume-title":"Proceedings of SPIE Conference on Internet Multimedia Management Systems","author":"Benitez","year":"2000"},{"key":"2026040314322783500_ref015","first-page":"395","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Bertini","year":"2005"},{"key":"2026040314322783500_ref016","volume-title":"Visual Information Retrieval","author":"Bimbo","year":"1999"},{"key":"2026040314322783500_ref017","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1109\/34.574790","article-title":"\u201cVisual image retrieval by elastic matching of user sketches\u201d","volume":"19","author":"Bimbo","year":"1997","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref018","volume-title":"Pattern Recognition and Machine Learning. Information Science and Statistics","author":"Bishop","year":"2006"},{"key":"2026040314322783500_ref019","doi-asserted-by":"crossref","DOI":"10.1007\/978-3-540-72895-5","volume-title":"Multimedia Retrieval","author":"Blanken","year":"2007"},{"key":"2026040314322783500_ref020","first-page":"359","volume-title":"Proceedings of the ACM SIGIR International Conference on Research and Development in Information Retrieval","author":"Bompada","year":"2007"},{"key":"2026040314322783500_ref021","doi-asserted-by":"crossref","first-page":"55","DOI":"10.1109\/34.41384","article-title":"\u201cMultichannel texture analysis using localized spatial filters\u201d","volume":"12","author":"Bovik","year":"1990","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref022","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Brown","year":"1995"},{"key":"2026040314322783500_ref023","doi-asserted-by":"crossref","first-page":"78","DOI":"10.1006\/jvci.1997.0404","article-title":"\u201cA survey on the automatic indexing of video data\u201d","volume":"10","author":"Brunelli","year":"1999","journal-title":"Journal of Visual Communication and Image Representation"},{"key":"2026040314322783500_ref024","doi-asserted-by":"crossref","first-page":"1520","DOI":"10.1109\/TPAMI.2007.70801","article-title":"\u201cDesign of multimodal dissimilarity spaces for retrieval of multimedia documents\u201d","volume":"30","author":"Bruno","year":"2008","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref025","first-page":"25","volume-title":"Proceedings of the ACM SIGIR International Conference on Research and Development in Information Retrieval","author":"Buckley","year":"2004"},{"key":"2026040314322783500_ref026","doi-asserted-by":"crossref","first-page":"13","DOI":"10.1162\/coli.2006.32.1.13","article-title":"\u201cEvaluating WordNet-based measures of lexical semantic relatedness\u201d","volume":"32","author":"Budanitsky","year":"2006","journal-title":"Computational Linguistics"},{"key":"2026040314322783500_ref027","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1023\/A:1009715923555","article-title":"\u201cA tutorial on support vector machines for pattern recognition\u201d","volume":"2","author":"Burges","year":"1998","journal-title":"Data Mining and Knowledge Discovery"},{"key":"2026040314322783500_ref028","doi-asserted-by":"crossref","first-page":"48","DOI":"10.1016\/j.cviu.2008.07.003","article-title":"\u201cPerformance evaluation of local color invariants\u201d","volume":"113","author":"Burghouts","year":"2009","journal-title":"Computer Vision and Image Understanding"},{"key":"2026040314322783500_ref029","first-page":"15","volume-title":"Proceedings International Conference on Semantics and Digital Media Technologies","author":"Byrne","year":"2008"},{"key":"2026040314322783500_ref030","volume-title":"Proceedings of the TRECVID Workshop","author":"Campbell","year":"2006"},{"key":"2026040314322783500_ref031","volume-title":"Proceedings of the TRECVID Workshop","author":"Cao","year":"2006"},{"key":"2026040314322783500_ref032","doi-asserted-by":"crossref","first-page":"1026","DOI":"10.1109\/TPAMI.2002.1023800","article-title":"\u201cBlobworld: Image segmentation using expectation-maximization and its application to image querying\u201d","volume":"24","author":"Carson","year":"2002","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref033","article-title":"\u201cLIBSVM: A library for support vector machines\u201d","author":"Chang","year":"2001"},{"key":"2026040314322783500_ref034","doi-asserted-by":"crossref","first-page":"602","DOI":"10.1109\/76.718507","article-title":"\u201cA fully automated content-based video search engine supporting spatio-temporal queries\u201d","volume":"8","author":"Chang","year":"1998","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"2026040314322783500_ref035","volume-title":"Proceedings of the TRECVID Workshop","author":"Chang","year":"2008"},{"key":"2026040314322783500_ref036","volume-title":"Proceedings of the TRECVID Workshop","author":"Chang","year":"2006"},{"key":"2026040314322783500_ref037","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/9780262033589.001.0001","volume-title":"Semi-Supervised Learning","author":"Chapelle","year":"2006"},{"key":"2026040314322783500_ref038","first-page":"902","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Chen","year":"2005"},{"key":"2026040314322783500_ref039","first-page":"212","volume-title":"Proceedings of the Joint Conference on Digital Libraries","author":"Chen","year":"2004"},{"key":"2026040314322783500_ref040","first-page":"21","volume-title":"CIVR","author":"Christel","year":"2006"},{"key":"2026040314322783500_ref041","first-page":"134","volume-title":"CIVR","author":"Christel","year":"2005"},{"key":"2026040314322783500_ref042","first-page":"561","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Christel","year":"2002"},{"key":"2026040314322783500_ref043","first-page":"1032","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Christel","year":"2004"},{"key":"2026040314322783500_ref044","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Chua","year":"2007"},{"key":"2026040314322783500_ref045","first-page":"656","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Chua","year":"2004"},{"key":"2026040314322783500_ref046","volume-title":"Proceedings of the TRECVID Workshop","author":"Chua","year":"2004"},{"key":"2026040314322783500_ref047","doi-asserted-by":"crossref","first-page":"1210","DOI":"10.1109\/TCSVT.2005.854238","article-title":"\u201cKnowledge-assisted semantic video object detection\u201d","volume":"15","author":"Dasiopoulou","year":"2005","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"2026040314322783500_ref048","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1348246.1348248","article-title":"\u201cImage retrieval: Ideas, influences and trends of the new age\u201d","volume":"40","author":"Datta","year":"2008","journal-title":"ACM Computing Surveys"},{"key":"2026040314322783500_ref049","doi-asserted-by":"crossref","first-page":"67","DOI":"10.1109\/38.126883","article-title":"\u201cCinematic principles for multimedia\u201d","volume":"11","author":"Davenport","year":"1991","journal-title":"IEEE Computer Graphics & Applications"},{"key":"2026040314322783500_ref050","doi-asserted-by":"crossref","first-page":"54","DOI":"10.1109\/MMUL.2003.1195161","article-title":"\u201cEditing out video editing\u201d","volume":"10","author":"Davis","year":"2003","journal-title":"IEEE MultiMedia"},{"key":"2026040314322783500_ref051","volume-title":"European Workshop on Content-Based Multimedia Indexing","author":"de Jong","year":"1999"},{"key":"2026040314322783500_ref052","doi-asserted-by":"crossref","first-page":"811","DOI":"10.1145\/1291233.1291417","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"de Rooij","year":"2007"},{"key":"2026040314322783500_ref053","first-page":"485","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"de Rooij","year":"2008"},{"key":"2026040314322783500_ref054","doi-asserted-by":"crossref","first-page":"800","DOI":"10.1109\/34.946985","article-title":"\u201cUnsupervised segmentation of color-texture regions in images and video\u201d","volume":"23","author":"Deng","year":"2001","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref055","first-page":"881","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Ebadollahi","year":"2006"},{"key":"2026040314322783500_ref056","doi-asserted-by":"crossref","first-page":"199","DOI":"10.1177\/016555150002600401","article-title":"\u201cVisual image retrieval: Seeking the alliance of concept-based and content-based paradigms\u201d","volume":"26","author":"Enser","year":"2000","journal-title":"Journal of Information Science"},{"key":"2026040314322783500_ref057","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1109\/TMM.2003.819583","article-title":"\u201cClassView: Hierarchical video shot classification, indexing and accessing\u201d","volume":"6","author":"Fan","year":"2004","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref058","doi-asserted-by":"crossref","first-page":"939","DOI":"10.1109\/TMM.2007.900143","article-title":"\u201cIncorporating concept ontology for hierarchical video classification, annotation and visualization\u201d","volume":"9","author":"Fan","year":"2007","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref059","doi-asserted-by":"crossref","DOI":"10.7551\/mitpress\/7287.001.0001","volume-title":"WordNet: An Electronic Lexical Database","author":"Fellbaum","year":"1998"},{"key":"2026040314322783500_ref060","volume-title":"Proceedings of the IEEE International Workshop on Performance Evaluation of Tracking and Surveillance","author":"Ferryman","year":"2007"},{"key":"2026040314322783500_ref061","doi-asserted-by":"crossref","first-page":"23","DOI":"10.1109\/2.410146","article-title":"\u201cQuery by image and video content: The QBIC system\u201d","volume":"28","author":"Flickner","year":"1995","journal-title":"IEEE Computer"},{"issue":"1-2","key":"2026040314322783500_ref062","doi-asserted-by":"crossref","first-page":"89","DOI":"10.1016\/S0167-6393(01)00061-9","article-title":"\u201cThe LIMSI broadcast news transcription system\u201d","volume":"37","author":"Gauvain","year":"2002","journal-title":"Speech Communication"},{"key":"2026040314322783500_ref063","doi-asserted-by":"crossref","first-page":"1338","DOI":"10.1109\/34.977559","article-title":"\u201cColor invariance\u201d","volume":"23","author":"Geusebroek","year":"2001","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"issue":"1-2","key":"2026040314322783500_ref064","doi-asserted-by":"crossref","first-page":"7","DOI":"10.1007\/s11263-005-4632-7","article-title":"\u201cA six-stimulus theory for stochastic texture\u201d","volume":"62","author":"Geusebroek","year":"2005","journal-title":"International Journal of Computer Vision"},{"key":"2026040314322783500_ref065","doi-asserted-by":"crossref","first-page":"848","DOI":"10.1109\/TPAMI.2002.1008391","article-title":"\u201cAdaptive image segmentation by combining photometric invariant region and edge information\u201d","volume":"24","author":"Gevers","year":"2002","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref066","doi-asserted-by":"crossref","first-page":"453","DOI":"10.1016\/S0031-3203(98)00036-3","article-title":"\u201cColor-based object recognition\u201d","volume":"32","author":"Gevers","year":"1999","journal-title":"Pattern Recognition"},{"key":"2026040314322783500_ref067","doi-asserted-by":"crossref","first-page":"102","DOI":"10.1109\/83.817602","article-title":"\u201cPicToSeek: Combining color and shape invariant features for image retrieval\u201d","volume":"9","author":"Gevers","year":"2000","journal-title":"IEEE Transactions on Image Processing"},{"key":"2026040314322783500_ref068","first-page":"564","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Goh","year":"2004"},{"key":"2026040314322783500_ref069","doi-asserted-by":"crossref","first-page":"198","DOI":"10.1177\/0165551506062337","article-title":"\u201cThe structure of collaborative tagging systems\u201d","volume":"32","author":"Golder","year":"2006","journal-title":"Journal of Information Science"},{"key":"2026040314322783500_ref070","doi-asserted-by":"crossref","first-page":"341","DOI":"10.1093\/applin\/11.4.341","article-title":"\u201cHow large can a receptive vocabulary be?\u201d","volume":"11","author":"Goulden","year":"1990","journal-title":"Applied Linguistics"},{"key":"2026040314322783500_ref071","doi-asserted-by":"crossref","first-page":"1605","DOI":"10.1109\/TMM.2008.2007290","article-title":"\u201cMulti-layer multi-instance learning for video concept detection\u201d","volume":"10","author":"Gu","year":"2008","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref072","doi-asserted-by":"crossref","first-page":"70","DOI":"10.1145\/253769.253798","article-title":"\u201cVisual information retrieval\u201d","volume":"40","author":"Gupta","year":"1997","journal-title":"Communications of the ACM"},{"key":"2026040314322783500_ref073","article-title":"\u201cFolksonomies: Tidying up tags?\u201d","volume":"12","author":"Guy","year":"2006","journal-title":"D-Lib Magazine"},{"key":"2026040314322783500_ref074","volume-title":"Content-Based Analysis of Digital Video","author":"Hanjalic","year":"2004"},{"key":"2026040314322783500_ref075","doi-asserted-by":"crossref","first-page":"541","DOI":"10.1109\/JPROC.2008.916338","article-title":"\u201cThe holy grail of multimedia information retrieval: So close or yet so far away?\u201d","volume":"96","author":"Hanjalic","year":"2008","journal-title":"Proceedings of the IEEE"},{"key":"2026040314322783500_ref076","doi-asserted-by":"crossref","first-page":"143","DOI":"10.1109\/TMM.2004.840618","article-title":"\u201cAffective video content representation and modeling\u201d","volume":"7","author":"Hanjalic","year":"2005","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref077","doi-asserted-by":"crossref","first-page":"41","DOI":"10.1145\/1282280.1282286","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Haubold","year":"2007"},{"key":"2026040314322783500_ref078","first-page":"437","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Haubold","year":"2008"},{"key":"2026040314322783500_ref079","volume-title":"Proceedings of the TRECVID Workshop","author":"Hauptmann","year":"2003"},{"key":"2026040314322783500_ref080","article-title":"\u201cLIBSCOM: Large analytics library and scalable concept ontology for multimedia research\u201d","author":"Hauptmann","year":"2009"},{"key":"2026040314322783500_ref081","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Hauptmann","year":"2004"},{"key":"2026040314322783500_ref082","doi-asserted-by":"crossref","first-page":"602","DOI":"10.1109\/JPROC.2008.916355","article-title":"\u201cVideo retrieval based on semantic concepts\u201d","volume":"96","author":"Hauptmann","year":"2008","journal-title":"Proceedings of the IEEE"},{"key":"2026040314322783500_ref083","first-page":"215","volume-title":"CIVR","author":"Hauptmann","year":"2005"},{"key":"2026040314322783500_ref084","doi-asserted-by":"crossref","first-page":"385","DOI":"10.1145\/1180639.1180721","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Hauptmann","year":"2006"},{"key":"2026040314322783500_ref085","doi-asserted-by":"crossref","first-page":"958","DOI":"10.1109\/TMM.2007.900150","article-title":"\u201cCan high-level concepts fill the semantic gap in video retrieval? A case study with broadcast news\u201d","volume":"9","author":"Hauptmann","year":"2007","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref086","doi-asserted-by":"crossref","first-page":"261","DOI":"10.1007\/s11042-008-0207-2","article-title":"\u201cA survey of browsing models for content based image retrieval\u201d","volume":"40","author":"Heesch","year":"2008","journal-title":"Multimedia Tools and Applications"},{"key":"2026040314322783500_ref087","first-page":"609","volume-title":"CIVR","author":"Heesch","year":"2005"},{"key":"2026040314322783500_ref088","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1109\/34.273716","article-title":"\u201cDecision combination in multiple classifier systems\u201d","volume":"16","author":"Ho","year":"1994","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref089","doi-asserted-by":"crossref","first-page":"265","DOI":"10.1016\/j.sigpro.2004.10.009","article-title":"\u201cColor texture measurement and segmentation\u201d","volume":"85","author":"Hoang","year":"2005","journal-title":"Signal Processing"},{"key":"2026040314322783500_ref090","first-page":"479","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Hollink","year":"2005"},{"key":"2026040314322783500_ref091","first-page":"327","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Hoogs","year":"2003"},{"key":"2026040314322783500_ref092","doi-asserted-by":"crossref","first-page":"45","DOI":"10.1109\/5254.796089","article-title":"\u201cNamed faces: Putting names to faces\u201d","volume":"14","author":"Houghton","year":"1999","journal-title":"IEEE Intelligent Systems"},{"key":"2026040314322783500_ref093","doi-asserted-by":"crossref","first-page":"14","DOI":"10.1109\/MMUL.2007.61","article-title":"\u201cReranking methods for visual search\u201d","volume":"14","author":"Hsu","year":"2007","journal-title":"IEEE MultiMedia"},{"key":"2026040314322783500_ref094","doi-asserted-by":"crossref","first-page":"179","DOI":"10.1109\/TIT.1962.1057692","article-title":"\u201cVisual pattern recognition by moment invariants\u201d","volume":"8","author":"Hu","year":"1962","journal-title":"IRE Transactions on Information Theory"},{"key":"2026040314322783500_ref095","doi-asserted-by":"crossref","first-page":"245","DOI":"10.1023\/A:1008108327226","article-title":"\u201cColor-spatial indexing and applications\u201d","volume":"35","author":"Huang","year":"1999","journal-title":"International Journal of Computer Vision"},{"key":"2026040314322783500_ref096","doi-asserted-by":"crossref","first-page":"648","DOI":"10.1109\/JPROC.2008.916364","article-title":"\u201cActive learning for interactive multimedia retrieval\u201d","volume":"96","author":"Huang","year":"2008","journal-title":"Proceedings of the IEEE"},{"key":"2026040314322783500_ref097","first-page":"78","volume-title":"Proceedings International Conference on Semantics and Digital Media Technologies","author":"Huijbregts","year":"2007"},{"key":"2026040314322783500_ref098","doi-asserted-by":"crossref","first-page":"76","DOI":"10.1109\/MMUL.2008.66","article-title":"\u201cVideo browsing on handheld devices \u2014 Interface designs for the next generation of mobile video players\u201d","volume":"15","author":"H\u00fcrst","year":"2008","journal-title":"IEEE MultiMedia"},{"key":"2026040314322783500_ref099","first-page":"177","volume-title":"Proceedings of the ACM SIGMM International Workshop on Multimedia Information Retrieval","author":"Huurnink","year":"2007"},{"key":"2026040314322783500_ref100","first-page":"459","volume-title":"Proceedings of the ACM International Conference on Multimedia Information Retrieval","author":"Huurnink","year":"2008"},{"key":"2026040314322783500_ref101","volume-title":"Proceedings of the International World Wide Web Conference","author":"Hyv\u00f6nen","year":"2003"},{"key":"2026040314322783500_ref102","first-page":"21","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Iyengar","year":"2005"},{"key":"2026040314322783500_ref103","doi-asserted-by":"crossref","first-page":"4","DOI":"10.1109\/34.824819","article-title":"\u201cStatistical pattern recognition: A review\u201d","volume":"22","author":"Jain","year":"2000","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref104","doi-asserted-by":"crossref","first-page":"1167","DOI":"10.1016\/0031-3203(91)90143-S","article-title":"\u201cUnsupervised texture segmentation using gabor filters\u201d","volume":"24","author":"Jain","year":"1991","journal-title":"Pattern Recognition"},{"key":"2026040314322783500_ref105","doi-asserted-by":"crossref","first-page":"1369","DOI":"10.1016\/S0031-3203(97)00131-3","article-title":"\u201cShape-based retrieval: A case study with trademark image databases\u201d","volume":"31","author":"Jain","year":"1998","journal-title":"Pattern Recognition"},{"key":"2026040314322783500_ref106","doi-asserted-by":"crossref","first-page":"27","DOI":"10.1145\/190627.190638","article-title":"\u201cMetadata in video databases\u201d","volume":"23","author":"Jain","year":"1994","journal-title":"ACM SIGMOD Record"},{"key":"2026040314322783500_ref107","first-page":"949","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Jiang","year":"2007"},{"key":"2026040314322783500_ref108","first-page":"161","volume-title":"Proceedings of the IEEE International Conference on Image Processing","author":"Jiang","year":"2008"},{"key":"2026040314322783500_ref109","doi-asserted-by":"crossref","first-page":"494","DOI":"10.1145\/1282280.1282352","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Jiang","year":"2007"},{"key":"2026040314322783500_ref110","volume-title":"Technical Report 223-2008-1","author":"Jiang","year":"2008"},{"key":"2026040314322783500_ref111","first-page":"169","volume-title":"Advances in Kernel Methods: Support Vector Learning","author":"Joachims","year":"1999"},{"key":"2026040314322783500_ref112","doi-asserted-by":"crossref","first-page":"712","DOI":"10.1109\/JPROC.2008.916383","article-title":"\u201cApplication potential of multimedia information retrieval\u201d","volume":"96","author":"Kankanhalli","year":"2008","journal-title":"Proceedings of the IEEE"},{"key":"2026040314322783500_ref113","doi-asserted-by":"crossref","first-page":"319","DOI":"10.1109\/TPAMI.2008.57","article-title":"\u201cFramework for performance evaluation of face, text and vehicle detection and tracking in video: Data, metrics and protocol\u201d","volume":"31","author":"Kasturi","year":"2009","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref114","first-page":"530","volume-title":"Proceedings of the International Conference on Pattern Recognition","author":"Kato","year":"1992"},{"key":"2026040314322783500_ref115","first-page":"1174","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Render","year":"2005"},{"key":"2026040314322783500_ref116","volume-title":"Technical Report 221-2006-7","author":"Kennedy","year":"2006"},{"key":"2026040314322783500_ref117","doi-asserted-by":"crossref","first-page":"333","DOI":"10.1145\/1282280.1282331","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Kennedy","year":"2007"},{"key":"2026040314322783500_ref118","doi-asserted-by":"crossref","first-page":"567","DOI":"10.1109\/JPROC.2008.916345","article-title":"\u201cQuery-adaptive fusion for multimodal search\u201d","volume":"96","author":"Kennedy","year":"2008","journal-title":"Proceedings of the IEEE"},{"key":"2026040314322783500_ref119","first-page":"882","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Kennedy","year":"2005"},{"key":"2026040314322783500_ref120","doi-asserted-by":"crossref","first-page":"489","DOI":"10.1109\/34.55109","article-title":"\u201cInvariant image recognition by zernike moments\u201d","volume":"12","author":"Khotanzad","year":"1990","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref121","doi-asserted-by":"crossref","first-page":"226","DOI":"10.1109\/34.667881","article-title":"\u201cOn combining classifiers\u201d","volume":"20","author":"Kittler","year":"1998","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref122","volume-title":"Working Notes for the Cross-Language Evaluation Forum Workshop","author":"Larson","year":"2008"},{"key":"2026040314322783500_ref123","first-page":"424","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Latecki","year":"2000"},{"key":"2026040314322783500_ref124","first-page":"2169","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Lazebnik","year":"2006"},{"key":"2026040314322783500_ref125","article-title":"\u201cDesigning the user-interface for the Ffschlar digital video library\u201d","volume":"2","author":"Lee","year":"2002","journal-title":"Journal of Digital Information"},{"key":"2026040314322783500_ref126","volume-title":"Building Large Knowledge-based Systems: Representation and Inference in the Cyc Project","author":"Lenat","year":"1990"},{"key":"2026040314322783500_ref127","first-page":"24","volume-title":"Proceedings of the International Conference on Systems Documentation","author":"Lesk","year":"1986"},{"key":"2026040314322783500_ref128","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4471-3702-3","volume-title":"Principles of Visual Information Retrieval","author":"Lew","year":"2001"},{"key":"2026040314322783500_ref129","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1126004.1126005","article-title":"\u201cContent-based multimedia information retrieval: State of the art and challenges\u201d","volume":"2","author":"Lew","year":"2006","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"2026040314322783500_ref130","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Li","year":"2008"},{"key":"2026040314322783500_ref131","volume-title":"Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing","author":"Li","year":"2009"},{"key":"2026040314322783500_ref132","doi-asserted-by":"crossref","first-page":"603","DOI":"10.1145\/1282280.1282366","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Li","year":"2007"},{"key":"2026040314322783500_ref133","volume-title":"Proceedings of the TRECVID Workshop","author":"Lin","year":"2003"},{"key":"2026040314322783500_ref134","doi-asserted-by":"crossref","first-page":"267","DOI":"10.1007\/s10994-007-5018-6","article-title":"\u201cA note on Platt\u2019s probabilistic outputs for support vector machines\u201d","volume":"68","author":"Lin","year":"2007","journal-title":"Machine Learning"},{"key":"2026040314322783500_ref135","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Lin","year":"2002"},{"key":"2026040314322783500_ref136","doi-asserted-by":"crossref","first-page":"211","DOI":"10.1023\/B:BTTJ.0000047600.45421.6d","article-title":"\u201cConceptNet: A practical commonsense reasoning toolkit\u201d","volume":"22","author":"Liu","year":"2004","journal-title":"BT Technology Journal"},{"key":"2026040314322783500_ref137","doi-asserted-by":"crossref","first-page":"208","DOI":"10.1145\/1291233.1291279","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Liu","year":"2007"},{"key":"2026040314322783500_ref138","doi-asserted-by":"crossref","first-page":"240","DOI":"10.1109\/TMM.2007.911826","article-title":"\u201cAssociation and temporal rule mining for post-filtering of semantic concept detection in video\u201d","volume":"10","author":"Liu","year":"2008","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref139","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1023\/B:VISI.0000029664.99615.94","article-title":"\u201cDistinctive image features from scale-invariant keypoints\u201d","volume":"60","author":"Lowe","year":"2004","journal-title":"International Journal of Computer Vision"},{"key":"2026040314322783500_ref140","doi-asserted-by":"crossref","first-page":"269","DOI":"10.1023\/A:1012491016871","article-title":"\u201cIndexing and retrieval of audio: A survey\u201d","volume":"15","author":"Lu","year":"2001","journal-title":"Multimedia Tools and Applications"},{"key":"2026040314322783500_ref141","doi-asserted-by":"crossref","first-page":"293","DOI":"10.1145\/1291233.1291295","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Luan","year":"2007"},{"key":"2026040314322783500_ref142","first-page":"107","volume-title":"IEEE Symposium on Visual Analytics Science and Technology","author":"Luo","year":"2007"},{"key":"2026040314322783500_ref143","doi-asserted-by":"crossref","first-page":"184","DOI":"10.1007\/s005300050121","article-title":"\u201cNeTra: A toolbox for navigating large image databases\u201d","volume":"7","author":"Ma","year":"1999","journal-title":"Multimedia Systems"},{"key":"2026040314322783500_ref144","doi-asserted-by":"crossref","first-page":"619","DOI":"10.1145\/1282280.1282368","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Magalh\u00e3es","year":"2007"},{"key":"2026040314322783500_ref145","doi-asserted-by":"crossref","first-page":"836","DOI":"10.1109\/34.531803","article-title":"\u201cTexture features for browsing and retrieval of image data\u201d","volume":"18","author":"Manjunath","year":"1996","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref146","volume-title":"Introduction to MPEG-7: Multimedia Content Description Interface","author":"Manjunath","year":"2002"},{"key":"2026040314322783500_ref147","doi-asserted-by":"crossref","DOI":"10.1017\/CBO9780511809071","volume-title":"Introduction to Information Retrieval","author":"Manning","year":"2008"},{"key":"2026040314322783500_ref148","first-page":"31","volume-title":"Proceedings ACM International Conference on Hypertext and Hypermedia","author":"Marlow","year":"2006"},{"key":"2026040314322783500_ref149","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1108\/10650750610706998","article-title":"\u201cTowards user-centered indexing in digital image collections\u201d","volume":"22","author":"Matusiak","year":"2006","journal-title":"OCLC Systems & Services"},{"key":"2026040314322783500_ref150","first-page":"61","volume-title":"CIVR","author":"McDonald","year":"2005"},{"key":"2026040314322783500_ref151","volume-title":"Proceedings of the TRECVID Workshop","author":"Mei","year":"2007"},{"issue":"1-3","key":"2026040314322783500_ref152","doi-asserted-by":"crossref","first-page":"3","DOI":"10.1016\/j.cviu.2003.10.011","article-title":"\u201cMoment invariants for recognition under changing viewpoint and illumination\u201d","volume":"94","author":"Mindru","year":"2004","journal-title":"Computer Vision and Image Understanding"},{"key":"2026040314322783500_ref153","doi-asserted-by":"crossref","first-page":"121","DOI":"10.1016\/j.jvcir.2007.04.002","article-title":"\u201cVideo summarisation: A conceptual framework and survey of the state of the art\u201d","volume":"19","author":"Money","year":"2008","journal-title":"Journal of Visual Communication and Image Representation"},{"key":"2026040314322783500_ref154","doi-asserted-by":"crossref","first-page":"1632","DOI":"10.1109\/TPAMI.2007.70822","article-title":"\u201cRandomized clustering forests for image classification\u201d","volume":"30","author":"Moosmann","year":"2008","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref155","doi-asserted-by":"crossref","first-page":"263","DOI":"10.1023\/B:MTAP.0000017031.26875.f7","article-title":"\u201cSaying what it means: Semi-automated (news) media annotation\u201d","volume":"22","author":"Nack","year":"2004","journal-title":"Multimedia Tools and Applications"},{"key":"2026040314322783500_ref156","doi-asserted-by":"crossref","first-page":"348","DOI":"10.1016\/j.jvcir.2004.04.010","article-title":"\u201cOn supervision and statistical learning for semantic multimedia analysis\u201d","volume":"15","author":"Naphade","year":"2004","journal-title":"Journal of Visual Communication and Image Representation"},{"key":"2026040314322783500_ref157","doi-asserted-by":"crossref","first-page":"141","DOI":"10.1109\/6046.909601","article-title":"\u201cA probabilistic framework for semantic video indexing, filtering and retrieval\u201d","volume":"3","author":"Naphade","year":"2001","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref158","doi-asserted-by":"crossref","first-page":"793","DOI":"10.1109\/TNN.2002.1021881","article-title":"\u201cExtracting semantics from audiovisual content: The final frontier in multimedia retrieval\u201d","volume":"13","author":"Naphade","year":"2002","journal-title":"IEEE Transactions on Neural Networks"},{"key":"2026040314322783500_ref159","volume-title":"Technical Report RC23612","author":"Naphade","year":"2005"},{"key":"2026040314322783500_ref160","doi-asserted-by":"crossref","first-page":"40","DOI":"10.1109\/76.981844","article-title":"\u201cA factor graph framework for semantic video indexing\u201d","volume":"12","author":"Naphade","year":"2002","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"2026040314322783500_ref161","first-page":"109","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Naphade","year":"2004"},{"key":"2026040314322783500_ref162","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Naphade","year":"2004"},{"key":"2026040314322783500_ref163","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1109\/MMUL.2006.63","article-title":"\u201cLarge-scale concept ontology for multimedia\u201d","volume":"13","author":"Naphade","year":"2006","journal-title":"IEEE MultiMedia"},{"key":"2026040314322783500_ref164","doi-asserted-by":"crossref","first-page":"991","DOI":"10.1145\/1291233.1291448","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Natsev","year":"2007"},{"key":"2026040314322783500_ref165","first-page":"598","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Natsev","year":"2005"},{"key":"2026040314322783500_ref166","first-page":"143","volume-title":"CIVR","author":"Neo","year":"2006"},{"key":"2026040314322783500_ref167","first-page":"476","volume-title":"Proceedings of the IEEE International Conference on Advanced Video and Signal Based Surveillance","author":"Nghiem","year":"2007"},{"key":"2026040314322783500_ref168","doi-asserted-by":"crossref","first-page":"203","DOI":"10.1016\/j.jvlc.2006.09.002","article-title":"\u201cInteractive access to large image collections using similarity-based visualization\u201d","volume":"19","author":"Nguyen","year":"2008","journal-title":"Journal of Visual Languages and Computing"},{"key":"2026040314322783500_ref169","doi-asserted-by":"crossref","first-page":"1404","DOI":"10.1109\/TMM.2007.906586","article-title":"\u201cInteractive search by direct manipulation of dissimilarity space\u201d","volume":"9","author":"Nguyen","year":"2007","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref170","doi-asserted-by":"crossref","first-page":"137","DOI":"10.1109\/83.817605","article-title":"\u201cDetection of moving objects in video using a robust motion similarity measure\u201d","volume":"9","author":"Nguyen","year":"2000","journal-title":"IEEE Transactions on Image Processing"},{"key":"2026040314322783500_ref171","article-title":"\u201cTRECVID video retrieval evaluation \u2014 Online proceedings\u201d","author":"NIST","year":"2001"},{"key":"2026040314322783500_ref172","article-title":"\u201cWhat is web 2.0\u201d","author":"O\u2019Reily","year":"2005"},{"key":"2026040314322783500_ref173","volume-title":"Proceedings of the TRECVID Workshop","author":"Over","year":"2008"},{"key":"2026040314322783500_ref174","volume-title":"Proceedings of the TRECVID Workshop","author":"Over","year":"2005"},{"key":"2026040314322783500_ref175","volume-title":"Proceedings of the TRECVID Workshop","author":"Over","year":"2006"},{"key":"2026040314322783500_ref176","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/1290031","volume-title":"Proceedings of the International Workshop on TRECVID Video Summarization","author":"Over","year":"2007"},{"key":"2026040314322783500_ref177","volume-title":"Vision Science: Photons to Phenomenology","author":"Palmer","year":"1999"},{"key":"2026040314322783500_ref178","first-page":"65","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Pass","year":"1996"},{"key":"2026040314322783500_ref179","volume-title":"Proceedings of the International Conference on Advances in Visual Information Systems","author":"Pecenovic","year":"2000"},{"key":"2026040314322783500_ref180","doi-asserted-by":"crossref","first-page":"233","DOI":"10.1007\/BF00123143","article-title":"\u201cPhotobook: Content-based manipulation of image databases\u201d","volume":"18","author":"Pentland","year":"1996","journal-title":"International Journal of Computer Vision"},{"key":"2026040314322783500_ref181","volume-title":"Proceedings of the TRECVID Workshop","author":"Petersohn","year":"2004"},{"key":"2026040314322783500_ref182","doi-asserted-by":"crossref","first-page":"1501","DOI":"10.1109\/TPAMI.2002.1046166","article-title":"\u201cMatching and retrieval of distorted and occluded shapes using dynamic programming\u201d","volume":"24","author":"Petrakis","year":"2002","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref183","doi-asserted-by":"crossref","first-page":"255","DOI":"10.1049\/ip-vis:20050059","article-title":"\u201cKnowledge representation and semantic annotation of multimedia content\u201d","volume":"153","author":"Petridis","year":"2006","journal-title":"IEE Proceedings of Vision, Image and Signal Processing"},{"key":"2026040314322783500_ref184","doi-asserted-by":"crossref","first-page":"61","DOI":"10.7551\/mitpress\/1113.003.0008","volume-title":"Advances in Large Margin Classifiers","author":"Platt","year":"2000"},{"key":"2026040314322783500_ref185","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Pogalin","year":"2008"},{"key":"2026040314322783500_ref186","article-title":"\u201cCorrelative multilabel video annotation with temporal kernels\u201d","volume":"5","author":"Qi","year":"2009","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"2026040314322783500_ref187","volume-title":"Proceedings of the 11th Text Retrieval Conference","author":"Qu\u00e9not","year":"2002"},{"key":"2026040314322783500_ref188","doi-asserted-by":"crossref","first-page":"257","DOI":"10.1109\/5.18626","article-title":"\u201cA tutorial on hidden markov models and selected applications in speech recognition\u201d","volume":"77","author":"Rabiner","year":"1989","journal-title":"Proceedings of the IEEE"},{"key":"2026040314322783500_ref189","doi-asserted-by":"crossref","first-page":"291","DOI":"10.1109\/34.761261","article-title":"\u201cFiltering for texture classification: A comparative study\u201d","volume":"21","author":"Randen","year":"1999","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref190","doi-asserted-by":"crossref","first-page":"923","DOI":"10.1109\/TMM.2007.900138","article-title":"\u201cBridging the gap: Query by semantic example\u201d","volume":"9","author":"Rasiwasia","year":"2007","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref191","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Rautiainen","year":"2004"},{"key":"2026040314322783500_ref192","first-page":"115","volume-title":"Proceedings of the Joint Workshop on Hands-Free Speech Communication and Microphone Arrays","author":"Renais","year":"2008"},{"key":"2026040314322783500_ref193","first-page":"448","volume-title":"International Joint Conference on Artificial Intelligence","author":"Resnik","year":"1995"},{"key":"2026040314322783500_ref194","first-page":"73","volume-title":"Proceedings of the Text Retrieval Conference","author":"Robertson","year":"1996"},{"key":"2026040314322783500_ref195","doi-asserted-by":"crossref","first-page":"147","DOI":"10.1145\/356551.356554","article-title":"\u201cPicture processing by computer\u201d","volume":"1","author":"Rosenfeld","year":"1969","journal-title":"ACM Computing Surveys"},{"key":"2026040314322783500_ref196","doi-asserted-by":"crossref","first-page":"99","DOI":"10.1023\/A:1026543900054","article-title":"\u201cThe earth mover\u2019s distance as a metric for image retrieval\u201d","volume":"40","author":"Rubner","year":"2000","journal-title":"International Journal of Computer Vision"},{"key":"2026040314322783500_ref197","doi-asserted-by":"crossref","first-page":"644","DOI":"10.1109\/76.718510","article-title":"\u201cRelevance feedback: A power tool in interactive content-based image retrieval\u201d","volume":"8","author":"Rui","year":"1998","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"issue":"1-3","key":"2026040314322783500_ref198","doi-asserted-by":"crossref","first-page":"157","DOI":"10.1007\/s11263-007-0090-8","article-title":"\u201cLabelMe: A database and web-based tool for image annotation\u201d","volume":"77","author":"Russell","year":"2008","journal-title":"International Journal of Computer Vision"},{"key":"2026040314322783500_ref199","volume-title":"Introduction to Modern Information Retrieval","author":"Salton","year":"1983"},{"key":"2026040314322783500_ref200","doi-asserted-by":"crossref","first-page":"22","DOI":"10.1109\/93.752960","article-title":"\u201cName-It: Naming and detecting faces in news videos\u201d","volume":"6","author":"Satoh","year":"1999","journal-title":"IEEE MultiMedia"},{"key":"2026040314322783500_ref201","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1109\/5254.940028","article-title":"\u201cOntology-based photo annotation\u201d","volume":"16","author":"Schreiber","year":"2001","journal-title":"IEEE Intelligent Systems"},{"key":"2026040314322783500_ref202","doi-asserted-by":"crossref","first-page":"51","DOI":"10.1007\/978-1-4471-3702-3_3","volume-title":"Principles of Visual Information Retrieval","author":"Sebe","year":"2001"},{"key":"2026040314322783500_ref203","doi-asserted-by":"crossref","first-page":"64","DOI":"10.1109\/MMUL.2007.74","article-title":"\u201cHigh-performance distributed image and video content analysis with parallel-horus\u201d","volume":"14","author":"Seinstra","year":"2007","journal-title":"IEEE MultiMedia"},{"key":"2026040314322783500_ref204","first-page":"275","volume-title":"Proceedings of the ACM SIGMM International Workshop on Multimedia Information Retrieval","author":"Shamma","year":"2007"},{"key":"2026040314322783500_ref205","first-page":"1","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Shotton","year":"2008"},{"key":"2026040314322783500_ref206","doi-asserted-by":"crossref","first-page":"189","DOI":"10.1007\/s11263-005-4264-y","article-title":"\u201cObject level grouping for video shots\u201d","volume":"67","author":"Sivic","year":"2006","journal-title":"International Journal of Computer Vision"},{"key":"2026040314322783500_ref207","doi-asserted-by":"crossref","first-page":"548","DOI":"10.1109\/JPROC.2008.916343","article-title":"\u201cEfficient visual search for objects in videos\u201d","volume":"96","author":"Sivic","year":"2008","journal-title":"Proceedings of the IEEE"},{"key":"2026040314322783500_ref208","first-page":"19","volume-title":"CIVR","author":"Smeaton","year":"2005"},{"key":"2026040314322783500_ref209","doi-asserted-by":"crossref","first-page":"545","DOI":"10.1016\/j.is.2006.09.001","article-title":"\u201cTechniques used and open challenges to the analysis, indexing and retrieval of digital video\u201d","volume":"32","author":"Smeaton","year":"2007","journal-title":"Information Systems"},{"key":"2026040314322783500_ref210","first-page":"547","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Smeaton","year":"2008"},{"key":"2026040314322783500_ref211","first-page":"321","article-title":"\u201cEvaluation campaigns and TRECVid\u201d","volume-title":"Proceedings of the ACM SIGMM International Workshop on Multimedia Information Retrieval","author":"Smeaton","year":"2006"},{"key":"2026040314322783500_ref212","volume-title":"Multimedia Content Analysis, Theory and Applications","author":"Smeaton","year":"2008"},{"key":"2026040314322783500_ref213","doi-asserted-by":"crossref","first-page":"1349","DOI":"10.1109\/34.895972","article-title":"\u201cContent-based image retrieval at the end of the early years\u201d","volume":"22","author":"Smeulders","year":"2000","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref214","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1109\/93.621578","article-title":"\u201cVisually searching the web for content\u201d","volume":"4","author":"Smith","year":"1997","journal-title":"IEEE MultiMedia"},{"key":"2026040314322783500_ref215","first-page":"445","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Smith","year":"2003"},{"key":"2026040314322783500_ref216","article-title":"\u201cIBM multimedia analysis and retrieval system\u201d","author":"Smith","year":"2008"},{"key":"2026040314322783500_ref217","doi-asserted-by":"crossref","first-page":"975","DOI":"10.1109\/TMM.2007.900156","article-title":"\u201cAdding semantics to detectors for video retrieval\u201d","volume":"9","author":"Snoek","year":"2007","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref218","volume-title":"Proceedings of the TRECVID Workshop","author":"Snoek","year":"2005"},{"key":"2026040314322783500_ref219","volume-title":"Proceedings of the TRECVID Workshop","author":"Snoek","year":"2006"},{"key":"2026040314322783500_ref220","doi-asserted-by":"crossref","first-page":"638","DOI":"10.1109\/TMM.2005.850966","article-title":"\u201cMultimedia event-based video indexing using time intervals\u201d","volume":"7","author":"Snoek","year":"2005","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref221","doi-asserted-by":"crossref","first-page":"5","DOI":"10.1023\/B:MTAP.0000046380.27575.a5","article-title":"\u201cMultimodal video indexing: A review of the state-of-the-art\u201d","volume":"25","author":"Snoek","year":"2005","journal-title":"Multimedia Tools and Applications"},{"key":"2026040314322783500_ref222","doi-asserted-by":"crossref","first-page":"86","DOI":"10.1109\/MMUL.2008.21","article-title":"\u201cVideOlympics: Real-time evaluation of multimedia retrieval systems\u201d","volume":"15","author":"Snoek","year":"2008","journal-title":"IEEE MultiMedia"},{"key":"2026040314322783500_ref223","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Snoek","year":"2005"},{"key":"2026040314322783500_ref224","doi-asserted-by":"crossref","first-page":"1678","DOI":"10.1109\/TPAMI.2006.212","article-title":"\u201cThe semantic pathfinder: Using an authoring metaphor for generic multimedia indexing\u201d","volume":"28","author":"Snoek","year":"2006","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref225","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Snoek","year":"2004"},{"key":"2026040314322783500_ref226","doi-asserted-by":"crossref","first-page":"91","DOI":"10.1145\/1142020.1142021","article-title":"\u201cLearning rich semantics from news video archives by style analysis\u201d","volume":"2","author":"Snoek","year":"2006","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"2026040314322783500_ref227","doi-asserted-by":"crossref","first-page":"280","DOI":"10.1109\/TMM.2006.886275","article-title":"\u201cA learned lexicon-driven paradigm for interactive video retrieval\u201d","volume":"9","author":"Snoek","year":"2007","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref228","first-page":"399","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Snoek","year":"2005"},{"key":"2026040314322783500_ref229","first-page":"252","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Snoek","year":"2007"},{"key":"2026040314322783500_ref230","doi-asserted-by":"crossref","first-page":"421","DOI":"10.1145\/1180639.1180727","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Snoek","year":"2006"},{"key":"2026040314322783500_ref231","volume-title":"Proceedings of the TRECVID Workshop","author":"Snoek","year":"2008"},{"key":"2026040314322783500_ref232","doi-asserted-by":"crossref","first-page":"11","DOI":"10.1007\/BF00130487","article-title":"\u201cColor Indexing\u201d","volume":"7","author":"Swain","year":"1991","journal-title":"International Journal of Computer Vision"},{"key":"2026040314322783500_ref233","volume-title":"IEEE International Workshop on Content-based Access of Image and Video Databases, in Conjunction with ICCV\u201998","author":"Szummer","year":"1998"},{"key":"2026040314322783500_ref234","doi-asserted-by":"crossref","first-page":"467","DOI":"10.1016\/0306-4573(92)90005-K","article-title":"\u201cThe pragmatics of information retrieval experimentation, revisited\u201d","volume":"28","author":"Tague-Sutcliffe","year":"1992","journal-title":"Information Processing & Management"},{"key":"2026040314322783500_ref235","volume-title":"Proceedings of the TRECVID Workshop","author":"Tang","year":"2008"},{"key":"2026040314322783500_ref236","doi-asserted-by":"crossref","first-page":"103","DOI":"10.1109\/TMM.2003.819783","article-title":"\u201cViBE: A compressed video database structured for active browsing and search\u201d","volume":"6","author":"Taskiran","year":"2004","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref237","doi-asserted-by":"crossref","first-page":"34","DOI":"10.1109\/MMUL.1994.318984","article-title":"\u201cStructured video computing\u201d","volume":"1","author":"Tonomura","year":"1994","journal-title":"IEEE MultiMedia"},{"key":"2026040314322783500_ref238","doi-asserted-by":"crossref","first-page":"1958","DOI":"10.1109\/TPAMI.2008.128","article-title":"\u201c80 million tiny images: A large data set for nonparametric object and scene recognition\u201d","volume":"30","author":"Torralba","year":"2008","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref239","doi-asserted-by":"crossref","DOI":"10.1145\/1198302.1198305","article-title":"\u201cVideo abstraction: A systematic review and classification\u201d","volume":"3","author":"Truong","year":"2007","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"2026040314322783500_ref240","first-page":"535","volume-title":"Proceedings of the IEEE International Conference on Image Processing","author":"Tseng","year":"2003"},{"key":"2026040314322783500_ref241","doi-asserted-by":"crossref","first-page":"177","DOI":"10.1561\/0600000017","article-title":"\u201cLocal invariant feature detectors: A survey\u201d","volume":"3","author":"Tuytelaars","year":"2008","journal-title":"Foundations and Trends in Computer Graphics and Vision"},{"key":"2026040314322783500_ref242","volume-title":"Proceedings of the TRECVID Workshop","author":"Urban","year":"2006"},{"key":"2026040314322783500_ref243","doi-asserted-by":"crossref","first-page":"117","DOI":"10.1109\/83.892448","article-title":"\u201cImage classification for content-based indexing\u201d","volume":"10","author":"Vailaya","year":"2001","journal-title":"IEEE Transactions on Image Processing"},{"key":"2026040314322783500_ref244","doi-asserted-by":"crossref","first-page":"1921","DOI":"10.1016\/S0031-3203(98)00079-X","article-title":"\u201cOn image classification: City images vs. landscapes\u201d","volume":"31","author":"Vailaya","year":"1998","journal-title":"Pattern Recognition"},{"key":"2026040314322783500_ref245","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"van de Sande","year":"2008"},{"key":"2026040314322783500_ref246","volume-title":"European Conference on Computer Vision","author":"van Gemert","year":"2008"},{"key":"2026040314322783500_ref247","volume-title":"International Workshop on Semantic Learning Applications in Multimedia, in Conjunction with CVPR\u201906","author":"van Gemert","year":"2006"},{"key":"2026040314322783500_ref248","doi-asserted-by":"crossref","first-page":"695","DOI":"10.1145\/1180639.1180786","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"van Gemert","year":"2006"},{"key":"2026040314322783500_ref249","doi-asserted-by":"crossref","DOI":"10.1007\/978-1-4757-3264-1","volume-title":"The Nature of Statistical Learning Theory","author":"Vapnik","year":"2000"},{"key":"2026040314322783500_ref250","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1007\/978-1-4471-3702-3_4","volume-title":"Principles of Visual Information Retrieval","author":"Veltkamp","year":"2001"},{"key":"2026040314322783500_ref251","first-page":"892","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Volkmer","year":"2005"},{"key":"2026040314322783500_ref252","doi-asserted-by":"crossref","first-page":"967","DOI":"10.1109\/TMM.2007.900153","article-title":"\u201cModelling human judgement of digital imagery for multimedia retrieval\u201d","volume":"9","author":"Volkmer","year":"2007","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref253","doi-asserted-by":"crossref","first-page":"92","DOI":"10.1109\/MC.2006.196","article-title":"\u201cGames with a purpose\u201d","volume":"39","author":"von Ahn","year":"2006","journal-title":"IEEE Computer"},{"key":"2026040314322783500_ref254","volume-title":"TREC: Experiment and Evaluation in Information Retrieval","author":"Voorhees","year":"2005"},{"key":"2026040314322783500_ref255","doi-asserted-by":"crossref","first-page":"66","DOI":"10.1109\/2.745722","article-title":"\u201cLessons learned from building a terabyte digital video library\u201d","volume":"32","author":"Wactlar","year":"1999","journal-title":"IEEE Computer"},{"key":"2026040314322783500_ref256","doi-asserted-by":"crossref","first-page":"285","DOI":"10.1145\/1291233.1291293","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Wang","year":"2007"},{"key":"2026040314322783500_ref257","first-page":"61","volume-title":"Proceedings of the ACM SIGMM International Workshop on Multimedia Information Retrieval","author":"Wang","year":"2007"},{"key":"2026040314322783500_ref258","doi-asserted-by":"crossref","first-page":"947","DOI":"10.1109\/34.955109","article-title":"\u201cSIMPLIcity: Semantics-sensitive integrated matching for picture libraries\u201d","volume":"23","author":"Wang","year":"2001","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"2026040314322783500_ref259","doi-asserted-by":"crossref","first-page":"862","DOI":"10.1145\/1291233.1291431","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Wang","year":"2007"},{"key":"2026040314322783500_ref260","first-page":"1483","volume-title":"Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition","author":"Wang","year":"2006"},{"key":"2026040314322783500_ref261","doi-asserted-by":"crossref","first-page":"12","DOI":"10.1109\/79.888862","article-title":"\u201cMultimedia content analysis using both audio and visual clues\u201d","volume":"17","author":"Wang","year":"2000","journal-title":"IEEE Signal Processing Magazine"},{"key":"2026040314322783500_ref262","doi-asserted-by":"crossref","first-page":"1085","DOI":"10.1109\/TMM.2008.2001382","article-title":"\u201cSelection of concept detectors for video search by ontology-enriched semantic spaces\u201d","volume":"10","author":"Wei","year":"2008","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref263","volume-title":"Everything is Miscellaneous","author":"Weinberger","year":"2007"},{"key":"2026040314322783500_ref264","doi-asserted-by":"crossref","first-page":"71","DOI":"10.1145\/1459359.1459370","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Weng","year":"2008"},{"key":"2026040314322783500_ref265","volume-title":"Technical Report","author":"Westerveld","year":"2004"},{"key":"2026040314322783500_ref266","first-page":"344","volume-title":"CIVR","author":"Westerveld","year":"2004"},{"issue":"8","key":"2026040314322783500_ref267","first-page":"635","article-title":"\u201cInexpensive fusion methods for enhancing feature detection\u201d","volume":"7","author":"Wilkins","year":"2007","journal-title":"Image Communication"},{"key":"2026040314322783500_ref268","first-page":"555","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Wilkins","year":"2008"},{"key":"2026040314322783500_ref269","volume-title":"Proceedings of the TRECVID Workshop","author":"Wilkins","year":"2008"},{"key":"2026040314322783500_ref270","doi-asserted-by":"crossref","first-page":"909","DOI":"10.1109\/TMM.2007.898913","article-title":"\u201cSemantic image and video indexing in broad domains\u201d","volume":"9","author":"Worring","year":"2007","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref271","doi-asserted-by":"crossref","first-page":"418","DOI":"10.1016\/j.cviu.2007.09.015","article-title":"\u201cNovelty and redundancy detection with multimodalities in cross-lingual broadcast domain\u201d","volume":"110","author":"Wu","year":"2008","journal-title":"Computer Vision and Image Understanding"},{"key":"2026040314322783500_ref272","first-page":"572","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Wu","year":"2004"},{"key":"2026040314322783500_ref273","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Wu","year":"2004"},{"key":"2026040314322783500_ref274","first-page":"297","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Xie","year":"2006"},{"key":"2026040314322783500_ref275","doi-asserted-by":"crossref","first-page":"623","DOI":"10.1109\/JPROC.2008.916362","article-title":"\u201cEvent mining in multimedia streams\u201d","volume":"96","author":"Xie","year":"2008","journal-title":"Proceedings of the IEEE"},{"key":"2026040314322783500_ref276","doi-asserted-by":"crossref","first-page":"589","DOI":"10.1109\/JPROC.2008.916351","article-title":"\u201cMobile search with multimodal queries\u201d","volume":"96","author":"Xie","year":"2008","journal-title":"Proceedings of the IEEE"},{"key":"2026040314322783500_ref277","doi-asserted-by":"crossref","first-page":"421","DOI":"10.1109\/TMM.2008.917346","article-title":"\u201cA novel framework for semantic annotation and personalized retrieval of sports video\u201d","volume":"10","author":"Xu","year":"2008","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref278","doi-asserted-by":"crossref","first-page":"1342","DOI":"10.1109\/TMM.2008.2004912","article-title":"\u201cUsing webcast text for semantic event detection in broadcast sports video\u201d","volume":"10","author":"Xu","year":"2008","journal-title":"IEEE Transactions on Multimedia"},{"key":"2026040314322783500_ref279","doi-asserted-by":"crossref","first-page":"44","DOI":"10.1145\/1126004.1126007","article-title":"\u201cFusion of AV features and external information sources for event detection in team sports video\u201d","volume":"2","author":"Xu","year":"2006","journal-title":"ACM Transactions on Multimedia Computing, Communications and Applications"},{"key":"2026040314322783500_ref280","first-page":"301","volume-title":"Proceedings of the IEEE International Conference on Multimedia & Expo","author":"Yan","year":"2006"},{"key":"2026040314322783500_ref281","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Yan","year":"2003"},{"key":"2026040314322783500_ref282","first-page":"324","volume-title":"Proceedings of the ACM SIGIR International Conference on Research and Development in Information Retrieval","author":"Yan","year":"2006"},{"issue":"4-5","key":"2026040314322783500_ref283","doi-asserted-by":"crossref","first-page":"445","DOI":"10.1007\/s10791-007-9031-y","article-title":"\u201cA review of text and image retrieval approaches for broadcast news video\u201d","volume":"10","author":"Yan","year":"2007","journal-title":"Information Retrieval"},{"key":"2026040314322783500_ref284","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Yan","year":"2003"},{"key":"2026040314322783500_ref285","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Yan","year":"2004"},{"key":"2026040314322783500_ref286","volume-title":"Technical Report 222-2006-8","author":"Yanagawa","year":"2007"},{"key":"2026040314322783500_ref287","first-page":"270","volume-title":"CIVR","author":"Yang","year":"2004"},{"key":"2026040314322783500_ref288","first-page":"85","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Yang","year":"2008"},{"key":"2026040314322783500_ref289","doi-asserted-by":"crossref","first-page":"188","DOI":"10.1145\/1291233.1291276","volume-title":"Proceedings of the ACM International Conference on Multimedia","author":"Yang","year":"2007"},{"key":"2026040314322783500_ref290","doi-asserted-by":"crossref","first-page":"173","DOI":"10.1007\/s10115-007-0101-7","article-title":"\u201cEstimating average precision when judgments are incomplete\u201d","volume":"16","author":"Yilmaz","year":"2008","journal-title":"Knowledge and Information Systems"},{"key":"2026040314322783500_ref291","volume-title":"Proceedings of the TRECVID Workshop","author":"Yuan","year":"2005"},{"key":"2026040314322783500_ref292","doi-asserted-by":"crossref","first-page":"168","DOI":"10.1109\/TCSVT.2006.888023","article-title":"\u201cA formal study of shot boundary detection\u201d","volume":"17","author":"Yuan","year":"2007","journal-title":"IEEE Transactions on Circuits and Systems for Video Technology"},{"key":"2026040314322783500_ref293","volume-title":"Proceedings of the TRECVID Workshop","author":"Yuan","year":"2007"},{"key":"2026040314322783500_ref294","first-page":"237","volume-title":"Proceedings of the ACM International Conference on Multimedia Information Retrieval","author":"Zavesky","year":"2008"},{"key":"2026040314322783500_ref295","first-page":"617","volume-title":"Proceedings of the ACM International Conference on Image and Video Retrieval","author":"Zavesky","year":"2008"},{"key":"2026040314322783500_ref296","first-page":"227","volume-title":"Proceedings of the ACM SIGMM International Workshop on Multimedia Information Retrieval","author":"Zha","year":"2007"},{"key":"2026040314322783500_ref297","doi-asserted-by":"crossref","first-page":"10","DOI":"10.1007\/BF01210504","article-title":"\u201cAutomatic partitioning of full-motion video\u201d","volume":"1","author":"Zhang","year":"1993","journal-title":"Multimedia Systems"},{"key":"2026040314322783500_ref298","doi-asserted-by":"crossref","first-page":"256","DOI":"10.1007\/BF01225243","article-title":"\u201cAutomatic parsing and indexing of news video\u201d","volume":"2","author":"Zhang","year":"1995","journal-title":"Multimedia Systems"},{"key":"2026040314322783500_ref299","doi-asserted-by":"crossref","first-page":"213","DOI":"10.1007\/s11263-006-9794-4","article-title":"\u201cLocal features and kernels for classification of texture and object categories: A comprehensive study\u201d","volume":"73","author":"Zhang","year":"2007","journal-title":"International Journal of Computer Vision"},{"key":"2026040314322783500_ref300","first-page":"193","volume-title":"Proceedings of the ACM SIGMM International Workshop on Multimedia Information Retrieval","author":"Zhang","year":"2006"},{"key":"2026040314322783500_ref301","article-title":"\u201cLIP-VIREO: Local interest point extraction toolkit\u201d","author":"Zhao","year":"2008"},{"key":"2026040314322783500_ref302","first-page":"536","article-title":"\u201cRelevance feedback in image retrieval: A comprehensive review\u201d","volume":"8","author":"Zhou","year":"2003","journal-title":"Multimedia Tools and Applications"},{"key":"2026040314322783500_ref303","volume-title":"Technical Report 1530","author":"Zhu","year":"2005"}],"container-title":["Foundations and Trends\u00ae in Information Retrieval"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.emerald.com\/ftinr\/article-pdf\/2\/4\/215\/11047834\/1500000014en.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"syndication"},{"URL":"https:\/\/www.emerald.com\/ftinr\/article-pdf\/2\/4\/215\/11047834\/1500000014en.pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,4,29]],"date-time":"2026-04-29T14:33:17Z","timestamp":1777473197000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.emerald.com\/ftinr\/article\/2\/4\/215\/1328659\/Concept-Based-Video-Retrieval"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2009,4,28]]},"references-count":303,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2009,4,28]]}},"URL":"https:\/\/doi.org\/10.1561\/1500000014","relation":{},"ISSN":["1554-0669","1554-0677"],"issn-type":[{"value":"1554-0669","type":"print"},{"value":"1554-0677","type":"electronic"}],"subject":[],"published":{"date-parts":[[2009,4,28]]}}}