{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,10]],"date-time":"2025-10-10T01:28:05Z","timestamp":1760059685906,"version":"build-2065373602"},"reference-count":97,"publisher":"MDPI AG","issue":"7","license":[{"start":{"date-parts":[[2025,6,30]],"date-time":"2025-06-30T00:00:00Z","timestamp":1751241600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["MTI"],"abstract":"<jats:p>Retrieving specific, often instantaneous, content from hours-long egocentric video footage based on hazily remembered details is challenging. Vision\u2013language models (VLMs) have been employed to enable zero-shot textual-based content retrieval from videos. But, they fall short if the textual query contains ambiguous terms or users fail to specify their queries enough, leading to vague semantic queries. Such queries can refer to several different video moments, not all of which can be relevant, making pinpointing content harder. We investigate the requirements for an egocentric video content retrieval framework that helps users handle vague queries. First, we narrow down vague query formulation factors and limit them to ambiguity and incompleteness. Second, we propose a zero-shot, user-centered video content retrieval framework that leverages a VLM to provide video data and query representations that users can incrementally combine to refine queries. Third, we compare our proposed framework to a baseline video player and analyze user strategies for answering vague video content retrieval scenarios in an experimental study. We report that both frameworks perform similarly, users favor our proposed framework, and, as far as navigation strategies go, users value classic interactions when initiating their search and rely on the abstract semantic video representation to refine their resulting moments.<\/jats:p>","DOI":"10.3390\/mti9070066","type":"journal-article","created":{"date-parts":[[2025,7,1]],"date-time":"2025-07-01T07:42:06Z","timestamp":1751355726000},"page":"66","update-policy":"https:\/\/doi.org\/10.3390\/mdpi_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Interactive Content Retrieval in Egocentric Videos Based on Vague Semantic Queries"],"prefix":"10.3390","volume":"9","author":[{"ORCID":"https:\/\/orcid.org\/0009-0002-8835-1376","authenticated-orcid":false,"given":"Linda","family":"Ablaoui","sequence":"first","affiliation":[{"name":"Ecole Nationale de l\u2019aviation Civile, 7 Avenue Edouard Belin, CS 54005, CEDEX 4, 31055 Toulouse, France"},{"name":"CNRS, IPAL, 15 Computing Drive, Singapore 117418, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-8580-2779","authenticated-orcid":false,"given":"Wilson Estecio","family":"Marcilio-Jr","sequence":"additional","affiliation":[{"name":"Faculty of Sciences and Technology, S\u00e3o Paulo State University (UNESP), Presidente Prudente 19060-900, SP, Brazil"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-5457-6289","authenticated-orcid":false,"given":"Lai Xing","family":"Ng","sequence":"additional","affiliation":[{"name":"CNRS, IPAL, 15 Computing Drive, Singapore 117418, Singapore"},{"name":"Institute for Infocomm Research, A*STAR, 1 Fusionopolis Way, #21-01, Connexis South Tower, Singapore 138632, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-0768-1019","authenticated-orcid":false,"given":"Christophe","family":"Jouffrais","sequence":"additional","affiliation":[{"name":"CNRS, IPAL, 15 Computing Drive, Singapore 117418, Singapore"},{"name":"Institut de Recherche en Informatique de Toulouse (IRIT), Centre National de la Recherche Scientifique (CNRS), 31062 Toulouse, France"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4318-6717","authenticated-orcid":false,"given":"Christophe","family":"Hurter","sequence":"additional","affiliation":[{"name":"Ecole Nationale de l\u2019aviation Civile, 7 Avenue Edouard Belin, CS 54005, CEDEX 4, 31055 Toulouse, France"},{"name":"CNRS, IPAL, 15 Computing Drive, Singapore 117418, Singapore"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1968","published-online":{"date-parts":[[2025,6,30]]},"reference":[{"key":"ref_1","doi-asserted-by":"crossref","first-page":"103252","DOI":"10.1016\/j.cviu.2021.103252","article-title":"Predicting the future from first person (egocentric) vision: A survey","volume":"211","author":"Rodin","year":"2021","journal-title":"Comput. Vis. Image Underst."},{"key":"ref_2","first-page":"1","article-title":"Egocentric hand detection via dynamic region growing","volume":"14","author":"Huang","year":"2017","journal-title":"ACM Trans. Multimed. Comput. Commun. Appl. (TOMM)"},{"key":"ref_3","doi-asserted-by":"crossref","unstructured":"Thakur, S.K., Beyan, C., Morerio, P., and Del Bue, A. (2021, January 18\u201322). Predicting gaze from egocentric social interaction videos and imu data. Proceedings of the 2021 International Conference on Multimodal Interaction, Ottawa, ON, Canada.","DOI":"10.1145\/3462244.3479954"},{"key":"ref_4","doi-asserted-by":"crossref","unstructured":"Legel, M., Deckers, S.R., Soto, G., Grove, N., Waller, A., Balkom, H.v., Spanjers, R., Norrie, C.S., and Steenbergen, B. (2025). Self-Created Film as a Resource in a Multimodal Conversational Narrative. Multimodal Technol. Interact., 9.","DOI":"10.3390\/mti9030025"},{"key":"ref_5","doi-asserted-by":"crossref","first-page":"10","DOI":"10.54254\/2753-7102\/2024.18644","article-title":"From Virality to Engagement: Examining the Transformative Impact of Social Media, Short Video Platforms, and Live Streaming on Information Dissemination and Audience Behavior in the Digital Age","volume":"14","author":"Yin","year":"2024","journal-title":"Adv. Soc. Behav. Res."},{"key":"ref_6","doi-asserted-by":"crossref","unstructured":"Patra, S., Aggarwal, H., Arora, H., Banerjee, S., and Arora, C. (2017, January 24\u201331). Computing Egomotion with Local Loop Closures for Egocentric Videos. Proceedings of the 2017 IEEE Winter Conference on Applications of Computer Vision (WACV), Santa Rosa, CA, USA.","DOI":"10.1109\/WACV.2017.57"},{"key":"ref_7","doi-asserted-by":"crossref","unstructured":"Zhang, L., Zhou, S., Stent, S., and Shi, J. (2022, January 23\u201327). Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications. Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel.","DOI":"10.1007\/978-3-031-19818-2_8"},{"key":"ref_8","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1007\/s11263-021-01531-2","article-title":"Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100","volume":"130","author":"Damen","year":"2022","journal-title":"Int. J. Comput. Vis. (IJCV)"},{"key":"ref_9","doi-asserted-by":"crossref","unstructured":"Mueller, F., Mehta, D., Sotnychenko, O., Sridhar, S., Casas, D., and Theobalt, C. (2017, January 22\u201327). Real-time hand tracking under occlusion from an egocentric rgb-d sensor. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.131"},{"key":"ref_10","doi-asserted-by":"crossref","first-page":"7774","DOI":"10.1109\/TCSVT.2023.3281671","article-title":"First-Person Video Domain Adaptation With Multi-Scene Cross-Site Datasets and Attention-Based Methods","volume":"33","author":"Liu","year":"2023","journal-title":"IEEE Trans. Circuits Syst. Video Technol."},{"key":"ref_11","unstructured":"Amsaleg, L., Guomundsson, G.B., Gurrin, C., J\u00f3nsson, B.B., and Satoh, S. (2017). An Annotation System for Egocentric Image Media. MultiMedia Modeling, Springer International Publishing."},{"key":"ref_12","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2333112.2333120","article-title":"Assistive tagging: A survey of multimedia tagging with human-computer joint exploration","volume":"44","author":"Wang","year":"2012","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"ref_13","first-page":"4051","article-title":"A review of generalized zero-shot learning methods","volume":"45","author":"Pourpanah","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_14","unstructured":"Grauman, K., Westbury, A., Byrne, E., Chavis, Z., Furnari, A., Girdhar, R., Hamburger, J., Jiang, H., Liu, M., and Liu, X. (2022, January 18\u201324). Ego4d: Around the world in 3000 hours of egocentric video. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA."},{"key":"ref_15","doi-asserted-by":"crossref","unstructured":"Zhu, C., Xiao, F., Alvarado, A., Babaei, Y., Hu, J., El-Mohri, H., Culatana, S., Sumbaly, R., and Yan, Z. (2023, January 4\u20136). Egoobjects: A large-scale egocentric dataset for fine-grained object understanding. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01840"},{"key":"ref_16","doi-asserted-by":"crossref","unstructured":"Xue, Z., Song, Y., Grauman, K., and Torresani, L. (2023, January 17\u201324). Egocentric video task translation. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada.","DOI":"10.1109\/CVPR52729.2023.00229"},{"key":"ref_17","doi-asserted-by":"crossref","first-page":"401","DOI":"10.1109\/TMM.2024.3521658","article-title":"GPT4Ego: Unleashing the potential of pre-trained models for zero-shot egocentric action recognition","volume":"27","author":"Dai","year":"2024","journal-title":"IEEE Trans. Multimed."},{"key":"ref_18","first-page":"7575","article-title":"Egocentric video-language pretraining","volume":"35","author":"Lin","year":"2022","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_19","doi-asserted-by":"crossref","unstructured":"Hong, H., Wang, S., Huang, Z., Wu, Q., and Liu, J. (2024, January 3\u20139). Why only text: Empowering vision-and-language navigation with multi-modal prompts. Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, Jeju, Republic of Korea.","DOI":"10.24963\/ijcai.2024\/93"},{"key":"ref_20","unstructured":"Zhang, Y., Li, Y., Cui, L., Cai, D., Liu, L., Fu, T., Huang, X., Zhao, E., Zhang, Y., and Chen, Y. (2023). Siren\u2019s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. arXiv."},{"key":"ref_21","doi-asserted-by":"crossref","unstructured":"Esfandiarpoor, R., Menghini, C., and Bach, S. (2024, January 12\u201316). If CLIP Could Talk: Understanding Vision-Language Model Representations Through Their Preferred Concept Descriptions. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA.","DOI":"10.18653\/v1\/2024.emnlp-main.547"},{"key":"ref_22","doi-asserted-by":"crossref","first-page":"964","DOI":"10.1145\/32206.32212","article-title":"The vocabulary problem in human-system communication","volume":"30","author":"Furnas","year":"1987","journal-title":"Commun. ACM"},{"key":"ref_23","doi-asserted-by":"crossref","unstructured":"Song, R., Luo, Z., Wen, J.R., Yu, Y., and Hon, H.W. (2007, January 8\u201312). Identifying ambiguous queries in web search. Proceedings of the 16th International Conference on World Wide Web, New York, NY, USA.","DOI":"10.1145\/1242572.1242749"},{"key":"ref_24","doi-asserted-by":"crossref","unstructured":"Asai, A., and Choi, E. (2021, January 1\u20136). Challenges in information-seeking QA: Unanswerable questions and paragraph retrieval. Proceedings of the ACL-IJCNLP 2021\u201459th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, Online.","DOI":"10.18653\/v1\/2021.acl-long.118"},{"key":"ref_25","doi-asserted-by":"crossref","first-page":"5719","DOI":"10.1007\/s11042-015-2537-1","article-title":"Fuzzy reasoning framework to improve semantic video interpretation","volume":"75","author":"Zarka","year":"2016","journal-title":"Multimed. Tools Appl."},{"key":"ref_26","doi-asserted-by":"crossref","unstructured":"Chae, J., and Kim, J. (2022, January 18\u201323). Uncertainty-based Visual Question Answering: Estimating Semantic Inconsistency between Image and Knowledge Base. Proceedings of the 2022 International Joint Conference on Neural Networks (IJCNN), Pajua, Italy.","DOI":"10.1109\/IJCNN55064.2022.9892787"},{"key":"ref_27","doi-asserted-by":"crossref","unstructured":"Chang, M., Huh, M., and Kim, J. (2021, January 8\u201313). Rubyslippers: Supporting content-based voice navigation for how-to videos. Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, Yokohama, Japan.","DOI":"10.1145\/3411764.3445131"},{"key":"ref_28","doi-asserted-by":"crossref","unstructured":"P\u00e9rez Ortiz, M., Bulathwela, S., Dormann, C., Verma, M., Kreitmayer, S., Noss, R., Shawe-Taylor, J., Rogers, Y., and Yilmaz, E. (2022, January 14\u201318). Watch less and uncover more: Could navigation tools help users search and explore videos?. Proceedings of the 2022 Conference on Human Information Interaction and Retrieval, New York, NY, USA.","DOI":"10.1145\/3498366.3505814"},{"key":"ref_29","doi-asserted-by":"crossref","first-page":"1354","DOI":"10.1109\/TVCG.2024.3361001","article-title":"Conceptthread: Visualizing threaded concepts in MOOC videos","volume":"31","author":"Zhou","year":"2024","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"ref_30","unstructured":"Fong, M., Miller, G., Zhang, X., Roll, I., Hendricks, C., and Fels, S.S. (2016, January 1\u20133). An Investigation of Textbook-Style Highlighting for Video. Proceedings of the Graphics Interface, Victoria, BC, Canada."},{"key":"ref_31","doi-asserted-by":"crossref","unstructured":"Dragicevic, P., Ramos, G., Bibliowitcz, J., Nowrouzezahrai, D., Balakrishnan, R., and Singh, K. (2008, January 5\u201310). Video browsing by direct manipulation. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, New York, NY, USA.","DOI":"10.1145\/1357054.1357096"},{"key":"ref_32","doi-asserted-by":"crossref","unstructured":"Clarke, C., Cavdir, D., Chiu, P., Denoue, L., and Kimber, D. (2020, January 20\u201323). Reactive video: Adaptive video playback based on user motion for supporting physical activity. Proceedings of the 33rd Annual ACM Symposium on User Interface Software and Technology, New York, NY, USA.","DOI":"10.1145\/3379337.3415591"},{"key":"ref_33","doi-asserted-by":"crossref","unstructured":"Lilija, K., Pohl, H., and Hornb\u00e6k, K. (2020, January 25\u201330). Who put that there? temporal navigation of spatial recordings by direct manipulation. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA.","DOI":"10.1145\/3313831.3376604"},{"key":"ref_34","doi-asserted-by":"crossref","first-page":"65","DOI":"10.1109\/TVCG.2018.2865041","article-title":"Forvizor: Visualizing spatio-temporal team formations in soccer","volume":"25","author":"Wu","year":"2018","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"ref_35","doi-asserted-by":"crossref","first-page":"981","DOI":"10.1109\/TVCG.2019.2934280","article-title":"Motion Browser: Visualizing and understanding complex upper limb movement under obstetrical brachial plexus injuries","volume":"26","author":"Chan","year":"2019","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"ref_36","doi-asserted-by":"crossref","first-page":"87","DOI":"10.1109\/TVCG.2023.3326586","article-title":"Videopro: A visual analytics approach for interactive video programming","volume":"30","author":"He","year":"2023","journal-title":"IEEE Trans. Vis. Comput. Graph."},{"key":"ref_37","doi-asserted-by":"crossref","unstructured":"Xian, Y., Schiele, B., and Akata, Z. (2017, January 21\u201326). Zero-shot learning-the good, the bad and the ugly. Proceedings of the IEEE conference on computer vision and pattern recognition, Honolulu, HI, USA.","DOI":"10.1109\/CVPR.2017.328"},{"key":"ref_38","unstructured":"Yang, A., Miech, A., Sivic, J., Laptev, I., and Schmid, C. (December, January 28). Zero-Shot Video Question Answering via Frozen Bidirectional Language Models Cross-modal Training. Proceedings of the 36th Conference on Neural Information Processing Systems (NeurIPS 2022), New Orleans, LA, USA."},{"key":"ref_39","doi-asserted-by":"crossref","first-page":"3202","DOI":"10.1109\/TPAMI.2022.3173208","article-title":"Learning to Answer Visual Questions from Web Videos","volume":"47","author":"Yang","year":"2022","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."},{"key":"ref_40","unstructured":"Wang, Y., Li, K., Li, X., Yu, J., He, Y., Chen, G., Pei, B., Zheng, R., Wang, Z., and Shi, Y. (October, January 29). InternVideo2: Scaling Foundation Models for Multimodal Video Understanding. Proceedings of the European Conference on Computer Vision, Milan, Italy."},{"key":"ref_41","doi-asserted-by":"crossref","unstructured":"Ma, K., Zang, X., Feng, Z., Fang, H., Ban, C., Wei, Y., He, Z., Li, Y., and Sun, H. (2023, January 4\u20136). LLaViLo: Boosting Video Moment Retrieval via Adapter-Based Multimodal Modeling. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCVW60793.2023.00297"},{"key":"ref_42","doi-asserted-by":"crossref","unstructured":"Panta, L., Shrestha, P., Sapkota, B., Bhattarai, A., Manandhar, S., and Sah, A.K. (2024, January 4\u20138). Cross-modal Contrastive Learning with Asymmetric Co-attention Network for Video Moment Retrieval. Proceedings of the 2024 IEEE\/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW), Waikoloa, HI, USA.","DOI":"10.1109\/WACVW60836.2024.00071"},{"key":"ref_43","doi-asserted-by":"crossref","unstructured":"Luo, D., Huang, J., Gong, S., Jin, H., and Liu, Y. (2024, January 4\u20138). Zero-Shot Video Moment Retrieval from Frozen Vision-Language Models. Proceedings of the 2024 IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA.","DOI":"10.1109\/WACV57701.2024.00538"},{"key":"ref_44","unstructured":"Jiang, X., Zhou, Z., Xu, X., Yang, Y., Wang, G., and Shen, H.T. (November, January 29). Faster video moment retrieval with point-level supervision. Proceedings of the 31st ACM International Conference on Multimedia, Ottawa, ON, Canada."},{"key":"ref_45","unstructured":"Ma, Y., Qing, L., Li, G., Qi, Y., Sheng, Q.Z., and Huang, Q. (2024). Retrieval Enhanced Zero-Shot Video Captioning. arXiv."},{"key":"ref_46","doi-asserted-by":"crossref","unstructured":"Wu, G., Lin, J., and Silva, C.T. (2022, January 19\u201324). IntentVizor: Towards Generic Query Guided Interactive Video Summarization. Proceedings of the 2022 IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.01025"},{"key":"ref_47","doi-asserted-by":"crossref","unstructured":"Pang, Z., Nakashima, Y., Otani, M., and Nagahara, H. (2024). Unleashing the Power of Contrastive Learning for Zero-Shot Video Summarization. J. Imaging, 10.","DOI":"10.3390\/jimaging10090229"},{"key":"ref_48","unstructured":"Ahmad, S., Chanda, S., and Rawat, Y.S. (2023). Ez-clip: Efficient zeroshot video action recognition. arXiv."},{"key":"ref_49","unstructured":"Yu, Y., Cao, C., Zhang, Y., Lv, Q., Min, L., and Zhang, Y. (March, January 25). Building a Multi-modal Spatiotemporal Expert for Zero-shot Action Recognition with CLIP. Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA."},{"key":"ref_50","doi-asserted-by":"crossref","first-page":"3232","DOI":"10.1007\/s11263-024-02024-8","article-title":"Adaptive multi-source predictor for zero-shot video object segmentation","volume":"132","author":"Zhao","year":"2024","journal-title":"Int. J. Comput. Vis."},{"key":"ref_51","doi-asserted-by":"crossref","unstructured":"Chu, W.H., Harley, A.W., Tokmakov, P., Dave, A., Guibas, L., and Fragkiadaki, K. (2024, January 13\u201317). Zero-Shot Open-Vocabulary Tracking with Large Pre-Trained Models. Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan.","DOI":"10.1109\/ICRA57147.2024.10611726"},{"key":"ref_52","unstructured":"Radford, A., Kim, J.W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., and Clark, J. (2021, January 18\u201324). Learning Transferable Visual Models From Natural Language Supervision. Proceedings of the 38th International Conference on Machine Learning, Westminster, UK. PMLR, Proceedings of Machine Learning Research."},{"key":"ref_53","first-page":"75716","article-title":"Egotracks: A long-term egocentric visual object tracking dataset","volume":"36","author":"Tang","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_54","doi-asserted-by":"crossref","unstructured":"Akiva, P., Huang, J., Liang, K.J., Kovvuri, R., Chen, X., Feiszli, M., Dana, K., and Hassner, T. (2023, January 4\u20136). Self-supervised object detection from egocentric videos. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00482"},{"key":"ref_55","doi-asserted-by":"crossref","unstructured":"Hao, S., Chai, W., Zhao, Z., Sun, M., Hu, W., Zhou, J., Zhao, Y., Li, Q., Wang, Y., and Li, X. (November, January 28). Ego3DT: Tracking Every 3D Object in Ego-centric Videos. Proceedings of the MM 2024 32nd ACM International Conference on Multimedia. Association for Computing Machinery, Melbourne, Australia.","DOI":"10.1145\/3664647.3680679"},{"key":"ref_56","unstructured":"Moreno-D\u00edaz, R., Pichler, F., and Quesada-Arencibia, A. (2018). Detecting Hands in Egocentric Videos: Towards Action Recognition. Computer Aided Systems Theory\u2014EUROCAST 2017, Springer International Publishing."},{"key":"ref_57","doi-asserted-by":"crossref","unstructured":"Liu, S., Tripathi, S., Majumdar, S., and Wang, X. (2022, January 18\u201324). Joint hand motion and interaction hotspots prediction from egocentric videos. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPR52688.2022.00328"},{"key":"ref_58","unstructured":"Xu, B., Wang, Z., Du, Y., Song, Z., Zheng, S., and Jin, Q. Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions? In Proceedings of The Thirteenth International Conference on Learning Representations, Singapore, 24\u201328 April 2025."},{"key":"ref_59","doi-asserted-by":"crossref","unstructured":"Xu, Y., Li, Y.L., Huang, Z., Liu, M.X., Lu, C., Tai, Y.W., and Tang, C.K. (2023, January 4\u20136). Egopca: A new framework for egocentric hand-object interaction understanding. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00486"},{"key":"ref_60","doi-asserted-by":"crossref","unstructured":"Ng, E., Xiang, D., Joo, H., and Grauman, K. (2020, January 13\u201319). You2me: Inferring body pose in egocentric video via first and second person interactions. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR42600.2020.00991"},{"key":"ref_61","doi-asserted-by":"crossref","unstructured":"Khirodkar, R., Bansal, A., Ma, L., Newcombe, R., Vo, M., and Kitani, K. (2023, January 4\u20136). Ego-humans: An ego-centric 3d multi-human benchmark. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.01814"},{"key":"ref_62","doi-asserted-by":"crossref","unstructured":"Kuehne, H., Jhuang, H., Garrote, E., Poggio, T., and Serre, T. (2011, January 6\u201313). HMDB: A large video database for human motion recognition. Proceedings of the IEEE 2011 International Conference on Computer Vision, Barcelona, Spain.","DOI":"10.1109\/ICCV.2011.6126543"},{"key":"ref_63","doi-asserted-by":"crossref","first-page":"1366","DOI":"10.1007\/s11263-022-01594-9","article-title":"Human action recognition and prediction: A survey","volume":"130","author":"Kong","year":"2022","journal-title":"Int. J. Comput. Vis."},{"key":"ref_64","unstructured":"Abreu, S., Do, T.D., Ahuja, K., Gonzalez, E.J., Payne, L., McDuff, D., and Gonzalez-Franco, M. (2024). Parse-ego4d: Personal action recommendation suggestions for egocentric videos. arXiv."},{"key":"ref_65","doi-asserted-by":"crossref","unstructured":"Fan, C. (2019, January 27\u201328). Egovqa-an egocentric video question answering benchmark dataset. Proceedings of the IEEE\/CVF International Conference on Computer Vision Workshops, Seoul, Republic of Korea.","DOI":"10.1109\/ICCVW.2019.00536"},{"key":"ref_66","doi-asserted-by":"crossref","unstructured":"B\u00e4rmann, L., and Waibel, A. (2022, January 18\u201324). Where did i leave my keys?-episodic-memory-based question answering on egocentric videos. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA.","DOI":"10.1109\/CVPRW56347.2022.00162"},{"key":"ref_67","doi-asserted-by":"crossref","unstructured":"Di, S., and Xie, W. (2024, January 17\u201321). Grounded question-answering in long egocentric videos. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01229"},{"key":"ref_68","first-page":"24143","article-title":"Single-stage visual query localization in egocentric videos","volume":"36","author":"Jiang","year":"2023","journal-title":"Adv. Neural Inf. Process. Syst."},{"key":"ref_69","doi-asserted-by":"crossref","unstructured":"Pramanick, S., Song, Y., Nag, S., Lin, K.Q., Shah, H., Shou, M.Z., Chellappa, R., and Zhang, P. (2023, January 4\u20136). Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00487"},{"key":"ref_70","unstructured":"Hummel, T., Karthik, S., Georgescu, M.I., and Akata, Z. (October, January 29). Egocvr: An egocentric benchmark for fine-grained composed video retrieval. Proceedings of the European Conference on Computer Vision, Milan, Italy."},{"key":"ref_71","doi-asserted-by":"crossref","unstructured":"Sanderson, M. (2008, January 20\u201324). Ambiguous queries: Test collections need more sense. Proceedings of the 31st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, Singapore.","DOI":"10.1145\/1390334.1390420"},{"key":"ref_72","doi-asserted-by":"crossref","unstructured":"Dhole, K.D., and Agichtein, E. (2024, January 24\u201328). Genqrensemble: Zero-shot llm ensemble prompting for generative query reformulation. Proceedings of the European Conference on Information Retrieval, Glasgow, UK.","DOI":"10.1007\/978-3-031-56063-7_24"},{"key":"ref_73","doi-asserted-by":"crossref","first-page":"76581","DOI":"10.1109\/ACCESS.2023.3295776","article-title":"Information Retrieval: Recent Advances and Beyond","volume":"11","author":"Hambarde","year":"2023","journal-title":"IEEE Access"},{"key":"ref_74","doi-asserted-by":"crossref","unstructured":"Woo, S., Jeon, S.Y., Park, J., Son, M., Lee, S., and Kim, C. (2024, January 4\u20138). Sketch-based video object localization. Proceedings of the IEEE\/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA.","DOI":"10.1109\/WACV57701.2024.00829"},{"key":"ref_75","doi-asserted-by":"crossref","unstructured":"Miyanishi, T., Hirayama, J.i., Kong, Q., Maekawa, T., Moriya, H., and Suyama, T. (2016, January 12\u201317). Egocentric video search via physical interactions. Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA.","DOI":"10.1609\/aaai.v30i1.10009"},{"key":"ref_76","doi-asserted-by":"crossref","unstructured":"Lee, D., Kim, S., Lee, M., Lee, H., Park, J., Lee, S.W., and Jung, K. (2023, January 6\u201310). Asking Clarification Questions to Handle Ambiguity in Open-Domain QA. Proceedings of The 2023 Conference on Empirical Methods in Natural Language Processing, Singapore.","DOI":"10.18653\/v1\/2023.findings-emnlp.772"},{"key":"ref_77","doi-asserted-by":"crossref","unstructured":"Nakano, Y., Kawano, S., Yoshino, K., Sudoh, K., and Nakamura, S. (2022, January 26). Pseudo ambiguous and clarifying questions based on sentence structures toward clarifying question answering system. Proceedings of the Second DialDoc Workshop on Document-grounded Dialogue and Conversational Question Answering, Dublin, Ireland.","DOI":"10.18653\/v1\/2022.dialdoc-1.4"},{"key":"ref_78","doi-asserted-by":"crossref","unstructured":"Gao, J., Sun, C., Yang, Z., and Nevatia, R. (2017, January 22\u201327). Tall: Temporal activity localization via language query. Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy.","DOI":"10.1109\/ICCV.2017.563"},{"key":"ref_79","doi-asserted-by":"crossref","unstructured":"Caba Heilbron, F., Escorcia, V., Ghanem, B., and Carlos Niebles, J. (2015, January 7\u201312). Activitynet: A large-scale video benchmark for human activity understanding. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA.","DOI":"10.1109\/CVPR.2015.7298698"},{"key":"ref_80","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1162\/tacl_a_00207","article-title":"Grounding action descriptions in videos","volume":"1","author":"Regneri","year":"2013","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"ref_81","doi-asserted-by":"crossref","first-page":"69","DOI":"10.1007\/s42113-018-0005-5","article-title":"Do people ask good questions?","volume":"1","author":"Rothe","year":"2018","journal-title":"Comput. Brain Behav."},{"key":"ref_82","unstructured":"Liu, S., Yu, C., and Meng, W. (November, January 31). Word sense disambiguation in queries. Proceedings of the 14th ACM International Conference on Information and Knowledge Management, CIKM \u201905, New York, NY, USA."},{"key":"ref_83","unstructured":"Wang, Y., and Agichtein, E. (2010, January 2\u20134). Query ambiguity revisited: Clickthrough measures for distinguishing informational and ambiguous queries. Proceedings of the Human Language Technologies: The 2010 Annual Conference of the North American Chapter of the Association for Computational Linguistics, Los Angeles, CA, USA."},{"key":"ref_84","doi-asserted-by":"crossref","unstructured":"Aliannejadi, M., Zamani, H., Crestani, F., and Croft, W.B. (2019, January 21\u201325). Asking clarifying questions in open-domain information-seeking conversations. Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval, Paris, France.","DOI":"10.1145\/3331184.3331265"},{"key":"ref_85","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3534965","article-title":"How to approach ambiguous queries in conversational search: A survey of techniques, approaches, tools, and challenges","volume":"55","author":"Keyvan","year":"2022","journal-title":"ACM Comput. Surv."},{"key":"ref_86","doi-asserted-by":"crossref","unstructured":"Hodges, S., Williams, L., Berry, E., Izadi, S., Srinivasan, J., Butler, A., Smyth, G., Kapur, N., and Wood, K. (2006, January 17\u201321). SenseCam: A retrospective memory aid. Proceedings of the UbiComp 2006: Ubiquitous Computing: 8th International Conference, UbiComp 2006, Orange County, CA, USA.","DOI":"10.1007\/11853565_11"},{"key":"ref_87","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1561\/1500000033","article-title":"Lifelogging: Personal big data","volume":"8","author":"Gurrin","year":"2014","journal-title":"Found. Trends\u00ae Inf. Retr."},{"key":"ref_88","doi-asserted-by":"crossref","first-page":"182","DOI":"10.1037\/0003-066X.54.3.182","article-title":"The seven sins of memory: Insights from psychology and cognitive neuroscience","volume":"54","author":"Schacter","year":"1999","journal-title":"Am. Psychol."},{"key":"ref_89","unstructured":"Diwan, A., Peng, P., and Mooney, R. Zero-shot Video Moment Retrieval with Off-the-Shelf Models. Proceedings of the Transfer Learning for Natural Language Processing Workshop, PMLR, Available online: https:\/\/proceedings.mlr.press\/v203\/diwan23a.html."},{"key":"ref_90","doi-asserted-by":"crossref","unstructured":"Sultan, M., Jacobs, L., Stylianou, A., and Pless, R. (2023, January 27\u201329). Exploring CLIP for Real World, Text-based Image Retrieval. Proceedings of the 2023 IEEE Applied Imagery Pattern Recognition Workshop (AIPR), Washington, DC, USA.","DOI":"10.1109\/AIPR60534.2023.10440710"},{"key":"ref_91","first-page":"2579","article-title":"Visualizing data using t-SNE","volume":"9","author":"Hinton","year":"2008","journal-title":"J. Mach. Learn. Res."},{"key":"ref_92","unstructured":"Ester, M., Kriegel, H.P., Sander, J., and Xu, X. (1996, January 2\u20134). A density-based algorithm for discovering clusters in large spatial databases with noise. Proceedings of the KDD, Portland, OR, USA."},{"key":"ref_93","doi-asserted-by":"crossref","unstructured":"Chen, J., Mao, J., Liu, Y., Zhang, F., Zhang, M., and Ma, S. (2021, January 19\u201323). Towards a better understanding of query reformulation behavior in web search. Proceedings of the Web Conference, Ljubljana, Slovenia.","DOI":"10.1145\/3442381.3450127"},{"key":"ref_94","doi-asserted-by":"crossref","unstructured":"Cheng, T., Song, L., Ge, Y., Liu, W., Wang, X., and Shan, Y. (2024, January 17\u201321). Yolo-world: Real-time open-vocabulary object detection. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA.","DOI":"10.1109\/CVPR52733.2024.01599"},{"key":"ref_95","doi-asserted-by":"crossref","unstructured":"Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A.C., and Lo, W.Y. (2023, January 4\u20136). Segment anything. Proceedings of the IEEE\/CVF International Conference on Computer Vision, Paris, France.","DOI":"10.1109\/ICCV51070.2023.00371"},{"key":"ref_96","unstructured":"Huang, X., Gao, Y., He, S., Xie, X., Wang, Y., Zhang, Y., Liu, X., Li, B., and Liu, Y. (2024, January 17\u201321). Exploring Clean-Label Backdoor Attacks and Defense in Language Models. Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA."},{"key":"ref_97","doi-asserted-by":"crossref","first-page":"259","DOI":"10.1016\/0169-7439(89)80095-4","article-title":"Analysis of variance (ANOVA)","volume":"6","author":"St","year":"1989","journal-title":"Chemom. Intell. Lab. Syst."}],"container-title":["Multimodal Technologies and Interaction"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/www.mdpi.com\/2414-4088\/9\/7\/66\/pdf","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,9]],"date-time":"2025-10-09T18:01:44Z","timestamp":1760032904000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.mdpi.com\/2414-4088\/9\/7\/66"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,6,30]]},"references-count":97,"journal-issue":{"issue":"7","published-online":{"date-parts":[[2025,7]]}},"alternative-id":["mti9070066"],"URL":"https:\/\/doi.org\/10.3390\/mti9070066","relation":{},"ISSN":["2414-4088"],"issn-type":[{"type":"electronic","value":"2414-4088"}],"subject":[],"published":{"date-parts":[[2025,6,30]]}}}