{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,3,27]],"date-time":"2025-03-27T08:06:33Z","timestamp":1743062793263,"version":"3.40.3"},"publisher-location":"Singapore","reference-count":28,"publisher":"Springer Nature Singapore","isbn-type":[{"type":"print","value":"9789819620708"},{"type":"electronic","value":"9789819620715"}],"license":[{"start":{"date-parts":[[2025,1,1]],"date-time":"2025-01-01T00:00:00Z","timestamp":1735689600000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2025,1,2]],"date-time":"2025-01-02T00:00:00Z","timestamp":1735776000000},"content-version":"vor","delay-in-days":1,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2025]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper presents a new approach to dense retrieval across multiple modalities, emphasizing the integration of images and sensor data. Traditional cross-modal retrieval techniques\u00a0face significant challenges, particularly in processing non-linguistic modalities and creating effective training datasets. To address these issues, we propose a method that uses a shared vector space, optimized with contrastive loss, to enable efficient and accurate retrieval across diverse modalities. A key innovation of\u00a0our approach is the introduction of a temporal closeness metric,\u00a0which evaluates the relationship between data points based on\u00a0their timestamps. This metric helps automatically extract positive\u00a0and hard negative samples related to the query, improving the training process and enhancing the retrieval model\u2019s performance. We validate our approach using the Lifelog Search Challenge 2024 (LSC\u201924) dataset, one of the largest multi-modal datasets, including non-linguistic data such as egocentric images, heart rate,\u00a0and location information. Our evaluation shows that incorporating temporal closeness into the dense retrieval process significantly improves retrieval accuracy and robustness in real-world, multi-modal scenarios. This paper\u2019s contributions include developing a novel dense retrieval framework, introducing the temporal closeness metric, and successfully applying these innovations to\u00a0a comprehensive multi-modal dataset.<\/jats:p>","DOI":"10.1007\/978-981-96-2071-5_13","type":"book-chapter","created":{"date-parts":[[2025,1,1]],"date-time":"2025-01-01T15:33:59Z","timestamp":1735745639000},"page":"170-183","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":0,"title":["Temporal Closeness for\u00a0Enhanced Cross-Modal Retrieval of\u00a0Sensor and\u00a0Image Data"],"prefix":"10.1007","author":[{"ORCID":"https:\/\/orcid.org\/0009-0001-1003-9107","authenticated-orcid":false,"given":"Shuhei","family":"Yamamoto","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-2133-0215","authenticated-orcid":false,"given":"Noriko","family":"Kando","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"297","published-online":{"date-parts":[[2025,1,2]]},"reference":[{"key":"13_CR1","doi-asserted-by":"crossref","unstructured":"Chang, Y., et al.: A survey on evaluation of large language models. ACM Trans. Intell. Syst. Technol. 15(3) (2024)","DOI":"10.1145\/3641289"},{"key":"13_CR2","unstructured":"Chen, T., Kornblith, S., Norouzi, M., Hinton, G.: A simple framework for contrastive learning of visual representations. In: ICML\u201920, pp. 1597\u20131607. PMLR (2020)"},{"key":"13_CR3","unstructured":"Faghri, F., Fleet, D.J., Kiros, J.R., Fidler, S.: VSE++: improving visual-semantic embeddings with hard negatives. In: BMVC\u201918 (2018)"},{"key":"13_CR4","doi-asserted-by":"crossref","unstructured":"Ge, Y., Zeng, X., Huffman, J.S., Lin, T.Y., Liu, M.Y., Cui, Y.: Visual fact checker: enabling high-fidelity detailed caption generation. In: CVPR\u201924, pp. 14033\u201314042 (2024)","DOI":"10.1109\/CVPR52733.2024.01331"},{"key":"13_CR5","doi-asserted-by":"crossref","unstructured":"Gurrin, C., et al.: Introduction to the seventh annual lifelog search challenge, lsc\u201924. In: ICMR\u201924, pp. 1334\u20141335 (2024)","DOI":"10.1145\/3652583.3658891"},{"key":"13_CR6","doi-asserted-by":"crossref","unstructured":"He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR\u201916, pp. 770\u2013778 (2016)","DOI":"10.1109\/CVPR.2016.90"},{"key":"13_CR7","doi-asserted-by":"crossref","unstructured":"Hu, W., Xu, Y., Li, Y., Li, W., Chen, Z., Tu, Z.: BLIVA: a simple multimodal LLM for better handling of text-rich visual questions. In: AAAI\u201924, pp. 2256\u20132264 (2024)","DOI":"10.1609\/aaai.v38i3.27999"},{"issue":"1","key":"13_CR8","doi-asserted-by":"publisher","first-page":"2","DOI":"10.3390\/technologies9010002","volume":"9","author":"A Jaiswal","year":"2020","unstructured":"Jaiswal, A., Babu, A.R., Zadeh, M.Z., Banerjee, D., Makedon, F.: A survey on contrastive self-supervised learning. Technologies 9(1), 2 (2020)","journal-title":"Technologies"},{"key":"13_CR9","unstructured":"Jia, C., et al.: Scaling up visual and vision-language representation learning with noisy text supervision. In: ICML\u201921, pp. 4904\u20134916. PMLR (2021)"},{"key":"13_CR10","doi-asserted-by":"crossref","unstructured":"Jia, Y., et al.: Caffe: convolutional architecture for fast feature embedding. In: ACM MM\u201914, pp. 675\u2013678 (2014)","DOI":"10.1145\/2647868.2654889"},{"key":"13_CR11","doi-asserted-by":"crossref","unstructured":"Karpukhin, V., et al.: Dense passage retrieval for open-domain question answering. In: EMNLP\u201920, pp. 6769\u20136781 (2020)","DOI":"10.18653\/v1\/2020.emnlp-main.550"},{"issue":"10s","key":"13_CR12","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3505244","volume":"54","author":"S Khan","year":"2022","unstructured":"Khan, S., Naseer, M., Hayat, M., Zamir, S.W., Khan, F.S., Shah, M.: Transformers in vision: a survey. ACM Comput. Surv. (CSUR) 54(10s), 1\u201341 (2022)","journal-title":"ACM Comput. Surv. (CSUR)"},{"key":"13_CR13","unstructured":"Kingma, D.P., Ba, J.: Adam: a method for stochastic optimization. In: ICLR\u201914, pp. 1\u201315 (2014)"},{"key":"13_CR14","unstructured":"Krizhevsky, A., Sutskever, I., Hinton, G.E.: ImageNet classification with deep convolutional neural networks. In: NIPS\u201912. vol.\u00a025 (2012)"},{"key":"13_CR15","doi-asserted-by":"crossref","unstructured":"Li, L.H., Yatskar, M., Yin, D., Hsieh, C.J., Chang, K.W.: What does BERT with vision look at? In: ACL\u201920, pp. 5265\u20135275 (2020)","DOI":"10.18653\/v1\/2020.acl-main.469"},{"key":"13_CR16","doi-asserted-by":"crossref","unstructured":"Lin, T.Y., et al.: Microsoft COCO: common objects in context. In: ECCV\u201914, pp. 740\u2013755. Springer (2014)","DOI":"10.1007\/978-3-319-10602-1_48"},{"key":"13_CR17","unstructured":"Radford, A., et\u00a0al.: Learning transferable visual models from natural language supervision. In: ICML\u201921, pp. 8748\u20138763 (2021)"},{"key":"13_CR18","unstructured":"Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et\u00a0al.: Improving language understanding by generative pre-training. https:\/\/s3-us-west-2.amazonaws.com\/openai-assets\/research-covers\/language-unsupervised\/language_understanding_paper.pdf (2018)"},{"issue":"1","key":"13_CR19","first-page":"1929","volume":"15","author":"N Srivastava","year":"2014","unstructured":"Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., Salakhutdinov, R.: Dropout: a simple way to prevent neural networks from overfitting. J. Mach. Learn. Res. 15(1), 1929\u20131958 (2014)","journal-title":"J. Mach. Learn. Res."},{"key":"13_CR20","doi-asserted-by":"crossref","unstructured":"Tan, H., Bansal, M.: LXMERT: learning cross-modality encoder representations from transformers. In: EMNLP-IJCNLP\u201919, pp. 5100\u20135111 (2019)","DOI":"10.18653\/v1\/D19-1514"},{"key":"13_CR21","unstructured":"Tian, Y., Sun, C., Poole, B., Krishnan, D., Schmid, C., Isola, P.: What makes for good views for contrastive learning. In: NeurIPS\u201920, pp. 6827\u20136839 (2020)"},{"key":"13_CR22","unstructured":"Vaswani, A., et al.: Attention is all you need. In: NIPS\u201917. vol.\u00a030. Curran Associates, Inc. (2017)"},{"key":"13_CR23","doi-asserted-by":"crossref","unstructured":"Voorhees, E.M., et\u00a0al.: The TREC-8 question answering track report. In: TREC. vol.\u00a099, pp. 77\u201382 (1999)","DOI":"10.6028\/NIST.SP.500-246.qa-overview"},{"key":"13_CR24","unstructured":"Wang, K., Yin, Q., Wang, W., Wu, S., Wang, L.: A comprehensive survey on cross-modal retrieval. arXiv preprint arXiv:1607.06215 (2016)"},{"key":"13_CR25","doi-asserted-by":"crossref","unstructured":"Wu, Z., Xiong, Y., Yu, S.X., Lin, D.: Unsupervised feature learning via non-parametric instance discrimination. In: CVPR\u201918, pp. 3733\u20133742 (2018)","DOI":"10.1109\/CVPR.2018.00393"},{"key":"13_CR26","doi-asserted-by":"publisher","first-page":"67","DOI":"10.1162\/tacl_a_00166","volume":"2","author":"P Young","year":"2014","unstructured":"Young, P., Lai, A., Hodosh, M., Hockenmaier, J.: From image descriptions to visual denotations: new similarity metrics for semantic inference over event descriptions. Trans. Assoc. Comput. Linguist. 2, 67\u201378 (2014)","journal-title":"Trans. Assoc. Comput. Linguist."},{"key":"13_CR27","doi-asserted-by":"crossref","unstructured":"Zhai, C.: Large language models and future of information retrieval: opportunities and challenges. In: SIGIR\u201924, pp. 481\u2013490 (2024)","DOI":"10.1145\/3626772.3657848"},{"issue":"6","key":"13_CR28","doi-asserted-by":"publisher","first-page":"1452","DOI":"10.1109\/TPAMI.2017.2723009","volume":"40","author":"B Zhou","year":"2017","unstructured":"Zhou, B., Lapedriza, A., Khosla, A., Oliva, A., Torralba, A.: Places: a 10 million image database for scene recognition. IEEE Trans. Pattern Anal. Mach. Intell. 40(6), 1452\u20131464 (2017)","journal-title":"IEEE Trans. Pattern Anal. Mach. Intell."}],"container-title":["Lecture Notes in Computer Science","MultiMedia Modeling"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/978-981-96-2071-5_13","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,1,1]],"date-time":"2025-01-01T16:04:30Z","timestamp":1735747470000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/978-981-96-2071-5_13"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025]]},"ISBN":["9789819620708","9789819620715"],"references-count":28,"URL":"https:\/\/doi.org\/10.1007\/978-981-96-2071-5_13","relation":{},"ISSN":["0302-9743","1611-3349"],"issn-type":[{"type":"print","value":"0302-9743"},{"type":"electronic","value":"1611-3349"}],"subject":[],"published":{"date-parts":[[2025]]},"assertion":[{"value":"2 January 2025","order":1,"name":"first_online","label":"First Online","group":{"name":"ChapterHistory","label":"Chapter History"}},{"value":"MMM","order":1,"name":"conference_acronym","label":"Conference Acronym","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"International Conference on Multimedia Modeling","order":2,"name":"conference_name","label":"Conference Name","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Nara","order":3,"name":"conference_city","label":"Conference City","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"Japan","order":4,"name":"conference_country","label":"Conference Country","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"2025","order":5,"name":"conference_year","label":"Conference Year","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"9 January 2025","order":7,"name":"conference_start_date","label":"Conference Start Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"11 January 2025","order":8,"name":"conference_end_date","label":"Conference End Date","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"31","order":9,"name":"conference_number","label":"Conference Number","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"mmm2025","order":10,"name":"conference_id","label":"Conference ID","group":{"name":"ConferenceInfo","label":"Conference Information"}},{"value":"https:\/\/mmm2025.net\/","order":11,"name":"conference_url","label":"Conference URL","group":{"name":"ConferenceInfo","label":"Conference Information"}}]}}