{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,24]],"date-time":"2026-06-24T19:00:57Z","timestamp":1782327657762,"version":"3.54.5"},"reference-count":43,"publisher":"Springer Science and Business Media LLC","issue":"24","license":[{"start":{"date-parts":[[2022,12,20]],"date-time":"2022-12-20T00:00:00Z","timestamp":1671494400000},"content-version":"tdm","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"},{"start":{"date-parts":[[2022,12,20]],"date-time":"2022-12-20T00:00:00Z","timestamp":1671494400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0"}],"funder":[{"DOI":"10.13039\/501100016386","name":"Conselleria de Innovaci\u00f3n, Universidades, Ciencia y Sociedad Digital, Generalitat Valenciana","doi-asserted-by":"publisher","award":["GVA- COVID19\/2021\/103"],"award-info":[{"award-number":["GVA- COVID19\/2021\/103"]}],"id":[{"id":"10.13039\/501100016386","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/501100008530","name":"European Regional Development Fund","doi-asserted-by":"crossref","award":["MCIN\/AEI\/10.13039\/501100011033"],"award-info":[{"award-number":["MCIN\/AEI\/10.13039\/501100011033"]}],"id":[{"id":"10.13039\/501100008530","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/100009092","name":"Universidad de Alicante","doi-asserted-by":"crossref","id":[{"id":"10.13039\/100009092","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["link.springer.com"],"crossmark-restriction":false},"short-container-title":["Multimed Tools Appl"],"published-print":{"date-parts":[[2023,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>This paper presents a machine learning-based classifier for detecting points of interest through the combined use of images and text from social networks. This model exploits the transfer learning capabilities of the neural network architecture CLIP (Contrastive Language-Image Pre-Training) in multimodal environments using image and text. Different methodologies based on multimodal information are explored for the geolocation of the places detected. To this end, pre-trained neural network models are used for the classification of images and their associated texts. The result is a system that allows creating new synergies between images and texts in order to detect and geolocate trending places that has not been previously tagged by any other means, providing potentially relevant information for tasks such as cataloging specific types of places in a city for the tourism industry. The experiments carried out reveal that, in general, textual information is more accurate and relevant than visual cues in this multimodal setting.<\/jats:p>","DOI":"10.1007\/s11042-022-14296-8","type":"journal-article","created":{"date-parts":[[2022,12,20]],"date-time":"2022-12-20T03:05:46Z","timestamp":1671505546000},"page":"38097-38116","update-policy":"https:\/\/doi.org\/10.1007\/springer_crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Detecting and locating trending places using multimodal social network data"],"prefix":"10.1007","volume":"82","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-8607-5147","authenticated-orcid":false,"given":"Luis","family":"Lucas","sequence":"first","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"David","family":"Tom\u00e1s","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Jose","family":"Garcia-Rodriguez","sequence":"additional","affiliation":[],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"297","published-online":{"date-parts":[[2022,12,20]]},"reference":[{"issue":"2021","key":"14296_CR1","doi-asserted-by":"publisher","first-page":"279","DOI":"10.1016\/j.inffus.2021.10.013","volume":"79","author":"I Afyouni","year":"2022","unstructured":"Afyouni I, Aghbari ZA, Razack RA (2022) Multi-feature, multi-modal, and multi-source social event detection: a comprehensive survey. Inf Fusion 79 (2021):279\u2013308. https:\/\/doi.org\/10.1016\/j.inffus.2021.10.013","journal-title":"Inf Fusion"},{"key":"14296_CR2","doi-asserted-by":"publisher","unstructured":"Arora G, Pavani PL, Kohli R, Bibhu V (2016) Multimodal biometrics for improvised security. 2016 1st Int Conf Innovation Challenges in Cyber Secur, ICICCS 2016 (Iciccs):1\u20135. https:\/\/doi.org\/10.1109\/ICICCS.2016.7542312https:\/\/doi.org\/10.1109\/ICICCS.2016.7542312","DOI":"10.1109\/ICICCS.2016.7542312 10.1109\/ICICCS.2016.7542312"},{"key":"14296_CR3","unstructured":"Chang M-W, Ratinov L, Roth D, Srikumar V (2008) Importance of semantic representation: dataless classification. In: Proceedings of the 23rd national conference on artificial intelligence - vol 2. AAAI\u201908. AAAI press, pp 830\u2013835"},{"key":"14296_CR4","doi-asserted-by":"publisher","unstructured":"Cheng J, Fostiropoulos I, Boehm B, Soleymani M (2021) Multimodal phased transformer for sentiment analysis. EMNLP 2021 - 2021 conference on empirical methods in natural language processing, proceedings, pp 2447\u20132458. https:\/\/doi.org\/10.18653\/v1\/2021.emnlp-main.189","DOI":"10.18653\/v1\/2021.emnlp-main.189"},{"key":"14296_CR5","unstructured":"Cho J, Lei J, Tan H, Bansal M (2021) Unifying vision-and-language tasks via text generation. arXiv:2102.02779"},{"issue":"2018","key":"14296_CR6","doi-asserted-by":"publisher","first-page":"259","DOI":"10.1016\/j.inffus.2019.02.010","volume":"51","author":"JH Choi","year":"2019","unstructured":"Choi JH, Lee JS (2019) Embracenet: a robust deep learning architecture for multimodal classification. Inf Fusion 51(2018):259\u2013270. arXiv:1904.09078. https:\/\/doi.org\/10.1016\/j.inffus.2019.02.010","journal-title":"Inf Fusion"},{"key":"14296_CR7","doi-asserted-by":"publisher","unstructured":"Devlin J, Chang M, Lee K, Toutanova K (2019) BERT: pre-training of deep bidirectional transformers for language understanding. In: Proceedings of the 2019 conference of the north american chapter of the association for computational linguistics: human language technologies, NAACL-HLT. Association for computational linguistics, pp 4171\u20134186. https:\/\/doi.org\/10.18653\/v1\/n19-1423","DOI":"10.18653\/v1\/n19-1423"},{"key":"14296_CR8","unstructured":"Dosovitskiy A, Beyer L, Kolesnikov A, Weissenborn D, Zhai X, Unterthiner T, Dehghani M, Minderer M, Heigold G, Gelly S, Uszkoreit J, Houlsby N (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv:2010.11929"},{"key":"14296_CR9","unstructured":"Duong CT, Lebret R, Aberer K (2017) Multimodal classification for analysing social media. arXiv:1708.02099"},{"key":"14296_CR10","doi-asserted-by":"publisher","unstructured":"Dzabraev M, Kalashnikov M, Komkov S, Petiushko A (2021) MDMMT: multidomain multimodal transformer for video retrieval. IEEE Comput Society Conf Comput Vis Pattern Recognit Workshops:3349\u20133358. https:\/\/doi.org\/10.1109\/CVPRW53098.2021.00374","DOI":"10.1109\/CVPRW53098.2021.00374"},{"key":"14296_CR11","unstructured":"Fan A, Grave E, Joulin A (2019) Reducing transformer depth on demand with structured dropout, vol 103, pp 1\u201315. arXiv:1909.11556"},{"key":"14296_CR12","doi-asserted-by":"publisher","first-page":"514","DOI":"10.1007\/978-3-030-11024-6_40","volume":"11134 LNCS","author":"R Gomez","year":"2019","unstructured":"Gomez R, Gomez L, Gibert J, Karatzas D (2019) Learning to learn from web data through deep semantic embeddings. Lect Notes Comput Sci (including subseries Lecture Notes Artif Intell Lecture Notes in Bioinformatics) 11134 LNCS:514\u2013529. arXiv:1808.06368. https:\/\/doi.org\/10.1007\/978-3-030-11024-6_40","journal-title":"Lect Notes Comput Sci (including subseries Lecture Notes Artif Intell Lecture Notes in Bioinformatics)"},{"key":"14296_CR13","doi-asserted-by":"publisher","unstructured":"He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. Proc IEEE Comput Society Conf Comput Vis Pattern Recognit:770\u2013778. https:\/\/doi.org\/10.1109\/CVPR.2016.90","DOI":"10.1109\/CVPR.2016.90"},{"key":"14296_CR14","doi-asserted-by":"publisher","unstructured":"Holzinger A (2021) The next frontier: AI we can really trust. In: Machine learning and principles and practice of knowledge discovery in databases. Springer, pp 427\u2013440. https:\/\/doi.org\/10.1007\/978-3-030-93736-2_33","DOI":"10.1007\/978-3-030-93736-2_33"},{"key":"14296_CR15","doi-asserted-by":"publisher","unstructured":"Huang J, Tao J, Liu B, Lian Z, Niu M (2020) Multimodal transformer fusion for continuous emotion recognition. In: ICASSP 2020 - 2020 IEEE international conference on acoustics, speech and signal processing (ICASSP), pp 3507\u20133511. https:\/\/doi.org\/10.1109\/ICASSP40776.2020.9053762","DOI":"10.1109\/ICASSP40776.2020.9053762"},{"key":"14296_CR16","unstructured":"Jaegle A, Gimeno F, Brock A, Vinyals O, Zisserman A, Carreira J (2021) Perceiver: general perception with iterative attention. In: Proceedings of the 38th international conference on machine learning. Proceedings of machine learning research. PMLR, vol 139, pp 4651\u20134664"},{"key":"14296_CR17","doi-asserted-by":"publisher","unstructured":"Kumar P, Ofli F, Imran M, Castillo C (2020) Detection of disaster-affected cultural heritage sites from social media images using deep learning techniques. J Comput Cultural Heritage, vol 13(3). https:\/\/doi.org\/10.1145\/3383314https:\/\/doi.org\/10.1145\/3383314","DOI":"10.1145\/3383314 10.1145\/3383314"},{"key":"14296_CR18","doi-asserted-by":"publisher","unstructured":"Kumar A, Singh JP, Dwivedi YK, Rana NP (2020) A deep multi-modal neural network for informative twitter content classification during emergencies. Annals Oper Res:(0123456789). https:\/\/doi.org\/10.1007\/s10479-020-03514-x","DOI":"10.1007\/s10479-020-03514-x"},{"key":"14296_CR19","doi-asserted-by":"publisher","first-page":"2476","DOI":"10.1109\/TASLP.2021.3065823","volume":"29","author":"Z Li","year":"2021","unstructured":"Li Z, Li Z, Zhang J, Feng Y, Zhou J (2021) Bridging text and video: a universal multimodal transformer for audio-visual scene-aware dialog. IEEE\/ACM Trans Audio Speech Language Process 29:2476\u20132483. https:\/\/doi.org\/10.1109\/TASLP.2021.3065823","journal-title":"IEEE\/ACM Trans Audio Speech Language Process"},{"issue":"6","key":"14296_CR20","doi-asserted-by":"publisher","first-page":"2111","DOI":"10.1109\/ICDE.2019.00250","volume":"2019-April","author":"P Li","year":"2019","unstructured":"Li P, Lu H, Kanhabua N, Zhao S, Pan G (2019) Location inference for non-geotagged tweets in user timelines [Extended Abstract]. Proc Int Conf Data Eng 2019-April(6):2111\u20132112. https:\/\/doi.org\/10.1109\/ICDE.2019.00250","journal-title":"Proc Int Conf Data Eng"},{"key":"14296_CR21","unstructured":"Li LH, Yatskar M, Yin D, Hsieh C-J, Chang K-W (2019) VisualBERT: a simple and performant baseline for vision and language:(2), pp 1\u201314. arXiv:1908.03557"},{"key":"14296_CR22","unstructured":"Liu Y, Ott M, Goyal N, Du J, Joshi M, Chen D, Levy O, Lewis M, Zettlemoyer L, Stoyanov V (2019) Roberta: a robustly optimized bert pretraining approach, pp 1\u201313, coRR arXiv:1907.11692. https:\/\/doi.org\/10.48550"},{"key":"14296_CR23","doi-asserted-by":"crossref","unstructured":"Lucas L, Tom\u00e1s D, Garcia-Rodriguez J (2022) Exploiting the relationship between visual and textual features in social networks for image classification with zero-shot deep learning. In: Sanjurjo gonzalez\u0301 H, Pastor L\u00f3pez I, Garc\u00eda Bringas P, Quinti\u00e1n H, Corchado E (eds) 16th International conference on soft computing models in industrial and environmental applications (SOCO 2021). Springer, pp 369\u2013378","DOI":"10.1007\/978-3-030-87869-6_35"},{"key":"14296_CR24","doi-asserted-by":"crossref","unstructured":"Lucas L, Tom\u00e1s D, Garcia-Rodriguez J (2022) Sentiment analysis and image classification in social networks with zero-shot deep learning: applications in tourism. In: Sanjurjo gonz\u00e1lez H, Pastor L\u00f3pez I, Garc\u00eda Bringas P, Quinti\u00e1n H, Corchado E (eds) 16th International conference on soft computing models in industrial and environmental applications (SOCO 2021). Springer, pp 419\u2013428","DOI":"10.1007\/978-3-030-87869-6_40"},{"key":"14296_CR25","unstructured":"Miller SJ, Howard J, Adams P, Schwan M, Slater R, Miller S, Howard J, Adams P, Schwan M, Slater R (2020) SMU data science review multi-modal classification using images and text multi-modal classification using images and text, vol 3(3)"},{"issue":"4","key":"14296_CR26","doi-asserted-by":"publisher","first-page":"510","DOI":"10.1016\/j.ipm.2014.07.011","volume":"51","author":"G Petz","year":"2015","unstructured":"Petz G, Karpowicz M, F\u00fcrschu\u00df H, Auinger A, St\u0159\u00edtesk\u00fd V, Holzinger A (2015) Reprint of: computational approaches for mining user\u2019s opinions on the web 2.0. Inf Process Manag 51(4):510\u2013519. https:\/\/doi.org\/10.1016\/j.ipm.2014.07.011","journal-title":"Inf Process Manag"},{"key":"14296_CR27","unstructured":"Radford A, Kim JW, Hallacy C, Ramesh A, Goh G, Agarwal S, Sastry G, Askell A, Mishkin P, Clark J, Krueger G, Sutskever I (2021) Learning transferable visual models from natural language supervision. In: Proceedings of the 38th international conference on machine learning. Proceedings of machine learning research. PMLR, vol 139, pp 8748\u20138763. Accessed Dec 2022. https:\/\/proceedings.mlr.press\/v139\/radford21a.html"},{"issue":"3","key":"14296_CR28","doi-asserted-by":"publisher","first-page":"211","DOI":"10.1007\/s11263-015-0816-y","volume":"115","author":"O Russakovsky","year":"2015","unstructured":"Russakovsky O, Deng J, Su H, Krause J, Satheesh S, Ma S, Huang Z, Karpathy A, Khosla A, Bernstein M, Berg AC, Fei-Fei L (2015) ImageNet large scale visual recognition challenge. Int J Comput Vis 115(3):211\u2013252. arXiv:1409.0575. https:\/\/doi.org\/10.1007\/s11263-015-0816-y","journal-title":"Int J Comput Vis"},{"key":"14296_CR29","unstructured":"Sabour S, Frosst N, Hinton GE (2017) Dynamic routing between capsules. In: Advances in neural information processing systems, vol 30, pp 3856-3866. Curran Associates, Inc., USA"},{"key":"14296_CR30","doi-asserted-by":"publisher","first-page":"112943","DOI":"10.1016\/j.eswa.2019.112943","volume":"141","author":"E Saquete","year":"2020","unstructured":"Saquete E, Tom\u00e1s D, Moreda P, Mart\u00ednez-Barco P, Palomar M (2020) Fighting post-truth using natural language processing: a review and open challenges. Expert Syst Appl 141:112943","journal-title":"Expert Syst Appl"},{"issue":"24","key":"14296_CR31","doi-asserted-by":"publisher","first-page":"21503","DOI":"10.1007\/s00521-021-06086-4","volume":"34","author":"B Singh","year":"2022","unstructured":"Singh B, Sharma DK (2022) Predicting image credibility in fake news over social media using multi-modal approach. Neural Comput Appl 34 (24):21503\u201321517. https:\/\/doi.org\/10.1007\/s00521-021-06086-4","journal-title":"Neural Comput Appl"},{"key":"14296_CR32","doi-asserted-by":"publisher","unstructured":"Tan H, Bansal M (2019) LXMErt: learning cross-modality encoder representations from transformers. In: EMNLP-IJCNLP 2019 - 2019 conference on empirical methods in natural language processing and 9th international joint conference on natural language processing, proceedings of the conference, pp 5100\u20135111. arXiv:1908.07490. https:\/\/doi.org\/10.18653\/v1\/d19-1514","DOI":"10.18653\/v1\/d19-1514"},{"key":"14296_CR33","doi-asserted-by":"publisher","unstructured":"Tom\u00e1s D, Ortega-Bueno R, Zhang G, Rosso P, Schifanella R (2022) Transformer-based models for multimodal irony detection. J Ambient Intell Humanized Comput:1\u201312. https:\/\/doi.org\/10.1007\/s12652-022-04447-yhttps:\/\/doi.org\/10.1007\/s12652-022-04447-y","DOI":"10.1007\/s12652-022-04447-y 10.1007\/s12652-022-04447-y"},{"key":"14296_CR34","unstructured":"Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017) Attention is all you need. In: Advances in neural information processing systems, vol 30, pp 5998-6008. Curran Associates, Inc., USA"},{"key":"14296_CR35","doi-asserted-by":"crossref","unstructured":"Wang L, Li Y, Lazebnik S (2015) Learning deep structure-preserving image-text embeddings. arXiv:1511.06078. https:\/\/doi.org\/10.48550","DOI":"10.1109\/CVPR.2016.541"},{"key":"14296_CR36","doi-asserted-by":"crossref","unstructured":"Xu P, Zhu X, Clifton DA (2022) Multimodal learning with transformers: a Survey:1\u201323. arXiv:2206.06488","DOI":"10.1109\/TPAMI.2023.3275156"},{"key":"14296_CR37","doi-asserted-by":"publisher","unstructured":"Yao S, Wan X (2020) Multimodal transformer for multimodal machine translation, pp 4346\u20134350. https:\/\/doi.org\/10.18653\/v1\/2020.acl-main.400","DOI":"10.18653\/v1\/2020.acl-main.400"},{"key":"14296_CR38","unstructured":"You Y, Li J, Reddi S, Hseu J, Kumar S, Bhojanapalli S, Song X, Demmel J, Keutzer K, Hsieh C-J (2019) Large batch optimization for deep learning: training BERT in 76 minutes. arXiv:1904.00962"},{"key":"14296_CR39","unstructured":"You K, Long M, Wang J, Jordan MI (2019) How does learning rate decay help modern neural networks. arXiv:1908.01878"},{"issue":"12","key":"14296_CR40","doi-asserted-by":"publisher","first-page":"4467","DOI":"10.1109\/TCSVT.2019.2947482","volume":"30","author":"J Yu","year":"2020","unstructured":"Yu J, Li J, Yu Z, Huang Q (2020) Multimodal transformer with Multi-View visual representation for image captioning. IEEE Trans Circuits Syst Video Technol 30(12):4467\u20134480. https:\/\/doi.org\/10.1109\/TCSVT.2019.2947482","journal-title":"IEEE Trans Circuits Syst Video Technol"},{"key":"14296_CR41","doi-asserted-by":"publisher","first-page":"360","DOI":"10.1016\/j.neucom.2021.10.039","volume":"468","author":"B Zhao","year":"2022","unstructured":"Zhao B, Gong M, Li X (2022) Hierarchical multimodal transformer to summarize videos. Neurocomputing 468:360\u2013369. https:\/\/doi.org\/10.1016\/j.neucom.2021.10.039","journal-title":"Neurocomputing"},{"key":"14296_CR42","first-page":"487","volume":"1","author":"B Zhou","year":"2014","unstructured":"Zhou B, Lapedriza A, Xiao J, Torralba A, Oliva A (2014) Learning deep features for scene recognition using places database - supplementary materials. NIPS\u201914 Proc 27th Int Conf Neural Inf Process Syst 1:487\u2013495","journal-title":"NIPS\u201914 Proc 27th Int Conf Neural Inf Process Syst"},{"key":"14296_CR43","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1109\/TNNLS.2022.3154204","volume":"1","author":"F Zhou","year":"2022","unstructured":"Zhou F, Qi X, Zhang K, Trajcevski G, Zhong T (2022) Metageo: a general framework for social user geolocation identification with few-shot learning. IEEE Trans Neural Netw Learn Syst 1:1\u201315. https:\/\/doi.org\/10.1109\/TNNLS.2022.3154204","journal-title":"IEEE Trans Neural Netw Learn Syst"}],"container-title":["Multimedia Tools and Applications"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-022-14296-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/article\/10.1007\/s11042-022-14296-8\/fulltext.html","content-type":"text\/html","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/link.springer.com\/content\/pdf\/10.1007\/s11042-022-14296-8.pdf","content-type":"application\/pdf","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,10,3]],"date-time":"2023-10-03T09:20:54Z","timestamp":1696324854000},"score":1,"resource":{"primary":{"URL":"https:\/\/link.springer.com\/10.1007\/s11042-022-14296-8"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2022,12,20]]},"references-count":43,"journal-issue":{"issue":"24","published-print":{"date-parts":[[2023,10]]}},"alternative-id":["14296"],"URL":"https:\/\/doi.org\/10.1007\/s11042-022-14296-8","relation":{},"ISSN":["1380-7501","1573-7721"],"issn-type":[{"value":"1380-7501","type":"print"},{"value":"1573-7721","type":"electronic"}],"subject":[],"published":{"date-parts":[[2022,12,20]]},"assertion":[{"value":"1 August 2022","order":1,"name":"received","label":"Received","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"10 November 2022","order":2,"name":"revised","label":"Revised","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"3 December 2022","order":3,"name":"accepted","label":"Accepted","group":{"name":"ArticleHistory","label":"Article History"}},{"value":"20 December 2022","order":4,"name":"first_online","label":"First Online","group":{"name":"ArticleHistory","label":"Article History"}},{"order":1,"name":"Ethics","group":{"name":"EthicsHeading","label":"Declarations"}},{"value":"The authors declare that they have no conflict of interest.","order":2,"name":"Ethics","group":{"name":"EthicsHeading","label":"<!--Emphasis Type='Bold' removed-->Conflict of Interests"}}]}}