{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T10:50:49Z","timestamp":1774263049415,"version":"3.50.1"},"reference-count":0,"publisher":"Slovenian Association Informatika","issue":"8","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["IJCAI"],"abstract":"<jats:p>This paper presents a computational framework for urban image modeling driven by multimodal discourse interaction, integrating cross-modal representation learning, graph-based semantic modeling, and temporal sequence prediction. Images and texts are embedded into a unified semantic space using CLIP and BLIP-2 models, enabling high-fidelity multimodal representation. A knowledge graph of urban discourse is constructed and modeled with a Graph Attention Network (GAT) to capture semantic relationships among multiple agents, while a Temporal Fusion Transformer (TFT) is employed to learn both long-term dependencies and local feature dynamics. To enhance interpretability, a variable selection network identifies the dominant multimodal features shaping urban image evolution. Experimental results based on Hefei\u2019s urban discourse demonstrate high semantic alignment between text\u2013image pairs (e.g., 0.87 for \u201cHefei Metro Expansion\u201d and \u201cMetro Station,\u201d 0.89 for \u201cUSTC Research Breakthrough\u201d and \u201cUSTC Campus\u201d), strong knowledge graph relations (e.g., 0.82 for USTC\u2013High-tech Zone linkage), and accurate temporal forecasting with RMSE reduced to 0.061. The dataset contains 18,426 text entries and 9,307 paired images, and the evaluation adopts a fixed 7:2:1 split with CLIP and BLIP-2 as embedding baselines. Comparative tests against LSTM and GRU yield RMSE values of 0.083 and 0.079 respectively. The findings confirm that the proposed graph-temporal multimodal framework provides an interpretable, data-driven methodology for quantifying and analyzing urban image formation.<\/jats:p>","DOI":"10.31449\/inf.v50i8.11578","type":"journal-article","created":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T09:50:39Z","timestamp":1774259439000},"source":"Crossref","is-referenced-by-count":0,"title":["Graph-Temporal Deep Learning for Urban Image Modeling From Multimodal Discourse Interaction: A Case Study of Hefei"],"prefix":"10.31449","volume":"50","author":[{"given":"Li-xia","family":"Xu","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"16141","published-online":{"date-parts":[[2026,3,23]]},"container-title":["Informatica"],"original-title":[],"link":[{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/11578\/6518","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/download\/11578\/6518","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2026,3,23]],"date-time":"2026-03-23T09:50:40Z","timestamp":1774259440000},"score":1,"resource":{"primary":{"URL":"https:\/\/www.informatica.si\/index.php\/informatica\/article\/view\/11578"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2026,3,23]]},"references-count":0,"journal-issue":{"issue":"8","published-online":{"date-parts":[[2026,2,21]]}},"URL":"https:\/\/doi.org\/10.31449\/inf.v50i8.11578","relation":{},"ISSN":["1854-3871","0350-5596"],"issn-type":[{"value":"1854-3871","type":"electronic"},{"value":"0350-5596","type":"print"}],"subject":[],"published":{"date-parts":[[2026,3,23]]}}}