{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,5]],"date-time":"2026-06-05T16:10:46Z","timestamp":1780675846370,"version":"3.54.1"},"reference-count":68,"publisher":"Association for Computing Machinery (ACM)","issue":"4","license":[{"start":{"date-parts":[[2023,5,8]],"date-time":"2023-05-08T00:00:00Z","timestamp":1683504000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"name":"National Key Research and Development Program of China","award":["2020YFB2104005"],"award-info":[{"award-number":["2020YFB2104005"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["U20B2060, 62272260, 62171260, and U21B2036"],"award-info":[{"award-number":["U20B2060, 62272260, 62171260, and U21B2036"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"International Postdoctoral Exchange Fellowship Program","award":["YJ20210274"],"award-info":[{"award-number":["YJ20210274"]}]},{"DOI":"10.13039\/501100002858","name":"China Postdoctoral Science Foundation","doi-asserted-by":"crossref","award":["2022M721891"],"award-info":[{"award-number":["2022M721891"]}],"id":[{"id":"10.13039\/501100002858","id-type":"DOI","asserted-by":"crossref"}]},{"DOI":"10.13039\/501100002341","name":"Academy of Finland","doi-asserted-by":"crossref","award":["319669, 319670, 325570, 326305, 325774, and 335934"],"award-info":[{"award-number":["319669, 319670, 325570, 326305, 325774, and 335934"]}],"id":[{"id":"10.13039\/501100002341","id-type":"DOI","asserted-by":"crossref"}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Intell. Syst. Technol."],"published-print":{"date-parts":[[2023,8,31]]},"abstract":"<jats:p>\n            Satellite imagery depicts the Earth\u2019s surface remotely and provides comprehensive information for many applications, such as land use monitoring and urban planning. Existing studies on unsupervised representation learning for satellite images only take into account the images\u2019 geographic information, ignoring human activity factors. To bridge this gap, we propose using the Point-of-Interest (POI) data to capture human factors and designing a contrastive learning-based framework to consolidate the representation of satellite imagery with POI information. Besides, we introduce a season-invariant representation learning model on satellite imagery, considering that human factors are mostly unchanging with respect to seasons. An attention model is designed at last to merge the representations from the geographic, seasonal, and POI perspectives adaptively. On the basis of real-world datasets collected from Beijing,\n            <jats:xref ref-type=\"fn\">\n              <jats:sup>1<\/jats:sup>\n            <\/jats:xref>\n            we evaluate our method for predicting socioeconomic indicators. The results show that the representation containing POI information outperforms the geographic representation in estimating commercial activity-related indicators. Our proposed attentional framework can estimate the socioeconomic indicators with\n            <jats:italic>R<\/jats:italic>\n            <jats:sup>2<\/jats:sup>\n            of 0.874 and outperforms the baseline methods. Furthermore, we explore the differences in the representations of satellite images with varying socioeconomic statuses. Finally, we investigate the impact of geographic and POI perspective information in the representation learning process, as well as the effect of satellite imagery on various spatial resolutions.\n          <\/jats:p>","DOI":"10.1145\/3589344","type":"journal-article","created":{"date-parts":[[2023,3,27]],"date-time":"2023-03-27T12:13:10Z","timestamp":1679919190000},"page":"1-32","update-policy":"https:\/\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":3,"title":["Learning Representations of Satellite Imagery by Leveraging Point-of-Interests"],"prefix":"10.1145","volume":"14","author":[{"ORCID":"https:\/\/orcid.org\/0000-0002-4343-703X","authenticated-orcid":false,"given":"Tong","family":"Li","sequence":"first","affiliation":[{"name":"Tsinghua University, China and University of Helsinki, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4715-2186","authenticated-orcid":false,"given":"Yanxin","family":"Xi","sequence":"additional","affiliation":[{"name":"University of Helsinki, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0002-6382-0861","authenticated-orcid":false,"given":"Huandong","family":"Wang","sequence":"additional","affiliation":[{"name":"Tsinghua University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-5617-1659","authenticated-orcid":false,"given":"Yong","family":"Li","sequence":"additional","affiliation":[{"name":"Tsinghua University, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0003-4220-3650","authenticated-orcid":false,"given":"Sasu","family":"Tarkoma","sequence":"additional","affiliation":[{"name":"University of Helsinki, Finland"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/orcid.org\/0000-0001-6026-1083","authenticated-orcid":false,"given":"Pan","family":"Hui","sequence":"additional","affiliation":[{"name":"Hong Kong University of Science and Technology, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2023,5,8]]},"reference":[{"key":"e_1_3_2_2_2","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-020-00243-5"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1145\/3097983.3098070"},{"key":"e_1_3_2_4_2","doi-asserted-by":"crossref","unstructured":"Kumar Ayush Burak Uzkent Marshall Burke David Lobell and Stefano Ermon. 2021. Generating interpretable poverty maps using object detection in satellite images. In Proceedings of the Twenty-Ninth International Conference on International Joint Conferences on Artificial Intelligence . 4410\u20134416.","DOI":"10.24963\/ijcai.2020\/608"},{"key":"e_1_3_2_5_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01002"},{"key":"e_1_3_2_6_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2013.50"},{"key":"e_1_3_2_7_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i17.17728"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i17.17728"},{"key":"e_1_3_2_9_2","doi-asserted-by":"publisher","DOI":"10.1145\/3347146.3359104"},{"issue":"1","key":"e_1_3_2_10_2","first-page":"1","article-title":"Analysis of regional economic development based on land use and land cover change information derived from Landsat imagery","volume":"10","author":"Chen Chao","year":"2020","unstructured":"Chao Chen, Xinyue He, Zhisong Liu, Weiwei Sun, Heng Dong, and Yanli Chu. 2020. Analysis of regional economic development based on land use and land cover change information derived from Landsat imagery. Sci. Rep. 10, 1 (2020), 1\u201316.","journal-title":"Sci. Rep."},{"key":"e_1_3_2_11_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3463495","article-title":"UVLens: Urban village boundary identification and population estimation leveraging open government data","author":"Chen Longbiao","year":"2021","unstructured":"Longbiao Chen, Chenhui Lu, Fangxu Yuan, Zhihan Jiang, Leye Wang, Daqing Zhang, Ruixiang Luo, Xiaoliang Fan, and Cheng Wang. 2021. UVLens: Urban village boundary identification and population estimation leveraging open government data. In Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies. 1\u201326.","journal-title":"Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies"},{"key":"e_1_3_2_12_2","first-page":"1597","volume-title":"International Conference on Machine Learning","author":"Chen Ting","year":"2020","unstructured":"Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. In International Conference on Machine Learning. PMLR, 1597\u20131607."},{"key":"e_1_3_2_13_2","article-title":"Improved baselines with momentum contrastive learning","author":"Chen Xinlei","year":"2020","unstructured":"Xinlei Chen, Haoqi Fan, Ross Girshick, and Kaiming He. 2020. Improved baselines with momentum contrastive learning. arXiv:2003.04297. Retrieved from https:\/\/arxiv.org\/abs\/2003.04297.","journal-title":"arXiv:2003.04297"},{"key":"e_1_3_2_14_2","doi-asserted-by":"publisher","DOI":"10.1109\/JSTARS.2021.3086139"},{"key":"e_1_3_2_15_2","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1903064116"},{"key":"e_1_3_2_16_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICDM.2019.00026"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1145\/3264916"},{"key":"e_1_3_2_18_2","doi-asserted-by":"publisher","DOI":"10.1145\/3477577"},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.3301906"},{"key":"e_1_3_2_20_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i01.5379"},{"key":"e_1_3_2_21_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394486.3403347"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00975"},{"key":"e_1_3_2_23_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_24_2","doi-asserted-by":"publisher","DOI":"10.1145\/3184558.3186353"},{"key":"e_1_3_2_25_2","doi-asserted-by":"publisher","DOI":"10.1145\/3136560.3136576"},{"key":"e_1_3_2_26_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2020.3015157"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.243"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1016\/S0893-6080(00)00026-5"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1126\/science.aaf7894"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v33i01.33013967"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs12193235"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2020.3007029"},{"key":"e_1_3_2_33_2","doi-asserted-by":"publisher","DOI":"10.1109\/IGARSS47720.2021.9553499"},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1145\/3486183.3491001"},{"key":"e_1_3_2_35_2","unstructured":"Diederik P. Kingma and Jimmy Ba. 2017. Adam: A Method for Stochastic Optimization. arxiv:cs.LG\/1412.6980. Retrieved from https:\/\/arxiv.org\/abs\/1412.6980."},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1002\/aic.690370209"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1145\/3342240"},{"key":"e_1_3_2_38_2","doi-asserted-by":"publisher","DOI":"10.3390\/land10060648"},{"key":"e_1_3_2_39_2","doi-asserted-by":"publisher","DOI":"10.1145\/3511808.3557153"},{"key":"e_1_3_2_40_2","doi-asserted-by":"publisher","DOI":"10.1145\/3369799"},{"key":"e_1_3_2_41_2","article-title":"Geo-Tile2Vec: A multi-modal and multi-stage embedding framework for urban analytics","author":"Luo Yan","year":"2022","unstructured":"Yan Luo, Chak-Tou Leong, Shuhai Jiao, Fu-Lai Chung, Wenjie Li, and Guoping Liu. 2022. Geo-Tile2Vec: A multi-modal and multi-stage embedding framework for urban analytics. ACM Trans. Spat. Algor. Syst. (2022).","journal-title":"ACM Trans. Spat. Algor. Syst."},{"key":"e_1_3_2_42_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00928"},{"key":"e_1_3_2_43_2","article-title":"Efficient estimation of word representations in vector space","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv:1301.3781. Retrieved from https:\/\/arxiv.org\/abs\/1301.3781.","journal-title":"arXiv:1301.3781"},{"key":"e_1_3_2_44_2","first-page":"3111","volume-title":"Advances in Neural Information Processing Systems","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S. Corrado, and Jeff Dean. 2013. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems. 3111\u20133119."},{"key":"e_1_3_2_45_2","doi-asserted-by":"publisher","DOI":"10.1111\/j.1467-8306.2004.09402005.x"},{"key":"e_1_3_2_46_2","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1919913118"},{"key":"e_1_3_2_47_2","first-page":"90","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops","author":"Piaggesi Simone","year":"2019","unstructured":"Simone Piaggesi, Laetitia Gauvin, Michele Tizzoni, Ciro Cattuto, Natalia Adler, Stefaan Verhulst, Andrew Young, Rhiannan Price, Leo Ferres, and Andr\u00e9 Panisson. 2019. Predicting city poverty using satellite imagery. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition Workshops. 90\u201396."},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs13183603"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1145\/3449257"},{"key":"e_1_3_2_50_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.rse.2021.112339"},{"key":"e_1_3_2_51_2","doi-asserted-by":"publisher","DOI":"10.1111\/1467-9868.00196"},{"key":"e_1_3_2_52_2","unstructured":"Aaron van den Oord Yazhe Li and Oriol Vinyals. 2019. Representation learning with contrastive predictive coding. arxiv:cs.LG\/1807.03748. Retrieved from https:\/\/arxiv.org\/abs\/1807.03748."},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1145\/3132847.3133006"},{"issue":"6","key":"e_1_3_2_54_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3209686","article-title":"Learning urban community structures: A collective embedding perspective with periodic spatial-temporal mobility graphs","volume":"9","author":"Wang Pengyang","year":"2018","unstructured":"Pengyang Wang, Yanjie Fu, Jiawei Zhang, Xiaolin Li, and Dan Lin. 2018. Learning urban community structures: A collective embedding perspective with periodic spatial-temporal mobility graphs. ACM Trans. Intell. Syst. Technol. 9, 6 (2018), 1\u201328.","journal-title":"ACM Trans. Intell. Syst. Technol."},{"key":"e_1_3_2_55_2","first-page":"34","article-title":"Measuring urban vibrancy of residential communities using big crowdsourced geotagged data","author":"Wang Pengyang","year":"2021","unstructured":"Pengyang Wang, Kunpeng Liu, Dongjie Wang, and Yanjie Fu. 2021. Measuring urban vibrancy of residential communities using big crowdsourced geotagged data. Front. Big Data (2021), 34.","journal-title":"Front. Big Data"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1145\/3184558.3186581"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i01.5450"},{"key":"e_1_3_2_58_2","doi-asserted-by":"publisher","DOI":"10.1145\/3513092"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00393"},{"key":"e_1_3_2_60_2","article-title":"Beyond the first law of geography: Learning representations of satellite imagery by leveraging point-of-interests","author":"Xi Yanxin","year":"2022","unstructured":"Yanxin Xi, Tong Li, Huandong Wang, Yong Li, Sasu Tarkoma, and Pan Hui. 2022. Beyond the first law of geography: Learning representations of satellite imagery by leveraging point-of-interests. In Proceedings of the Web Conference (2022).","journal-title":"Proceedings of the Web Conference"},{"key":"e_1_3_2_61_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v30i1.9906"},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2017.634"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1145\/3372406"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.3390\/rs11050574"},{"key":"e_1_3_2_65_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41467-020-16185-w"},{"key":"e_1_3_2_66_2","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2020\/611"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.3390\/s18113717"},{"key":"e_1_3_2_68_2","doi-asserted-by":"publisher","DOI":"10.1145\/3292500.3330972"},{"issue":"3","key":"e_1_3_2_69_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/2629592","article-title":"Urban computing: Concepts, methodologies, and applications","volume":"5","author":"Zheng Yu","year":"2014","unstructured":"Yu Zheng, Licia Capra, Ouri Wolfson, and Hai Yang. 2014. Urban computing: Concepts, methodologies, and applications. ACM Trans. Intell. Syst. Technol. 5, 3 (2014), 1\u201355.","journal-title":"ACM Trans. Intell. Syst. Technol."}],"container-title":["ACM Transactions on Intelligent Systems and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589344","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.1145\/3589344","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T00:03:16Z","timestamp":1750291396000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.1145\/3589344"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2023,5,8]]},"references-count":68,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2023,8,31]]}},"alternative-id":["10.1145\/3589344"],"URL":"https:\/\/doi.org\/10.1145\/3589344","relation":{},"ISSN":["2157-6904","2157-6912"],"issn-type":[{"value":"2157-6904","type":"print"},{"value":"2157-6912","type":"electronic"}],"subject":[],"published":{"date-parts":[[2023,5,8]]},"assertion":[{"value":"2022-06-22","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-03-11","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-05-08","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}