{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,8,23]],"date-time":"2025-08-23T05:24:22Z","timestamp":1755926662743},"reference-count":34,"publisher":"Association for Computing Machinery (ACM)","issue":"4","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["Proc. VLDB Endow."],"published-print":{"date-parts":[[2021,12]]},"abstract":"<jats:p>Graph embedding aims at learning a vector-based representation of vertices that incorporates the structure of the graph. This representation then enables inference of graph properties. Existing graph embedding techniques, however, do not scale well to large graphs. While several techniques to scale graph embedding using compute clusters have been proposed, they require continuous communication between the compute nodes and cannot handle node failure. We therefore propose a framework for scalable and robust graph embedding based on the MapReduce model, which can distribute any existing embedding technique. Our method splits a graph into subgraphs to learn their embeddings in isolation and subsequently reconciles the embedding spaces derived for the subgraphs. We realize this idea through a novel distributed graph decomposition algorithm. In addition, we show how to implement our framework in Spark to enable efficient learning of effective embeddings. Experimental results illustrate that our approach scales well, while largely maintaining the embedding quality.<\/jats:p>","DOI":"10.14778\/3503585.3503599","type":"journal-article","created":{"date-parts":[[2022,4,14]],"date-time":"2022-04-14T22:18:07Z","timestamp":1649974687000},"page":"914-922","source":"Crossref","is-referenced-by-count":3,"title":["Scalable robust graph embedding with Spark"],"prefix":"10.14778","volume":"15","author":[{"given":"Chi Thang","family":"Duong","sequence":"first","affiliation":[{"name":"EPFL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Trung Dung","family":"Hoang","sequence":"additional","affiliation":[{"name":"EPFL"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Hongzhi","family":"Yin","sequence":"additional","affiliation":[{"name":"The University of Queensland"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Matthias","family":"Weidlich","sequence":"additional","affiliation":[{"name":"Humboldt-Universit\u00e4t zu Berlin"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Quoc Viet Hung","family":"Nguyen","sequence":"additional","affiliation":[{"name":"Griffith University"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Karl","family":"Aberer","sequence":"additional","affiliation":[{"name":"EPFL"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2022,4,14]]},"reference":[{"key":"e_1_2_1_1_1","doi-asserted-by":"crossref","unstructured":"Mikel Artetxe Gorka Labaka and Eneko Agirre. 2016. Learning principled bilingual mappings of word embeddings while preserving monolingual invariance. In ACL. 2289--2294.  Mikel Artetxe Gorka Labaka and Eneko Agirre. 2016. Learning principled bilingual mappings of word embeddings while preserving monolingual invariance. In ACL. 2289--2294.","DOI":"10.18653\/v1\/D16-1250"},{"key":"e_1_2_1_2_1","volume-title":"A comprehensive survey of graph embedding: problems, techniques and applications. TKDE","author":"Cai Hongyun","year":"2018","unstructured":"Hongyun Cai , Vincent W Zheng , and Kevin Chang . 2018. A comprehensive survey of graph embedding: problems, techniques and applications. TKDE ( 2018 ). Hongyun Cai, Vincent W Zheng, and Kevin Chang. 2018. A comprehensive survey of graph embedding: problems, techniques and applications. TKDE (2018)."},{"key":"e_1_2_1_3_1","doi-asserted-by":"crossref","unstructured":"Wei-Lin Chiang Xuanqing Liu Si Si Yang Li Samy Bengio and Cho-Jui Hsieh. 2019. Cluster-GCN: An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks. In KDD. 257--266.  Wei-Lin Chiang Xuanqing Liu Si Si Yang Li Samy Bengio and Cho-Jui Hsieh. 2019. Cluster-GCN: An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks. In KDD. 257--266.","DOI":"10.1145\/3292500.3330925"},{"key":"e_1_2_1_4_1","doi-asserted-by":"publisher","DOI":"10.1145\/1327452.1327492"},{"key":"e_1_2_1_5_1","unstructured":"Micha\u00ebl Defferrard Xavier Bresson and Pierre Vandergheynst. 2016. Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering. In NIPS. 3837--3845.  Micha\u00ebl Defferrard Xavier Bresson and Pierre Vandergheynst. 2016. Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering. In NIPS. 3837--3845."},{"key":"e_1_2_1_6_1","doi-asserted-by":"publisher","DOI":"10.1007\/s00453-013-9802-3"},{"key":"e_1_2_1_7_1","volume-title":"Fast Graph Representation Learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds.","author":"Fey Matthias","unstructured":"Matthias Fey and Jan E. Lenssen . 2019 . Fast Graph Representation Learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds. Matthias Fey and Jan E. Lenssen. 2019. Fast Graph Representation Learning with PyTorch Geometric. In ICLR Workshop on Representation Learning on Graphs and Manifolds."},{"key":"e_1_2_1_8_1","volume-title":"Dahl","author":"Gilmer Justin","year":"2017","unstructured":"Justin Gilmer , Samuel S. Schoenholz , Patrick F. Riley , Oriol Vinyals , and George E . Dahl . 2017 . Neural Message Passing for Quantum Chemistry. In ICML. 1263--1272. Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. 2017. Neural Message Passing for Quantum Chemistry. In ICML. 1263--1272."},{"key":"e_1_2_1_9_1","doi-asserted-by":"crossref","unstructured":"Aditya Grover and Jure Leskovec. 2016. node2vec: Scalable Feature Learning for Networks. In KDD. 855--864.  Aditya Grover and Jure Leskovec. 2016. node2vec: Scalable Feature Learning for Networks. In KDD. 855--864.","DOI":"10.1145\/2939672.2939754"},{"key":"e_1_2_1_10_1","volume-title":"Representation learning on graphs: Methods and applications","author":"Hamilton William L","year":"2017","unstructured":"William L Hamilton , Rex Ying , and Jure Leskovec . 2017. Representation learning on graphs: Methods and applications . IEEE Data Engineering Bulletin ( 2017 ). William L Hamilton, Rex Ying, and Jure Leskovec. 2017. Representation learning on graphs: Methods and applications. IEEE Data Engineering Bulletin (2017)."},{"key":"e_1_2_1_11_1","unstructured":"William L. Hamilton Zhitao Ying and Jure Leskovec. 2017. Inductive Representation Learning on Large Graphs. In NIPS. 1024--1034.  William L. Hamilton Zhitao Ying and Jure Leskovec. 2017. Inductive Representation Learning on Large Graphs. In NIPS. 1024--1034."},{"key":"e_1_2_1_12_1","unstructured":"Weihua Hu Matthias Fey Marinka Zitnik Yuxiao Dong Hongyu Ren Bowen Liu Michele Catasta and Jure Leskovec. 2020. Open Graph Benchmark: Datasets for Machine Learning on Graphs. In NIPS.  Weihua Hu Matthias Fey Marinka Zitnik Yuxiao Dong Hongyu Ren Bowen Liu Michele Catasta and Jure Leskovec. 2020. Open Graph Benchmark: Datasets for Machine Learning on Graphs. In NIPS."},{"key":"e_1_2_1_13_1","unstructured":"George Karypis and Vipin Kumar. 1995. METIS-unstructured graph partitioning and sparse matrix ordering system version 2.0. (1995).  George Karypis and Vipin Kumar. 1995. METIS-unstructured graph partitioning and sparse matrix ordering system version 2.0. (1995)."},{"key":"e_1_2_1_14_1","volume-title":"A software package for partitioning unstructured graphs, partitioning meshes, and computing fill-reducing orderings of sparse matrices","author":"Karypis George","year":"1998","unstructured":"George Karypis and Vipin Kumar . 1998. A software package for partitioning unstructured graphs, partitioning meshes, and computing fill-reducing orderings of sparse matrices . University of Minnesota, Department of Computer Science and Engineering, Army HPC Research Center , Minneapolis, MN ( 1998 ). George Karypis and Vipin Kumar. 1998. A software package for partitioning unstructured graphs, partitioning meshes, and computing fill-reducing orderings of sparse matrices. University of Minnesota, Department of Computer Science and Engineering, Army HPC Research Center, Minneapolis, MN (1998)."},{"key":"e_1_2_1_15_1","unstructured":"Guillaume Lample Alexis Conneau Marc'Aurelio Ranzato Ludovic Denoyer and Herv\u00e9 J\u00e9gou. 2018. Word translation without parallel data. In ICLR.  Guillaume Lample Alexis Conneau Marc'Aurelio Ranzato Ludovic Denoyer and Herv\u00e9 J\u00e9gou. 2018. Word translation without parallel data. In ICLR."},{"key":"e_1_2_1_16_1","unstructured":"Adam Lerer Ledell Wu Jiajun Shen Timoth\u00e9e Lacroix Luca Wehrstedt Abhijit Bose and Alex Peysakhovich. 2019. Pytorch-BigGraph: A Large Scale Graph Embedding System. In MLSys.  Adam Lerer Ledell Wu Jiajun Shen Timoth\u00e9e Lacroix Luca Wehrstedt Abhijit Bose and Alex Peysakhovich. 2019. Pytorch-BigGraph: A Large Scale Graph Embedding System. In MLSys."},{"key":"e_1_2_1_17_1","doi-asserted-by":"publisher","DOI":"10.1145\/1217299.1217301"},{"key":"e_1_2_1_18_1","volume-title":"Mile: A multi-level framework for scalable graph embedding. arXiv preprint arXiv:1802.09612","author":"Liang Jiongqian","year":"2018","unstructured":"Jiongqian Liang , Saket Gurukar , and Srinivasan Parthasarathy . 2018 . Mile: A multi-level framework for scalable graph embedding. arXiv preprint arXiv:1802.09612 (2018). Jiongqian Liang, Saket Gurukar, and Srinivasan Parthasarathy. 2018. Mile: A multi-level framework for scalable graph embedding. arXiv preprint arXiv:1802.09612 (2018)."},{"key":"e_1_2_1_19_1","doi-asserted-by":"crossref","unstructured":"Grzegorz Malewicz Matthew H. Austern Aart J. C. Bik James C. Dehnert Ilan Horn Naty Leiser and Grzegorz Czajkowski. 2010. Pregel: a system for large-scale graph processing. In SIGMOD. 135--146.  Grzegorz Malewicz Matthew H. Austern Aart J. C. Bik James C. Dehnert Ilan Horn Naty Leiser and Grzegorz Czajkowski. 2010. Pregel: a system for large-scale graph processing. In SIGMOD. 135--146.","DOI":"10.1145\/1807167.1807184"},{"key":"e_1_2_1_20_1","unstructured":"Tong Man Huawei Shen Shenghua Liu Xiaolong Jin and Xueqi Cheng. 2016. Predict Anchor Links across Social Networks via an Embedding Approach. In IJCAI. 1823--1829.  Tong Man Huawei Shen Shenghua Liu Xiaolong Jin and Xueqi Cheng. 2016. Predict Anchor Links across Social Networks via an Embedding Approach. In IJCAI. 1823--1829."},{"key":"e_1_2_1_21_1","volume-title":"Spinner: Scalable graph partitioning in the cloud. In ICDE. 1083--1094.","author":"Martella Claudio","year":"2017","unstructured":"Claudio Martella , Dionysios Logothetis , Andreas Loukas , and Georgos Siganos . 2017 . Spinner: Scalable graph partitioning in the cloud. In ICDE. 1083--1094. Claudio Martella, Dionysios Logothetis, Andreas Loukas, and Georgos Siganos. 2017. Spinner: Scalable graph partitioning in the cloud. In ICDE. 1083--1094."},{"key":"e_1_2_1_22_1","volume-title":"Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781","author":"Mikolov Tomas","year":"2013","unstructured":"Tomas Mikolov , Kai Chen , Greg Corrado , and Jeffrey Dean . 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 ( 2013 ). Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient estimation of word representations in vector space. arXiv preprint arXiv:1301.3781 (2013)."},{"key":"e_1_2_1_23_1","doi-asserted-by":"publisher","DOI":"10.1145\/2939672.2939751"},{"key":"e_1_2_1_24_1","volume-title":"PyTorch: An Imperative Style","author":"Paszke Adam","unstructured":"Adam Paszke , Sam Gross , Francisco Massa , Adam Lerer , James Bradbury , Gregory Chanan , Trevor Killeen , Zeming Lin , Natalia Gimelshein , Luca Antiga , Alban Desmaison , Andreas K\u00f6pf , Edward Yang , Zachary DeVito , Martin Raison , Alykhan Tejani , Sasank Chilamkurthy , Benoit Steiner , Lu Fang , Junjie Bai , and Soumith Chintala . 2019. PyTorch: An Imperative Style , High-Performance Deep Learning Library . In NIPS. 8024--8035. Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas K\u00f6pf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019. PyTorch: An Imperative Style, High-Performance Deep Learning Library. In NIPS. 8024--8035."},{"key":"e_1_2_1_25_1","doi-asserted-by":"crossref","unstructured":"Bryan Perozzi Rami Al-Rfou and Steven Skiena. 2014. DeepWalk: online learning of social representations. In KDD. 701--710.  Bryan Perozzi Rami Al-Rfou and Steven Skiena. 2014. DeepWalk: online learning of social representations. In KDD. 701--710.","DOI":"10.1145\/2623330.2623732"},{"key":"e_1_2_1_26_1","unstructured":"Jiezhong Qiu Yuxiao Dong Hao Ma Jian Li Kuansan Wang and Jie Tang. 2018. Network Embedding as Matrix Factorization: Unifying DeepWalk LINE PTE and node2vec. In WSDM. 459--467.  Jiezhong Qiu Yuxiao Dong Hao Ma Jian Li Kuansan Wang and Jie Tang. 2018. Network Embedding as Matrix Factorization: Unifying DeepWalk LINE PTE and node2vec. In WSDM. 459--467."},{"key":"e_1_2_1_27_1","volume-title":"Sign: Scalable inception graph neural networks. arXiv preprint arXiv:2004.11198","author":"Rossi Emanuele","year":"2020","unstructured":"Emanuele Rossi , Fabrizio Frasca , Ben Chamberlain , Davide Eynard , Michael Bronstein , and Federico Monti . 2020 . Sign: Scalable inception graph neural networks. arXiv preprint arXiv:2004.11198 (2020). Emanuele Rossi, Fabrizio Frasca, Ben Chamberlain, Davide Eynard, Michael Bronstein, and Federico Monti. 2020. Sign: Scalable inception graph neural networks. arXiv preprint arXiv:2004.11198 (2020)."},{"key":"e_1_2_1_28_1","doi-asserted-by":"publisher","DOI":"10.1109\/72.80266"},{"key":"e_1_2_1_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/2736277.2741093"},{"key":"e_1_2_1_30_1","doi-asserted-by":"crossref","unstructured":"Lei Tang and Huan Liu. 2009. Relational learning via latent social dimensions. In KDD. 817--826.  Lei Tang and Huan Liu. 2009. Relational learning via latent social dimensions. In KDD. 817--826.","DOI":"10.1145\/1557019.1557109"},{"key":"e_1_2_1_31_1","volume-title":"Graph attention networks. arXiv preprint arXiv:1710.10903","author":"Veli\u010dkovi\u0107 Petar","year":"2017","unstructured":"Petar Veli\u010dkovi\u0107 , Guillem Cucurull , Arantxa Casanova , Adriana Romero , Pietro Lio , and Yoshua Bengio . 2017. Graph attention networks. arXiv preprint arXiv:1710.10903 ( 2017 ). Petar Veli\u010dkovi\u0107, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017. Graph attention networks. arXiv preprint arXiv:1710.10903 (2017)."},{"key":"e_1_2_1_32_1","first-page":"6861","article-title":"Simplifying Graph Convolutional Networks","volume":"97","author":"Wu Felix","year":"2019","unstructured":"Felix Wu , Amauri H. Souza Jr ., Tianyi Zhang , Christopher Fifty , Tao Yu , and Kilian Q. Weinberger . 2019 . Simplifying Graph Convolutional Networks . In ICML , Vol. 97. 6861 -- 6871 . Felix Wu, Amauri H. Souza Jr., Tianyi Zhang, Christopher Fifty, Tao Yu, and Kilian Q. Weinberger. 2019. Simplifying Graph Convolutional Networks. In ICML, Vol. 97. 6861--6871.","journal-title":"ICML"},{"key":"e_1_2_1_33_1","doi-asserted-by":"publisher","DOI":"10.1145\/2934664"},{"key":"e_1_2_1_34_1","volume-title":"DistDGL: Distributed Graph Neural Network Training for Billion-Scale Graphs. arXiv preprint arXiv:2010.05337","author":"Zheng Da","year":"2020","unstructured":"Da Zheng , Chao Ma , Minjie Wang , Jinjing Zhou , Qidong Su , Xiang Song , Quan Gan , Zheng Zhang , and George Karypis . 2020. DistDGL: Distributed Graph Neural Network Training for Billion-Scale Graphs. arXiv preprint arXiv:2010.05337 ( 2020 ). Da Zheng, Chao Ma, Minjie Wang, Jinjing Zhou, Qidong Su, Xiang Song, Quan Gan, Zheng Zhang, and George Karypis. 2020. DistDGL: Distributed Graph Neural Network Training for Billion-Scale Graphs. arXiv preprint arXiv:2010.05337 (2020)."}],"container-title":["Proceedings of the VLDB Endowment"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/dl.acm.org\/doi\/pdf\/10.14778\/3503585.3503599","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,12,28]],"date-time":"2022-12-28T10:30:37Z","timestamp":1672223437000},"score":1,"resource":{"primary":{"URL":"https:\/\/dl.acm.org\/doi\/10.14778\/3503585.3503599"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2021,12]]},"references-count":34,"journal-issue":{"issue":"4","published-print":{"date-parts":[[2021,12]]}},"alternative-id":["10.14778\/3503585.3503599"],"URL":"https:\/\/doi.org\/10.14778\/3503585.3503599","relation":{},"ISSN":["2150-8097"],"issn-type":[{"value":"2150-8097","type":"print"}],"subject":[],"published":{"date-parts":[[2021,12]]}}}