Large-Scale Graphs Community Detection using Spark GraphFrames
Fuente:
arXiv
Guardado en:
| Autores principales: | , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866917744039952384 |
|---|---|
| author | Apostol, Elena-Simona Cojocaru, Adrian-Cosmin Truică, Ciprian-Octavian |
| author_facet | Apostol, Elena-Simona Cojocaru, Adrian-Cosmin Truică, Ciprian-Octavian |
| contents | With the emergence of social networks, online platforms dedicated to different use cases, and sensor networks, the emergence of large-scale graph community detection has become a steady field of research with real-world applications. Community detection algorithms have numerous practical applications, particularly due to their scalability with data size. Nonetheless, a notable drawback of community detection algorithms is their computational intensity~\cite{Apostol2014}, resulting in decreasing performance as data size increases. For this purpose, new frameworks that employ distributed systems such as Apache Hadoop and Apache Spark which can seamlessly handle large-scale graphs must be developed. In this paper, we propose a novel framework for community detection algorithms, i.e., K-Cliques, Louvain, and Fast Greedy, developed using Apache Spark GraphFrames. We test their performance and scalability on two real-world datasets. The experimental results prove the feasibility of developing graph mining algorithms using Apache Spark GraphFrames. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2408_03966 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Large-Scale Graphs Community Detection using Spark GraphFrames Apostol, Elena-Simona Cojocaru, Adrian-Cosmin Truică, Ciprian-Octavian Social and Information Networks Distributed, Parallel, and Cluster Computing With the emergence of social networks, online platforms dedicated to different use cases, and sensor networks, the emergence of large-scale graph community detection has become a steady field of research with real-world applications. Community detection algorithms have numerous practical applications, particularly due to their scalability with data size. Nonetheless, a notable drawback of community detection algorithms is their computational intensity~\cite{Apostol2014}, resulting in decreasing performance as data size increases. For this purpose, new frameworks that employ distributed systems such as Apache Hadoop and Apache Spark which can seamlessly handle large-scale graphs must be developed. In this paper, we propose a novel framework for community detection algorithms, i.e., K-Cliques, Louvain, and Fast Greedy, developed using Apache Spark GraphFrames. We test their performance and scalability on two real-world datasets. The experimental results prove the feasibility of developing graph mining algorithms using Apache Spark GraphFrames. |
| title | Large-Scale Graphs Community Detection using Spark GraphFrames |
| topic | Social and Information Networks Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2408.03966 |