Bridging Academia and Industry: A Comprehensive Benchmark for Attributed Graph Clustering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yunhui, Qiu, Pengyu, Xing, Yu, Liu, Yongchao, Du, Peng, Hong, Chuntao, Zheng, Jiajun, Zheng, Tao, He, Tieke
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911434268475392
author Liu, Yunhui
Qiu, Pengyu
Xing, Yu
Liu, Yongchao
Du, Peng
Hong, Chuntao
Zheng, Jiajun
Zheng, Tao
He, Tieke
author_facet Liu, Yunhui
Qiu, Pengyu
Xing, Yu
Liu, Yongchao
Du, Peng
Hong, Chuntao
Zheng, Jiajun
Zheng, Tao
He, Tieke
contents Attributed Graph Clustering (AGC) is a fundamental unsupervised task that integrates structural topology and node attributes to uncover latent patterns in graph-structured data. Despite its significance in industrial applications such as fraud detection and user segmentation, a significant chasm persists between academic research and real-world deployment. Current evaluation protocols suffer from the small-scale, high-homophily citation datasets, non-scalable full-batch training paradigms, and a reliance on supervised metrics that fail to reflect performance in label-scarce environments. To bridge these gaps, we present PyAGC, a comprehensive, production-ready benchmark and library designed to stress-test AGC methods across diverse scales and structural properties. We unify existing methodologies into a modular Encode-Cluster-Optimize framework and, for the first time, provide memory-efficient, mini-batch implementations for a wide array of state-of-the-art AGC algorithms. Our benchmark curates 12 diverse datasets, ranging from 2.7K to 111M nodes, specifically incorporating industrial graphs with complex tabular features and low homophily. Furthermore, we advocate for a holistic evaluation protocol that mandates unsupervised structural metrics and efficiency profiling alongside traditional supervised metrics. Battle-tested in high-stakes industrial workflows at Ant Group, this benchmark offers the community a robust, reproducible, and scalable platform to advance AGC research towards realistic deployment. The code and resources are publicly available via GitHub (https://github.com/Cloudy1225/PyAGC), PyPI (https://pypi.org/project/pyagc), and Documentation (https://pyagc.readthedocs.io).
format Preprint
id arxiv_https___arxiv_org_abs_2602_08519
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Bridging Academia and Industry: A Comprehensive Benchmark for Attributed Graph Clustering
Liu, Yunhui
Qiu, Pengyu
Xing, Yu
Liu, Yongchao
Du, Peng
Hong, Chuntao
Zheng, Jiajun
Zheng, Tao
He, Tieke
Machine Learning
Attributed Graph Clustering (AGC) is a fundamental unsupervised task that integrates structural topology and node attributes to uncover latent patterns in graph-structured data. Despite its significance in industrial applications such as fraud detection and user segmentation, a significant chasm persists between academic research and real-world deployment. Current evaluation protocols suffer from the small-scale, high-homophily citation datasets, non-scalable full-batch training paradigms, and a reliance on supervised metrics that fail to reflect performance in label-scarce environments. To bridge these gaps, we present PyAGC, a comprehensive, production-ready benchmark and library designed to stress-test AGC methods across diverse scales and structural properties. We unify existing methodologies into a modular Encode-Cluster-Optimize framework and, for the first time, provide memory-efficient, mini-batch implementations for a wide array of state-of-the-art AGC algorithms. Our benchmark curates 12 diverse datasets, ranging from 2.7K to 111M nodes, specifically incorporating industrial graphs with complex tabular features and low homophily. Furthermore, we advocate for a holistic evaluation protocol that mandates unsupervised structural metrics and efficiency profiling alongside traditional supervised metrics. Battle-tested in high-stakes industrial workflows at Ant Group, this benchmark offers the community a robust, reproducible, and scalable platform to advance AGC research towards realistic deployment. The code and resources are publicly available via GitHub (https://github.com/Cloudy1225/PyAGC), PyPI (https://pypi.org/project/pyagc), and Documentation (https://pyagc.readthedocs.io).
title Bridging Academia and Industry: A Comprehensive Benchmark for Attributed Graph Clustering
topic Machine Learning
url https://arxiv.org/abs/2602.08519