CC-GSEO-Bench: A Content-Centric Benchmark for Measuring Source Influence in Generative Search Engines

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Qiyuan, Chen, Jiahe, Huang, Hongsen, Shao, Qian, Chen, Jintai, Hua, Renjie, Xu, Hongxia, Wu, Ruijia, Chuan, Ren, Wu, Jian
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918263955390464
author Chen, Qiyuan
Chen, Jiahe
Huang, Hongsen
Shao, Qian
Chen, Jintai
Hua, Renjie
Xu, Hongxia
Wu, Ruijia
Chuan, Ren
Wu, Jian
author_facet Chen, Qiyuan
Chen, Jiahe
Huang, Hongsen
Shao, Qian
Chen, Jintai
Hua, Renjie
Xu, Hongxia
Wu, Ruijia
Chuan, Ren
Wu, Jian
contents Generative Search Engines (GSEs) synthesize conversational answers from multiple sources, weakening the long-standing link between search ranking and digital visibility. This shift raises a central question for content creators: How can we reliably quantify a source article's influence on a GSE's synthesized answer across diverse intents and follow-up questions? We introduce CC-GSEO-Bench, a content-centric benchmark that couples a large-scale dataset with a creator-centered evaluation framework. The dataset contains over 1,000 source articles and over 5,000 query-article pairs, organized in a one-to-many structure for article-level evaluation. We ground construction in realistic retrieval by combining seed queries from public QA datasets with limited synthesized augmentation and retaining only queries whose paired source reappears in a follow-up retrieval step. On top of this dataset, we operationalize influence along three core dimensions: Exposure, Faithful Credit, and Causal Impact, and two content-quality dimensions: Readability and Structure, and Trustworthiness and Safety. We aggregate query-level signals over each article's query cluster to summarize influence strength, coverage, and stability, and empirically characterize influence dynamics across representative content patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2509_05607
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CC-GSEO-Bench: A Content-Centric Benchmark for Measuring Source Influence in Generative Search Engines
Chen, Qiyuan
Chen, Jiahe
Huang, Hongsen
Shao, Qian
Chen, Jintai
Hua, Renjie
Xu, Hongxia
Wu, Ruijia
Chuan, Ren
Wu, Jian
Computation and Language
Generative Search Engines (GSEs) synthesize conversational answers from multiple sources, weakening the long-standing link between search ranking and digital visibility. This shift raises a central question for content creators: How can we reliably quantify a source article's influence on a GSE's synthesized answer across diverse intents and follow-up questions? We introduce CC-GSEO-Bench, a content-centric benchmark that couples a large-scale dataset with a creator-centered evaluation framework. The dataset contains over 1,000 source articles and over 5,000 query-article pairs, organized in a one-to-many structure for article-level evaluation. We ground construction in realistic retrieval by combining seed queries from public QA datasets with limited synthesized augmentation and retaining only queries whose paired source reappears in a follow-up retrieval step. On top of this dataset, we operationalize influence along three core dimensions: Exposure, Faithful Credit, and Causal Impact, and two content-quality dimensions: Readability and Structure, and Trustworthiness and Safety. We aggregate query-level signals over each article's query cluster to summarize influence strength, coverage, and stability, and empirically characterize influence dynamics across representative content patterns.
title CC-GSEO-Bench: A Content-Centric Benchmark for Measuring Source Influence in Generative Search Engines
topic Computation and Language
url https://arxiv.org/abs/2509.05607