GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Silan, Zhang, Shiqi, Shi, Yimin, Xiao, Xiaokui |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
by: Hu, Caiyu, et al.
Published: (2025)
by: Hu, Caiyu, et al.
Published: (2025)
You Are What You Bought: Generating Customer Personas for E-commerce Applications
by: Shi, Yimin, et al.
Published: (2025)
by: Shi, Yimin, et al.
Published: (2025)
Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025)
by: Tang, Minghao, et al.
Published: (2025)
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
by: Chen, Jianlyu, et al.
Published: (2024)
by: Chen, Jianlyu, et al.
Published: (2024)
Evaluating Generative Ad Hoc Information Retrieval
by: Gienapp, Lukas, et al.
Published: (2023)
by: Gienapp, Lukas, et al.
Published: (2023)
Diagnosing and Repairing Citation Failures in Generative Engine Optimization
by: Tian, Zhihua, et al.
Published: (2026)
by: Tian, Zhihua, et al.
Published: (2026)
KET-RAG: A Cost-Efficient Multi-Granular Indexing Framework for Graph-RAG
by: Huang, Yiqian, et al.
Published: (2025)
by: Huang, Yiqian, et al.
Published: (2025)
From Relevance to Authority: Authority-aware Generative Retrieval in Web Search Engines
by: Lee, Sunkyung, et al.
Published: (2026)
by: Lee, Sunkyung, et al.
Published: (2026)
Retrieval Augmented Generation with Collaborative Filtering for Personalized Text Generation
by: Shi, Teng, et al.
Published: (2025)
by: Shi, Teng, et al.
Published: (2025)
Exposing Citation Vulnerabilities in Generative Engines
by: Mochizuki, Riku, et al.
Published: (2025)
by: Mochizuki, Riku, et al.
Published: (2025)
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
by: Du, Mingxuan, et al.
Published: (2025)
by: Du, Mingxuan, et al.
Published: (2025)
FAB-Bench: A Framework for Adaptive RAG Benchmarking in Semiconductor Manufacturing
by: Qian, Jingbin, et al.
Published: (2026)
by: Qian, Jingbin, et al.
Published: (2026)
ClarifyMT-Bench: Benchmarking and Improving Multi-Turn Clarification for Conversational Large Language Models
by: Luo, Sichun, et al.
Published: (2025)
by: Luo, Sichun, et al.
Published: (2025)
Evaluating Robustness of Generative Search Engine on Adversarial Factual Questions
by: Hu, Xuming, et al.
Published: (2024)
by: Hu, Xuming, et al.
Published: (2024)
RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment
by: Jin, Zhuoran, et al.
Published: (2024)
by: Jin, Zhuoran, et al.
Published: (2024)
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
by: Xi, Yunjia, et al.
Published: (2025)
by: Xi, Yunjia, et al.
Published: (2025)
BERGEN: A Benchmarking Library for Retrieval-Augmented Generation
by: Rau, David, et al.
Published: (2024)
by: Rau, David, et al.
Published: (2024)
Controlling Output Rankings in Generative Engines for LLM-based Search
by: Jin, Haibo, et al.
Published: (2026)
by: Jin, Haibo, et al.
Published: (2026)
LLaDA-Rec: Discrete Diffusion for Parallel Semantic ID Generation in Generative Recommendation
by: Shi, Teng, et al.
Published: (2025)
by: Shi, Teng, et al.
Published: (2025)
LexRAG: Benchmarking Retrieval-Augmented Generation in Multi-Turn Legal Consultation Conversation
by: Li, Haitao, et al.
Published: (2025)
by: Li, Haitao, et al.
Published: (2025)
Generating Diverse Q&A Benchmarks for RAG Evaluation with DataMorgana
by: Filice, Simone, et al.
Published: (2025)
by: Filice, Simone, et al.
Published: (2025)
GINGER: Grounded Information Nugget-Based Generation of Responses
by: Łajewska, Weronika, et al.
Published: (2025)
by: Łajewska, Weronika, et al.
Published: (2025)
Cocktail: A Comprehensive Information Retrieval Benchmark with LLM-Generated Documents Integration
by: Dai, Sunhao, et al.
Published: (2024)
by: Dai, Sunhao, et al.
Published: (2024)
CORAL: Benchmarking Multi-turn Conversational Retrieval-Augmentation Generation
by: Cheng, Yiruo, et al.
Published: (2024)
by: Cheng, Yiruo, et al.
Published: (2024)
DomainRAG: A Chinese Benchmark for Evaluating Domain-specific Retrieval-Augmented Generation
by: Wang, Shuting, et al.
Published: (2024)
by: Wang, Shuting, et al.
Published: (2024)
Spacerini: Plug-and-play Search Engines with Pyserini and Hugging Face
by: Akiki, Christopher, et al.
Published: (2023)
by: Akiki, Christopher, et al.
Published: (2023)
GISA: A Benchmark for General Information-Seeking Assistant
by: Zhu, Yutao, et al.
Published: (2026)
by: Zhu, Yutao, et al.
Published: (2026)
Generate-then-Ground in Retrieval-Augmented Generation for Multi-hop Question Answering
by: Shi, Zhengliang, et al.
Published: (2024)
by: Shi, Zhengliang, et al.
Published: (2024)
UniConv: Unifying Retrieval and Response Generation for Large Language Models in Conversations
by: Mo, Fengran, et al.
Published: (2025)
by: Mo, Fengran, et al.
Published: (2025)
QP-OneModel: A Unified Generative LLM for Multi-Task Query Understanding in Xiaohongshu Search
by: Huang, Jianzhao, et al.
Published: (2026)
by: Huang, Jianzhao, et al.
Published: (2026)
Yesterday's News: Benchmarking Multi-Dimensional Out-of-Distribution Generalization of Misinformation Detection Models
by: Verhoeven, Ivo, et al.
Published: (2024)
by: Verhoeven, Ivo, et al.
Published: (2024)
JuDGE: Benchmarking Judgment Document Generation for Chinese Legal System
by: Su, Weihang, et al.
Published: (2025)
by: Su, Weihang, et al.
Published: (2025)
Towards Reliable and Factual Response Generation: Detecting Unanswerable Questions in Information-Seeking Conversations
by: Łajewska, Weronika, et al.
Published: (2024)
by: Łajewska, Weronika, et al.
Published: (2024)
SAGE: Benchmarking and Improving Retrieval for Deep Research Agents
by: Hu, Tiansheng, et al.
Published: (2026)
by: Hu, Tiansheng, et al.
Published: (2026)
HetaRAG: Hybrid Deep Retrieval-Augmented Generation across Heterogeneous Data Stores
by: Yan, Guohang, et al.
Published: (2025)
by: Yan, Guohang, et al.
Published: (2025)
LegalAgentBench: Evaluating LLM Agents in Legal Domain
by: Li, Haitao, et al.
Published: (2024)
by: Li, Haitao, et al.
Published: (2024)
NLP-Powered Repository and Search Engine for Academic Papers: A Case Study on Cyber Risk Literature with CyLit
by: Zhang, Linfeng, et al.
Published: (2024)
by: Zhang, Linfeng, et al.
Published: (2024)
Environmental large language model Evaluation (ELLE) dataset: A Benchmark for Evaluating Generative AI applications in Eco-environment Domain
by: Guo, Jing, et al.
Published: (2025)
by: Guo, Jing, et al.
Published: (2025)
RAR-b: Reasoning as Retrieval Benchmark
by: Xiao, Chenghao, et al.
Published: (2024)
by: Xiao, Chenghao, et al.
Published: (2024)
FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation
by: Zhang, Zhuocheng, et al.
Published: (2025)
by: Zhang, Zhuocheng, et al.
Published: (2025)
Similar Items
-
MCiteBench: A Multimodal Benchmark for Generating Text with Citations
by: Hu, Caiyu, et al.
Published: (2025) -
You Are What You Bought: Generating Customer Personas for E-commerce Applications
by: Shi, Yimin, et al.
Published: (2025) -
Injecting External Knowledge into the Reasoning Process Enhances Retrieval-Augmented Generation
by: Tang, Minghao, et al.
Published: (2025) -
AIR-Bench: Automated Heterogeneous Information Retrieval Benchmark
by: Chen, Jianlyu, et al.
Published: (2024) -
Evaluating Generative Ad Hoc Information Retrieval
by: Gienapp, Lukas, et al.
Published: (2023)