UserSumBench: A Benchmark Framework for Evaluating User Summarization Approaches
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Chao, Wu, Neo, Ning, Lin, Wu, Jiaxing, Liu, Luyang, Xie, Jun, O'Banion, Shawn, Green, Bradley |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
by: Wu, Jiaxing, et al.
Published: (2024)
by: Wu, Jiaxing, et al.
Published: (2024)
User-LLM: Efficient LLM Contextualization with User Embeddings
by: Ning, Lin, et al.
Published: (2024)
by: Ning, Lin, et al.
Published: (2024)
Enhancing Math through Literature.
by: O'Banion, Carie
Published: (1997)
by: O'Banion, Carie
Published: (1997)
CI-Bench: Benchmarking Contextual Integrity of AI Assistants on Synthetic Data
by: Cheng, Zhao, et al.
Published: (2024)
by: Cheng, Zhao, et al.
Published: (2024)
BookAsSumQA: An Evaluation Framework for Aspect-Based Book Summarization via Question Answering
by: Miyazato, Ryuhei, et al.
Published: (2025)
by: Miyazato, Ryuhei, et al.
Published: (2025)
Deliberation in Latent Space via Differentiable Cache Augmentation
by: Liu, Luyang, et al.
Published: (2024)
by: Liu, Luyang, et al.
Published: (2024)
PersonalSum: A User-Subjective Guided Personalized Summarization Dataset for Large Language Models
by: Zhang, Lemei, et al.
Published: (2024)
by: Zhang, Lemei, et al.
Published: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
by: Lin, Bill Yuchen, et al.
Published: (2024)
by: Lin, Bill Yuchen, et al.
Published: (2024)
Enhancing User Sequence Modeling through Barlow Twins-based Self-Supervised Learning
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
LLM Benchmark-User Need Misalignment for Climate Change
by: Liu, Oucheng, et al.
Published: (2026)
by: Liu, Oucheng, et al.
Published: (2026)
EmbSum: Leveraging the Summarization Capabilities of Large Language Models for Content-Based Recommendations
by: Zhang, Chiyu, et al.
Published: (2024)
by: Zhang, Chiyu, et al.
Published: (2024)
UserBench: An Interactive Gym Environment for User-Centric Agents
by: Qian, Cheng, et al.
Published: (2025)
by: Qian, Cheng, et al.
Published: (2025)
Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations
by: Liu, Yu Lu, et al.
Published: (2026)
by: Liu, Yu Lu, et al.
Published: (2026)
FAB-Bench: A Framework for Adaptive RAG Benchmarking in Semiconductor Manufacturing
by: Qian, Jingbin, et al.
Published: (2026)
by: Qian, Jingbin, et al.
Published: (2026)
PosterSum: A Multimodal Benchmark for Scientific Poster Summarization
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
CroCoSum: A Benchmark Dataset for Cross-Lingual Code-Switched Summarization
by: Zhang, Ruochen, et al.
Published: (2023)
by: Zhang, Ruochen, et al.
Published: (2023)
From 1,000,000 Users to Every User: Scaling Up Personalized Preference for User-level Alignment
by: Li, Jia-Nan, et al.
Published: (2025)
by: Li, Jia-Nan, et al.
Published: (2025)
ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
by: Wu, Xixi, et al.
Published: (2025)
by: Wu, Xixi, et al.
Published: (2025)
I Need Help! Evaluating LLM's Ability to Ask for Users' Support: A Case Study on Text-to-SQL Generation
by: Wu, Cheng-Kuang, et al.
Published: (2024)
by: Wu, Cheng-Kuang, et al.
Published: (2024)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
by: Bao, Forrest Sheng, et al.
Published: (2024)
by: Bao, Forrest Sheng, et al.
Published: (2024)
Eustachian Tube Dysfunction Questionnaire Score Changes at 6 and 12 Weeks: Follow‐Up Implications
by: Alexander R. Gomez‐Lara, et al.
Published: (2026)
by: Alexander R. Gomez‐Lara, et al.
Published: (2026)
DomainSum: A Hierarchical Benchmark for Fine-Grained Domain Shift in Abstractive Text Summarization
by: Yuan, Haohan, et al.
Published: (2024)
by: Yuan, Haohan, et al.
Published: (2024)
StreamProfileBench: A Benchmark for Fine-Grained User Profile Inference in Real-World Streaming Scenarios
by: Wang, Sizhe, et al.
Published: (2026)
by: Wang, Sizhe, et al.
Published: (2026)
SUMIE: A Synthetic Benchmark for Incremental Entity Summarization
by: Hwang, Eunjeong, et al.
Published: (2024)
by: Hwang, Eunjeong, et al.
Published: (2024)
BenchBench: Benchmarking Automated Benchmark Generation
by: Zheng, Yandan, et al.
Published: (2026)
by: Zheng, Yandan, et al.
Published: (2026)
SalesSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators
by: Pruksachatkun, Yada, et al.
Published: (2026)
by: Pruksachatkun, Yada, et al.
Published: (2026)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
by: Perlitz, Yotam, et al.
Published: (2024)
by: Perlitz, Yotam, et al.
Published: (2024)
CRAB-Bench: Evaluating LLM Agents under Complex Task Dependencies and Human-aligned User Simulation
by: Wang, Danqing, et al.
Published: (2026)
by: Wang, Danqing, et al.
Published: (2026)
AgenticSum: An Agentic Inference-Time Framework for Faithful Clinical Text Summarization
by: Piya, Fahmida Liza, et al.
Published: (2026)
by: Piya, Fahmida Liza, et al.
Published: (2026)
SPECTRA: Revealing the Full Spectrum of User Preferences via Distributional LLM Inference
by: Zhang, Luyang, et al.
Published: (2025)
by: Zhang, Luyang, et al.
Published: (2025)
TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domain
by: Chu, Bohao, et al.
Published: (2025)
by: Chu, Bohao, et al.
Published: (2025)
A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models
by: Wang, Jiayin, et al.
Published: (2024)
by: Wang, Jiayin, et al.
Published: (2024)
Extract-and-Abstract: Unifying Extractive and Abstractive Summarization within Single Encoder-Decoder Framework
by: Wu, Yuping, et al.
Published: (2024)
by: Wu, Yuping, et al.
Published: (2024)
CS-Sum: A Benchmark for Code-Switching Dialogue Summarization and the Limits of Large Language Models
by: Suresh, Sathya Krishnan, et al.
Published: (2025)
by: Suresh, Sathya Krishnan, et al.
Published: (2025)
OrderSum: Semantic Sentence Ordering for Extractive Summarization
by: Kwon, Taewan, et al.
Published: (2025)
by: Kwon, Taewan, et al.
Published: (2025)
SumHiS: Extractive Summarization Exploiting Hidden Structure
by: Pavel, Tikhonov, et al.
Published: (2024)
by: Pavel, Tikhonov, et al.
Published: (2024)
Diversity Augmentation of Dynamic User Preference Data for Boosting Personalized Text Summarizers
by: Chatterjee, Parthiv, et al.
Published: (2025)
by: Chatterjee, Parthiv, et al.
Published: (2025)
Beyond One-Size-Fits-All Summarization: Customizing Summaries for Diverse Users
by: Duran, Mehmet Samet, et al.
Published: (2025)
by: Duran, Mehmet Samet, et al.
Published: (2025)
Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows
by: Chen, Hardy, et al.
Published: (2026)
by: Chen, Hardy, et al.
Published: (2026)
Similar Items
-
RLPF: Reinforcement Learning from Prediction Feedback for User Summarization with LLMs
by: Wu, Jiaxing, et al.
Published: (2024) -
User-LLM: Efficient LLM Contextualization with User Embeddings
by: Ning, Lin, et al.
Published: (2024) -
Enhancing Math through Literature.
by: O'Banion, Carie
Published: (1997) -
CI-Bench: Benchmarking Contextual Integrity of AI Assistants on Synthetic Data
by: Cheng, Zhao, et al.
Published: (2024) -
BookAsSumQA: An Evaluation Framework for Aspect-Based Book Summarization via Question Answering
by: Miyazato, Ryuhei, et al.
Published: (2025)