Personalized Turn-Level User Conversation Satisfaction Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Zhefan, Guo, Zhiqiang, Ma, Weizhi, Zhang, Min, Yan, Quanjia, Luo, Hengliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
by: Wan, Luanbo, et al.
Published: (2025)
by: Wan, Luanbo, et al.
Published: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026)
by: Zheng, Mingqian, et al.
Published: (2026)
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
by: Li, Li, et al.
Published: (2025)
by: Li, Li, et al.
Published: (2025)
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
by: Mao, Hanyi, et al.
Published: (2025)
by: Mao, Hanyi, et al.
Published: (2025)
MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents
by: Hu, Tianyu, et al.
Published: (2026)
by: Hu, Tianyu, et al.
Published: (2026)
Adaptive Layer-skipping in Pre-trained LLMs
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
Direct Multi-Token Decoding
by: Luo, Xuan, et al.
Published: (2025)
by: Luo, Xuan, et al.
Published: (2025)
OP-Bench: Benchmarking Over-Personalization for Memory-Augmented Personalized Conversational Agents
by: Hu, Yulin, et al.
Published: (2026)
by: Hu, Yulin, et al.
Published: (2026)
Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
by: Liu, Geng, et al.
Published: (2026)
by: Liu, Geng, et al.
Published: (2026)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
by: Katsis, Yannis, et al.
Published: (2025)
by: Katsis, Yannis, et al.
Published: (2025)
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
by: Liu, Junyu, et al.
Published: (2026)
by: Liu, Junyu, et al.
Published: (2026)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
by: Sirdeshmukh, Ved, et al.
Published: (2025)
by: Sirdeshmukh, Ved, et al.
Published: (2025)
In2x at WMT25 Translation Task
by: Pang, Lei, et al.
Published: (2025)
by: Pang, Lei, et al.
Published: (2025)
Do Ethical AI Principles Matter to Users? A Large-Scale Analysis of User Sentiment and Satisfaction
by: Pasch, Stefan, et al.
Published: (2025)
by: Pasch, Stefan, et al.
Published: (2025)
ES-MemEval: Benchmarking Conversational Agents on Personalized Long-Term Emotional Support
by: Chen, Tiantian, et al.
Published: (2026)
by: Chen, Tiantian, et al.
Published: (2026)
LTLBench: Towards Benchmarks for Evaluating Temporal Reasoning in Large Language Models
by: Tang, Weizhi, et al.
Published: (2024)
by: Tang, Weizhi, et al.
Published: (2024)
StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning
by: Yu, Yuanqing, et al.
Published: (2024)
by: Yu, Yuanqing, et al.
Published: (2024)
C3: A Bilingual Benchmark for Spoken Dialogue Models Exploring Challenges in Complex Conversations
by: Ma, Chengqian, et al.
Published: (2025)
by: Ma, Chengqian, et al.
Published: (2025)
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
by: Wang, Zhebo, et al.
Published: (2026)
by: Wang, Zhebo, et al.
Published: (2026)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
by: Guan, Shengyue, et al.
Published: (2025)
by: Guan, Shengyue, et al.
Published: (2025)
PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
by: Chai, Huacan, et al.
Published: (2025)
by: Chai, Huacan, et al.
Published: (2025)
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
by: Wan, Yanming, et al.
Published: (2025)
by: Wan, Yanming, et al.
Published: (2025)
PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems
by: Yu, Jiongchi, et al.
Published: (2026)
by: Yu, Jiongchi, et al.
Published: (2026)
From 1,000,000 Users to Every User: Scaling Up Personalized Preference for User-level Alignment
by: Li, Jia-Nan, et al.
Published: (2025)
by: Li, Jia-Nan, et al.
Published: (2025)
Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History
by: Kim, Serin, et al.
Published: (2026)
by: Kim, Serin, et al.
Published: (2026)
TD-EVAL: Revisiting Task-Oriented Dialogue Evaluation by Combining Turn-Level Precision with Dialogue-Level Comparisons
by: Acikgoz, Emre Can, et al.
Published: (2025)
by: Acikgoz, Emre Can, et al.
Published: (2025)
One Size doesn't Fit All: A Personalized Conversational Tutoring Agent for Mathematics Instruction
by: Liu, Ben, et al.
Published: (2025)
by: Liu, Ben, et al.
Published: (2025)
On Memory Construction and Retrieval for Personalized Conversational Agents
by: Pan, Zhuoshi, et al.
Published: (2025)
by: Pan, Zhuoshi, et al.
Published: (2025)
Human Latency Conversational Turns for Spoken Avatar Systems
by: Jacoby, Derek, et al.
Published: (2024)
by: Jacoby, Derek, et al.
Published: (2024)
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
by: Chen, Xingwu, et al.
Published: (2026)
by: Chen, Xingwu, et al.
Published: (2026)
T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
by: Chakraborty, Amartya, et al.
Published: (2025)
by: Chakraborty, Amartya, et al.
Published: (2025)
Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching
by: Savage, Thomas
Published: (2025)
by: Savage, Thomas
Published: (2025)
Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization
by: Xiao, Quanjia, et al.
Published: (2026)
by: Xiao, Quanjia, et al.
Published: (2026)
Unveiling the Impact of Multi-Modal Interactions on User Engagement: A Comprehensive Evaluation in AI-driven Conversations
by: Zhang, Lichao, et al.
Published: (2024)
by: Zhang, Lichao, et al.
Published: (2024)
Reasoning-Based Personalized Generation for Users with Sparse Data
by: Ni, Bo, et al.
Published: (2026)
by: Ni, Bo, et al.
Published: (2026)
OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations
by: Chen, Jiangwang, et al.
Published: (2026)
by: Chen, Jiangwang, et al.
Published: (2026)
AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents
by: Kim, Takyoung, et al.
Published: (2025)
by: Kim, Takyoung, et al.
Published: (2025)
PersonaLens: A Benchmark for Personalization Evaluation in Conversational AI Assistants
by: Zhao, Zheng, et al.
Published: (2025)
by: Zhao, Zheng, et al.
Published: (2025)
ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
by: Lei, Yiming, et al.
Published: (2025)
by: Lei, Yiming, et al.
Published: (2025)
Similar Items
-
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
by: Wan, Luanbo, et al.
Published: (2025) -
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
by: Zheng, Mingqian, et al.
Published: (2026) -
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
by: Li, Li, et al.
Published: (2025) -
Clip Your Sequences Fairly: Enforcing Length Fairness for Sequence-Level RL
by: Mao, Hanyi, et al.
Published: (2025) -
MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents
by: Hu, Tianyu, et al.
Published: (2026)