Amulet: Putting Complex Multi-Turn Conversations on the Stand with LLM Juries
Fuente:
arXiv
Saved in:
| Main Authors: | Ramnath, Sahana, Mudgil, Anurag, Joshi, Brihi, Hallinan, Skyler, Ren, Xiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tailoring Self-Rationalizers with Multi-Reward Distillation
by: Ramnath, Sahana, et al.
Published: (2023)
by: Ramnath, Sahana, et al.
Published: (2023)
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
by: Hallinan, Skyler, et al.
Published: (2025)
by: Hallinan, Skyler, et al.
Published: (2025)
CAVE: Controllable Authorship Verification Explanations
by: Ramnath, Sahana, et al.
Published: (2024)
by: Ramnath, Sahana, et al.
Published: (2024)
Improving Language Model Personas via Rationalization with Psychological Scaffolds
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
PrimeX: A Dataset of Worldview, Opinion, and Explanation
by: Koncel-Kedziorski, Rik, et al.
Published: (2025)
by: Koncel-Kedziorski, Rik, et al.
Published: (2025)
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
by: He, Keyu, et al.
Published: (2025)
by: He, Keyu, et al.
Published: (2025)
OpaqueToolsBench: Learning Nuances of Tool Behavior Through Interaction
by: Hallinan, Skyler, et al.
Published: (2026)
by: Hallinan, Skyler, et al.
Published: (2026)
CoKe: Customizable Fine-Grained Story Evaluation via Chain-of-Keyword Rationalization
by: Joshi, Brihi, et al.
Published: (2025)
by: Joshi, Brihi, et al.
Published: (2025)
StyleRemix: Interpretable Authorship Obfuscation via Distillation and Perturbation of Style Elements
by: Fisher, Jillian, et al.
Published: (2024)
by: Fisher, Jillian, et al.
Published: (2024)
'Put the Car on the Stand': SMT-based Oracles for Investigating Decisions
by: Judson, Samuel, et al.
Published: (2023)
by: Judson, Samuel, et al.
Published: (2023)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
by: Guan, Shengyue, et al.
Published: (2025)
by: Guan, Shengyue, et al.
Published: (2025)
OATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM Assistants
by: Ranjit, Jaspreet, et al.
Published: (2024)
by: Ranjit, Jaspreet, et al.
Published: (2024)
Eliciting Behaviors in Multi-Turn Conversations
by: Huang, Jing, et al.
Published: (2025)
by: Huang, Jing, et al.
Published: (2025)
LLMs Get Lost In Multi-Turn Conversation
by: Laban, Philippe, et al.
Published: (2025)
by: Laban, Philippe, et al.
Published: (2025)
Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation
by: Luo, Jiani, et al.
Published: (2026)
by: Luo, Jiani, et al.
Published: (2026)
ToolWeave: Structured Synthesis of Complex Multi-Turn Tool-Calling Dialogues
by: Khandelwal, Dinesh, et al.
Published: (2026)
by: Khandelwal, Dinesh, et al.
Published: (2026)
Enabling Precise Topic Alignment in Large Language Models Via Sparse Autoencoders
by: Joshi, Ananya, et al.
Published: (2025)
by: Joshi, Ananya, et al.
Published: (2025)
FlowKV: Enhancing Multi-Turn Conversational Coherence in LLMs via Isolated Key-Value Cache Management
by: Liu, Xiang, et al.
Published: (2025)
by: Liu, Xiang, et al.
Published: (2025)
MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations
by: Lumer, Elias, et al.
Published: (2025)
by: Lumer, Elias, et al.
Published: (2025)
Proactive Guidance of Multi-Turn Conversation in Industrial Search
by: Li, Xiaoyu, et al.
Published: (2025)
by: Li, Xiaoyu, et al.
Published: (2025)
Expectation Confirmation Preference Optimization for Multi-Turn Conversational Recommendation Agent
by: Feng, Xueyang, et al.
Published: (2025)
by: Feng, Xueyang, et al.
Published: (2025)
Found in Conversation: LLMs Teach Themselves to Close the Multi-Turn Gap
by: Chen, Tianlang, et al.
Published: (2026)
by: Chen, Tianlang, et al.
Published: (2026)
MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation
by: Singh, Jyotika, et al.
Published: (2026)
by: Singh, Jyotika, et al.
Published: (2026)
Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn Conversations
by: Juneja, Prerna, et al.
Published: (2026)
by: Juneja, Prerna, et al.
Published: (2026)
MTP: A Dataset for Multi-Modal Turning Points in Casual Conversations
by: Ho, Gia-Bao Dinh, et al.
Published: (2024)
by: Ho, Gia-Bao Dinh, et al.
Published: (2024)
Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs
by: Zhang, Zhaowei, et al.
Published: (2025)
by: Zhang, Zhaowei, et al.
Published: (2025)
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
by: Jung, Jaehun, et al.
Published: (2025)
by: Jung, Jaehun, et al.
Published: (2025)
TurnBench-MS: A Benchmark for Evaluating Multi-Turn, Multi-Step Reasoning in Large Language Models
by: Zhang, Yiran, et al.
Published: (2025)
by: Zhang, Yiran, et al.
Published: (2025)
C-MTCSD: A Chinese Multi-Turn Conversational Stance Detection Dataset
by: Niu, Fuqiang, et al.
Published: (2025)
by: Niu, Fuqiang, et al.
Published: (2025)
MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
by: Rosenthal, Sara, et al.
Published: (2026)
by: Rosenthal, Sara, et al.
Published: (2026)
Quantifying Conversational Reliability of Large Language Models under Multi-Turn Interaction
by: Myung, Jiyoon
Published: (2026)
by: Myung, Jiyoon
Published: (2026)
Temporal Graph Network: Hallucination Detection in Multi-Turn Conversation
by: Rathore, Vidhi, et al.
Published: (2026)
by: Rathore, Vidhi, et al.
Published: (2026)
ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
by: Wang, Zhebo, et al.
Published: (2026)
by: Wang, Zhebo, et al.
Published: (2026)
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
by: Verga, Pat, et al.
Published: (2024)
by: Verga, Pat, et al.
Published: (2024)
NaturalTurn: A Method to Segment Speech into Psychologically Meaningful Conversational Turns
by: Cooney, Gus, et al.
Published: (2024)
by: Cooney, Gus, et al.
Published: (2024)
Adaptive Stopping for Multi-Turn LLM Reasoning
by: Zhou, Xiaofan, et al.
Published: (2026)
by: Zhou, Xiaofan, et al.
Published: (2026)
Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries
by: Ullah, Muhammad Aziz, et al.
Published: (2026)
by: Ullah, Muhammad Aziz, et al.
Published: (2026)
Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations
by: Chandra, Mohit, et al.
Published: (2025)
by: Chandra, Mohit, et al.
Published: (2025)
Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations
by: Liu, Yu Lu, et al.
Published: (2026)
by: Liu, Yu Lu, et al.
Published: (2026)
Similar Items
-
Tailoring Self-Rationalizers with Multi-Reward Distillation
by: Ramnath, Sahana, et al.
Published: (2023) -
ELI-Why: Evaluating the Pedagogical Utility of Language Model Explanations
by: Joshi, Brihi, et al.
Published: (2025) -
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
by: Hallinan, Skyler, et al.
Published: (2025) -
CAVE: Controllable Authorship Verification Explanations
by: Ramnath, Sahana, et al.
Published: (2024) -
Improving Language Model Personas via Rationalization with Psychological Scaffolds
by: Joshi, Brihi, et al.
Published: (2025)