Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lachenmaier, Clara, Bultmann, Hannah, Zarrieß, Sina |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
von: Lachenmaier, Clara, et al.
Veröffentlicht: (2025)
von: Lachenmaier, Clara, et al.
Veröffentlicht: (2025)
LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
von: Sieker, Judith, et al.
Veröffentlicht: (2025)
von: Sieker, Judith, et al.
Veröffentlicht: (2025)
GerPS-Compare: Comparing NER methods for legal norm analysis
von: Bachinger, Sarah T., et al.
Veröffentlicht: (2024)
von: Bachinger, Sarah T., et al.
Veröffentlicht: (2024)
SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
von: Brinner, Marc, et al.
Veröffentlicht: (2025)
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
von: Dongre, Vardhan, et al.
Veröffentlicht: (2026)
von: Dongre, Vardhan, et al.
Veröffentlicht: (2026)
Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets
von: Momen, Omar, et al.
Veröffentlicht: (2026)
von: Momen, Omar, et al.
Veröffentlicht: (2026)
Evaluating Students' Open-ended Written Responses with LLMs: Using the RAG Framework for GPT-3.5, GPT-4, Claude-3, and Mistral-Large
von: Jauhiainen, Jussi S., et al.
Veröffentlicht: (2024)
von: Jauhiainen, Jussi S., et al.
Veröffentlicht: (2024)
You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
von: Shu, Bangzhao, et al.
Veröffentlicht: (2023)
von: Shu, Bangzhao, et al.
Veröffentlicht: (2023)
Beyond Continuity: Challenges of Context Switching in Multi-Turn Dialogue with LLMs
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
von: Liu, Geng, et al.
Veröffentlicht: (2026)
von: Liu, Geng, et al.
Veröffentlicht: (2026)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
von: Jaipersaud, Brandon, et al.
Veröffentlicht: (2025)
KnowGPT: Knowledge Graph based Prompting for Large Language Models
von: Zhang, Qinggang, et al.
Veröffentlicht: (2023)
von: Zhang, Qinggang, et al.
Veröffentlicht: (2023)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
Not All Thoughts are Generated Equal: Efficient LLM Reasoning via Multi-Turn Reinforcement Learning
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
von: Ning, Yansong, et al.
Veröffentlicht: (2025)
Knowing What LLMs DO NOT Know: A Simple Yet Effective Self-Detection Method
von: Zhao, Yukun, et al.
Veröffentlicht: (2023)
von: Zhao, Yukun, et al.
Veröffentlicht: (2023)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
von: Lee, Joosung, et al.
Veröffentlicht: (2026)
von: Lee, Joosung, et al.
Veröffentlicht: (2026)
Thinker: Training LLMs in Hierarchical Thinking for Deep Search via Multi-Turn Interaction
von: Xu, Jun, et al.
Veröffentlicht: (2025)
von: Xu, Jun, et al.
Veröffentlicht: (2025)
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
von: Varshney, Prasoon, et al.
Veröffentlicht: (2025)
von: Varshney, Prasoon, et al.
Veröffentlicht: (2025)
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
The Chameleon Nature of LLMs: Quantifying Multi-Turn Stance Instability in Search-Enabled Language Models
von: Ratnakar, Shivam, et al.
Veröffentlicht: (2025)
von: Ratnakar, Shivam, et al.
Veröffentlicht: (2025)
How Emotion Shapes the Behavior of LLMs and Agents: A Mechanistic Study
von: Sun, Moran, et al.
Veröffentlicht: (2026)
von: Sun, Moran, et al.
Veröffentlicht: (2026)
Models That Know How Evaluations Are Designed Score Safer
von: Deckenbach, Katharina, et al.
Veröffentlicht: (2026)
von: Deckenbach, Katharina, et al.
Veröffentlicht: (2026)
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
von: Wang, Jian, et al.
Veröffentlicht: (2025)
von: Wang, Jian, et al.
Veröffentlicht: (2025)
Mining for Species, Locations, Habitats, and Ecosystems from Scientific Papers in Invasion Biology: A Large-Scale Exploratory Study with Large Language Models
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
von: D'Souza, Jennifer, et al.
Veröffentlicht: (2025)
Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
von: Machcha, Sravanthi, et al.
Veröffentlicht: (2026)
von: Machcha, Sravanthi, et al.
Veröffentlicht: (2026)
MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
von: Yang, Lin, et al.
Veröffentlicht: (2026)
von: Yang, Lin, et al.
Veröffentlicht: (2026)
What Do LLMs Know About Alzheimer's Disease? Multi-loss Fine-Tuning and Probing for AD Detection
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
von: Jiang, Lei, et al.
Veröffentlicht: (2026)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
von: Cohen-Inger, Nurit, et al.
Veröffentlicht: (2025)
von: Cohen-Inger, Nurit, et al.
Veröffentlicht: (2025)
NanoKnow: How to Know What Your Language Model Knows
von: Gu, Lingwei, et al.
Veröffentlicht: (2026)
von: Gu, Lingwei, et al.
Veröffentlicht: (2026)
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
von: Phute, Mansi, et al.
Veröffentlicht: (2023)
DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
von: Mousavi, Seyed Mahed, et al.
Veröffentlicht: (2024)
Beyond Single-Turn: A Survey on Multi-Turn Interactions with Large Language Models
von: Li, Yubo, et al.
Veröffentlicht: (2025)
von: Li, Yubo, et al.
Veröffentlicht: (2025)
KnowRL: Teaching Language Models to Know What They Know
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs
von: Fan, Siqi, et al.
Veröffentlicht: (2025)
von: Fan, Siqi, et al.
Veröffentlicht: (2025)
Self-Anchoring Calibration Drift in Large Language Models: How Multi-Turn Conversations Reshape Model Confidence
von: Harshavardhan
Veröffentlicht: (2026)
von: Harshavardhan
Veröffentlicht: (2026)
How Hypocritical Is Your LLM judge? Listener-Speaker Asymmetries in the Pragmatic Competence of Large Language Models
von: Sieker, Judith, et al.
Veröffentlicht: (2026)
von: Sieker, Judith, et al.
Veröffentlicht: (2026)
Breaking Contextual Inertia: Reinforcement Learning with Single-Turn Anchors for Stable Multi-Turn Interaction
von: Chen, Xingwu, et al.
Veröffentlicht: (2026)
von: Chen, Xingwu, et al.
Veröffentlicht: (2026)
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
von: Şenol, Ali, et al.
Veröffentlicht: (2026)
von: Şenol, Ali, et al.
Veröffentlicht: (2026)
Discourse Diversity in Multi-Turn Empathic Dialogue
von: Zhan, Hongli, et al.
Veröffentlicht: (2026)
von: Zhan, Hongli, et al.
Veröffentlicht: (2026)
Adaptive Stopping for Multi-Turn LLM Reasoning
von: Zhou, Xiaofan, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaofan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can LLMs Ground when they (Don't) Know: A Study on Direct and Loaded Political Questions
von: Lachenmaier, Clara, et al.
Veröffentlicht: (2025) -
LLMs Struggle to Reject False Presuppositions when Misinformation Stakes are High
von: Sieker, Judith, et al.
Veröffentlicht: (2025) -
GerPS-Compare: Comparing NER methods for legal norm analysis
von: Bachinger, Sarah T., et al.
Veröffentlicht: (2024) -
SemCSE-Multi: Multifaceted and Decodable Embeddings for Aspect-Specific and Interpretable Scientific Domain Mapping
von: Brinner, Marc, et al.
Veröffentlicht: (2025) -
When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction
von: Dongre, Vardhan, et al.
Veröffentlicht: (2026)