Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Abdulhai, Marwa, Cheng, Ryan, Shrivastava, Aryansh, Jaques, Natasha, Gal, Yarin, Levine, Sergey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025)
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
von: Wan, Yanming, et al.
Veröffentlicht: (2025)
von: Wan, Yanming, et al.
Veröffentlicht: (2025)
How LLMs Distort Our Written Language
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2026)
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2026)
Language Models Change Facts Based on the Way You Talk
von: Kearney, Matthew, et al.
Veröffentlicht: (2025)
von: Kearney, Matthew, et al.
Veröffentlicht: (2025)
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)
Speak Out of Turn: Safety Vulnerability of Large Language Models in Multi-turn Dialogue
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenhong, et al.
Veröffentlicht: (2024)
Data Selection for Multi-turn Dialogue Instruction Tuning
von: Li, Bo, et al.
Veröffentlicht: (2026)
von: Li, Bo, et al.
Veröffentlicht: (2026)
Virtual Personas for Language Models via an Anthology of Backstories
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
von: Moon, Suhong, et al.
Veröffentlicht: (2024)
Planning without Search: Refining Frontier LLMs with Offline Goal-Conditioned RL
von: Hong, Joey, et al.
Veröffentlicht: (2025)
von: Hong, Joey, et al.
Veröffentlicht: (2025)
Improving Multi-turn Dialogue Consistency with Self-Recall Thinking
von: Pang, Renning, et al.
Veröffentlicht: (2026)
von: Pang, Renning, et al.
Veröffentlicht: (2026)
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
von: Kautsar, Muhammad Dehan Al, et al.
Veröffentlicht: (2025)
Maastricht University at AMIYA: Adapting LLMs for Dialectal Arabic using Fine-tuning and MBR Decoding
von: Alali, Abdulhai, et al.
Veröffentlicht: (2026)
von: Alali, Abdulhai, et al.
Veröffentlicht: (2026)
Kernel Language Entropy: Fine-grained Uncertainty Quantification for LLMs from Semantic Similarities
von: Nikitin, Alexander, et al.
Veröffentlicht: (2024)
von: Nikitin, Alexander, et al.
Veröffentlicht: (2024)
Do Multilingual LLMs Think In English?
von: Schut, Lisa, et al.
Veröffentlicht: (2025)
von: Schut, Lisa, et al.
Veröffentlicht: (2025)
In-Context Learning Learns Label Relationships but Is Not Conventional Learning
von: Kossen, Jannik, et al.
Veröffentlicht: (2023)
von: Kossen, Jannik, et al.
Veröffentlicht: (2023)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2026)
von: Olson, Matthew Lyle, et al.
Veröffentlicht: (2026)
A Survey on Recent Advances in LLM-Based Multi-turn Dialogue Systems
von: Yi, Zihao, et al.
Veröffentlicht: (2024)
von: Yi, Zihao, et al.
Veröffentlicht: (2024)
A State-Update Prompting Strategy for Efficient and Robust Multi-turn Dialogue
von: Liu, Ziyi
Veröffentlicht: (2025)
von: Liu, Ziyi
Veröffentlicht: (2025)
Existing Large Language Model Unlearning Evaluations Are Inconclusive
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
von: Feng, Zhili, et al.
Veröffentlicht: (2025)
SOMA: Efficient Multi-turn LLM Serving via Small Language Model
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
von: Cheng, Xueqi, et al.
Veröffentlicht: (2026)
Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
von: Nam, Hyunji, et al.
Veröffentlicht: (2026)
From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Gao, Jiaxuan, et al.
Veröffentlicht: (2026)
SafeMT: Multi-turn Safety for Multimodal Language Models
von: Zhu, Han, et al.
Veröffentlicht: (2025)
von: Zhu, Han, et al.
Veröffentlicht: (2025)
MemeCMD: An Automatically Generated Chinese Multi-turn Dialogue Dataset with Contextually Retrieved Memes
von: Wang, Yuheng, et al.
Veröffentlicht: (2025)
von: Wang, Yuheng, et al.
Veröffentlicht: (2025)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
von: Chopra, Harshita, et al.
Veröffentlicht: (2026)
von: Chopra, Harshita, et al.
Veröffentlicht: (2026)
To Tell The Truth: Language of Deception and Language Models
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
von: Hazra, Sanchaita, et al.
Veröffentlicht: (2023)
CPsyCoun: A Report-based Multi-turn Dialogue Reconstruction and Evaluation Framework for Chinese Psychological Counseling
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
von: Zhang, Chenhao, et al.
Veröffentlicht: (2024)
From Deception to Detection: The Dual Roles of Large Language Models in Fake News
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2024)
von: Sallami, Dorsaf, et al.
Veröffentlicht: (2024)
Infer Human's Intentions Before Following Natural Language Instructions
von: Wan, Yanming, et al.
Veröffentlicht: (2024)
von: Wan, Yanming, et al.
Veröffentlicht: (2024)
CoSafe: Evaluating Large Language Model Safety in Multi-Turn Dialogue Coreference
von: Yu, Erxin, et al.
Veröffentlicht: (2024)
von: Yu, Erxin, et al.
Veröffentlicht: (2024)
Self-Challenging Language Model Agents
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
von: Zhou, Yifei, et al.
Veröffentlicht: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
von: Wang, Xingyao, et al.
Veröffentlicht: (2023)
von: Wang, Xingyao, et al.
Veröffentlicht: (2023)
Q-SFT: Q-Learning for Language Models via Supervised Fine-Tuning
von: Hong, Joey, et al.
Veröffentlicht: (2024)
von: Hong, Joey, et al.
Veröffentlicht: (2024)
TurnWise: The Gap between Single- and Multi-turn Language Model Capabilities
von: Graf, Victoria, et al.
Veröffentlicht: (2026)
von: Graf, Victoria, et al.
Veröffentlicht: (2026)
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
von: Pombal, José, et al.
Veröffentlicht: (2025)
von: Pombal, José, et al.
Veröffentlicht: (2025)
MTMCS-Bench: Evaluating Contextual Safety of Multimodal Large Language Models in Multi-Turn Dialogues
von: Liu, Zheyuan, et al.
Veröffentlicht: (2026)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2026)
Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
von: Lu, Miao, et al.
Veröffentlicht: (2025)
von: Lu, Miao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Consistently Simulating Human Personas with Multi-Turn Reinforcement Learning
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2025) -
Enhancing Personalized Multi-Turn Dialogue with Curiosity Reward
von: Wan, Yanming, et al.
Veröffentlicht: (2025) -
How LLMs Distort Our Written Language
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2026) -
Language Models Change Facts Based on the Way You Talk
von: Kearney, Matthew, et al.
Veröffentlicht: (2025) -
ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
von: Zhou, Yifei, et al.
Veröffentlicht: (2024)