ICPO: Illocution-Calibrated Policy Optimization for Multi-Turn Conversation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhebo, Mu, Xiaohu, Zhou, Zijie, Li, Mohan, Xing, Wenpeng, Kong, Dezhang, Han, Meng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
von: Li, Jin, et al.
Veröffentlicht: (2025)
von: Li, Jin, et al.
Veröffentlicht: (2025)
Fingerprint Vector: Enabling Scalable and Efficient Model Fingerprint Transfer via Vector Addition
von: Xu, Zhenhua, et al.
Veröffentlicht: (2024)
von: Xu, Zhenhua, et al.
Veröffentlicht: (2024)
Policy of Thoughts: Scaling LLM Reasoning via Test-time Policy Evolution
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026)
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026)
NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
von: Zhou, Yi, et al.
Veröffentlicht: (2025)
von: Zhou, Yi, et al.
Veröffentlicht: (2025)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
von: Liu, Geng, et al.
Veröffentlicht: (2026)
von: Liu, Geng, et al.
Veröffentlicht: (2026)
SproutBench: A Benchmark for Safe and Ethical Large Language Models for Youth
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
PARL-MT: Learning to Call Functions in Multi-Turn Conversation with Progress Awareness
von: Chai, Huacan, et al.
Veröffentlicht: (2025)
von: Chai, Huacan, et al.
Veröffentlicht: (2025)
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
Self-Anchoring Calibration Drift in Large Language Models: How Multi-Turn Conversations Reshape Model Confidence
von: Harshavardhan
Veröffentlicht: (2026)
von: Harshavardhan
Veröffentlicht: (2026)
AT$^2$PO: Agentic Turn-based Policy Optimization via Tree Search
von: Zong, Zefang, et al.
Veröffentlicht: (2026)
von: Zong, Zefang, et al.
Veröffentlicht: (2026)
ForgetMark: Stealthy Fingerprint Embedding via Targeted Unlearning in Language Models
von: Xu, Zhenhua, et al.
Veröffentlicht: (2026)
von: Xu, Zhenhua, et al.
Veröffentlicht: (2026)
MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
von: Sirdeshmukh, Ved, et al.
Veröffentlicht: (2025)
Human Latency Conversational Turns for Spoken Avatar Systems
von: Jacoby, Derek, et al.
Veröffentlicht: (2024)
von: Jacoby, Derek, et al.
Veröffentlicht: (2024)
Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey
von: Guan, Shengyue, et al.
Veröffentlicht: (2025)
von: Guan, Shengyue, et al.
Veröffentlicht: (2025)
Same Evidence, Different Answers: Canonical-Context On-Policy Distillation for Multi-Turn Language Models
von: Lin, Zizhuo, et al.
Veröffentlicht: (2026)
von: Lin, Zizhuo, et al.
Veröffentlicht: (2026)
Towards Robust and Secure Embodied AI: A Survey on Vulnerabilities and Attacks
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
von: Xing, Wenpeng, et al.
Veröffentlicht: (2025)
ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
von: Lei, Yiming, et al.
Veröffentlicht: (2025)
von: Lei, Yiming, et al.
Veröffentlicht: (2025)
Multimodal Policy Internalization for Conversational Agents
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
von: Wang, Zhenhailong, et al.
Veröffentlicht: (2025)
RC-GRPO: Reward-Conditioned Group Relative Policy Optimization for Multi-Turn Tool Calling Agents
von: Zhong, Haitian, et al.
Veröffentlicht: (2026)
von: Zhong, Haitian, et al.
Veröffentlicht: (2026)
Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning
von: Meng, Zhaohan, et al.
Veröffentlicht: (2025)
von: Meng, Zhaohan, et al.
Veröffentlicht: (2025)
Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching
von: Savage, Thomas
Veröffentlicht: (2025)
von: Savage, Thomas
Veröffentlicht: (2025)
ICPO: Intrinsic Confidence-Driven Group Relative Preference Optimization for Efficient Reinforcement Learning
von: Wang, Jinpeng, et al.
Veröffentlicht: (2025)
von: Wang, Jinpeng, et al.
Veröffentlicht: (2025)
T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning
von: Chakraborty, Amartya, et al.
Veröffentlicht: (2025)
von: Chakraborty, Amartya, et al.
Veröffentlicht: (2025)
Personalized Turn-Level User Conversation Satisfaction Benchmark
von: Wang, Zhefan, et al.
Veröffentlicht: (2026)
von: Wang, Zhefan, et al.
Veröffentlicht: (2026)
Anti-Overestimation Dialogue Policy Learning for Task-Completion Dialogue System
von: Tian, Chang, et al.
Veröffentlicht: (2022)
von: Tian, Chang, et al.
Veröffentlicht: (2022)
ReviewInstruct: A Review-Driven Multi-Turn Conversations Generation Method for Large Language Models
von: Wu, Jiangxu, et al.
Veröffentlicht: (2025)
von: Wu, Jiangxu, et al.
Veröffentlicht: (2025)
MedMT-Bench: Can LLMs Memorize and Understand Long Multi-Turn Conversations in Medical Scenarios?
von: Yang, Lin, et al.
Veröffentlicht: (2026)
von: Yang, Lin, et al.
Veröffentlicht: (2026)
Information Gain-based Policy Optimization: A Simple and Effective Approach for Multi-Turn Search Agents
von: Wang, Guoqing, et al.
Veröffentlicht: (2025)
von: Wang, Guoqing, et al.
Veröffentlicht: (2025)
SeDT: Sentence-Transformer Decision-Transformer Conditioning for Multi-Turn Conversation Reliability
von: Setti, Ramakrishna Vamsi, et al.
Veröffentlicht: (2026)
von: Setti, Ramakrishna Vamsi, et al.
Veröffentlicht: (2026)
Uncertainty as a Planning Signal: Multi-Turn Decision Making for Goal-Oriented Conversation
von: Ling, Xinyi, et al.
Veröffentlicht: (2026)
von: Ling, Xinyi, et al.
Veröffentlicht: (2026)
MTRAG: A Multi-Turn Conversational Benchmark for Evaluating Retrieval-Augmented Generation Systems
von: Katsis, Yannis, et al.
Veröffentlicht: (2025)
von: Katsis, Yannis, et al.
Veröffentlicht: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
von: Zheng, Mingqian, et al.
Veröffentlicht: (2026)
von: Zheng, Mingqian, et al.
Veröffentlicht: (2026)
Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users
von: Ozolcer, Melik, et al.
Veröffentlicht: (2025)
von: Ozolcer, Melik, et al.
Veröffentlicht: (2025)
JMedEthicBench: A Multi-Turn Conversational Benchmark for Evaluating Medical Safety in Japanese Large Language Models
von: Liu, Junyu, et al.
Veröffentlicht: (2026)
von: Liu, Junyu, et al.
Veröffentlicht: (2026)
OnePred: Next-Query Prediction via Recursive Intent Memory in Multi-Turn Conversations
von: Chen, Jiangwang, et al.
Veröffentlicht: (2026)
von: Chen, Jiangwang, et al.
Veröffentlicht: (2026)
From Intents to Conversations: Generating Intent-Driven Dialogues with Contrastive Learning for Multi-Turn Classification
von: Liu, Junhua, et al.
Veröffentlicht: (2024)
von: Liu, Junhua, et al.
Veröffentlicht: (2024)
Hindsight-Anchored Policy Optimization: Turning Failure into Feedback in Sparse Reward Settings
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
von: Wu, Yuning, et al.
Veröffentlicht: (2026)
COMPACT: Common-token Optimized Model Pruning Across Channels and Tokens
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
von: Kwek, Eugene, et al.
Veröffentlicht: (2025)
Adaptive Stopping for Multi-Turn LLM Reasoning
von: Zhou, Xiaofan, et al.
Veröffentlicht: (2026)
von: Zhou, Xiaofan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
von: Li, Jin, et al.
Veröffentlicht: (2025) -
Fingerprint Vector: Enabling Scalable and Efficient Model Fingerprint Transfer via Vector Addition
von: Xu, Zhenhua, et al.
Veröffentlicht: (2024) -
Policy of Thoughts: Scaling LLM Reasoning via Test-time Policy Evolution
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026) -
NeuRel-Attack: Neuron Relearning for Safety Disalignment in Large Language Models
von: Zhou, Yi, et al.
Veröffentlicht: (2025) -
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)