Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-Training
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Maximillian, Sun, Ruoxi, Pfister, Tomas, Arık, Sercan Ö. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
by: Chen, Maximillian, et al.
Published: (2024)
by: Chen, Maximillian, et al.
Published: (2024)
SQL-GEN: Bridging the Dialect Gap for Text-to-SQL Via Synthetic Data And Model Merging
by: Pourreza, Mohammadreza, et al.
Published: (2024)
by: Pourreza, Mohammadreza, et al.
Published: (2024)
Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization
by: Wan, Xingchen, et al.
Published: (2024)
by: Wan, Xingchen, et al.
Published: (2024)
Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models
by: Wang, Fei, et al.
Published: (2024)
by: Wang, Fei, et al.
Published: (2024)
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
by: Wang, Fei, et al.
Published: (2025)
by: Wang, Fei, et al.
Published: (2025)
LANISTR: Multimodal Learning from Structured and Unstructured Data
by: Ebrahimi, Sayna, et al.
Published: (2023)
by: Ebrahimi, Sayna, et al.
Published: (2023)
Multi-Agent Design: Optimizing Agents with Better Prompts and Topologies
by: Zhou, Han, et al.
Published: (2025)
by: Zhou, Han, et al.
Published: (2025)
Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG
by: Jin, Bowen, et al.
Published: (2024)
by: Jin, Bowen, et al.
Published: (2024)
Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding Dimensions
by: Yoon, Jinsung, et al.
Published: (2024)
by: Yoon, Jinsung, et al.
Published: (2024)
CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQL
by: Pourreza, Mohammadreza, et al.
Published: (2024)
by: Pourreza, Mohammadreza, et al.
Published: (2024)
Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments
by: Su, Hongjin, et al.
Published: (2025)
by: Su, Hongjin, et al.
Published: (2025)
SETS: Leveraging Self-Verification and Self-Correction for Improved Test-Time Scaling
by: Chen, Jiefeng, et al.
Published: (2025)
by: Chen, Jiefeng, et al.
Published: (2025)
CROME: Cross-Modal Adapters for Efficient Multimodal LLM
by: Ebrahimi, Sayna, et al.
Published: (2024)
by: Ebrahimi, Sayna, et al.
Published: (2024)
Reasoning-SQL: Reinforcement Learning with SQL Tailored Partial Rewards for Reasoning-Enhanced Text-to-SQL
by: Pourreza, Mohammadreza, et al.
Published: (2025)
by: Pourreza, Mohammadreza, et al.
Published: (2025)
Effective Large Language Model Adaptation for Improved Grounding and Citation Generation
by: Ye, Xi, et al.
Published: (2023)
by: Ye, Xi, et al.
Published: (2023)
From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation
by: Wan, Xingchen, et al.
Published: (2025)
by: Wan, Xingchen, et al.
Published: (2025)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
by: Baan, Joris, et al.
Published: (2026)
by: Baan, Joris, et al.
Published: (2026)
Stabilizing Long-term Multi-turn Reinforcement Learning with Gated Rewards
by: Sun, Zetian, et al.
Published: (2025)
by: Sun, Zetian, et al.
Published: (2025)
Mitigating Lost in Multi-turn Conversation via Curriculum RL with Verifiable Accuracy and Abstention Rewards
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Stop Listening to Me! How Multi-turn Conversations Can Degrade LLM Reliability
by: Guo, Kevin H., et al.
Published: (2026)
by: Guo, Kevin H., et al.
Published: (2026)
Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
CaLM: Contrasting Large and Small Language Models to Verify Grounded Generation
by: Hsu, I-Hung, et al.
Published: (2024)
by: Hsu, I-Hung, et al.
Published: (2024)
Securing Multi-turn Conversational Language Models From Distributed Backdoor Triggers
by: Tong, Terry, et al.
Published: (2024)
by: Tong, Terry, et al.
Published: (2024)
A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning
by: Wang, Ruiyi, et al.
Published: (2025)
by: Wang, Ruiyi, et al.
Published: (2025)
Scaling LLM Multi-turn RL with End-to-end Summarization-based Context Management
by: Lu, Miao, et al.
Published: (2025)
by: Lu, Miao, et al.
Published: (2025)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
by: Gao, Zhaolin, et al.
Published: (2024)
by: Gao, Zhaolin, et al.
Published: (2024)
TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting
by: Cao, Defu, et al.
Published: (2023)
by: Cao, Defu, et al.
Published: (2023)
SQL-PaLM: Improved Large Language Model Adaptation for Text-to-SQL (extended)
by: Sun, Ruoxi, et al.
Published: (2023)
by: Sun, Ruoxi, et al.
Published: (2023)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
by: Wang, Xingyao, et al.
Published: (2023)
by: Wang, Xingyao, et al.
Published: (2023)
Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain
by: Cheng, Gang, et al.
Published: (2025)
by: Cheng, Gang, et al.
Published: (2025)
FLAIRR-TS -- Forecasting LLM-Agents with Iterative Refinement and Retrieval for Time Series
by: Jalori, Gunjan, et al.
Published: (2025)
by: Jalori, Gunjan, et al.
Published: (2025)
HISR: Hindsight Information Modulated Segmental Process Rewards For Multi-turn Agentic Reinforcement Learning
by: Lu, Zhicong, et al.
Published: (2026)
by: Lu, Zhicong, et al.
Published: (2026)
Clarify: Improving Model Robustness With Natural Language Corrections
by: Lee, Yoonho, et al.
Published: (2024)
by: Lee, Yoonho, et al.
Published: (2024)
Resolving Intent Ambiguities by Retrieving Discriminative Clarifying Questions
by: Dhole, Kaustubh D.
Published: (2020)
by: Dhole, Kaustubh D.
Published: (2020)
Self-Trained Verification for Training- and Test-Time Self-Improvement
by: Wu, Chen Henry, et al.
Published: (2026)
by: Wu, Chen Henry, et al.
Published: (2026)
AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents
by: Ma, Chang, et al.
Published: (2024)
by: Ma, Chang, et al.
Published: (2024)
Learning to Reason as Action Abstractions with Scalable Mid-Training RL
by: Zhang, Shenao, et al.
Published: (2025)
by: Zhang, Shenao, et al.
Published: (2025)
Towards Compute-Optimal Many-Shot In-Context Learning
by: Golchin, Shahriar, et al.
Published: (2025)
by: Golchin, Shahriar, et al.
Published: (2025)
Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
by: Abdulhai, Marwa, et al.
Published: (2025)
by: Abdulhai, Marwa, et al.
Published: (2025)
Similar Items
-
Data-Centric Improvements for Enhancing Multi-Modal Understanding in Spoken Conversation Modeling
by: Chen, Maximillian, et al.
Published: (2024) -
SQL-GEN: Bridging the Dialect Gap for Text-to-SQL Via Synthetic Data And Model Merging
by: Pourreza, Mohammadreza, et al.
Published: (2024) -
Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt Optimization
by: Wan, Xingchen, et al.
Published: (2024) -
Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language Models
by: Wang, Fei, et al.
Published: (2024) -
DynScaling: Efficient Verifier-free Inference Scaling via Dynamic and Integrated Sampling
by: Wang, Fei, et al.
Published: (2025)