Instruct Once, Chat Consistently in Multiple Rounds: An Efficient Tuning Framework for Dialogue
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Jian, Leong, Chak Tou, Wang, Jiashuo, Lin, Dongding, Li, Wenjie, Wei, Xiao-Yong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Target-constrained Bidirectional Planning for Generation of Target-oriented Proactive Dialogue
by: Wang, Jian, et al.
Published: (2024)
by: Wang, Jian, et al.
Published: (2024)
E2CL: Exploration-based Error Correction Learning for Embodied Agents
by: Wang, Hanlin, et al.
Published: (2024)
by: Wang, Hanlin, et al.
Published: (2024)
Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents
by: Wang, Hanlin, et al.
Published: (2026)
by: Wang, Hanlin, et al.
Published: (2026)
STeCa: Step-level Trajectory Calibration for LLM Agent Learning
by: Wang, Hanlin, et al.
Published: (2025)
by: Wang, Hanlin, et al.
Published: (2025)
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
by: Leong, Chak Tou, et al.
Published: (2025)
by: Leong, Chak Tou, et al.
Published: (2025)
TokenSkip: Controllable Chain-of-Thought Compression in LLMs
by: Xia, Heming, et al.
Published: (2025)
by: Xia, Heming, et al.
Published: (2025)
Scaling over Scaling: Exploring Test-Time Scaling Plateau in Large Reasoning Models
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Finding RELIEF: Shaping Reasoning Behavior without Reasoning Supervision via Belief Engineering
by: Leong, Chak Tou, et al.
Published: (2026)
by: Leong, Chak Tou, et al.
Published: (2026)
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution
by: Wang, Hanlin, et al.
Published: (2025)
by: Wang, Hanlin, et al.
Published: (2025)
Induction Head Toxicity Mechanistically Explains Repetition Curse in Large Language Models
by: Wang, Shuxun, et al.
Published: (2025)
by: Wang, Shuxun, et al.
Published: (2025)
Where and What: Reasoning Dynamic and Implicit Preferences in Situated Conversational Recommendation
by: Lin, Dongding, et al.
Published: (2026)
by: Lin, Dongding, et al.
Published: (2026)
Mitigating Unhelpfulness in Emotional Support Conversations with Multifaceted AI Feedback
by: Wang, Jiashuo, et al.
Published: (2024)
by: Wang, Jiashuo, et al.
Published: (2024)
ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents
by: Li, Zhigen, et al.
Published: (2024)
by: Li, Zhigen, et al.
Published: (2024)
Constrain Alignment with Sparse Autoencoders
by: Yin, Qingyu, et al.
Published: (2024)
by: Yin, Qingyu, et al.
Published: (2024)
AutoPal: Autonomous Adaptation to Users for Personal AI Companionship
by: Cheng, Yi, et al.
Published: (2024)
by: Cheng, Yi, et al.
Published: (2024)
MRJ-Agent: An Effective Jailbreak Agent for Multi-Round Dialogue
by: Wang, Fengxiang, et al.
Published: (2024)
by: Wang, Fengxiang, et al.
Published: (2024)
Subtle Errors in Reasoning: Preference Learning via Error-injected Self-editing
by: Xu, Kaishuai, et al.
Published: (2024)
by: Xu, Kaishuai, et al.
Published: (2024)
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
YAYI-UIE: A Chat-Enhanced Instruction Tuning Framework for Universal Information Extraction
by: Xiao, Xinglin, et al.
Published: (2023)
by: Xiao, Xinglin, et al.
Published: (2023)
Probing the Difficulty Perception Mechanism of Large Language Models
by: Lee, Sunbowen, et al.
Published: (2025)
by: Lee, Sunbowen, et al.
Published: (2025)
Medical Dialogue Generation via Intuitive-then-Analytical Differential Diagnosis
by: Xu, Kaishuai, et al.
Published: (2024)
by: Xu, Kaishuai, et al.
Published: (2024)
Dream to Chat: Model-based Reinforcement Learning on Dialogues with User Belief Modeling
by: Zhao, Yue, et al.
Published: (2025)
by: Zhao, Yue, et al.
Published: (2025)
ClimateChat: Designing Data and Methods for Instruction Tuning LLMs to Answer Climate Change Queries
by: Chen, Zhou, et al.
Published: (2025)
by: Chen, Zhou, et al.
Published: (2025)
PlatoLM: Teaching LLMs in Multi-Round Dialogue via a User Simulator
by: Kong, Chuyi, et al.
Published: (2023)
by: Kong, Chuyi, et al.
Published: (2023)
SCALE: Selective Resource Allocation for Overcoming Performance Bottlenecks in Mathematical Test-time Scaling
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
To Retrieve or To Think? An Agentic Approach for Context Evolution
by: Chen, Rubing, et al.
Published: (2026)
by: Chen, Rubing, et al.
Published: (2026)
Dialogue is Better Than Monologue: Instructing Medical LLMs via Strategical Conversations
by: Liu, Zijie, et al.
Published: (2025)
by: Liu, Zijie, et al.
Published: (2025)
Multi-Dimensional Prompt Chaining to Improve Open-Domain Dialogue Generation
by: Teng, Livia Leong Hui
Published: (2026)
by: Teng, Livia Leong Hui
Published: (2026)
InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
by: Wang, Boxin, et al.
Published: (2023)
by: Wang, Boxin, et al.
Published: (2023)
Training Turn-by-Turn Verifiers for Dialogue Tutoring Agents: The Curious Case of LLMs as Your Coding Tutors
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
ChatASU: Evoking LLM's Reflexion to Truly Understand Aspect Sentiment in Dialogues
by: Liu, Yiding, et al.
Published: (2024)
by: Liu, Yiding, et al.
Published: (2024)
Data Selection for Multi-turn Dialogue Instruction Tuning
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
by: Philippy, Fred, et al.
Published: (2025)
by: Philippy, Fred, et al.
Published: (2025)
AlpsBench: An LLM Personalization Benchmark for Real-Dialogue Memorization and Preference Alignment
by: Xiao, Jianfei, et al.
Published: (2026)
by: Xiao, Jianfei, et al.
Published: (2026)
Generating Realistic, Protocol-Compliant Maritime Radio Dialogues using Self-Instruct and Low-Rank Adaptation
by: Akdeniz, Gürsel, et al.
Published: (2026)
by: Akdeniz, Gürsel, et al.
Published: (2026)
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
InverseCoder: Self-improving Instruction-Tuned Code LLMs with Inverse-Instruct
by: Wu, Yutong, et al.
Published: (2024)
by: Wu, Yutong, et al.
Published: (2024)
Decomposing the Basic Abilities of Large Language Models: Mitigating Cross-Task Interference in Multi-Task Instruct-Tuning
by: Wang, Bing, et al.
Published: (2026)
by: Wang, Bing, et al.
Published: (2026)
Shadow-FT: Tuning Instruct Model via Training on Paired Base Model
by: Wu, Taiqiang, et al.
Published: (2025)
by: Wu, Taiqiang, et al.
Published: (2025)
No Two Devils Alike: Unveiling Distinct Mechanisms of Fine-tuning Attacks
by: Leong, Chak Tou, et al.
Published: (2024)
by: Leong, Chak Tou, et al.
Published: (2024)
Similar Items
-
Target-constrained Bidirectional Planning for Generation of Target-oriented Proactive Dialogue
by: Wang, Jian, et al.
Published: (2024) -
E2CL: Exploration-based Error Correction Learning for Embodied Agents
by: Wang, Hanlin, et al.
Published: (2024) -
Seeing Isn't Believing: Mitigating Belief Inertia via Active Intervention in Embodied Agents
by: Wang, Hanlin, et al.
Published: (2026) -
STeCa: Step-level Trajectory Calibration for LLM Agent Learning
by: Wang, Hanlin, et al.
Published: (2025) -
Why Safeguarded Ships Run Aground? Aligned Large Language Models' Safety Mechanisms Tend to Be Anchored in The Template Region
by: Leong, Chak Tou, et al.
Published: (2025)