Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Kaiyuan, Zhuang, Ziyuan, Bai, Yang, Wang, Bing, Weng, Rongxiang, Ye, Jieping |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
por: Bai, Yang, et al.
Publicado: (2026)
por: Bai, Yang, et al.
Publicado: (2026)
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
por: Liu, Kaiyuan, et al.
Publicado: (2025)
por: Liu, Kaiyuan, et al.
Publicado: (2025)
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
por: Yan, Shaotian, et al.
Publicado: (2026)
por: Yan, Shaotian, et al.
Publicado: (2026)
Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller
por: Cai, Min, et al.
Publicado: (2024)
por: Cai, Min, et al.
Publicado: (2024)
Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning
por: Zhai, Naixin, et al.
Publicado: (2026)
por: Zhai, Naixin, et al.
Publicado: (2026)
Strong Teacher Not Needed? On Distillation in LLM Pretraining
por: Lu, Taiming, et al.
Publicado: (2026)
por: Lu, Taiming, et al.
Publicado: (2026)
Long-Chain Reasoning Distillation via Adaptive Prefix Alignment
por: Liu, Zhenghao, et al.
Publicado: (2026)
por: Liu, Zhenghao, et al.
Publicado: (2026)
Natural Language Communication with a Teachable Agent
por: Love, Rachel, et al.
Publicado: (2022)
por: Love, Rachel, et al.
Publicado: (2022)
LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
por: Zhang, Xuemiao, et al.
Publicado: (2025)
por: Zhang, Xuemiao, et al.
Publicado: (2025)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
Libra: Assessing and Improving Reward Model by Learning to Think
por: Zhou, Meng, et al.
Publicado: (2025)
por: Zhou, Meng, et al.
Publicado: (2025)
Weak-to-Strong Reasoning
por: Yang, Yuqing, et al.
Publicado: (2024)
por: Yang, Yuqing, et al.
Publicado: (2024)
Teaching LLM to be Persuasive: Reward-Enhanced Policy Optimization for Alignment from Heterogeneous Rewards
por: Zeng, Xia, et al.
Publicado: (2025)
por: Zeng, Xia, et al.
Publicado: (2025)
Advancing Adversarial Suffix Transfer Learning on Aligned Large Language Models
por: Liu, Hongfu, et al.
Publicado: (2024)
por: Liu, Hongfu, et al.
Publicado: (2024)
Length Desensitization in Direct Preference Optimization
por: Liu, Wei, et al.
Publicado: (2024)
por: Liu, Wei, et al.
Publicado: (2024)
On the Step Length Confounding in LLM Reasoning Data Selection
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
Weak-to-Strong Jailbreaking on Large Language Models
por: Zhao, Xuandong, et al.
Publicado: (2024)
por: Zhao, Xuandong, et al.
Publicado: (2024)
Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model
por: Zhu, Wenhong, et al.
Publicado: (2024)
por: Zhu, Wenhong, et al.
Publicado: (2024)
PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention
por: Wang, Haonan, et al.
Publicado: (2025)
por: Wang, Haonan, et al.
Publicado: (2025)
Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
Efficient Reasoning Through Suppression of Self-Affirmation Reflections in Large Reasoning Models
por: Liu, Kaiyuan, et al.
Publicado: (2025)
por: Liu, Kaiyuan, et al.
Publicado: (2025)
Unlearning Backdoor Attacks for LLMs with Weak-to-Strong Knowledge Distillation
por: Zhao, Shuai, et al.
Publicado: (2024)
por: Zhao, Shuai, et al.
Publicado: (2024)
A Survey on LLM Mid-Training
por: Tu, Chengying, et al.
Publicado: (2025)
por: Tu, Chengying, et al.
Publicado: (2025)
PrefixQuant: Eliminating Outliers by Prefixed Tokens for Large Language Models Quantization
por: Chen, Mengzhao, et al.
Publicado: (2024)
por: Chen, Mengzhao, et al.
Publicado: (2024)
Concept Distillation from Strong to Weak Models via Hypotheses-to-Theories Prompting
por: Boateng, Emmanuel Aboah, et al.
Publicado: (2024)
por: Boateng, Emmanuel Aboah, et al.
Publicado: (2024)
Rumor Detection by Multi-task Suffix Learning based on Time-series Dual Sentiments
por: Liu, Zhiwei, et al.
Publicado: (2025)
por: Liu, Zhiwei, et al.
Publicado: (2025)
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
por: Yang, Wenkai, et al.
Publicado: (2024)
por: Yang, Wenkai, et al.
Publicado: (2024)
Improving Weak-to-Strong Generalization with Scalable Oversight and Ensemble Learning
por: Sang, Jitao, et al.
Publicado: (2024)
por: Sang, Jitao, et al.
Publicado: (2024)
MoPE: Mixture of Prefix Experts for Zero-Shot Dialogue State Tracking
por: Tang, Tianwen, et al.
Publicado: (2024)
por: Tang, Tianwen, et al.
Publicado: (2024)
AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts
por: Kumar, Vishal, et al.
Publicado: (2024)
por: Kumar, Vishal, et al.
Publicado: (2024)
Improving Weak-to-Strong Generalization with Reliability-Aware Alignment
por: Guo, Yue, et al.
Publicado: (2024)
por: Guo, Yue, et al.
Publicado: (2024)
On-Policy Context Distillation for Language Models
por: Ye, Tianzhu, et al.
Publicado: (2026)
por: Ye, Tianzhu, et al.
Publicado: (2026)
SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
por: Li, Hao, et al.
Publicado: (2026)
por: Li, Hao, et al.
Publicado: (2026)
Weak-to-Strong Generalization beyond Accuracy: a Pilot Study in Safety, Toxicity, and Legal Reasoning
por: Ye, Ruimeng, et al.
Publicado: (2024)
por: Ye, Ruimeng, et al.
Publicado: (2024)
Synthesizing Text-to-SQL Data from Weak and Strong LLMs
por: Yang, Jiaxi, et al.
Publicado: (2024)
por: Yang, Jiaxi, et al.
Publicado: (2024)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
por: Zhang, Xing, et al.
Publicado: (2026)
por: Zhang, Xing, et al.
Publicado: (2026)
ROSA-Tuning: Enhancing Long-Context Modeling via Suffix Matching
por: Zheng, Yunao, et al.
Publicado: (2026)
por: Zheng, Yunao, et al.
Publicado: (2026)
FIRE: Flexible Integration of Data Quality Ratings for Effective Pre-Training
por: Xu, Liangyu, et al.
Publicado: (2025)
por: Xu, Liangyu, et al.
Publicado: (2025)
Unlocking Implicit Experience: Synthesizing Tool-Use Trajectories from Text
por: Xu, Zhihao, et al.
Publicado: (2026)
por: Xu, Zhihao, et al.
Publicado: (2026)
MAIGO: Mitigating Lost-in-Conversation with History-Cleaned On-Policy Self-Distillation
por: Zheng, Haoyu, et al.
Publicado: (2026)
por: Zheng, Haoyu, et al.
Publicado: (2026)
Ejemplares similares
-
Multi-Objective and Mixed-Reward Reinforcement Learning via Reward-Decorrelated Policy Optimization
por: Bai, Yang, et al.
Publicado: (2026) -
Where Did This Sentence Come From? Tracing Provenance in LLM Reasoning Distillation
por: Liu, Kaiyuan, et al.
Publicado: (2025) -
Distribution-Aligned Sequence Distillation for Superior Long-CoT Reasoning
por: Yan, Shaotian, et al.
Publicado: (2026) -
Self-Control of LLM Behaviors by Compressing Suffix Gradient into Prefix Controller
por: Cai, Min, et al.
Publicado: (2024) -
Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning
por: Zhai, Naixin, et al.
Publicado: (2026)