Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Yanjiang, Lou, Jie, Guan, Xinyan, Ji, Yuqiu, Lin, Hongyu, He, Ben, Han, Xianpei, Sun, Le, Yu, Xing, Lu, Yaojie |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation
por: Wen, Xueru, et al.
Publicado: (2024)
por: Wen, Xueru, et al.
Publicado: (2024)
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
por: Yuan, Qianhao, et al.
Publicado: (2026)
por: Yuan, Qianhao, et al.
Publicado: (2026)
You Can't Fight in Here! This is BBS!
por: Futrell, Richard, et al.
Publicado: (2026)
por: Futrell, Richard, et al.
Publicado: (2026)
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
por: Guan, Xinyan, et al.
Publicado: (2024)
por: Guan, Xinyan, et al.
Publicado: (2024)
Your Teaching Can't Help
por: Rebecca Weaver
Publicado: (2024)
por: Rebecca Weaver
Publicado: (2024)
REInstruct: Building Instruction Data from Unlabeled Corpus
por: Chen, Shu, et al.
Publicado: (2024)
por: Chen, Shu, et al.
Publicado: (2024)
SAISA: Towards Multimodal Large Language Models with Both Training and Inference Efficiency
por: Yuan, Qianhao, et al.
Publicado: (2025)
por: Yuan, Qianhao, et al.
Publicado: (2025)
Coupled Variational Reinforcement Learning for Language Model General Reasoning
por: Wen, Xueru, et al.
Publicado: (2025)
por: Wen, Xueru, et al.
Publicado: (2025)
Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch
por: Wen, Xueru, et al.
Publicado: (2025)
por: Wen, Xueru, et al.
Publicado: (2025)
The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models
por: Li, Zichao, et al.
Publicado: (2025)
por: Li, Zichao, et al.
Publicado: (2025)
Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic
por: Zheng, Xin, et al.
Publicado: (2024)
por: Zheng, Xin, et al.
Publicado: (2024)
Learning from Failures: Correction-Oriented Policy Optimization with Verifiable Rewards
por: Ren, Mengjie, et al.
Publicado: (2026)
por: Ren, Mengjie, et al.
Publicado: (2026)
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
por: Liu, Yanjiang, et al.
Publicado: (2025)
por: Liu, Yanjiang, et al.
Publicado: (2025)
You Can't Get a Library Card if You're Homeless.
Publicado: (1989)
Publicado: (1989)
You Can Help Your Country
por: Mayall, Berry, et al.
Publicado: (2021)
por: Mayall, Berry, et al.
Publicado: (2021)
PPTAgent: Generating and Evaluating Presentations Beyond Text-to-Slides
por: Zheng, Hao, et al.
Publicado: (2025)
por: Zheng, Hao, et al.
Publicado: (2025)
ConsistentChat: Building Skeleton-Guided Consistent Multi-Turn Dialogues for Large Language Models from Scratch
por: Chen, Jiawei, et al.
Publicado: (2025)
por: Chen, Jiawei, et al.
Publicado: (2025)
Rule or Story, Which is a Better Commonsense Expression for Talking with Large Language Models?
por: Bian, Ning, et al.
Publicado: (2024)
por: Bian, Ning, et al.
Publicado: (2024)
You Can't Get There From Here: Redefining Information Science to address our sociotechnical futures
por: Humr, Scott, et al.
Publicado: (2025)
por: Humr, Scott, et al.
Publicado: (2025)
DeepRAG: Thinking to Retrieve Step by Step for Large Language Models
por: Guan, Xinyan, et al.
Publicado: (2025)
por: Guan, Xinyan, et al.
Publicado: (2025)
When Models Outthink Their Safety: Unveiling and Mitigating Self-Jailbreak in Large Reasoning Models
por: Mao, Yingzhi, et al.
Publicado: (2025)
por: Mao, Yingzhi, et al.
Publicado: (2025)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
por: Liang, Qiao, et al.
Publicado: (2025)
por: Liang, Qiao, et al.
Publicado: (2025)
LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
por: Kambhampati, Subbarao, et al.
Publicado: (2024)
por: Kambhampati, Subbarao, et al.
Publicado: (2024)
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
por: Yuan, Qianhao, et al.
Publicado: (2025)
por: Yuan, Qianhao, et al.
Publicado: (2025)
You're Not from Around Here, Are You?
por: Blum, Louise A.
Publicado: (2023)
por: Blum, Louise A.
Publicado: (2023)
The Struggle You Can’t See
por: Lierman, Ash
Publicado: (2024)
por: Lierman, Ash
Publicado: (2024)
Rethinking Reward Model Evaluation: Are We Barking up the Wrong Tree?
por: Wen, Xueru, et al.
Publicado: (2024)
por: Wen, Xueru, et al.
Publicado: (2024)
"If It's Not Here, I Can't Be Bothered...": Limiting Searches to In-House Journals.
por: Davidoff, Donna J., et al.
Publicado: (1991)
por: Davidoff, Donna J., et al.
Publicado: (1991)
You Can't Get There from Here: Issues in Remote Access to Electronic Journals for a Health Sciences Library.
por: Krieb, Dennis
Publicado: (1999)
por: Krieb, Dennis
Publicado: (1999)
LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
por: Mo, Guozhao, et al.
Publicado: (2025)
por: Mo, Guozhao, et al.
Publicado: (2025)
Executing Natural Language-Described Algorithms with Large Language Models: An Investigation
por: Zheng, Xin, et al.
Publicado: (2024)
por: Zheng, Xin, et al.
Publicado: (2024)
Meta-Cognitive Analysis: Evaluating Declarative and Procedural Knowledge in Datasets and Large Language Models
por: Li, Zhuoqun, et al.
Publicado: (2024)
por: Li, Zhuoqun, et al.
Publicado: (2024)
Open Grounded Planning: Challenges and Benchmark Construction
por: Guo, Shiguang, et al.
Publicado: (2024)
por: Guo, Shiguang, et al.
Publicado: (2024)
Distance Education Students Speak to the Library: Here's How You Can Help Even More.
por: Kazmer, Michelle M.
Publicado: (2002)
por: Kazmer, Michelle M.
Publicado: (2002)
You Can't Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense
por: Mai, Wuyuao, et al.
Publicado: (2025)
por: Mai, Wuyuao, et al.
Publicado: (2025)
Beyond Isolated Dots: Benchmarking Structured Table Construction as Deep Knowledge Extraction
por: Zhong, Tianyun, et al.
Publicado: (2025)
por: Zhong, Tianyun, et al.
Publicado: (2025)
P^2O: Joint Policy and Prompt Optimization
por: Lu, Xinyu, et al.
Publicado: (2026)
por: Lu, Xinyu, et al.
Publicado: (2026)
ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
por: Men, Xin, et al.
Publicado: (2024)
por: Men, Xin, et al.
Publicado: (2024)
You Are Your Best Teacher: Semi-Supervised Surgical Point Tracking with Cycle-Consistent Self-Distillation
por: Bundele, Valay, et al.
Publicado: (2025)
por: Bundele, Valay, et al.
Publicado: (2025)
You Can't Spell Tragedy Without Rage
por: Elaine McGirr
Publicado: (2025)
por: Elaine McGirr
Publicado: (2025)
Ejemplares similares
-
On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation
por: Wen, Xueru, et al.
Publicado: (2024) -
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
por: Yuan, Qianhao, et al.
Publicado: (2026) -
You Can't Fight in Here! This is BBS!
por: Futrell, Richard, et al.
Publicado: (2026) -
Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
por: Guan, Xinyan, et al.
Publicado: (2024) -
Your Teaching Can't Help
por: Rebecca Weaver
Publicado: (2024)