Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
Fuente:
arXiv
Guardado en:
| Autores principales: | He, Yinghui, Kaur, Simran, Bhaskar, Adithya, Yang, Yongjin, Liu, Jiarui, Ri, Narutatsu, Fowl, Liam, Panigrahi, Abhishek, Chen, Danqi, Arora, Sanjeev |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
por: He, Yinghui, et al.
Publicado: (2025)
por: He, Yinghui, et al.
Publicado: (2025)
Skill-Targeted Adaptive Training
por: He, Yinghui, et al.
Publicado: (2025)
por: He, Yinghui, et al.
Publicado: (2025)
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
por: Razin, Noam, et al.
Publicado: (2024)
por: Razin, Noam, et al.
Publicado: (2024)
How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
por: Park, Simon, et al.
Publicado: (2025)
por: Park, Simon, et al.
Publicado: (2025)
On the Power of Context-Enhanced Learning in LLMs
por: Zhu, Xingyu, et al.
Publicado: (2025)
por: Zhu, Xingyu, et al.
Publicado: (2025)
The Heuristic Core: Understanding Subnetwork Generalization in Pretrained Language Models
por: Bhaskar, Adithya, et al.
Publicado: (2024)
por: Bhaskar, Adithya, et al.
Publicado: (2024)
Language Models that Think, Chat Better
por: Bhaskar, Adithya, et al.
Publicado: (2025)
por: Bhaskar, Adithya, et al.
Publicado: (2025)
MixSD: Mixed Contextual Self-Distillation for Knowledge Injection
por: Liu, Jiarui, et al.
Publicado: (2026)
por: Liu, Jiarui, et al.
Publicado: (2026)
The unregulated plant‐based ‘milk’ industry: A threat to nutrition, health and safety?
por: Simran Kaur Arora
Publicado: (2024)
por: Simran Kaur Arora
Publicado: (2024)
Escaping the Cognitive Well: Efficient Competition Math with Off-the-Shelf Models
por: Dang, Xingyu, et al.
Publicado: (2026)
por: Dang, Xingyu, et al.
Publicado: (2026)
Instruct-SkillMix: A Powerful Pipeline for LLM Instruction Tuning
por: Kaur, Simran, et al.
Publicado: (2024)
por: Kaur, Simran, et al.
Publicado: (2024)
Trainable Transformer in Transformer
por: Panigrahi, Abhishek, et al.
Publicado: (2023)
por: Panigrahi, Abhishek, et al.
Publicado: (2023)
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms
por: Malladi, Sadhika, et al.
Publicado: (2022)
por: Malladi, Sadhika, et al.
Publicado: (2022)
Representing Rule-based Chatbots with Transformers
por: Friedman, Dan, et al.
Publicado: (2024)
por: Friedman, Dan, et al.
Publicado: (2024)
Reranking-based Generation for Unbiased Perspective Summarization
por: Ri, Narutatsu, et al.
Publicado: (2025)
por: Ri, Narutatsu, et al.
Publicado: (2025)
Finding Transformer Circuits with Edge Pruning
por: Bhaskar, Adithya, et al.
Publicado: (2024)
por: Bhaskar, Adithya, et al.
Publicado: (2024)
Improving Language Understanding from Screenshots
por: Gao, Tianyu, et al.
Publicado: (2024)
por: Gao, Tianyu, et al.
Publicado: (2024)
Extracting Rule-based Descriptions of Attention Features in Transformers
por: Friedman, Dan, et al.
Publicado: (2025)
por: Friedman, Dan, et al.
Publicado: (2025)
Why is Your Language Model a Poor Implicit Reward Model?
por: Razin, Noam, et al.
Publicado: (2025)
por: Razin, Noam, et al.
Publicado: (2025)
Can Models Learn Skill Composition from Examples?
por: Zhao, Haoyu, et al.
Publicado: (2024)
por: Zhao, Haoyu, et al.
Publicado: (2024)
3D Gaussian Splatting with Normal Information for Mesh Extraction and Improved Rendering
por: Krishnan, Meenakshi, et al.
Publicado: (2025)
por: Krishnan, Meenakshi, et al.
Publicado: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
por: Chan, Yik Siu, et al.
Publicado: (2025)
por: Chan, Yik Siu, et al.
Publicado: (2025)
Zero-shot Active Learning Using Self Supervised Learning
por: Sinha, Abhishek, et al.
Publicado: (2024)
por: Sinha, Abhishek, et al.
Publicado: (2024)
Continual Memorization of Factoids in Language Models
por: Chen, Howard, et al.
Publicado: (2024)
por: Chen, Howard, et al.
Publicado: (2024)
Cache Me If You Can: How Many KVs Do You Need for Effective Long-Context LMs?
por: Bhaskar, Adithya, et al.
Publicado: (2025)
por: Bhaskar, Adithya, et al.
Publicado: (2025)
CEC-Zero: Zero-Supervision Character Error Correction with Self-Generated Rewards
por: Lin, Zhiming, et al.
Publicado: (2025)
por: Lin, Zhiming, et al.
Publicado: (2025)
To See the Unseen: on the Generalization Ability of Transformers in Symbolic Reasoning
por: Lazić, Nevena, et al.
Publicado: (2026)
por: Lazić, Nevena, et al.
Publicado: (2026)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
por: Park, Simon, et al.
Publicado: (2025)
por: Park, Simon, et al.
Publicado: (2025)
Late Time Acceleration with Observational Constraints in Modified Theories of Gravity
por: Arora, Simran
Publicado: (2023)
por: Arora, Simran
Publicado: (2023)
Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation
por: Wang, Longwen, et al.
Publicado: (2026)
por: Wang, Longwen, et al.
Publicado: (2026)
Embedding Generalized CP Symmetry in One Zero Texture Neutrino Mass Models
por: Priya, et al.
Publicado: (2025)
por: Priya, et al.
Publicado: (2025)
GenVC: Self-Supervised Zero-Shot Voice Conversion
por: Cai, Zexin, et al.
Publicado: (2025)
por: Cai, Zexin, et al.
Publicado: (2025)
LESS: Selecting Influential Data for Targeted Instruction Tuning
por: Xia, Mengzhou, et al.
Publicado: (2024)
por: Xia, Mengzhou, et al.
Publicado: (2024)
On the Impossibility of Retrain Equivalence in Machine Unlearning
por: Yu, Jiatong, et al.
Publicado: (2025)
por: Yu, Jiatong, et al.
Publicado: (2025)
Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
por: Didolkar, Aniket, et al.
Publicado: (2025)
por: Didolkar, Aniket, et al.
Publicado: (2025)
Occlusion-Aware Self-Supervised Monocular Depth Estimation for Weak-Texture Endoscopic Images
por: Huang, Zebo, et al.
Publicado: (2025)
por: Huang, Zebo, et al.
Publicado: (2025)
Tailoring Self-Rationalizers with Multi-Reward Distillation
por: Ramnath, Sahana, et al.
Publicado: (2023)
por: Ramnath, Sahana, et al.
Publicado: (2023)
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
por: Tomihari, Akiyoshi, et al.
Publicado: (2026)
por: Tomihari, Akiyoshi, et al.
Publicado: (2026)
p-(001)NiO/n-(0001)ZnO heterostructures grown by pulsed laser deposition technique
por: Sahu, Bhabani Prasad, et al.
Publicado: (2024)
por: Sahu, Bhabani Prasad, et al.
Publicado: (2024)
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
por: Lowe, Scott C., et al.
Publicado: (2026)
por: Lowe, Scott C., et al.
Publicado: (2026)
Ejemplares similares
-
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models
por: He, Yinghui, et al.
Publicado: (2025) -
Skill-Targeted Adaptive Training
por: He, Yinghui, et al.
Publicado: (2025) -
Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
por: Razin, Noam, et al.
Publicado: (2024) -
How Does RL Post-training Induce Skill Composition? A Case Study on Countdown
por: Park, Simon, et al.
Publicado: (2025) -
On the Power of Context-Enhanced Learning in LLMs
por: Zhu, Xingyu, et al.
Publicado: (2025)