Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zehao, Cao, Yuanpu, Chen, Jinghui, Honavar, Vasant G. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EsaCL: Efficient Continual Learning of Sparse Models
von: Ren, Weijieying, et al.
Veröffentlicht: (2024)
von: Ren, Weijieying, et al.
Veröffentlicht: (2024)
TruthFlow: Truthful LLM Generation via Representation Flow Correction
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
von: Wang, Hanyu, et al.
Veröffentlicht: (2025)
A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models
von: Ren, Weijieying, et al.
Veröffentlicht: (2025)
von: Ren, Weijieying, et al.
Veröffentlicht: (2025)
Hyperdimensional Representation Learning for Node Classification and Link Prediction
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2024)
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2024)
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
von: Lan, Yifan, et al.
Veröffentlicht: (2026)
von: Lan, Yifan, et al.
Veröffentlicht: (2026)
Hyperdimensional Cross-Modal Alignment of Frozen Language and Image Models for Efficient Image Captioning
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2026)
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2026)
Defending Against Alignment-Breaking Attacks via Robustly Aligned LLM
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
von: Cao, Bochuan, et al.
Veröffentlicht: (2023)
Deep Learning within Tabular Data: Foundations, Challenges, Advances and Future Directions
von: Ren, Weijieying, et al.
Veröffentlicht: (2025)
von: Ren, Weijieying, et al.
Veröffentlicht: (2025)
Brains vs. Bytes: Evaluating LLM Proficiency in Olympiad Mathematics
von: Mahdavi, Hamed, et al.
Veröffentlicht: (2025)
von: Mahdavi, Hamed, et al.
Veröffentlicht: (2025)
Finding the Sweet Spot: Trading Quality, Cost, and Speed During Inference-Time LLM Reflection
von: Butler, Jack, et al.
Veröffentlicht: (2025)
von: Butler, Jack, et al.
Veröffentlicht: (2025)
When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
von: Liu, Xiaogeng, et al.
Veröffentlicht: (2026)
OGLS-SD: On-Policy Self-Distillation with Outcome-Guided Logit Steering for LLM Reasoning
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
von: Yang, Yuxiao, et al.
Veröffentlicht: (2026)
Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
Reinforcement-aware Knowledge Distillation for LLM Reasoning
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Zhang, Zhaoyang, et al.
Veröffentlicht: (2026)
RefGrader: Automated Grading of Mathematical Competition Proofs using Agentic Workflows
von: Mahdavi, Hamed, et al.
Veröffentlicht: (2025)
von: Mahdavi, Hamed, et al.
Veröffentlicht: (2025)
When Planning Fails Despite Correct Execution: On Epistemic Calibration for LLM-Based Multi-Agent Systems
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
von: Wang, Zehao, et al.
Veröffentlicht: (2026)
Self-Evolving Curriculum for LLM Reasoning
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyin, et al.
Veröffentlicht: (2025)
Stealthy and Persistent Unalignment on Large Language Models via Backdoor Injections
von: Cao, Yuanpu, et al.
Veröffentlicht: (2023)
von: Cao, Yuanpu, et al.
Veröffentlicht: (2023)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
von: Liu, Hugh Xuechen, et al.
Veröffentlicht: (2026)
von: Liu, Hugh Xuechen, et al.
Veröffentlicht: (2026)
Scalable LLM Reasoning Acceleration with Low-rank Distillation
von: Dong, Harry, et al.
Veröffentlicht: (2025)
von: Dong, Harry, et al.
Veröffentlicht: (2025)
Architecture Is All You Need: Diversity-Enabled Sweet Spots for Robust Humanoid Locomotion
von: Werner, Blake, et al.
Veröffentlicht: (2025)
von: Werner, Blake, et al.
Veröffentlicht: (2025)
Weight-of-Thought Reasoning: Exploring Neural Network Weights for Enhanced LLM Reasoning
von: Punjwani, Saif, et al.
Veröffentlicht: (2025)
von: Punjwani, Saif, et al.
Veröffentlicht: (2025)
Not All Instances Are Equally Valuable: Towards Influence-Weighted Dataset Distillation
von: Deng, Qiyan, et al.
Veröffentlicht: (2025)
von: Deng, Qiyan, et al.
Veröffentlicht: (2025)
Causal Effect Estimation Using Random Hyperplane Tessellations
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2024)
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2024)
ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks
von: Saini, Dhruv, et al.
Veröffentlicht: (2026)
von: Saini, Dhruv, et al.
Veröffentlicht: (2026)
Learning from Partial Chain-of-Thought via Truncated-Reasoning Self-Distillation
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2026)
von: Silvestri, Gianluigi, et al.
Veröffentlicht: (2026)
Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-training
von: Barakat, Anas, et al.
Veröffentlicht: (2026)
von: Barakat, Anas, et al.
Veröffentlicht: (2026)
Skin-R1: Toward Trustworthy Clinical Reasoning for Dermatological Diagnosis
von: Liu, Zehao, et al.
Veröffentlicht: (2025)
von: Liu, Zehao, et al.
Veröffentlicht: (2025)
HEAL: Hindsight Entropy-Assisted Learning for Reasoning Distillation
von: Zhang, Wenjing, et al.
Veröffentlicht: (2026)
von: Zhang, Wenjing, et al.
Veröffentlicht: (2026)
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
PreFlect: From Retrospective to Prospective Reflection in Large Language Model Agents
von: Wang, Hanyu, et al.
Veröffentlicht: (2026)
von: Wang, Hanyu, et al.
Veröffentlicht: (2026)
Minimax Rates and Spectral Distillation for Tree Ensembles
von: Vu, Binh Duc, et al.
Veröffentlicht: (2026)
von: Vu, Binh Duc, et al.
Veröffentlicht: (2026)
Validity-Calibrated Reasoning Distillation
von: Saadi, Khouloud, et al.
Veröffentlicht: (2026)
von: Saadi, Khouloud, et al.
Veröffentlicht: (2026)
Enhancing LLM Agents for Code Generation with Possibility and Pass-rate Prioritized Experience Replay
von: Chen, Yuyang, et al.
Veröffentlicht: (2024)
von: Chen, Yuyang, et al.
Veröffentlicht: (2024)
FedSDR: Federated Self-Distillation with Rectification
von: Ren, Ziheng, et al.
Veröffentlicht: (2026)
von: Ren, Ziheng, et al.
Veröffentlicht: (2026)
Adversarially Robust Industrial Anomaly Detection Through Diffusion Model
von: Cao, Yuanpu, et al.
Veröffentlicht: (2024)
von: Cao, Yuanpu, et al.
Veröffentlicht: (2024)
ParaBlock: Communication-Computation Parallel Block Coordinate Federated Learning for Large Language Models
von: Wang, Yujia, et al.
Veröffentlicht: (2025)
von: Wang, Yujia, et al.
Veröffentlicht: (2025)
Distillation Traps and Guards: A Calibration Knob for LLM Distillability
von: Zhan, Weixiao, et al.
Veröffentlicht: (2026)
von: Zhan, Weixiao, et al.
Veröffentlicht: (2026)
C-HDNet: Hyperdimensional Computing for Causal Effect Estimation from Observational Data Under Network Interference
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2025)
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2025)
TED: Training-Free Experience Distillation for Multimodal Reasoning
von: Yuan, Shuozhi, et al.
Veröffentlicht: (2026)
von: Yuan, Shuozhi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
EsaCL: Efficient Continual Learning of Sparse Models
von: Ren, Weijieying, et al.
Veröffentlicht: (2024) -
TruthFlow: Truthful LLM Generation via Representation Flow Correction
von: Wang, Hanyu, et al.
Veröffentlicht: (2025) -
A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models
von: Ren, Weijieying, et al.
Veröffentlicht: (2025) -
Hyperdimensional Representation Learning for Node Classification and Link Prediction
von: Dalvi, Abhishek, et al.
Veröffentlicht: (2024) -
The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation
von: Lan, Yifan, et al.
Veröffentlicht: (2026)