Critique-Guided Distillation for Robust Reasoning via Refinement
Fuente:
arXiv
Saved in:
| Main Authors: | Kapusuzoglu, Berkcan, Chakraborty, Supriyo, Sarwar, Zain, Lee, Chia-Hsuan, Sahu, Sambit |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation
by: Kapusuzoglu, Berkcan, et al.
Published: (2025)
by: Kapusuzoglu, Berkcan, et al.
Published: (2025)
Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
by: Panda, Ashwinee, et al.
Published: (2025)
by: Panda, Ashwinee, et al.
Published: (2025)
Your Model Diversity, Not Method, Determines Reasoning Strategy
by: Choraria, Moulik, et al.
Published: (2026)
by: Choraria, Moulik, et al.
Published: (2026)
LLM Surgery: Efficient Knowledge Unlearning and Editing in Large Language Models
by: Veldanda, Akshaj Kumar, et al.
Published: (2024)
by: Veldanda, Akshaj Kumar, et al.
Published: (2024)
Continual Pre-training of MoEs: How robust is your router?
by: Thérien, Benjamin, et al.
Published: (2025)
by: Thérien, Benjamin, et al.
Published: (2025)
Influence Functions for Efficient Data Selection in Reasoning
by: Humane, Prateek, et al.
Published: (2025)
by: Humane, Prateek, et al.
Published: (2025)
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
by: Niu, Tianyi, et al.
Published: (2026)
by: Niu, Tianyi, et al.
Published: (2026)
Distilled Self-Critique of LLMs with Synthetic Data: a Bayesian Perspective
by: Gallego, Victor
Published: (2023)
by: Gallego, Victor
Published: (2023)
Refining Dimensions for Improving Clustering-based Cross-lingual Topic Models
by: Chang, Chia-Hsuan, et al.
Published: (2024)
by: Chang, Chia-Hsuan, et al.
Published: (2024)
AVSD: Adaptive-View Self-Distillation by Balancing Consensus and Teacher-Specific Privileged Signals
by: Nguyen, Duy, et al.
Published: (2026)
by: Nguyen, Duy, et al.
Published: (2026)
Self-Refining Language Model Anonymizers via Adversarial Distillation
by: Kim, Kyuyoung, et al.
Published: (2025)
by: Kim, Kyuyoung, et al.
Published: (2025)
Decomposing the Delta: What Do Models Actually Learn from Preference Pairs?
by: Lee, Chia-Hsuan, et al.
Published: (2026)
by: Lee, Chia-Hsuan, et al.
Published: (2026)
Beyond Correctness: Learning Robust Reasoning via Transfer
by: Lee, Hyunseok, et al.
Published: (2026)
by: Lee, Hyunseok, et al.
Published: (2026)
Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
by: Ko, Jongwoo, et al.
Published: (2026)
by: Ko, Jongwoo, et al.
Published: (2026)
Probing to Refine: Reinforcement Distillation of LLMs via Explanatory Inversion
by: Tan, Zhen, et al.
Published: (2026)
by: Tan, Zhen, et al.
Published: (2026)
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
CriticBench: Benchmarking LLMs for Critique-Correct Reasoning
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
by: Xi, Zhiheng, et al.
Published: (2024)
by: Xi, Zhiheng, et al.
Published: (2024)
Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
by: Padarha, Shreyansh
Published: (2025)
by: Padarha, Shreyansh
Published: (2025)
CoT-Guard: Small Models for Strong Monitoring
by: Diwan, Nirav, et al.
Published: (2026)
by: Diwan, Nirav, et al.
Published: (2026)
Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
by: Zhao, Siyan, et al.
Published: (2026)
by: Zhao, Siyan, et al.
Published: (2026)
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
by: Kim, Jeonghye, et al.
Published: (2026)
by: Kim, Jeonghye, et al.
Published: (2026)
LLM Self-Correction with DeCRIM: Decompose, Critique, and Refine for Enhanced Following of Instructions with Multiple Constraints
by: Ferraz, Thomas Palmeira, et al.
Published: (2024)
by: Ferraz, Thomas Palmeira, et al.
Published: (2024)
MYCROFT: Towards Effective and Efficient External Data Augmentation
by: Sarwar, Zain, et al.
Published: (2024)
by: Sarwar, Zain, et al.
Published: (2024)
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
by: Havrilla, Alex, et al.
Published: (2024)
by: Havrilla, Alex, et al.
Published: (2024)
Stepwise Perplexity-Guided Refinement for Efficient Chain-of-Thought Reasoning in Large Language Models
by: Cui, Yingqian, et al.
Published: (2025)
by: Cui, Yingqian, et al.
Published: (2025)
Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
by: Akpinar, Nil-Jana, et al.
Published: (2025)
by: Akpinar, Nil-Jana, et al.
Published: (2025)
TT-XAI: Trustworthy Clinical Text Explanations via Keyword Distillation and LLM Reasoning
by: Miok, Kristian, et al.
Published: (2025)
by: Miok, Kristian, et al.
Published: (2025)
Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization
by: He, Junlin, et al.
Published: (2026)
by: He, Junlin, et al.
Published: (2026)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
by: Wang, Qibin, et al.
Published: (2025)
by: Wang, Qibin, et al.
Published: (2025)
When Actions Teach You to Think: Reasoning-Action Synergy via Reinforcement Learning in Conversational Agents
by: Rawat, Mrinal, et al.
Published: (2025)
by: Rawat, Mrinal, et al.
Published: (2025)
Structural Rationale Distillation via Reasoning Space Compression
by: Yang, Jialin, et al.
Published: (2026)
by: Yang, Jialin, et al.
Published: (2026)
Critique of Impure Reason: Unveiling the reasoning behaviour of medical Large Language Models
by: Sim, Shamus, et al.
Published: (2024)
by: Sim, Shamus, et al.
Published: (2024)
Teaching Language Models to Critique via Reinforcement Learning
by: Xie, Zhihui, et al.
Published: (2025)
by: Xie, Zhihui, et al.
Published: (2025)
CoG: Controllable Graph Reasoning via Relational Blueprints and Failure-Aware Refinement over Knowledge Graphs
by: Liu, Yuanxiang, et al.
Published: (2026)
by: Liu, Yuanxiang, et al.
Published: (2026)
ReVISE: Learning to Refine at Test-Time via Intrinsic Self-Verification
by: Lee, Hyunseok, et al.
Published: (2025)
by: Lee, Hyunseok, et al.
Published: (2025)
Find Your Optimal Teacher: Personalized Data Synthesis via Router-Guided Multi-Teacher Distillation
by: Zhang, Hengyuan, et al.
Published: (2025)
by: Zhang, Hengyuan, et al.
Published: (2025)
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA
by: Sarwar, Nobin
Published: (2025)
by: Sarwar, Nobin
Published: (2025)
Similar Items
-
SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation
by: Kapusuzoglu, Berkcan, et al.
Published: (2025) -
Optimizing Reasoning Efficiency through Prompt Difficulty Prediction
by: Zhao, Bo, et al.
Published: (2025) -
Leveraging Parameter Space Symmetries for Reasoning Skill Transfer in LLMs
by: Horoi, Stefan, et al.
Published: (2025) -
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
by: Panda, Ashwinee, et al.
Published: (2025) -
Your Model Diversity, Not Method, Determines Reasoning Strategy
by: Choraria, Moulik, et al.
Published: (2026)