Reinforcement Learning-based Knowledge Distillation with LLM-as-a-Judge
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shen, Yiyang, Tu, Lifu, Wang, Weiran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Quantitative LLM Judges
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025)
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
Self-Distilled Agentic Reinforcement Learning
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
One Token to Fool LLM-as-a-Judge
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
von: Sheng, Haotian, et al.
Veröffentlicht: (2026)
von: Sheng, Haotian, et al.
Veröffentlicht: (2026)
Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clustering
von: Zhang, Nonghai, et al.
Veröffentlicht: (2026)
von: Zhang, Nonghai, et al.
Veröffentlicht: (2026)
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
von: Li, Chengye, et al.
Veröffentlicht: (2025)
von: Li, Chengye, et al.
Veröffentlicht: (2025)
Can I understand what I create? Self-Knowledge Evaluation of Large Language Models
von: Tan, Zhiquan, et al.
Veröffentlicht: (2024)
von: Tan, Zhiquan, et al.
Veröffentlicht: (2024)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
von: Pangakis, Nicholas, et al.
Veröffentlicht: (2024)
von: Pangakis, Nicholas, et al.
Veröffentlicht: (2024)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Tailoring Instructions to Student's Learning Levels Boosts Knowledge Distillation
von: Ren, Yuxin, et al.
Veröffentlicht: (2023)
von: Ren, Yuxin, et al.
Veröffentlicht: (2023)
Enhancing LLM Knowledge Learning through Generalization
von: Zhu, Mingkang, et al.
Veröffentlicht: (2025)
von: Zhu, Mingkang, et al.
Veröffentlicht: (2025)
StagePilot: A Deep Reinforcement Learning Agent for Stage-Controlled Cybergrooming Simulation
von: An, Heajun, et al.
Veröffentlicht: (2026)
von: An, Heajun, et al.
Veröffentlicht: (2026)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
Knowledge Distillation with Training Wheels
von: Liu, Guanlin, et al.
Veröffentlicht: (2025)
von: Liu, Guanlin, et al.
Veröffentlicht: (2025)
Self-Supervised Learning for Neural Topic Models with Variance-Invariance-Covariance Regularization
von: Xu, Weiran, et al.
Veröffentlicht: (2025)
von: Xu, Weiran, et al.
Veröffentlicht: (2025)
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
von: Qi, Jingyuan, et al.
Veröffentlicht: (2023)
von: Qi, Jingyuan, et al.
Veröffentlicht: (2023)
How to Correctly Report LLM-as-a-Judge Evaluations
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
Sinkhorn Distance Minimization for Knowledge Distillation
von: Cui, Xiao, et al.
Veröffentlicht: (2024)
von: Cui, Xiao, et al.
Veröffentlicht: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
von: Chan, Chi-Min, et al.
Veröffentlicht: (2025)
von: Chan, Chi-Min, et al.
Veröffentlicht: (2025)
DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
Investigating Non-Transitivity in LLM-as-a-Judge
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
LoCa: Logit Calibration for Knowledge Distillation
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
Confidence Preservation Property in Knowledge Distillation Abstractions
von: Vengertsev, Dmitry, et al.
Veröffentlicht: (2024)
von: Vengertsev, Dmitry, et al.
Veröffentlicht: (2024)
Knowledge Distillation and Dataset Distillation of Large Language Models: Emerging Trends, Challenges, and Future Directions
von: Fang, Luyang, et al.
Veröffentlicht: (2025)
von: Fang, Luyang, et al.
Veröffentlicht: (2025)
Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM Detection
von: Wei, Zhipeng, et al.
Veröffentlicht: (2024)
von: Wei, Zhipeng, et al.
Veröffentlicht: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
DAPO: An Open-Source LLM Reinforcement Learning System at Scale
von: Yu, Qiying, et al.
Veröffentlicht: (2025)
von: Yu, Qiying, et al.
Veröffentlicht: (2025)
Judge Circuits
von: Feldhus, Nils, et al.
Veröffentlicht: (2026)
von: Feldhus, Nils, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025) -
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024) -
Quantitative LLM Judges
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025) -
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025) -
Self-Distilled Agentic Reinforcement Learning
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)