Gespeichert in:
| Hauptverfasser: | Shen, Yiyang, Tu, Lifu, Wang, Weiran |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2604.02621 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Self-Distilled Agentic Reinforcement Learning
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
Quantitative LLM Judges
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025)
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
von: Zhou, Yilun, et al.
Veröffentlicht: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
von: Xu, Hongling, et al.
Veröffentlicht: (2025)
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
von: Shi, Taiwei, et al.
Veröffentlicht: (2025)
One Token to Fool LLM-as-a-Judge
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
von: Zhao, Yulai, et al.
Veröffentlicht: (2025)
Can I understand what I create? Self-Knowledge Evaluation of Large Language Models
von: Tan, Zhiquan, et al.
Veröffentlicht: (2024)
von: Tan, Zhiquan, et al.
Veröffentlicht: (2024)
Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
von: Nguyen, Hieu, et al.
Veröffentlicht: (2025)
MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
von: Qi, Jingyuan, et al.
Veröffentlicht: (2023)
von: Qi, Jingyuan, et al.
Veröffentlicht: (2023)
Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clustering
von: Zhang, Nonghai, et al.
Veröffentlicht: (2026)
von: Zhang, Nonghai, et al.
Veröffentlicht: (2026)
StagePilot: A Deep Reinforcement Learning Agent for Stage-Controlled Cybergrooming Simulation
von: An, Heajun, et al.
Veröffentlicht: (2026)
von: An, Heajun, et al.
Veröffentlicht: (2026)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
von: Sheng, Haotian, et al.
Veröffentlicht: (2026)
von: Sheng, Haotian, et al.
Veröffentlicht: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
von: Xu, Austin, et al.
Veröffentlicht: (2025)
von: Xu, Austin, et al.
Veröffentlicht: (2025)
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
von: Li, Chengye, et al.
Veröffentlicht: (2025)
von: Li, Chengye, et al.
Veröffentlicht: (2025)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
von: Yang, Xuewei, et al.
Veröffentlicht: (2026)
Self-Supervised Learning for Neural Topic Models with Variance-Invariance-Covariance Regularization
von: Xu, Weiran, et al.
Veröffentlicht: (2025)
von: Xu, Weiran, et al.
Veröffentlicht: (2025)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
von: Yang, Runming, et al.
Veröffentlicht: (2024)
von: Yang, Runming, et al.
Veröffentlicht: (2024)
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
von: Chen, Jennifer, et al.
Veröffentlicht: (2025)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
von: Pangakis, Nicholas, et al.
Veröffentlicht: (2024)
von: Pangakis, Nicholas, et al.
Veröffentlicht: (2024)
Tailoring Instructions to Student's Learning Levels Boosts Knowledge Distillation
von: Ren, Yuxin, et al.
Veröffentlicht: (2023)
von: Ren, Yuxin, et al.
Veröffentlicht: (2023)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
von: Li, Zhuochun, et al.
Veröffentlicht: (2026)
Knowledge Distillation with Training Wheels
von: Liu, Guanlin, et al.
Veröffentlicht: (2025)
von: Liu, Guanlin, et al.
Veröffentlicht: (2025)
Enhancing LLM Knowledge Learning through Generalization
von: Zhu, Mingkang, et al.
Veröffentlicht: (2025)
von: Zhu, Mingkang, et al.
Veröffentlicht: (2025)
How to Correctly Report LLM-as-a-Judge Evaluations
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
von: Liang, Xiao, et al.
Veröffentlicht: (2025)
Investigating Non-Transitivity in LLM-as-a-Judge
von: Xu, Yi, et al.
Veröffentlicht: (2025)
von: Xu, Yi, et al.
Veröffentlicht: (2025)
DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
von: Li, Pingzhi, et al.
Veröffentlicht: (2025)
Sinkhorn Distance Minimization for Knowledge Distillation
von: Cui, Xiao, et al.
Veröffentlicht: (2024)
von: Cui, Xiao, et al.
Veröffentlicht: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
von: Hong, Yihan, et al.
Veröffentlicht: (2026)
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
von: Chan, Chi-Min, et al.
Veröffentlicht: (2025)
von: Chan, Chi-Min, et al.
Veröffentlicht: (2025)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
von: Jung, Jaehun, et al.
Veröffentlicht: (2024)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
von: Zhang, Zhengze, et al.
Veröffentlicht: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
von: Liu, Yiyang, et al.
Veröffentlicht: (2025)
A Semi-supervised Generative Model for Incomplete Multi-view Data Integration with Missing Labels
von: Shen, Yiyang, et al.
Veröffentlicht: (2025)
von: Shen, Yiyang, et al.
Veröffentlicht: (2025)
Knowledge Editing on Black-box Large Language Models
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2024)
von: Song, Xiaoshuai, et al.
Veröffentlicht: (2024)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
von: Li, Tung-Ling, et al.
Veröffentlicht: (2025)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
von: Wei, Lai, et al.
Veröffentlicht: (2025)
von: Wei, Lai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
von: Whitehouse, Chenxi, et al.
Veröffentlicht: (2025) -
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024) -
Self-Distilled Agentic Reinforcement Learning
von: Lu, Zhengxi, et al.
Veröffentlicht: (2026) -
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
von: Xu, Ran, et al.
Veröffentlicht: (2025) -
Quantitative LLM Judges
von: Sahoo, Aishwarya, et al.
Veröffentlicht: (2025)