Salvato in:
| Autori principali: | Shen, Yiyang, Tu, Lifu, Wang, Weiran |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2604.02621 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024)
di: Tan, Sijun, et al.
Pubblicazione: (2024)
Self-Distilled Agentic Reinforcement Learning
di: Lu, Zhengxi, et al.
Pubblicazione: (2026)
di: Lu, Zhengxi, et al.
Pubblicazione: (2026)
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
di: Xu, Ran, et al.
Pubblicazione: (2025)
di: Xu, Ran, et al.
Pubblicazione: (2025)
Quantitative LLM Judges
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
di: Zhou, Yilun, et al.
Pubblicazione: (2025)
KDRL: Post-Training Reasoning LLMs via Unified Knowledge Distillation and Reinforcement Learning
di: Xu, Hongling, et al.
Pubblicazione: (2025)
di: Xu, Hongling, et al.
Pubblicazione: (2025)
Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
di: Shi, Taiwei, et al.
Pubblicazione: (2025)
di: Shi, Taiwei, et al.
Pubblicazione: (2025)
One Token to Fool LLM-as-a-Judge
di: Zhao, Yulai, et al.
Pubblicazione: (2025)
di: Zhao, Yulai, et al.
Pubblicazione: (2025)
Can I understand what I create? Self-Knowledge Evaluation of Large Language Models
di: Tan, Zhiquan, et al.
Pubblicazione: (2024)
di: Tan, Zhiquan, et al.
Pubblicazione: (2024)
Smoothing Out Hallucinations: Mitigating LLM Hallucination with Smoothed Knowledge Distillation
di: Nguyen, Hieu, et al.
Pubblicazione: (2025)
di: Nguyen, Hieu, et al.
Pubblicazione: (2025)
MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
di: Qi, Jingyuan, et al.
Pubblicazione: (2023)
di: Qi, Jingyuan, et al.
Pubblicazione: (2023)
Silence the Judge: Reinforcement Learning with Self-Verifier via Latent Geometric Clustering
di: Zhang, Nonghai, et al.
Pubblicazione: (2026)
di: Zhang, Nonghai, et al.
Pubblicazione: (2026)
StagePilot: A Deep Reinforcement Learning Agent for Stage-Controlled Cybergrooming Simulation
di: An, Heajun, et al.
Pubblicazione: (2026)
di: An, Heajun, et al.
Pubblicazione: (2026)
Deep Learning-based Method for Expressing Knowledge Boundary of Black-Box LLM
di: Sheng, Haotian, et al.
Pubblicazione: (2026)
di: Sheng, Haotian, et al.
Pubblicazione: (2026)
Does Context Matter? ContextualJudgeBench for Evaluating LLM-based Judges in Contextual Settings
di: Xu, Austin, et al.
Pubblicazione: (2025)
di: Xu, Austin, et al.
Pubblicazione: (2025)
Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning
di: Li, Chengye, et al.
Pubblicazione: (2025)
di: Li, Chengye, et al.
Pubblicazione: (2025)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
di: Yang, Xuewei, et al.
Pubblicazione: (2026)
di: Yang, Xuewei, et al.
Pubblicazione: (2026)
Self-Supervised Learning for Neural Topic Models with Variance-Invariance-Covariance Regularization
di: Xu, Weiran, et al.
Pubblicazione: (2025)
di: Xu, Weiran, et al.
Pubblicazione: (2025)
LLM-NEO: Parameter Efficient Knowledge Distillation for Large Language Models
di: Yang, Runming, et al.
Pubblicazione: (2024)
di: Yang, Runming, et al.
Pubblicazione: (2024)
DRAG: Distilling RAG for SLMs from LLMs to Transfer Knowledge and Mitigate Hallucination via Evidence and Graph-based Distillation
di: Chen, Jennifer, et al.
Pubblicazione: (2025)
di: Chen, Jennifer, et al.
Pubblicazione: (2025)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
di: Pangakis, Nicholas, et al.
Pubblicazione: (2024)
di: Pangakis, Nicholas, et al.
Pubblicazione: (2024)
Tailoring Instructions to Student's Learning Levels Boosts Knowledge Distillation
di: Ren, Yuxin, et al.
Pubblicazione: (2023)
di: Ren, Yuxin, et al.
Pubblicazione: (2023)
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
di: Li, Zhuochun, et al.
Pubblicazione: (2026)
Knowledge Distillation with Training Wheels
di: Liu, Guanlin, et al.
Pubblicazione: (2025)
di: Liu, Guanlin, et al.
Pubblicazione: (2025)
Enhancing LLM Knowledge Learning through Generalization
di: Zhu, Mingkang, et al.
Pubblicazione: (2025)
di: Zhu, Mingkang, et al.
Pubblicazione: (2025)
How to Correctly Report LLM-as-a-Judge Evaluations
di: Lee, Chungpa, et al.
Pubblicazione: (2025)
di: Lee, Chungpa, et al.
Pubblicazione: (2025)
SwS: Self-aware Weakness-driven Problem Synthesis in Reinforcement Learning for LLM Reasoning
di: Liang, Xiao, et al.
Pubblicazione: (2025)
di: Liang, Xiao, et al.
Pubblicazione: (2025)
Investigating Non-Transitivity in LLM-as-a-Judge
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
DOGe: Defensive Output Generation for LLM Protection Against Knowledge Distillation
di: Li, Pingzhi, et al.
Pubblicazione: (2025)
di: Li, Pingzhi, et al.
Pubblicazione: (2025)
Sinkhorn Distance Minimization for Knowledge Distillation
di: Cui, Xiao, et al.
Pubblicazione: (2024)
di: Cui, Xiao, et al.
Pubblicazione: (2024)
From Rubrics to Reliable Scores: Evidence-Grounded Text Evaluation with LLM Judges
di: Hong, Yihan, et al.
Pubblicazione: (2026)
di: Hong, Yihan, et al.
Pubblicazione: (2026)
J1: Exploring Simple Test-Time Scaling for LLM-as-a-Judge
di: Chan, Chi-Min, et al.
Pubblicazione: (2025)
di: Chan, Chi-Min, et al.
Pubblicazione: (2025)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
daDPO: Distribution-Aware DPO for Distilling Conversational Abilities
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
di: Zhang, Zhengze, et al.
Pubblicazione: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
di: Liu, Yiyang, et al.
Pubblicazione: (2025)
di: Liu, Yiyang, et al.
Pubblicazione: (2025)
A Semi-supervised Generative Model for Incomplete Multi-view Data Integration with Missing Labels
di: Shen, Yiyang, et al.
Pubblicazione: (2025)
di: Shen, Yiyang, et al.
Pubblicazione: (2025)
Knowledge Editing on Black-box Large Language Models
di: Song, Xiaoshuai, et al.
Pubblicazione: (2024)
di: Song, Xiaoshuai, et al.
Pubblicazione: (2024)
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
di: Li, Tung-Ling, et al.
Pubblicazione: (2025)
di: Li, Tung-Ling, et al.
Pubblicazione: (2025)
Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
di: Wei, Lai, et al.
Pubblicazione: (2025)
di: Wei, Lai, et al.
Pubblicazione: (2025)
Documenti analoghi
-
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
di: Whitehouse, Chenxi, et al.
Pubblicazione: (2025) -
JudgeBench: A Benchmark for Evaluating LLM-based Judges
di: Tan, Sijun, et al.
Pubblicazione: (2024) -
Self-Distilled Agentic Reinforcement Learning
di: Lu, Zhengxi, et al.
Pubblicazione: (2026) -
Incentivizing Agentic Reasoning in LLM Judges via Tool-Integrated Reinforcement Learning
di: Xu, Ran, et al.
Pubblicazione: (2025) -
Quantitative LLM Judges
di: Sahoo, Aishwarya, et al.
Pubblicazione: (2025)