Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agnihotri, Rudransh, Pandey, Ananya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Plug-and-Play Parameter-Efficient Tuning of Embeddings for Federated Recommendation
von: Yuan, Haochen, et al.
Veröffentlicht: (2025)
von: Yuan, Haochen, et al.
Veröffentlicht: (2025)
Plug-and-Play Transformer Modules for Test-Time Adaptation
von: Chang, Xiangyu, et al.
Veröffentlicht: (2024)
von: Chang, Xiangyu, et al.
Veröffentlicht: (2024)
Enhancing Online Continual Learning with Plug-and-Play State Space Model and Class-Conditional Mixture of Discretization
von: Liu, Sihao, et al.
Veröffentlicht: (2024)
von: Liu, Sihao, et al.
Veröffentlicht: (2024)
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
von: Chai, Jiajun, et al.
Veröffentlicht: (2025)
von: Chai, Jiajun, et al.
Veröffentlicht: (2025)
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
von: Caccia, Lucas, et al.
Veröffentlicht: (2025)
von: Caccia, Lucas, et al.
Veröffentlicht: (2025)
TS-Memory: Plug-and-Play Memory for Time Series Foundation Models
von: Lyu, Sisuo, et al.
Veröffentlicht: (2026)
von: Lyu, Sisuo, et al.
Veröffentlicht: (2026)
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
von: Chen, Zhaorun, et al.
Veröffentlicht: (2025)
von: Chen, Zhaorun, et al.
Veröffentlicht: (2025)
FitLight: Federated Imitation Learning for Plug-and-Play Autonomous Traffic Signal Control
von: Ye, Yutong, et al.
Veröffentlicht: (2025)
von: Ye, Yutong, et al.
Veröffentlicht: (2025)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2025)
von: Kumarappan, Adarsh, et al.
Veröffentlicht: (2025)
Adaptive Multimodal Protein Plug-and-Play with Diffusion-Based Priors
von: Banerjee, Amartya, et al.
Veröffentlicht: (2025)
von: Banerjee, Amartya, et al.
Veröffentlicht: (2025)
ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
von: Riera, Carlos Boned, et al.
Veröffentlicht: (2025)
von: Riera, Carlos Boned, et al.
Veröffentlicht: (2025)
CrossLinear: Plug-and-Play Cross-Correlation Embedding for Time Series Forecasting with Exogenous Variables
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2025)
Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores
von: Panda, Shevya, et al.
Veröffentlicht: (2026)
von: Panda, Shevya, et al.
Veröffentlicht: (2026)
Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems
von: Karnam, Meghana, et al.
Veröffentlicht: (2026)
von: Karnam, Meghana, et al.
Veröffentlicht: (2026)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
von: Chen, Sixu, et al.
Veröffentlicht: (2026)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2025)
Position: State-of-the-Art Claims Require State-of-the-Art Evidence
von: Oh, YongKyung
Veröffentlicht: (2026)
von: Oh, YongKyung
Veröffentlicht: (2026)
Auto-Prompt Ensemble for LLM Judge
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
von: Li, Jiajie, et al.
Veröffentlicht: (2025)
Classical Planning with LLM-Generated Heuristics: Challenging the State of the Art with Python Code
von: Corrêa, Augusto B., et al.
Veröffentlicht: (2025)
von: Corrêa, Augusto B., et al.
Veröffentlicht: (2025)
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
von: Kim, Jaemin, et al.
Veröffentlicht: (2025)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
von: Wang, Yutong, et al.
Veröffentlicht: (2025)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2023)
von: Agnihotri, Akhil, et al.
Veröffentlicht: (2023)
Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions
von: Tsai, Chen Feng, et al.
Veröffentlicht: (2023)
von: Tsai, Chen Feng, et al.
Veröffentlicht: (2023)
DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery
von: Zhang, Yikun, et al.
Veröffentlicht: (2026)
von: Zhang, Yikun, et al.
Veröffentlicht: (2026)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
von: Tan, Zhewen, et al.
Veröffentlicht: (2026)
von: Tan, Zhewen, et al.
Veröffentlicht: (2026)
CARE-RFT: Confidence-Anchored Reinforcement Finetuning for Reliable Reasoning in Large Language Models
von: Li, Shuozhe, et al.
Veröffentlicht: (2026)
von: Li, Shuozhe, et al.
Veröffentlicht: (2026)
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
von: Shahout, Rana, et al.
Veröffentlicht: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
von: Tan, Sijun, et al.
Veröffentlicht: (2024)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
von: Huang, Tzu-Heng, et al.
Veröffentlicht: (2025)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
von: Wagner, Nico, et al.
Veröffentlicht: (2024)
von: Wagner, Nico, et al.
Veröffentlicht: (2024)
REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
von: Zhang, Yasi, et al.
Veröffentlicht: (2026)
von: Zhang, Yasi, et al.
Veröffentlicht: (2026)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
von: Zhang, Wenbo, et al.
Veröffentlicht: (2026)
von: Zhang, Wenbo, et al.
Veröffentlicht: (2026)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
von: Duo, Jiangshan, et al.
Veröffentlicht: (2026)
HiGS: History-Guided Sampling for Plug-and-Play Enhancement of Diffusion Models
von: Sadat, Seyedmorteza, et al.
Veröffentlicht: (2025)
von: Sadat, Seyedmorteza, et al.
Veröffentlicht: (2025)
BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance
von: Ye, Xin, et al.
Veröffentlicht: (2025)
von: Ye, Xin, et al.
Veröffentlicht: (2025)
Taming Score-Based Denoisers in ADMM: A Convergent Plug-and-Play Framework
von: Shrestha, Rajesh, et al.
Veröffentlicht: (2026)
von: Shrestha, Rajesh, et al.
Veröffentlicht: (2026)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
von: Dadkhahi, Hamid, et al.
Veröffentlicht: (2025)
von: Dadkhahi, Hamid, et al.
Veröffentlicht: (2025)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
von: Radharapu, Bhaktipriya, et al.
Veröffentlicht: (2025)
von: Radharapu, Bhaktipriya, et al.
Veröffentlicht: (2025)
C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation
von: Chen, Guoxin, et al.
Veröffentlicht: (2025)
von: Chen, Guoxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Plug-and-Play Parameter-Efficient Tuning of Embeddings for Federated Recommendation
von: Yuan, Haochen, et al.
Veröffentlicht: (2025) -
Plug-and-Play Transformer Modules for Test-Time Adaptation
von: Chang, Xiangyu, et al.
Veröffentlicht: (2024) -
Enhancing Online Continual Learning with Plug-and-Play State Space Model and Class-Conditional Mixture of Discretization
von: Liu, Sihao, et al.
Veröffentlicht: (2024) -
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
von: Chai, Jiajun, et al.
Veröffentlicht: (2025) -
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
von: Caccia, Lucas, et al.
Veröffentlicht: (2025)