Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance
Fuente:
arXiv
Saved in:
| Main Authors: | Agnihotri, Rudransh, Pandey, Ananya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Plug-and-Play Parameter-Efficient Tuning of Embeddings for Federated Recommendation
by: Yuan, Haochen, et al.
Published: (2025)
by: Yuan, Haochen, et al.
Published: (2025)
Plug-and-Play Transformer Modules for Test-Time Adaptation
by: Chang, Xiangyu, et al.
Published: (2024)
by: Chang, Xiangyu, et al.
Published: (2024)
Enhancing Online Continual Learning with Plug-and-Play State Space Model and Class-Conditional Mixture of Discretization
by: Liu, Sihao, et al.
Published: (2024)
by: Liu, Sihao, et al.
Published: (2024)
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
by: Chai, Jiajun, et al.
Published: (2025)
by: Chai, Jiajun, et al.
Published: (2025)
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
by: Caccia, Lucas, et al.
Published: (2025)
by: Caccia, Lucas, et al.
Published: (2025)
TS-Memory: Plug-and-Play Memory for Time Series Foundation Models
by: Lyu, Sisuo, et al.
Published: (2026)
by: Lyu, Sisuo, et al.
Published: (2026)
ARMs: Adaptive Red-Teaming Agent against Multimodal Models with Plug-and-Play Attacks
by: Chen, Zhaorun, et al.
Published: (2025)
by: Chen, Zhaorun, et al.
Published: (2025)
FitLight: Federated Imitation Learning for Plug-and-Play Autonomous Traffic Signal Control
by: Ye, Yutong, et al.
Published: (2025)
by: Ye, Yutong, et al.
Published: (2025)
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
by: Kumarappan, Adarsh, et al.
Published: (2025)
by: Kumarappan, Adarsh, et al.
Published: (2025)
Adaptive Multimodal Protein Plug-and-Play with Diffusion-Based Priors
by: Banerjee, Amartya, et al.
Published: (2025)
by: Banerjee, Amartya, et al.
Published: (2025)
ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
by: Riera, Carlos Boned, et al.
Published: (2025)
by: Riera, Carlos Boned, et al.
Published: (2025)
CrossLinear: Plug-and-Play Cross-Correlation Embedding for Time Series Forecasting with Exogenous Variables
by: Zhou, Pengfei, et al.
Published: (2025)
by: Zhou, Pengfei, et al.
Published: (2025)
Reliability Auditing for Downstream LLM tasks in Psychiatry: LLM-Generated Hospitalization Risk Scores
by: Panda, Shevya, et al.
Published: (2026)
by: Panda, Shevya, et al.
Published: (2026)
Reliable Self-Harm Risk Screening via Adaptive Multi-Agent LLM Systems
by: Karnam, Meghana, et al.
Published: (2026)
by: Karnam, Meghana, et al.
Published: (2026)
Prompt2Fingerprint: Plug-and-Play LLM Fingerprinting via Text-to-Weight Generation
by: Chen, Sixu, et al.
Published: (2026)
by: Chen, Sixu, et al.
Published: (2026)
Who Judges the Judge? LLM Jury-on-Demand: Building Trustworthy LLM Evaluation Systems
by: Li, Xiaochuan, et al.
Published: (2025)
by: Li, Xiaochuan, et al.
Published: (2025)
Position: State-of-the-Art Claims Require State-of-the-Art Evidence
by: Oh, YongKyung
Published: (2026)
by: Oh, YongKyung
Published: (2026)
Auto-Prompt Ensemble for LLM Judge
by: Li, Jiajie, et al.
Published: (2025)
by: Li, Jiajie, et al.
Published: (2025)
Classical Planning with LLM-Generated Heuristics: Challenging the State of the Art with Python Code
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
by: Kim, Jaemin, et al.
Published: (2025)
by: Kim, Jaemin, et al.
Published: (2025)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
by: Ajwani, Rohan Deepak, et al.
Published: (2024)
by: Ajwani, Rohan Deepak, et al.
Published: (2024)
MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation
by: Wang, Yutong, et al.
Published: (2025)
by: Wang, Yutong, et al.
Published: (2025)
ACPO: A Policy Optimization Algorithm for Average MDPs with Constraints
by: Agnihotri, Akhil, et al.
Published: (2023)
by: Agnihotri, Akhil, et al.
Published: (2023)
Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions
by: Tsai, Chen Feng, et al.
Published: (2023)
by: Tsai, Chen Feng, et al.
Published: (2023)
DrugSAGE:Self-evolving Agent Experience for Efficient State-of-the-Art Drug Discovery
by: Zhang, Yikun, et al.
Published: (2026)
by: Zhang, Yikun, et al.
Published: (2026)
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
by: Tan, Zhewen, et al.
Published: (2026)
by: Tan, Zhewen, et al.
Published: (2026)
CARE-RFT: Confidence-Anchored Reinforcement Finetuning for Reliable Reasoning in Large Language Models
by: Li, Shuozhe, et al.
Published: (2026)
by: Li, Shuozhe, et al.
Published: (2026)
From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
by: Shahout, Rana, et al.
Published: (2025)
by: Shahout, Rana, et al.
Published: (2025)
JudgeBench: A Benchmark for Evaluating LLM-based Judges
by: Tan, Sijun, et al.
Published: (2024)
by: Tan, Sijun, et al.
Published: (2024)
Time To Impeach LLM-as-a-Judge: Programs are the Future of Evaluation
by: Huang, Tzu-Heng, et al.
Published: (2025)
by: Huang, Tzu-Heng, et al.
Published: (2025)
Black-box Uncertainty Quantification Method for LLM-as-a-Judge
by: Wagner, Nico, et al.
Published: (2024)
by: Wagner, Nico, et al.
Published: (2024)
REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge
by: Zhang, Yasi, et al.
Published: (2026)
by: Zhang, Yasi, et al.
Published: (2026)
Reasoning Is Not Free: Robust Adaptive Cost-Efficient Routing for LLM-as-a-Judge
by: Zhang, Wenbo, et al.
Published: (2026)
by: Zhang, Wenbo, et al.
Published: (2026)
JudgeRLVR: Judge First, Generate Second for Efficient Reasoning
by: Duo, Jiangshan, et al.
Published: (2026)
by: Duo, Jiangshan, et al.
Published: (2026)
HiGS: History-Guided Sampling for Plug-and-Play Enhancement of Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2025)
by: Sadat, Seyedmorteza, et al.
Published: (2025)
BEVDiffuser: Plug-and-Play Diffusion Model for BEV Denoising with Ground-Truth Guidance
by: Ye, Xin, et al.
Published: (2025)
by: Ye, Xin, et al.
Published: (2025)
Taming Score-Based Denoisers in ADMM: A Convergent Plug-and-Play Framework
by: Shrestha, Rajesh, et al.
Published: (2026)
by: Shrestha, Rajesh, et al.
Published: (2026)
Distribution-Calibrated Inference time compute for Thinking LLM-as-a-Judge
by: Dadkhahi, Hamid, et al.
Published: (2025)
by: Dadkhahi, Hamid, et al.
Published: (2025)
Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
by: Radharapu, Bhaktipriya, et al.
Published: (2025)
C-3PO: Compact Plug-and-Play Proxy Optimization to Achieve Human-like Retrieval-Augmented Generation
by: Chen, Guoxin, et al.
Published: (2025)
by: Chen, Guoxin, et al.
Published: (2025)
Similar Items
-
Plug-and-Play Parameter-Efficient Tuning of Embeddings for Federated Recommendation
by: Yuan, Haochen, et al.
Published: (2025) -
Plug-and-Play Transformer Modules for Test-Time Adaptation
by: Chang, Xiangyu, et al.
Published: (2024) -
Enhancing Online Continual Learning with Plug-and-Play State Space Model and Class-Conditional Mixture of Discretization
by: Liu, Sihao, et al.
Published: (2024) -
RLFactory: A Plug-and-Play Reinforcement Learning Post-Training Framework for LLM Multi-Turn Tool-Use
by: Chai, Jiajun, et al.
Published: (2025) -
Training Plug-n-Play Knowledge Modules with Deep Context Distillation
by: Caccia, Lucas, et al.
Published: (2025)