Agentic Rubrics as Contextual Verifiers for SWE Agents
Fuente:
arXiv
Guardado en:
| Autores principales: | Raghavendra, Mohit, Gunjal, Anisha, Liu, Bing, He, Yunzhong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025)
por: Gunjal, Anisha, et al.
Publicado: (2025)
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
por: Raghavendra, Mohit, et al.
Publicado: (2026)
por: Raghavendra, Mohit, et al.
Publicado: (2026)
Reward Hacking in Rubric-Based Reinforcement Learning
por: Mahmoud, Anas, et al.
Publicado: (2026)
por: Mahmoud, Anas, et al.
Publicado: (2026)
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
por: Huang, Jiawei, et al.
Publicado: (2026)
por: Huang, Jiawei, et al.
Publicado: (2026)
Online Rubrics Elicitation from Pairwise Comparisons
por: Rezaei, MohammadHossein, et al.
Publicado: (2025)
por: Rezaei, MohammadHossein, et al.
Publicado: (2025)
Detecting and Preventing Hallucinations in Large Vision Language Models
por: Gunjal, Anisha, et al.
Publicado: (2023)
por: Gunjal, Anisha, et al.
Publicado: (2023)
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards
por: Li, Gaotang, et al.
Publicado: (2026)
por: Li, Gaotang, et al.
Publicado: (2026)
Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
por: Zhang, Junkai, et al.
Publicado: (2025)
por: Zhang, Junkai, et al.
Publicado: (2025)
TutorBench: A Benchmark To Assess Tutoring Capabilities Of Large Language Models
por: Srinivasa, Rakshith S, et al.
Publicado: (2025)
por: Srinivasa, Rakshith S, et al.
Publicado: (2025)
Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning
por: Raghavendra, Mohit, et al.
Publicado: (2025)
por: Raghavendra, Mohit, et al.
Publicado: (2025)
ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
por: Sharma, Manasi, et al.
Publicado: (2025)
por: Sharma, Manasi, et al.
Publicado: (2025)
R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
por: Jain, Naman, et al.
Publicado: (2025)
por: Jain, Naman, et al.
Publicado: (2025)
SWE-Adept: An LLM-Based Agentic Framework for Deep Codebase Analysis and Structured Issue Resolution
por: He, Kang, et al.
Publicado: (2026)
por: He, Kang, et al.
Publicado: (2026)
Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models
por: Nath, Vaskar, et al.
Publicado: (2025)
por: Nath, Vaskar, et al.
Publicado: (2025)
Alternating Reinforcement Learning for Rubric-Based Reward Modeling in Non-Verifiable LLM Post-Training
por: Xu, Ran, et al.
Publicado: (2026)
por: Xu, Ran, et al.
Publicado: (2026)
Revisiting the Superficial Alignment Hypothesis
por: Raghavendra, Mohit, et al.
Publicado: (2024)
por: Raghavendra, Mohit, et al.
Publicado: (2024)
Verifiable Agentic Infrastructure: Proof-Derived Authorization for Sovereign AI Systems
por: He, Jun, et al.
Publicado: (2026)
por: He, Jun, et al.
Publicado: (2026)
SWE-Bench-CL: Continual Learning for Coding Agents
por: Joshi, Thomas, et al.
Publicado: (2025)
por: Joshi, Thomas, et al.
Publicado: (2025)
GSO: Challenging Software Optimization Tasks for Evaluating SWE-Agents
por: Shetty, Manish, et al.
Publicado: (2025)
por: Shetty, Manish, et al.
Publicado: (2025)
Auto-Rubric: Learning From Implicit Weights to Explicit Rubrics for Reward Modeling
por: Xie, Lipeng, et al.
Publicado: (2025)
por: Xie, Lipeng, et al.
Publicado: (2025)
Contextual Rollout Bandits for Reinforcement Learning with Verifiable Rewards
por: Lu, Xiaodong, et al.
Publicado: (2026)
por: Lu, Xiaodong, et al.
Publicado: (2026)
SHARP: A Self-Evolving Human-Auditable Rubric Policy for Financial Trading Agents
por: Chen, Xiwen, et al.
Publicado: (2026)
por: Chen, Xiwen, et al.
Publicado: (2026)
SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents
por: Ding, Yifeng, et al.
Publicado: (2026)
por: Ding, Yifeng, et al.
Publicado: (2026)
SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
por: Li, Han, et al.
Publicado: (2025)
por: Li, Han, et al.
Publicado: (2025)
Hindsight Hint Distillation: Scaffolded Reasoning for SWE Agents from CoT-free Answers
por: Wang, Shengjie, et al.
Publicado: (2026)
por: Wang, Shengjie, et al.
Publicado: (2026)
Molecular Facts: Desiderata for Decontextualization in LLM Fact Verification
por: Gunjal, Anisha, et al.
Publicado: (2024)
por: Gunjal, Anisha, et al.
Publicado: (2024)
Investigating Test Overfitting on SWE-bench
por: Ahmed, Toufique, et al.
Publicado: (2025)
por: Ahmed, Toufique, et al.
Publicado: (2025)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
por: Lan, Guangchen, et al.
Publicado: (2026)
por: Lan, Guangchen, et al.
Publicado: (2026)
Step-wise Rubric Rewards for LLM Reasoning
por: Xie, Weichu, et al.
Publicado: (2026)
por: Xie, Weichu, et al.
Publicado: (2026)
MemGuard-Alpha: Detecting and Filtering Memorization-Contaminated Signals in LLM-Based Financial Forecasting via Membership Inference and Cross-Model Disagreement
por: Roy, Anisha, et al.
Publicado: (2026)
por: Roy, Anisha, et al.
Publicado: (2026)
Robust Reward Modeling via Causal Rubrics
por: Srivastava, Pragya, et al.
Publicado: (2025)
por: Srivastava, Pragya, et al.
Publicado: (2025)
Rubric-based On-policy Distillation
por: Fang, Junfeng, et al.
Publicado: (2026)
por: Fang, Junfeng, et al.
Publicado: (2026)
SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents
por: Yuan, Danlong, et al.
Publicado: (2026)
por: Yuan, Danlong, et al.
Publicado: (2026)
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
por: Wei, Yuxiang, et al.
Publicado: (2025)
por: Wei, Yuxiang, et al.
Publicado: (2025)
Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?
por: Xia, Chunqiu Steven, et al.
Publicado: (2025)
por: Xia, Chunqiu Steven, et al.
Publicado: (2025)
Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
por: Tyagi, Utkarsh, et al.
Publicado: (2026)
por: Tyagi, Utkarsh, et al.
Publicado: (2026)
RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
por: LeVine, Will, et al.
Publicado: (2026)
por: LeVine, Will, et al.
Publicado: (2026)
Reinforcement Learning with Rubric Anchors
por: Huang, Zenan, et al.
Publicado: (2025)
por: Huang, Zenan, et al.
Publicado: (2025)
AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning
por: Wu, Peilin, et al.
Publicado: (2026)
por: Wu, Peilin, et al.
Publicado: (2026)
SWE-Exp: Experience-Driven Software Issue Resolution
por: Chen, Silin, et al.
Publicado: (2025)
por: Chen, Silin, et al.
Publicado: (2025)
Ejemplares similares
-
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
por: Gunjal, Anisha, et al.
Publicado: (2025) -
SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution
por: Raghavendra, Mohit, et al.
Publicado: (2026) -
Reward Hacking in Rubric-Based Reinforcement Learning
por: Mahmoud, Anas, et al.
Publicado: (2026) -
Beyond Verifiable Rewards: Rubric-Based GRM for Reinforced Fine-Tuning SWE Agents
por: Huang, Jiawei, et al.
Publicado: (2026) -
Online Rubrics Elicitation from Pairwise Comparisons
por: Rezaei, MohammadHossein, et al.
Publicado: (2025)