Aletheia: What Makes RLVR For Code Verifiers Tick?
Fuente:
arXiv
Saved in:
| Main Authors: | Venkatkrishna, Vatsal, Paul, Indraneil, Gurevych, Iryna |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2025)
by: Orel, Daniil, et al.
Published: (2025)
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
by: Paul, Indraneil, et al.
Published: (2025)
by: Paul, Indraneil, et al.
Published: (2025)
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
by: Paul, Indraneil, et al.
Published: (2026)
by: Paul, Indraneil, et al.
Published: (2026)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2026)
by: Orel, Daniil, et al.
Published: (2026)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
by: Paul, Indraneil, et al.
Published: (2024)
by: Paul, Indraneil, et al.
Published: (2024)
What Makes Code Generation Ethically Sourced?
by: Xu, Zhuolin, et al.
Published: (2025)
by: Xu, Zhuolin, et al.
Published: (2025)
WybeCoder: Verified Imperative Code Generation
by: Gloeckle, Fabian, et al.
Published: (2026)
by: Gloeckle, Fabian, et al.
Published: (2026)
SIEVE: Towards Verifiable Certification for Code-datasets
by: Mbodji, Fatou Ndiaye, et al.
Published: (2025)
by: Mbodji, Fatou Ndiaye, et al.
Published: (2025)
Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks
by: Ebrahimi, Amir M., et al.
Published: (2026)
by: Ebrahimi, Amir M., et al.
Published: (2026)
An Execution-Verified Multi-Language Benchmark for Code Semantic Reasoning
by: Li, Yikun, et al.
Published: (2026)
by: Li, Yikun, et al.
Published: (2026)
Verifying LLM-Generated Code in the Context of Software Verification with Ada/SPARK
by: Cramer, Marcos, et al.
Published: (2025)
by: Cramer, Marcos, et al.
Published: (2025)
CVeDRL: An Efficient Code Verifier via Difficulty-aware Reinforcement Learning
by: Shi, Ji, et al.
Published: (2026)
by: Shi, Ji, et al.
Published: (2026)
Uncovering Systematic Failures of LLMs in Verifying Code Against Natural Language Specifications
by: Jin, Haolin, et al.
Published: (2025)
by: Jin, Haolin, et al.
Published: (2025)
From Natural Language to Verified Code: Toward AI Assisted Problem-to-Code Generation with Dafny-Based Formal Verification
by: Erfan, Md, et al.
Published: (2026)
by: Erfan, Md, et al.
Published: (2026)
AutoICE: Automatically Synthesizing Verifiable C Code via LLM-driven Evolution
by: Luo, Weilin, et al.
Published: (2025)
by: Luo, Weilin, et al.
Published: (2025)
GeoContra: From Fluent GIS Code to Verifiable Spatial Analysis with Geography-Grounded Repair
by: Xiao, Yinhao, et al.
Published: (2026)
by: Xiao, Yinhao, et al.
Published: (2026)
The Code Barrier: What LLMs Actually Understand?
by: Nikiema, Serge Lionel, et al.
Published: (2025)
by: Nikiema, Serge Lionel, et al.
Published: (2025)
WhatsCode: Large-Scale GenAI Deployment for Developer Efficiency at WhatsApp
by: Mao, Ke, et al.
Published: (2025)
by: Mao, Ke, et al.
Published: (2025)
SWE-Fuse: Empowering Software Agents via Issue-free Trajectory Learning and Entropy-aware RLVR Training
by: Wen, Xin-Cheng, et al.
Published: (2026)
by: Wen, Xin-Cheng, et al.
Published: (2026)
Clover: Closed-Loop Verifiable Code Generation
by: Sun, Chuyue, et al.
Published: (2023)
by: Sun, Chuyue, et al.
Published: (2023)
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
by: Xu, Mingde, et al.
Published: (2025)
by: Xu, Mingde, et al.
Published: (2025)
A Deep Dive Into Large Language Model Code Generation Mistakes: What and Why?
by: Chen, QiHong, et al.
Published: (2024)
by: Chen, QiHong, et al.
Published: (2024)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
by: Bai, Yifan, et al.
Published: (2026)
by: Bai, Yifan, et al.
Published: (2026)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
by: Xie, Zichen, et al.
Published: (2026)
by: Xie, Zichen, et al.
Published: (2026)
AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms
by: Zhao, Haoyu, et al.
Published: (2026)
by: Zhao, Haoyu, et al.
Published: (2026)
Investigating The Smells of LLM Generated Code
by: Paul, Debalina Ghosh, et al.
Published: (2025)
by: Paul, Debalina Ghosh, et al.
Published: (2025)
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations
by: Rani, Pooja, et al.
Published: (2025)
by: Rani, Pooja, et al.
Published: (2025)
BugSpotter: Automated Generation of Code Debugging Exercises
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
by: Pădurean, Victor-Alexandru, et al.
Published: (2024)
SWE-Universe: Scale Real-World Verifiable Environments to Millions
by: Chen, Mouxiang, et al.
Published: (2026)
by: Chen, Mouxiang, et al.
Published: (2026)
OpenComputer: Verifiable Software Worlds for Computer-Use Agents
by: Wei, Jinbiao, et al.
Published: (2026)
by: Wei, Jinbiao, et al.
Published: (2026)
Benchmarks and Metrics for Evaluations of Code Generation: A Critical Review
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
Agentic Agile-V: From Vibe Coding to Verified Engineering in Software and Hardware Development
by: Koch, Christopher
Published: (2026)
by: Koch, Christopher
Published: (2026)
ScenEval: A Benchmark for Scenario-Based Evaluation of Code Generation
by: Paul, Debalina Ghosh, et al.
Published: (2024)
by: Paul, Debalina Ghosh, et al.
Published: (2024)
MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software Engineering
by: Guo, Chuanzhe, et al.
Published: (2026)
by: Guo, Chuanzhe, et al.
Published: (2026)
Hypothesize-Then-Verify: Speculative Root Cause Analysis for Microservices with Pathwise Parallelism
by: Zhang, Lingzhe, et al.
Published: (2026)
by: Zhang, Lingzhe, et al.
Published: (2026)
AgentHub: A Registry for Discoverable, Verifiable, and Reproducible AI Agents
by: Pautsch, Erik, et al.
Published: (2025)
by: Pautsch, Erik, et al.
Published: (2025)
VeriAct: Beyond Verifiability -- Agentic Synthesis of Correct and Complete Formal Specifications
by: Misu, Md Rakib Hossain, et al.
Published: (2026)
by: Misu, Md Rakib Hossain, et al.
Published: (2026)
Faver: Boosting LLM-based RTL Generation with Function Abstracted Verifiable Middleware
by: Mu, Jianan, et al.
Published: (2025)
by: Mu, Jianan, et al.
Published: (2025)
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
by: Khati, Dipin, et al.
Published: (2026)
by: Khati, Dipin, et al.
Published: (2026)
NeuroStrata: Harnessing Neurosymbolic Paradigms for Improved Design, Testability, and Verifiability of Autonomous CPS
by: Zheng, Xi, et al.
Published: (2025)
by: Zheng, Xi, et al.
Published: (2025)
Similar Items
-
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2025) -
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
by: Paul, Indraneil, et al.
Published: (2025) -
Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
by: Paul, Indraneil, et al.
Published: (2026) -
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2026) -
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
by: Paul, Indraneil, et al.
Published: (2024)