Rethinking Verification for LLM Code Generation: From Generation to Testing
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Zihan, Zhang, Taolin, Cao, Maosong, Liu, Junnan, Zhang, Wenwei, Luo, Minnan, Zhang, Songyang, Chen, Kai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Coding Triangle: How Does Large Language Model Understand Code?
by: Zhang, Taolin, et al.
Published: (2025)
by: Zhang, Taolin, et al.
Published: (2025)
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
by: Ma, Zihan, et al.
Published: (2025)
by: Ma, Zihan, et al.
Published: (2025)
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
by: Zhang, Taolin, et al.
Published: (2025)
by: Zhang, Taolin, et al.
Published: (2025)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
by: Cao, Maosong, et al.
Published: (2025)
by: Cao, Maosong, et al.
Published: (2025)
Rectifying LLM Thought from Lens of Optimization
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
by: Li, Mo, et al.
Published: (2024)
by: Li, Mo, et al.
Published: (2024)
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024)
by: Liu, Junnan, et al.
Published: (2024)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
by: Liu, Shudong, et al.
Published: (2025)
by: Liu, Shudong, et al.
Published: (2025)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
by: Cao, Maosong, et al.
Published: (2024)
by: Cao, Maosong, et al.
Published: (2024)
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
by: Zhang, Chuyu, et al.
Published: (2024)
by: Zhang, Chuyu, et al.
Published: (2024)
InternLM-Law: An Open Source Chinese Legal Large Language Model
by: Fei, Zhiwei, et al.
Published: (2024)
by: Fei, Zhiwei, et al.
Published: (2024)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
by: Zhao, Yufeng, et al.
Published: (2025)
by: Zhao, Yufeng, et al.
Published: (2025)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
by: Wan, Herun, et al.
Published: (2024)
by: Wan, Herun, et al.
Published: (2024)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
by: Lyu, Chengqi, et al.
Published: (2025)
by: Lyu, Chengqi, et al.
Published: (2025)
GTA: A Benchmark for General Tool Agents
by: Wang, Jize, et al.
Published: (2024)
by: Wang, Jize, et al.
Published: (2024)
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
by: Hua, Zhouqi, et al.
Published: (2025)
by: Hua, Zhouqi, et al.
Published: (2025)
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
by: Wu, Zijian, et al.
Published: (2025)
by: Wu, Zijian, et al.
Published: (2025)
HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring
by: Su, Zhixiong, et al.
Published: (2025)
by: Su, Zhixiong, et al.
Published: (2025)
DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales
by: Wan, Herun, et al.
Published: (2025)
by: Wan, Herun, et al.
Published: (2025)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
by: Liu, Hongwei, et al.
Published: (2024)
by: Liu, Hongwei, et al.
Published: (2024)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
by: Chen, Zehui, et al.
Published: (2023)
by: Chen, Zehui, et al.
Published: (2023)
Personalized LLM Response Generation with Parameterized Memory Injection
by: Zhang, Kai, et al.
Published: (2024)
by: Zhang, Kai, et al.
Published: (2024)
HuixiangDou: Overcoming Group Chat Scenarios with LLM-based Technical Assistance
by: Kong, Huanjun, et al.
Published: (2024)
by: Kong, Huanjun, et al.
Published: (2024)
Retrieving, Rethinking and Revising: The Chain-of-Verification Can Improve Retrieval Augmented Generation
by: He, Bolei, et al.
Published: (2024)
by: He, Bolei, et al.
Published: (2024)
LLM$\times$MapReduce-V3: Enabling Interactive In-Depth Survey Generation through a MCP-Driven Hierarchically Modular Agent System
by: Chao, Yu, et al.
Published: (2025)
by: Chao, Yu, et al.
Published: (2025)
Monocle: Hybrid Local-Global In-Context Evaluation for Long-Text Generation with Uncertainty-Based Active Learning
by: Wang, Xiaorong, et al.
Published: (2025)
by: Wang, Xiaorong, et al.
Published: (2025)
Retrieve-Plan-Generation: An Iterative Planning and Answering Framework for Knowledge-Intensive LLM Generation
by: Lyu, Yuanjie, et al.
Published: (2024)
by: Lyu, Yuanjie, et al.
Published: (2024)
Dafny as Verification-Aware Intermediate Language for Code Generation
by: Li, Yue Chen, et al.
Published: (2025)
by: Li, Yue Chen, et al.
Published: (2025)
Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL
by: Hong, Zijin, et al.
Published: (2024)
by: Hong, Zijin, et al.
Published: (2024)
RepoAgent: An LLM-Powered Open-Source Framework for Repository-level Code Documentation Generation
by: Luo, Qinyu, et al.
Published: (2024)
by: Luo, Qinyu, et al.
Published: (2024)
General-Reasoner: Advancing LLM Reasoning Across All Domains
by: Ma, Xueguang, et al.
Published: (2025)
by: Ma, Xueguang, et al.
Published: (2025)
LLM$\times$MapReduce-V2: Entropy-Driven Convolutional Test-Time Scaling for Generating Long-Form Articles from Extremely Long Resources
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
HardTests: Synthesizing High-Quality Test Cases for LLM Coding
by: He, Zhongmou, et al.
Published: (2025)
by: He, Zhongmou, et al.
Published: (2025)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
by: Guo, Jiaxun, et al.
Published: (2026)
by: Guo, Jiaxun, et al.
Published: (2026)
LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models
by: Chen, Xi, et al.
Published: (2024)
by: Chen, Xi, et al.
Published: (2024)
LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification
by: Li, Xuan, et al.
Published: (2026)
by: Li, Xuan, et al.
Published: (2026)
Measuring the Influence of Incorrect Code on Test Generation
by: Huang, Dong, et al.
Published: (2024)
by: Huang, Dong, et al.
Published: (2024)
Dynamic Scaling of Unit Tests for Code Reward Modeling
by: Ma, Zeyao, et al.
Published: (2025)
by: Ma, Zeyao, et al.
Published: (2025)
Similar Items
-
Coding Triangle: How Does Large Language Model Understand Code?
by: Zhang, Taolin, et al.
Published: (2025) -
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
by: Ma, Zihan, et al.
Published: (2025) -
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
by: Liu, Junnan, et al.
Published: (2025) -
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
by: Zhang, Taolin, et al.
Published: (2025) -
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
by: Cao, Maosong, et al.
Published: (2025)