Rethinking Verification for LLM Code Generation: From Generation to Testing
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Zihan, Zhang, Taolin, Cao, Maosong, Liu, Junnan, Zhang, Wenwei, Luo, Minnan, Zhang, Songyang, Chen, Kai |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Coding Triangle: How Does Large Language Model Understand Code?
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
di: Ma, Zihan, et al.
Pubblicazione: (2025)
di: Ma, Zihan, et al.
Pubblicazione: (2025)
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
di: Liu, Junnan, et al.
Pubblicazione: (2025)
di: Liu, Junnan, et al.
Pubblicazione: (2025)
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
di: Zhang, Taolin, et al.
Pubblicazione: (2025)
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
di: Cao, Maosong, et al.
Pubblicazione: (2025)
di: Cao, Maosong, et al.
Pubblicazione: (2025)
Rectifying LLM Thought from Lens of Optimization
di: Liu, Junnan, et al.
Pubblicazione: (2025)
di: Liu, Junnan, et al.
Pubblicazione: (2025)
NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
di: Li, Mo, et al.
Pubblicazione: (2024)
di: Li, Mo, et al.
Pubblicazione: (2024)
Are Your LLMs Capable of Stable Reasoning?
di: Liu, Junnan, et al.
Pubblicazione: (2024)
di: Liu, Junnan, et al.
Pubblicazione: (2024)
CompassVerifier: A Unified and Robust Verifier for LLMs Evaluation and Outcome Reward
di: Liu, Shudong, et al.
Pubblicazione: (2025)
di: Liu, Shudong, et al.
Pubblicazione: (2025)
CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution
di: Cao, Maosong, et al.
Pubblicazione: (2024)
di: Cao, Maosong, et al.
Pubblicazione: (2024)
CIBench: Evaluating Your LLMs with a Code Interpreter Plugin
di: Zhang, Chuyu, et al.
Pubblicazione: (2024)
di: Zhang, Chuyu, et al.
Pubblicazione: (2024)
InternLM-Law: An Open Source Chinese Legal Large Language Model
di: Fei, Zhiwei, et al.
Pubblicazione: (2024)
di: Fei, Zhiwei, et al.
Pubblicazione: (2024)
Dissecting Tool-Integrated Reasoning: An Empirical Study and Analysis
di: Zhao, Yufeng, et al.
Pubblicazione: (2025)
di: Zhao, Yufeng, et al.
Pubblicazione: (2025)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
di: Wang, Zihan, et al.
Pubblicazione: (2025)
di: Wang, Zihan, et al.
Pubblicazione: (2025)
DELL: Generating Reactions and Explanations for LLM-Based Misinformation Detection
di: Wan, Herun, et al.
Pubblicazione: (2024)
di: Wan, Herun, et al.
Pubblicazione: (2024)
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
di: Lyu, Chengqi, et al.
Pubblicazione: (2025)
di: Lyu, Chengqi, et al.
Pubblicazione: (2025)
GTA: A Benchmark for General Tool Agents
di: Wang, Jize, et al.
Pubblicazione: (2024)
di: Wang, Jize, et al.
Pubblicazione: (2024)
The Imitation Game: Turing Machine Imitator is Length Generalizable Reasoner
di: Hua, Zhouqi, et al.
Pubblicazione: (2025)
di: Hua, Zhouqi, et al.
Pubblicazione: (2025)
OPV: Outcome-based Process Verifier for Efficient Long Chain-of-Thought Verification
di: Wu, Zijian, et al.
Pubblicazione: (2025)
di: Wu, Zijian, et al.
Pubblicazione: (2025)
HACo-Det: A Study Towards Fine-Grained Machine-Generated Text Detection under Human-AI Coauthoring
di: Su, Zhixiong, et al.
Pubblicazione: (2025)
di: Su, Zhixiong, et al.
Pubblicazione: (2025)
DiFaR: Enhancing Multimodal Misinformation Detection with Diverse, Factual, and Relevant Rationales
di: Wan, Herun, et al.
Pubblicazione: (2025)
di: Wan, Herun, et al.
Pubblicazione: (2025)
MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
di: Liu, Hongwei, et al.
Pubblicazione: (2024)
di: Liu, Hongwei, et al.
Pubblicazione: (2024)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
di: Chen, Zehui, et al.
Pubblicazione: (2023)
di: Chen, Zehui, et al.
Pubblicazione: (2023)
Personalized LLM Response Generation with Parameterized Memory Injection
di: Zhang, Kai, et al.
Pubblicazione: (2024)
di: Zhang, Kai, et al.
Pubblicazione: (2024)
HuixiangDou: Overcoming Group Chat Scenarios with LLM-based Technical Assistance
di: Kong, Huanjun, et al.
Pubblicazione: (2024)
di: Kong, Huanjun, et al.
Pubblicazione: (2024)
Retrieving, Rethinking and Revising: The Chain-of-Verification Can Improve Retrieval Augmented Generation
di: He, Bolei, et al.
Pubblicazione: (2024)
di: He, Bolei, et al.
Pubblicazione: (2024)
LLM$\times$MapReduce-V3: Enabling Interactive In-Depth Survey Generation through a MCP-Driven Hierarchically Modular Agent System
di: Chao, Yu, et al.
Pubblicazione: (2025)
di: Chao, Yu, et al.
Pubblicazione: (2025)
Monocle: Hybrid Local-Global In-Context Evaluation for Long-Text Generation with Uncertainty-Based Active Learning
di: Wang, Xiaorong, et al.
Pubblicazione: (2025)
di: Wang, Xiaorong, et al.
Pubblicazione: (2025)
Retrieve-Plan-Generation: An Iterative Planning and Answering Framework for Knowledge-Intensive LLM Generation
di: Lyu, Yuanjie, et al.
Pubblicazione: (2024)
di: Lyu, Yuanjie, et al.
Pubblicazione: (2024)
Dafny as Verification-Aware Intermediate Language for Code Generation
di: Li, Yue Chen, et al.
Pubblicazione: (2025)
di: Li, Yue Chen, et al.
Pubblicazione: (2025)
Next-Generation Database Interfaces: A Survey of LLM-based Text-to-SQL
di: Hong, Zijin, et al.
Pubblicazione: (2024)
di: Hong, Zijin, et al.
Pubblicazione: (2024)
RepoAgent: An LLM-Powered Open-Source Framework for Repository-level Code Documentation Generation
di: Luo, Qinyu, et al.
Pubblicazione: (2024)
di: Luo, Qinyu, et al.
Pubblicazione: (2024)
General-Reasoner: Advancing LLM Reasoning Across All Domains
di: Ma, Xueguang, et al.
Pubblicazione: (2025)
di: Ma, Xueguang, et al.
Pubblicazione: (2025)
LLM$\times$MapReduce-V2: Entropy-Driven Convolutional Test-Time Scaling for Generating Long-Form Articles from Extremely Long Resources
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
di: Wang, Haoyu, et al.
Pubblicazione: (2025)
HardTests: Synthesizing High-Quality Test Cases for LLM Coding
di: He, Zhongmou, et al.
Pubblicazione: (2025)
di: He, Zhongmou, et al.
Pubblicazione: (2025)
Code Fingerprints: Disentangled Attribution of LLM-Generated Code
di: Guo, Jiaxun, et al.
Pubblicazione: (2026)
di: Guo, Jiaxun, et al.
Pubblicazione: (2026)
LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models
di: Chen, Xi, et al.
Pubblicazione: (2024)
di: Chen, Xi, et al.
Pubblicazione: (2024)
LaTER: Efficient Test-Time Reasoning via Latent Exploration and Explicit Verification
di: Li, Xuan, et al.
Pubblicazione: (2026)
di: Li, Xuan, et al.
Pubblicazione: (2026)
Measuring the Influence of Incorrect Code on Test Generation
di: Huang, Dong, et al.
Pubblicazione: (2024)
di: Huang, Dong, et al.
Pubblicazione: (2024)
Dynamic Scaling of Unit Tests for Code Reward Modeling
di: Ma, Zeyao, et al.
Pubblicazione: (2025)
di: Ma, Zeyao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Coding Triangle: How Does Large Language Model Understand Code?
di: Zhang, Taolin, et al.
Pubblicazione: (2025) -
How Brittle is Agent Safety? Rethinking Agent Risk under Intent Concealment and Task Complexity
di: Ma, Zihan, et al.
Pubblicazione: (2025) -
Deciphering Trajectory-Aided LLM Reasoning: An Optimization Perspective
di: Liu, Junnan, et al.
Pubblicazione: (2025) -
CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards
di: Zhang, Taolin, et al.
Pubblicazione: (2025) -
Condor: Enhance LLM Alignment with Knowledge-Driven Data Synthesis and Refinement
di: Cao, Maosong, et al.
Pubblicazione: (2025)