Semantic Voting: Execution-Grounded Consensus for LLM Code Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Shan, Yi, Zijian, Zhu, Chenguang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sketch-and-Verify: Structured Inference-Time Scaling via Program Sketching
by: Jiang, Shan, et al.
Published: (2026)
by: Jiang, Shan, et al.
Published: (2026)
How Robustly do LLMs Understand Execution Semantics?
by: Spiess, Claudio, et al.
Published: (2026)
by: Spiess, Claudio, et al.
Published: (2026)
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
by: Gu, Alex, et al.
Published: (2024)
by: Gu, Alex, et al.
Published: (2024)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
by: Yu, Zhuohao, et al.
Published: (2024)
by: Yu, Zhuohao, et al.
Published: (2024)
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
by: Liu, Hugh Xuechen, et al.
Published: (2026)
by: Liu, Hugh Xuechen, et al.
Published: (2026)
Generating executable oracles to check conformance of client code to requirements of JDK Javadocs using LLMs
by: Jiang, Shan, et al.
Published: (2024)
by: Jiang, Shan, et al.
Published: (2024)
Enhancing LLM-Based Test Generation by Eliminating Covered Code
by: Xu, WeiZhe, et al.
Published: (2026)
by: Xu, WeiZhe, et al.
Published: (2026)
Feedback Over Form: Why Execution Feedback Matters More Than Pipeline Topology in 1-3B Code Generation
by: McAndrews, Charles Junichi
Published: (2026)
by: McAndrews, Charles Junichi
Published: (2026)
Operational Robustness of LLMs on Code Generation
by: Paul, Debalina Ghosh, et al.
Published: (2026)
by: Paul, Debalina Ghosh, et al.
Published: (2026)
How Efficient is LLM-Generated Code? A Rigorous & High-Standard Benchmark
by: Qiu, Ruizhong, et al.
Published: (2024)
by: Qiu, Ruizhong, et al.
Published: (2024)
Promise and Peril of Collaborative Code Generation Models: Balancing Effectiveness and Memorization
by: Chen, Zhi, et al.
Published: (2024)
by: Chen, Zhi, et al.
Published: (2024)
scicode-lint: Detecting Methodology Bugs in Scientific Python Code with LLM-Generated Patterns
by: Samsonau, Sergey V.
Published: (2026)
by: Samsonau, Sergey V.
Published: (2026)
Automatic Generation of Executable BPMN Models from Medical Guidelines
by: Sekar, Praveen Kumar Menaka, et al.
Published: (2026)
by: Sekar, Praveen Kumar Menaka, et al.
Published: (2026)
OBsmith: LLM-Powered JavaScript Obfuscator Testing
by: Jiang, Shan, et al.
Published: (2025)
by: Jiang, Shan, et al.
Published: (2025)
A Survey on Code Generation with LLM-based Agents
by: Dong, Yihong, et al.
Published: (2025)
by: Dong, Yihong, et al.
Published: (2025)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
by: Zheng, Qinkai, et al.
Published: (2023)
by: Zheng, Qinkai, et al.
Published: (2023)
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
by: Bisztray, Tamas, et al.
Published: (2025)
by: Bisztray, Tamas, et al.
Published: (2025)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
by: Palacio, David N., et al.
Published: (2024)
by: Palacio, David N., et al.
Published: (2024)
Your Simulation Runs but Solves the Wrong Physics: PDE-Grounded Intent Verification for LLM-Generated Multiphysics Simulation Code
by: Song, Zhenghan, et al.
Published: (2026)
by: Song, Zhenghan, et al.
Published: (2026)
Automatic Detection of LLM-Generated Code: A Comparative Case Study of Contemporary Models Across Function and Class Granularities
by: Rahman, Musfiqur, et al.
Published: (2024)
by: Rahman, Musfiqur, et al.
Published: (2024)
A Stochastic Differential Equation Framework for Multi-Objective LLM Interactions: Dynamical Systems Analysis with Code Generation Applications
by: Shukla, Shivani, et al.
Published: (2025)
by: Shukla, Shivani, et al.
Published: (2025)
CodeTaste: Can LLMs Generate Human-Level Code Refactorings?
by: Thillen, Alex, et al.
Published: (2026)
by: Thillen, Alex, et al.
Published: (2026)
Can Coding Agents Be General Agents?
by: Ivanov, Maksim, et al.
Published: (2026)
by: Ivanov, Maksim, et al.
Published: (2026)
Optimizing AI-Assisted Code Generation
by: Torka, Simon, et al.
Published: (2024)
by: Torka, Simon, et al.
Published: (2024)
CAPE: Capability Achievement via Policy Execution
by: Ball, David
Published: (2025)
by: Ball, David
Published: (2025)
Code Generation by Differential Test Time Scaling
by: He, Yifeng, et al.
Published: (2026)
by: He, Yifeng, et al.
Published: (2026)
Clover: Closed-Loop Verifiable Code Generation
by: Sun, Chuyue, et al.
Published: (2023)
by: Sun, Chuyue, et al.
Published: (2023)
Functional Overlap Reranking for Neural Code Generation
by: To, Hung Quoc, et al.
Published: (2023)
by: To, Hung Quoc, et al.
Published: (2023)
Your Code Agent Can Grow Alongside You with Structured Memory
by: Deng, Yi-Xuan, et al.
Published: (2026)
by: Deng, Yi-Xuan, et al.
Published: (2026)
Verify Before You Fix: Agentic Execution Grounding for Trustworthy Cross-Language Code Analysis
by: Gajjar, Jugal
Published: (2026)
by: Gajjar, Jugal
Published: (2026)
A Theoretical Analysis of Test-Driven Code Generation
by: Menet, Nicolas, et al.
Published: (2026)
by: Menet, Nicolas, et al.
Published: (2026)
Protocode: Prototype-Driven Interpretability for Code Generation in LLMs
by: Bodla, Krishna Vamshi, et al.
Published: (2025)
by: Bodla, Krishna Vamshi, et al.
Published: (2025)
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
by: Li, Xin-Ye, et al.
Published: (2026)
by: Li, Xin-Ye, et al.
Published: (2026)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
Drawing Pandas: A Benchmark for LLMs in Generating Plotting Code
by: Galimzyanov, Timur, et al.
Published: (2024)
by: Galimzyanov, Timur, et al.
Published: (2024)
Supersonic: Learning to Generate Source Code Optimizations in C/C++
by: Chen, Zimin, et al.
Published: (2023)
by: Chen, Zimin, et al.
Published: (2023)
Co-Located Tests, Better AI Code: How Test Syntax Structure Affects Foundation Model Code Generation
by: Jacopin, Éric
Published: (2026)
by: Jacopin, Éric
Published: (2026)
LLM Benchmarking with LLaMA2: Evaluating Code Development Performance Across Multiple Programming Languages
by: Diehl, Patrick, et al.
Published: (2025)
by: Diehl, Patrick, et al.
Published: (2025)
Understanding LLM-Driven Test Oracle Generation
by: Bodicoat, Adam, et al.
Published: (2026)
by: Bodicoat, Adam, et al.
Published: (2026)
Similar Items
-
Sketch-and-Verify: Structured Inference-Time Scaling via Program Sketching
by: Jiang, Shan, et al.
Published: (2026) -
How Robustly do LLMs Understand Execution Semantics?
by: Spiess, Claudio, et al.
Published: (2026) -
CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution
by: Gu, Alex, et al.
Published: (2024) -
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
by: Yu, Zhuohao, et al.
Published: (2024) -
Mage: Multi-Axis Evaluation of LLM-Generated Executable Game Scenes Beyond Compile-Pass Rate
by: Liu, Hugh Xuechen, et al.
Published: (2026)