Ensemble-Based Uncertainty Estimation for Code Correctness Estimation
Fuente:
arXiv
Saved in:
| Main Authors: | Wei, Yunxiang, Li, Tianlin, Zheng, Yuwei, Dong, Yanni, Liu, Aishan, Hu, Qiang, Zhang, Xiaoyu, Cheng, Mingfei, Yang, Jian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling
by: Wang, Kaixin, et al.
Published: (2025)
by: Wang, Kaixin, et al.
Published: (2025)
Assessing Correctness in LLM-Based Code Generation via Uncertainty Estimation
by: Sharma, Arindam, et al.
Published: (2025)
by: Sharma, Arindam, et al.
Published: (2025)
Investigating Training Data Detection in AI Coders
by: Li, Tianlin, et al.
Published: (2025)
by: Li, Tianlin, et al.
Published: (2025)
From Context to Intent: Reasoning-Guided Function-Level Code Completion
by: Li, Yanzhou, et al.
Published: (2025)
by: Li, Yanzhou, et al.
Published: (2025)
Exploring the Power of Diffusion Large Language Models for Software Engineering: An Empirical Investigation
by: Zhang, Jingyao, et al.
Published: (2025)
by: Zhang, Jingyao, et al.
Published: (2025)
Latent Imitator: Generating Natural Individual Discriminatory Instances for Black-Box Fairness Testing
by: Xiao, Yisong, et al.
Published: (2023)
by: Xiao, Yisong, et al.
Published: (2023)
Software Development Life Cycle Perspective: A Survey of Benchmarks for Code Large Language Models and Agents
by: Wang, Kaixin, et al.
Published: (2025)
by: Wang, Kaixin, et al.
Published: (2025)
Teaching Code LLMs to Use Autocompletion Tools in Repository-Level Code Generation
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
Using Semantic Distance to Estimate Uncertainty in LLM-Based Code Generation
by: He, Weilin, et al.
Published: (2026)
by: He, Weilin, et al.
Published: (2026)
BDefects4NN: A Backdoor Defect Database for Controlled Localization Studies in Neural Networks
by: Xiao, Yisong, et al.
Published: (2024)
by: Xiao, Yisong, et al.
Published: (2024)
Benchmarking and Revisiting Code Generation Assessment: A Mutation-Based Approach
by: Wang, Longtian, et al.
Published: (2025)
by: Wang, Longtian, et al.
Published: (2025)
Rethinking Technology Stack Selection with AI Coding Proficiency
by: Zhang, Xiaoyu, et al.
Published: (2025)
by: Zhang, Xiaoyu, et al.
Published: (2025)
On the Effectiveness of Code Representation in Deep Learning-Based Automated Patch Correctness Assessment
by: Zhang, Quanjun, et al.
Published: (2026)
by: Zhang, Quanjun, et al.
Published: (2026)
Towards Trustworthy LLMs for Code: A Data-Centric Synergistic Auditing Framework
by: Wang, Chong, et al.
Published: (2024)
by: Wang, Chong, et al.
Published: (2024)
Knowledge-Guided Multi-Agent Framework for Application-Level Software Code Generation
by: Xiong, Qian, et al.
Published: (2025)
by: Xiong, Qian, et al.
Published: (2025)
AP2O-Coder: Adaptively Progressive Preference Optimization for Reducing Compilation and Runtime Errors in LLM-Generated Code
by: Zhang, Jianqing, et al.
Published: (2025)
by: Zhang, Jianqing, et al.
Published: (2025)
Moral Testing of Autonomous Driving Systems
by: Tang, Wenbing, et al.
Published: (2025)
by: Tang, Wenbing, et al.
Published: (2025)
Do Code Semantics Help? A Comprehensive Study on Execution Trace-Based Information for Code Large Language Models
by: Wang, Jian, et al.
Published: (2025)
by: Wang, Jian, et al.
Published: (2025)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
by: Zhang, Zhe, et al.
Published: (2025)
by: Zhang, Zhe, et al.
Published: (2025)
Enhancing LLM Code Generation with Ensembles: A Similarity-Based Selection Approach
by: Mahmud, Tarek, et al.
Published: (2025)
by: Mahmud, Tarek, et al.
Published: (2025)
In-Context Learning as an Effective Estimator of Functional Correctness of LLM-Generated Code
by: Das, Susmita, et al.
Published: (2025)
by: Das, Susmita, et al.
Published: (2025)
SemOpt: LLM-Driven Code Optimization via Rule-Based Analysis
by: Zhao, Yuwei, et al.
Published: (2025)
by: Zhao, Yuwei, et al.
Published: (2025)
Ensembling Large Language Models for Code Vulnerability Detection: An Empirical Evaluation
by: Sun, Zhihong, et al.
Published: (2025)
by: Sun, Zhihong, et al.
Published: (2025)
Drivora: A Unified and Extensible Infrastructure for Search-based Autonomous Driving Testing
by: Cheng, Mingfei, et al.
Published: (2026)
by: Cheng, Mingfei, et al.
Published: (2026)
DriveTester: A Unified Platform for Simulation-Based Autonomous Driving Testing
by: Cheng, Mingfei, et al.
Published: (2024)
by: Cheng, Mingfei, et al.
Published: (2024)
ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case Pairs
by: Kong, Jiaolong, et al.
Published: (2024)
by: Kong, Jiaolong, et al.
Published: (2024)
Code Reviewer Recommendation Based on a Hypergraph with Multiplex Relationships
by: Qiao, Yu, et al.
Published: (2024)
by: Qiao, Yu, et al.
Published: (2024)
An Empirical Study of Bugs in Modern LLM Agent Frameworks
by: Zhu, Xinxue, et al.
Published: (2026)
by: Zhu, Xinxue, et al.
Published: (2026)
CodeScore-R: An Automated Robustness Metric for Assessing the FunctionalCorrectness of Code Synthesis
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
Uncertainty-Guided Chain-of-Thought for Code Generation with LLMs
by: Zhu, Yuqi, et al.
Published: (2025)
by: Zhu, Yuqi, et al.
Published: (2025)
CKGFuzzer: LLM-Based Fuzz Driver Generation Enhanced By Code Knowledge Graph
by: Xu, Hanxiang, et al.
Published: (2024)
by: Xu, Hanxiang, et al.
Published: (2024)
Causality-aware Safety Testing for Autonomous Driving Systems
by: Tang, Wenbing, et al.
Published: (2025)
by: Tang, Wenbing, et al.
Published: (2025)
Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language Models
by: Xiao, Yisong, et al.
Published: (2025)
by: Xiao, Yisong, et al.
Published: (2025)
EnCoDe: Energy Estimation of Source Code At Design-Time
by: Goyal, Shailender, et al.
Published: (2026)
by: Goyal, Shailender, et al.
Published: (2026)
GraphCodeAgent: Dual Graph-Guided LLM Agent for Retrieval-Augmented Repo-Level Code Generation
by: Li, Jia, et al.
Published: (2025)
by: Li, Jia, et al.
Published: (2025)
Evaluating and Achieving Controllable Code Completion in Code LLM
by: Zhang, Jiajun, et al.
Published: (2026)
by: Zhang, Jiajun, et al.
Published: (2026)
CITYWALK: Enhancing LLM-Based C++ Unit Test Generation via Project-Dependency Awareness and Language-Specific Knowledge
by: Zhang, Yuwei, et al.
Published: (2025)
by: Zhang, Yuwei, et al.
Published: (2025)
Unveiling Project-Specific Bias in Neural Code Models
by: Li, Zhiming, et al.
Published: (2022)
by: Li, Zhiming, et al.
Published: (2022)
Bootstrapping Code Translation with Weighted Multilanguage Exploration
by: Wu, Yuhan, et al.
Published: (2026)
by: Wu, Yuhan, et al.
Published: (2026)
An Effective Docker Image Slimming Approach Based on Source Code Data Dependency Analysis
by: Han, Jiaxuan, et al.
Published: (2025)
by: Han, Jiaxuan, et al.
Published: (2025)
Similar Items
-
CodeChemist: Functional Knowledge Transfer for Low-Resource Code Generation via Test-Time Scaling
by: Wang, Kaixin, et al.
Published: (2025) -
Assessing Correctness in LLM-Based Code Generation via Uncertainty Estimation
by: Sharma, Arindam, et al.
Published: (2025) -
Investigating Training Data Detection in AI Coders
by: Li, Tianlin, et al.
Published: (2025) -
From Context to Intent: Reasoning-Guided Function-Level Code Completion
by: Li, Yanzhou, et al.
Published: (2025) -
Exploring the Power of Diffusion Large Language Models for Software Engineering: An Empirical Investigation
by: Zhang, Jingyao, et al.
Published: (2025)