Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain
Fuente:
arXiv
Saved in:
| Main Authors: | Min, Marcus J., Ding, Yangruibo, Buratti, Luca, Pujar, Saurabh, Kaiser, Gail, Jana, Suman, Ray, Baishakhi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CYCLE: Learning to Self-Refine the Code Generation
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study
by: Ceka, Ira, et al.
Published: (2025)
by: Ceka, Ira, et al.
Published: (2025)
Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action
by: Pujar, Saurabh, et al.
Published: (2025)
by: Pujar, Saurabh, et al.
Published: (2025)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
by: Berman, Shmuel, et al.
Published: (2024)
by: Berman, Shmuel, et al.
Published: (2024)
AdvFusion: Adapter-based Knowledge Transfer for Code Summarization on Code Language Models
by: Saberi, Iman, et al.
Published: (2023)
by: Saberi, Iman, et al.
Published: (2023)
TDAD: Test-Driven Agentic Development - Reducing Code Regressions in AI Coding Agents via Graph-Based Impact Analysis
by: Alonso, Pepe, et al.
Published: (2026)
by: Alonso, Pepe, et al.
Published: (2026)
Automated Code Editing with Search-Generate-Modify
by: Liu, Changshu, et al.
Published: (2023)
by: Liu, Changshu, et al.
Published: (2023)
Beyond Accuracy: Decomposing the Reasoning Efficiency of LLMs
by: Kaiser, Daniel, et al.
Published: (2026)
by: Kaiser, Daniel, et al.
Published: (2026)
A Prompt Learning Framework for Source Code Summarization
by: Xu, Tingting, et al.
Published: (2023)
by: Xu, Tingting, et al.
Published: (2023)
Insights from the Usage of the Ansible Lightspeed Code Completion Service
by: Sahoo, Priyam, et al.
Published: (2024)
by: Sahoo, Priyam, et al.
Published: (2024)
SeaView: Software Engineering Agent Visual Interface for Enhanced Workflow
by: Bula, Timothy, et al.
Published: (2025)
by: Bula, Timothy, et al.
Published: (2025)
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
by: Basit, Abdul, et al.
Published: (2025)
by: Basit, Abdul, et al.
Published: (2025)
React-ing to Grace Hopper 200: Five Open-Weights Coding Models, One React Native App, One GH200, One Weekend
by: Potanin, Alex
Published: (2026)
by: Potanin, Alex
Published: (2026)
FEM-Bench: A Structured Scientific Reasoning Benchmark for Evaluating Code-Generating LLMs
by: Mohammadzadeh, Saeed, et al.
Published: (2025)
by: Mohammadzadeh, Saeed, et al.
Published: (2025)
Agile Effort Estimation: Comparing the Accuracy and Efficiency of Planning Poker, Bucket System, and Affinity Estimation methods
by: Poženel, Marko, et al.
Published: (2024)
by: Poženel, Marko, et al.
Published: (2024)
Commenting Higher-level Code Unit: Full Code, Reduced Code, or Hierarchical Code Summarization
by: Sun, Weisong, et al.
Published: (2025)
by: Sun, Weisong, et al.
Published: (2025)
A RAG Method for Source Code Inquiry Tailored to Long-Context LLMs
by: Kamiya, Toshihiro
Published: (2024)
by: Kamiya, Toshihiro
Published: (2024)
The Right Prompts for the Job: Repair Code-Review Defects with Large Language Model
by: Zhao, Zelin, et al.
Published: (2023)
by: Zhao, Zelin, et al.
Published: (2023)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
by: Haseeb, Muhammad
Published: (2025)
by: Haseeb, Muhammad
Published: (2025)
SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
by: Vargas, Matheus J. T.
Published: (2025)
by: Vargas, Matheus J. T.
Published: (2025)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
by: Xia, Bowei, et al.
Published: (2026)
by: Xia, Bowei, et al.
Published: (2026)
Eliminating Backdoors in Neural Code Models for Secure Code Understanding
by: Sun, Weisong, et al.
Published: (2024)
by: Sun, Weisong, et al.
Published: (2024)
ESALE: Enhancing Code-Summary Alignment Learning for Source Code Summarization
by: Fang, Chunrong, et al.
Published: (2024)
by: Fang, Chunrong, et al.
Published: (2024)
GraphSense: Graph Embedding Based Code Suggestion Framework
by: Peiris, H. R Navod Thisura
Published: (2025)
by: Peiris, H. R Navod Thisura
Published: (2025)
Anchor Attention, Small Cache: Code Generation with Large Language Models
by: Zhang, Xiangyu, et al.
Published: (2024)
by: Zhang, Xiangyu, et al.
Published: (2024)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
See-Saw Generative Mechanism for Scalable Recursive Code Generation with Generative AI
by: Vsevolodovna, Ruslan Idelfonso Magaña
Published: (2024)
by: Vsevolodovna, Ruslan Idelfonso Magaña
Published: (2024)
Red Teaming Program Repair Agents: When Correct Patches can Hide Vulnerabilities
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Source Code Summarization in the Era of Large Language Models
by: Sun, Weisong, et al.
Published: (2024)
by: Sun, Weisong, et al.
Published: (2024)
The Integration of Agile Methodologies in DevOps Practices within the Information Technology Industry
by: Hourigan, Ashley, et al.
Published: (2025)
by: Hourigan, Ashley, et al.
Published: (2025)
Chain-Oriented Objective Logic with Neural Network Feedback Control and Cascade Filtering for Dynamic Multi-DSL Regulation
by: Han, Jipeng
Published: (2024)
by: Han, Jipeng
Published: (2024)
Optimizing Large Language Models for OpenAPI Code Completion
by: Petryshyn, Bohdan, et al.
Published: (2024)
by: Petryshyn, Bohdan, et al.
Published: (2024)
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
by: Hahn, Joonghyuk, et al.
Published: (2025)
by: Hahn, Joonghyuk, et al.
Published: (2025)
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
by: Lim, Soohan, et al.
Published: (2025)
by: Lim, Soohan, et al.
Published: (2025)
Show Me Your Code! Kill Code Poisoning: A Lightweight Method Based on Code Naturalness
by: Sun, Weisong, et al.
Published: (2025)
by: Sun, Weisong, et al.
Published: (2025)
Securing the Dark Matter: A Semantic-Enhanced Neuro-Symbolic Framework for Supply Chain Analysis of Opaque Industrial Software
by: Ning, Bowei, et al.
Published: (2026)
by: Ning, Bowei, et al.
Published: (2026)
Validation of an analyzability model for quantum software: a family of experiments
by: Díaz-Muñoz, Ana, et al.
Published: (2026)
by: Díaz-Muñoz, Ana, et al.
Published: (2026)
GraphSkill: Documentation-Guided Hierarchical Retrieval-Augmented Coding for Complex Graph Reasoning
by: Wang, Fali, et al.
Published: (2026)
by: Wang, Fali, et al.
Published: (2026)
Mind the Metrics: Patterns for Telemetry-Aware In-IDE AI Application Development using the Model Context Protocol (MCP)
by: Koc, Vincent, et al.
Published: (2025)
by: Koc, Vincent, et al.
Published: (2025)
Similar Items
-
CYCLE: Learning to Self-Refine the Code Generation
by: Ding, Yangruibo, et al.
Published: (2024) -
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
by: Ding, Yangruibo, et al.
Published: (2024) -
Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study
by: Ceka, Ira, et al.
Published: (2025) -
Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action
by: Pujar, Saurabh, et al.
Published: (2025) -
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
by: Berman, Shmuel, et al.
Published: (2024)