When to Answer and When to Defer: A Decision Framework for Reliable Code Predictions
Fuente:
arXiv
Saved in:
| Main Authors: | Rathnasuriya, Ravishka, Yang, Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On-the-Fly Input Adaptation for Reliable Code Intelligence
by: Rathnasuriya, Ravishka, et al.
Published: (2026)
by: Rathnasuriya, Ravishka, et al.
Published: (2026)
Framework for On the Fly Input Refinement for Deep Learning Models
by: Rathnasuriya, Ravishka
Published: (2025)
by: Rathnasuriya, Ravishka
Published: (2025)
CodeImprove: Program Adaptation for Deep Code Models
by: Rathnasuriya, Ravishka, et al.
Published: (2025)
by: Rathnasuriya, Ravishka, et al.
Published: (2025)
Characterizing Real-World Bugs in Tile Programs for Automated Bug Detection
by: Rathnasuriya, Ravishka, et al.
Published: (2026)
by: Rathnasuriya, Ravishka, et al.
Published: (2026)
Can You Mimic Me? Exploring the Use of Android Record & Replay Tools in Debugging
by: Song, Zihe, et al.
Published: (2025)
by: Song, Zihe, et al.
Published: (2025)
HateModerate: Testing Hate Speech Detectors against Content Moderation Policies
by: Zheng, Jiangrui, et al.
Published: (2023)
by: Zheng, Jiangrui, et al.
Published: (2023)
Coding Agents Don't Know When to Act
by: Gloaguen, Thibaud, et al.
Published: (2026)
by: Gloaguen, Thibaud, et al.
Published: (2026)
When Retriever Meets Generator: A Joint Model for Code Comment Generation
by: Le, Tien P. T., et al.
Published: (2025)
by: Le, Tien P. T., et al.
Published: (2025)
When More Retrieval Hurts: Retrieval-Augmented Code Review Generation
by: Meng, Qianru, et al.
Published: (2025)
by: Meng, Qianru, et al.
Published: (2025)
When is Generated Code Difficult to Comprehend? Assessing AI Agent Python Code Proficiency in the Wild
by: Temkulkiat, Nanthit, et al.
Published: (2026)
by: Temkulkiat, Nanthit, et al.
Published: (2026)
What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants
by: Hasan, Alif Al, et al.
Published: (2026)
by: Hasan, Alif Al, et al.
Published: (2026)
A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and How
by: Wang, Chaozheng, et al.
Published: (2024)
by: Wang, Chaozheng, et al.
Published: (2024)
When Code Becomes Abundant: Redefining Software Engineering Around Orchestration and Verification
by: Kohl, Karina, et al.
Published: (2026)
by: Kohl, Karina, et al.
Published: (2026)
When to Stop? Towards Efficient Code Generation in LLMs with Excess Token Prevention
by: Guo, Lianghong, et al.
Published: (2024)
by: Guo, Lianghong, et al.
Published: (2024)
Programming Language Confusion: When Code LLMs Can't Keep their Languages Straight
by: Moumoula, Micheline Bénédicte, et al.
Published: (2025)
by: Moumoula, Micheline Bénédicte, et al.
Published: (2025)
On the Possibility of Breaking Copyleft Licenses When Reusing Code Generated by ChatGPT
by: Colombo, Gaia, et al.
Published: (2025)
by: Colombo, Gaia, et al.
Published: (2025)
When LLMs Lag Behind: Knowledge Conflicts from Evolving APIs in Code Generation
by: Ashik, Ahmed Nusayer, et al.
Published: (2026)
by: Ashik, Ahmed Nusayer, et al.
Published: (2026)
When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation
by: AKLI, Amal, et al.
Published: (2026)
by: AKLI, Amal, et al.
Published: (2026)
SpareCodeSearch: Searching for Code Context When You Have No Spare GPU
by: Nguyen, Minh
Published: (2025)
by: Nguyen, Minh
Published: (2025)
When LLMs Meet API Documentation: Can Retrieval Augmentation Aid Code Generation Just as It Helps Developers?
by: Chen, Jingyi, et al.
Published: (2025)
by: Chen, Jingyi, et al.
Published: (2025)
Fixing Your Own Smells: Adding a Mistake-Based Familiarisation Step When Teaching Code Refactoring
by: Tan, Ivan, et al.
Published: (2024)
by: Tan, Ivan, et al.
Published: (2024)
On the Reliability of Code Comprehension Proxies
by: Arvan, Erfan, et al.
Published: (2026)
by: Arvan, Erfan, et al.
Published: (2026)
ConAIR:Consistency-Augmented Iterative Interaction Framework to Enhance the Reliability of Code Generation
by: Dong, Jinhao, et al.
Published: (2024)
by: Dong, Jinhao, et al.
Published: (2024)
When Labels Are Scarce: A Systematic Mapping of Label-Efficient Code Vulnerability Detection
by: Khalal, Noor, et al.
Published: (2026)
by: Khalal, Noor, et al.
Published: (2026)
LLM-as-a-Judge for Human-AI Co-Creation: A Reliability-Aware Evaluation Framework for Coding
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
by: Amin, Md Faizul Ibne, et al.
Published: (2026)
When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
by: Peng, Yibo, et al.
Published: (2025)
by: Peng, Yibo, et al.
Published: (2025)
An Empirical Study of Java Code Improvements Based on Stack Overflow Answer Edits
by: Wiratsin, In-on, et al.
Published: (2025)
by: Wiratsin, In-on, et al.
Published: (2025)
When the Specification Emerges: Benchmarking Faithfulness Loss in Long-Horizon Coding Agents
by: Yan, Lu, et al.
Published: (2026)
by: Yan, Lu, et al.
Published: (2026)
When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning
by: Morabito, Roberto, et al.
Published: (2025)
by: Morabito, Roberto, et al.
Published: (2025)
When Names Disappear: Revealing What LLMs Actually Understand About Code
by: Le, Cuong Chi, et al.
Published: (2025)
by: Le, Cuong Chi, et al.
Published: (2025)
Multi-Agent Code-Orchestrated Generation for Reliable Infrastructure-as-Code
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
by: Khan, Rana Nameer Hussain, et al.
Published: (2025)
Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
by: Lu, Xu, et al.
Published: (2025)
by: Lu, Xu, et al.
Published: (2025)
AVX / NEON Intrinsic Functions: When Should They Be Used?
by: Boivin, Théo, et al.
Published: (2026)
by: Boivin, Théo, et al.
Published: (2026)
Perplexed: Understanding When Large Language Models are Confused
by: Cooper, Nathan, et al.
Published: (2024)
by: Cooper, Nathan, et al.
Published: (2024)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
ChatGPT: Friend or Foe When Comprehending and Changing Unfamiliar Code
by: Anderson, Norman, et al.
Published: (2026)
by: Anderson, Norman, et al.
Published: (2026)
CodeArt: Better Code Models by Attention Regularization When Symbols Are Lacking
by: Su, Zian, et al.
Published: (2024)
by: Su, Zian, et al.
Published: (2024)
Manifestations of Empathy in Software Engineering: How, Why, and When It Matters
by: Gunatilake, Hashini, et al.
Published: (2025)
by: Gunatilake, Hashini, et al.
Published: (2025)
Is LLM-Generated Code More Maintainable \& Reliable than Human-Written Code?
by: Molison, Alfred Santa, et al.
Published: (2025)
by: Molison, Alfred Santa, et al.
Published: (2025)
A Tale of Two DL Cities: When Library Tests Meet Compiler
by: Shen, Qingchao, et al.
Published: (2024)
by: Shen, Qingchao, et al.
Published: (2024)
Similar Items
-
On-the-Fly Input Adaptation for Reliable Code Intelligence
by: Rathnasuriya, Ravishka, et al.
Published: (2026) -
Framework for On the Fly Input Refinement for Deep Learning Models
by: Rathnasuriya, Ravishka
Published: (2025) -
CodeImprove: Program Adaptation for Deep Code Models
by: Rathnasuriya, Ravishka, et al.
Published: (2025) -
Characterizing Real-World Bugs in Tile Programs for Automated Bug Detection
by: Rathnasuriya, Ravishka, et al.
Published: (2026) -
Can You Mimic Me? Exploring the Use of Android Record & Replay Tools in Debugging
by: Song, Zihe, et al.
Published: (2025)