Saved in:
| Main Authors: | Khan, Saadiq Rauf, Chandak, Vinit, Mukherjea, Sougata |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.10996 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Evaluating LLMs for Visualization Generation and Understanding
by: Khan, Saadiq Rauf, et al.
Published: (2025)
by: Khan, Saadiq Rauf, et al.
Published: (2025)
TaskEval: Synthesised Evaluation for Foundation-Model Tasks
by: Widanapathiranage, Dilani, et al.
Published: (2025)
by: Widanapathiranage, Dilani, et al.
Published: (2025)
Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities
by: Garg, Aayush, et al.
Published: (2025)
by: Garg, Aayush, et al.
Published: (2025)
Mind the Prompt: Self-adaptive Generation of Task Plan Explanations via LLMs
by: Vázquez, Gricel, et al.
Published: (2026)
by: Vázquez, Gricel, et al.
Published: (2026)
Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs
by: Chen, Yujia, et al.
Published: (2026)
by: Chen, Yujia, et al.
Published: (2026)
Evaluating the Energy-Efficiency of the Code Generated by LLMs
by: Islam, Md Arman, et al.
Published: (2025)
by: Islam, Md Arman, et al.
Published: (2025)
Evaluating the Generalizability of LLMs in Automated Program Repair
by: Li, Fengjie, et al.
Published: (2025)
by: Li, Fengjie, et al.
Published: (2025)
CoRe: Benchmarking LLMs Code Reasoning Capabilities through Static Analysis Tasks
by: Xie, Danning, et al.
Published: (2025)
by: Xie, Danning, et al.
Published: (2025)
Holistic Evaluation of State-of-the-Art LLMs for Code Generation
by: Zhang, Le, et al.
Published: (2025)
by: Zhang, Le, et al.
Published: (2025)
Using LLMs in Software Requirements Specifications: An Empirical Evaluation
by: Krishna, Madhava, et al.
Published: (2024)
by: Krishna, Madhava, et al.
Published: (2024)
InteractiveGNNExplainer: A Visual Analytics Framework for Multi-Faceted Understanding and Probing of Graph Neural Network Predictions
by: Singh, TC, et al.
Published: (2025)
by: Singh, TC, et al.
Published: (2025)
A Multi-Dataset Evaluation of Models for Automated Vulnerability Repair
by: Khan, Zanis Ali, et al.
Published: (2025)
by: Khan, Zanis Ali, et al.
Published: (2025)
Benchmark Dataset Generation and Evaluation for Excel Formula Repair with LLMs
by: Singha, Ananya, et al.
Published: (2025)
by: Singha, Ananya, et al.
Published: (2025)
Evaluating LLM-Based Test Generation Under Software Evolution
by: Haroon, Sabaat, et al.
Published: (2026)
by: Haroon, Sabaat, et al.
Published: (2026)
TAM-Eval: Evaluating LLMs for Automated Unit Test Maintenance
by: Bruches, Elena, et al.
Published: (2026)
by: Bruches, Elena, et al.
Published: (2026)
FullStack Bench: Evaluating LLMs as Full Stack Coders
by: Bytedance-Seed-Foundation-Code-Team, et al.
Published: (2024)
by: Bytedance-Seed-Foundation-Code-Team, et al.
Published: (2024)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
by: Zhu, Hongda, et al.
Published: (2025)
by: Zhu, Hongda, et al.
Published: (2025)
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
by: Garg, Spandan, et al.
Published: (2026)
by: Garg, Spandan, et al.
Published: (2026)
MergeRepair: An Exploratory Study on Merging Task-Specific Adapters in Code LLMs for Automated Program Repair
by: Dehghan, Meghdad, et al.
Published: (2024)
by: Dehghan, Meghdad, et al.
Published: (2024)
Evaluating the Effectiveness of LLMs in Fixing Maintainability Issues in Real-World Projects
by: Nunes, Henrique, et al.
Published: (2025)
by: Nunes, Henrique, et al.
Published: (2025)
Open the Oyster: Empirical Evaluation and Improvement of Code Reasoning Confidence in LLMs
by: Wang, Shufan, et al.
Published: (2025)
by: Wang, Shufan, et al.
Published: (2025)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
by: Gao, Shuzheng, et al.
Published: (2025)
by: Gao, Shuzheng, et al.
Published: (2025)
GBQA: A Game Benchmark for Evaluating LLMs as Quality Assurance Engineers
by: Jiang, Shufan, et al.
Published: (2026)
by: Jiang, Shufan, et al.
Published: (2026)
Mutation-based Consistency Testing for Evaluating the Code Understanding Capability of LLMs
by: Li, Ziyu, et al.
Published: (2024)
by: Li, Ziyu, et al.
Published: (2024)
Adaptive Hierarchical Evaluation of LLMs and SAST tools for CWE Prediction in Python
by: Adnan, Muntasir, et al.
Published: (2026)
by: Adnan, Muntasir, et al.
Published: (2026)
LLMs: A Game-Changer for Software Engineers?
by: Haque, Md Asraful
Published: (2024)
by: Haque, Md Asraful
Published: (2024)
Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
by: Gong, Linyuan, et al.
Published: (2024)
by: Gong, Linyuan, et al.
Published: (2024)
From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
by: Ho, Anh, et al.
Published: (2025)
by: Ho, Anh, et al.
Published: (2025)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
by: Li, Chunyang, et al.
Published: (2025)
by: Li, Chunyang, et al.
Published: (2025)
Beyond Isolated Tasks: A Framework for Evaluating Coding Agents on Sequential Software Evolution
by: Shastry, KN Ajay, et al.
Published: (2026)
by: Shastry, KN Ajay, et al.
Published: (2026)
Evaluating the Use of LLMs for Automated DOM-Level Resolution of Web Performance Issues
by: Peters, Gideon, et al.
Published: (2026)
by: Peters, Gideon, et al.
Published: (2026)
PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
by: Ravishankara, Mayank
Published: (2026)
by: Ravishankara, Mayank
Published: (2026)
Using LLMs in Software Design: An Empirical Study of GitHub and A Practitioner Survey
by: Wang, Yifei, et al.
Published: (2026)
by: Wang, Yifei, et al.
Published: (2026)
Capture the Flags: Family-Based Evaluation of Agentic LLMs via Semantics-Preserving Transformations
by: Honarvar, Shahin, et al.
Published: (2026)
by: Honarvar, Shahin, et al.
Published: (2026)
LLMs in Web Development: Evaluating LLM-Generated PHP Code Unveiling Vulnerabilities and Limitations
by: Tóth, Rebeka, et al.
Published: (2024)
by: Tóth, Rebeka, et al.
Published: (2024)
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
by: Liu, Jingjing, et al.
Published: (2025)
by: Liu, Jingjing, et al.
Published: (2025)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
by: Larbi, Maya, et al.
Published: (2025)
by: Larbi, Maya, et al.
Published: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Coding in a Bubble? Evaluating LLMs in Resolving Context Adaptation Bugs During Code Adaptation
by: Zhang, Tanghaoran, et al.
Published: (2026)
by: Zhang, Tanghaoran, et al.
Published: (2026)
InteractScience: Programmatic and Visually-Grounded Evaluation of Interactive Scientific Demonstration Code Generation
by: Chen, Qiaosheng, et al.
Published: (2025)
by: Chen, Qiaosheng, et al.
Published: (2025)
Similar Items
-
Evaluating LLMs for Visualization Generation and Understanding
by: Khan, Saadiq Rauf, et al.
Published: (2025) -
TaskEval: Synthesised Evaluation for Foundation-Model Tasks
by: Widanapathiranage, Dilani, et al.
Published: (2025) -
Evaluating LLMs for One-Shot Patching of Real and Artificial Vulnerabilities
by: Garg, Aayush, et al.
Published: (2025) -
Mind the Prompt: Self-adaptive Generation of Task Plan Explanations via LLMs
by: Vázquez, Gricel, et al.
Published: (2026) -
Schedule-and-Calibrate: Utility-Guided Multi-Task Reinforcement Learning for Code LLMs
by: Chen, Yujia, et al.
Published: (2026)