Tricky$^2$: Towards a Benchmark for Evaluating Human and LLM Error Interactions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Granger, Cole, Khati, Dipin, Rodriguez-Cardenas, Daniel, Poshyvanyk, Denys |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
par: Khati, Dipin, et autres
Publié: (2026)
par: Khati, Dipin, et autres
Publié: (2026)
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
par: Palacio, David N., et autres
Publié: (2024)
par: Palacio, David N., et autres
Publié: (2024)
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
par: Khati, Dipin, et autres
Publié: (2025)
par: Khati, Dipin, et autres
Publié: (2025)
A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code
par: Velasco, Alejandro, et autres
Publié: (2025)
par: Velasco, Alejandro, et autres
Publié: (2025)
Enabling Global, Human-Centered Explanations for LLMs:From Tokens to Interpretable Code and Test Generation
par: Khati, Dipin, et autres
Publié: (2025)
par: Khati, Dipin, et autres
Publié: (2025)
SnipGen: A Mining Repository Framework for Evaluating LLMs for Code
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
How Propense Are Large Language Models at Producing Code Smells? A Benchmarking Study
par: Velasco, Alejandro, et autres
Publié: (2024)
par: Velasco, Alejandro, et autres
Publié: (2024)
On Interpreting the Effectiveness of Unsupervised Software Traceability with Information Theory
par: Palacio, David N., et autres
Publié: (2024)
par: Palacio, David N., et autres
Publié: (2024)
Toward Neurosymbolic Program Comprehension
par: Velasco, Alejandro, et autres
Publié: (2025)
par: Velasco, Alejandro, et autres
Publié: (2025)
Towards Enabling An Artificial Self-Construction Software Life-cycle via Autopoietic Architectures
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026)
"Don't Be Afraid, Just Learn": Insights from Industry Practitioners to Prepare Software Engineers in the Age of Generative AI
par: Otten, Daniel, et autres
Publié: (2026)
par: Otten, Daniel, et autres
Publié: (2026)
Toward a Theory of Causation for Interpreting Neural Code Models
par: Palacio, David N., et autres
Publié: (2023)
par: Palacio, David N., et autres
Publié: (2023)
Toward Explaining Large Language Models in Software Engineering Tasks
par: Vitale, Antonio, et autres
Publié: (2025)
par: Vitale, Antonio, et autres
Publié: (2025)
Understanding Privacy Risks in Code Models Through Training Dynamics: A Causal Approach
par: Yang, Hua, et autres
Publié: (2025)
par: Yang, Hua, et autres
Publié: (2025)
Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development
par: Stalnaker, Trevor, et autres
Publié: (2024)
par: Stalnaker, Trevor, et autres
Publié: (2024)
A Path Less Traveled: Reimagining Software Engineering Automation via a Neurosymbolic Paradigm
par: Mastropaolo, Antonio, et autres
Publié: (2025)
par: Mastropaolo, Antonio, et autres
Publié: (2025)
Which Syntactic Capabilities Are Statistically Learned by Masked Language Models for Code?
par: Velasco, Alejandro, et autres
Publié: (2024)
par: Velasco, Alejandro, et autres
Publié: (2024)
How Do Semantically Equivalent Code Transformations Impact Membership Inference on LLMs for Code?
par: Yang, Hua, et autres
Publié: (2025)
par: Yang, Hua, et autres
Publié: (2025)
Re-Evaluating Code LLM Benchmarks Under Semantic Mutation
par: Pan, Zhiyuan, et autres
Publié: (2025)
par: Pan, Zhiyuan, et autres
Publié: (2025)
Towards More Trustworthy Deep Code Models by Enabling Out-of-Distribution Detection
par: Yan, Yanfu, et autres
Publié: (2025)
par: Yan, Yanfu, et autres
Publié: (2025)
LoCoBench-Agent: An Interactive Benchmark for LLM Agents in Long-Context Software Engineering
par: Qiu, Jielin, et autres
Publié: (2025)
par: Qiu, Jielin, et autres
Publié: (2025)
Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We?
par: O'Brien, Conor, et autres
Publié: (2024)
par: O'Brien, Conor, et autres
Publié: (2024)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
par: Wang, Ruiqi, et autres
Publié: (2025)
par: Wang, Ruiqi, et autres
Publié: (2025)
EvoCodeBench: A Human-Performance Benchmark for Self-Evolving LLM-Driven Coding Systems
par: Zhang, Wentao, et autres
Publié: (2026)
par: Zhang, Wentao, et autres
Publié: (2026)
ToolScan: A Benchmark for Characterizing Errors in Tool-Use LLMs
par: Kokane, Shirley, et autres
Publié: (2024)
par: Kokane, Shirley, et autres
Publié: (2024)
CodeFuse-CommitEval: Towards Benchmarking LLM's Power on Commit Message and Code Change Inconsistency Detection
par: Zhang, Qingyu, et autres
Publié: (2025)
par: Zhang, Qingyu, et autres
Publié: (2025)
Eliminating Hallucination-Induced Errors in LLM Code Generation with Functional Clustering
par: Ravuri, Chaitanya, et autres
Publié: (2025)
par: Ravuri, Chaitanya, et autres
Publié: (2025)
Testing Practices, Challenges, and Developer Perspectives in Open-Source IoT Platforms
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2025)
ComBench: A Repo-level Real-world Benchmark for Compilation Error Repair
par: Li, Jia, et autres
Publié: (2026)
par: Li, Jia, et autres
Publié: (2026)
RESTestBench: A Benchmark for Evaluating the Effectiveness of LLM-Generated REST API Test Cases from NL Requirements
par: Kogler, Leon, et autres
Publié: (2026)
par: Kogler, Leon, et autres
Publié: (2026)
On the Effectiveness of LLM-as-a-judge for Code Generation and Summarization
par: Crupi, Giuseppe, et autres
Publié: (2025)
par: Crupi, Giuseppe, et autres
Publié: (2025)
RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation
par: Liang, Linxi, et autres
Publié: (2025)
par: Liang, Linxi, et autres
Publié: (2025)
Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages
par: Wu, Fan, et autres
Publié: (2026)
par: Wu, Fan, et autres
Publié: (2026)
From Empirical Evaluation to Context-Aware Enhancement: Repairing Regression Errors with LLMs
par: Ho, Anh, et autres
Publié: (2025)
par: Ho, Anh, et autres
Publié: (2025)
LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps
par: Zhao, Shanhui, et autres
Publié: (2025)
par: Zhao, Shanhui, et autres
Publié: (2025)
Rethinking Testing for LLM Applications: Characteristics, Challenges, and a Lightweight Interaction Protocol
par: Ma, Wei, et autres
Publié: (2025)
par: Ma, Wei, et autres
Publié: (2025)
Copilot Evaluation Harness: Evaluating LLM-Guided Software Programming
par: Agarwal, Anisha, et autres
Publié: (2024)
par: Agarwal, Anisha, et autres
Publié: (2024)
Evaluating the effectiveness of LLM-based interoperability
par: Falcão, Rodrigo, et autres
Publié: (2025)
par: Falcão, Rodrigo, et autres
Publié: (2025)
LLM Code Customization with Visual Results: A Benchmark on TikZ
par: Reux, Charly, et autres
Publié: (2025)
par: Reux, Charly, et autres
Publié: (2025)
Documents similaires
-
Detecting and Correcting Hallucinations in LLM-Generated Code via Deterministic AST Analysis
par: Khati, Dipin, et autres
Publié: (2026) -
Towards Comprehensive Benchmarking Infrastructure for LLMs In Software Engineering
par: Rodriguez-Cardenas, Daniel, et autres
Publié: (2026) -
Towards More Trustworthy and Interpretable LLMs for Code through Syntax-Grounded Explanations
par: Palacio, David N., et autres
Publié: (2024) -
Mapping the Trust Terrain: LLMs in Software Engineering -- Insights and Perspectives
par: Khati, Dipin, et autres
Publié: (2025) -
A Causal Perspective on Measuring, Explaining and Mitigating Smells in LLM-Generated Code
par: Velasco, Alejandro, et autres
Publié: (2025)