ThrowBench: Benchmarking LLMs by Predicting Runtime Exceptions
Fuente:
arXiv
Saved in:
| Main Authors: | Prenner, Julian Aron, Robbes, Romain |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Extracting Fix Ingredients using Language Models
by: Prenner, Julian Aron, et al.
Published: (2025)
by: Prenner, Julian Aron, et al.
Published: (2025)
Bogus Bugs, Duplicates, and Revealing Comments: Data Quality Issues in NPR
by: Prenner, Julian Aron, et al.
Published: (2025)
by: Prenner, Julian Aron, et al.
Published: (2025)
Simple Fault Localization using Execution Traces
by: Prenner, Julian Aron, et al.
Published: (2025)
by: Prenner, Julian Aron, et al.
Published: (2025)
OSS-Bench: Benchmark Generator for Coding LLMs
by: Jiang, Yuancheng, et al.
Published: (2025)
by: Jiang, Yuancheng, et al.
Published: (2025)
DRAGON: Robust Classification for Very Large Collections of Software Repositories
by: Balla, Stefano, et al.
Published: (2026)
by: Balla, Stefano, et al.
Published: (2026)
Are Coding Agents Generating Over-Mocked Tests? An Empirical Study
by: Hora, Andre, et al.
Published: (2026)
by: Hora, Andre, et al.
Published: (2026)
AI Policy, Disclosure, and Human in the Loop: How Are Contribution Guidelines Adapting to GenAI?
by: Hora, Andre, et al.
Published: (2026)
by: Hora, Andre, et al.
Published: (2026)
Themisto: Jupyter-Based Runtime Benchmark
by: Grotov, Konstantin, et al.
Published: (2025)
by: Grotov, Konstantin, et al.
Published: (2025)
EnvBench: A Benchmark for Automated Environment Setup
by: Eliseeva, Aleksandra, et al.
Published: (2025)
by: Eliseeva, Aleksandra, et al.
Published: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
by: Orel, Daniil, et al.
Published: (2026)
by: Orel, Daniil, et al.
Published: (2026)
CPP-UT-Bench: Can LLMs Write Complex Unit Tests in C++?
by: Bhargava, Vaishnavi, et al.
Published: (2024)
by: Bhargava, Vaishnavi, et al.
Published: (2024)
AssertionBench: A Benchmark to Evaluate Large-Language Models for Assertion Generation
by: Pulavarthi, Vaishnavi, et al.
Published: (2024)
by: Pulavarthi, Vaishnavi, et al.
Published: (2024)
MobileDev-Bench: A Benchmark for Issue Resolution in Mobile Application Development
by: Fakorede, Moshood A., et al.
Published: (2026)
by: Fakorede, Moshood A., et al.
Published: (2026)
LLM-Guided Runtime Parameter Optimization for Energy-Efficient Model Inference
by: Crumpacker, Katelyn, et al.
Published: (2026)
by: Crumpacker, Katelyn, et al.
Published: (2026)
Cerberus: Multi-Agent Reasoning and Coverage-Guided Exploration for Static Detection of Runtime Errors
by: Dhulipala, Hridya, et al.
Published: (2025)
by: Dhulipala, Hridya, et al.
Published: (2025)
Where Do LLMs Still Struggle? An In-Depth Analysis of Code Generation Benchmarks
by: Sharifloo, Amir Molzam, et al.
Published: (2025)
by: Sharifloo, Amir Molzam, et al.
Published: (2025)
Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
by: McGregor, Sean, et al.
Published: (2025)
by: McGregor, Sean, et al.
Published: (2025)
Deep-Bench: Deep Learning Benchmark Dataset for Code Generation
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
by: Daghighfarsoodeh, Alireza, et al.
Published: (2025)
SoundnessBench: A Soundness Benchmark for Neural Network Verifiers
by: Zhou, Xingjian, et al.
Published: (2024)
by: Zhou, Xingjian, et al.
Published: (2024)
Green AI: A Preliminary Empirical Study on Energy Consumption in DL Models Across Different Runtime Infrastructures
by: Alizadeh, Negar, et al.
Published: (2024)
by: Alizadeh, Negar, et al.
Published: (2024)
CRUST-Bench: A Comprehensive Benchmark for C-to-safe-Rust Transpilation
by: Khatry, Anirudh, et al.
Published: (2025)
by: Khatry, Anirudh, et al.
Published: (2025)
Runtime Anomaly Detection for Drones: An Integrated Rule-Mining and Unsupervised-Learning Approach
by: Tan, Ivan, et al.
Published: (2025)
by: Tan, Ivan, et al.
Published: (2025)
Client--Library Compatibility Testing with API Interaction Snapshots
by: Monce, Gustave, et al.
Published: (2025)
by: Monce, Gustave, et al.
Published: (2025)
Pimp My LLM: Leveraging Variability Modeling to Tune Inference Hyperparameters
by: Zine, Nada, et al.
Published: (2026)
by: Zine, Nada, et al.
Published: (2026)
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
by: Shbita, Basel, et al.
Published: (2025)
by: Shbita, Basel, et al.
Published: (2025)
CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research Repositories
by: Xiao, Yijia, et al.
Published: (2025)
by: Xiao, Yijia, et al.
Published: (2025)
DevBench: A Realistic, Developer-Informed Benchmark for Code Generation Models
by: Kumarappan, Adarsh, et al.
Published: (2026)
by: Kumarappan, Adarsh, et al.
Published: (2026)
RepairBench: Leaderboard of Frontier Models for Program Repair
by: Silva, André, et al.
Published: (2024)
by: Silva, André, et al.
Published: (2024)
TeleResilienceBench: Quantifying Resilience for LLM Reasoning in Telecommunications
by: Gajjar, Pranshav, et al.
Published: (2026)
by: Gajjar, Pranshav, et al.
Published: (2026)
TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
by: Ahmed, Toufique, et al.
Published: (2024)
by: Ahmed, Toufique, et al.
Published: (2024)
CONCUR: Benchmarking LLMs for Concurrent Code Generation
by: Huang, Jue, et al.
Published: (2026)
by: Huang, Jue, et al.
Published: (2026)
StackEval: Benchmarking LLMs in Coding Assistance
by: Shah, Nidhish, et al.
Published: (2024)
by: Shah, Nidhish, et al.
Published: (2024)
KernelBench: Can LLMs Write Efficient GPU Kernels?
by: Ouyang, Anne, et al.
Published: (2025)
by: Ouyang, Anne, et al.
Published: (2025)
CoDocBench: A Dataset for Code-Documentation Alignment in Software Maintenance
by: Pai, Kunal, et al.
Published: (2025)
by: Pai, Kunal, et al.
Published: (2025)
InfiBench: Evaluating the Question-Answering Capabilities of Code Large Language Models
by: Li, Linyi, et al.
Published: (2024)
by: Li, Linyi, et al.
Published: (2024)
Promises, Perils, and (Timely) Heuristics for Mining Coding Agent Activity
by: Robbes, Romain, et al.
Published: (2026)
by: Robbes, Romain, et al.
Published: (2026)
Package-Aware Approach for Repository-Level Code Completion in Pharo
by: Abedelkader, Omar, et al.
Published: (2026)
by: Abedelkader, Omar, et al.
Published: (2026)
Agentic Much? Adoption of Coding Agents on GitHub
by: Robbes, Romain, et al.
Published: (2026)
by: Robbes, Romain, et al.
Published: (2026)
Debugging and Runtime Analysis of Neural Networks with VLMs (A Case Study)
by: Hu, Boyue Caroline, et al.
Published: (2025)
by: Hu, Boyue Caroline, et al.
Published: (2025)
Automating Code Adaptation for MLOps -- A Benchmarking Study on LLMs
by: Patel, Harsh, et al.
Published: (2024)
by: Patel, Harsh, et al.
Published: (2024)
Similar Items
-
Extracting Fix Ingredients using Language Models
by: Prenner, Julian Aron, et al.
Published: (2025) -
Bogus Bugs, Duplicates, and Revealing Comments: Data Quality Issues in NPR
by: Prenner, Julian Aron, et al.
Published: (2025) -
Simple Fault Localization using Execution Traces
by: Prenner, Julian Aron, et al.
Published: (2025) -
OSS-Bench: Benchmark Generator for Coding LLMs
by: Jiang, Yuancheng, et al.
Published: (2025) -
DRAGON: Robust Classification for Very Large Collections of Software Repositories
by: Balla, Stefano, et al.
Published: (2026)