LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
Fuente:
Zenodo
Saved in:
| Main Author: | Anonymous, Anonymous |
|---|---|
| Format: | Recurso digital |
| Published: |
Zenodo
2025
|
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
by: Anonymous, Anonymous
Published: (2024)
by: Anonymous, Anonymous
Published: (2024)
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
by: Anonymous, Anonymous
Published: (2024)
by: Anonymous, Anonymous
Published: (2024)
Replication kit for: Can LLMs Make Software Testing Greener? An Empirical Study on JUnit Test Energy Reengineering
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Supplementary data for Perspectives, Needs and Challenges for Sustainable Software Engineering Teams: A FinServ Case Study
by: Anonymous, Anonymous
Published: (2025)
by: Anonymous, Anonymous
Published: (2025)
Connecting Generations Through Code: GABI, A Community-Driven Framework for Engineering Inclusive Financial Software for the Elderly
by: Anonymous, Anonymous
Published: (2025)
by: Anonymous, Anonymous
Published: (2025)
[Tool Demo] Leveraging LLMs to Judge Student-Authored Software Requirements and User Stories
by: Anonymous
Published: (2026)
by: Anonymous
Published: (2026)
How do GitHub Repositories Support Software Engineering Candidates to Prepare for Interviews?
by: Anonymous Researcher
Published: (2026)
by: Anonymous Researcher
Published: (2026)
Supplementary Material for the Article "Leadership Styles Across Civilizational Contexts: A Cross-Cultural Empirical Study"
by: Anonymous
Published: (2026)
by: Anonymous
Published: (2026)
Replication kit of: On the Energy Cost of Static Analysis Precision An Empirical Study of SpotBugs Effort Levels
by: Anonymous
Published: (2026)
by: Anonymous
Published: (2026)
Supplementary Material - Requirements of a Guide to Cloud Migration Strategies for Legacy Software Assets in Software Ecosystems
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
How to Get Insight from Disagreeing Post-hoc Explanations in ML-based Defect Prediction Methods? An Empirical Study on Defect Predictors
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Failure Modes and Effects Analysis: An Experience from the E-Bike Domain
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Replication Package for "Green AI: Cost of LLM-based Code Completion"
by: Anonymous, Anonymous
Published: (2026)
by: Anonymous, Anonymous
Published: (2026)
LLM-based Structure Extraction Prompt for the Mapping Study
by: Anonymous Authors
Published: (2026)
by: Anonymous Authors
Published: (2026)
On Automatically Generating Assertions for Validating Reported Failures in Android Apps
by: Anonymous Authors
Published: (2025)
by: Anonymous Authors
Published: (2025)
Dataset For: From Disruptions to Discussions: How GenAI Impacts Human Interactions in Software Development
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Dataset For: From Disruptions to Discussions: How GenAI Impacts Human Interactions in Software Development
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Longitudinal Analyses of Static Analysis Tools: A CodeQL Case Study
by: Anonymous, Anonymous
Published: (2026)
by: Anonymous, Anonymous
Published: (2026)
Exploring Generative AI Tools for Software Quality: Insights from a Rapid Multivocal Literature Review
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Does Good Design Prevent Code Clones? An Exploratory Study on SOLID Principles in Java Systems
by: Anonymous, Anonymous
Published: (2025)
by: Anonymous, Anonymous
Published: (2025)
Artifact for "Are LLMs all we need in code comparison tasks?"
by: Anonymous
Published: (2024)
by: Anonymous
Published: (2024)
Incorporating convergent enzyme thermal evolution into species distribution models unveils future global marine fish distribution under climate change
by: Anonymous, Anonymous, et al.
Published: (2025)
by: Anonymous, Anonymous, et al.
Published: (2025)
Dataset for "What loss functions do humans optimize when they perform regression and classification"
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Apêndices - Diagnóstico da Maturidade em Gestão do Conhecimento: Um Estudo em uma Empresa Desenvolvedora de Software de Pequeno Porte
by: Anonymous
Published: (2026)
by: Anonymous
Published: (2026)
Appendix for paper "Is RAG All You Need for Effective Real-World Vulnerability Detection with LLMs?"
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Qualitative Content Analysis on Seminar Paper: What Fosters the Acceptance of a Campus App? – Insights into German Student's Perspective
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Replication Package of "Mind the SBOM Gap: Adoption and Compliance in Open Source Software"
by: Anonymous, Author
Published: (2025)
by: Anonymous, Author
Published: (2025)
Can LLMs Replace Human Evaluators? An Empirical Study of LLM-as-a-Judge in Software Engineering
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
AgentBelt: Deterministic Runtime Guardrails for Execution-Level Security in LLM Agents
by: Anonymous
Published: (2026)
by: Anonymous
Published: (2026)
Replication Package for "Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLMs Cost-efficiency "
by: Anonymous
Published: (2026)
by: Anonymous
Published: (2026)
AgentBelt: Runtime Guardrails for LLM Agent Tool Calls — ASE 2026 Artifact
by: Anonymous
Published: (2026)
by: Anonymous
Published: (2026)
When Faster Isn't Greener: The Hidden Costs of LLM-Based Code Optimization - Replication Package
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
RamFuzz: LLM-Guided Greybox Fuzzing for Spatial Memory Corruption via Valid Range Violation
by: Anonymous
Published: (2026)
by: Anonymous
Published: (2026)
LiveBench benchmark ordered by reasoning - Snapshot - November 20, 2025
by: Anonymous, et al.
Published: (2025)
by: Anonymous, et al.
Published: (2025)
Screen2AX-Tree_old
by: Anonymous, et al.
Published: (2025)
by: Anonymous, et al.
Published: (2025)
BLIP for icon captioning
by: Anonymous, et al.
Published: (2025)
by: Anonymous, et al.
Published: (2025)
Artifact: GraphDroid: Asynchronous LLM-Based Mobile App GUI Testing via History-Aware Exploration and Hybrid Intent Fulfillment
by: Anonymous
Published: (2026)
by: Anonymous
Published: (2026)
Ethereum Transaction Datasets for Training and Evaluation LLMs, ML, and DL models.
by: Anonymous Author(s).
Published: (2025)
by: Anonymous Author(s).
Published: (2025)
An Empirical Study on Usage and Perceptions of LLMs in a Software Engineering Project
by: Rasnayaka, Sanka, et al.
Published: (2024)
by: Rasnayaka, Sanka, et al.
Published: (2024)
A Dual Case Study Evaluation of LINDDUN (Supporting Materials)
by: Anonymous
Published: (2025)
by: Anonymous
Published: (2025)
Similar Items
-
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
by: Anonymous, Anonymous
Published: (2024) -
LLMs are Imperfect, Then What? An Empirical Study on LLM Failures in Software Engineering
by: Anonymous, Anonymous
Published: (2024) -
Replication kit for: Can LLMs Make Software Testing Greener? An Empirical Study on JUnit Test Energy Reengineering
by: Anonymous
Published: (2025) -
Supplementary data for Perspectives, Needs and Challenges for Sustainable Software Engineering Teams: A FinServ Case Study
by: Anonymous, Anonymous
Published: (2025) -
Connecting Generations Through Code: GABI, A Community-Driven Framework for Engineering Inclusive Financial Software for the Elderly
by: Anonymous, Anonymous
Published: (2025)