What Is Wrong with My Model? Identifying Systematic Problems with Semantic Data Slicing
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Chenyang, Hong, Yining, Lewis, Grace A., Wu, Tongshuang, Kästner, Christian |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025)
by: Yang, Chenyang, et al.
Published: (2025)
(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
by: Ma, Wanqin, et al.
Published: (2023)
by: Ma, Wanqin, et al.
Published: (2023)
From Hazard Identification to Controller Design: Proactive and LLM-Supported Safety Engineering for ML-Powered Systems
by: Hong, Yining, et al.
Published: (2025)
by: Hong, Yining, et al.
Published: (2025)
cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
by: Zhang, Yilin, et al.
Published: (2025)
by: Zhang, Yilin, et al.
Published: (2025)
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
Static Program Slicing Using Language Models With Dataflow-Aware Pretraining and Constrained Decoding
by: He, Pengfei, et al.
Published: (2026)
by: He, Pengfei, et al.
Published: (2026)
SemCoder: Training Code Language Models with Comprehensive Semantics Reasoning
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
DebugBench: Evaluating Debugging Capability of Large Language Models
by: Tian, Runchu, et al.
Published: (2024)
by: Tian, Runchu, et al.
Published: (2024)
Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Grammar-Constrained Refinement of Safety Operational Rules Using Language in the Loop: What Could Go Wrong
by: Gaaloul, Khouloud, et al.
Published: (2026)
by: Gaaloul, Khouloud, et al.
Published: (2026)
PPM: Automated Generation of Diverse Programming Problems for Benchmarking Code Generation Models
by: Chen, Simin, et al.
Published: (2024)
by: Chen, Simin, et al.
Published: (2024)
REINFOREST: Reinforcing Semantic Code Similarity for Cross-Lingual Code Search Models
by: Saieva, Anthony, et al.
Published: (2023)
by: Saieva, Anthony, et al.
Published: (2023)
Solving Data-centric Tasks using Large Language Models
by: Barke, Shraddha, et al.
Published: (2024)
by: Barke, Shraddha, et al.
Published: (2024)
How Diversely Can Language Models Solve Problems? Exploring the Algorithmic Diversity of Model-Generated Code
by: Lee, Seonghyeon, et al.
Published: (2025)
by: Lee, Seonghyeon, et al.
Published: (2025)
LLMON: An LLM-native Markup Language to Leverage Structure and Semantics at the LLM Interface
by: Hind, Michael, et al.
Published: (2026)
by: Hind, Michael, et al.
Published: (2026)
Evaluation of General Large Language Models in Contextually Assessing Semantic Concepts Extracted from Adult Critical Care Electronic Health Record Notes
by: Liu, Darren, et al.
Published: (2024)
by: Liu, Darren, et al.
Published: (2024)
What's Wrong with Your Code Generated by Large Language Models? An Extensive Study
by: Dou, Shihan, et al.
Published: (2024)
by: Dou, Shihan, et al.
Published: (2024)
A Critical Study of What Code-LLMs (Do Not) Learn
by: Anand, Abhinav, et al.
Published: (2024)
by: Anand, Abhinav, et al.
Published: (2024)
Semantically Aligned Question and Code Generation for Automated Insight Generation
by: Singha, Ananya, et al.
Published: (2024)
by: Singha, Ananya, et al.
Published: (2024)
Your Simulation Runs but Solves the Wrong Physics: PDE-Grounded Intent Verification for LLM-Generated Multiphysics Simulation Code
by: Song, Zhenghan, et al.
Published: (2026)
by: Song, Zhenghan, et al.
Published: (2026)
SWE-smith: Scaling Data for Software Engineering Agents
by: Yang, John, et al.
Published: (2025)
by: Yang, John, et al.
Published: (2025)
Exploring Data-Efficient Adaptation of Large Language Models for Code Generation
by: Jiang, Xue, et al.
Published: (2024)
by: Jiang, Xue, et al.
Published: (2024)
Self-Improving Code Generation via Semantic Entropy and Behavioral Consensus
by: Zhang, Huan, et al.
Published: (2026)
by: Zhang, Huan, et al.
Published: (2026)
Hardness, Structural Knowledge, and Opportunity: An Analytical Framework for Modular Performance Modeling
by: Gheibi, Omid, et al.
Published: (2025)
by: Gheibi, Omid, et al.
Published: (2025)
A Problem-Oriented Perspective and Anchor Verification for Code Optimization
by: Ye, Tong, et al.
Published: (2024)
by: Ye, Tong, et al.
Published: (2024)
Can LLMs Generate Reliable Test Case Generators? A Study on Competition-Level Programming Problems
by: Cao, Yuhan, et al.
Published: (2025)
by: Cao, Yuhan, et al.
Published: (2025)
ABC-Bench: Benchmarking Agentic Backend Coding in Real-World Development
by: Yang, Jie, et al.
Published: (2026)
by: Yang, Jie, et al.
Published: (2026)
CodeRL+: Improving Code Generation via Reinforcement with Execution Semantics Alignment
by: Jiang, Xue, et al.
Published: (2025)
by: Jiang, Xue, et al.
Published: (2025)
IRepair: An Intent-Aware Approach to Repair Data-Driven Errors in Large Language Models
by: Imtiaz, Sayem Mohammad, et al.
Published: (2025)
by: Imtiaz, Sayem Mohammad, et al.
Published: (2025)
Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
by: Chen, Simin, et al.
Published: (2025)
by: Chen, Simin, et al.
Published: (2025)
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety
by: Lee, Seongmin, et al.
Published: (2025)
by: Lee, Seongmin, et al.
Published: (2025)
Groot: Adversarial Testing for Generative Text-to-Image Models with Tree-based Semantic Transformation
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
LLMs Lean on Priors, Not Programming Language Semantics
by: Thimmaiah, Aditya, et al.
Published: (2025)
by: Thimmaiah, Aditya, et al.
Published: (2025)
Granite Code Models: A Family of Open Foundation Models for Code Intelligence
by: Mishra, Mayank, et al.
Published: (2024)
by: Mishra, Mayank, et al.
Published: (2024)
CodeV: Issue Resolving with Visual Data
by: Zhang, Linhao, et al.
Published: (2024)
by: Zhang, Linhao, et al.
Published: (2024)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
by: Larbi, Maya, et al.
Published: (2025)
by: Larbi, Maya, et al.
Published: (2025)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
by: Le-Cong, Thanh, et al.
Published: (2025)
by: Le-Cong, Thanh, et al.
Published: (2025)
What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects
by: Robredo, Mikel, et al.
Published: (2025)
by: Robredo, Mikel, et al.
Published: (2025)
CODESIM: Multi-Agent Code Generation and Problem Solving through Simulation-Driven Planning and Debugging
by: Islam, Md. Ashraful, et al.
Published: (2025)
by: Islam, Md. Ashraful, et al.
Published: (2025)
ZnTrack -- Data as Code
by: Zills, Fabian, et al.
Published: (2024)
by: Zills, Fabian, et al.
Published: (2024)
Similar Items
-
What Prompts Don't Say: Understanding and Managing Underspecification in LLM Prompts
by: Yang, Chenyang, et al.
Published: (2025) -
(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
by: Ma, Wanqin, et al.
Published: (2023) -
From Hazard Identification to Controller Design: Proactive and LLM-Supported Safety Engineering for ML-Powered Systems
by: Hong, Yining, et al.
Published: (2025) -
cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
by: Zhang, Yilin, et al.
Published: (2025) -
Symbolic Guardrails for Domain-Specific Agents: Stronger Safety and Security Guarantees Without Sacrificing Utility
by: Hong, Yining, et al.
Published: (2026)