When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
Fuente:
arXiv
Salvato in:
| Autori principali: | Larbi, Maya, Akli, Amal, Papadakis, Mike, Bouyousfi, Rihab, Cordy, Maxime, Sarro, Federica, Traon, Yves Le |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
di: Akli, Amal, et al.
Pubblicazione: (2026)
di: Akli, Amal, et al.
Pubblicazione: (2026)
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
di: Akli, Amal, et al.
Pubblicazione: (2026)
di: Akli, Amal, et al.
Pubblicazione: (2026)
Software Fairness: An Analysis and Survey
di: Soremekun, Ezekiel, et al.
Pubblicazione: (2022)
di: Soremekun, Ezekiel, et al.
Pubblicazione: (2022)
Going Further: Flatness at the Rescue of Early Stopping for Adversarial Example Transferability
di: Gubri, Martin, et al.
Pubblicazione: (2023)
di: Gubri, Martin, et al.
Pubblicazione: (2023)
GenCode: A Generic Data Augmentation Framework for Boosting Deep Learning-Based Code Understanding
di: Dong, Zeming, et al.
Pubblicazione: (2024)
di: Dong, Zeming, et al.
Pubblicazione: (2024)
Leveraging External Factors in Household-Level Electrical Consumption Forecasting using Hypernetworks
di: Bernier, Fabien, et al.
Pubblicazione: (2025)
di: Bernier, Fabien, et al.
Pubblicazione: (2025)
Boosting Source Code Learning with Text-Oriented Data Augmentation: An Empirical Study
di: Dong, Zeming, et al.
Pubblicazione: (2023)
di: Dong, Zeming, et al.
Pubblicazione: (2023)
When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation
di: AKLI, Amal, et al.
Pubblicazione: (2026)
di: AKLI, Amal, et al.
Pubblicazione: (2026)
On the Effectiveness of Hybrid Pooling in Mixup-Based Graph Learning for Language Processing
di: Dong, Zeming, et al.
Pubblicazione: (2022)
di: Dong, Zeming, et al.
Pubblicazione: (2022)
Automated Repair of Ambiguous Problem Descriptions for LLM-Based Code Generation
di: Jia, Haoxiang, et al.
Pubblicazione: (2025)
di: Jia, Haoxiang, et al.
Pubblicazione: (2025)
Robustness Analysis of AI Models in Critical Energy Systems
di: Dogoulis, Pantelis, et al.
Pubblicazione: (2024)
di: Dogoulis, Pantelis, et al.
Pubblicazione: (2024)
HInter: Exposing Hidden Intersectional Bias in Large Language Models
di: Souani, Badr, et al.
Pubblicazione: (2025)
di: Souani, Badr, et al.
Pubblicazione: (2025)
DNN Testing, Regression Testing and Software Reliability Prediction
di: Yves Le Traon, et al.
Pubblicazione: (2024)
di: Yves Le Traon, et al.
Pubblicazione: (2024)
Fault tolerance and metamorphic relation prediction
di: Yves Le Traon, et al.
Pubblicazione: (2024)
di: Yves Le Traon, et al.
Pubblicazione: (2024)
Test code evolution and mutation testing
di: Yves Le Traon, et al.
Pubblicazione: (2024)
di: Yves Le Traon, et al.
Pubblicazione: (2024)
Metamorphic Testing and Web Element Localization
di: Yves Le Traon, et al.
Pubblicazione: (2024)
di: Yves Le Traon, et al.
Pubblicazione: (2024)
Investigating fault injection techniques in hardware‐based deep neural networks and mutation‐based fault localization
di: Yves Le Traon, et al.
Pubblicazione: (2024)
di: Yves Le Traon, et al.
Pubblicazione: (2024)
On the Robustness of Tabular Foundation Models: Test-Time Attacks and In-Context Defenses
di: Djilani, Mohamed, et al.
Pubblicazione: (2025)
di: Djilani, Mohamed, et al.
Pubblicazione: (2025)
Comparing AI Coding Agents: A Task-Stratified Analysis of Pull Request Acceptance
di: Pinna, Giovanni, et al.
Pubblicazione: (2026)
di: Pinna, Giovanni, et al.
Pubblicazione: (2026)
Inferring Code Correctness from Specification
di: Florian, Tambon, et al.
Pubblicazione: (2026)
di: Florian, Tambon, et al.
Pubblicazione: (2026)
Physics Informed Reinforcement Learning with Gibbs Priors for Topology Control in Power Grids
di: Dogoulis, Pantelis, et al.
Pubblicazione: (2026)
di: Dogoulis, Pantelis, et al.
Pubblicazione: (2026)
Learning Generalizable Multimodal Representations for Software Vulnerability Detection
di: Dong, Zeming, et al.
Pubblicazione: (2026)
di: Dong, Zeming, et al.
Pubblicazione: (2026)
Image Generation from Contextually-Contradictory Prompts
di: Huberman, Saar, et al.
Pubblicazione: (2025)
di: Huberman, Saar, et al.
Pubblicazione: (2025)
What Can Go Wrong During Caplet Stripping ?
di: Floc'h, Fabien Le
Pubblicazione: (2026)
di: Floc'h, Fabien Le
Pubblicazione: (2026)
The Wrong Way to Go.
di: Wright, H. Curtis
Pubblicazione: (1979)
di: Wright, H. Curtis
Pubblicazione: (1979)
Foundation Models for Autonomous Driving System: An Initial Roadmap
di: Wu, Xiongfei, et al.
Pubblicazione: (2025)
di: Wu, Xiongfei, et al.
Pubblicazione: (2025)
Initial-boundary value problem for second order hyperbolic operator with mixed boundary conditions
di: Ait-Akli, Djamel
Pubblicazione: (2024)
di: Ait-Akli, Djamel
Pubblicazione: (2024)
When PINNs Go Wrong: Pseudo-Time Stepping Against Spurious Solutions
di: Wang, Sifan, et al.
Pubblicazione: (2026)
di: Wang, Sifan, et al.
Pubblicazione: (2026)
Why and When LLM-Based Assistants Can Go Wrong: Investigating the Effectiveness of Prompt-Based Interactions for Software Help-Seeking
di: Khurana, Anjali, et al.
Pubblicazione: (2024)
di: Khurana, Anjali, et al.
Pubblicazione: (2024)
On the Compression of Language Models for Code: An Empirical Study on CodeBERT
di: d'Aloisio, Giordano, et al.
Pubblicazione: (2024)
di: d'Aloisio, Giordano, et al.
Pubblicazione: (2024)
Post‐Artesunate Delayed Hemolysis: Anything That Can Go Wrong Will Go Wrong—Murphy's Law
di: Beliza Chemutai, et al.
Pubblicazione: (2024)
di: Beliza Chemutai, et al.
Pubblicazione: (2024)
Contradictory Theology
di: Jc Beall, et al.
Pubblicazione: (2025)
di: Jc Beall, et al.
Pubblicazione: (2025)
When Learning Rates Go Wrong: Early Structural Signals in PPO Actor-Critic
di: Fernández-Hernández, Alberto, et al.
Pubblicazione: (2026)
di: Fernández-Hernández, Alberto, et al.
Pubblicazione: (2026)
TabularBench: Benchmarking Adversarial Robustness for Tabular Deep Learning in Real-world Use-cases
di: Simonetto, Thibault, et al.
Pubblicazione: (2024)
di: Simonetto, Thibault, et al.
Pubblicazione: (2024)
Constrained Adaptive Attack: Effective Adversarial Attack Against Deep Neural Networks for Tabular Data
di: Simonetto, Thibault, et al.
Pubblicazione: (2024)
di: Simonetto, Thibault, et al.
Pubblicazione: (2024)
RobustBlack: Challenging Black-Box Adversarial Attacks on State-of-the-Art Defenses
di: Djilani, Mohamed, et al.
Pubblicazione: (2024)
di: Djilani, Mohamed, et al.
Pubblicazione: (2024)
KCLNet: Physics-Informed Power Flow Prediction via Constraints Projections
di: Dogoulis, Pantelis, et al.
Pubblicazione: (2025)
di: Dogoulis, Pantelis, et al.
Pubblicazione: (2025)
Centralized vs Decentralized Federated Learning: A trade-off performance analysis
di: Medjadji, Chaimaa, et al.
Pubblicazione: (2026)
di: Medjadji, Chaimaa, et al.
Pubblicazione: (2026)
An Empirical Study of the Imbalance Issue in Software Vulnerability Detection
di: Guo, Yuejun, et al.
Pubblicazione: (2026)
di: Guo, Yuejun, et al.
Pubblicazione: (2026)
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
di: Fleisig, Eve, et al.
Pubblicazione: (2023)
di: Fleisig, Eve, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Defective Task Descriptions in LLM-Based Code Generation: Detection and Analysis
di: Akli, Amal, et al.
Pubblicazione: (2026) -
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
di: Akli, Amal, et al.
Pubblicazione: (2026) -
Software Fairness: An Analysis and Survey
di: Soremekun, Ezekiel, et al.
Pubblicazione: (2022) -
Going Further: Flatness at the Rescue of Early Stopping for Adversarial Example Transferability
di: Gubri, Martin, et al.
Pubblicazione: (2023) -
GenCode: A Generic Data Augmentation Framework for Boosting Deep Learning-Based Code Understanding
di: Dong, Zeming, et al.
Pubblicazione: (2024)