When "Better" Prompts Hurt: Evaluation-Driven Iteration for LLM Applications
Fuente:
arXiv
Saved in:
| Main Author: | Commey, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Agents Improve Semantic Code Search
by: Jain, Sarthak, et al.
Published: (2024)
by: Jain, Sarthak, et al.
Published: (2024)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
by: Zhang, Haonan, et al.
Published: (2025)
by: Zhang, Haonan, et al.
Published: (2025)
Towards AI Evaluation in Domain-Specific RAG Systems: The AgriHubi Case Study
by: Hasan, Md. Toufique, et al.
Published: (2026)
by: Hasan, Md. Toufique, et al.
Published: (2026)
SemLink: A Semantic-Aware Automated Test Oracle for Hyperlink Verification using Siamese Sentence-BERT
by: Yang, Guan-Yan, et al.
Published: (2026)
by: Yang, Guan-Yan, et al.
Published: (2026)
Domain-Specific Retrieval-Augmented Generation Using Vector Stores, Knowledge Graphs, and Tensor Factorization
by: Barron, Ryan C., et al.
Published: (2024)
by: Barron, Ryan C., et al.
Published: (2024)
cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree
by: Zhang, Yilin, et al.
Published: (2025)
by: Zhang, Yilin, et al.
Published: (2025)
Automating Categorization of Scientific Texts with In-Context Learning and Prompt-Chaining in Large Language Models
by: Shahi, Gautam Kishore, et al.
Published: (2026)
by: Shahi, Gautam Kishore, et al.
Published: (2026)
The Invisible Hand of AI Libraries Shaping Open Source Projects and Communities
by: Esposito, Matteo, et al.
Published: (2026)
by: Esposito, Matteo, et al.
Published: (2026)
Automating Database-Native Function Code Synthesis with LLMs
by: Zhou, Wei, et al.
Published: (2026)
by: Zhou, Wei, et al.
Published: (2026)
Iterative Self-Training for Code Generation via Reinforced Re-Ranking
by: Sorokin, Nikita, et al.
Published: (2025)
by: Sorokin, Nikita, et al.
Published: (2025)
Uncertainty and Fairness Awareness in LLM-Based Recommendation Systems
by: Sah, Chandan Kumar, et al.
Published: (2026)
by: Sah, Chandan Kumar, et al.
Published: (2026)
LLM-Cure: LLM-based Competitor User Review Analysis for Feature Enhancement
by: Assi, Maram, et al.
Published: (2024)
by: Assi, Maram, et al.
Published: (2024)
CodeKGC: Code Language Model for Generative Knowledge Graph Construction
by: Bi, Zhen, et al.
Published: (2023)
by: Bi, Zhen, et al.
Published: (2023)
SQuaD: The Software Quality Dataset
by: Robredo, Mikel, et al.
Published: (2025)
by: Robredo, Mikel, et al.
Published: (2025)
ReCode: Updating Code API Knowledge with Reinforcement Learning
by: Wu, Haoze, et al.
Published: (2025)
by: Wu, Haoze, et al.
Published: (2025)
Do Deployment Constraints Make LLMs Hallucinate Citations? An Empirical Study across Four Models and Five Prompting Regimes
by: Zhao, Chen, et al.
Published: (2026)
by: Zhao, Chen, et al.
Published: (2026)
Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets
by: Gurioli, Andrea, et al.
Published: (2026)
by: Gurioli, Andrea, et al.
Published: (2026)
Developing Retrieval Augmented Generation (RAG) based LLM Systems from PDFs: An Experience Report
by: Khan, Ayman Asad, et al.
Published: (2024)
by: Khan, Ayman Asad, et al.
Published: (2024)
Can Code Evaluation Metrics Detect Code Plagiarism?
by: Ebrahim, Fahad, et al.
Published: (2026)
by: Ebrahim, Fahad, et al.
Published: (2026)
Unveiling Bias in Fairness Evaluations of Large Language Models: A Critical Literature Review of Music and Movie Recommendation Systems
by: Sah, Chandan Kumar, et al.
Published: (2024)
by: Sah, Chandan Kumar, et al.
Published: (2024)
Toward building next-generation Geocoding systems: a systematic review
by: Yin, Zhengcong, et al.
Published: (2025)
by: Yin, Zhengcong, et al.
Published: (2025)
Large Language Models to Enhance Business Process Modeling: Past, Present, and Future Trends
by: Bettencourt, João, et al.
Published: (2026)
by: Bettencourt, João, et al.
Published: (2026)
Taxonomy of the Retrieval System Framework: Pitfalls and Paradigms
by: Shah, Deep, et al.
Published: (2026)
by: Shah, Deep, et al.
Published: (2026)
Context-Augmented Code Generation Using Programming Knowledge Graphs
by: Saberi, Iman, et al.
Published: (2024)
by: Saberi, Iman, et al.
Published: (2024)
Repairing Language Model Pipelines by Meta Self-Refining Competing Constraints at Runtime
by: Eshghie, Mojtaba
Published: (2025)
by: Eshghie, Mojtaba
Published: (2025)
Formal Concept Analysis: a Structural Framework for Variability Extraction and Analysis
by: Galasso, Jessie
Published: (2025)
by: Galasso, Jessie
Published: (2025)
SweRank: Software Issue Localization with Code Ranking
by: Reddy, Revanth Gangi, et al.
Published: (2025)
by: Reddy, Revanth Gangi, et al.
Published: (2025)
Language Model Powered Digital Biology with BRAD
by: Pickard, Joshua, et al.
Published: (2024)
by: Pickard, Joshua, et al.
Published: (2024)
Dynamic ReAct: Scalable Tool Selection for Large-Scale MCP Environments
by: Gaurav, Nishant, et al.
Published: (2025)
by: Gaurav, Nishant, et al.
Published: (2025)
CPRet: A Dataset, Benchmark, and Model for Retrieval in Competitive Programming
by: Deng, Han, et al.
Published: (2025)
by: Deng, Han, et al.
Published: (2025)
Examining the Usage of Generative AI Models in Student Learning Activities for Software Programming
by: Chen, Rufeng, et al.
Published: (2025)
by: Chen, Rufeng, et al.
Published: (2025)
Reformulate, Retrieve, Localize: Agents for Repository-Level Bug Localization
by: Caumartin, Genevieve, et al.
Published: (2025)
by: Caumartin, Genevieve, et al.
Published: (2025)
MRWeb: An Exploration of Generating Multi-Page Resource-Aware Web Code from UI Designs
by: Wan, Yuxuan, et al.
Published: (2024)
by: Wan, Yuxuan, et al.
Published: (2024)
Enhancing Open-Domain Task-Solving Capability of LLMs via Autonomous Tool Integration from GitHub
by: Lyu, Bohan, et al.
Published: (2023)
by: Lyu, Bohan, et al.
Published: (2023)
DeepCodeSeek: Real-Time API Retrieval for Context-Aware Code Generation
by: Esakkiraja, Esakkivel, et al.
Published: (2025)
by: Esakkiraja, Esakkivel, et al.
Published: (2025)
TreeRanker: Fast and Model-agnostic Ranking System for Code Suggestions in IDEs
by: Cipollone, Daniele, et al.
Published: (2025)
by: Cipollone, Daniele, et al.
Published: (2025)
Issue Localization via LLM-Driven Iterative Code Graph Searching
by: Jiang, Zhonghao, et al.
Published: (2025)
by: Jiang, Zhonghao, et al.
Published: (2025)
Evidence Absence Is Not Evidence Insufficiency: Diagnosing NEI Construction Artifacts in Fact Verification
by: Qiu, Jingxi, et al.
Published: (2026)
by: Qiu, Jingxi, et al.
Published: (2026)
Assessing the Ability of ChatGPT to Screen Articles for Systematic Reviews
by: Syriani, Eugene, et al.
Published: (2023)
by: Syriani, Eugene, et al.
Published: (2023)
CodeRAG: Finding Relevant and Necessary Knowledge for Retrieval-Augmented Repository-Level Code Completion
by: Zhang, Sheng, et al.
Published: (2025)
by: Zhang, Sheng, et al.
Published: (2025)
Similar Items
-
LLM Agents Improve Semantic Code Search
by: Jain, Sarthak, et al.
Published: (2024) -
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
by: Zhang, Haonan, et al.
Published: (2025) -
Towards AI Evaluation in Domain-Specific RAG Systems: The AgriHubi Case Study
by: Hasan, Md. Toufique, et al.
Published: (2026) -
SemLink: A Semantic-Aware Automated Test Oracle for Hyperlink Verification using Siamese Sentence-BERT
by: Yang, Guan-Yan, et al.
Published: (2026) -
Domain-Specific Retrieval-Augmented Generation Using Vector Stores, Knowledge Graphs, and Tensor Factorization
by: Barron, Ryan C., et al.
Published: (2024)