Prompt Optimization Is a Coin Flip: Diagnosing When It Helps in Compound AI Systems
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xing, Wang, Guanghui, Cui, Yanwei, Qiu, Wei, Li, Ziyuan, Zhu, Bing, He, Peiyang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Guardrails Beat Guidance: A Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Hindsight Preference Optimization for Financial Time Series Advisory
by: Cui, Yanwei, et al.
Published: (2026)
by: Cui, Yanwei, et al.
Published: (2026)
Verified Multi-Agent Orchestration: A Plan-Execute-Verify-Replan Framework for Complex Query Resolution
by: Zhang, Xing, et al.
Published: (2026)
by: Zhang, Xing, et al.
Published: (2026)
Enough Coin Flips Can Make LLMs Act Bayesian
by: Gupta, Ritwik, et al.
Published: (2025)
by: Gupta, Ritwik, et al.
Published: (2025)
A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness
by: Schwinn, Leo, et al.
Published: (2026)
by: Schwinn, Leo, et al.
Published: (2026)
STED and Consistency Scoring: A Framework for Evaluating LLM Structured Output Reliability
by: Wang, Guanghui, et al.
Published: (2025)
by: Wang, Guanghui, et al.
Published: (2025)
DelvePO: Direction-Guided Self-Evolving Framework for Flexible Prompt Optimization
by: Tao, Tao, et al.
Published: (2025)
by: Tao, Tao, et al.
Published: (2025)
Do Physicians Know How to Prompt? The Need for Automatic Prompt Optimization Help in Clinical Note Generation
by: Yao, Zonghai, et al.
Published: (2023)
by: Yao, Zonghai, et al.
Published: (2023)
Meta Prompting for AI Systems
by: Zhang, Yifan, et al.
Published: (2023)
by: Zhang, Yifan, et al.
Published: (2023)
Compound AI Systems Optimization: A Survey of Methods, Challenges, and Future Directions
by: Lee, Yu-Ang, et al.
Published: (2025)
by: Lee, Yu-Ang, et al.
Published: (2025)
Continual Learning Using Only Large Language Model Prompting
by: Qiu, Jiabao, et al.
Published: (2024)
by: Qiu, Jiabao, et al.
Published: (2024)
Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict
by: Chen, Yihang, et al.
Published: (2026)
by: Chen, Yihang, et al.
Published: (2026)
Do LLMs Know When to Flip a Coin? Strategic Randomization through Reasoning and Experience
by: Yang, Lingyu
Published: (2025)
by: Yang, Lingyu
Published: (2025)
Optimizing Model Selection for Compound AI Systems
by: Chen, Lingjiao, et al.
Published: (2025)
by: Chen, Lingjiao, et al.
Published: (2025)
Less Is More: Elevating RAG via Performance-Driven Context Compression
by: Cui, Ziqiang, et al.
Published: (2025)
by: Cui, Ziqiang, et al.
Published: (2025)
FlipGuard: Defending Preference Alignment against Update Regression with Constrained Optimization
by: Zhu, Mingye, et al.
Published: (2024)
by: Zhu, Mingye, et al.
Published: (2024)
Evaluating Social Bias in RAG Systems: When External Context Helps and Reasoning Hurts
by: Parihar, Shweta, et al.
Published: (2026)
by: Parihar, Shweta, et al.
Published: (2026)
Small Language Model Helps Resolve Semantic Ambiguity of LLM Prompt
by: Huang, Zhenzhen, et al.
Published: (2026)
by: Huang, Zhenzhen, et al.
Published: (2026)
Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion Decoding
by: Xiang, Yanzheng, et al.
Published: (2026)
by: Xiang, Yanzheng, et al.
Published: (2026)
When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
by: Zheng, Mingqian, et al.
Published: (2023)
by: Zheng, Mingqian, et al.
Published: (2023)
Fortifying Ethical Boundaries in AI: Advanced Strategies for Enhancing Security in Large Language Models
by: He, Yunhong, et al.
Published: (2024)
by: He, Yunhong, et al.
Published: (2024)
The Other Side of the Coin: Exploring Fairness in Retrieval-Augmented Generation
by: Zhang, Zheng, et al.
Published: (2025)
by: Zhang, Zheng, et al.
Published: (2025)
When Prompt Optimization Becomes Jailbreaking: Adaptive Red-Teaming of Large Language Models
by: Shamsi, Zafir, et al.
Published: (2026)
by: Shamsi, Zafir, et al.
Published: (2026)
Promptomatix: An Automatic Prompt Optimization Framework for Large Language Models
by: Murthy, Rithesh, et al.
Published: (2025)
by: Murthy, Rithesh, et al.
Published: (2025)
ELPO: Ensemble Learning Based Prompt Optimization for Large Language Models
by: Zhang, Qing, et al.
Published: (2025)
by: Zhang, Qing, et al.
Published: (2025)
MARS: Multi-Agent Adaptive Reasoning with Socratic Guidance for Automated Prompt Optimization
by: Zhang, Jian, et al.
Published: (2025)
by: Zhang, Jian, et al.
Published: (2025)
RAGChecker: A Fine-grained Framework for Diagnosing Retrieval-Augmented Generation
by: Ru, Dongyu, et al.
Published: (2024)
by: Ru, Dongyu, et al.
Published: (2024)
RAGViz: Diagnose and Visualize Retrieval-Augmented Generation
by: Wang, Tevin, et al.
Published: (2024)
by: Wang, Tevin, et al.
Published: (2024)
When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models
by: Choi, Dasol, et al.
Published: (2026)
by: Choi, Dasol, et al.
Published: (2026)
Position: The Current AI Conference Model is Unsustainable! Diagnosing the Crisis of Centralized AI Conference
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
The Two Sides of the Coin: Hallucination Generation and Detection with LLMs as Evaluators for LLMs
by: Bui, Anh Thu Maria, et al.
Published: (2024)
by: Bui, Anh Thu Maria, et al.
Published: (2024)
Chimera: Diagnosing Shortcut Learning in Visual-Language Understanding
by: Chi, Ziheng, et al.
Published: (2025)
by: Chi, Ziheng, et al.
Published: (2025)
A Prompt-Based Knowledge Graph Foundation Model for Universal In-Context Reasoning
by: Cui, Yuanning, et al.
Published: (2024)
by: Cui, Yuanning, et al.
Published: (2024)
When Does Language Transfer Help? Sequential Fine-Tuning for Cross-Lingual Euphemism Detection
by: Sammartino, Julia, et al.
Published: (2025)
by: Sammartino, Julia, et al.
Published: (2025)
Prompt Stability Matters: Evaluating and Optimizing Auto-Generated Prompt in General-Purpose Systems
by: Chen, Ke, et al.
Published: (2025)
by: Chen, Ke, et al.
Published: (2025)
PGSO: Prompt-based Generative Sequence Optimization Network for Aspect-based Sentiment Analysis
by: Dong, Hao, et al.
Published: (2024)
by: Dong, Hao, et al.
Published: (2024)
Similar Items
-
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
by: Zhang, Xing, et al.
Published: (2026) -
Guardrails Beat Guidance: A Large-Scale Study of Rules, Skills, and Persistent Configuration for Coding Agents
by: Zhang, Xing, et al.
Published: (2026) -
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
by: Zhang, Xing, et al.
Published: (2026) -
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
by: Zhang, Xing, et al.
Published: (2026) -
The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs
by: Zhang, Xing, et al.
Published: (2026)