Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
Fuente:
arXiv
Saved in:
| Main Authors: | Opsahl-Ong, Krista, Ryan, Michael J, Purtell, Josh, Broman, David, Potts, Christopher, Zaharia, Matei, Khattab, Omar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
by: Saad-Falcon, Jon, et al.
Published: (2023)
by: Saad-Falcon, Jon, et al.
Published: (2023)
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
by: Singhvi, Arnav, et al.
Published: (2023)
by: Singhvi, Arnav, et al.
Published: (2023)
LangProBe: a Language Programs Benchmark
by: Tan, Shangyin, et al.
Published: (2025)
by: Tan, Shangyin, et al.
Published: (2025)
Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together
by: Soylu, Dilara, et al.
Published: (2024)
by: Soylu, Dilara, et al.
Published: (2024)
Composing Policy Gradients and Prompt Optimization for Language Model Programs
by: Ziems, Noah, et al.
Published: (2025)
by: Ziems, Noah, et al.
Published: (2025)
OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
by: Opsahl-Ong, Krista, et al.
Published: (2026)
by: Opsahl-Ong, Krista, et al.
Published: (2026)
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
by: Agrawal, Lakshya A, et al.
Published: (2025)
by: Agrawal, Lakshya A, et al.
Published: (2025)
WARP: An Efficient Engine for Multi-Vector Retrieval
by: Scheerer, Jan Luca, et al.
Published: (2025)
by: Scheerer, Jan Luca, et al.
Published: (2025)
In-Context Learning for Extreme Multi-Label Classification
by: D'Oosterlinck, Karel, et al.
Published: (2024)
by: D'Oosterlinck, Karel, et al.
Published: (2024)
Drowning in Documents: Consequences of Scaling Reranker Inference
by: Jacob, Mathew, et al.
Published: (2024)
by: Jacob, Mathew, et al.
Published: (2024)
Recursive Language Models
by: Zhang, Alex L., et al.
Published: (2025)
by: Zhang, Alex L., et al.
Published: (2025)
A State-of-the-Art SQL Reasoning Model using RLVR
by: Ali, Alnur, et al.
Published: (2025)
by: Ali, Alnur, et al.
Published: (2025)
RAG over Thinking Traces Can Improve Reasoning Tasks
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
by: Kolasani, Sai, et al.
Published: (2025)
by: Kolasani, Sai, et al.
Published: (2025)
RAFT: Adapting Language Model to Domain Specific RAG
by: Zhang, Tianjun, et al.
Published: (2024)
by: Zhang, Tianjun, et al.
Published: (2024)
Reasoning Models Can Be Effective Without Thinking
by: Ma, Wenjie, et al.
Published: (2025)
by: Ma, Wenjie, et al.
Published: (2025)
Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP
by: Rozner, Josh, et al.
Published: (2021)
by: Rozner, Josh, et al.
Published: (2021)
Fact or Fiction? Improving Fact Verification with Knowledge Graphs through Simplified Subgraph Retrievals
by: Opsahl, Tobias A.
Published: (2024)
by: Opsahl, Tobias A.
Published: (2024)
ScenicNL: Generating Probabilistic Scenario Programs from Natural Language
by: Elmaaroufi, Karim, et al.
Published: (2024)
by: Elmaaroufi, Karim, et al.
Published: (2024)
Optimizing Model Selection for Compound AI Systems
by: Chen, Lingjiao, et al.
Published: (2025)
by: Chen, Lingjiao, et al.
Published: (2025)
$L^*LM$: Learning Automata from Examples using Natural Language Oracles
by: Vazquez-Chanlatte, Marcell, et al.
Published: (2024)
by: Vazquez-Chanlatte, Marcell, et al.
Published: (2024)
Long Context RAG Performance of Large Language Models
by: Leng, Quinn, et al.
Published: (2024)
by: Leng, Quinn, et al.
Published: (2024)
Automating the Enterprise with Foundation Models
by: Wornow, Michael, et al.
Published: (2024)
by: Wornow, Michael, et al.
Published: (2024)
Semantic Operators: A Declarative Model for Rich, AI-based Data Processing
by: Patel, Liana, et al.
Published: (2024)
by: Patel, Liana, et al.
Published: (2024)
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation
by: Zhu, Alan, et al.
Published: (2025)
by: Zhu, Alan, et al.
Published: (2025)
Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models
by: Shao, Yijia, et al.
Published: (2024)
by: Shao, Yijia, et al.
Published: (2024)
Building Efficient and Effective OpenQA Systems for Low-Resource Languages
by: Budur, Emrah, et al.
Published: (2024)
by: Budur, Emrah, et al.
Published: (2024)
Large Language Models Are Self-Taught Reasoners: Enhancing LLM Applications via Tailored Problem-Solving Demonstrations
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
by: Ong, Kai Tzu-iunn, et al.
Published: (2024)
SIEVE: Sample-Efficient Parametric Learning from Natural Language
by: Asawa, Parth, et al.
Published: (2026)
by: Asawa, Parth, et al.
Published: (2026)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
by: Ma, Wenjie, et al.
Published: (2025)
by: Ma, Wenjie, et al.
Published: (2025)
DeepScholar-Bench: A Live Benchmark and Automated Evaluation for Generative Research Synthesis
by: Patel, Liana, et al.
Published: (2025)
by: Patel, Liana, et al.
Published: (2025)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
by: Asawa, Parth, et al.
Published: (2025)
by: Asawa, Parth, et al.
Published: (2025)
ColBERT-serve: Efficient Multi-Stage Memory-Mapped Scoring
by: Huang, Kaili, et al.
Published: (2025)
by: Huang, Kaili, et al.
Published: (2025)
Real-Time Probabilistic Programming
by: Hummelgren, Lars, et al.
Published: (2023)
by: Hummelgren, Lars, et al.
Published: (2023)
The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More
by: Chen, Lingjiao, et al.
Published: (2026)
by: Chen, Lingjiao, et al.
Published: (2026)
Can QPP Choose the Right Query Variant? Evaluating Query Variant Selection for RAG Pipelines
by: Arabzadeh, Negar, et al.
Published: (2026)
by: Arabzadeh, Negar, et al.
Published: (2026)
Instruction Agent: Enhancing Agent with Expert Demonstration
by: Li, Yinheng, et al.
Published: (2025)
by: Li, Yinheng, et al.
Published: (2025)
PEEK: Context Map as an Orientation Cache for Long-Context LLM Agents
by: Gu, Zhuohan, et al.
Published: (2026)
by: Gu, Zhuohan, et al.
Published: (2026)
I am a Strange Dataset: Metalinguistic Tests for Language Models
by: Thrush, Tristan, et al.
Published: (2024)
by: Thrush, Tristan, et al.
Published: (2024)
DS SERVE: A Framework for Efficient and Scalable Neural Retrieval
by: Liu, Jinjian, et al.
Published: (2025)
by: Liu, Jinjian, et al.
Published: (2025)
Similar Items
-
ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems
by: Saad-Falcon, Jon, et al.
Published: (2023) -
DSPy Assertions: Computational Constraints for Self-Refining Language Model Pipelines
by: Singhvi, Arnav, et al.
Published: (2023) -
LangProBe: a Language Programs Benchmark
by: Tan, Shangyin, et al.
Published: (2025) -
Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together
by: Soylu, Dilara, et al.
Published: (2024) -
Composing Policy Gradients and Prompt Optimization for Language Model Programs
by: Ziems, Noah, et al.
Published: (2025)