The Context Gathering Decision Process: A POMDP Framework for Agentic Search
Fuente:
arXiv
Saved in:
| Main Authors: | Kausik, Chinmaya, Swaminathan, Adith, Kallus, Nathan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Theoretical Framework for Partially Observed Reward-States in RLHF
by: Kausik, Chinmaya, et al.
Published: (2024)
by: Kausik, Chinmaya, et al.
Published: (2024)
Leveraging Offline Data in Linear Latent Contextual Bandits
by: Kausik, Chinmaya, et al.
Published: (2024)
by: Kausik, Chinmaya, et al.
Published: (2024)
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
by: Cheng, Ching-An, et al.
Published: (2024)
by: Cheng, Ching-An, et al.
Published: (2024)
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
by: Bennett, Andrew, et al.
Published: (2024)
by: Bennett, Andrew, et al.
Published: (2024)
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
by: Kallus, Nathan
Published: (2025)
by: Kallus, Nathan
Published: (2025)
Low-Rank MDPs with Continuous Action Spaces
by: Bennett, Andrew, et al.
Published: (2023)
by: Bennett, Andrew, et al.
Published: (2023)
Reasoning about Reasoning: BAPO Bounds on Chain-of-Thought Token Complexity in LLMs
by: Tomlinson, Kiran, et al.
Published: (2026)
by: Tomlinson, Kiran, et al.
Published: (2026)
On Overcoming Miscalibrated Conversational Priors in LLM-based Chatbots
by: Herlihy, Christine, et al.
Published: (2024)
by: Herlihy, Christine, et al.
Published: (2024)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
by: Schnabel, Tobias, et al.
Published: (2025)
by: Schnabel, Tobias, et al.
Published: (2025)
Combining Open-box Simulation and Importance Sampling for Tuning Large-Scale Recommenders
by: Paneri, Kaushal, et al.
Published: (2024)
by: Paneri, Kaushal, et al.
Published: (2024)
Adjusting Regression Models for Conditional Uncertainty Calibration
by: Gao, Ruijiang, et al.
Published: (2024)
by: Gao, Ruijiang, et al.
Published: (2024)
Rao-Blackwellized POMDP Planning
by: Lee, Jiho, et al.
Published: (2024)
by: Lee, Jiho, et al.
Published: (2024)
Deep Belief Markov Models for POMDP Inference
by: Arcieri, Giacomo, et al.
Published: (2025)
by: Arcieri, Giacomo, et al.
Published: (2025)
DiFFPO: Training Diffusion LLMs to Reason Fast and Furious via Reinforcement Learning
by: Zhao, Hanyang, et al.
Published: (2025)
by: Zhao, Hanyang, et al.
Published: (2025)
Value-Guided Search for Efficient Chain-of-Thought Reasoning
by: Wang, Kaiwen, et al.
Published: (2025)
by: Wang, Kaiwen, et al.
Published: (2025)
Explanation through Reward Model Reconciliation using POMDP Tree Search
by: Kraske, Benjamin D., et al.
Published: (2023)
by: Kraske, Benjamin D., et al.
Published: (2023)
Expert-Guided POMDP Learning for Data-Efficient Modeling in Healthcare
by: Locatelli, Marco, et al.
Published: (2025)
by: Locatelli, Marco, et al.
Published: (2025)
GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support
by: Sheikh, Muhammad Umer, et al.
Published: (2026)
by: Sheikh, Muhammad Umer, et al.
Published: (2026)
Transformer-Gather, Fuzzy-Reconsider: A Scalable Hybrid Framework for Entity Resolution
by: Sharifi, Mohammadreza, et al.
Published: (2025)
by: Sharifi, Mohammadreza, et al.
Published: (2025)
Anytime Incremental $ρ$POMDP Planning in Continuous Spaces
by: Benchetrit, Ron, et al.
Published: (2025)
by: Benchetrit, Ron, et al.
Published: (2025)
Deep Reinforcement Multi-agent Learning framework for Information Gathering with Local Gaussian Processes for Water Monitoring
by: Luis, Samuel Yanes, et al.
Published: (2024)
by: Luis, Samuel Yanes, et al.
Published: (2024)
Learning Optimal Defender Strategies for CAGE-2 using a POMDP Model
by: Le, Duc Huy, et al.
Published: (2025)
by: Le, Duc Huy, et al.
Published: (2025)
SOAP-RL: Sequential Option Advantage Propagation for Reinforcement Learning in POMDP Environments
by: Ishida, Shu, et al.
Published: (2024)
by: Ishida, Shu, et al.
Published: (2024)
On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating
by: Widener, Michael, et al.
Published: (2025)
by: Widener, Michael, et al.
Published: (2025)
Smooth Non-Stationary Bandits
by: Jia, Su, et al.
Published: (2023)
by: Jia, Su, et al.
Published: (2023)
Optimal Decision Tree Policies for Markov Decision Processes
by: Vos, Daniël, et al.
Published: (2023)
by: Vos, Daniël, et al.
Published: (2023)
Learning Explainable and Better Performing Representations of POMDP Strategies
by: Bork, Alexander, et al.
Published: (2024)
by: Bork, Alexander, et al.
Published: (2024)
Double Descent and Overfitting under Noisy Inputs and Distribution Shift for Linear Denoisers
by: Kausik, Chinmaya, et al.
Published: (2023)
by: Kausik, Chinmaya, et al.
Published: (2023)
Understanding the Challenges in Iterative Generative Optimization with LLMs
by: Nie, Allen, et al.
Published: (2026)
by: Nie, Allen, et al.
Published: (2026)
Markov Decision Processes under External Temporal Processes
by: Ayyagari, Ranga Shaarad, et al.
Published: (2023)
by: Ayyagari, Ranga Shaarad, et al.
Published: (2023)
Levin Tree Search with Context Models
by: Orseau, Laurent, et al.
Published: (2023)
by: Orseau, Laurent, et al.
Published: (2023)
STX-Search: Explanation Search for Continuous Dynamic Spatio-Temporal Models
by: Anwar, Saif, et al.
Published: (2025)
by: Anwar, Saif, et al.
Published: (2025)
Scaling Agentic Capabilities, Not Context: Efficient Reinforcement Finetuning for Large Toolspaces
by: Gupta, Karan, et al.
Published: (2026)
by: Gupta, Karan, et al.
Published: (2026)
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
by: Bick, Aviv, et al.
Published: (2025)
by: Bick, Aviv, et al.
Published: (2025)
SPAARS: Safer RL Policy Alignment through Abstract Exploration and Refined Exploitation of Action Space
by: K, Swaminathan S, et al.
Published: (2026)
by: K, Swaminathan S, et al.
Published: (2026)
In Search of Trees: Decision-Tree Policy Synthesis for Black-Box Systems via Search
by: Demirović, Emir, et al.
Published: (2024)
by: Demirović, Emir, et al.
Published: (2024)
Context, Reasoning, and Hierarchy: A Cost-Performance Study of Compound LLM Agent Design in an Adversarial POMDP
by: Bogdanov, Igor, et al.
Published: (2026)
by: Bogdanov, Igor, et al.
Published: (2026)
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
by: Jin, Jiahe, et al.
Published: (2025)
by: Jin, Jiahe, et al.
Published: (2025)
Hierarchical Object-Oriented POMDP Planning for Object Rearrangement
by: Mangannavar, Rajesh, et al.
Published: (2024)
by: Mangannavar, Rajesh, et al.
Published: (2024)
Policy Gradient for Robust Markov Decision Processes
by: Wang, Qiuhao, et al.
Published: (2024)
by: Wang, Qiuhao, et al.
Published: (2024)
Similar Items
-
A Theoretical Framework for Partially Observed Reward-States in RLHF
by: Kausik, Chinmaya, et al.
Published: (2024) -
Leveraging Offline Data in Linear Latent Contextual Bandits
by: Kausik, Chinmaya, et al.
Published: (2024) -
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
by: Cheng, Ching-An, et al.
Published: (2024) -
Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision Processes
by: Bennett, Andrew, et al.
Published: (2024) -
Semiparametric Preference Optimization: Your Language Model is Secretly a Single-Index Model
by: Kallus, Nathan
Published: (2025)