Reflective Prompted Policy Optimization: Trajectory-Grounded Revision and Salience Bias
Fuente:
arXiv
Saved in:
| Main Authors: | Hara, Rahaf Abu, Murarri, Vaibbhav, Zito, Claudio |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Modeling Saliency Dataset Bias
by: Kümmerer, Matthias, et al.
Published: (2025)
by: Kümmerer, Matthias, et al.
Published: (2025)
Reflective Policy Optimization
by: Gan, Yaozhong, et al.
Published: (2024)
by: Gan, Yaozhong, et al.
Published: (2024)
Trajectory-Oriented Policy Optimization with Sparse Rewards
by: Wang, Guojian, et al.
Published: (2024)
by: Wang, Guojian, et al.
Published: (2024)
Interactive Prompt Debugging with Sequence Salience
by: Tenney, Ian, et al.
Published: (2024)
by: Tenney, Ian, et al.
Published: (2024)
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
by: Zeris, Athanasios
Published: (2026)
by: Zeris, Athanasios
Published: (2026)
Optimizing Backward Policies in GFlowNets via Trajectory Likelihood Maximization
by: Gritsaev, Timofei, et al.
Published: (2024)
by: Gritsaev, Timofei, et al.
Published: (2024)
From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models
by: Gumbsch, Christian, et al.
Published: (2026)
by: Gumbsch, Christian, et al.
Published: (2026)
Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
by: Zhao, Zihui, et al.
Published: (2025)
by: Zhao, Zihui, et al.
Published: (2025)
When Data Falls Short: Grokking Below the Critical Threshold
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
Dual-Phase Continual Learning: Supervised Adaptation Meets Unsupervised Retention
by: Singh, Vaibhav, et al.
Published: (2024)
by: Singh, Vaibhav, et al.
Published: (2024)
Tempered Sequential Monte Carlo for Trajectory and Policy Optimization with Differentiable Dynamics
by: Yang, Heng
Published: (2026)
by: Yang, Heng
Published: (2026)
P^2O: Joint Policy and Prompt Optimization
by: Lu, Xinyu, et al.
Published: (2026)
by: Lu, Xinyu, et al.
Published: (2026)
Flow-based Policy With Distributional Reinforcement Learning in Trajectory Optimization
by: Hao, Ruijie, et al.
Published: (2026)
by: Hao, Ruijie, et al.
Published: (2026)
SAFE-RL: Saliency-Aware Counterfactual Explainer for Deep Reinforcement Learning Policies
by: Samadi, Amir, et al.
Published: (2024)
by: Samadi, Amir, et al.
Published: (2024)
ReflectivePrompt: Reflective evolution in autoprompting algorithms
by: Zhuravlev, Viktor N., et al.
Published: (2025)
by: Zhuravlev, Viktor N., et al.
Published: (2025)
How Prompts Move Language Model Behavior: Frames, Salience, and Construal as Semantic Control
by: Kim, Dongseok, et al.
Published: (2025)
by: Kim, Dongseok, et al.
Published: (2025)
On the Reuse Bias in Off-Policy Reinforcement Learning
by: Ying, Chengyang, et al.
Published: (2022)
by: Ying, Chengyang, et al.
Published: (2022)
STEP: Success-Rate-Aware Trajectory-Efficient Policy Optimization
by: Chen, Yuhan, et al.
Published: (2025)
by: Chen, Yuhan, et al.
Published: (2025)
Enhancing PPO with Trajectory-Aware Hybrid Policies
by: Liu, Qisai, et al.
Published: (2025)
by: Liu, Qisai, et al.
Published: (2025)
Saliency-Aware Regularized Graph Neural Network
by: Pei, Wenjie, et al.
Published: (2024)
by: Pei, Wenjie, et al.
Published: (2024)
On the Convergence of Experience Replay in Policy Optimization: Characterizing Bias, Variance, and Finite-Time Convergence
by: Zheng, Hua, et al.
Published: (2021)
by: Zheng, Hua, et al.
Published: (2021)
AutoPDL: Automatic Prompt Optimization for LLM Agents
by: Spiess, Claudio, et al.
Published: (2025)
by: Spiess, Claudio, et al.
Published: (2025)
SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization
by: Wang, Zhengcheng, et al.
Published: (2025)
by: Wang, Zhengcheng, et al.
Published: (2025)
Learning an Efficient Optimizer via Hybrid-Policy Sub-Trajectory Balance
by: Guan, Yunchuan, et al.
Published: (2025)
by: Guan, Yunchuan, et al.
Published: (2025)
Reusing Trajectories in Policy Gradients Enables Fast Convergence
by: Montenegro, Alessandro, et al.
Published: (2025)
by: Montenegro, Alessandro, et al.
Published: (2025)
Best Policy Learning from Trajectory Preference Feedback
by: Agnihotri, Akhil, et al.
Published: (2025)
by: Agnihotri, Akhil, et al.
Published: (2025)
Neural at ArchEHR-QA 2025: Agentic Prompt Optimization for Evidence-Grounded Clinical Question Answering
by: Bogireddy, Sai Prasanna Teja Reddy, et al.
Published: (2025)
by: Bogireddy, Sai Prasanna Teja Reddy, et al.
Published: (2025)
Provable Robust Saliency-based Explanations
by: Chen, Chao, et al.
Published: (2022)
by: Chen, Chao, et al.
Published: (2022)
Can We Optimize Deep RL Policy Weights as Trajectory Modeling?
by: Tang, Hongyao
Published: (2025)
by: Tang, Hongyao
Published: (2025)
TMS: Trajectory-Mixed Supervision for Reward-Free, On-Policy SFT
by: Khan, Rana Muhammad Shahroz, et al.
Published: (2026)
by: Khan, Rana Muhammad Shahroz, et al.
Published: (2026)
Truncating Trajectories in Monte Carlo Policy Evaluation: an Adaptive Approach
by: Poiani, Riccardo, et al.
Published: (2024)
by: Poiani, Riccardo, et al.
Published: (2024)
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
by: Suau, Miguel, et al.
Published: (2023)
by: Suau, Miguel, et al.
Published: (2023)
Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning
by: Daley, Brett, et al.
Published: (2023)
by: Daley, Brett, et al.
Published: (2023)
Info-CELS: Informative Saliency Map Guided Counterfactual Explanation
by: Li, Peiyu, et al.
Published: (2024)
by: Li, Peiyu, et al.
Published: (2024)
A Minimalist Prompt for Zero-Shot Policy Learning
by: Song, Meng, et al.
Published: (2024)
by: Song, Meng, et al.
Published: (2024)
Trajectory First: A Curriculum for Discovering Diverse Policies
by: Braun, Cornelius V., et al.
Published: (2025)
by: Braun, Cornelius V., et al.
Published: (2025)
Bias after Prompting: Persistent Discrimination in Large Language Models
by: Sivakumar, Nivedha, et al.
Published: (2025)
by: Sivakumar, Nivedha, et al.
Published: (2025)
Pareto optimal proxy metrics
by: Zito, Alessandro, et al.
Published: (2023)
by: Zito, Alessandro, et al.
Published: (2023)
T2IBias: Uncovering Societal Bias Encoded in the Latent Space of Text-to-Image Generative Models
by: Sufian, Abu, et al.
Published: (2025)
by: Sufian, Abu, et al.
Published: (2025)
Offline Reinforcement Learning with Generative Trajectory Policies
by: Feng, Xinsong, et al.
Published: (2025)
by: Feng, Xinsong, et al.
Published: (2025)
Similar Items
-
Modeling Saliency Dataset Bias
by: Kümmerer, Matthias, et al.
Published: (2025) -
Reflective Policy Optimization
by: Gan, Yaozhong, et al.
Published: (2024) -
Trajectory-Oriented Policy Optimization with Sparse Rewards
by: Wang, Guojian, et al.
Published: (2024) -
Interactive Prompt Debugging with Sequence Salience
by: Tenney, Ian, et al.
Published: (2024) -
Energy-Gated Attention: Spectral Salience as an Inductive Bias for Transformer Attention
by: Zeris, Athanasios
Published: (2026)