Agentic Planning with Reasoning for Image Styling via Offline RL
Fuente:
arXiv
Saved in:
| Main Authors: | Mukherjee, Subhojyoti, Petrangeli, Stefano, Kveton, Branislav, Bui, Trung, Dernoncourt, Franck, Mukherjee, Arko |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Efficient and Interpretable Bandit Algorithms
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
by: Mukherjee, Subhojyoti, et al.
Published: (2025)
by: Mukherjee, Subhojyoti, et al.
Published: (2025)
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026)
by: Mathur, Puneet, et al.
Published: (2026)
Off-Policy Evaluation from Logged Human Feedback
by: Bhargava, Aniruddha, et al.
Published: (2024)
by: Bhargava, Aniruddha, et al.
Published: (2024)
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
AdvantageFlow: Advantage-Weighted Least Squares for RL in Flow Models
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
Learning to Reason in LLMs by Expectation Maximization
by: Lee, Junghyun, et al.
Published: (2025)
by: Lee, Junghyun, et al.
Published: (2025)
Experimental Design for Active Transductive Inference in Large Language Models
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
Optimal Design for Human Preference Elicitation
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
by: Lee, Daeun, et al.
Published: (2025)
by: Lee, Daeun, et al.
Published: (2025)
Selective Uncertainty Propagation in Offline RL
by: Krishnamurthy, Sanath Kumar, et al.
Published: (2023)
by: Krishnamurthy, Sanath Kumar, et al.
Published: (2023)
Stepwise Credit Assignment for GRPO on Flow-Matching Models
by: Savani, Yash, et al.
Published: (2026)
by: Savani, Yash, et al.
Published: (2026)
Learning from a single labeled face and a stream of unlabeled data
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
SaVeR: Optimal Data Collection Strategy for Safe Policy Evaluation in Tabular MDP
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
SPEED: Experimental Design for Policy Evaluation in Linear Heteroscedastic Bandits
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
by: Mukherjee, Subhojyoti, et al.
Published: (2023)
LLM-as-Judge on a Budget
by: Saha, Aadirupa, et al.
Published: (2026)
by: Saha, Aadirupa, et al.
Published: (2026)
Cross-Validated Off-Policy Evaluation
by: Cief, Matej, et al.
Published: (2024)
by: Cief, Matej, et al.
Published: (2024)
Pretraining Decision Transformers with Reward Prediction for In-Context Multi-task Structured Bandit Learning
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
RADAR: Reasoning-Ability and Difficulty-Aware Routing for Reasoning LLMs
by: Fernandez, Nigel, et al.
Published: (2025)
by: Fernandez, Nigel, et al.
Published: (2025)
ML-Tool-Bench: Tool-Augmented Planning for ML Tasks
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
Steering MoE LLMs via Expert (De)Activation
by: Fayyaz, Mohsen, et al.
Published: (2025)
by: Fayyaz, Mohsen, et al.
Published: (2025)
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
by: Tanjim, Md Mehrab, et al.
Published: (2026)
by: Tanjim, Md Mehrab, et al.
Published: (2026)
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)
by: Hiranandani, Gaurush, et al.
Published: (2025)
Pessimistic Off-Policy Optimization for Learning to Rank
by: Cief, Matej, et al.
Published: (2022)
by: Cief, Matej, et al.
Published: (2022)
Semi-supervised learning with max-margin graph cuts
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
Spectral bandits for smooth graph functions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Online semi-supervised perception: Real-time learning without explicit feedback
by: Kveton, Branislav, et al.
Published: (2026)
by: Kveton, Branislav, et al.
Published: (2026)
Fine-tuning CLIP Text Encoders with Two-step Paraphrasing
by: Kim, Hyunjae, et al.
Published: (2024)
by: Kim, Hyunjae, et al.
Published: (2024)
Spectral bandits for smooth graph functions with applications in recommender systems
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Evidence-based anomaly detection in clinical domains
by: Hauskrecht, Milos, et al.
Published: (2026)
by: Hauskrecht, Milos, et al.
Published: (2026)
Online Posterior Sampling with a Diffusion Prior
by: Kveton, Branislav, et al.
Published: (2024)
by: Kveton, Branislav, et al.
Published: (2024)
Finite-Time Logarithmic Bayes Regret Upper Bounds
by: Atsidakou, Alexia, et al.
Published: (2023)
by: Atsidakou, Alexia, et al.
Published: (2023)
RAGEN-2: Reasoning Collapse in Agentic RL
by: Wang, Zihan, et al.
Published: (2026)
by: Wang, Zihan, et al.
Published: (2026)
Reconstructing the Hubble parameter with future Gravitational Wave missions using Machine Learning
by: Mukherjee, Purba, et al.
Published: (2023)
by: Mukherjee, Purba, et al.
Published: (2023)
Conditional anomaly detection with soft harmonic functions
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Conditional anomaly detection using soft harmonic functions: An application to clinical alerting
by: Valko, Michal, et al.
Published: (2026)
by: Valko, Michal, et al.
Published: (2026)
Quantitative LLM Judges
by: Sahoo, Aishwarya, et al.
Published: (2025)
by: Sahoo, Aishwarya, et al.
Published: (2025)
Spectral bandits
by: Kocák, Tomáš, et al.
Published: (2026)
by: Kocák, Tomáš, et al.
Published: (2026)
Offline RL via Feature-Occupancy Gradient Ascent
by: Neu, Gergely, et al.
Published: (2024)
by: Neu, Gergely, et al.
Published: (2024)
Out-of-Distribution Adaptation in Offline RL: Counterfactual Reasoning via Causal Normalizing Flows
by: Cho, Minjae, et al.
Published: (2024)
by: Cho, Minjae, et al.
Published: (2024)
Similar Items
-
Efficient and Interpretable Bandit Algorithms
by: Mukherjee, Subhojyoti, et al.
Published: (2023) -
Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization
by: Mukherjee, Subhojyoti, et al.
Published: (2025) -
Partial Policy Gradients for RL in LLMs
by: Mathur, Puneet, et al.
Published: (2026) -
Off-Policy Evaluation from Logged Human Feedback
by: Bhargava, Aniruddha, et al.
Published: (2024) -
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
by: Mukherjee, Subhojyoti, et al.
Published: (2024)