Pure Exploration Beyond Reward Feedback: The Role of Post-Action Context
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shahverdikondori, Mohammad, Abouei, Amir Mohammad, Rezaeimoghadam, Alireza, Kiyavash, Negar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Active Context Selection Improves Simple Regret in Contextual Bandits
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2026)
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2026)
Graph-Dependent Regret Bounds in Multi-Armed Bandits with Interference
von: Jamshidi, Fateme, et al.
Veröffentlicht: (2025)
von: Jamshidi, Fateme, et al.
Veröffentlicht: (2025)
Graph Learning Is Suboptimal in Causal Bandits
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2025)
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2025)
Best Group Identification in Multi-Objective Bandits
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2025)
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2025)
s-ID: Causal Effect Identification in a Sub-Population
von: Abouei, Amir Mohammad, et al.
Veröffentlicht: (2023)
von: Abouei, Amir Mohammad, et al.
Veröffentlicht: (2023)
QWO: Speeding Up Permutation-Based Causal Discovery in LiGAMs
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2024)
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2024)
Causal Effect Identification in a Sub-Population with Latent Variables
von: Abouei, Amir Mohammad, et al.
Veröffentlicht: (2024)
von: Abouei, Amir Mohammad, et al.
Veröffentlicht: (2024)
Neighborhood-Aware Graph Labeling Problem
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2026)
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2026)
Fusing Rewards and Preferences in Reinforcement Learning
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025)
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025)
Beyond Optimism: Exploration With Partially Observable Rewards
von: Parisi, Simone, et al.
Veröffentlicht: (2024)
von: Parisi, Simone, et al.
Veröffentlicht: (2024)
Sample Complexity of Nonparametric Closeness Testing for Continuous Distributions and Its Application to Causal Discovery with Hidden Confounding
von: Jamshidi, Fateme, et al.
Veröffentlicht: (2025)
von: Jamshidi, Fateme, et al.
Veröffentlicht: (2025)
Confounded Budgeted Causal Bandits
von: Jamshidi, Fateme, et al.
Veröffentlicht: (2024)
von: Jamshidi, Fateme, et al.
Veröffentlicht: (2024)
Pure Exploration with Feedback Graphs
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Pure Exploration under Mediators' Feedback
von: Poiani, Riccardo, et al.
Veröffentlicht: (2023)
von: Poiani, Riccardo, et al.
Veröffentlicht: (2023)
Near-Optimal Experiment Design in Linear non-Gaussian Cyclic Models
von: Sharifian, Ehsan, et al.
Veröffentlicht: (2025)
von: Sharifian, Ehsan, et al.
Veröffentlicht: (2025)
Learning Peer Influence Probabilities with Linear Contextual Bandits
von: Faruk, Ahmed Sayeed, et al.
Veröffentlicht: (2025)
von: Faruk, Ahmed Sayeed, et al.
Veröffentlicht: (2025)
Multi-Domain Causal Discovery in Bijective Causal Models
von: Jalaldoust, Kasra, et al.
Veröffentlicht: (2025)
von: Jalaldoust, Kasra, et al.
Veröffentlicht: (2025)
Learning Unknown Intervention Targets in Structural Causal Models from Heterogeneous Data
von: Yang, Yuqin, et al.
Veröffentlicht: (2023)
von: Yang, Yuqin, et al.
Veröffentlicht: (2023)
In-Context Learning for Pure Exploration
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
von: Russo, Alessio, et al.
Veröffentlicht: (2025)
Recursive Causal Discovery
von: Mokhtarian, Ehsan, et al.
Veröffentlicht: (2024)
von: Mokhtarian, Ehsan, et al.
Veröffentlicht: (2024)
Multi-armed Bandits with Missing Outcome
von: Mahrooghi, Ilia, et al.
Veröffentlicht: (2024)
von: Mahrooghi, Ilia, et al.
Veröffentlicht: (2024)
Causal Discovery in Linear Models with Unobserved Variables and Measurement Error
von: Yang, Yuqin, et al.
Veröffentlicht: (2024)
von: Yang, Yuqin, et al.
Veröffentlicht: (2024)
Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026)
von: Afsharrad, Amirhossein, et al.
Veröffentlicht: (2026)
Causal Effect Identification in Heterogeneous Environments from Higher-Order Moments
von: Kivva, Yaroslav, et al.
Veröffentlicht: (2025)
von: Kivva, Yaroslav, et al.
Veröffentlicht: (2025)
Hierarchical Reinforcement Learning with Targeted Causal Interventions
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025)
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2025)
In-Context Learning for Pure Exploration in Continuous Spaces
von: Russo, Alessio, et al.
Veröffentlicht: (2026)
von: Russo, Alessio, et al.
Veröffentlicht: (2026)
Debate as Reward: A Multi-Agent Reward System for Scientific Ideation via RL Post-Training
von: Salimi, Moein, et al.
Veröffentlicht: (2026)
von: Salimi, Moein, et al.
Veröffentlicht: (2026)
Pure Exploration for a Good Policy in Reinforcement Learning with Bandit Feedback
von: Li, Zitian, et al.
Veröffentlicht: (2026)
von: Li, Zitian, et al.
Veröffentlicht: (2026)
Zeroth-Order Stackelberg Control in Combinatorial Congestion Games
von: Masiha, Saeed, et al.
Veröffentlicht: (2026)
von: Masiha, Saeed, et al.
Veröffentlicht: (2026)
UCB for Large-Scale Pure Exploration: Beyond Sub-Gaussianity
von: Li, Zaile, et al.
Veröffentlicht: (2025)
von: Li, Zaile, et al.
Veröffentlicht: (2025)
Causal Effect Identification in lvLiNGAM from Higher-Order Cumulants
von: Tramontano, Daniele, et al.
Veröffentlicht: (2025)
von: Tramontano, Daniele, et al.
Veröffentlicht: (2025)
Causal Effect Identification in LiNGAM Models with Latent Confounders
von: Tramontano, Daniele, et al.
Veröffentlicht: (2024)
von: Tramontano, Daniele, et al.
Veröffentlicht: (2024)
Zero-Shot LLMs in Human-in-the-Loop RL: Replacing Human Feedback for Reward Shaping
von: Nazir, Mohammad Saif, et al.
Veröffentlicht: (2025)
von: Nazir, Mohammad Saif, et al.
Veröffentlicht: (2025)
Optimal Local Convergence Rates of Stochastic First-Order Methods under Local $α$-PL
von: Masiha, Saeed, et al.
Veröffentlicht: (2024)
von: Masiha, Saeed, et al.
Veröffentlicht: (2024)
Fast Proxy Experiment Design for Causal Effect Identification
von: Elahi, Sepehr, et al.
Veröffentlicht: (2024)
von: Elahi, Sepehr, et al.
Veröffentlicht: (2024)
Provably Efficient Exploration in Reward Machines with Low Regret
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
von: Bourel, Hippolyte, et al.
Veröffentlicht: (2024)
Fusing Reward and Dueling Feedback in Stochastic Bandits
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
von: Wang, Xuchuang, et al.
Veröffentlicht: (2025)
Pure Exploration with Infinite Answers
von: Poiani, Riccardo, et al.
Veröffentlicht: (2025)
von: Poiani, Riccardo, et al.
Veröffentlicht: (2025)
Preference-based Pure Exploration
von: Shukla, Apurv, et al.
Veröffentlicht: (2024)
von: Shukla, Apurv, et al.
Veröffentlicht: (2024)
Efficiently Escaping Saddle Points for Policy Optimization
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2023)
von: Khorasani, Sadegh, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Active Context Selection Improves Simple Regret in Contextual Bandits
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2026) -
Graph-Dependent Regret Bounds in Multi-Armed Bandits with Interference
von: Jamshidi, Fateme, et al.
Veröffentlicht: (2025) -
Graph Learning Is Suboptimal in Causal Bandits
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2025) -
Best Group Identification in Multi-Objective Bandits
von: Shahverdikondori, Mohammad, et al.
Veröffentlicht: (2025) -
s-ID: Causal Effect Identification in a Sub-Population
von: Abouei, Amir Mohammad, et al.
Veröffentlicht: (2023)