q-exponential family for policy optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Lingwei, Shah, Haseeb, Wang, Han, Nagai, Yukie, White, Martha |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Symmetric Behavior Regularized Policy Optimization
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
Towards Physiologically Sensible Predictions via the Rule-based Reinforcement Learning Layer
by: Zhu, Lingwei, et al.
Published: (2025)
by: Zhu, Lingwei, et al.
Published: (2025)
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
by: Zhu, Lingwei, et al.
Published: (2023)
by: Zhu, Lingwei, et al.
Published: (2023)
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
by: Herbuela, Von Ralph Dane Marquez, et al.
Published: (2025)
by: Herbuela, Von Ralph Dane Marquez, et al.
Published: (2025)
A Generalized Projected Bellman Error for Off-policy Value Estimation in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2021)
by: Patterson, Andrew, et al.
Published: (2021)
Fine-Tuning without Performance Degradation
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Correspondence learning between morphologically different robots via task demonstrations
by: Aktas, Hakan, et al.
Published: (2023)
by: Aktas, Hakan, et al.
Published: (2023)
Laws of thermodynamics for exponential families
by: Balsubramani, Akshay
Published: (2025)
by: Balsubramani, Akshay
Published: (2025)
Cross-Embodied Affordance Transfer through Learning Affordance Equivalences
by: Aktas, Hakan, et al.
Published: (2024)
by: Aktas, Hakan, et al.
Published: (2024)
Extracting Money Laundering Transactions from Quasi-Temporal Graph Representation
by: Tariq, Haseeb, et al.
Published: (2026)
by: Tariq, Haseeb, et al.
Published: (2026)
What to Do When Your Discrete Optimization Is the Size of a Neural Network?
by: Silva, Hugo, et al.
Published: (2024)
by: Silva, Hugo, et al.
Published: (2024)
Adam with model exponential moving average is effective for nonconvex optimization
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Divergences induced by dual subtractive and divisive normalizations of exponential families and their convex deformations
by: Nielsen, Frank
Published: (2023)
by: Nielsen, Frank
Published: (2023)
Memory-Augmented Architecture for Long-Term Context Handling in Large Language Models
by: Shinwari, Haseeb Ullah Khan, et al.
Published: (2025)
by: Shinwari, Haseeb Ullah Khan, et al.
Published: (2025)
Detecting Complex Money Laundering Patterns with Incremental and Distributed Graph Modeling
by: Tariq, Haseeb, et al.
Published: (2026)
by: Tariq, Haseeb, et al.
Published: (2026)
Investigating Action Encodings in Recurrent Neural Networks in Reinforcement Learning
by: Schlegel, Matthew, et al.
Published: (2026)
by: Schlegel, Matthew, et al.
Published: (2026)
Regret of exploratory policy improvement and $q$-learning
by: Tang, Wenpin, et al.
Published: (2024)
by: Tang, Wenpin, et al.
Published: (2024)
Deriving Lehmer and Hölder means as maximum weighted likelihood estimates for the multivariate exponential family
by: Ziou, Djemel, et al.
Published: (2024)
by: Ziou, Djemel, et al.
Published: (2024)
Selective Denoising Diffusion Model for Time Series Anomaly Detection
by: Obata, Kohei, et al.
Published: (2026)
by: Obata, Kohei, et al.
Published: (2026)
Mechanistic Analysis of Circuit Preservation in Federated Learning
by: Haseeb, Muhammad, et al.
Published: (2025)
by: Haseeb, Muhammad, et al.
Published: (2025)
Risk and optimal policies in bandit experiments
by: Adusumilli, Karun
Published: (2021)
by: Adusumilli, Karun
Published: (2021)
Rethinking the Role of Temperature in Large Language Model Distillation
by: Luong, Hoang-Chau, et al.
Published: (2026)
by: Luong, Hoang-Chau, et al.
Published: (2026)
The Cross-environment Hyperparameter Setting Benchmark for Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2024)
by: Patterson, Andrew, et al.
Published: (2024)
Some remarks on gradient dominance and LQR policy optimization
by: Sontag, Eduardo D.
Published: (2025)
by: Sontag, Eduardo D.
Published: (2025)
Topology-Agnostic Detection of Temporal Money Laundering Flows in Billion-Scale Transactions
by: Tariq, Haseeb, et al.
Published: (2023)
by: Tariq, Haseeb, et al.
Published: (2023)
ARD-LoRA: Dynamic Rank Allocation for Parameter-Efficient Fine-Tuning of Foundation Models with Heterogeneous Adaptation Needs
by: Shinwari, Haseeb Ullah Khan, et al.
Published: (2025)
by: Shinwari, Haseeb Ullah Khan, et al.
Published: (2025)
Real-Time Recurrent Learning using Trace Units in Reinforcement Learning
by: Elelimy, Esraa, et al.
Published: (2024)
by: Elelimy, Esraa, et al.
Published: (2024)
Exact characterization of ε-Safe Decision Regions for exponential family distributions and Multi Cost SVM approximation
by: Carlevaro, Alberto, et al.
Published: (2025)
by: Carlevaro, Alberto, et al.
Published: (2025)
Empirical Design in Reinforcement Learning
by: Patterson, Andrew, et al.
Published: (2023)
by: Patterson, Andrew, et al.
Published: (2023)
qPOTS: Efficient batch multiobjective Bayesian optimization via Pareto optimal Thompson sampling
by: Renganathan, Ashwin, et al.
Published: (2023)
by: Renganathan, Ashwin, et al.
Published: (2023)
A policy gradient approach for optimization of smooth risk measures
by: Vijayan, Nithia, et al.
Published: (2022)
by: Vijayan, Nithia, et al.
Published: (2022)
Averaging $n$-step Returns Reduces Variance in Reinforcement Learning
by: Daley, Brett, et al.
Published: (2024)
by: Daley, Brett, et al.
Published: (2024)
Accelerating trajectory optimization with Sobolev-trained diffusion policies
by: Hellard, Théotime Le, et al.
Published: (2026)
by: Hellard, Théotime Le, et al.
Published: (2026)
Investigating the Interplay of Prioritized Replay and Generalization
by: Panahi, Parham Mohammad, et al.
Published: (2024)
by: Panahi, Parham Mohammad, et al.
Published: (2024)
A note on convergence of Wasserstein policy optimization
by: Šiška, David, et al.
Published: (2026)
by: Šiška, David, et al.
Published: (2026)
From exponential to finite/fixed-time stability: Applications to optimization
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
by: Ozaslan, Ibrahim K., et al.
Published: (2024)
Position: Lifetime tuning is incompatible with continual reinforcement learning
by: Mesbahi, Golnaz, et al.
Published: (2024)
by: Mesbahi, Golnaz, et al.
Published: (2024)
Transformer-based CoVaR: Systemic Risk in Textual Information
by: Chen, Junyu, et al.
Published: (2026)
by: Chen, Junyu, et al.
Published: (2026)
Value Bonuses using Ensemble Errors for Exploration in Reinforcement Learning
by: Wahab, Abdul, et al.
Published: (2026)
by: Wahab, Abdul, et al.
Published: (2026)
Similar Items
-
Symmetric Behavior Regularized Policy Optimization
by: Zhu, Lingwei, et al.
Published: (2025) -
Fat-to-Thin Policy Optimization: Offline RL with Sparse Policies
by: Zhu, Lingwei, et al.
Published: (2025) -
Towards Physiologically Sensible Predictions via the Rule-based Reinforcement Learning Layer
by: Zhu, Lingwei, et al.
Published: (2025) -
Generalized Munchausen Reinforcement Learning using Tsallis KL Divergence
by: Zhu, Lingwei, et al.
Published: (2023) -
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
by: Herbuela, Von Ralph Dane Marquez, et al.
Published: (2025)