Distributional Bellman Operators over Mean Embeddings
Fuente:
arXiv
Guardado en:
| Autores principales: | Wenliang, Li Kevin, Delétang, Grégoire, Aitchison, Matthew, Hutter, Marcus, Ruoss, Anian, Gretton, Arthur, Rowland, Mark |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning Universal Predictors
por: Grau-Moya, Jordi, et al.
Publicado: (2024)
por: Grau-Moya, Jordi, et al.
Publicado: (2024)
Why is prompting hard? Understanding prompts on binary sequence predictors
por: Wenliang, Li Kevin, et al.
Publicado: (2025)
por: Wenliang, Li Kevin, et al.
Publicado: (2025)
Language Modeling Is Compression
por: Delétang, Grégoire, et al.
Publicado: (2023)
por: Delétang, Grégoire, et al.
Publicado: (2023)
Understanding Prompt Tuning and In-Context Learning via Meta-Learning
por: Genewein, Tim, et al.
Publicado: (2025)
por: Genewein, Tim, et al.
Publicado: (2025)
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
por: Ruoss, Anian, et al.
Publicado: (2024)
por: Ruoss, Anian, et al.
Publicado: (2024)
Foundations of Multivariate Distributional Reinforcement Learning
por: Wiltzer, Harley, et al.
Publicado: (2024)
por: Wiltzer, Harley, et al.
Publicado: (2024)
Doubly-Robust Estimation of Counterfactual Policy Mean Embeddings
por: Zenati, Houssam, et al.
Publicado: (2025)
por: Zenati, Houssam, et al.
Publicado: (2025)
Compression via Pre-trained Transformers: A Study on Byte-Level Multimodal Data
por: Heurtel-Depeiges, David, et al.
Publicado: (2024)
por: Heurtel-Depeiges, David, et al.
Publicado: (2024)
Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative Model
por: Rowland, Mark, et al.
Publicado: (2024)
por: Rowland, Mark, et al.
Publicado: (2024)
Semiparametric Efficient Test for Interpretable Distributional Treatment Effects
por: Zenati, Houssam, et al.
Publicado: (2026)
por: Zenati, Houssam, et al.
Publicado: (2026)
Doubly Robust Proxy Causal Learning with Neural Mean Embeddings
por: Bozkurt, Bariscan, et al.
Publicado: (2026)
por: Bozkurt, Bariscan, et al.
Publicado: (2026)
LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations
por: Ruoss, Anian, et al.
Publicado: (2024)
por: Ruoss, Anian, et al.
Publicado: (2024)
A Distributional Analogue to the Successor Representation
por: Wiltzer, Harley, et al.
Publicado: (2024)
por: Wiltzer, Harley, et al.
Publicado: (2024)
Evaluating Frontier Models for Dangerous Capabilities
por: Phuong, Mary, et al.
Publicado: (2024)
por: Phuong, Mary, et al.
Publicado: (2024)
Sequential Kernel Embedding for Mediated and Time-Varying Dose Response Curves
por: Singh, Rahul, et al.
Publicado: (2021)
por: Singh, Rahul, et al.
Publicado: (2021)
On the Wasserstein Gradient Flow Interpretation of Drifting Models
por: Gretton, Arthur, et al.
Publicado: (2026)
por: Gretton, Arthur, et al.
Publicado: (2026)
Near-Optimality of Contrastive Divergence Algorithms
por: Glaser, Pierre, et al.
Publicado: (2025)
por: Glaser, Pierre, et al.
Publicado: (2025)
Kernel Single Proxy Control for Deterministic Confounding
por: Xu, Liyuan, et al.
Publicado: (2023)
por: Xu, Liyuan, et al.
Publicado: (2023)
Perturbative methods for non-parametric instrumental variable
por: Bu, Wei, et al.
Publicado: (2026)
por: Bu, Wei, et al.
Publicado: (2026)
A Spectral Revisit of the Distributional Bellman Operator under the Cramér Metric
por: Wang, Keru, et al.
Publicado: (2026)
por: Wang, Keru, et al.
Publicado: (2026)
Bellman Diffusion: Generative Modeling as Learning a Linear Operator in the Distribution Space
por: Li, Yangming, et al.
Publicado: (2024)
por: Li, Yangming, et al.
Publicado: (2024)
(De)-regularized Maximum Mean Discrepancy Gradient Flow
por: Chen, Zonghao, et al.
Publicado: (2024)
por: Chen, Zonghao, et al.
Publicado: (2024)
Parameterized Projected Bellman Operator
por: Vincent, Théo, et al.
Publicado: (2023)
por: Vincent, Théo, et al.
Publicado: (2023)
Distributional Diffusion Models with Scoring Rules
por: De Bortoli, Valentin, et al.
Publicado: (2025)
por: De Bortoli, Valentin, et al.
Publicado: (2025)
Regularized $f$-Divergence Kernel Tests
por: Ribero, Mónica, et al.
Publicado: (2026)
por: Ribero, Mónica, et al.
Publicado: (2026)
Deep Proxy Causal Learning and its Application to Confounded Bandit Policy Evaluation
por: Xu, Liyuan, et al.
Publicado: (2021)
por: Xu, Liyuan, et al.
Publicado: (2021)
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
por: Omura, Motoki, et al.
Publicado: (2025)
por: Omura, Motoki, et al.
Publicado: (2025)
Interventional Processes for Causal Uncertainty Quantification
por: Dance, Hugh, et al.
Publicado: (2024)
por: Dance, Hugh, et al.
Publicado: (2024)
Kernel Treatment Effects with Adaptively Collected Data
por: Zenati, Houssam, et al.
Publicado: (2025)
por: Zenati, Houssam, et al.
Publicado: (2025)
Your Policy Regularizer is Secretly an Adversary
por: Brekelmans, Rob, et al.
Publicado: (2022)
por: Brekelmans, Rob, et al.
Publicado: (2022)
Towards Optimal Sobolev Norm Rates for the Vector-Valued Regularized Least-Squares Algorithm
por: Li, Zhu, et al.
Publicado: (2023)
por: Li, Zhu, et al.
Publicado: (2023)
Nonparametric Instrumental Regression via Kernel Methods is Minimax Optimal
por: Meunier, Dimitri, et al.
Publicado: (2024)
por: Meunier, Dimitri, et al.
Publicado: (2024)
Why you don't overfit, and don't need Bayes if you only train for one epoch
por: Aitchison, Laurence
Publicado: (2024)
por: Aitchison, Laurence
Publicado: (2024)
Nonlinear Meta-Learning Can Guarantee Faster Rates
por: Meunier, Dimitri, et al.
Publicado: (2023)
por: Meunier, Dimitri, et al.
Publicado: (2023)
Fast and Scalable Score-Based Kernel Calibration Tests
por: Glaser, Pierre, et al.
Publicado: (2025)
por: Glaser, Pierre, et al.
Publicado: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
por: Chen, Zonghao, et al.
Publicado: (2025)
por: Chen, Zonghao, et al.
Publicado: (2025)
Semiparametrically Efficient Inference for Kernel Measures of Noise Heterogeneity
por: Wornbard, Jakub, et al.
Publicado: (2026)
por: Wornbard, Jakub, et al.
Publicado: (2026)
Distributional Off-policy Evaluation with Bellman Residual Minimization
por: Hong, Sungee, et al.
Publicado: (2024)
por: Hong, Sungee, et al.
Publicado: (2024)
Deep MMD Gradient Flow without adversarial training
por: Galashov, Alexandre, et al.
Publicado: (2024)
por: Galashov, Alexandre, et al.
Publicado: (2024)
Optimal Rates for Vector-Valued Spectral Regularization Learning Algorithms
por: Meunier, Dimitri, et al.
Publicado: (2024)
por: Meunier, Dimitri, et al.
Publicado: (2024)
Ejemplares similares
-
Learning Universal Predictors
por: Grau-Moya, Jordi, et al.
Publicado: (2024) -
Why is prompting hard? Understanding prompts on binary sequence predictors
por: Wenliang, Li Kevin, et al.
Publicado: (2025) -
Language Modeling Is Compression
por: Delétang, Grégoire, et al.
Publicado: (2023) -
Understanding Prompt Tuning and In-Context Learning via Meta-Learning
por: Genewein, Tim, et al.
Publicado: (2025) -
Amortized Planning with Large-Scale Transformers: A Case Study on Chess
por: Ruoss, Anian, et al.
Publicado: (2024)