Do Transformer World Models Give Better Policy Gradients?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ma, Michel, Ni, Tianwei, Gehring, Clement, D'Oro, Pierluca, Bacon, Pierre-Luc |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bridging State and History Representations: Understanding Self-Predictive RL
von: Ni, Tianwei, et al.
Veröffentlicht: (2024)
von: Ni, Tianwei, et al.
Veröffentlicht: (2024)
Maxwell's Demon at Work: Efficient Pruning by Leveraging Saturation of Neurons
von: Dufort-Labbé, Simon, et al.
Veröffentlicht: (2024)
von: Dufort-Labbé, Simon, et al.
Veröffentlicht: (2024)
Mol-MoE: Training Preference-Guided Routers for Molecule Generation
von: Calanzone, Diego, et al.
Veröffentlicht: (2025)
von: Calanzone, Diego, et al.
Veröffentlicht: (2025)
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
von: Onoda, Ku, et al.
Veröffentlicht: (2026)
von: Onoda, Ku, et al.
Veröffentlicht: (2026)
Towards General-Purpose Model-Free Reinforcement Learning
von: Fujimoto, Scott, et al.
Veröffentlicht: (2025)
von: Fujimoto, Scott, et al.
Veröffentlicht: (2025)
Policy Optimization in a Noisy Neighborhood: On Return Landscapes in Continuous Control
von: Rahn, Nate, et al.
Veröffentlicht: (2023)
von: Rahn, Nate, et al.
Veröffentlicht: (2023)
ADEPTS: A Capability Framework for Human-Centered Agent Design
von: D'Oro, Pierluca, et al.
Veröffentlicht: (2025)
von: D'Oro, Pierluca, et al.
Veröffentlicht: (2025)
The Three Regimes of Offline-to-Online Reinforcement Learning
von: Li, Lu, et al.
Veröffentlicht: (2025)
von: Li, Lu, et al.
Veröffentlicht: (2025)
MaestroMotif: Skill Design from Artificial Intelligence Feedback
von: Klissarov, Martin, et al.
Veröffentlicht: (2024)
von: Klissarov, Martin, et al.
Veröffentlicht: (2024)
Understanding Behavioral Metric Learning: A Large-Scale Study on Distracting Reinforcement Learning Environments
von: Luo, Ziyan, et al.
Veröffentlicht: (2025)
von: Luo, Ziyan, et al.
Veröffentlicht: (2025)
Hierarchical Behaviour Spaces
von: Matthews, Michael Tryfan, et al.
Veröffentlicht: (2026)
von: Matthews, Michael Tryfan, et al.
Veröffentlicht: (2026)
Rotation-Preserving Supervised Fine-Tuning
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
von: Jin, Hangzhan, et al.
Veröffentlicht: (2026)
Controlling Multimodal LLMs via Reward-guided Decoding
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
von: Mañas, Oscar, et al.
Veröffentlicht: (2025)
The Curse of Diversity in Ensemble-Based Exploration
von: Lin, Zhixuan, et al.
Veröffentlicht: (2024)
von: Lin, Zhixuan, et al.
Veröffentlicht: (2024)
Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement Learning
von: Ma, Guozheng, et al.
Veröffentlicht: (2025)
von: Ma, Guozheng, et al.
Veröffentlicht: (2025)
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
von: Muni, Aneri, et al.
Veröffentlicht: (2026)
von: Muni, Aneri, et al.
Veröffentlicht: (2026)
The Gradient of Algebraic Model Counting
von: Maene, Jaron, et al.
Veröffentlicht: (2025)
von: Maene, Jaron, et al.
Veröffentlicht: (2025)
Better World Models Can Lead to Better Post-Training Performance
von: Gupta, Prakhar, et al.
Veröffentlicht: (2025)
von: Gupta, Prakhar, et al.
Veröffentlicht: (2025)
Layerwise LQR for Geometry-Aware Optimization of Deep Networks
von: Dufort-Labbé, Simon, et al.
Veröffentlicht: (2026)
von: Dufort-Labbé, Simon, et al.
Veröffentlicht: (2026)
Learning with Differentially Private (Sliced) Wasserstein Gradients
von: Rodríguez-Vítores, David, et al.
Veröffentlicht: (2025)
von: Rodríguez-Vítores, David, et al.
Veröffentlicht: (2025)
Better Decisions through the Right Causal World Model
von: Dillies, Elisabeth, et al.
Veröffentlicht: (2025)
von: Dillies, Elisabeth, et al.
Veröffentlicht: (2025)
GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
Do pretrained Transformers Learn In-Context by Gradient Descent?
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
von: Shen, Lingfeng, et al.
Veröffentlicht: (2023)
Do Large Language Models Reason Causally Like Us? Even Better?
von: Dettki, Hanna M., et al.
Veröffentlicht: (2025)
von: Dettki, Hanna M., et al.
Veröffentlicht: (2025)
nnterp: A Standardized Interface for Mechanistic Interpretability of Transformers
von: Dumas, Clément
Veröffentlicht: (2025)
von: Dumas, Clément
Veröffentlicht: (2025)
Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct Backpropagation
von: Sun, Mingfei
Veröffentlicht: (2026)
von: Sun, Mingfei
Veröffentlicht: (2026)
Asymptotically Stable Quaternion-valued Hopfield-structured Neural Network with Periodic Projection-based Supervised Learning Rules
von: Wang, Tianwei, et al.
Veröffentlicht: (2025)
von: Wang, Tianwei, et al.
Veröffentlicht: (2025)
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
von: Hassid, Michael, et al.
Veröffentlicht: (2024)
von: Hassid, Michael, et al.
Veröffentlicht: (2024)
WorldLLM: Improving LLMs' world modeling using curiosity-driven theory-making
von: Levy, Guillaume, et al.
Veröffentlicht: (2025)
von: Levy, Guillaume, et al.
Veröffentlicht: (2025)
Learning General Policies with Policy Gradient Methods
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2025)
von: Ståhlberg, Simon, et al.
Veröffentlicht: (2025)
Generalized Policy Gradient with History-Aware Decision Transformer for Reliable Routing over Graph Signals
von: Wei, Xing, et al.
Veröffentlicht: (2025)
von: Wei, Xing, et al.
Veröffentlicht: (2025)
Policy Gradient with Kernel Quadrature
von: Hayakawa, Satoshi, et al.
Veröffentlicht: (2023)
von: Hayakawa, Satoshi, et al.
Veröffentlicht: (2023)
Policy Gradient with Tree Expansion
von: Dalal, Gal, et al.
Veröffentlicht: (2023)
von: Dalal, Gal, et al.
Veröffentlicht: (2023)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
von: Kurtic, Eldar, et al.
Veröffentlicht: (2024)
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
von: Mousavi-Hosseini, Alireza, et al.
Veröffentlicht: (2026)
Robust Finite-Memory Policy Gradients for Hidden-Model POMDPs
von: Galesloot, Maris F. L., et al.
Veröffentlicht: (2025)
von: Galesloot, Maris F. L., et al.
Veröffentlicht: (2025)
$K$-Level Policy Gradients for Multi-Agent Reinforcement Learning
von: Reddi, Aryaman, et al.
Veröffentlicht: (2025)
von: Reddi, Aryaman, et al.
Veröffentlicht: (2025)
Gradient Extrapolation-Based Policy Optimization
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2026)
von: Swapnil, Ismam Nur, et al.
Veröffentlicht: (2026)
Partial Policy Gradients for RL in LLMs
von: Mathur, Puneet, et al.
Veröffentlicht: (2026)
von: Mathur, Puneet, et al.
Veröffentlicht: (2026)
Training Large Language Models to Reason via EM Policy Gradient
von: Xu, Tianbing
Veröffentlicht: (2025)
von: Xu, Tianbing
Veröffentlicht: (2025)
Ähnliche Einträge
-
Bridging State and History Representations: Understanding Self-Predictive RL
von: Ni, Tianwei, et al.
Veröffentlicht: (2024) -
Maxwell's Demon at Work: Efficient Pruning by Leveraging Saturation of Neurons
von: Dufort-Labbé, Simon, et al.
Veröffentlicht: (2024) -
Mol-MoE: Training Preference-Guided Routers for Molecule Generation
von: Calanzone, Diego, et al.
Veröffentlicht: (2025) -
Does "Do Differentiable Simulators Give Better Policy Gradients?'' Give Better Policy Gradients?
von: Onoda, Ku, et al.
Veröffentlicht: (2026) -
Towards General-Purpose Model-Free Reinforcement Learning
von: Fujimoto, Scott, et al.
Veröffentlicht: (2025)