Good Actions Succeed, Bad Actions Generalize: A Case Study on Why RL Generalizes Better
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Song, Meng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When and Why Does Unsupervised RL Succeed in Mathematical Reasoning? A Manifold Envelopment Perspective
von: Zhang, Zelin, et al.
Veröffentlicht: (2026)
von: Zhang, Zelin, et al.
Veröffentlicht: (2026)
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
von: Suau, Miguel, et al.
Veröffentlicht: (2023)
von: Suau, Miguel, et al.
Veröffentlicht: (2023)
FADE: Why Bad Descriptions Happen to Good Features
von: Puri, Bruno, et al.
Veröffentlicht: (2025)
von: Puri, Bruno, et al.
Veröffentlicht: (2025)
Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning
von: Wu, Xiefeng, et al.
Veröffentlicht: (2025)
von: Wu, Xiefeng, et al.
Veröffentlicht: (2025)
Why LLMs Are Bad at Synthetic Table Generation (and what to do about it)
von: Xu, Shengzhe, et al.
Veröffentlicht: (2024)
von: Xu, Shengzhe, et al.
Veröffentlicht: (2024)
VAM: Verbalized Action Masking for Controllable Exploration in RL Post-Training -- A Chess Case Study
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2026)
Mimicking Better by Matching the Approximate Action Distribution
von: Ramos, João A. Cândido, et al.
Veröffentlicht: (2023)
von: Ramos, João A. Cândido, et al.
Veröffentlicht: (2023)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
von: Driess, Danny, et al.
Veröffentlicht: (2025)
von: Driess, Danny, et al.
Veröffentlicht: (2025)
RL Token: Bootstrapping Online RL with Vision-Language-Action Models
von: Xu, Charles, et al.
Veröffentlicht: (2026)
von: Xu, Charles, et al.
Veröffentlicht: (2026)
When Good Equations Get Bad Scores: Improving Symbolic Regression Through Better Parameter Optimization
von: Wang, Boxiao, et al.
Veröffentlicht: (2026)
von: Wang, Boxiao, et al.
Veröffentlicht: (2026)
Good Allocations from Bad Estimates
von: Casacuberta, Sílvia, et al.
Veröffentlicht: (2026)
von: Casacuberta, Sílvia, et al.
Veröffentlicht: (2026)
ActionPiece: Contextually Tokenizing Action Sequences for Generative Recommendation
von: Hou, Yupeng, et al.
Veröffentlicht: (2025)
von: Hou, Yupeng, et al.
Veröffentlicht: (2025)
The Infinite-Dimensional Nature of Spectroscopy and Why Models Succeed, Fail, and Mislead
von: Michelucci, Umberto, et al.
Veröffentlicht: (2026)
von: Michelucci, Umberto, et al.
Veröffentlicht: (2026)
Scalable Offline Model-Based RL with Action Chunks
von: Park, Kwanyoung, et al.
Veröffentlicht: (2025)
von: Park, Kwanyoung, et al.
Veröffentlicht: (2025)
Better Generalization with Semantic IDs: A Case Study in Ranking for Recommendations
von: Singh, Anima, et al.
Veröffentlicht: (2023)
von: Singh, Anima, et al.
Veröffentlicht: (2023)
$π_\texttt{RL}$: Online RL Fine-tuning for Flow-based Vision-Language-Action Models
von: Chen, Kang, et al.
Veröffentlicht: (2025)
von: Chen, Kang, et al.
Veröffentlicht: (2025)
From Actions to Words: Towards Abstractive-Textual Policy Summarization in RL
von: Admoni, Sahar, et al.
Veröffentlicht: (2025)
von: Admoni, Sahar, et al.
Veröffentlicht: (2025)
SALSA-RL: Stability Analysis in the Latent Space of Actions for Reinforcement Learning
von: Li, Xuyang, et al.
Veröffentlicht: (2025)
von: Li, Xuyang, et al.
Veröffentlicht: (2025)
Benchmarking the Generality of Vision-Language-Action Models
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
von: Guruprasad, Pranav, et al.
Veröffentlicht: (2025)
Stochastic Gradient Succeeds for Bandits
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
von: Mei, Jincheng, et al.
Veröffentlicht: (2024)
Behavior Generation with Latent Actions
von: Lee, Seungjae, et al.
Veröffentlicht: (2024)
von: Lee, Seungjae, et al.
Veröffentlicht: (2024)
Towards Understanding Why FixMatch Generalizes Better Than Supervised Learning
von: Li, Jingyang, et al.
Veröffentlicht: (2024)
von: Li, Jingyang, et al.
Veröffentlicht: (2024)
Learning to Generate All Feasible Actions
von: Theile, Mirco, et al.
Veröffentlicht: (2023)
von: Theile, Mirco, et al.
Veröffentlicht: (2023)
Assessing the Zero-Shot Capabilities of LLMs for Action Evaluation in RL
von: Pignatelli, Eduardo, et al.
Veröffentlicht: (2024)
von: Pignatelli, Eduardo, et al.
Veröffentlicht: (2024)
Reasoning through Verifiable Forecast Actions: Consistency-Grounded RL for Financial LLMs
von: Chen, Jialin, et al.
Veröffentlicht: (2026)
von: Chen, Jialin, et al.
Veröffentlicht: (2026)
Off-Policy Maximum Entropy RL with Future State and Action Visitation Measures
von: Bolland, Adrien, et al.
Veröffentlicht: (2024)
von: Bolland, Adrien, et al.
Veröffentlicht: (2024)
Recommender Systems for Good (RS4Good): Survey of Use Cases and a Call to Action for Research that Matters
von: Jannach, Dietmar, et al.
Veröffentlicht: (2024)
von: Jannach, Dietmar, et al.
Veröffentlicht: (2024)
$\varepsilon$-Good Action Identification in Fixed-Budget Monte Carlo Tree Search
von: Li, Yinan, et al.
Veröffentlicht: (2026)
von: Li, Yinan, et al.
Veröffentlicht: (2026)
The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
von: Sahoo, Subramanyam
Veröffentlicht: (2025)
von: Sahoo, Subramanyam
Veröffentlicht: (2025)
Action-Adaptive Continual Learning: Enabling Policy Generalization under Dynamic Action Spaces
von: Pan, Chaofan, et al.
Veröffentlicht: (2025)
von: Pan, Chaofan, et al.
Veröffentlicht: (2025)
Vertical Federated Learning in Practice: The Good, the Bad, and the Ugly
von: Wu, Zhaomin, et al.
Veröffentlicht: (2025)
von: Wu, Zhaomin, et al.
Veröffentlicht: (2025)
Sequence-Aware Inline Measurement Attribution for Good-Bad Wafer Diagnosis
von: Miyaguchi, Kohei, et al.
Veröffentlicht: (2025)
von: Miyaguchi, Kohei, et al.
Veröffentlicht: (2025)
Agent Performing Autonomous Stock Trading under Good and Bad Situations
von: Luo, Yunfei, et al.
Veröffentlicht: (2023)
von: Luo, Yunfei, et al.
Veröffentlicht: (2023)
HIQL: Offline Goal-Conditioned RL with Latent States as Actions
von: Park, Seohong, et al.
Veröffentlicht: (2023)
von: Park, Seohong, et al.
Veröffentlicht: (2023)
When Bad Data Leads to Good Models
von: Li, Kenneth, et al.
Veröffentlicht: (2025)
von: Li, Kenneth, et al.
Veröffentlicht: (2025)
Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
von: Gupta, Isha, et al.
Veröffentlicht: (2025)
Action-Free Offline-to-Online RL via Discretised State Policies
von: Neggatu, Natinael Solomon, et al.
Veröffentlicht: (2026)
von: Neggatu, Natinael Solomon, et al.
Veröffentlicht: (2026)
Improving and Accelerating Offline RL in Large Discrete Action Spaces with Structured Policy Initialization
von: Landers, Matthew, et al.
Veröffentlicht: (2026)
von: Landers, Matthew, et al.
Veröffentlicht: (2026)
GEM: Guided Expectation-Maximization for Behavior-Normalized Candidate Action Selection in Offline RL
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
von: Wang, Haoyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
When and Why Does Unsupervised RL Succeed in Mathematical Reasoning? A Manifold Envelopment Perspective
von: Zhang, Zelin, et al.
Veröffentlicht: (2026) -
Bad Habits: Policy Confounding and Out-of-Trajectory Generalization in RL
von: Suau, Miguel, et al.
Veröffentlicht: (2023) -
FADE: Why Bad Descriptions Happen to Good Features
von: Puri, Bruno, et al.
Veröffentlicht: (2025) -
Teaching RL Agents to Act Better: VLM as Action Advisor for Online Reinforcement Learning
von: Wu, Xiefeng, et al.
Veröffentlicht: (2025) -
Why LLMs Are Bad at Synthetic Table Generation (and what to do about it)
von: Xu, Shengzhe, et al.
Veröffentlicht: (2024)