CAST: Non-Privileged Clipped Asymmetric Self-Teaching with Advantage Flipping for GRPO
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Yang, Xue, Gongle, Guo, Yijia, Yuan, Yuheng, Hu, Liwen, Ma, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025)
DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage
von: Gao, Haowen, et al.
Veröffentlicht: (2026)
von: Gao, Haowen, et al.
Veröffentlicht: (2026)
Negative Advantages Is a Double-Edged Sword: Calibrating advantages in GRPO for Search Agents
von: Wu, Jiayi, et al.
Veröffentlicht: (2026)
von: Wu, Jiayi, et al.
Veröffentlicht: (2026)
Adaptive-Boundary-Clipping GRPO: Ensuring Bounded Ratios for Stable and Generalizable Training
von: Liu, Chi, et al.
Veröffentlicht: (2026)
von: Liu, Chi, et al.
Veröffentlicht: (2026)
Privileged Self-Access Matters for Introspection in AI
von: Song, Siyuan, et al.
Veröffentlicht: (2025)
von: Song, Siyuan, et al.
Veröffentlicht: (2025)
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
von: Ding, Zheng, et al.
Veröffentlicht: (2025)
von: Ding, Zheng, et al.
Veröffentlicht: (2025)
Unveiling Implicit Advantage Symmetry: Why GRPO Struggles with Exploration and Difficulty Adaptation
von: Yu, Zhiqi, et al.
Veröffentlicht: (2026)
von: Yu, Zhiqi, et al.
Veröffentlicht: (2026)
Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
von: Wang, Cheng, et al.
Veröffentlicht: (2026)
FORTIS: Benchmarking Over-Privilege in Agent Skills
von: Li, Shawn, et al.
Veröffentlicht: (2026)
von: Li, Shawn, et al.
Veröffentlicht: (2026)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
von: Chen, Minghan, et al.
Veröffentlicht: (2025)
von: Chen, Minghan, et al.
Veröffentlicht: (2025)
iGRPO: Self-Feedback-Driven LLM Reasoning
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2026)
MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
von: Li, Junzhe, et al.
Veröffentlicht: (2025)
von: Li, Junzhe, et al.
Veröffentlicht: (2025)
Emergence: Overcoming Privileged Information Bias in Asymmetric Embodied Agents via Active Querying
von: Baek, Shaun, et al.
Veröffentlicht: (2025)
von: Baek, Shaun, et al.
Veröffentlicht: (2025)
Progent: Securing AI Agents with Privilege Control
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
von: Shi, Tianneng, et al.
Veröffentlicht: (2025)
FlipAttack: Jailbreak LLMs via Flipping
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
AMIR-GRPO: Inducing Implicit Preference Signals into GRPO
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
von: Yari, Amir Hossein, et al.
Veröffentlicht: (2026)
Privileged Sensing Scaffolds Reinforcement Learning
von: Hu, Edward S., et al.
Veröffentlicht: (2024)
von: Hu, Edward S., et al.
Veröffentlicht: (2024)
Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
von: Hong, Haoyang, et al.
Veröffentlicht: (2025)
Do Coding Agents Understand Least-Privilege Authorization?
von: Yan, Zheng, et al.
Veröffentlicht: (2026)
von: Yan, Zheng, et al.
Veröffentlicht: (2026)
Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
von: Gu, Hengrui, et al.
Veröffentlicht: (2026)
HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
von: Ding, Ken
Veröffentlicht: (2026)
von: Ding, Ken
Veröffentlicht: (2026)
DCPO: Dynamic Clipping Policy Optimization
von: Yang, Shihui, et al.
Veröffentlicht: (2025)
von: Yang, Shihui, et al.
Veröffentlicht: (2025)
Dialogue Model Optimization via Agent Game and Adaptive Tree-based GRPO
von: Peng, Kun, et al.
Veröffentlicht: (2026)
von: Peng, Kun, et al.
Veröffentlicht: (2026)
Skeleton-based Action Recognition with Non-linear Dependency Modeling and Hilbert-Schmidt Independence Criterion
von: Yang, Yuheng
Veröffentlicht: (2024)
von: Yang, Yuheng
Veröffentlicht: (2024)
DRUPI: Dataset Reduction Using Privileged Information
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
CAST: Achieving Stable LLM-based Text Analysis for Data Analytics
von: Xie, Jinxiang, et al.
Veröffentlicht: (2026)
von: Xie, Jinxiang, et al.
Veröffentlicht: (2026)
PPO-Clip Attains Global Optimality: Towards Deeper Understandings of Clipping
von: Huang, Nai-Chieh, et al.
Veröffentlicht: (2023)
von: Huang, Nai-Chieh, et al.
Veröffentlicht: (2023)
EvoFSM: Controllable Self-Evolution for Deep Research with Finite State Machines
von: Zhang, Shuo, et al.
Veröffentlicht: (2026)
von: Zhang, Shuo, et al.
Veröffentlicht: (2026)
HiCAST: Highly Customized Arbitrary Style Transfer with Adapter Enhanced Diffusion Models
von: Wang, Hanzhang, et al.
Veröffentlicht: (2024)
von: Wang, Hanzhang, et al.
Veröffentlicht: (2024)
A Safe and Efficient Self-evolving Algorithm for Decision-making and Control of Autonomous Driving Systems
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
von: Yang, Shuo, et al.
Veröffentlicht: (2024)
CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation
von: Taghavi, Pardis, et al.
Veröffentlicht: (2025)
von: Taghavi, Pardis, et al.
Veröffentlicht: (2025)
Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables
von: Chen, Yanzuo, et al.
Veröffentlicht: (2023)
von: Chen, Yanzuo, et al.
Veröffentlicht: (2023)
VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory
von: Lei, Yuheng, et al.
Veröffentlicht: (2026)
von: Lei, Yuheng, et al.
Veröffentlicht: (2026)
MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
von: Jia, Weitao, et al.
Veröffentlicht: (2025)
von: Jia, Weitao, et al.
Veröffentlicht: (2025)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
von: Liu, Henglin, et al.
Veröffentlicht: (2025)
von: Liu, Henglin, et al.
Veröffentlicht: (2025)
Signal Reshaping for GRPO in Weak-Feedback Agentic Code Repair
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
Improving LLM-Generated Code Quality with GRPO
von: Robeyns, Maxime, et al.
Veröffentlicht: (2025)
von: Robeyns, Maxime, et al.
Veröffentlicht: (2025)
PAPO: Stabilizing Rubric Integration Training via Decoupled Advantage Normalization
von: Tan, Zelin, et al.
Veröffentlicht: (2026)
von: Tan, Zelin, et al.
Veröffentlicht: (2026)
What is the Alignment Objective of GRPO?
von: Vojnovic, Milan, et al.
Veröffentlicht: (2025)
von: Vojnovic, Milan, et al.
Veröffentlicht: (2025)
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
von: Fei, Wu, et al.
Veröffentlicht: (2025)
von: Fei, Wu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
von: Zhang, Xingjian, et al.
Veröffentlicht: (2025) -
DIVA-GRPO: Enhancing Multimodal Reasoning through Difficulty-Adaptive Variant Advantage
von: Gao, Haowen, et al.
Veröffentlicht: (2026) -
Negative Advantages Is a Double-Edged Sword: Calibrating advantages in GRPO for Search Agents
von: Wu, Jiayi, et al.
Veröffentlicht: (2026) -
Adaptive-Boundary-Clipping GRPO: Ensuring Bounded Ratios for Stable and Generalizable Training
von: Liu, Chi, et al.
Veröffentlicht: (2026) -
Privileged Self-Access Matters for Introspection in AI
von: Song, Siyuan, et al.
Veröffentlicht: (2025)