Gespeichert in:
| Hauptverfasser: | Simoni, Marco, Fontana, Aleksandar, Rossolini, Giulio, Saracino, Andrea, Mori, Paolo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2508.03772 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
On the Hidden Objective Biases of Group-based Reinforcement Learning
von: Fontana, Aleksandar, et al.
Veröffentlicht: (2026)
von: Fontana, Aleksandar, et al.
Veröffentlicht: (2026)
Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
TITAN: Graph-Executable Reasoning for Cyber Threat Intelligence
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
KGQuest: Template-Driven QA Generation from Knowledge Graphs with LLM-Based Refinement
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
von: Nayab, Sania, et al.
Veröffentlicht: (2024)
von: Nayab, Sania, et al.
Veröffentlicht: (2024)
Leveraging Knowledge Graphs and LLMs for Structured Generation of Misinformation
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
How Worst-Case Are Adversarial Attacks? Linking Adversarial and Perturbation Robustness
von: Rossolini, Giulio
Veröffentlicht: (2026)
von: Rossolini, Giulio
Veröffentlicht: (2026)
EBPO: Empirical Bayes Shrinkage for Stabilizing Group-Relative Policy Optimization
von: Han, Kevin, et al.
Veröffentlicht: (2026)
von: Han, Kevin, et al.
Veröffentlicht: (2026)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
von: Tan, Hongze, et al.
Veröffentlicht: (2025)
Personalized Group Relative Policy Optimization for Heterogenous Preference Alignment
von: Wang, Jialu, et al.
Veröffentlicht: (2026)
von: Wang, Jialu, et al.
Veröffentlicht: (2026)
Increasing the Confidence of Deep Neural Networks by Coverage Analysis
von: Rossolini, Giulio, et al.
Veröffentlicht: (2021)
von: Rossolini, Giulio, et al.
Veröffentlicht: (2021)
Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
von: Zhang, Xichen, et al.
Veröffentlicht: (2025)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
Group Sequence Policy Optimization
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
von: Zheng, Chujie, et al.
Veröffentlicht: (2025)
Leveraging Group Relative Policy Optimization to Advance Large Language Models in Traditional Chinese Medicine
von: Xie, Jiacheng, et al.
Veröffentlicht: (2025)
von: Xie, Jiacheng, et al.
Veröffentlicht: (2025)
Flexible Entropy Control in RLVR with a Gradient-Preserving Perspective
von: Chen, Kun, et al.
Veröffentlicht: (2026)
von: Chen, Kun, et al.
Veröffentlicht: (2026)
Entropy Controllable Direct Preference Optimization
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
von: Omura, Motoki, et al.
Veröffentlicht: (2024)
Klear-Reasoner: Advancing Reasoning Capability via Gradient-Preserving Clipping Policy Optimization
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
von: Su, Zhenpeng, et al.
Veröffentlicht: (2025)
BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2025)
Agentic Entropy-Balanced Policy Optimization
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
von: Dong, Guanting, et al.
Veröffentlicht: (2025)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
von: Li, Changming, et al.
Veröffentlicht: (2026)
von: Li, Changming, et al.
Veröffentlicht: (2026)
GRAPH-GRPO-LEX: Contract Graph Modeling and Reinforcement Learning with Group Relative Policy Optimization
von: Dechtiar, Moriya, et al.
Veröffentlicht: (2025)
von: Dechtiar, Moriya, et al.
Veröffentlicht: (2025)
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
von: Liu, Shih-Yang, et al.
Veröffentlicht: (2026)
C$^2$GSPG: Confidence-calibrated Group Sequence Policy Gradient towards Self-aware Reasoning
von: Liu, Haotian, et al.
Veröffentlicht: (2025)
von: Liu, Haotian, et al.
Veröffentlicht: (2025)
PPO-BR: Dual-Signal Entropy-Reward Adaptation for Trust Region Policy Optimization
von: Rahman, Ben
Veröffentlicht: (2025)
von: Rahman, Ben
Veröffentlicht: (2025)
Exploiting Edge Features for Transferable Adversarial Attacks in Distributed Machine Learning
von: Rossolini, Giulio, et al.
Veröffentlicht: (2025)
von: Rossolini, Giulio, et al.
Veröffentlicht: (2025)
Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
von: Li, Chengao, et al.
Veröffentlicht: (2025)
von: Li, Chengao, et al.
Veröffentlicht: (2025)
Agentic Policy Optimization via Instruction-Policy Co-Evolution
von: Zhou, Han, et al.
Veröffentlicht: (2025)
von: Zhou, Han, et al.
Veröffentlicht: (2025)
Regressing the Relative Future: Efficient Policy Optimization for Multi-turn RLHF
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
von: Gao, Zhaolin, et al.
Veröffentlicht: (2024)
GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
von: Mao, Hangyu, et al.
Veröffentlicht: (2025)
EP-GRPO: Entropy-Progress Aligned Group Relative Policy Optimization with Implicit Process Guidance
von: Yu, Song, et al.
Veröffentlicht: (2026)
von: Yu, Song, et al.
Veröffentlicht: (2026)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
von: Ding, Yifeng, et al.
Veröffentlicht: (2025)
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
Group-Relative REINFORCE Is Secretly an Off-Policy Algorithm: Demystifying Some Myths About GRPO and Its Friends
von: Yao, Chaorui, et al.
Veröffentlicht: (2025)
von: Yao, Chaorui, et al.
Veröffentlicht: (2025)
Optimize Weight Rounding via Signed Gradient Descent for the Quantization of LLMs
von: Cheng, Wenhua, et al.
Veröffentlicht: (2023)
von: Cheng, Wenhua, et al.
Veröffentlicht: (2023)
Seek in the Dark: Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space
von: Li, Hengli, et al.
Veröffentlicht: (2025)
von: Li, Hengli, et al.
Veröffentlicht: (2025)
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization
von: Zhang, Jingyi, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2025)
Simple Policy Gradients for Reasoning with Diffusion Language Models
von: Zhan, Anthony
Veröffentlicht: (2025)
von: Zhan, Anthony
Veröffentlicht: (2025)
Fine-Tuning Discrete Diffusion Models with Policy Gradient Methods
von: Zekri, Oussama, et al.
Veröffentlicht: (2025)
von: Zekri, Oussama, et al.
Veröffentlicht: (2025)
On the Design of KL-Regularized Policy Gradient Algorithms for LLM Reasoning
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
von: Zhang, Yifan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
On the Hidden Objective Biases of Group-based Reinforcement Learning
von: Fontana, Aleksandar, et al.
Veröffentlicht: (2026) -
Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization
von: Simoni, Marco, et al.
Veröffentlicht: (2025) -
TITAN: Graph-Executable Reasoning for Cyber Threat Intelligence
von: Simoni, Marco, et al.
Veröffentlicht: (2025) -
KGQuest: Template-Driven QA Generation from Knowledge Graphs with LLM-Based Refinement
von: Nayab, Sania, et al.
Veröffentlicht: (2025) -
Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
von: Nayab, Sania, et al.
Veröffentlicht: (2024)