On the Hidden Objective Biases of Group-based Reinforcement Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fontana, Aleksandar, Simoni, Marco, Rossolini, Giulio, Saracino, Andrea, Mori, Paolo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
TITAN: Graph-Executable Reasoning for Cyber Threat Intelligence
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
von: Simoni, Marco, et al.
Veröffentlicht: (2025)
KGQuest: Template-Driven QA Generation from Knowledge Graphs with LLM-Based Refinement
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
von: Nayab, Sania, et al.
Veröffentlicht: (2024)
von: Nayab, Sania, et al.
Veröffentlicht: (2024)
Leveraging Knowledge Graphs and LLMs for Structured Generation of Misinformation
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
von: Nayab, Sania, et al.
Veröffentlicht: (2025)
How Worst-Case Are Adversarial Attacks? Linking Adversarial and Perturbation Robustness
von: Rossolini, Giulio
Veröffentlicht: (2026)
von: Rossolini, Giulio
Veröffentlicht: (2026)
Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
von: Luo, Haozheng, et al.
Veröffentlicht: (2026)
Inductive Biases for Zero-shot Systematic Generalization in Language-informed Reinforcement Learning
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
von: Dijujin, Negin Hashemi, et al.
Veröffentlicht: (2025)
Large Language Models are Biased Reinforcement Learners
von: Hayes, William M., et al.
Veröffentlicht: (2024)
von: Hayes, William M., et al.
Veröffentlicht: (2024)
RLP: Reinforcement as a Pretraining Objective
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
von: Hatamizadeh, Ali, et al.
Veröffentlicht: (2025)
Alignment Tampering: How Reinforcement Learning from Human Feedback Is Exploited to Optimize Misaligned Biases
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
von: Hahm, Dongyoon, et al.
Veröffentlicht: (2026)
EMORL: Ensemble Multi-Objective Reinforcement Learning for Efficient and Flexible LLM Fine-Tuning
von: Kong, Lingxiao, et al.
Veröffentlicht: (2025)
von: Kong, Lingxiao, et al.
Veröffentlicht: (2025)
Group Distributionally Robust Optimization-Driven Reinforcement Learning for LLM Reasoning
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
von: Panaganti, Kishan, et al.
Veröffentlicht: (2026)
Group-Aware Reinforcement Learning for Output Diversity in Large Language Models
von: Anschel, Oron, et al.
Veröffentlicht: (2025)
von: Anschel, Oron, et al.
Veröffentlicht: (2025)
Reinforce-Ada: An Adaptive Sampling Framework under Non-linear RL Objectives
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
von: Xiong, Wei, et al.
Veröffentlicht: (2025)
FairFlow: Mitigating Dataset Biases through Undecided Learning
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
von: Cheng, Jiali, et al.
Veröffentlicht: (2025)
Do Large Language Models Show Biases in Causal Learning?
von: Carro, Maria Victoria, et al.
Veröffentlicht: (2024)
von: Carro, Maria Victoria, et al.
Veröffentlicht: (2024)
Increasing the Confidence of Deep Neural Networks by Coverage Analysis
von: Rossolini, Giulio, et al.
Veröffentlicht: (2021)
von: Rossolini, Giulio, et al.
Veröffentlicht: (2021)
Learning to (Learn at Test Time): RNNs with Expressive Hidden States
von: Sun, Yu, et al.
Veröffentlicht: (2024)
von: Sun, Yu, et al.
Veröffentlicht: (2024)
GRADIEND: Feature Learning within Neural Networks Exemplified through Biases
von: Drechsel, Jonathan, et al.
Veröffentlicht: (2025)
von: Drechsel, Jonathan, et al.
Veröffentlicht: (2025)
Multi-Objective Reinforcement Learning for Large Language Model Optimization: Visionary Perspective
von: Kong, Lingxiao, et al.
Veröffentlicht: (2025)
von: Kong, Lingxiao, et al.
Veröffentlicht: (2025)
Hidden State Poisoning Attacks against Mamba-based Language Models
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
von: Mercier, Alexandre Le, et al.
Veröffentlicht: (2026)
Limits of Transformer Language Models on Learning to Compose Algorithms
von: Thomm, Jonathan, et al.
Veröffentlicht: (2024)
von: Thomm, Jonathan, et al.
Veröffentlicht: (2024)
Relative Value Biases in Large Language Models
von: Hayes, William M., et al.
Veröffentlicht: (2024)
von: Hayes, William M., et al.
Veröffentlicht: (2024)
Efficient Preference-based Reinforcement Learning via Aligned Experience Estimation
von: Bai, Fengshuo, et al.
Veröffentlicht: (2024)
von: Bai, Fengshuo, et al.
Veröffentlicht: (2024)
Conditional Language Policy: A General Framework for Steerable Multi-Objective Finetuning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2024)
Efficient Reasoning with Hidden Thinking
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
Learning to Hint for Reinforcement Learning
von: Xia, Yu, et al.
Veröffentlicht: (2026)
von: Xia, Yu, et al.
Veröffentlicht: (2026)
Exploiting Synergistic Cognitive Biases to Bypass Safety in LLMs
von: Yang, Xikang, et al.
Veröffentlicht: (2025)
von: Yang, Xikang, et al.
Veröffentlicht: (2025)
Self-Speculative Biased Decoding for Faster Re-Translation
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
von: Zeng, Linxiao, et al.
Veröffentlicht: (2025)
More is not always better? Enhancing Many-Shot In-Context Learning with Differentiated and Reweighting Objectives
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoqing, et al.
Veröffentlicht: (2025)
Bridging Kolmogorov Complexity and Deep Learning: Asymptotically Optimal Description Length Objectives for Transformers
von: Shaw, Peter, et al.
Veröffentlicht: (2025)
von: Shaw, Peter, et al.
Veröffentlicht: (2025)
Heuristics and Biases in AI Decision-Making: Implications for Responsible AGI
von: Saeedi, Payam, et al.
Veröffentlicht: (2024)
von: Saeedi, Payam, et al.
Veröffentlicht: (2024)
Reward-free Alignment for Conflicting Objectives
von: Chen, Peter, et al.
Veröffentlicht: (2026)
von: Chen, Peter, et al.
Veröffentlicht: (2026)
Beyond RAG: Task-Aware KV Cache Compression for Comprehensive Knowledge Reasoning
von: Corallo, Giulio, et al.
Veröffentlicht: (2025)
von: Corallo, Giulio, et al.
Veröffentlicht: (2025)
Exploiting Edge Features for Transferable Adversarial Attacks in Distributed Machine Learning
von: Rossolini, Giulio, et al.
Veröffentlicht: (2025)
von: Rossolini, Giulio, et al.
Veröffentlicht: (2025)
Reinforcement Learning with Backtracking Feedback
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
von: Sel, Bilgehan, et al.
Veröffentlicht: (2026)
Natural Language Reinforcement Learning
von: Feng, Xidong, et al.
Veröffentlicht: (2024)
von: Feng, Xidong, et al.
Veröffentlicht: (2024)
Reinforcement Learning with Rubric Anchors
von: Huang, Zenan, et al.
Veröffentlicht: (2025)
von: Huang, Zenan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
von: Simoni, Marco, et al.
Veröffentlicht: (2025) -
Improving LLM Reasoning for Vulnerability Detection via Group Relative Policy Optimization
von: Simoni, Marco, et al.
Veröffentlicht: (2025) -
TITAN: Graph-Executable Reasoning for Cyber Threat Intelligence
von: Simoni, Marco, et al.
Veröffentlicht: (2025) -
KGQuest: Template-Driven QA Generation from Knowledge Graphs with LLM-Based Refinement
von: Nayab, Sania, et al.
Veröffentlicht: (2025) -
Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost
von: Nayab, Sania, et al.
Veröffentlicht: (2024)