VeriGate: Verifier-Gated Step-Level Supervision for GRPO
Fuente:
arXiv
Salvato in:
| Autori principali: | Agrawal, Aakriti, Liu, Minghui, Huang, Furong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
GATES: Self-Distillation under Privileged Context with Consensus Gating
di: Stein, Alex, et al.
Pubblicazione: (2026)
di: Stein, Alex, et al.
Pubblicazione: (2026)
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025)
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025)
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
di: Agrawal, Aakriti, et al.
Pubblicazione: (2024)
di: Agrawal, Aakriti, et al.
Pubblicazione: (2024)
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
di: Wang, Jingyi, et al.
Pubblicazione: (2026)
di: Wang, Jingyi, et al.
Pubblicazione: (2026)
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025)
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025)
PoisonedParrot: Subtle Data Poisoning Attacks to Elicit Copyright-Infringing Content from Large Language Models
di: Panaitescu-Liess, Michael-Andrei, et al.
Pubblicazione: (2025)
di: Panaitescu-Liess, Michael-Andrei, et al.
Pubblicazione: (2025)
GatedFWA: Linear Flash Windowed Attention with Gated Associative Memory
di: Liu, Jiaxu, et al.
Pubblicazione: (2025)
di: Liu, Jiaxu, et al.
Pubblicazione: (2025)
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?
di: Ho, Sy-Tuyen, et al.
Pubblicazione: (2026)
di: Ho, Sy-Tuyen, et al.
Pubblicazione: (2026)
FineGates: LLMs Finetuning with Compression using Stochastic Gates
di: Svirsky, Jonathan, et al.
Pubblicazione: (2024)
di: Svirsky, Jonathan, et al.
Pubblicazione: (2024)
MAGIC: Multi-Step Advantage-Gated Causal Influence for Multi-agent Reinforcement Learning
di: Yu, Haohan, et al.
Pubblicazione: (2026)
di: Yu, Haohan, et al.
Pubblicazione: (2026)
Reinforcement Learning with Verifiable Rewards: GRPO's Effective Loss, Dynamics, and Success Amplification
di: Mroueh, Youssef
Pubblicazione: (2025)
di: Mroueh, Youssef
Pubblicazione: (2025)
SigGate: Enhancing Recurrent Neural Networks with Signature-Based Gating Mechanisms
di: Genet, Rémi, et al.
Pubblicazione: (2025)
di: Genet, Rémi, et al.
Pubblicazione: (2025)
Sigmoid Gating is More Sample Efficient than Softmax Gating in Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
di: Nguyen, Huy, et al.
Pubblicazione: (2024)
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
di: Jia, Zeyu, et al.
Pubblicazione: (2025)
ArcGate: Adaptive Arctangent Gated Activation
di: Bhattacharya, Avik, et al.
Pubblicazione: (2026)
di: Bhattacharya, Avik, et al.
Pubblicazione: (2026)
VeriThinker: Learning to Verify Makes Reasoning Model Efficient
di: Chen, Zigeng, et al.
Pubblicazione: (2025)
di: Chen, Zigeng, et al.
Pubblicazione: (2025)
Gaussian Process-Gated Hierarchical Mixtures of Experts
di: Liu, Yuhao, et al.
Pubblicazione: (2023)
di: Liu, Yuhao, et al.
Pubblicazione: (2023)
Uncertainty-Gated Generative Modeling
di: Gu, Xingrui, et al.
Pubblicazione: (2026)
di: Gu, Xingrui, et al.
Pubblicazione: (2026)
Differential Gated Self-Attention
di: Lygizou, Elpiniki Maria, et al.
Pubblicazione: (2025)
di: Lygizou, Elpiniki Maria, et al.
Pubblicazione: (2025)
AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
di: Huang, Runhui, et al.
Pubblicazione: (2026)
di: Huang, Runhui, et al.
Pubblicazione: (2026)
WS-GRPO: Weakly-Supervised Group-Relative Policy Optimization for Rollout-Efficient Reasoning
di: Mundada, Gagan, et al.
Pubblicazione: (2026)
di: Mundada, Gagan, et al.
Pubblicazione: (2026)
Expanded Gating Ranges Improve Activation Functions
di: Huang, Allen Hao
Pubblicazione: (2024)
di: Huang, Allen Hao
Pubblicazione: (2024)
Spectral Gating Networks
di: Zhang, Jusheng, et al.
Pubblicazione: (2026)
di: Zhang, Jusheng, et al.
Pubblicazione: (2026)
Gated-SwinRMT: Unifying Swin Windowed Attention with Retentive Manhattan Decay via Input-Dependent Gating
di: Maity, Dipan, et al.
Pubblicazione: (2026)
di: Maity, Dipan, et al.
Pubblicazione: (2026)
Imagine, Verify, Execute: Memory-guided Agentic Exploration with Vision-Language Models
di: Lee, Seungjae, et al.
Pubblicazione: (2025)
di: Lee, Seungjae, et al.
Pubblicazione: (2025)
Easy2Hard-Bench: Standardized Difficulty Labels for Profiling LLM Performance and Generalization
di: Ding, Mucong, et al.
Pubblicazione: (2024)
di: Ding, Mucong, et al.
Pubblicazione: (2024)
Masked Gated Linear Unit
di: Tajima, Yukito, et al.
Pubblicazione: (2025)
di: Tajima, Yukito, et al.
Pubblicazione: (2025)
LithoGRPO: Fast Inverse Lithography via GRPO Reinforced Flow Matching
di: Lai, Yao, et al.
Pubblicazione: (2026)
di: Lai, Yao, et al.
Pubblicazione: (2026)
Mixture-of-Experts for Distributed Edge Computing with Channel-Aware Gating Function
di: Song, Qiuchen, et al.
Pubblicazione: (2025)
di: Song, Qiuchen, et al.
Pubblicazione: (2025)
Entropy-Gated Selective Policy Optimization:Token-Level Gradient Allocation for Hybrid Training of Large Language Models
di: Hu, Yuelin, et al.
Pubblicazione: (2026)
di: Hu, Yuelin, et al.
Pubblicazione: (2026)
Light Differentiable Logic Gate Networks
di: Rüttgers, Lukas, et al.
Pubblicazione: (2025)
di: Rüttgers, Lukas, et al.
Pubblicazione: (2025)
Gated Graph Attention Networks with Learnable Temperature
di: Ma, Zhongtian, et al.
Pubblicazione: (2026)
di: Ma, Zhongtian, et al.
Pubblicazione: (2026)
Remember to Forget: Gated Adaptive Positional Encoding
di: Ali, Riccardo, et al.
Pubblicazione: (2026)
di: Ali, Riccardo, et al.
Pubblicazione: (2026)
Recurrent Deep Differentiable Logic Gate Networks
di: Bührer, Simon, et al.
Pubblicazione: (2025)
di: Bührer, Simon, et al.
Pubblicazione: (2025)
Convergence Rates for Softmax Gating Mixture of Experts
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
di: Nguyen, Huy, et al.
Pubblicazione: (2025)
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
di: Ren, Yiming, et al.
Pubblicazione: (2026)
di: Ren, Yiming, et al.
Pubblicazione: (2026)
WAVES: Benchmarking the Robustness of Image Watermarks
di: An, Bang, et al.
Pubblicazione: (2024)
di: An, Bang, et al.
Pubblicazione: (2024)
VeriContest: A Competitive-Programming Benchmark for Verifiable Code Generation
di: Xie, Zichen, et al.
Pubblicazione: (2026)
di: Xie, Zichen, et al.
Pubblicazione: (2026)
Rethinking Gating Mechanism in Sparse MoE: Handling Arbitrary Modality Inputs with Confidence-Guided Gate
di: Zheng, Liangwei Nathan, et al.
Pubblicazione: (2025)
di: Zheng, Liangwei Nathan, et al.
Pubblicazione: (2025)
VeriScale: Adversarial Test-Suite Scaling for Verifiable Code Generation
di: Bai, Yifan, et al.
Pubblicazione: (2026)
di: Bai, Yifan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
GATES: Self-Distillation under Privileged Context with Consensus Gating
di: Stein, Alex, et al.
Pubblicazione: (2026) -
Uncertainty-Aware Answer Selection for Improved Reasoning in Multi-LLM Systems
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025) -
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
di: Agrawal, Aakriti, et al.
Pubblicazione: (2024) -
GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
di: Wang, Jingyi, et al.
Pubblicazione: (2026) -
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
di: Agrawal, Aakriti, et al.
Pubblicazione: (2025)