Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Yiming, Xu, Yiran, Lin, Zicheng, Shi, Chufan, Chen, Yukang, Wang, Dingdong, Wu, Tianhe, Wang, Junjie, Yang, Yujiu, Qiao, Yu, Chu, Ruihang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
by: Ren, Yiming, et al.
Published: (2025)
by: Ren, Yiming, et al.
Published: (2025)
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
by: Ren, Yiming, et al.
Published: (2026)
by: Ren, Yiming, et al.
Published: (2026)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
by: Luo, Ruilin, et al.
Published: (2025)
by: Luo, Ruilin, et al.
Published: (2025)
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
by: Lin, Zicheng, et al.
Published: (2024)
by: Lin, Zicheng, et al.
Published: (2024)
LiFi: Lightweight Controlled Text Generation with Fine-Grained Control Codes
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
by: Chen, Yuqing, et al.
Published: (2025)
by: Chen, Yuqing, et al.
Published: (2025)
Velocity-Space 3D Asset Editing
by: Liu, Hao, et al.
Published: (2026)
by: Liu, Hao, et al.
Published: (2026)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
ContextVis: Envision Contextual Learning and Interaction with Generative Models
by: Shui, Bo, et al.
Published: (2024)
by: Shui, Bo, et al.
Published: (2024)
GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy Entropy
by: Tan, Hongze, et al.
Published: (2025)
by: Tan, Hongze, et al.
Published: (2025)
MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery
by: Yao, Nanjie, et al.
Published: (2026)
by: Yao, Nanjie, et al.
Published: (2026)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025)
by: Ding, Yang, et al.
Published: (2025)
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
by: Luo, Ruilin, et al.
Published: (2026)
by: Luo, Ruilin, et al.
Published: (2026)
Advances in Engineered Virus‐Like Particles for Applications in Nanomedicine
by: Qingxia Shi, et al.
Published: (2026)
by: Qingxia Shi, et al.
Published: (2026)
Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
by: Yang, Xuewei, et al.
Published: (2026)
by: Yang, Xuewei, et al.
Published: (2026)
Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
by: Chu, Ruihang, et al.
Published: (2025)
by: Chu, Ruihang, et al.
Published: (2025)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
by: Liu, Henglin, et al.
Published: (2025)
by: Liu, Henglin, et al.
Published: (2025)
DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
by: Chen, Xiwen, et al.
Published: (2025)
by: Chen, Xiwen, et al.
Published: (2025)
A Comprehensive Study of Multimodal Large Language Models for Image Quality Assessment
by: Wu, Tianhe, et al.
Published: (2024)
by: Wu, Tianhe, et al.
Published: (2024)
PTD-SQL: Partitioning and Targeted Drilling with LLMs in Text-to-SQL
by: Luo, Ruilin, et al.
Published: (2024)
by: Luo, Ruilin, et al.
Published: (2024)
MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation
by: Ma, Xiaoxiao, et al.
Published: (2026)
by: Ma, Xiaoxiao, et al.
Published: (2026)
InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions
by: Wang, Yifan, et al.
Published: (2024)
by: Wang, Yifan, et al.
Published: (2024)
Hint-enhanced In-Context Learning wakes Large Language Models up for knowledge-intensive tasks
by: Wang, Yifan, et al.
Published: (2023)
by: Wang, Yifan, et al.
Published: (2023)
A Thorough Examination of Decoding Methods in the Era of LLMs
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
EDGE-GRPO: Entropy-Driven GRPO with Guided Error Correction for Advantage Diversity
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
TTM-RE: Memory-Augmented Document-Level Relation Extraction
by: Gao, Chufan, et al.
Published: (2024)
by: Gao, Chufan, et al.
Published: (2024)
Nav-$R^2$ Dual-Relation Reasoning for Generalizable Open-Vocabulary Object-Goal Navigation
by: Xiang, Wentao, et al.
Published: (2025)
by: Xiang, Wentao, et al.
Published: (2025)
LLM2: Let Large Language Models Harness System 2 Reasoning
by: Yang, Cheng, et al.
Published: (2024)
by: Yang, Cheng, et al.
Published: (2024)
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization
by: Chen, Minghan, et al.
Published: (2025)
by: Chen, Minghan, et al.
Published: (2025)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
by: Yang, Jialiang, et al.
Published: (2026)
by: Yang, Jialiang, et al.
Published: (2026)
Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance
by: Yu, Jiachen, et al.
Published: (2026)
by: Yu, Jiachen, et al.
Published: (2026)
Computer‐Aided Technology for Bioactive Protein Design and Clinical Application
by: Chufan Wang, et al.
Published: (2025)
by: Chufan Wang, et al.
Published: (2025)
STUDY ON PREPARATION AND PROPERTIES OF PHENOL-FORMALDEHYDE-CHINESE FIR LIQUEFACTION COPOLYMER RESIN
by: Ruihang Lin
Published: (2014)
by: Ruihang Lin
Published: (2014)
LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
by: Li, Yu, et al.
Published: (2024)
by: Li, Yu, et al.
Published: (2024)
Video-Zero: Self-Evolution Video Understanding
by: Zhang, Ruixu, et al.
Published: (2026)
by: Zhang, Ruixu, et al.
Published: (2026)
Generative Universal Verifier as Multimodal Meta-Reasoner
by: Zhang, Xinchen, et al.
Published: (2025)
by: Zhang, Xinchen, et al.
Published: (2025)
TTS-VAR: A Test-Time Scaling Framework for Visual Auto-Regressive Generation
by: Chen, Zhekai, et al.
Published: (2025)
by: Chen, Zhekai, et al.
Published: (2025)
Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast
by: Shi, Chufan, et al.
Published: (2024)
by: Shi, Chufan, et al.
Published: (2024)
HoLLMwood: Unleashing the Creativity of Large Language Models in Screenwriting via Role Playing
by: Chen, Jing, et al.
Published: (2024)
by: Chen, Jing, et al.
Published: (2024)
Similar Items
-
AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
by: Ren, Yiming, et al.
Published: (2025) -
Mitigating the Reasoning Tax in Vision-Language Fine-Tuning with Input-Adaptive Depth Aggregation
by: Ren, Yiming, et al.
Published: (2026) -
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
by: Ren, Yiming, et al.
Published: (2026) -
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
by: Luo, Ruilin, et al.
Published: (2025) -
Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability
by: Lin, Zicheng, et al.
Published: (2024)