DRA-GRPO: Your GRPO Needs to Know Diverse Reasoning Paths for Mathematical Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Xiwen, Zhu, Wenhui, Qiu, Peijie, Dong, Xuanzhao, Wang, Hao, Wu, Haiyu, Li, Huayu, Sotiras, Aristeidis, Wang, Yalin, Razi, Abolfazl |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DGR-MIL: Exploring Diverse Global Representation in Multiple Instance Learning for Whole Slide Image Classification
di: Zhu, Wenhui, et al.
Pubblicazione: (2024)
di: Zhu, Wenhui, et al.
Pubblicazione: (2024)
Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model Adaptation
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
TimeMIL: Advancing Multivariate Time Series Classification via a Time-aware Multiple Instance Learning
di: Chen, Xiwen, et al.
Pubblicazione: (2024)
di: Chen, Xiwen, et al.
Pubblicazione: (2024)
Sequence Complementor: Complementing Transformers For Time Series Forecasting with Learnable Sequences
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image Analysis
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
SelfReg-UNet: Self-Regularized UNet for Medical Image Segmentation
di: Zhu, Wenhui, et al.
Pubblicazione: (2024)
di: Zhu, Wenhui, et al.
Pubblicazione: (2024)
How Effective Can Dropout Be in Multiple Instance Learning ?
di: Zhu, Wenhui, et al.
Pubblicazione: (2025)
di: Zhu, Wenhui, et al.
Pubblicazione: (2025)
FIC-TSC: Learning Time Series Classification with Fisher Information Constraint
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
di: Zhu, Wenhui, et al.
Pubblicazione: (2025)
di: Zhu, Wenhui, et al.
Pubblicazione: (2025)
Multimodal Variational Autoencoder: a Barycentric View
di: Qiu, Peijie, et al.
Pubblicazione: (2024)
di: Qiu, Peijie, et al.
Pubblicazione: (2024)
SC-MIL: Sparsely Coded Multiple Instance Learning for Whole Slide Image Classification
di: Qiu, Peijie, et al.
Pubblicazione: (2023)
di: Qiu, Peijie, et al.
Pubblicazione: (2023)
Imaging Signal Recovery Using Neural Network Priors Under Uncertain Forward Model Parameters
di: Chen, Xiwen, et al.
Pubblicazione: (2024)
di: Chen, Xiwen, et al.
Pubblicazione: (2024)
STA-Unet: Rethink the semantic redundant for Medical Imaging Segmentation
di: Vasa, Vamsi Krishna, et al.
Pubblicazione: (2024)
di: Vasa, Vamsi Krishna, et al.
Pubblicazione: (2024)
OTPrune: Distribution-Aligned Visual Token Pruning via Optimal Transport
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
Talk Before You Retrieve: Agent-Led Discussions for Better RAG in Medical QA
di: Dong, Xuanzhao, et al.
Pubblicazione: (2025)
di: Dong, Xuanzhao, et al.
Pubblicazione: (2025)
Multimodal normative modeling in Alzheimers Disease with introspective variational autoencoders
di: Kumar, Sayantan, et al.
Pubblicazione: (2026)
di: Kumar, Sayantan, et al.
Pubblicazione: (2026)
Your Agent is More Brittle Than You Think: Uncovering Indirect Injection Vulnerabilities in Agentic LLMs
di: Zhu, Wenhui, et al.
Pubblicazione: (2026)
di: Zhu, Wenhui, et al.
Pubblicazione: (2026)
LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
di: Dong, Xuanzhao, et al.
Pubblicazione: (2025)
di: Dong, Xuanzhao, et al.
Pubblicazione: (2025)
AHA: Aligning Large Audio-Language Models for Reasoning Hallucinations via Counterfactual Hard Negatives
di: Chen, Yanxi, et al.
Pubblicazione: (2025)
di: Chen, Yanxi, et al.
Pubblicazione: (2025)
Many-MobileNet: Multi-Model Augmentation for Robust Retinal Disease Classification
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
GRPO and Reflection Reward for Mathematical Reasoning in Large Language Models
di: Wang, Zhijie
Pubblicazione: (2026)
di: Wang, Zhijie
Pubblicazione: (2026)
Fast 2DGS: Efficient Image Representation with Deep Gaussian Prior
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
EZBlender: Efficient 3D Editing with Plan-and-ReAct Agent
di: Wang, Hao, et al.
Pubblicazione: (2026)
di: Wang, Hao, et al.
Pubblicazione: (2026)
RBAD: A Dataset and Benchmark for Retinal Vessels Branching Angle Detection
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
Mags-RL: Wearing Multimodal LLMs a Magnifying Glass via Agentic Reinforcement Learning For Complex Scene Reasoning
di: Dong, Xuanzhao, et al.
Pubblicazione: (2026)
di: Dong, Xuanzhao, et al.
Pubblicazione: (2026)
PDL: Regularizing Multiple Instance Learning with Progressive Dropout Layers
di: Zhu, Wenhui, et al.
Pubblicazione: (2023)
di: Zhu, Wenhui, et al.
Pubblicazione: (2023)
Context Pruning for Coding Agents via Multi-Rubric Latent Reasoning
di: Wang, Jingjing, et al.
Pubblicazione: (2026)
di: Wang, Jingjing, et al.
Pubblicazione: (2026)
FLAME Diffuser: Wildfire Image Synthesis using Mask Guided Diffusion
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery
di: Yao, Nanjie, et al.
Pubblicazione: (2026)
di: Yao, Nanjie, et al.
Pubblicazione: (2026)
DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
di: Liu, Henglin, et al.
Pubblicazione: (2025)
di: Liu, Henglin, et al.
Pubblicazione: (2025)
ExGRPO: Learning to Reason from Experience
di: Zhan, Runzhe, et al.
Pubblicazione: (2025)
di: Zhan, Runzhe, et al.
Pubblicazione: (2025)
Diffusion Prism: Enhancing Diversity and Morphology Consistency in Mask-to-Image Diffusion
di: Wang, Hao, et al.
Pubblicazione: (2025)
di: Wang, Hao, et al.
Pubblicazione: (2025)
CUNSB-RFIE: Context-aware Unpaired Neural Schrödinger Bridge in Retinal Fundus Image Enhancement
di: Dong, Xuanzhao, et al.
Pubblicazione: (2024)
di: Dong, Xuanzhao, et al.
Pubblicazione: (2024)
Implicit Hierarchical GRPO: Decoupling Tool Invocation from Execution for Tool-Integrated Mathematical Reasoning
di: Wang, Li, et al.
Pubblicazione: (2026)
di: Wang, Li, et al.
Pubblicazione: (2026)
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
di: Dipta, Shubhashis Roy, et al.
Pubblicazione: (2026)
Boosting MLLM Reasoning with Text-Debiased Hint-GRPO
di: Huang, Qihan, et al.
Pubblicazione: (2025)
di: Huang, Qihan, et al.
Pubblicazione: (2025)
D-Net: Dynamic Large Kernel with Dynamic Feature Fusion for Volumetric Medical Image Segmentation
di: Yang, Jin, et al.
Pubblicazione: (2024)
di: Yang, Jin, et al.
Pubblicazione: (2024)
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
di: Chen, Xiwen, et al.
Pubblicazione: (2026)
Step-GRPO: Internalizing Dynamic Early Exit for Efficient Reasoning
di: Chen, Benteng, et al.
Pubblicazione: (2026)
di: Chen, Benteng, et al.
Pubblicazione: (2026)
FedMIL: Federated-Multiple Instance Learning for Video Analysis with Optimized DPP Scheduling
di: Bastola, Ashish, et al.
Pubblicazione: (2024)
di: Bastola, Ashish, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DGR-MIL: Exploring Diverse Global Representation in Multiple Instance Learning for Whole Slide Image Classification
di: Zhu, Wenhui, et al.
Pubblicazione: (2024) -
Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model Adaptation
di: Chen, Xiwen, et al.
Pubblicazione: (2025) -
TimeMIL: Advancing Multivariate Time Series Classification via a Time-aware Multiple Instance Learning
di: Chen, Xiwen, et al.
Pubblicazione: (2024) -
Sequence Complementor: Complementing Transformers For Time Series Forecasting with Learnable Sequences
di: Chen, Xiwen, et al.
Pubblicazione: (2025) -
Cracking Instance Jigsaw Puzzles: An Alternative to Multiple Instance Learning for Whole Slide Image Analysis
di: Chen, Xiwen, et al.
Pubblicazione: (2025)