Adaptive Reinforcement for Open-ended Medical Reasoning via Semantic-Guided Reward Collapse Mitigation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yizhou, Yang, Dingkang, Chen, Zizhi, Han, Minghao, Zhang, Xukun, Liu, Keliang, Wei, Jingwei, Zhang, Lihua |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
by: Chen, Zizhi, et al.
Published: (2025)
by: Chen, Zizhi, et al.
Published: (2025)
Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics
by: Han, Minghao, et al.
Published: (2024)
by: Han, Minghao, et al.
Published: (2024)
Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
by: Han, Minghao, et al.
Published: (2025)
by: Han, Minghao, et al.
Published: (2025)
VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining
by: Chen, Zizhi, et al.
Published: (2025)
by: Chen, Zizhi, et al.
Published: (2025)
Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
by: Liu, Keliang, et al.
Published: (2025)
by: Liu, Keliang, et al.
Published: (2025)
FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
by: Jiang, Yue, et al.
Published: (2025)
by: Jiang, Yue, et al.
Published: (2025)
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization
by: Han, Minghao, et al.
Published: (2026)
by: Han, Minghao, et al.
Published: (2026)
VGAT: A Cancer Survival Analysis Framework Transitioning from Generative Visual Question Answering to Genomic Reconstruction
by: Chen, Zizhi, et al.
Published: (2025)
by: Chen, Zizhi, et al.
Published: (2025)
MSCPT: Few-shot Whole Slide Image Classification with Multi-scale and Context-focused Prompt Tuning
by: Han, Minghao, et al.
Published: (2024)
by: Han, Minghao, et al.
Published: (2024)
Multi-Scale Heterogeneity-Aware Hypergraph Representation for Histopathology Whole Slide Images
by: Han, Minghao, et al.
Published: (2024)
by: Han, Minghao, et al.
Published: (2024)
Fusing Pixels and Genes: Spatially-Aware Learning in Computational Pathology
by: Han, Minghao, et al.
Published: (2026)
by: Han, Minghao, et al.
Published: (2026)
PersonaAnimator: Personalized Motion Transfer from Unconstrained Videos
by: Qian, Ziyun, et al.
Published: (2025)
by: Qian, Ziyun, et al.
Published: (2025)
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
by: Liu, Keliang, et al.
Published: (2025)
by: Liu, Keliang, et al.
Published: (2025)
Improving Multimodal Sentiment Analysis via Modality Optimization and Dynamic Primary Modality Selection
by: Yang, Dingkang, et al.
Published: (2025)
by: Yang, Dingkang, et al.
Published: (2025)
Skip and Skip: Segmenting Medical Images with Prompts
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
MaskBEV: Towards A Unified Framework for BEV Detection and Map Segmentation
by: Zhao, Xiao, et al.
Published: (2024)
by: Zhao, Xiao, et al.
Published: (2024)
Reward-Guided Semantic Evolution for Test-time Adaptive Object Detection
by: Zhou, Lihua, et al.
Published: (2026)
by: Zhou, Lihua, et al.
Published: (2026)
MedAide: Information Fusion and Anatomy of Medical Intents via LLM-based Agent Collaboration
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
The Choice of Divergence: A Neglected Key to Mitigating Diversity Collapse in Reinforcement Learning with Verifiable Reward
by: Li, Long, et al.
Published: (2025)
by: Li, Long, et al.
Published: (2025)
ProFocus: Proactive Perception and Focused Reasoning in Vision-and-Language Navigation
by: Xue, Wei, et al.
Published: (2026)
by: Xue, Wei, et al.
Published: (2026)
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
by: Ye, Zhiling, et al.
Published: (2025)
by: Ye, Zhiling, et al.
Published: (2025)
Can LLMs' Tuning Methods Work in Medical Multimodal Domain?
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
HybridOcc: NeRF Enhanced Transformer-based Multi-Camera 3D Occupancy Prediction
by: Zhao, Xiao, et al.
Published: (2024)
by: Zhao, Xiao, et al.
Published: (2024)
Semantic Reward Collapse and the Preservation of Epistemic Integrity in Adaptive AI Systems
by: Parris, William
Published: (2026)
by: Parris, William
Published: (2026)
Evaluating and Mitigating Social Bias for Large Language Models in Open-ended Settings
by: Liu, Zhao, et al.
Published: (2024)
by: Liu, Zhao, et al.
Published: (2024)
Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models
by: Chen, Jiawei, et al.
Published: (2024)
by: Chen, Jiawei, et al.
Published: (2024)
SatireDecoder: Visual Cascaded Decoupling for Enhancing Satirical Image Comprehension
by: Jiang, Yue, et al.
Published: (2025)
by: Jiang, Yue, et al.
Published: (2025)
Beyond Semantic Priors: Mitigating Optimization Collapse for Generalizable Visual Forensics
by: Liu, Jipeng, et al.
Published: (2026)
by: Liu, Jipeng, et al.
Published: (2026)
Adaptive Test-Time Reasoning via Reward-Guided Dual-Phase Search
by: Cui, Yingqian, et al.
Published: (2025)
by: Cui, Yingqian, et al.
Published: (2025)
CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
From Verifiable Dot to Reward Chain: Harnessing Verifiable Reference-based Rewards for Reinforcement Learning of Open-ended Generation
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
Reinforced Context Order Recovery for Adaptive Reasoning and Planning
by: Ma, Long, et al.
Published: (2025)
by: Ma, Long, et al.
Published: (2025)
Learn to Reason Efficiently with Adaptive Length-based Reward Shaping
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Adaptive Draft-Verification for Efficient Large Language Model Decoding
by: Liu, Xukun, et al.
Published: (2024)
by: Liu, Xukun, et al.
Published: (2024)
When Sharpening Becomes Collapse: Sampling Bias and Semantic Coupling in RL with Verifiable Rewards
by: Fan, Mingyuan, et al.
Published: (2026)
by: Fan, Mingyuan, et al.
Published: (2026)
MM-Snowball: Evaluating and Mitigating Hallucination Snowballing in Multimodal Multi-Turn Dialogue
by: Jiang, Yue, et al.
Published: (2026)
by: Jiang, Yue, et al.
Published: (2026)
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models
by: Zhang, Yuheng, et al.
Published: (2026)
by: Zhang, Yuheng, et al.
Published: (2026)
Stable Reasoning, Unstable Responses: Mitigating LLM Deception via Stability Asymmetry
by: Zhang, Guoxi, et al.
Published: (2026)
by: Zhang, Guoxi, et al.
Published: (2026)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
by: Fan, Lishui, et al.
Published: (2025)
by: Fan, Lishui, et al.
Published: (2025)
Towards Context-Aware Emotion Recognition Debiasing from a Causal Demystification Perspective via De-confounded Training
by: Yang, Dingkang, et al.
Published: (2024)
by: Yang, Dingkang, et al.
Published: (2024)
Similar Items
-
Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
by: Chen, Zizhi, et al.
Published: (2025) -
Towards Unified Molecule-Enhanced Pathology Image Representation Learning via Integrating Spatial Transcriptomics
by: Han, Minghao, et al.
Published: (2024) -
Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
by: Han, Minghao, et al.
Published: (2025) -
VLM-based Prompts as the Optimal Assistant for Unpaired Histopathology Virtual Staining
by: Chen, Zizhi, et al.
Published: (2025) -
Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
by: Liu, Keliang, et al.
Published: (2025)