Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jinghan, Fang, Junfeng, Lu, Jinda, Wang, Yuan, Guo, Xiaoyan, Zhang, Tianyu, Wang, Xiang, He, Xiangnan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
Boosting Few-Shot Learning via Attentive Feature Regularization
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Enhancing Tree Type Detection in Forest Fire Risk Assessment: Multi-Stage Approach and Color Encoding with Forest Fire Risk Evaluation Framework for UAV Imagery
von: Zhang, Jinda
Veröffentlicht: (2024)
von: Zhang, Jinda
Veröffentlicht: (2024)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Reliable Multi-Modal Object Re-Identification via Modality-Aware Graph Reasoning
von: Wan, Xixi, et al.
Veröffentlicht: (2025)
von: Wan, Xixi, et al.
Veröffentlicht: (2025)
Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts
von: He, Junwen, et al.
Veröffentlicht: (2024)
von: He, Junwen, et al.
Veröffentlicht: (2024)
Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Label What Matters: Modality-Balanced and Difficulty-Aware Multimodal Active Learning
von: Zeng, Yuqiao, et al.
Veröffentlicht: (2026)
von: Zeng, Yuqiao, et al.
Veröffentlicht: (2026)
EMIT: Enhancing MLLMs for Industrial Anomaly Detection via Difficulty-Aware GRPO
von: Guan, Wei, et al.
Veröffentlicht: (2025)
von: Guan, Wei, et al.
Veröffentlicht: (2025)
Enhance Image Classification via Inter-Class Image Mixup with Diffusion Model
von: Wang, Zhicai, et al.
Veröffentlicht: (2024)
von: Wang, Zhicai, et al.
Veröffentlicht: (2024)
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation
von: Lu, Jinda, et al.
Veröffentlicht: (2024)
von: Lu, Jinda, et al.
Veröffentlicht: (2024)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
von: Hua, Zhenglin, et al.
Veröffentlicht: (2025)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
von: Huang, Jiehui, et al.
Veröffentlicht: (2025)
von: Huang, Jiehui, et al.
Veröffentlicht: (2025)
PositionOCR: Augmenting Positional Awareness in Multi-Modal Models via Hybrid Specialist Integration
von: Duan, Chen, et al.
Veröffentlicht: (2026)
von: Duan, Chen, et al.
Veröffentlicht: (2026)
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
von: Hao, Yunzhuo, et al.
Veröffentlicht: (2025)
Adaptive Multi-Modal Cross-Entropy Loss for Stereo Matching
von: Xu, Peng, et al.
Veröffentlicht: (2023)
von: Xu, Peng, et al.
Veröffentlicht: (2023)
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
von: He, Jinghan, et al.
Veröffentlicht: (2026)
von: He, Jinghan, et al.
Veröffentlicht: (2026)
Accelerating Diffusion Transformer via Gradient-Optimized Cache
von: Qiu, Junxiang, et al.
Veröffentlicht: (2025)
von: Qiu, Junxiang, et al.
Veröffentlicht: (2025)
MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
von: Zheng, Xuhui, et al.
Veröffentlicht: (2025)
von: Zheng, Xuhui, et al.
Veröffentlicht: (2025)
Modality-Aware Infrared and Visible Image Fusion with Target-Aware Supervision
von: Sun, Tianyao, et al.
Veröffentlicht: (2025)
von: Sun, Tianyao, et al.
Veröffentlicht: (2025)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
von: Guo, Jiajie, et al.
Veröffentlicht: (2025)
von: Guo, Jiajie, et al.
Veröffentlicht: (2025)
GUI-C$^2$: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement Learning
von: Li, Junlong, et al.
Veröffentlicht: (2026)
von: Li, Junlong, et al.
Veröffentlicht: (2026)
SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
von: Sheng, Yuan, et al.
Veröffentlicht: (2025)
von: Sheng, Yuan, et al.
Veröffentlicht: (2025)
GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
von: Guo, Shasha, et al.
Veröffentlicht: (2025)
von: Guo, Shasha, et al.
Veröffentlicht: (2025)
Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning
von: Chen, Mingrui, et al.
Veröffentlicht: (2025)
von: Chen, Mingrui, et al.
Veröffentlicht: (2025)
A Cascading Cooperative Multi-agent Framework for On-ramp Merging Control Integrating Large Language Models
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
von: Zhang, Miao, et al.
Veröffentlicht: (2025)
CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
von: Zhang, Guanghao, et al.
Veröffentlicht: (2025)
von: Zhang, Guanghao, et al.
Veröffentlicht: (2025)
Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters
von: Wang, Yuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuan, et al.
Veröffentlicht: (2024)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
von: He, Jinghan, et al.
Veröffentlicht: (2024)
von: He, Jinghan, et al.
Veröffentlicht: (2024)
Accelerating Diffusion Transformer via Error-Optimized Cache
von: Qiu, Junxiang, et al.
Veröffentlicht: (2025)
von: Qiu, Junxiang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
von: Lu, Jinda, et al.
Veröffentlicht: (2025) -
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2025) -
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2026) -
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026) -
Boosting Few-Shot Learning via Attentive Feature Regularization
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)