VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Fuente:
arXiv
Saved in:
| Main Authors: | Shen, Haozhan, Liu, Peng, Li, Jingcheng, Fang, Chunxin, Ma, Yibo, Liao, Jiajia, Shen, Qiaoli, Zhang, Zilun, Zhao, Kangjia, Zhang, Qianqian, Xu, Ruochen, Zhao, Tiancheng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
by: Shen, Haozhan, et al.
Published: (2024)
by: Shen, Haozhan, et al.
Published: (2024)
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research
by: Zhang, Qianqian, et al.
Published: (2025)
by: Zhang, Qianqian, et al.
Published: (2025)
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
by: Shen, Haozhan, et al.
Published: (2026)
by: Shen, Haozhan, et al.
Published: (2026)
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
by: Zhao, Tiancheng, et al.
Published: (2024)
by: Zhao, Tiancheng, et al.
Published: (2024)
Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning
by: Zhang, Zilun, et al.
Published: (2025)
by: Zhang, Zilun, et al.
Published: (2025)
RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
by: Zhang, Zilun, et al.
Published: (2023)
by: Zhang, Zilun, et al.
Published: (2023)
Talking to Yourself: Defying Forgetting in Large Language Models
by: Sun, Yutao, et al.
Published: (2026)
by: Sun, Yutao, et al.
Published: (2026)
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing
by: Zhang, Zilun, et al.
Published: (2025)
by: Zhang, Zilun, et al.
Published: (2025)
The Self-Improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities without External Scaffolding?
by: Sun, Yutao, et al.
Published: (2025)
by: Sun, Yutao, et al.
Published: (2025)
Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression
by: Zhang, Zilun, et al.
Published: (2024)
by: Zhang, Zilun, et al.
Published: (2024)
DetailCLIP: Injecting Image Details into CLIP's Feature Space
by: Zhang, Zilun, et al.
Published: (2022)
by: Zhang, Zilun, et al.
Published: (2022)
ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG
by: Zhang, Zilun, et al.
Published: (2024)
by: Zhang, Zilun, et al.
Published: (2024)
MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
by: Shen, Haozhan, et al.
Published: (2026)
by: Shen, Haozhan, et al.
Published: (2026)
Med-R1: Reinforcement Learning for Generalizable Medical Reasoning in Vision-Language Models
by: Lai, Yuxiang, et al.
Published: (2025)
by: Lai, Yuxiang, et al.
Published: (2025)
GUI Testing Arena: A Unified Benchmark for Advancing Autonomous GUI Testing Agent
by: Zhao, Kangjia, et al.
Published: (2024)
by: Zhao, Kangjia, et al.
Published: (2024)
CrowdVLM-R1: Expanding R1 Ability to Vision Language Model for Crowd Counting using Fuzzy Group Relative Policy Reward
by: Wang, Zhiqiang, et al.
Published: (2025)
by: Wang, Zhiqiang, et al.
Published: (2025)
SRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation
by: Guo, Yulong, et al.
Published: (2025)
by: Guo, Yulong, et al.
Published: (2025)
OmAgent: A Multi-modal Agent Framework for Complex Video Understanding with Task Divide-and-Conquer
by: Zhang, Lu, et al.
Published: (2024)
by: Zhang, Lu, et al.
Published: (2024)
Unilaw-R1: A Large Language Model for Legal Reasoning with Reinforcement Learning and Iterative Inference
by: Cai, Hua, et al.
Published: (2025)
by: Cai, Hua, et al.
Published: (2025)
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
by: Fan, Zhiwen, et al.
Published: (2025)
by: Fan, Zhiwen, et al.
Published: (2025)
R-VLM: Region-Aware Vision Language Model for Precise GUI Grounding
by: Park, Joonhyung, et al.
Published: (2025)
by: Park, Joonhyung, et al.
Published: (2025)
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
by: Tian, Beitong, et al.
Published: (2025)
by: Tian, Beitong, et al.
Published: (2025)
Bridge Thinking and Acting: Unleashing Physical Potential of VLM with Generalizable Action Expert
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
Skywork-R1V3 Technical Report
by: Shen, Wei, et al.
Published: (2025)
by: Shen, Wei, et al.
Published: (2025)
MFVLR: Multi-domain Fine-grained Vision-Language Reconstruction for Generalizable Diffusion Face Forgery Detection and Localization
by: Zhang, Yaning, et al.
Published: (2026)
by: Zhang, Yaning, et al.
Published: (2026)
GephiForR: An R package for creating Gephi-style network visualizations
by: Manso, Julia
Published: (2024)
by: Manso, Julia
Published: (2024)
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
by: Zhai, Skylar, et al.
Published: (2026)
by: Zhai, Skylar, et al.
Published: (2026)
VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
by: Qi, Zhangyang, et al.
Published: (2025)
by: Qi, Zhangyang, et al.
Published: (2025)
HanMoVLM: Large Vision-Language Models for Professional Artistic Painting Evaluation
by: Yang, Hongji, et al.
Published: (2026)
by: Yang, Hongji, et al.
Published: (2026)
DriveAgent-R1: Advancing VLM-based Autonomous Driving with Active Perception and Hybrid Thinking
by: Zheng, Weicheng, et al.
Published: (2025)
by: Zheng, Weicheng, et al.
Published: (2025)
VTC-R1: Vision-Text Compression for Efficient Long-Context Reasoning
by: Wang, Yibo, et al.
Published: (2026)
by: Wang, Yibo, et al.
Published: (2026)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
by: Zhao, Tiancheng, et al.
Published: (2022)
by: Zhao, Tiancheng, et al.
Published: (2022)
PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks
by: Wu, Jianyu, et al.
Published: (2025)
by: Wu, Jianyu, et al.
Published: (2025)
R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
by: Yao, Huanjin, et al.
Published: (2025)
by: Yao, Huanjin, et al.
Published: (2025)
FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
by: Cai, Hengxing, et al.
Published: (2025)
by: Cai, Hengxing, et al.
Published: (2025)
Generalizable Entity Grounding via Assistance of Large Language Model
by: Qi, Lu, et al.
Published: (2024)
by: Qi, Lu, et al.
Published: (2024)
Stable Knowledge Editing in Large Language Models
by: Wei, Zihao, et al.
Published: (2024)
by: Wei, Zihao, et al.
Published: (2024)
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
Domain Adaptation of VLM for Soccer Video Understanding
by: Jiang, Tiancheng, et al.
Published: (2025)
by: Jiang, Tiancheng, et al.
Published: (2025)
Similar Items
-
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025) -
ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
by: Shen, Haozhan, et al.
Published: (2024) -
Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research
by: Zhang, Qianqian, et al.
Published: (2025) -
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models
by: Shen, Haozhan, et al.
Published: (2026) -
OmChat: A Recipe to Train Multimodal Language Models with Strong Long Context and Video Understanding
by: Zhao, Tiancheng, et al.
Published: (2024)