DeepEyesV2: Toward Agentic Multimodal Model
Fuente:
arXiv
Saved in:
| Main Authors: | Hong, Jack, Zhao, Chenxiao, Zhu, ChengLin, Lu, Weiheng, Xu, Guohai, Yu, Xing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
by: Zheng, Ziwei, et al.
Published: (2025)
by: Zheng, Ziwei, et al.
Published: (2025)
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
by: Wang, Pengfei, et al.
Published: (2025)
by: Wang, Pengfei, et al.
Published: (2025)
Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
by: Chen, Guo, et al.
Published: (2026)
by: Chen, Guo, et al.
Published: (2026)
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
by: V Team, et al.
Published: (2026)
by: V Team, et al.
Published: (2026)
ReDiPrune: Relevance-Diversity Pre-Projection Token Pruning for Efficient Multimodal LLMs
by: Yu, An, et al.
Published: (2026)
by: Yu, An, et al.
Published: (2026)
Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
by: Zhang, Wenchuan, et al.
Published: (2025)
by: Zhang, Wenchuan, et al.
Published: (2025)
DiffThinker: Towards Generative Multimodal Reasoning with Diffusion Models
by: He, Zefeng, et al.
Published: (2025)
by: He, Zefeng, et al.
Published: (2025)
Towards Long-horizon Agentic Multimodal Search
by: Du, Yifan, et al.
Published: (2026)
by: Du, Yifan, et al.
Published: (2026)
Eyes Closed, Safety On: Protecting Multimodal LLMs via Image-to-Text Transformation
by: Gou, Yunhao, et al.
Published: (2024)
by: Gou, Yunhao, et al.
Published: (2024)
ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
by: Shen, Haozhan, et al.
Published: (2024)
by: Shen, Haozhan, et al.
Published: (2024)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
by: Huang, Haoyu, et al.
Published: (2026)
by: Huang, Haoyu, et al.
Published: (2026)
ModalPrompt: Towards Efficient Multimodal Continual Instruction Tuning with Dual-Modality Guided Prompt
by: Zeng, Fanhu, et al.
Published: (2024)
by: Zeng, Fanhu, et al.
Published: (2024)
POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
by: Tran, Tuyen, et al.
Published: (2025)
by: Tran, Tuyen, et al.
Published: (2025)
GeoMix: Towards Geometry-Aware Data Augmentation
by: Zhao, Wentao, et al.
Published: (2024)
by: Zhao, Wentao, et al.
Published: (2024)
TalkingEyes: Pluralistic Speech-Driven 3D Eye Gaze Animation
by: Zhuang, Yixiang, et al.
Published: (2025)
by: Zhuang, Yixiang, et al.
Published: (2025)
V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval
by: Chen, Dongyang, et al.
Published: (2026)
by: Chen, Dongyang, et al.
Published: (2026)
Towards Unsupervised Eye-Region Segmentation for Eye Tracking
by: Deng, Jiangfan, et al.
Published: (2024)
by: Deng, Jiangfan, et al.
Published: (2024)
Toward Native Multimodal Modeling: A Roadmap
by: An, Siyu, et al.
Published: (2026)
by: An, Siyu, et al.
Published: (2026)
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
by: Guo, Yifu, et al.
Published: (2025)
by: Guo, Yifu, et al.
Published: (2025)
SenseNova-MARS: Empowering Multimodal Agentic Reasoning and Search via Reinforcement Learning
by: Chng, Yong Xien, et al.
Published: (2025)
by: Chng, Yong Xien, et al.
Published: (2025)
The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
by: Qi, Daiqing, et al.
Published: (2025)
by: Qi, Daiqing, et al.
Published: (2025)
SuperCLIP: CLIP with Simple Classification Supervision
by: Zhao, Weiheng, et al.
Published: (2025)
by: Zhao, Weiheng, et al.
Published: (2025)
AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment
by: Zhu, Hanwei, et al.
Published: (2025)
by: Zhu, Hanwei, et al.
Published: (2025)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
by: Ding, Shengyuan, et al.
Published: (2025)
by: Ding, Shengyuan, et al.
Published: (2025)
Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization
by: Yang, Wenhao, et al.
Published: (2026)
by: Yang, Wenhao, et al.
Published: (2026)
MM-DeepResearch: A Simple and Effective Multimodal Agentic Search Baseline
by: Yao, Huanjin, et al.
Published: (2026)
by: Yao, Huanjin, et al.
Published: (2026)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
by: Yu, An, et al.
Published: (2025)
by: Yu, An, et al.
Published: (2025)
FAPNet: An Effective Frequency Adaptive Point-based Eye Tracker
by: Lin, Xiaopeng, et al.
Published: (2024)
by: Lin, Xiaopeng, et al.
Published: (2024)
EchoAgent: Towards Reliable Echocardiography Interpretation with "Eyes","Hands" and "Minds"
by: Wang, Qin, et al.
Published: (2026)
by: Wang, Qin, et al.
Published: (2026)
Just Noticeable Difference Modeling for Deep Visual Features
by: Zhao, Rui, et al.
Published: (2026)
by: Zhao, Rui, et al.
Published: (2026)
iPay: Integrated Payment Action Recognition via Multimodal Networks and Adaptive Spatial Prior Learning
by: Huang, Kaicong, et al.
Published: (2026)
by: Huang, Kaicong, et al.
Published: (2026)
ProGuard: Towards Proactive Multimodal Safeguard
by: Yu, Shaohan, et al.
Published: (2025)
by: Yu, Shaohan, et al.
Published: (2025)
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval
by: Lu, Weiheng, et al.
Published: (2024)
by: Lu, Weiheng, et al.
Published: (2024)
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
by: Luo, Gen, et al.
Published: (2024)
by: Luo, Gen, et al.
Published: (2024)
The Eye of Sherlock Holmes: Uncovering User Private Attribute Profiling via Vision-Language Model Agentic Framework
by: Liu, Feiran, et al.
Published: (2025)
by: Liu, Feiran, et al.
Published: (2025)
SegLocNet: Multimodal Localization Network for Autonomous Driving via Bird's-Eye-View Segmentation
by: Zhou, Zijie, et al.
Published: (2025)
by: Zhou, Zijie, et al.
Published: (2025)
Empowering Functional Neuroimaging: A Pre-trained Generative Framework for Unified Representation of Neural Signals
by: Yao, Weiheng, et al.
Published: (2025)
by: Yao, Weiheng, et al.
Published: (2025)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
by: Zhao, Jiahe, et al.
Published: (2025)
by: Zhao, Jiahe, et al.
Published: (2025)
Similar Items
-
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
by: Zheng, Ziwei, et al.
Published: (2025) -
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
by: Wang, Pengfei, et al.
Published: (2025) -
Skywork-R1V4: Toward Agentic Multimodal Intelligence through Interleaved Thinking with Images and DeepResearch
by: Zhang, Yifan, et al.
Published: (2025) -
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
by: Chen, Guo, et al.
Published: (2026) -
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents
by: V Team, et al.
Published: (2026)