Region-Level Context-Aware Multimodal Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wei, Hongliang, Zhang, Xianqi, Wang, Xingtao, Fan, Xiaopeng, Zhao, Debin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LG-HCC: Local Geometry-Aware Hierarchical Context Compression for 3D Gaussian Splatting
von: Deng, Xuan, et al.
Veröffentlicht: (2026)
von: Deng, Xuan, et al.
Veröffentlicht: (2026)
Region Prompt Tuning: Fine-grained Scene Text Detection Utilizing Region Text Prompt
von: Lin, Xingtao, et al.
Veröffentlicht: (2024)
von: Lin, Xingtao, et al.
Veröffentlicht: (2024)
Bridging the Visual-to-Physical Gap: Physically Aligned Representations for Fall Risk Analysis
von: Zhang, Xianqi
Veröffentlicht: (2026)
von: Zhang, Xianqi
Veröffentlicht: (2026)
PVINet: Point-Voxel Interlaced Network for Point Cloud Compression
von: Deng, Xuan, et al.
Veröffentlicht: (2025)
von: Deng, Xuan, et al.
Veröffentlicht: (2025)
Hyperbolic Distillation: Geometry-Guided Cross-Modal Transfer for Robust 3D Object Detection
von: Ning, Kanglin, et al.
Veröffentlicht: (2026)
von: Ning, Kanglin, et al.
Veröffentlicht: (2026)
PairDropGS: Paired Dropout-Induced Consistency Regularization for Sparse-View Gaussian Splatting
von: Li, Hantang, et al.
Veröffentlicht: (2026)
von: Li, Hantang, et al.
Veröffentlicht: (2026)
T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates
von: Wang, Zhitao, et al.
Veröffentlicht: (2025)
von: Wang, Zhitao, et al.
Veröffentlicht: (2025)
RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding
von: Fang, Tianchen, et al.
Veröffentlicht: (2025)
von: Fang, Tianchen, et al.
Veröffentlicht: (2025)
Bidirectional Feature-aligned Motion Transformation for Efficient Dynamic Point Cloud Compression
von: Deng, Xuan, et al.
Veröffentlicht: (2025)
von: Deng, Xuan, et al.
Veröffentlicht: (2025)
PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
von: Blume, Ansel, et al.
Veröffentlicht: (2025)
von: Blume, Ansel, et al.
Veröffentlicht: (2025)
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
von: Rahman, Ben
Veröffentlicht: (2025)
von: Rahman, Ben
Veröffentlicht: (2025)
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs
von: Wang, Wei-Yao, et al.
Veröffentlicht: (2025)
von: Wang, Wei-Yao, et al.
Veröffentlicht: (2025)
DynamicPAE: Generating Scene-Aware Physical Adversarial Examples in Real-Time
von: Hu, Jin, et al.
Veröffentlicht: (2024)
von: Hu, Jin, et al.
Veröffentlicht: (2024)
iGVLM: Dynamic Instruction-Guided Vision Encoding for Question-Aware Multimodal Understanding
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
von: Liu, Hanpeng, et al.
Veröffentlicht: (2026)
Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
von: Jang, Lawrence, et al.
Veröffentlicht: (2024)
von: Jang, Lawrence, et al.
Veröffentlicht: (2024)
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
von: Chen, Pengtao, et al.
Veröffentlicht: (2025)
von: Chen, Pengtao, et al.
Veröffentlicht: (2025)
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding
von: Nguyen-Truong, Hai, et al.
Veröffentlicht: (2024)
von: Nguyen-Truong, Hai, et al.
Veröffentlicht: (2024)
Deep Network for Image Compressed Sensing Coding Using Local Structural Sampling
von: Cui, Wenxue, et al.
Veröffentlicht: (2024)
von: Cui, Wenxue, et al.
Veröffentlicht: (2024)
GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought Reasoning
von: Liu, Bo, et al.
Veröffentlicht: (2025)
von: Liu, Bo, et al.
Veröffentlicht: (2025)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
von: Zhang, Yue, et al.
Veröffentlicht: (2024)
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
von: Guo, Zhongbin, et al.
Veröffentlicht: (2025)
von: Guo, Zhongbin, et al.
Veröffentlicht: (2025)
Personalized Video Summarization by Multimodal Video Understanding
von: Chen, Brian, et al.
Veröffentlicht: (2024)
von: Chen, Brian, et al.
Veröffentlicht: (2024)
ContextDet: Temporal Action Detection with Adaptive Context Aggregation
von: Wang, Ning, et al.
Veröffentlicht: (2024)
von: Wang, Ning, et al.
Veröffentlicht: (2024)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
von: Peng, Tianhao, et al.
Veröffentlicht: (2025)
von: Peng, Tianhao, et al.
Veröffentlicht: (2025)
VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhihong, et al.
Veröffentlicht: (2025)
IAD-Unify: A Region-Grounded Unified Model for Industrial Anomaly Segmentation, Understanding, and Generation
von: Zheng, Haoyu, et al.
Veröffentlicht: (2026)
von: Zheng, Haoyu, et al.
Veröffentlicht: (2026)
BEVWorld: A Multimodal World Simulator for Autonomous Driving via Scene-Level BEV Latents
von: Zhang, Yumeng, et al.
Veröffentlicht: (2024)
von: Zhang, Yumeng, et al.
Veröffentlicht: (2024)
ARMOR: Empowering Multimodal Understanding Model with Interleaved Multimodal Generation Capability
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
von: Sun, Jianwen, et al.
Veröffentlicht: (2025)
MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding
von: Li, Zekun, et al.
Veröffentlicht: (2024)
von: Li, Zekun, et al.
Veröffentlicht: (2024)
SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
Task-Agnostic Learning to Accomplish New Tasks
von: Zhang, Xianqi, et al.
Veröffentlicht: (2022)
von: Zhang, Xianqi, et al.
Veröffentlicht: (2022)
MARS: Multimodal Active Robotic Sensing for Articulated Characterization
von: Zeng, Hongliang, et al.
Veröffentlicht: (2024)
von: Zeng, Hongliang, et al.
Veröffentlicht: (2024)
Robust-R1: Degradation-Aware Reasoning for Robust Visual Understanding
von: Tang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Tang, Jiaqi, et al.
Veröffentlicht: (2025)
A-MESS: Anchor based Multimodal Embedding with Semantic Synchronization for Multimodal Intent Recognition
von: Shen, Yaomin, et al.
Veröffentlicht: (2025)
von: Shen, Yaomin, et al.
Veröffentlicht: (2025)
Region-to-Region: Enhancing Generative Image Harmonization with Adaptive Regional Injection
von: Zhang, Zhiqiu, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiqiu, et al.
Veröffentlicht: (2025)
ContextGS: Compact 3D Gaussian Splatting with Anchor Level Context Model
von: Wang, Yufei, et al.
Veröffentlicht: (2024)
von: Wang, Yufei, et al.
Veröffentlicht: (2024)
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?
von: Wen, Zimo, et al.
Veröffentlicht: (2026)
von: Wen, Zimo, et al.
Veröffentlicht: (2026)
Exploring Task-Level Optimal Prompts for Visual In-Context Learning
von: Zhu, Yan, et al.
Veröffentlicht: (2025)
von: Zhu, Yan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LG-HCC: Local Geometry-Aware Hierarchical Context Compression for 3D Gaussian Splatting
von: Deng, Xuan, et al.
Veröffentlicht: (2026) -
Region Prompt Tuning: Fine-grained Scene Text Detection Utilizing Region Text Prompt
von: Lin, Xingtao, et al.
Veröffentlicht: (2024) -
Bridging the Visual-to-Physical Gap: Physically Aligned Representations for Fall Risk Analysis
von: Zhang, Xianqi
Veröffentlicht: (2026) -
PVINet: Point-Voxel Interlaced Network for Point Cloud Compression
von: Deng, Xuan, et al.
Veröffentlicht: (2025) -
Hyperbolic Distillation: Geometry-Guided Cross-Modal Transfer for Robust 3D Object Detection
von: Ning, Kanglin, et al.
Veröffentlicht: (2026)