VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kang, Weitai, Kuen, Jason, Ren, Mengwei, Wei, Zijun, Yan, Yan, Liu, Kangning |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ACTRESS: Active Retraining for Semi-supervised Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
Inline Critic Steers Image Editing
von: Kang, Weitai, et al.
Veröffentlicht: (2026)
von: Kang, Weitai, et al.
Veröffentlicht: (2026)
Visual Grounding with Attention-Driven Constraint Balancing
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
von: Kang, Weitai, et al.
Veröffentlicht: (2025)
von: Kang, Weitai, et al.
Veröffentlicht: (2025)
Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models
von: Li, Shufan, et al.
Veröffentlicht: (2025)
von: Li, Shufan, et al.
Veröffentlicht: (2025)
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
von: Li, Shufan, et al.
Veröffentlicht: (2025)
von: Li, Shufan, et al.
Veröffentlicht: (2025)
On the Faithfulness of Vision Transformer Explanations
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
Token Transformation Matters: Towards Faithful Post-hoc Explanation for Vision Transformer
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
von: Wu, Junyi, et al.
Veröffentlicht: (2024)
3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation
von: Chen, Wenxin, et al.
Veröffentlicht: (2025)
von: Chen, Wenxin, et al.
Veröffentlicht: (2025)
Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
Refer to Any Segmentation Mask Group With Vision-Language Prompts
von: Cao, Shengcao, et al.
Veröffentlicht: (2025)
von: Cao, Shengcao, et al.
Veröffentlicht: (2025)
SNCE: Geometry-Aware Supervision for Scalable Discrete Image Generation
von: Li, Shufan, et al.
Veröffentlicht: (2026)
von: Li, Shufan, et al.
Veröffentlicht: (2026)
Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models
von: Li, Shufan, et al.
Veröffentlicht: (2026)
von: Li, Shufan, et al.
Veröffentlicht: (2026)
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
von: Shu, Yan, et al.
Veröffentlicht: (2026)
von: Shu, Yan, et al.
Veröffentlicht: (2026)
ReDiStory: Region-Disentangled Diffusion for Consistent Visual Story Generation
von: Sarkar, Ayushman, et al.
Veröffentlicht: (2026)
von: Sarkar, Ayushman, et al.
Veröffentlicht: (2026)
Predicting Visual Attention in Graphic Design Documents
von: Chakraborty, Souradeep, et al.
Veröffentlicht: (2024)
von: Chakraborty, Souradeep, et al.
Veröffentlicht: (2024)
GroundFlow: A Plug-in Module for Temporal Reasoning on 3D Point Cloud Sequential Grounding
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
von: Lin, Zijun, et al.
Veröffentlicht: (2025)
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction
von: Yan, Xiaoyang, et al.
Veröffentlicht: (2026)
von: Yan, Xiaoyang, et al.
Veröffentlicht: (2026)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
von: Zhang, Bob, et al.
Veröffentlicht: (2025)
von: Zhang, Bob, et al.
Veröffentlicht: (2025)
Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
von: Li, Jun, et al.
Veröffentlicht: (2025)
von: Li, Jun, et al.
Veröffentlicht: (2025)
Grounded Reinforcement Learning for Visual Reasoning
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
von: Sarch, Gabriel, et al.
Veröffentlicht: (2025)
Multimodal Latent Reasoning via Hierarchical Visual Cues Injection
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
von: Huang, Xinyu, et al.
Veröffentlicht: (2025)
von: Huang, Xinyu, et al.
Veröffentlicht: (2025)
VGR: Visual Grounded Reasoning
von: Wang, Jiacong, et al.
Veröffentlicht: (2025)
von: Wang, Jiacong, et al.
Veröffentlicht: (2025)
Connecting the Dots: Training-Free Visual Grounding via Agentic Reasoning
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
von: Luo, Liqin, et al.
Veröffentlicht: (2025)
MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
von: Shen, Haozhan, et al.
Veröffentlicht: (2026)
von: Shen, Haozhan, et al.
Veröffentlicht: (2026)
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
von: Yu, Xiao, et al.
Veröffentlicht: (2025)
Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
von: Yuan, Haobo, et al.
Veröffentlicht: (2025)
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
von: Sun, Wenfang, et al.
Veröffentlicht: (2026)
von: Sun, Wenfang, et al.
Veröffentlicht: (2026)
Reasoning in Space via Grounding in the World
von: Chen, Yiming, et al.
Veröffentlicht: (2025)
von: Chen, Yiming, et al.
Veröffentlicht: (2025)
Guiding the Inner Eye: A Framework for Hierarchical and Flexible Visual Grounded Reasoning
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2025)
von: Wei, Zhaoyang, et al.
Veröffentlicht: (2025)
Disentangled Representation Learning via Modular Compositional Bias
von: Jung, Whie, et al.
Veröffentlicht: (2025)
von: Jung, Whie, et al.
Veröffentlicht: (2025)
When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
von: Zhou, Yiyang, et al.
Veröffentlicht: (2025)
ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition
von: Huang, Ronggang, et al.
Veröffentlicht: (2025)
von: Huang, Ronggang, et al.
Veröffentlicht: (2025)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
Enhancing Visual Programming for Visual Reasoning via Probabilistic Graphs
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
Chain-of-Ground: Improving GUI Grounding via Iterative Reasoning and Reference Feedback
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2025)
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ACTRESS: Active Retraining for Semi-supervised Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024) -
SegVG: Transferring Object Bounding Box to Segmentation for Visual Grounding
von: Kang, Weitai, et al.
Veröffentlicht: (2024) -
Inline Critic Steers Image Editing
von: Kang, Weitai, et al.
Veröffentlicht: (2026) -
Visual Grounding with Attention-Driven Constraint Balancing
von: Kang, Weitai, et al.
Veröffentlicht: (2024) -
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
von: Kang, Weitai, et al.
Veröffentlicht: (2025)