VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zhaozhi, Zhang, Tong, Guo, Mingyue, Wang, Yaowei, Ye, Qixiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mixture of Physical Priors Adapter for Parameter-Efficient Fine-Tuning
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2024)
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2024)
Building Vision Models upon Heat Conduction
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2024)
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2024)
Virtual Classification: Modulating Domain-Specific Knowledge for Multidomain Crowd Counting
von: Guo, Mingyue, et al.
Veröffentlicht: (2024)
von: Guo, Mingyue, et al.
Veröffentlicht: (2024)
Regressor-Segmenter Mutual Prompt Learning for Crowd Counting
von: Guo, Mingyue, et al.
Veröffentlicht: (2023)
von: Guo, Mingyue, et al.
Veröffentlicht: (2023)
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)
VMamba: Visual State Space Model
von: Liu, Yue, et al.
Veröffentlicht: (2024)
von: Liu, Yue, et al.
Veröffentlicht: (2024)
VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
von: Swetha, Sirnam, et al.
Veröffentlicht: (2025)
RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation Model
von: Hu, Huiyang, et al.
Veröffentlicht: (2024)
von: Hu, Huiyang, et al.
Veröffentlicht: (2024)
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
von: Tao, Ming, et al.
Veröffentlicht: (2024)
von: Tao, Ming, et al.
Veröffentlicht: (2024)
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
von: Wu, Zhiheng, et al.
Veröffentlicht: (2026)
von: Wu, Zhiheng, et al.
Veröffentlicht: (2026)
Multimodal Latent Reasoning via Hierarchical Visual Cues Injection
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
InViC: Intent-aware Visual Cues for Medical Visual Question Answering
von: Wang, Zhisong, et al.
Veröffentlicht: (2026)
von: Wang, Zhisong, et al.
Veröffentlicht: (2026)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
BEV$^2$PR: BEV-Enhanced Visual Place Recognition with Structural Cues
von: Ge, Fudong, et al.
Veröffentlicht: (2024)
von: Ge, Fudong, et al.
Veröffentlicht: (2024)
ReDDiT: Rehashing Noise for Discrete Visual Generation
von: Ma, Tianren, et al.
Veröffentlicht: (2025)
von: Ma, Tianren, et al.
Veröffentlicht: (2025)
Neurons: Emulating the Human Visual Cortex Improves Fidelity and Interpretability in fMRI-to-Video Reconstruction
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
von: Qi, Yu, et al.
Veröffentlicht: (2026)
von: Qi, Yu, et al.
Veröffentlicht: (2026)
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning
von: Gong, Xuan, et al.
Veröffentlicht: (2026)
von: Gong, Xuan, et al.
Veröffentlicht: (2026)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
von: Ma, Tianren, et al.
Veröffentlicht: (2024)
von: Ma, Tianren, et al.
Veröffentlicht: (2024)
TagFog: Textual Anchor Guidance and Fake Outlier Generation for Visual Out-of-Distribution Detection
von: Chen, Jiankang, et al.
Veröffentlicht: (2024)
von: Chen, Jiankang, et al.
Veröffentlicht: (2024)
Towards Visual Grounding: A Survey
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
Reinforcing Multimodal Reasoning Against Visual Degradation
von: Liu, Rui, et al.
Veröffentlicht: (2026)
von: Liu, Rui, et al.
Veröffentlicht: (2026)
Video-ToC: Video Tree-of-Cue Reasoning
von: Tan, Qizhong, et al.
Veröffentlicht: (2026)
von: Tan, Qizhong, et al.
Veröffentlicht: (2026)
VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
von: Yan, Ziang, et al.
Veröffentlicht: (2025)
von: Yan, Ziang, et al.
Veröffentlicht: (2025)
Enhancing Spatial Reasoning through Visual and Textual Thinking
von: Liang, Xun, et al.
Veröffentlicht: (2025)
von: Liang, Xun, et al.
Veröffentlicht: (2025)
CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis
von: Zhang, Mu, et al.
Veröffentlicht: (2024)
von: Zhang, Mu, et al.
Veröffentlicht: (2024)
Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
von: Feng, X., et al.
Veröffentlicht: (2024)
von: Feng, X., et al.
Veröffentlicht: (2024)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning
von: Luo, Zhanpeng, et al.
Veröffentlicht: (2026)
von: Luo, Zhanpeng, et al.
Veröffentlicht: (2026)
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
von: Wu, Junfei, et al.
Veröffentlicht: (2025)
von: Wu, Junfei, et al.
Veröffentlicht: (2025)
Visual Reasoning through Tool-supervised Reinforcement Learning
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
von: Dong, Qihua, et al.
Veröffentlicht: (2026)
Beyond Visual Cues: Leveraging General Semantics as Support for Few-Shot Segmentation
von: Wang, Jin, et al.
Veröffentlicht: (2025)
von: Wang, Jin, et al.
Veröffentlicht: (2025)
Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Qizhe, et al.
Veröffentlicht: (2024)
Video-R1: Reinforcing Video Reasoning in MLLMs
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
von: Feng, Kaituo, et al.
Veröffentlicht: (2025)
Structure Over Scale: Learning Visual Reasoning from Pedagogical Video
von: Galoaa, Bishoy, et al.
Veröffentlicht: (2026)
von: Galoaa, Bishoy, et al.
Veröffentlicht: (2026)
SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
von: Sun, Peiwen, et al.
Veröffentlicht: (2025)
von: Sun, Peiwen, et al.
Veröffentlicht: (2025)
Triage: Hierarchical Visual Budgeting for Efficient Video Reasoning in Vision-Language Models
von: Wang, Anmin, et al.
Veröffentlicht: (2026)
von: Wang, Anmin, et al.
Veröffentlicht: (2026)
NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
von: Liu, Xiangyan, et al.
Veröffentlicht: (2025)
von: Liu, Xiangyan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Mixture of Physical Priors Adapter for Parameter-Efficient Fine-Tuning
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2024) -
Building Vision Models upon Heat Conduction
von: Wang, Zhaozhi, et al.
Veröffentlicht: (2024) -
Virtual Classification: Modulating Domain-Specific Knowledge for Multidomain Crowd Counting
von: Guo, Mingyue, et al.
Veröffentlicht: (2024) -
Regressor-Segmenter Mutual Prompt Learning for Crowd Counting
von: Guo, Mingyue, et al.
Veröffentlicht: (2023) -
From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
von: Wu, Mingrui, et al.
Veröffentlicht: (2025)