Composition-Grounded Data Synthesis for Visual Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gu, Xinyi, Mao, Jiayuan, Hong, Zhang-Wei, Yu, Zhuoran, Li, Pengyuan, Joshi, Dhiraj, Feris, Rogerio, He, Zexue |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Generating Fine Details of Entity Interactions
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
von: Gu, Xinyi, et al.
Veröffentlicht: (2025)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
von: Yu, Zhuoran, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoran, et al.
Veröffentlicht: (2025)
LVCHAT: Facilitating Long Video Comprehension
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
von: Kondic, Jovana, et al.
Veröffentlicht: (2025)
von: Kondic, Jovana, et al.
Veröffentlicht: (2025)
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
von: Zhao, Yu, et al.
Veröffentlicht: (2025)
ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
von: Kondic, Jovana, et al.
Veröffentlicht: (2026)
von: Kondic, Jovana, et al.
Veröffentlicht: (2026)
Latent Implicit Visual Reasoning
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
von: Li, Kelvin, et al.
Veröffentlicht: (2025)
On Data Synthesis and Post-training for Visual Abstract Reasoning
von: Zhu, Ke, et al.
Veröffentlicht: (2025)
von: Zhu, Ke, et al.
Veröffentlicht: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
von: Chung, Jiwan, et al.
Veröffentlicht: (2025)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2026)
VGR: Visual Grounded Reasoning
von: Wang, Jiacong, et al.
Veröffentlicht: (2025)
von: Wang, Jiacong, et al.
Veröffentlicht: (2025)
Learning from Synthetic Data for Visual Grounding
von: He, Ruozhen, et al.
Veröffentlicht: (2024)
von: He, Ruozhen, et al.
Veröffentlicht: (2024)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
Learning to Correction: Explainable Feedback Generation for Visual Commonsense Reasoning Distractor
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
von: Chen, Jiali, et al.
Veröffentlicht: (2024)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
von: Luo, Jiayun, et al.
Veröffentlicht: (2024)
von: Luo, Jiayun, et al.
Veröffentlicht: (2024)
SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis
von: Wu, Zijian, et al.
Veröffentlicht: (2025)
von: Wu, Zijian, et al.
Veröffentlicht: (2025)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
von: Acuna, David, et al.
Veröffentlicht: (2025)
von: Acuna, David, et al.
Veröffentlicht: (2025)
Learning Compositional Behaviors from Demonstration and Language
von: Liu, Weiyu, et al.
Veröffentlicht: (2025)
von: Liu, Weiyu, et al.
Veröffentlicht: (2025)
ScanReason: Empowering 3D Visual Grounding with Reasoning Capabilities
von: Zhu, Chenming, et al.
Veröffentlicht: (2024)
von: Zhu, Chenming, et al.
Veröffentlicht: (2024)
SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2026)
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2026)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
von: Shrestha, Robik, et al.
Veröffentlicht: (2020)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
von: Bai, Tianyi, et al.
Veröffentlicht: (2025)
UI-E2I-Synth: Advancing GUI Grounding with Large-Scale Instruction Synthesis
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
von: Ma, David, et al.
Veröffentlicht: (2025)
von: Ma, David, et al.
Veröffentlicht: (2025)
Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
von: Yang, Minglai, et al.
Veröffentlicht: (2026)
Latent Visual Reasoning
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Symbolic Grounding Reveals Representational Bottlenecks in Abstract Visual Reasoning
von: Vaishnav, Mohit, et al.
Veröffentlicht: (2026)
von: Vaishnav, Mohit, et al.
Veröffentlicht: (2026)
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
von: Guo, Pinxue, et al.
Veröffentlicht: (2025)
von: Guo, Pinxue, et al.
Veröffentlicht: (2025)
Adaptive Memory Replay for Continual Learning
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
von: Smith, James Seale, et al.
Veröffentlicht: (2024)
VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models
von: Li, Zejun, et al.
Veröffentlicht: (2024)
von: Li, Zejun, et al.
Veröffentlicht: (2024)
Seeing Culture: A Benchmark for Visual Reasoning and Grounding
von: Satar, Burak, et al.
Veröffentlicht: (2025)
von: Satar, Burak, et al.
Veröffentlicht: (2025)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
von: Tang, Yolo Y., et al.
Veröffentlicht: (2025)
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
VEGGIE: Instructional Editing and Reasoning Video Concepts with Grounded Generation
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
von: Yu, Shoubin, et al.
Veröffentlicht: (2025)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
von: Wei, Yana, et al.
Veröffentlicht: (2025)
von: Wei, Yana, et al.
Veröffentlicht: (2025)
Refining Skewed Perceptions in Vision-Language Contrastive Models through Visual Representations
von: Dai, Haocheng, et al.
Veröffentlicht: (2024)
von: Dai, Haocheng, et al.
Veröffentlicht: (2024)
How Far Are We from Intelligent Visual Deductive Reasoning?
von: Zhang, Yizhe, et al.
Veröffentlicht: (2024)
von: Zhang, Yizhe, et al.
Veröffentlicht: (2024)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
von: Li, Chengzu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Generating Fine Details of Entity Interactions
von: Gu, Xinyi, et al.
Veröffentlicht: (2025) -
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
von: Yu, Zhuoran, et al.
Veröffentlicht: (2025) -
LVCHAT: Facilitating Long Video Comprehension
von: Wang, Yu, et al.
Veröffentlicht: (2024) -
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
von: Kondic, Jovana, et al.
Veröffentlicht: (2025) -
Learning GUI Grounding with Spatial Reasoning from Visual Feedback
von: Zhao, Yu, et al.
Veröffentlicht: (2025)