ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yuxuan, Yuille, Alan, Li, Zhuowan, Zheng, Zilong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
by: Li, Zhuowan, et al.
Published: (2022)
by: Li, Zhuowan, et al.
Published: (2022)
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
by: Li, Zhuowan, et al.
Published: (2024)
by: Li, Zhuowan, et al.
Published: (2024)
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning
by: Wang, Jianghui, et al.
Published: (2023)
by: Wang, Jianghui, et al.
Published: (2023)
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
by: Yang, Timing, et al.
Published: (2025)
by: Yang, Timing, et al.
Published: (2025)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
ViT-5: Vision Transformers for The Mid-2020s
by: Wang, Feng, et al.
Published: (2026)
by: Wang, Feng, et al.
Published: (2026)
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
by: Thawakar, Omkar, et al.
Published: (2025)
by: Thawakar, Omkar, et al.
Published: (2025)
ViTamin: Designing Scalable Vision Models in the Vision-Language Era
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
by: Yang, Timing, et al.
Published: (2025)
by: Yang, Timing, et al.
Published: (2025)
Large Language Models are Universal Reasoners for Visual Generation
by: Ren, Sucheng, et al.
Published: (2026)
by: Ren, Sucheng, et al.
Published: (2026)
Computer Vision and Its Relationship to Cognitive Science: A perspective from Bayes Decision Theory
by: Yuille, Alan, et al.
Published: (2026)
by: Yuille, Alan, et al.
Published: (2026)
ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
by: Wang, Lihong, et al.
Published: (2025)
by: Wang, Lihong, et al.
Published: (2025)
Reasoning Palette: Modulating Reasoning via Latent Contextualization for Controllable Exploration for (V)LMs
by: Long, Rujiao, et al.
Published: (2025)
by: Long, Rujiao, et al.
Published: (2025)
Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark
by: Liu, Xinxin, et al.
Published: (2025)
by: Liu, Xinxin, et al.
Published: (2025)
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
by: Wang, Feng, et al.
Published: (2023)
by: Wang, Feng, et al.
Published: (2023)
Rejuvenating image-GPT as Strong Visual Representation Learners
by: Ren, Sucheng, et al.
Published: (2023)
by: Ren, Sucheng, et al.
Published: (2023)
Slow Perception: Let's Perceive Geometric Figures Step-by-step
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
PaLM2-VAdapter: Progressively Aligned Language Model Makes a Strong Vision-language Adapter
by: Xiao, Junfei, et al.
Published: (2024)
by: Xiao, Junfei, et al.
Published: (2024)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
by: Ma, Wufei, et al.
Published: (2025)
by: Ma, Wufei, et al.
Published: (2025)
Thinking with Spatial Code for Physical-World Video Reasoning
by: Chen, Jieneng, et al.
Published: (2026)
by: Chen, Jieneng, et al.
Published: (2026)
Efficient Large Multi-modal Models via Visual Context Compression
by: Chen, Jieneng, et al.
Published: (2024)
by: Chen, Jieneng, et al.
Published: (2024)
ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
by: Liao, Wenjie, et al.
Published: (2025)
by: Liao, Wenjie, et al.
Published: (2025)
Gaussian Scenes: Pose-Free Sparse-View Scene Reconstruction using Depth-Enhanced Diffusion Priors
by: Paul, Soumava, et al.
Published: (2024)
by: Paul, Soumava, et al.
Published: (2024)
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
by: Paul, Soumava, et al.
Published: (2026)
by: Paul, Soumava, et al.
Published: (2026)
Quality Sentinel: Estimating Label Quality and Errors in Medical Segmentation Datasets
by: Chen, Yixiong, et al.
Published: (2024)
by: Chen, Yixiong, et al.
Published: (2024)
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
ViLLa: Video Reasoning Segmentation with Large Language Model
by: Zheng, Rongkun, et al.
Published: (2024)
by: Zheng, Rongkun, et al.
Published: (2024)
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
by: Xu, Jilan, et al.
Published: (2025)
by: Xu, Jilan, et al.
Published: (2025)
EgoExo-WM: Unlocking Exo Video for Ego World Models
by: Tran, Danny, et al.
Published: (2026)
by: Tran, Danny, et al.
Published: (2026)
HECTOR: Hybrid Editable Compositional Object References for Video Generation
by: Zhang, Guofeng, et al.
Published: (2026)
by: Zhang, Guofeng, et al.
Published: (2026)
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos
by: Liu, Ruiping, et al.
Published: (2026)
by: Liu, Ruiping, et al.
Published: (2026)
KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
by: Li, Zhiyang, et al.
Published: (2026)
by: Li, Zhiyang, et al.
Published: (2026)
InViC: Intent-aware Visual Cues for Medical Visual Question Answering
by: Wang, Zhisong, et al.
Published: (2026)
by: Wang, Zhisong, et al.
Published: (2026)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
by: Wang, Yuxuan, et al.
Published: (2024)
by: Wang, Yuxuan, et al.
Published: (2024)
RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology
by: Li, Wenxuan, et al.
Published: (2026)
by: Li, Wenxuan, et al.
Published: (2026)
OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
by: Han, Zongyan, et al.
Published: (2025)
by: Han, Zongyan, et al.
Published: (2025)
Beyond Next-Token: Next-X Prediction for Autoregressive Visual Generation
by: Ren, Sucheng, et al.
Published: (2025)
by: Ren, Sucheng, et al.
Published: (2025)
Play to Generalize: Learning to Reason Through Game Play
by: Xie, Yunfei, et al.
Published: (2025)
by: Xie, Yunfei, et al.
Published: (2025)
CamFreeDiff: Camera-free Image to Panorama Generation with Diffusion Model
by: Yuan, Xiaoding, et al.
Published: (2024)
by: Yuan, Xiaoding, et al.
Published: (2024)
Beyond Masks: The Case for Medical Image Parsing
by: Gupta, Siddharth, et al.
Published: (2026)
by: Gupta, Siddharth, et al.
Published: (2026)
Similar Items
-
Localization vs. Semantics: Visual Representations in Unimodal and Multimodal Models
by: Li, Zhuowan, et al.
Published: (2022) -
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
by: Li, Zhuowan, et al.
Published: (2024) -
MoviePuzzle: Visual Narrative Reasoning through Multimodal Order Learning
by: Wang, Jianghui, et al.
Published: (2023) -
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
by: Yang, Timing, et al.
Published: (2025) -
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
by: Bai, Tianyi, et al.
Published: (2025)