Explain Before You Answer: A Survey on Compositional Visual Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Ke, Fucai, Hsu, Joy, Cai, Zhixi, Ma, Zixian, Zheng, Xin, Wu, Xindi, Huang, Sukai, Wang, Weiqing, Haghighi, Pari Delir, Haffari, Gholamreza, Krishna, Ranjay, Wu, Jiajun, Rezatofighi, Hamid |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning
by: Ke, Fucai, et al.
Published: (2024)
by: Ke, Fucai, et al.
Published: (2024)
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
by: Ke, Fucai, et al.
Published: (2026)
by: Ke, Fucai, et al.
Published: (2026)
Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents
by: Huang, Sukai, et al.
Published: (2026)
by: Huang, Sukai, et al.
Published: (2026)
DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning
by: Ke, Fucai, et al.
Published: (2025)
by: Ke, Fucai, et al.
Published: (2025)
AR-Facilitated Safety Inspection and Fall Hazard Detection on Construction Sites
by: Liu, Jiazhou, et al.
Published: (2024)
by: Liu, Jiazhou, et al.
Published: (2024)
NAVER: A Neuro-Symbolic Compositional Automaton for Visual Grounding with Explicit Logic Reasoning
by: Cai, Zhixi, et al.
Published: (2025)
by: Cai, Zhixi, et al.
Published: (2025)
MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
by: Cai, Zhixi, et al.
Published: (2026)
by: Cai, Zhixi, et al.
Published: (2026)
ARIS: Agentic and Relationship Intelligence System for Social Robots
by: Datta, Stavya, et al.
Published: (2026)
by: Datta, Stavya, et al.
Published: (2026)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
by: Lou, Haowei, et al.
Published: (2024)
by: Lou, Haowei, et al.
Published: (2024)
JRDB-Social: A Multifaceted Robotic Dataset for Understanding of Context and Dynamics of Human Interactions Within Social Groups
by: Jahangard, Simindokht, et al.
Published: (2024)
by: Jahangard, Simindokht, et al.
Published: (2024)
Physics-Grounded Motion Forecasting via Equation Discovery for Trajectory-Guided Image-to-Video Generation
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
by: Yang, Yinuo, et al.
Published: (2026)
by: Yang, Yinuo, et al.
Published: (2026)
JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
by: Jahangard, Simindokht, et al.
Published: (2025)
by: Jahangard, Simindokht, et al.
Published: (2025)
Hier-SLAM: Scaling-up Semantics in SLAM with a Hierarchically Categorical Gaussian Splatting
by: Li, Boying, et al.
Published: (2024)
by: Li, Boying, et al.
Published: (2024)
Predicate Hierarchies Improve Few-Shot State Classification
by: Jin, Emily, et al.
Published: (2025)
by: Jin, Emily, et al.
Published: (2025)
The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
Assistive Large Language Model Agents for Socially-Aware Negotiation Dialogues
by: Hua, Yuncheng, et al.
Published: (2024)
by: Hua, Yuncheng, et al.
Published: (2024)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
by: Ma, Zixian, et al.
Published: (2024)
by: Ma, Zixian, et al.
Published: (2024)
TFS-NeRF: Template-Free NeRF for Semantic 3D Reconstruction of Dynamic Scene
by: Biswas, Sandika, et al.
Published: (2024)
by: Biswas, Sandika, et al.
Published: (2024)
Asset-Driven Sematic Reconstruction of Dynamic Scene with Multi-Human-Object Interactions
by: Biswas, Sandika, et al.
Published: (2025)
by: Biswas, Sandika, et al.
Published: (2025)
Importance-Aware Data Augmentation for Document-Level Neural Machine Translation
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
by: Feng, Chun, et al.
Published: (2024)
by: Feng, Chun, et al.
Published: (2024)
Neuro-Symbolic Decoding of Neural Activity
by: Wang, Yanchen, et al.
Published: (2026)
by: Wang, Yanchen, et al.
Published: (2026)
Watch Before You Answer: Learning from Visually Grounded Post-Training
by: Zhang, Yuxuan, et al.
Published: (2026)
by: Zhang, Yuxuan, et al.
Published: (2026)
Adapting Large Language Models for Document-Level Machine Translation
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
Rethinking Human Preference Evaluation of LLM Rationales
by: Li, Ziang, et al.
Published: (2025)
by: Li, Ziang, et al.
Published: (2025)
dinov3.seg: Open-Vocabulary Semantic Segmentation with DINOv3
by: Dutta, Saikat, et al.
Published: (2026)
by: Dutta, Saikat, et al.
Published: (2026)
JRDB-Pose3D: A Multi-person 3D Human Pose and Shape Estimation Dataset for Robotics
by: Biswas, Sandika, et al.
Published: (2026)
by: Biswas, Sandika, et al.
Published: (2026)
Social-MAE: Social Masked Autoencoder for Multi-person Motion Representation Learning
by: Ehsanpour, Mahsa, et al.
Published: (2024)
by: Ehsanpour, Mahsa, et al.
Published: (2024)
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Zero-Shot Privacy-Aware Text Rewriting via Iterative Tree Search
by: Huang, Shuo, et al.
Published: (2025)
by: Huang, Shuo, et al.
Published: (2025)
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
IRIS: An Iterative and Integrated Framework for Verifiable Causal Discovery in the Absence of Tabular Data
by: Feng, Tao, et al.
Published: (2025)
by: Feng, Tao, et al.
Published: (2025)
CausalScore: An Automatic Reference-Free Metric for Assessing Response Relevance in Open-Domain Dialogue Systems
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries
by: Hsu, Joy, et al.
Published: (2025)
by: Hsu, Joy, et al.
Published: (2025)
Normal-GS: 3D Gaussian Splatting with Normal-Involved Rendering
by: Wei, Meng, et al.
Published: (2024)
by: Wei, Meng, et al.
Published: (2024)
Similar Items
-
HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning
by: Ke, Fucai, et al.
Published: (2024) -
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations
by: Ke, Fucai, et al.
Published: (2026) -
Mini-BEHAVIOR-Gran: Revealing U-Shaped Effects of Instruction Granularity on Language-Guided Embodied Agents
by: Huang, Sukai, et al.
Published: (2026) -
DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning
by: Ke, Fucai, et al.
Published: (2025) -
AR-Facilitated Safety Inspection and Fall Hazard Detection on Construction Sites
by: Liu, Jiazhou, et al.
Published: (2024)