Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
Fuente:
arXiv
Saved in:
| Main Authors: | Jaiswal, Shantanu, Roy, Debaditya, Fernando, Basura, Tan, Cheston |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
by: Nagar, Aishik, et al.
Published: (2024)
by: Nagar, Aishik, et al.
Published: (2024)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023)
by: Rawal, Ishaan Singh, et al.
Published: (2023)
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Learning to Visually Connect Actions and their Effects
by: Parmar, Paritosh, et al.
Published: (2024)
by: Parmar, Paritosh, et al.
Published: (2024)
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
by: Parmar, Paritosh, et al.
Published: (2025)
by: Parmar, Paritosh, et al.
Published: (2025)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022)
by: Tan, Clement, et al.
Published: (2022)
Iterative Refinement Improves Compositional Image Generation
by: Jaiswal, Shantanu, et al.
Published: (2026)
by: Jaiswal, Shantanu, et al.
Published: (2026)
CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes
by: Parmar, Paritosh, et al.
Published: (2024)
by: Parmar, Paritosh, et al.
Published: (2024)
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
by: Rajendiran, Ramanathan, et al.
Published: (2025)
by: Rajendiran, Ramanathan, et al.
Published: (2025)
Multi-Label Contrastive Learning for Abstract Visual Reasoning
by: Małkiński, Mikołaj, et al.
Published: (2020)
by: Małkiński, Mikołaj, et al.
Published: (2020)
Learning Differentiable Logic Programs for Abstract Visual Reasoning
by: Shindo, Hikaru, et al.
Published: (2023)
by: Shindo, Hikaru, et al.
Published: (2023)
Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
by: Tan, Zhaorui, et al.
Published: (2024)
by: Tan, Zhaorui, et al.
Published: (2024)
Learning Visual Abstract Reasoning through Dual-Stream Networks
by: Zhao, Kai, et al.
Published: (2024)
by: Zhao, Kai, et al.
Published: (2024)
Weakly Supervised Concept Learning for Object-centric Visual Reasoning
by: Tiwari, Sparsh, et al.
Published: (2026)
by: Tiwari, Sparsh, et al.
Published: (2026)
Learning Concept-Based Causal Transition and Symbolic Reasoning for Visual Planning
by: Qian, Yilue, et al.
Published: (2023)
by: Qian, Yilue, et al.
Published: (2023)
MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
by: Kumar, Somnath, et al.
Published: (2024)
by: Kumar, Somnath, et al.
Published: (2024)
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
by: Verma, Dhruv, et al.
Published: (2024)
by: Verma, Dhruv, et al.
Published: (2024)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023)
by: Rajendiran, Ramanathan, et al.
Published: (2023)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
A Unified View of Abstract Visual Reasoning Problems
by: Małkiński, Mikołaj, et al.
Published: (2024)
by: Małkiński, Mikołaj, et al.
Published: (2024)
SPHINX: A Synthetic Environment for Visual Perception and Reasoning
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
PhysInOne: Visual Physics Learning and Reasoning in One Suite
by: Zhou, Siyuan, et al.
Published: (2026)
by: Zhou, Siyuan, et al.
Published: (2026)
Learning to Play Video Games with Intuitive Physics Priors
by: Jaiswal, Abhishek, et al.
Published: (2024)
by: Jaiswal, Abhishek, et al.
Published: (2024)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
RECODE: Reasoning Through Code Generation for Visual Question Answering
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
by: Reddy, N Dinesh, et al.
Published: (2025)
by: Reddy, N Dinesh, et al.
Published: (2025)
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
by: Qiao, Runqi, et al.
Published: (2025)
by: Qiao, Runqi, et al.
Published: (2025)
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers
by: Stanić, Aleksandar, et al.
Published: (2024)
by: Stanić, Aleksandar, et al.
Published: (2024)
Advancing Generalization Across a Variety of Abstract Visual Reasoning Tasks
by: Małkiński, Mikołaj, et al.
Published: (2025)
by: Małkiński, Mikołaj, et al.
Published: (2025)
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
by: Fan, Chenrui, et al.
Published: (2025)
by: Fan, Chenrui, et al.
Published: (2025)
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
by: Sun, Hai-Long, et al.
Published: (2025)
by: Sun, Hai-Long, et al.
Published: (2025)
Reasoning-Enhanced Object-Centric Learning for Videos
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
What's Holding Back Latent Visual Reasoning?
by: Viveiros, André G., et al.
Published: (2026)
by: Viveiros, André G., et al.
Published: (2026)
UniBench: Visual Reasoning Requires Rethinking Vision-Language Beyond Scaling
by: Al-Tahan, Haider, et al.
Published: (2024)
by: Al-Tahan, Haider, et al.
Published: (2024)
IPR-1: Interactive Physical Reasoner
by: Zhang, Mingyu, et al.
Published: (2025)
by: Zhang, Mingyu, et al.
Published: (2025)
MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes
by: Pande, Nilay, et al.
Published: (2025)
by: Pande, Nilay, et al.
Published: (2025)
One Self-Configurable Model to Solve Many Abstract Visual Reasoning Problems
by: Małkiński, Mikołaj, et al.
Published: (2023)
by: Małkiński, Mikołaj, et al.
Published: (2023)
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
by: Kumar, Sunil, et al.
Published: (2025)
by: Kumar, Sunil, et al.
Published: (2025)
Similar Items
-
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
by: Nagar, Aishik, et al.
Published: (2024) -
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
by: Rawal, Ishaan Singh, et al.
Published: (2023) -
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022) -
Learning to Visually Connect Actions and their Effects
by: Parmar, Paritosh, et al.
Published: (2024) -
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
by: Parmar, Paritosh, et al.
Published: (2025)