Saved in:
| Main Authors: | Zhao, Kai, Xu, Chang, Si, Bailu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.19451 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multi-Label Contrastive Learning for Abstract Visual Reasoning
by: Małkiński, Mikołaj, et al.
Published: (2020)
by: Małkiński, Mikołaj, et al.
Published: (2020)
Learning Differentiable Logic Programs for Abstract Visual Reasoning
by: Shindo, Hikaru, et al.
Published: (2023)
by: Shindo, Hikaru, et al.
Published: (2023)
A Unified View of Abstract Visual Reasoning Problems
by: Małkiński, Mikołaj, et al.
Published: (2024)
by: Małkiński, Mikołaj, et al.
Published: (2024)
Advancing Generalization Across a Variety of Abstract Visual Reasoning Tasks
by: Małkiński, Mikołaj, et al.
Published: (2025)
by: Małkiński, Mikołaj, et al.
Published: (2025)
One Self-Configurable Model to Solve Many Abstract Visual Reasoning Problems
by: Małkiński, Mikołaj, et al.
Published: (2023)
by: Małkiński, Mikołaj, et al.
Published: (2023)
ConTextual: Evaluating Context-Sensitive Text-Rich Visual Reasoning in Large Multimodal Models
by: Wadhawan, Rohan, et al.
Published: (2024)
by: Wadhawan, Rohan, et al.
Published: (2024)
A-I-RAVEN and I-RAVEN-Mesh: Two New Benchmarks for Abstract Visual Reasoning
by: Małkiński, Mikołaj, et al.
Published: (2024)
by: Małkiński, Mikołaj, et al.
Published: (2024)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios
by: Jaiswal, Shantanu, et al.
Published: (2024)
by: Jaiswal, Shantanu, et al.
Published: (2024)
Streaming 4D Visual Geometry Transformer
by: Zhuo, Dong, et al.
Published: (2025)
by: Zhuo, Dong, et al.
Published: (2025)
ViT-2SPN: Vision Transformer-based Dual-Stream Self-Supervised Pretraining Networks for Retinal OCT Classification
by: Saraei, Mohammadreza, et al.
Published: (2025)
by: Saraei, Mohammadreza, et al.
Published: (2025)
VisualPredicator: Learning Abstract World Models with Neuro-Symbolic Predicates for Robot Planning
by: Liang, Yichao, et al.
Published: (2024)
by: Liang, Yichao, et al.
Published: (2024)
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
by: Lu, Pan, et al.
Published: (2023)
by: Lu, Pan, et al.
Published: (2023)
Weakly Supervised Concept Learning for Object-centric Visual Reasoning
by: Tiwari, Sparsh, et al.
Published: (2026)
by: Tiwari, Sparsh, et al.
Published: (2026)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
by: Zheng, Naishan, et al.
Published: (2025)
by: Zheng, Naishan, et al.
Published: (2025)
Learning Concept-Based Causal Transition and Symbolic Reasoning for Visual Planning
by: Qian, Yilue, et al.
Published: (2023)
by: Qian, Yilue, et al.
Published: (2023)
Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
by: Kumar, Sunil, et al.
Published: (2025)
by: Kumar, Sunil, et al.
Published: (2025)
CAPM: Fast and Robust Verification on Maxpool-based CNN via Dual Network
by: Bai, Jia-Hau, et al.
Published: (2024)
by: Bai, Jia-Hau, et al.
Published: (2024)
Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
by: Yang, Xiaoyu, et al.
Published: (2025)
by: Yang, Xiaoyu, et al.
Published: (2025)
PhysInOne: Visual Physics Learning and Reasoning in One Suite
by: Zhou, Siyuan, et al.
Published: (2026)
by: Zhou, Siyuan, et al.
Published: (2026)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding
by: Fan, Zezhong, et al.
Published: (2024)
by: Fan, Zezhong, et al.
Published: (2024)
Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
by: Reddy, N Dinesh, et al.
Published: (2025)
by: Reddy, N Dinesh, et al.
Published: (2025)
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
by: Xiao, Tong, et al.
Published: (2025)
by: Xiao, Tong, et al.
Published: (2025)
SPHINX: A Synthetic Environment for Visual Perception and Reasoning
by: Alam, Md Tanvirul, et al.
Published: (2025)
by: Alam, Md Tanvirul, et al.
Published: (2025)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
by: Miao, Yanting, et al.
Published: (2026)
by: Miao, Yanting, et al.
Published: (2026)
PROL : Rehearsal Free Continual Learning in Streaming Data via Prompt Online Learning
by: Ma'sum, M. Anwar, et al.
Published: (2025)
by: Ma'sum, M. Anwar, et al.
Published: (2025)
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
by: Zou, Chang, et al.
Published: (2024)
by: Zou, Chang, et al.
Published: (2024)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
RECODE: Reasoning Through Code Generation for Visual Question Answering
by: Shen, Junhong, et al.
Published: (2025)
by: Shen, Junhong, et al.
Published: (2025)
Colour and Brush Stroke Pattern Recognition in Abstract Art using Modified Deep Convolutional Generative Adversarial Networks
by: Srinivasan, Srinitish, et al.
Published: (2024)
by: Srinivasan, Srinitish, et al.
Published: (2024)
VERSE: Virtual-Gradient Aware Streaming Lifelong Learning with Anytime Inference
by: Banerjee, Soumya, et al.
Published: (2023)
by: Banerjee, Soumya, et al.
Published: (2023)
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
by: Sun, Hai-Long, et al.
Published: (2025)
by: Sun, Hai-Long, et al.
Published: (2025)
Abstract Art Interpretation Using ControlNet
by: Srivastava, Rishabh, et al.
Published: (2024)
by: Srivastava, Rishabh, et al.
Published: (2024)
Interpret Your Decision: Logical Reasoning Regularization for Generalization in Visual Classification
by: Tan, Zhaorui, et al.
Published: (2024)
by: Tan, Zhaorui, et al.
Published: (2024)
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers
by: Stanić, Aleksandar, et al.
Published: (2024)
by: Stanić, Aleksandar, et al.
Published: (2024)
V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
by: Fan, Chenrui, et al.
Published: (2025)
by: Fan, Chenrui, et al.
Published: (2025)
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
by: Qiao, Runqi, et al.
Published: (2025)
by: Qiao, Runqi, et al.
Published: (2025)
Visual Generation Without Guidance
by: Chen, Huayu, et al.
Published: (2025)
by: Chen, Huayu, et al.
Published: (2025)
Similar Items
-
Multi-Label Contrastive Learning for Abstract Visual Reasoning
by: Małkiński, Mikołaj, et al.
Published: (2020) -
Learning Differentiable Logic Programs for Abstract Visual Reasoning
by: Shindo, Hikaru, et al.
Published: (2023) -
A Unified View of Abstract Visual Reasoning Problems
by: Małkiński, Mikołaj, et al.
Published: (2024) -
Advancing Generalization Across a Variety of Abstract Visual Reasoning Tasks
by: Małkiński, Mikołaj, et al.
Published: (2025) -
One Self-Configurable Model to Solve Many Abstract Visual Reasoning Problems
by: Małkiński, Mikołaj, et al.
Published: (2023)