ARC Is a Vision Problem!
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Keya, Cy, Ali, Qiu, Linlu, Ding, Xiaoman Delores, Wang, Runqian, Zhu, Yeyin Eva, Andreas, Jacob, He, Kaiming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diffuse and Disperse: Image Generation with Representation Regularization
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
von: Wang, Runqian, et al.
Veröffentlicht: (2025)
Video-based Vehicle Surveillance in the Wild: License Plate, Make, and Model Recognition with Self Reflective Vision-Language Models
von: Parsa, Pouya, et al.
Veröffentlicht: (2025)
von: Parsa, Pouya, et al.
Veröffentlicht: (2025)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
Bongards at the Boundary of Perception and Reasoning: Programs or Language?
von: Langenfeld, Cassidy, et al.
Veröffentlicht: (2026)
von: Langenfeld, Cassidy, et al.
Veröffentlicht: (2026)
3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models
von: Sambara, Sraavya, et al.
Veröffentlicht: (2025)
von: Sambara, Sraavya, et al.
Veröffentlicht: (2025)
ELF: Embedded Language Flows
von: Hu, Keya, et al.
Veröffentlicht: (2026)
von: Hu, Keya, et al.
Veröffentlicht: (2026)
Does Visual Information Play a Decisive Role in Vision-Language-Action Model Driving Behavior?
von: He, Jingtao, et al.
Veröffentlicht: (2026)
von: He, Jingtao, et al.
Veröffentlicht: (2026)
Surface Vision Mamba: Leveraging Bidirectional State Space Model for Efficient Spherical Manifold Representation
von: He, Rongzhao, et al.
Veröffentlicht: (2025)
von: He, Rongzhao, et al.
Veröffentlicht: (2025)
Large Vision Models Can Solve Mental Rotation Problems
von: Mason, Sebastian Ray, et al.
Veröffentlicht: (2025)
von: Mason, Sebastian Ray, et al.
Veröffentlicht: (2025)
Image Generators are Generalist Vision Learners
von: Gabeur, Valentin, et al.
Veröffentlicht: (2026)
von: Gabeur, Valentin, et al.
Veröffentlicht: (2026)
Diffusion-Based Offline RL for Improved Decision-Making in Augmented ARC Task
von: Kim, Yunho, et al.
Veröffentlicht: (2024)
von: Kim, Yunho, et al.
Veröffentlicht: (2024)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
Nested-TNT: Hierarchical Vision Transformers with Multi-Scale Feature Processing
von: Liu, Yuang, et al.
Veröffentlicht: (2024)
von: Liu, Yuang, et al.
Veröffentlicht: (2024)
LPT: Less-overfitting Prompt Tuning for Vision-Language Model
von: Ding, Chenhao, et al.
Veröffentlicht: (2024)
von: Ding, Chenhao, et al.
Veröffentlicht: (2024)
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
von: Ding, Guodong, et al.
Veröffentlicht: (2026)
von: Ding, Guodong, et al.
Veröffentlicht: (2026)
Vision-Language Models in Remote Sensing: Current Progress and Future Trends
von: Li, Xiang, et al.
Veröffentlicht: (2023)
von: Li, Xiang, et al.
Veröffentlicht: (2023)
How Modality Shapes Perception and Reasoning: A Study of Error Propagation in ARC-AGI
von: Wen, Bo, et al.
Veröffentlicht: (2025)
von: Wen, Bo, et al.
Veröffentlicht: (2025)
Binary Verification for Zero-Shot Vision
von: Hu, Rongbin, et al.
Veröffentlicht: (2025)
von: Hu, Rongbin, et al.
Veröffentlicht: (2025)
Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
von: Zhang, Jie, et al.
Veröffentlicht: (2025)
Vision Language Model-Empowered Contract Theory for AIGC Task Allocation in Teleoperation
von: Zhan, Zijun, et al.
Veröffentlicht: (2024)
von: Zhan, Zijun, et al.
Veröffentlicht: (2024)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
von: Wu, Xiangyang, et al.
Veröffentlicht: (2025)
von: Wu, Xiangyang, et al.
Veröffentlicht: (2025)
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
von: Ding, Sihao, et al.
Veröffentlicht: (2025)
von: Ding, Sihao, et al.
Veröffentlicht: (2025)
Subspace Alignment for Vision-Language Model Test-time Adaptation
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
When Visuals Aren't the Problem: Evaluating Vision-Language Models on Misleading Data Visualizations
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2026)
von: Lalai, Harsh Nishant, et al.
Veröffentlicht: (2026)
Transfer Learning Applied to Computer Vision Problems: Survey on Current Progress, Limitations, and Opportunities
von: Panda, Aaryan, et al.
Veröffentlicht: (2024)
von: Panda, Aaryan, et al.
Veröffentlicht: (2024)
Unsupervised Training of Vision Transformers with Synthetic Negatives
von: Giakoumoglou, Nikolaos, et al.
Veröffentlicht: (2025)
von: Giakoumoglou, Nikolaos, et al.
Veröffentlicht: (2025)
Transformers without Normalization
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
von: Zhu, Jiachen, et al.
Veröffentlicht: (2025)
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
von: Lu, Zhuqiang, et al.
Veröffentlicht: (2024)
von: Lu, Zhuqiang, et al.
Veröffentlicht: (2024)
Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
von: Jiang, Hao, et al.
Veröffentlicht: (2025)
ReasonEdit: Editing Vision-Language Models using Human Reasoning
von: Qiu, Jiaxing, et al.
Veröffentlicht: (2026)
von: Qiu, Jiaxing, et al.
Veröffentlicht: (2026)
Line of Sight: On Linear Representations in VLLMs
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2025)
von: Rajaram, Achyuta, et al.
Veröffentlicht: (2025)
ROCKET: Residual-Oriented Multi-Layer Alignment for Spatially-Aware Vision-Language-Action Models
von: Sun, Guoheng, et al.
Veröffentlicht: (2026)
von: Sun, Guoheng, et al.
Veröffentlicht: (2026)
From Fewer Samples to Fewer Bits: Reframing Dataset Distillation as Joint Optimization of Precision and Compactness
von: Dinh, My H., et al.
Veröffentlicht: (2026)
von: Dinh, My H., et al.
Veröffentlicht: (2026)
Adversarial Prompt Tuning for Vision-Language Models
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
Template-Guided Reconstruction of Pulmonary Segments with Neural Implicit Functions
von: Xie, Kangxian, et al.
Veröffentlicht: (2025)
von: Xie, Kangxian, et al.
Veröffentlicht: (2025)
(1D) Ordered Tokens Enable Efficient Test-Time Search
von: Gao, Zhitong, et al.
Veröffentlicht: (2026)
von: Gao, Zhitong, et al.
Veröffentlicht: (2026)
Animate Your Thoughts: Decoupled Reconstruction of Dynamic Natural Vision from Slow Brain Activity
von: Lu, Yizhuo, et al.
Veröffentlicht: (2024)
von: Lu, Yizhuo, et al.
Veröffentlicht: (2024)
Weierstrass Positional Encoding for Vision Transformers
von: Xin, Zhihang, et al.
Veröffentlicht: (2026)
von: Xin, Zhihang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Diffuse and Disperse: Image Generation with Representation Regularization
von: Wang, Runqian, et al.
Veröffentlicht: (2025) -
Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
von: Wang, Runqian, et al.
Veröffentlicht: (2025) -
Video-based Vehicle Surveillance in the Wild: License Plate, Make, and Model Recognition with Self Reflective Vision-Language Models
von: Parsa, Pouya, et al.
Veröffentlicht: (2025) -
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025) -
Bongards at the Boundary of Perception and Reasoning: Programs or Language?
von: Langenfeld, Cassidy, et al.
Veröffentlicht: (2026)