Tackling the Abstraction and Reasoning Corpus with Vision Transformers: the Importance of 2D Representation, Positions, and Objects
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Wenhao, Xu, Yudong, Sanner, Scott, Khalil, Elias Boutros |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LLMs and the Abstraction and Reasoning Corpus: Successes, Failures, and the Importance of Object-based Representations
por: Xu, Yudong, et al.
Publicado: (2023)
por: Xu, Yudong, et al.
Publicado: (2023)
A 2D Semantic-Aware Position Encoding for Vision Transformers
por: Chen, Xi, et al.
Publicado: (2025)
por: Chen, Xi, et al.
Publicado: (2025)
Two-stage Vision Transformers and Hard Masking offer Robust Object Representations
por: Aniraj, Ananthu, et al.
Publicado: (2025)
por: Aniraj, Ananthu, et al.
Publicado: (2025)
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
por: Xu, Xuwei, et al.
Publicado: (2025)
por: Xu, Xuwei, et al.
Publicado: (2025)
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
por: Xu, Xuwei, et al.
Publicado: (2023)
por: Xu, Xuwei, et al.
Publicado: (2023)
Self-Supervised Transformers as Iterative Solution Improvers for Constraint Satisfaction
por: Xu, Yudong W., et al.
Publicado: (2025)
por: Xu, Yudong W., et al.
Publicado: (2025)
C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning
por: Ye, Guanting, et al.
Publicado: (2026)
por: Ye, Guanting, et al.
Publicado: (2026)
Weierstrass Positional Encoding for Vision Transformers
por: Xin, Zhihang, et al.
Publicado: (2026)
por: Xin, Zhihang, et al.
Publicado: (2026)
Representation Learning with Adaptive Superpixel Coding
por: Khalil, Mahmoud, et al.
Publicado: (2025)
por: Khalil, Mahmoud, et al.
Publicado: (2025)
Hybrid Spiking Vision Transformer for Object Detection with Event Cameras
por: Xu, Qi, et al.
Publicado: (2025)
por: Xu, Qi, et al.
Publicado: (2025)
Human-Like Coarse Object Representations in Vision Models
por: Gizdov, Andrey, et al.
Publicado: (2026)
por: Gizdov, Andrey, et al.
Publicado: (2026)
ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
por: Khalil, Ahmad, et al.
Publicado: (2025)
por: Khalil, Ahmad, et al.
Publicado: (2025)
CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
por: Pothiraj, Atin, et al.
Publicado: (2025)
por: Pothiraj, Atin, et al.
Publicado: (2025)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
por: Yu, Jiangyong, et al.
Publicado: (2025)
por: Yu, Jiangyong, et al.
Publicado: (2025)
Representation Separation for Semantic Segmentation with Vision Transformers
por: Hong, Yuanduo, et al.
Publicado: (2022)
por: Hong, Yuanduo, et al.
Publicado: (2022)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
por: Kumar, Yogesh, et al.
Publicado: (2025)
por: Kumar, Yogesh, et al.
Publicado: (2025)
On Extending Semantic Abstraction for Efficient Search of Hidden Objects
por: Pais, Tasha, et al.
Publicado: (2025)
por: Pais, Tasha, et al.
Publicado: (2025)
OCTrack: Benchmarking the Open-Corpus Multi-Object Tracking
por: Qian, Zekun, et al.
Publicado: (2024)
por: Qian, Zekun, et al.
Publicado: (2024)
Synthesizing 3D Abstractions by Inverting Procedural Buildings with Transformers
por: Dax, Maximilian, et al.
Publicado: (2025)
por: Dax, Maximilian, et al.
Publicado: (2025)
MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models
por: Cai, Huanqia, et al.
Publicado: (2025)
por: Cai, Huanqia, et al.
Publicado: (2025)
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
por: Chowdhury, Md Abtahi Majeed, et al.
Publicado: (2025)
por: Chowdhury, Md Abtahi Majeed, et al.
Publicado: (2025)
Mitigating Bias with Words: Inducing Demographic Ambiguity in Face Recognition Templates by Text Encoding
por: Chettaoui, Tahar, et al.
Publicado: (2025)
por: Chettaoui, Tahar, et al.
Publicado: (2025)
Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
por: Lepori, Michael A., et al.
Publicado: (2024)
por: Lepori, Michael A., et al.
Publicado: (2024)
Keypoint Abstraction using Large Models for Object-Relative Imitation Learning
por: Fang, Xiaolin, et al.
Publicado: (2024)
por: Fang, Xiaolin, et al.
Publicado: (2024)
Mind's Eye: A Benchmark of Visual Abstraction, Transformation and Composition for Multimodal LLMs
por: Sinha, Rohit, et al.
Publicado: (2026)
por: Sinha, Rohit, et al.
Publicado: (2026)
Quantum Inverse Contextual Vision Transformers (Q-ICVT): A New Frontier in 3D Object Detection for AVs
por: Dharavath, Sanjay Bhargav, et al.
Publicado: (2024)
por: Dharavath, Sanjay Bhargav, et al.
Publicado: (2024)
The Progression of Transformers from Language to Vision to MOT: A Literature Review on Multi-Object Tracking with Transformers
por: Kamboj, Abhi
Publicado: (2024)
por: Kamboj, Abhi
Publicado: (2024)
AYDIV: Adaptable Yielding 3D Object Detection via Integrated Contextual Vision Transformer
por: Dam, Tanmoy, et al.
Publicado: (2024)
por: Dam, Tanmoy, et al.
Publicado: (2024)
CityCube: Benchmarking Cross-view Spatial Reasoning on Vision-Language Models in Urban Environments
por: Xu, Haotian, et al.
Publicado: (2026)
por: Xu, Haotian, et al.
Publicado: (2026)
CRD: Collaborative Representation Distance for Practical Anomaly Detection
por: Han, Chao, et al.
Publicado: (2023)
por: Han, Chao, et al.
Publicado: (2023)
COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking
por: Zhang, Chunhui, et al.
Publicado: (2025)
por: Zhang, Chunhui, et al.
Publicado: (2025)
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
por: Li, Wenxi, et al.
Publicado: (2025)
por: Li, Wenxi, et al.
Publicado: (2025)
Surformer v1: Transformer-Based Surface Classification Using Tactile and Vision Features
por: Kansana, Manish, et al.
Publicado: (2025)
por: Kansana, Manish, et al.
Publicado: (2025)
CoRE: Concept-Reasoning Expansion for Continual Brain Lesion Segmentation
por: Chen, Qianqian, et al.
Publicado: (2026)
por: Chen, Qianqian, et al.
Publicado: (2026)
Cross-Axis Transformer with 3D Rotary Positional Embeddings
por: Erickson, Lily
Publicado: (2023)
por: Erickson, Lily
Publicado: (2023)
Extending 6D Object Pose Estimators for Stereo Vision
por: Pöllabauer, Thomas, et al.
Publicado: (2024)
por: Pöllabauer, Thomas, et al.
Publicado: (2024)
Learning Disentangled Representation in Object-Centric Models for Visual Dynamics Prediction via Transformers
por: Gandhi, Sanket, et al.
Publicado: (2024)
por: Gandhi, Sanket, et al.
Publicado: (2024)
Are Vision Transformer Representations Semantically Meaningful? A Case Study in Medical Imaging
por: Shams, Montasir, et al.
Publicado: (2025)
por: Shams, Montasir, et al.
Publicado: (2025)
Symbolic Rule Extraction from Attention-Guided Sparse Representations in Vision Transformers
por: Padalkar, Parth, et al.
Publicado: (2025)
por: Padalkar, Parth, et al.
Publicado: (2025)
IFViT: Interpretable Fixed-Length Representation for Fingerprint Matching via Vision Transformer
por: Qiu, Yuhang, et al.
Publicado: (2024)
por: Qiu, Yuhang, et al.
Publicado: (2024)
Ejemplares similares
-
LLMs and the Abstraction and Reasoning Corpus: Successes, Failures, and the Importance of Object-based Representations
por: Xu, Yudong, et al.
Publicado: (2023) -
A 2D Semantic-Aware Position Encoding for Vision Transformers
por: Chen, Xi, et al.
Publicado: (2025) -
Two-stage Vision Transformers and Hard Masking offer Robust Object Representations
por: Aniraj, Ananthu, et al.
Publicado: (2025) -
RePaViT: Scalable Vision Transformer Acceleration via Structural Reparameterization on Feedforward Network Layers
por: Xu, Xuwei, et al.
Publicado: (2025) -
GTP-ViT: Efficient Vision Transformers via Graph-based Token Propagation
por: Xu, Xuwei, et al.
Publicado: (2023)