Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
Fuente:
arXiv
Salvato in:
| Autori principali: | Lepori, Michael A., Tartaglini, Alexa R., Vong, Wai Keen, Serre, Thomas, Lake, Brenden M., Pavlick, Ellie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Deep Neural Networks Can Learn Generalizable Same-Different Visual Relations
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2023)
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2023)
I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2026)
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2026)
Uncovering Intermediate Variables in Transformers using Circuit Probing
di: Lepori, Michael A., et al.
Pubblicazione: (2023)
di: Lepori, Michael A., et al.
Pubblicazione: (2023)
On the robustness of modeling grounded word learning through a child's egocentric input
di: Vong, Wai Keen, et al.
Pubblicazione: (2025)
di: Vong, Wai Keen, et al.
Pubblicazione: (2025)
Video Finetuning Improves Reasoning Between Frames
di: Yang, Ruiqi, et al.
Pubblicazione: (2025)
di: Yang, Ruiqi, et al.
Pubblicazione: (2025)
How Do Vision-Language Models Process Conflicting Information Across Modalities?
di: Hua, Tianze, et al.
Pubblicazione: (2025)
di: Hua, Tianze, et al.
Pubblicazione: (2025)
From Prediction to Understanding: Will AI Foundation Models Transform Brain Science?
di: Serre, Thomas, et al.
Pubblicazione: (2025)
di: Serre, Thomas, et al.
Pubblicazione: (2025)
Representative Attention For Vision Transformers
di: Li, Yuntong, et al.
Pubblicazione: (2026)
di: Li, Yuntong, et al.
Pubblicazione: (2026)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2025)
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2025)
Capability $\neq$ Interpretability: Human Interpretability of Vision Foundation Models
di: Colin, Julien, et al.
Pubblicazione: (2026)
di: Colin, Julien, et al.
Pubblicazione: (2026)
PLOOD: Partial Label Learning with Out-of-distribution Objects
di: Huang, Jintao, et al.
Pubblicazione: (2024)
di: Huang, Jintao, et al.
Pubblicazione: (2024)
ObjectTransforms for Uncertainty Quantification and Reduction in Vision-Based Perception for Autonomous Vehicles
di: Sahu, Nishad, et al.
Pubblicazione: (2025)
di: Sahu, Nishad, et al.
Pubblicazione: (2025)
AnyDoor: Zero-shot Object-level Image Customization
di: Chen, Xi, et al.
Pubblicazione: (2023)
di: Chen, Xi, et al.
Pubblicazione: (2023)
FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph Diffusion
di: Feng, Chen-Bin, et al.
Pubblicazione: (2026)
di: Feng, Chen-Bin, et al.
Pubblicazione: (2026)
The Components of Collaborative Joint Perception and Prediction -- A Conceptual Framework
di: Wan, Lei, et al.
Pubblicazione: (2025)
di: Wan, Lei, et al.
Pubblicazione: (2025)
Evaluating Graphical Perception Capabilities of Vision Transformers
di: Poonam, Poonam, et al.
Pubblicazione: (2026)
di: Poonam, Poonam, et al.
Pubblicazione: (2026)
Weakly-supervised Semantic Segmentation via Dual-stream Contrastive Learning of Cross-image Contextual Information
di: Lai, Qi, et al.
Pubblicazione: (2024)
di: Lai, Qi, et al.
Pubblicazione: (2024)
Systematic Literature Review on Vehicular Collaborative Perception -- A Computer Vision Perspective
di: Wan, Lei, et al.
Pubblicazione: (2025)
di: Wan, Lei, et al.
Pubblicazione: (2025)
Beyond First-Order: Learning Riemannian Geometries for Invariant Visual Place Recognition
di: Cheng, Jintao, et al.
Pubblicazione: (2026)
di: Cheng, Jintao, et al.
Pubblicazione: (2026)
Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric Vision
di: Jeong, David C., et al.
Pubblicazione: (2025)
di: Jeong, David C., et al.
Pubblicazione: (2025)
ViTOC: Vision Transformer and Object-aware Captioner
di: Huang, Feiyang
Pubblicazione: (2024)
di: Huang, Feiyang
Pubblicazione: (2024)
Object Referring-Guided Scanpath Prediction with Perception-Enhanced Vision-Language Models
di: Quan, Rong, et al.
Pubblicazione: (2026)
di: Quan, Rong, et al.
Pubblicazione: (2026)
Temporal-Spatial Object Relations Modeling for Vision-and-Language Navigation
di: Huang, Bowen, et al.
Pubblicazione: (2024)
di: Huang, Bowen, et al.
Pubblicazione: (2024)
DoorDet: Semi-Automated Multi-Class Door Detection Dataset via Object Detection and Large Language Models
di: Zhang, Licheng, et al.
Pubblicazione: (2025)
di: Zhang, Licheng, et al.
Pubblicazione: (2025)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
di: Lu, Jialin, et al.
Pubblicazione: (2024)
di: Lu, Jialin, et al.
Pubblicazione: (2024)
AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection
di: Lu, Jialin, et al.
Pubblicazione: (2025)
di: Lu, Jialin, et al.
Pubblicazione: (2025)
Beyond Grids: Exploring Elastic Input Sampling for Vision Transformers
di: Pardyl, Adam, et al.
Pubblicazione: (2023)
di: Pardyl, Adam, et al.
Pubblicazione: (2023)
R-LiViT: A LiDAR-Visual-Thermal Dataset Enabling Vulnerable Road User Focused Roadside Perception
di: Mirlach, Jonas, et al.
Pubblicazione: (2025)
di: Mirlach, Jonas, et al.
Pubblicazione: (2025)
DOPE: Dual Object Perception-Enhancement Network for Vision-and-Language Navigation
di: Yu, Yinfeng, et al.
Pubblicazione: (2025)
di: Yu, Yinfeng, et al.
Pubblicazione: (2025)
Seeing Beyond the Scene: Analyzing and Mitigating Background Bias in Action Recognition
di: Zhou, Ellie, et al.
Pubblicazione: (2025)
di: Zhou, Ellie, et al.
Pubblicazione: (2025)
Synchronized Object Detection for Autonomous Sorting, Mapping, and Quantification of Materials in Circular Healthcare
di: Zocco, Federico, et al.
Pubblicazione: (2024)
di: Zocco, Federico, et al.
Pubblicazione: (2024)
Can Transformers Capture Spatial Relations between Objects?
di: Wen, Chuan, et al.
Pubblicazione: (2024)
di: Wen, Chuan, et al.
Pubblicazione: (2024)
Retina Vision Transformer (RetinaViT): Introducing Scaled Patches into Vision Transformers
di: Shu, Yuyang, et al.
Pubblicazione: (2024)
di: Shu, Yuyang, et al.
Pubblicazione: (2024)
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
di: Gee, Leonidas, et al.
Pubblicazione: (2024)
A Simple yet Effective Network based on Vision Transformer for Camouflaged Object and Salient Object Detection
di: Hao, Chao, et al.
Pubblicazione: (2024)
di: Hao, Chao, et al.
Pubblicazione: (2024)
Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?
di: Cohen, Itay, et al.
Pubblicazione: (2025)
di: Cohen, Itay, et al.
Pubblicazione: (2025)
Choosing the right basis for interpretability: Psychophysical comparison between neuron-based and dictionary-based representations
di: Colin, Julien, et al.
Pubblicazione: (2024)
di: Colin, Julien, et al.
Pubblicazione: (2024)
3D Hand Mesh-Guided AI-Generated Malformed Hand Refinement with Hand Pose Transformation via Diffusion Model
di: Feng, Chen-Bin, et al.
Pubblicazione: (2025)
di: Feng, Chen-Bin, et al.
Pubblicazione: (2025)
Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
di: Traub, Manuel, et al.
Pubblicazione: (2025)
di: Traub, Manuel, et al.
Pubblicazione: (2025)
GUMBEL-NERF: Representing Unseen Objects as Part-Compositional Neural Radiance Fields
di: Sekikawa, Yusuke, et al.
Pubblicazione: (2024)
di: Sekikawa, Yusuke, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Deep Neural Networks Can Learn Generalizable Same-Different Visual Relations
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2023) -
I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers
di: Tartaglini, Alexa R., et al.
Pubblicazione: (2026) -
Uncovering Intermediate Variables in Transformers using Circuit Probing
di: Lepori, Michael A., et al.
Pubblicazione: (2023) -
On the robustness of modeling grounded word learning through a child's egocentric input
di: Vong, Wai Keen, et al.
Pubblicazione: (2025) -
Video Finetuning Improves Reasoning Between Frames
di: Yang, Ruiqi, et al.
Pubblicazione: (2025)