I Walk the Line: Examining the Role of Gestalt Continuity in Object Binding for Vision Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Tartaglini, Alexa R., Lepori, Michael A. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
by: Lepori, Michael A., et al.
Published: (2024)
by: Lepori, Michael A., et al.
Published: (2024)
Deep Neural Networks Can Learn Generalizable Same-Different Visual Relations
by: Tartaglini, Alexa R., et al.
Published: (2023)
by: Tartaglini, Alexa R., et al.
Published: (2023)
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
by: Tartaglini, Alexa R., et al.
Published: (2025)
by: Tartaglini, Alexa R., et al.
Published: (2025)
Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
by: Li, Yihao, et al.
Published: (2025)
by: Li, Yihao, et al.
Published: (2025)
From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models
by: Li, Tianqin, et al.
Published: (2025)
by: Li, Tianqin, et al.
Published: (2025)
WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
by: Yuan, Zhiqiang, et al.
Published: (2024)
by: Yuan, Zhiqiang, et al.
Published: (2024)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
Attention Retention for Continual Learning with Vision Transformers
by: Lu, Yue, et al.
Published: (2026)
by: Lu, Yue, et al.
Published: (2026)
Hybrid Spiking Vision Transformer for Object Detection with Event Cameras
by: Xu, Qi, et al.
Published: (2025)
by: Xu, Qi, et al.
Published: (2025)
Object-centric Binding in Contrastive Language-Image Pretraining
by: Assouel, Rim, et al.
Published: (2025)
by: Assouel, Rim, et al.
Published: (2025)
Investigating Mechanisms for In-Context Vision Language Binding
by: Saravanan, Darshana, et al.
Published: (2025)
by: Saravanan, Darshana, et al.
Published: (2025)
The Progression of Transformers from Language to Vision to MOT: A Literature Review on Multi-Object Tracking with Transformers
by: Kamboj, Abhi
Published: (2024)
by: Kamboj, Abhi
Published: (2024)
COST: Contrastive One-Stage Transformer for Vision-Language Small Object Tracking
by: Zhang, Chunhui, et al.
Published: (2025)
by: Zhang, Chunhui, et al.
Published: (2025)
SparseFormer: Detecting Objects in HRW Shots via Sparse Vision Transformer
by: Li, Wenxi, et al.
Published: (2025)
by: Li, Wenxi, et al.
Published: (2025)
Two-stage Vision Transformers and Hard Masking offer Robust Object Representations
by: Aniraj, Ananthu, et al.
Published: (2025)
by: Aniraj, Ananthu, et al.
Published: (2025)
Tackling the Abstraction and Reasoning Corpus with Vision Transformers: the Importance of 2D Representation, Positions, and Objects
by: Li, Wenhao, et al.
Published: (2024)
by: Li, Wenhao, et al.
Published: (2024)
Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
by: Vu, Tuan-Anh, et al.
Published: (2025)
by: Vu, Tuan-Anh, et al.
Published: (2025)
Walk and Read Less: Improving the Efficiency of Vision-and-Language Navigation via Tuning-Free Multimodal Token Pruning
by: Qin, Wenda, et al.
Published: (2025)
by: Qin, Wenda, et al.
Published: (2025)
Object-Centric Vision Token Pruning for Vision Language Models
by: Li, Guangyuan, et al.
Published: (2025)
by: Li, Guangyuan, et al.
Published: (2025)
Continual Adaptation of Vision Transformers for Federated Learning
by: Halbe, Shaunak, et al.
Published: (2023)
by: Halbe, Shaunak, et al.
Published: (2023)
Bi-Orthogonal Factor Decomposition for Vision Transformers
by: Doshi, Fenil R., et al.
Published: (2026)
by: Doshi, Fenil R., et al.
Published: (2026)
Dynamic Object Queries for Transformer-based Incremental Object Detection
by: Zhang, Jichuan, et al.
Published: (2024)
by: Zhang, Jichuan, et al.
Published: (2024)
Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
by: Vlachogiannis, Dimitrios N., et al.
Published: (2025)
by: Vlachogiannis, Dimitrios N., et al.
Published: (2025)
physfusion: A Transformer-based Dual-Stream Radar and Vision Fusion Framework for Open Water Surface Object Detection
by: Wan, Yuting, et al.
Published: (2026)
by: Wan, Yuting, et al.
Published: (2026)
Quantum Inverse Contextual Vision Transformers (Q-ICVT): A New Frontier in 3D Object Detection for AVs
by: Dharavath, Sanjay Bhargav, et al.
Published: (2024)
by: Dharavath, Sanjay Bhargav, et al.
Published: (2024)
Efficient Parameter Mining and Freezing for Continual Object Detection
by: Menezes, Angelo G., et al.
Published: (2024)
by: Menezes, Angelo G., et al.
Published: (2024)
Continual Vision-and-Language Navigation
by: Jeong, Seongjun, et al.
Published: (2024)
by: Jeong, Seongjun, et al.
Published: (2024)
Human-like Object Grouping in Self-supervised Vision Transformers
by: Adeli, Hossein, et al.
Published: (2026)
by: Adeli, Hossein, et al.
Published: (2026)
Knowledge Amalgamation for Object Detection with Transformers
by: Zhang, Haofei, et al.
Published: (2022)
by: Zhang, Haofei, et al.
Published: (2022)
VLM's Eye Examination: Instruct and Inspect Visual Competency of Vision Language Models
by: Hyeon-Woo, Nam, et al.
Published: (2024)
by: Hyeon-Woo, Nam, et al.
Published: (2024)
Leveraging Computer Vision in the Intensive Care Unit (ICU) for Examining Visitation and Mobility
by: Siegel, Scott, et al.
Published: (2024)
by: Siegel, Scott, et al.
Published: (2024)
Benchmarking Unlearning for Vision Transformers
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Latent Distillation for Continual Object Detection at the Edge
by: Pasti, Francesco, et al.
Published: (2024)
by: Pasti, Francesco, et al.
Published: (2024)
Vision Bridge Transformer at Scale
by: Tan, Zhenxiong, et al.
Published: (2025)
by: Tan, Zhenxiong, et al.
Published: (2025)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
AYDIV: Adaptable Yielding 3D Object Detection via Integrated Contextual Vision Transformer
by: Dam, Tanmoy, et al.
Published: (2024)
by: Dam, Tanmoy, et al.
Published: (2024)
An Examination of Offline-Trained Encoders in Vision-Based Deep Reinforcement Learning for Autonomous Driving
by: Mohammed, Shawan, et al.
Published: (2024)
by: Mohammed, Shawan, et al.
Published: (2024)
DinoTwins: Combining DINO and Barlow Twins for Robust, Label-Efficient Vision Transformers
by: Podsiadly, Michael, et al.
Published: (2025)
by: Podsiadly, Michael, et al.
Published: (2025)
An Examination of the Compositionality of Large Generative Vision-Language Models
by: Ma, Teli, et al.
Published: (2023)
by: Ma, Teli, et al.
Published: (2023)
Source-Free Object Detection with Detection Transformer
by: Yao, Huizai, et al.
Published: (2025)
by: Yao, Huizai, et al.
Published: (2025)
Similar Items
-
Beyond the Doors of Perception: Vision Transformers Represent Relations Between Objects
by: Lepori, Michael A., et al.
Published: (2024) -
Deep Neural Networks Can Learn Generalizable Same-Different Visual Relations
by: Tartaglini, Alexa R., et al.
Published: (2023) -
Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
by: Tartaglini, Alexa R., et al.
Published: (2025) -
Does Object Binding Naturally Emerge in Large Pretrained Vision Transformers?
by: Li, Yihao, et al.
Published: (2025) -
From Local Cues to Global Percepts: Emergent Gestalt Organization in Self-Supervised Vision Models
by: Li, Tianqin, et al.
Published: (2025)