Looking Locally: Object-Centric Vision Transformers as Foundation Models for Efficient Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Traub, Manuel, Butz, Martin V. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Object Permanence from Videos via Latent Imaginations
von: Traub, Manuel, et al.
Veröffentlicht: (2023)
von: Traub, Manuel, et al.
Veröffentlicht: (2023)
Loci-Segmented: Improving Scene Segmentation Learning
von: Traub, Manuel, et al.
Veröffentlicht: (2023)
von: Traub, Manuel, et al.
Veröffentlicht: (2023)
Vector-Quantized Vision Foundation Models for Object-Centric Learning
von: Zhao, Rongzhen, et al.
Veröffentlicht: (2025)
von: Zhao, Rongzhen, et al.
Veröffentlicht: (2025)
Bootstrap Segmentation Foundation Model under Distribution Shift via Object-Centric Learning
von: Tang, Luyao, et al.
Veröffentlicht: (2024)
von: Tang, Luyao, et al.
Veröffentlicht: (2024)
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
von: Huang, Shaofei, et al.
Veröffentlicht: (2025)
von: Huang, Shaofei, et al.
Veröffentlicht: (2025)
PanSR: An Object-Centric Mask Transformer for Panoptic Segmentation
von: Žust, Lojze, et al.
Veröffentlicht: (2024)
von: Žust, Lojze, et al.
Veröffentlicht: (2024)
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2026)
Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
von: Mirjalili, Vahid, et al.
Veröffentlicht: (2025)
von: Mirjalili, Vahid, et al.
Veröffentlicht: (2025)
Rein++: Efficient Generalization and Adaptation for Semantic Segmentation with Vision Foundation Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
Look Before Acting: Enhancing Vision Foundation Representations for Vision-Language-Action Models
von: Luo, Yulin, et al.
Veröffentlicht: (2026)
von: Luo, Yulin, et al.
Veröffentlicht: (2026)
RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
von: Dong, Guanfang, et al.
Veröffentlicht: (2025)
von: Dong, Guanfang, et al.
Veröffentlicht: (2025)
Are Vision Foundation Models Foundational for Electron Microscopy Image Segmentation?
von: Fuster-Barceló, Caterina, et al.
Veröffentlicht: (2026)
von: Fuster-Barceló, Caterina, et al.
Veröffentlicht: (2026)
Appearance-Based Refinement for Object-Centric Motion Segmentation
von: Xie, Junyu, et al.
Veröffentlicht: (2023)
von: Xie, Junyu, et al.
Veröffentlicht: (2023)
SimPLR: A Simple and Plain Transformer for Efficient Object Detection and Segmentation
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023)
von: Nguyen, Duy-Kien, et al.
Veröffentlicht: (2023)
Object-Centric Vision Token Pruning for Vision Language Models
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
von: Li, Guangyuan, et al.
Veröffentlicht: (2025)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
von: Norouzi, Narges, et al.
Veröffentlicht: (2024)
von: Norouzi, Narges, et al.
Veröffentlicht: (2024)
Recasting Generic Pretrained Vision Transformers As Object-Centric Scene Encoders For Manipulation Policies
von: Qian, Jianing, et al.
Veröffentlicht: (2024)
von: Qian, Jianing, et al.
Veröffentlicht: (2024)
ZISVFM: Zero-Shot Object Instance Segmentation in Indoor Robotic Environments with Vision Foundation Models
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
von: Zhang, Ying, et al.
Veröffentlicht: (2025)
Annotation Free Semantic Segmentation with Vision Foundation Models
von: Seifi, Soroush, et al.
Veröffentlicht: (2024)
von: Seifi, Soroush, et al.
Veröffentlicht: (2024)
Mechanisms of Object Localization in Vision-Language Models
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2026)
von: Schaumlöffel, Timothy, et al.
Veröffentlicht: (2026)
Can Modern Vision Models Understand the Difference Between an Object and a Look-alike?
von: Cohen, Itay, et al.
Veröffentlicht: (2025)
von: Cohen, Itay, et al.
Veröffentlicht: (2025)
Classifier-Centric Adaptive Framework for Open-Vocabulary Camouflaged Object Segmentation
von: Zhang, Hanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Hanyu, et al.
Veröffentlicht: (2025)
EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation
von: Kim, Chanyoung, et al.
Veröffentlicht: (2024)
von: Kim, Chanyoung, et al.
Veröffentlicht: (2024)
Scalable Object Detection in the Car Interior With Vision Foundation Models
von: Schmidt, Sebastian, et al.
Veröffentlicht: (2025)
von: Schmidt, Sebastian, et al.
Veröffentlicht: (2025)
Evaluating Vision Foundation Models for Pixel and Object Classification in Microscopy
von: Teuber, Carolin, et al.
Veröffentlicht: (2026)
von: Teuber, Carolin, et al.
Veröffentlicht: (2026)
SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models
von: Dünkel, Olaf, et al.
Veröffentlicht: (2026)
von: Dünkel, Olaf, et al.
Veröffentlicht: (2026)
Token-Space Mask Prediction for Efficient Vision Transformer Segmentation
von: Galagain, Calvin, et al.
Veröffentlicht: (2026)
von: Galagain, Calvin, et al.
Veröffentlicht: (2026)
Zero-shot Object-Centric Instruction Following: Integrating Foundation Models with Traditional Navigation
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2024)
von: Raychaudhuri, Sonia, et al.
Veröffentlicht: (2024)
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
von: Hannan, Darryl, et al.
Veröffentlicht: (2025)
von: Hannan, Darryl, et al.
Veröffentlicht: (2025)
Self-Supervised Vision Transformers Are Efficient Segmentation Learners for Imperfect Labels
von: Lee, Seungho, et al.
Veröffentlicht: (2024)
von: Lee, Seungho, et al.
Veröffentlicht: (2024)
LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate
von: Fuller, Anthony, et al.
Veröffentlicht: (2024)
von: Fuller, Anthony, et al.
Veröffentlicht: (2024)
HeightFormer: Learning Height Prediction in Voxel Features for Roadside Vision Centric 3D Object Detection via Transformer
von: Zhang, Zhang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhang, et al.
Veröffentlicht: (2025)
Analysis of Object Detection Models for Tiny Object in Satellite Imagery: A Dataset-Centric Approach
von: PS, Kailas, et al.
Veröffentlicht: (2024)
von: PS, Kailas, et al.
Veröffentlicht: (2024)
Unbiased Semantic Decoding with Vision Foundation Models for Few-shot Segmentation
von: Wang, Jin, et al.
Veröffentlicht: (2025)
von: Wang, Jin, et al.
Veröffentlicht: (2025)
Adapting Vision Foundation Models for Real-time Ultrasound Image Segmentation
von: Zhang, Xiaoran, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoran, et al.
Veröffentlicht: (2025)
Set Pivot Learning: Redefining Generalized Segmentation with Vision Foundation Models
von: Li, Xinhui, et al.
Veröffentlicht: (2025)
von: Li, Xinhui, et al.
Veröffentlicht: (2025)
GuiDINO: Rethinking Vision Foundation Model in Medical Image Segmentation
von: Liang, Zhuonan, et al.
Veröffentlicht: (2026)
von: Liang, Zhuonan, et al.
Veröffentlicht: (2026)
Generalizable Knowledge Distillation from Vision Foundation Models for Semantic Segmentation
von: Lv, Chonghua, et al.
Veröffentlicht: (2026)
von: Lv, Chonghua, et al.
Veröffentlicht: (2026)
A Novel Vision Transformer for Camera-LiDAR Fusion based Traffic Object Segmentation
von: Tahves, Toomas, et al.
Veröffentlicht: (2025)
von: Tahves, Toomas, et al.
Veröffentlicht: (2025)
Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
von: Zhao, Kai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning Object Permanence from Videos via Latent Imaginations
von: Traub, Manuel, et al.
Veröffentlicht: (2023) -
Loci-Segmented: Improving Scene Segmentation Learning
von: Traub, Manuel, et al.
Veröffentlicht: (2023) -
Vector-Quantized Vision Foundation Models for Object-Centric Learning
von: Zhao, Rongzhen, et al.
Veröffentlicht: (2025) -
Bootstrap Segmentation Foundation Model under Distribution Shift via Object-Centric Learning
von: Tang, Luyao, et al.
Veröffentlicht: (2024) -
Revisiting Audio-Visual Segmentation with Vision-Centric Transformer
von: Huang, Shaofei, et al.
Veröffentlicht: (2025)