VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Norouzi, Narges, Zulfikar, Idil Esen, Cavagnero, Niccolò, Kerssies, Tommie, Leibe, Bastian, Dubbelman, Gijs, de Geus, Daan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Your ViT is Secretly an Image Segmentation Model
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2026)
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2026)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
von: Norouzi, Narges, et al.
Veröffentlicht: (2024)
von: Norouzi, Narges, et al.
Veröffentlicht: (2024)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
von: Englert, Brunó B., et al.
Veröffentlicht: (2024)
von: Englert, Brunó B., et al.
Veröffentlicht: (2024)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
von: de Geus, Daan, et al.
Veröffentlicht: (2024)
von: de Geus, Daan, et al.
Veröffentlicht: (2024)
Point-VOS: Pointing Up Video Object Segmentation
von: Zulfikar, Idil Esen, et al.
Veröffentlicht: (2024)
von: Zulfikar, Idil Esen, et al.
Veröffentlicht: (2024)
What is the Added Value of UDA in the VFM Era?
von: Englert, Brunó B., et al.
Veröffentlicht: (2025)
von: Englert, Brunó B., et al.
Veröffentlicht: (2025)
Interactive4D: Interactive 4D LiDAR Segmentation
von: Fradlin, Ilya, et al.
Veröffentlicht: (2024)
von: Fradlin, Ilya, et al.
Veröffentlicht: (2024)
Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
von: Orlova, Svetlana, et al.
Veröffentlicht: (2026)
von: Orlova, Svetlana, et al.
Veröffentlicht: (2026)
Orion-Lite: Distilling LLM Reasoning into Efficient Vision-Only Driving Models
von: Gu, Jing, et al.
Veröffentlicht: (2026)
von: Gu, Jing, et al.
Veröffentlicht: (2026)
Simplifying Traffic Anomaly Detection with Video Foundation Models
von: Orlova, Svetlana, et al.
Veröffentlicht: (2025)
von: Orlova, Svetlana, et al.
Veröffentlicht: (2025)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
von: Kerssies, Tommie, et al.
Veröffentlicht: (2026)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2026)
DONUT: A Decoder-Only Model for Trajectory Prediction
von: Knoche, Markus, et al.
Veröffentlicht: (2025)
von: Knoche, Markus, et al.
Veröffentlicht: (2025)
Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2025)
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2025)
How Important are Videos for Training Video LLMs?
von: Lydakis, George, et al.
Veröffentlicht: (2025)
von: Lydakis, George, et al.
Veröffentlicht: (2025)
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
von: Yilmaz, Kadir, et al.
Veröffentlicht: (2026)
von: Yilmaz, Kadir, et al.
Veröffentlicht: (2026)
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
von: Knaebel, Karim, et al.
Veröffentlicht: (2025)
von: Knaebel, Karim, et al.
Veröffentlicht: (2025)
The BRAVO Semantic Segmentation Challenge Results in UNCV2024
von: Vu, Tuan-Hung, et al.
Veröffentlicht: (2024)
von: Vu, Tuan-Hung, et al.
Veröffentlicht: (2024)
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
von: Garcia, Gonzalo Martin, et al.
Veröffentlicht: (2024)
von: Garcia, Gonzalo Martin, et al.
Veröffentlicht: (2024)
VFM-UDA++: Improving Network Architectures and Data Strategies for Unsupervised Domain Adaptive Semantic Segmentation
von: Englert, Brunó B., et al.
Veröffentlicht: (2025)
von: Englert, Brunó B., et al.
Veröffentlicht: (2025)
SurGe: Improved Surface Geometry in Point Maps
von: Knaebel, Karim, et al.
Veröffentlicht: (2026)
von: Knaebel, Karim, et al.
Veröffentlicht: (2026)
ViT Registers and Fractal ViT
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2026)
von: Chou, Jason Chuan-Chih, et al.
Veröffentlicht: (2026)
The revenge of BiSeNet: Efficient Multi-Task Image Segmentation
von: Rosi, Gabriele, et al.
Veröffentlicht: (2024)
von: Rosi, Gabriele, et al.
Veröffentlicht: (2024)
Vanilla ViT for Automotive Point Cloud Semantic Segmentation
von: Puy, Gilles, et al.
Veröffentlicht: (2026)
von: Puy, Gilles, et al.
Veröffentlicht: (2026)
Applying ViT in Generalized Few-shot Semantic Segmentation
von: Geng, Liyuan, et al.
Veröffentlicht: (2024)
von: Geng, Liyuan, et al.
Veröffentlicht: (2024)
Mask4Former: Mask Transformer for 4D Panoptic Segmentation
von: Yilmaz, Kadir, et al.
Veröffentlicht: (2023)
von: Yilmaz, Kadir, et al.
Veröffentlicht: (2023)
SE-Enhanced ViT and BiLSTM-Based Intrusion Detection for Secure IIoT and IoMT Environments
von: Gueriani, Afrah, et al.
Veröffentlicht: (2026)
von: Gueriani, Afrah, et al.
Veröffentlicht: (2026)
Searching on a Budget: HW-NAS with 10 Latency Probes
von: Capuano, Francesco, et al.
Veröffentlicht: (2025)
von: Capuano, Francesco, et al.
Veröffentlicht: (2025)
Deeper Inside Deep ViT
von: Hong, Sungrae
Veröffentlicht: (2025)
von: Hong, Sungrae
Veröffentlicht: (2025)
La experiencia directiva en tiempo de virus: memoria para construir lo que sigue
von: Ana Cavagnero
Veröffentlicht: (2022)
von: Ana Cavagnero
Veröffentlicht: (2022)
SPAR: Single-Pass Any-Resolution ViT for Open-vocabulary Segmentation
von: Kombol, Naomi, et al.
Veröffentlicht: (2026)
von: Kombol, Naomi, et al.
Veröffentlicht: (2026)
I&S-ViT: An Inclusive & Stable Method for Pushing the Limit of Post-Training ViTs Quantization
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
von: Zhong, Yunshan, et al.
Veröffentlicht: (2023)
CLAMP-ViT: Contrastive Data-Free Learning for Adaptive Post-Training Quantization of ViTs
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
ODE-ViT: Plug & Play Attention Layer from the Generalization of the ViT as an Ordinary Differential Equation
von: Riera, Carlos Boned, et al.
Veröffentlicht: (2025)
von: Riera, Carlos Boned, et al.
Veröffentlicht: (2025)
STRAP-ViT: Segregated Tokens with Randomized -- Transformations for Defense against Adversarial Patches in ViTs
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
von: Chattopadhyay, Nandish, et al.
Veröffentlicht: (2026)
A Resource Efficient Fusion Network for Object Detection in Bird's-Eye View using Camera and Raw Radar Data
von: Chandrasekaran, Kavin, et al.
Veröffentlicht: (2024)
von: Chandrasekaran, Kavin, et al.
Veröffentlicht: (2024)
REFNet++: Multi-Task Efficient Fusion of Camera and Radar Sensor Data in Bird's-Eye Polar View
von: Chandrasekaran, Kavin, et al.
Veröffentlicht: (2026)
von: Chandrasekaran, Kavin, et al.
Veröffentlicht: (2026)
ViTCAE: ViT-based Class-conditioned Autoencoder
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
von: Jebraeeli, Vahid, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Your ViT is Secretly an Image Segmentation Model
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025) -
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2026) -
How to Benchmark Vision Foundation Models for Semantic Segmentation?
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024) -
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024) -
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
von: Norouzi, Narges, et al.
Veröffentlicht: (2024)