How to Benchmark Vision Foundation Models for Semantic Segmentation?
Fuente:
arXiv
Salvato in:
| Autori principali: | Kerssies, Tommie, de Geus, Daan, Dubbelman, Gijs |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
di: Kerssies, Tommie, et al.
Pubblicazione: (2024)
di: Kerssies, Tommie, et al.
Pubblicazione: (2024)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
di: Englert, Brunó B., et al.
Pubblicazione: (2024)
di: Englert, Brunó B., et al.
Pubblicazione: (2024)
The BRAVO Semantic Segmentation Challenge Results in UNCV2024
di: Vu, Tuan-Hung, et al.
Pubblicazione: (2024)
di: Vu, Tuan-Hung, et al.
Pubblicazione: (2024)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
di: de Geus, Daan, et al.
Pubblicazione: (2024)
di: de Geus, Daan, et al.
Pubblicazione: (2024)
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
di: Norouzi, Narges, et al.
Pubblicazione: (2026)
di: Norouzi, Narges, et al.
Pubblicazione: (2026)
Your ViT is Secretly an Image Segmentation Model
di: Kerssies, Tommie, et al.
Pubblicazione: (2025)
di: Kerssies, Tommie, et al.
Pubblicazione: (2025)
Hyperspectral Adapter for Semantic Segmentation with Vision Foundation Models
di: Hurtado, Juana Valeria, et al.
Pubblicazione: (2025)
di: Hurtado, Juana Valeria, et al.
Pubblicazione: (2025)
Simplifying Traffic Anomaly Detection with Video Foundation Models
di: Orlova, Svetlana, et al.
Pubblicazione: (2025)
di: Orlova, Svetlana, et al.
Pubblicazione: (2025)
Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
di: Nezhurina, Marianna, et al.
Pubblicazione: (2025)
di: Nezhurina, Marianna, et al.
Pubblicazione: (2025)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
di: Norouzi, Narges, et al.
Pubblicazione: (2024)
di: Norouzi, Narges, et al.
Pubblicazione: (2024)
What is the Added Value of UDA in the VFM Era?
di: Englert, Brunó B., et al.
Pubblicazione: (2025)
di: Englert, Brunó B., et al.
Pubblicazione: (2025)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
di: Cavagnero, Niccolò, et al.
Pubblicazione: (2026)
di: Cavagnero, Niccolò, et al.
Pubblicazione: (2026)
A Frame is Worth One Token: Efficient Generative World Modeling with Delta Tokens
di: Kerssies, Tommie, et al.
Pubblicazione: (2026)
di: Kerssies, Tommie, et al.
Pubblicazione: (2026)
Theia: Distilling Diverse Vision Foundation Models for Robot Learning
di: Shang, Jinghuan, et al.
Pubblicazione: (2024)
di: Shang, Jinghuan, et al.
Pubblicazione: (2024)
BelHouse3D: A Benchmark Dataset for Assessing Occlusion Robustness in 3D Point Cloud Semantic Segmentation
di: Kumar, Umamaheswaran Raman, et al.
Pubblicazione: (2024)
di: Kumar, Umamaheswaran Raman, et al.
Pubblicazione: (2024)
DistortBench: Benchmarking Vision Language Models on Image Distortion Identification
di: Goyal, Divyanshu, et al.
Pubblicazione: (2026)
di: Goyal, Divyanshu, et al.
Pubblicazione: (2026)
Grounding Foundational Vision Models with 3D Human Poses for Robust Action Recognition
di: Babey, Nicholas, et al.
Pubblicazione: (2025)
di: Babey, Nicholas, et al.
Pubblicazione: (2025)
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
di: Zhang, Zhengshen, et al.
Pubblicazione: (2025)
di: Zhang, Zhengshen, et al.
Pubblicazione: (2025)
FoundObj: Self-supervised Foundation Models as Rewards for Label-free 3D Object Segmentation
di: Zhang, Zihui, et al.
Pubblicazione: (2026)
di: Zhang, Zihui, et al.
Pubblicazione: (2026)
ForecastOcc: Vision-based Semantic Occupancy Forecasting
di: Mohan, Riya, et al.
Pubblicazione: (2026)
di: Mohan, Riya, et al.
Pubblicazione: (2026)
SegXAL: Explainable Active Learning for Semantic Segmentation in Driving Scene Scenarios
di: Mandalika, Sriram, et al.
Pubblicazione: (2024)
di: Mandalika, Sriram, et al.
Pubblicazione: (2024)
MESSI: A Multi-Elevation Semantic Segmentation Image Dataset of an Urban Environment
di: Pinkovich, Barak, et al.
Pubblicazione: (2025)
di: Pinkovich, Barak, et al.
Pubblicazione: (2025)
DarkQA: Benchmarking Vision-Language Models on Visual-Primitive Question Answering in Low-Light Indoor Scenes
di: Park, Yohan, et al.
Pubblicazione: (2025)
di: Park, Yohan, et al.
Pubblicazione: (2025)
LIX: Implicitly Infusing Spatial Geometric Prior Knowledge into Visual Semantic Segmentation for Autonomous Driving
di: Guo, Sicen, et al.
Pubblicazione: (2024)
di: Guo, Sicen, et al.
Pubblicazione: (2024)
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
di: Tomilin, Tristan, et al.
Pubblicazione: (2025)
di: Tomilin, Tristan, et al.
Pubblicazione: (2025)
TAVIS: A Benchmark for Egocentric Active Vision and Anticipatory Gaze in Imitation Learning
di: Spigler, Giacomo
Pubblicazione: (2026)
di: Spigler, Giacomo
Pubblicazione: (2026)
LogoSP: Local-global Grouping of Superpoints for Unsupervised Semantic Segmentation of 3D Point Clouds
di: Zhang, Zihui, et al.
Pubblicazione: (2025)
di: Zhang, Zihui, et al.
Pubblicazione: (2025)
Benchmarking Feature Upsampling Methods for Vision Foundation Models using Interactive Segmentation
di: Havrylov, Volodymyr, et al.
Pubblicazione: (2025)
di: Havrylov, Volodymyr, et al.
Pubblicazione: (2025)
Cosmos World Foundation Model Platform for Physical AI
di: NVIDIA, et al.
Pubblicazione: (2025)
di: NVIDIA, et al.
Pubblicazione: (2025)
World Simulation with Video Foundation Models for Physical AI
di: NVIDIA, et al.
Pubblicazione: (2025)
di: NVIDIA, et al.
Pubblicazione: (2025)
Pedestrian Intention Prediction via Vision-Language Foundation Models
di: Azarmi, Mohsen, et al.
Pubblicazione: (2025)
di: Azarmi, Mohsen, et al.
Pubblicazione: (2025)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
di: Chow, Wei, et al.
Pubblicazione: (2025)
di: Chow, Wei, et al.
Pubblicazione: (2025)
Real-World Robot Applications of Foundation Models: A Review
di: Kawaharazuka, Kento, et al.
Pubblicazione: (2024)
di: Kawaharazuka, Kento, et al.
Pubblicazione: (2024)
Towards Natural Language-Driven Assembly Using Foundation Models
di: Joglekar, Omkar, et al.
Pubblicazione: (2024)
di: Joglekar, Omkar, et al.
Pubblicazione: (2024)
AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
di: Vasa, Santosh, et al.
Pubblicazione: (2025)
di: Vasa, Santosh, et al.
Pubblicazione: (2025)
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
di: Li, Qixiu, et al.
Pubblicazione: (2024)
di: Li, Qixiu, et al.
Pubblicazione: (2024)
RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
di: Liu, Songming, et al.
Pubblicazione: (2024)
di: Liu, Songming, et al.
Pubblicazione: (2024)
Advances in Multimodal Adaptation and Generalization: From Traditional Approaches to Foundation Models
di: Dong, Hao, et al.
Pubblicazione: (2025)
di: Dong, Hao, et al.
Pubblicazione: (2025)
Hybrid Training for Vision-Language-Action Models
di: Mazzaglia, Pietro, et al.
Pubblicazione: (2025)
di: Mazzaglia, Pietro, et al.
Pubblicazione: (2025)
Scaling Robot Policy Learning via Zero-Shot Labeling with Foundation Models
di: Blank, Nils, et al.
Pubblicazione: (2024)
di: Blank, Nils, et al.
Pubblicazione: (2024)
Documenti analoghi
-
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
di: Kerssies, Tommie, et al.
Pubblicazione: (2024) -
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
di: Englert, Brunó B., et al.
Pubblicazione: (2024) -
The BRAVO Semantic Segmentation Challenge Results in UNCV2024
di: Vu, Tuan-Hung, et al.
Pubblicazione: (2024) -
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
di: de Geus, Daan, et al.
Pubblicazione: (2024) -
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
di: Norouzi, Narges, et al.
Pubblicazione: (2026)