DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Knaebel, Karim, Yilmaz, Kadir, de Geus, Daan, Hermans, Alexander, Adrian, David, Linder, Timm, Leibe, Bastian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
von: Garcia, Gonzalo Martin, et al.
Veröffentlicht: (2024)
von: Garcia, Gonzalo Martin, et al.
Veröffentlicht: (2024)
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
von: Yilmaz, Kadir, et al.
Veröffentlicht: (2026)
von: Yilmaz, Kadir, et al.
Veröffentlicht: (2026)
Point2Vec for Self-Supervised Representation Learning on Point Clouds
von: Knaebel, Karim, et al.
Veröffentlicht: (2023)
von: Knaebel, Karim, et al.
Veröffentlicht: (2023)
Acquisition of high-quality images for camera calibration in robotics applications via speech prompts
von: Linder, Timm, et al.
Veröffentlicht: (2025)
von: Linder, Timm, et al.
Veröffentlicht: (2025)
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
von: Piekenbrinck, Jens, et al.
Veröffentlicht: (2025)
von: Piekenbrinck, Jens, et al.
Veröffentlicht: (2025)
Sa2VA-i: Improving Sa2VA Results with Consistent Training and Inference
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2025)
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2025)
How Important are Videos for Training Video LLMs?
von: Lydakis, George, et al.
Veröffentlicht: (2025)
von: Lydakis, George, et al.
Veröffentlicht: (2025)
DONUT: A Decoder-Only Model for Trajectory Prediction
von: Knoche, Markus, et al.
Veröffentlicht: (2025)
von: Knoche, Markus, et al.
Veröffentlicht: (2025)
Mask4Former: Mask Transformer for 4D Panoptic Segmentation
von: Yilmaz, Kadir, et al.
Veröffentlicht: (2023)
von: Yilmaz, Kadir, et al.
Veröffentlicht: (2023)
SurGe: Improved Surface Geometry in Point Maps
von: Knaebel, Karim, et al.
Veröffentlicht: (2026)
von: Knaebel, Karim, et al.
Veröffentlicht: (2026)
Your ViT is Secretly an Image Segmentation Model
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2025)
Interactive4D: Interactive 4D LiDAR Segmentation
von: Fradlin, Ilya, et al.
Veröffentlicht: (2024)
von: Fradlin, Ilya, et al.
Veröffentlicht: (2024)
VidEoMT: Your ViT is Secretly Also a Video Segmentation Model
von: Norouzi, Narges, et al.
Veröffentlicht: (2026)
von: Norouzi, Narges, et al.
Veröffentlicht: (2026)
OoDIS: Anomaly Instance Segmentation and Detection Benchmark
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2024)
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2024)
How to Benchmark Vision Foundation Models for Semantic Segmentation?
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
MaskTerial: A Foundation Model for Automated 2D Material Flake Detection
von: Uslu, Jan-Lucas, et al.
Veröffentlicht: (2024)
von: Uslu, Jan-Lucas, et al.
Veröffentlicht: (2024)
Task-aligned Part-aware Panoptic Segmentation through Joint Object-Part Representations
von: de Geus, Daan, et al.
Veröffentlicht: (2024)
von: de Geus, Daan, et al.
Veröffentlicht: (2024)
First Place Solution to the ECCV 2024 BRAVO Challenge: Evaluating Robustness of Vision Foundation Models for Semantic Segmentation
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
von: Kerssies, Tommie, et al.
Veröffentlicht: (2024)
Systematic Comparison of Projection Methods for Monocular 3D Human Pose Estimation on Fisheye Images
von: Käs, Stephanie, et al.
Veröffentlicht: (2025)
von: Käs, Stephanie, et al.
Veröffentlicht: (2025)
Spotting the Unexpected (STU): A 3D LiDAR Dataset for Anomaly Segmentation in Autonomous Driving
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2025)
von: Nekrasov, Alexey, et al.
Veröffentlicht: (2025)
AGILE3D: Attention Guided Interactive Multi-object 3D Segmentation
von: Yue, Yuanwen, et al.
Veröffentlicht: (2023)
von: Yue, Yuanwen, et al.
Veröffentlicht: (2023)
OCCUQ: Exploring Efficient Uncertainty Quantification for 3D Occupancy Prediction
von: Heidrich, Severin, et al.
Veröffentlicht: (2025)
von: Heidrich, Severin, et al.
Veröffentlicht: (2025)
UPTor: Unified 3D Human Pose Dynamics and Trajectory Prediction for Human-Robot Interaction
von: Nilavadi, Nisarga, et al.
Veröffentlicht: (2025)
von: Nilavadi, Nisarga, et al.
Veröffentlicht: (2025)
PMT: Plain Mask Transformer for Image and Video Segmentation with Frozen Vision Encoders
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2026)
von: Cavagnero, Niccolò, et al.
Veröffentlicht: (2026)
ALGM: Adaptive Local-then-Global Token Merging for Efficient Semantic Segmentation with Plain Vision Transformers
von: Norouzi, Narges, et al.
Veröffentlicht: (2024)
von: Norouzi, Narges, et al.
Veröffentlicht: (2024)
OpenMaskDINO3D : Reasoning 3D Segmentation via Large Language Model
von: Zhang, Kunshen
Veröffentlicht: (2025)
von: Zhang, Kunshen
Veröffentlicht: (2025)
Exploring the Benefits of Vision Foundation Models for Unsupervised Domain Adaptation
von: Englert, Brunó B., et al.
Veröffentlicht: (2024)
von: Englert, Brunó B., et al.
Veröffentlicht: (2024)
Point-VOS: Pointing Up Video Object Segmentation
von: Zulfikar, Idil Esen, et al.
Veröffentlicht: (2024)
von: Zulfikar, Idil Esen, et al.
Veröffentlicht: (2024)
GuiDINO: Rethinking Vision Foundation Model in Medical Image Segmentation
von: Liang, Zhuonan, et al.
Veröffentlicht: (2026)
von: Liang, Zhuonan, et al.
Veröffentlicht: (2026)
An Ordinal Regression Framework for a Deep Learning Based Severity Assessment for Chest Radiographs
von: Wienholt, Patrick, et al.
Veröffentlicht: (2024)
von: Wienholt, Patrick, et al.
Veröffentlicht: (2024)
MonoDINO-DETR: Depth-Enhanced Monocular 3D Object Detection Using a Vision Foundation Model
von: Kim, Jihyeok, et al.
Veröffentlicht: (2025)
von: Kim, Jihyeok, et al.
Veröffentlicht: (2025)
Look Gauss, No Pose: Novel View Synthesis using Gaussian Splatting without Accurate Pose Initialization
von: Schmidt, Christian, et al.
Veröffentlicht: (2024)
von: Schmidt, Christian, et al.
Veröffentlicht: (2024)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features
von: Qu, Jinyuan, et al.
Veröffentlicht: (2025)
von: Qu, Jinyuan, et al.
Veröffentlicht: (2025)
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
von: Yang, Sicheng, et al.
Veröffentlicht: (2025)
von: Yang, Sicheng, et al.
Veröffentlicht: (2025)
DINO-VO: A Feature-based Visual Odometry Leveraging a Visual Foundation Model
von: Azhari, Maulana Bisyir, et al.
Veröffentlicht: (2025)
von: Azhari, Maulana Bisyir, et al.
Veröffentlicht: (2025)
How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction?
von: Käs, Stephanie, et al.
Veröffentlicht: (2025)
von: Käs, Stephanie, et al.
Veröffentlicht: (2025)
vesselFM: A Foundation Model for Universal 3D Blood Vessel Segmentation
von: Wittmann, Bastian, et al.
Veröffentlicht: (2024)
von: Wittmann, Bastian, et al.
Veröffentlicht: (2024)
SegFormer3D: an Efficient Transformer for 3D Medical Image Segmentation
von: Perera, Shehan, et al.
Veröffentlicht: (2024)
von: Perera, Shehan, et al.
Veröffentlicht: (2024)
DINO-SLAM: DINO-informed RGB-D SLAM for Neural Implicit and Explicit Representations
von: Gong, Ziren, et al.
Veröffentlicht: (2025)
von: Gong, Ziren, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fine-Tuning Image-Conditional Diffusion Models is Easier than You Think
von: Garcia, Gonzalo Martin, et al.
Veröffentlicht: (2024) -
Volume Transformer: Revisiting Vanilla Transformers for 3D Scene Understanding
von: Yilmaz, Kadir, et al.
Veröffentlicht: (2026) -
Point2Vec for Self-Supervised Representation Learning on Point Clouds
von: Knaebel, Karim, et al.
Veröffentlicht: (2023) -
Acquisition of high-quality images for camera calibration in robotics applications via speech prompts
von: Linder, Timm, et al.
Veröffentlicht: (2025) -
OpenSplat3D: Open-Vocabulary 3D Instance Segmentation using Gaussian Splatting
von: Piekenbrinck, Jens, et al.
Veröffentlicht: (2025)