A Mixed Diet Makes DINO An Omnivorous Vision Encoder
Fuente:
arXiv
Saved in:
| Main Authors: | Kabra, Rishabh, Ovsjanikov, Maks, Hudson, Drew A., Xia, Ye, Koppula, Skanda, Araujo, Andre, Carreira, Joao, Mitra, Niloy J. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How to Spin an Object: First, Get the Shape Right
by: Kabra, Rishabh, et al.
Published: (2024)
by: Kabra, Rishabh, et al.
Published: (2024)
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
by: Attaiki, Souhaib, et al.
Published: (2024)
by: Attaiki, Souhaib, et al.
Published: (2024)
Leveraging VLM-Based Pipelines to Annotate 3D Objects
by: Kabra, Rishabh, et al.
Published: (2023)
by: Kabra, Rishabh, et al.
Published: (2023)
Shape Non-rigid Kinematics (SNK): A Zero-Shot Method for Non-Rigid Shape Matching via Unsupervised Functional Map Regularized Reconstruction
by: Attaiki, Souhaib, et al.
Published: (2024)
by: Attaiki, Souhaib, et al.
Published: (2024)
FILTR: Extracting Topological Features from Pretrained 3D Models
by: Martinez, Louis, et al.
Published: (2026)
by: Martinez, Louis, et al.
Published: (2026)
Memory-Scalable and Simplified Functional Map Learning
by: Magnet, Robin, et al.
Published: (2024)
by: Magnet, Robin, et al.
Published: (2024)
Frozen Forecasting: A Unified Evaluation
by: Walker, Jacob C, et al.
Published: (2025)
by: Walker, Jacob C, et al.
Published: (2025)
Self-Supervised Dual Contouring
by: Sundararaman, Ramana, et al.
Published: (2024)
by: Sundararaman, Ramana, et al.
Published: (2024)
FourieRF: Few-Shot NeRFs via Progressive Fourier Frequency Control
by: Gomez, Diego, et al.
Published: (2025)
by: Gomez, Diego, et al.
Published: (2025)
To Supervise or Not to Supervise: Understanding and Addressing the Key Challenges of Point Cloud Transfer Learning
by: Hadgi, Souhail, et al.
Published: (2024)
by: Hadgi, Souhail, et al.
Published: (2024)
Back to 3D: Few-Shot 3D Keypoint Detection with Back-Projected 2D Features
by: Wimmer, Thomas, et al.
Published: (2023)
by: Wimmer, Thomas, et al.
Published: (2023)
TAPVid-3D: A Benchmark for Tracking Any Point in 3D
by: Koppula, Skanda, et al.
Published: (2024)
by: Koppula, Skanda, et al.
Published: (2024)
OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language Prompts
by: Xiao, Shiting, et al.
Published: (2025)
by: Xiao, Shiting, et al.
Published: (2025)
Graph Alignment via Dual-Pass Spectral Encoding and Latent Space Communication
by: Behmanesh, Maysam, et al.
Published: (2025)
by: Behmanesh, Maysam, et al.
Published: (2025)
DiffuMatch: Category-Agnostic Spectral Diffusion Priors for Robust Non-rigid Shape Matching
by: Pierson, Emery, et al.
Published: (2025)
by: Pierson, Emery, et al.
Published: (2025)
PoNQ: a Neural QEM-based Mesh Representation
by: Maruani, Nissim, et al.
Published: (2024)
by: Maruani, Nissim, et al.
Published: (2024)
Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
by: Chen, Victoria Yue, et al.
Published: (2026)
by: Chen, Victoria Yue, et al.
Published: (2026)
DeBaRA: Denoising-Based 3D Room Arrangement Generation
by: Maillard, Léopold, et al.
Published: (2024)
by: Maillard, Léopold, et al.
Published: (2024)
LACONIC: A 3D Layout Adapter for Controllable Image Creation
by: Maillard, Léopold, et al.
Published: (2025)
by: Maillard, Léopold, et al.
Published: (2025)
SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
by: Hasson, Yana, et al.
Published: (2025)
by: Hasson, Yana, et al.
Published: (2025)
From Blobs to Spokes: High-Fidelity Surface Reconstruction via Oriented Gaussians
by: Gomez, Diego, et al.
Published: (2026)
by: Gomez, Diego, et al.
Published: (2026)
PaNDaS: Learnable Deformation Modeling with Localized Control
by: Besnier, Thomas, et al.
Published: (2024)
by: Besnier, Thomas, et al.
Published: (2024)
Dynamic Reflections: Probing Video Representations with Text Alignment
by: Zhu, Tyler, et al.
Published: (2025)
by: Zhu, Tyler, et al.
Published: (2025)
A Simple Recipe for Contrastively Pre-training Video-First Encoders Beyond 16 Frames
by: Papalampidi, Pinelopi, et al.
Published: (2023)
by: Papalampidi, Pinelopi, et al.
Published: (2023)
Memory Consolidation Enables Long-Context Video Understanding
by: Balažević, Ivana, et al.
Published: (2024)
by: Balažević, Ivana, et al.
Published: (2024)
BootsTAP: Bootstrapped Training for Tracking-Any-Point
by: Doersch, Carl, et al.
Published: (2024)
by: Doersch, Carl, et al.
Published: (2024)
Recurrent Video Masked Autoencoders
by: Zoran, Daniel, et al.
Published: (2025)
by: Zoran, Daniel, et al.
Published: (2025)
Unique Lives, Shared World: Learning from Single-Life Videos
by: Han, Tengda, et al.
Published: (2025)
by: Han, Tengda, et al.
Published: (2025)
MILo: Mesh-In-the-Loop Gaussian Splatting for Detailed and Efficient Surface Reconstruction
by: Guédon, Antoine, et al.
Published: (2025)
by: Guédon, Antoine, et al.
Published: (2025)
PatchAlign3D: Local Feature Alignment for Dense 3D Shape understanding
by: Hadgi, Souhail, et al.
Published: (2026)
by: Hadgi, Souhail, et al.
Published: (2026)
ZeroKey: Point-Level Reasoning and Zero-Shot 3D Keypoint Detection from Large Language Models
by: Gong, Bingchen, et al.
Published: (2024)
by: Gong, Bingchen, et al.
Published: (2024)
SceneTeract: Agentic Functional Affordances and VLM Grounding in 3D Scenes
by: Maillard, Léopold, et al.
Published: (2026)
by: Maillard, Léopold, et al.
Published: (2026)
SAGE: Structure-Aware Generative Video Transitions between Diverse Clips
by: Kan, Mia, et al.
Published: (2025)
by: Kan, Mia, et al.
Published: (2025)
From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models
by: Jiang, Dongsheng, et al.
Published: (2023)
by: Jiang, Dongsheng, et al.
Published: (2023)
DINO-Foresight: Looking into the Future with DINO
by: Karypidis, Efstathios, et al.
Published: (2024)
by: Karypidis, Efstathios, et al.
Published: (2024)
LayerLock: Non-collapsing Representation Learning with Progressive Freezing
by: Erdogan, Goker, et al.
Published: (2025)
by: Erdogan, Goker, et al.
Published: (2025)
Escaping Plato's Cave: Towards the Alignment of 3D and Text Latent Spaces
by: Hadgi, Souhail, et al.
Published: (2025)
by: Hadgi, Souhail, et al.
Published: (2025)
DINO-Tok: Adapting DINO for Visual Tokenizers
by: Jia, Mingkai, et al.
Published: (2025)
by: Jia, Mingkai, et al.
Published: (2025)
Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Models
by: Wu, Ziyi, et al.
Published: (2024)
by: Wu, Ziyi, et al.
Published: (2024)
Animal Avatars: Reconstructing Animatable 3D Animals from Casual Videos
by: Sabathier, Remy, et al.
Published: (2024)
by: Sabathier, Remy, et al.
Published: (2024)
Similar Items
-
How to Spin an Object: First, Get the Shape Right
by: Kabra, Rishabh, et al.
Published: (2024) -
GANFusion: Feed-Forward Text-to-3D with Diffusion in GAN Space
by: Attaiki, Souhaib, et al.
Published: (2024) -
Leveraging VLM-Based Pipelines to Annotate 3D Objects
by: Kabra, Rishabh, et al.
Published: (2023) -
Shape Non-rigid Kinematics (SNK): A Zero-Shot Method for Non-Rigid Shape Matching via Unsupervised Functional Map Regularized Reconstruction
by: Attaiki, Souhaib, et al.
Published: (2024) -
FILTR: Extracting Topological Features from Pretrained 3D Models
by: Martinez, Louis, et al.
Published: (2026)