Probing the 3D Awareness of Visual Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Banani, Mohamed El, Raj, Amit, Maninis, Kevis-Kokitsi, Kar, Abhishek, Li, Yuanzhen, Rubinstein, Michael, Sun, Deqing, Guibas, Leonidas, Johnson, Justin, Jampani, Varun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OmniNOCS: A unified NOCS dataset and model for 3D lifting of 2D objects
by: Krishnan, Akshay, et al.
Published: (2024)
by: Krishnan, Akshay, et al.
Published: (2024)
EgoCast: Forecasting Egocentric Human Pose in the Wild
by: Escobar, Maria, et al.
Published: (2024)
by: Escobar, Maria, et al.
Published: (2024)
3D Congealing: 3D-Aware Image Alignment in the Wild
by: Zhang, Yunzhi, et al.
Published: (2024)
by: Zhang, Yunzhi, et al.
Published: (2024)
SHINOBI: Shape and Illumination using Neural Object Decomposition via BRDF Optimization In-the-wild
by: Engelhardt, Andreas, et al.
Published: (2024)
by: Engelhardt, Andreas, et al.
Published: (2024)
ConDense: Consistent 2D/3D Pre-training for Dense and Sparse Features from Multi-View Images
by: Zhang, Xiaoshuai, et al.
Published: (2024)
by: Zhang, Xiaoshuai, et al.
Published: (2024)
WordRobe: Text-Guided Generation of Textured 3D Garments
by: Srivastava, Astitva, et al.
Published: (2024)
by: Srivastava, Astitva, et al.
Published: (2024)
DiffusionLight-Turbo: Accelerated Light Probes for Free via Single-Pass Chrome Ball Inpainting
by: Chinchuthakun, Worameth, et al.
Published: (2025)
by: Chinchuthakun, Worameth, et al.
Published: (2025)
CamCtrl3D: Single-Image Scene Exploration with Precise 3D Camera Control
by: Popov, Stefan, et al.
Published: (2025)
by: Popov, Stefan, et al.
Published: (2025)
DiffusionLight: Light Probes for Free by Painting a Chrome Ball
by: Phongthawee, Pakkapon, et al.
Published: (2023)
by: Phongthawee, Pakkapon, et al.
Published: (2023)
SMooDi: Stylized Motion Diffusion Model
by: Zhong, Lei, et al.
Published: (2024)
by: Zhong, Lei, et al.
Published: (2024)
OmniControl: Control Any Joint at Any Time for Human Motion Generation
by: Xie, Yiming, et al.
Published: (2023)
by: Xie, Yiming, et al.
Published: (2023)
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
by: Peng, Xiaogang, et al.
Published: (2023)
by: Peng, Xiaogang, et al.
Published: (2023)
ContactArt: Learning 3D Interaction Priors for Category-level Articulated Object and Hand Poses Estimation
by: Zhu, Zehao, et al.
Published: (2023)
by: Zhu, Zehao, et al.
Published: (2023)
Telling Left from Right: Identifying Geometry-Aware Semantic Correspondence
by: Zhang, Junyi, et al.
Published: (2023)
by: Zhang, Junyi, et al.
Published: (2023)
HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models
by: Ruiz, Nataniel, et al.
Published: (2023)
by: Ruiz, Nataniel, et al.
Published: (2023)
PASTA: Controllable Part-Aware Shape Generation with Autoregressive Transformers
by: Li, Songlin, et al.
Published: (2024)
by: Li, Songlin, et al.
Published: (2024)
ICE-G: Image Conditional Editing of 3D Gaussian Splats
by: Jaganathan, Vishnu, et al.
Published: (2024)
by: Jaganathan, Vishnu, et al.
Published: (2024)
Dress-Me-Up: A Dataset & Method for Self-Supervised 3D Garment Retargeting
by: Naik, Shanthika, et al.
Published: (2024)
by: Naik, Shanthika, et al.
Published: (2024)
HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation
by: Hu, Tao, et al.
Published: (2026)
by: Hu, Tao, et al.
Published: (2026)
BlenderAlchemy: Editing 3D Graphics with Vision-Language Models
by: Huang, Ian, et al.
Published: (2024)
by: Huang, Ian, et al.
Published: (2024)
TIPS: Text-Image Pretraining with Spatial awareness
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024)
by: Maninis, Kevis-Kokitsi, et al.
Published: (2024)
LightHeadEd: Relightable & Editable Head Avatars from a Smartphone
by: Manu, Pranav, et al.
Published: (2025)
by: Manu, Pranav, et al.
Published: (2025)
Dynamic Reflections: Probing Video Representations with Text Alignment
by: Zhu, Tyler, et al.
Published: (2025)
by: Zhu, Tyler, et al.
Published: (2025)
InfoGaussian: Structure-Aware Dynamic Gaussians through Lightweight Information Shaping
by: Zhang, Yunchao, et al.
Published: (2024)
by: Zhang, Yunchao, et al.
Published: (2024)
SPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single Images
by: Huang, Zixuan, et al.
Published: (2025)
by: Huang, Zixuan, et al.
Published: (2025)
ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs
by: Shah, Viraj, et al.
Published: (2023)
by: Shah, Viraj, et al.
Published: (2023)
BlenderGym: Benchmarking Foundational Model Systems for Graphics Editing
by: Gu, Yunqi, et al.
Published: (2025)
by: Gu, Yunqi, et al.
Published: (2025)
SuperDec: 3D Scene Decomposition with Superquadric Primitives
by: Fedele, Elisabetta, et al.
Published: (2025)
by: Fedele, Elisabetta, et al.
Published: (2025)
SF3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination Disentanglement
by: Boss, Mark, et al.
Published: (2024)
by: Boss, Mark, et al.
Published: (2024)
Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation Models
by: Michalkiewicz, Mateusz, et al.
Published: (2024)
by: Michalkiewicz, Mateusz, et al.
Published: (2024)
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2024)
by: Fan, Lijie, et al.
Published: (2024)
OCH3R: Object-Centric Holistic 3D Reconstruction
by: Du, Yi, et al.
Published: (2026)
by: Du, Yi, et al.
Published: (2026)
Refining Pre-Trained Motion Models
by: Sun, Xinglong, et al.
Published: (2024)
by: Sun, Xinglong, et al.
Published: (2024)
MoMaps: Semantics-Aware Scene Motion Generation with Motion Maps
by: Lei, Jiahui, et al.
Published: (2025)
by: Lei, Jiahui, et al.
Published: (2025)
MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation
by: Hu, Hanzhe, et al.
Published: (2024)
by: Hu, Hanzhe, et al.
Published: (2024)
Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction
by: Kilian, Maciej, et al.
Published: (2024)
by: Kilian, Maciej, et al.
Published: (2024)
GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
by: Michalkiewicz, Mateusz, et al.
Published: (2025)
by: Michalkiewicz, Mateusz, et al.
Published: (2025)
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
by: Zhang, Junyi, et al.
Published: (2024)
by: Zhang, Junyi, et al.
Published: (2024)
Unified Autoregressive Visual Generation and Understanding with Continuous Tokens
by: Fan, Lijie, et al.
Published: (2025)
by: Fan, Lijie, et al.
Published: (2025)
HouseCrafter: Lifting Floorplans to 3D Scenes with 2D Diffusion Model
by: Nguyen, Hieu T., et al.
Published: (2024)
by: Nguyen, Hieu T., et al.
Published: (2024)
Similar Items
-
OmniNOCS: A unified NOCS dataset and model for 3D lifting of 2D objects
by: Krishnan, Akshay, et al.
Published: (2024) -
EgoCast: Forecasting Egocentric Human Pose in the Wild
by: Escobar, Maria, et al.
Published: (2024) -
3D Congealing: 3D-Aware Image Alignment in the Wild
by: Zhang, Yunzhi, et al.
Published: (2024) -
SHINOBI: Shape and Illumination using Neural Object Decomposition via BRDF Optimization In-the-wild
by: Engelhardt, Andreas, et al.
Published: (2024) -
ConDense: Consistent 2D/3D Pre-training for Dense and Sparse Features from Multi-View Images
by: Zhang, Xiaoshuai, et al.
Published: (2024)