Emergent Extreme-View Geometry in 3D Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yiwen, Tung, Joseph, Cai, Ruojin, Fouhey, David, Averbuch-Elor, Hadar |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Extreme Rotation Estimation in the Wild
von: Bezalel, Hana, et al.
Veröffentlicht: (2024)
von: Bezalel, Hana, et al.
Veröffentlicht: (2024)
Emergent Visual-Semantic Hierarchies in Image-Text Representations
von: Alper, Morris, et al.
Veröffentlicht: (2024)
von: Alper, Morris, et al.
Veröffentlicht: (2024)
Long-tail Internet photo reconstruction
von: Li, Yuan, et al.
Veröffentlicht: (2026)
von: Li, Yuan, et al.
Veröffentlicht: (2026)
Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction
von: Phung, Hao, et al.
Veröffentlicht: (2026)
von: Phung, Hao, et al.
Veröffentlicht: (2026)
WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild
von: Alper, Morris, et al.
Veröffentlicht: (2025)
von: Alper, Morris, et al.
Veröffentlicht: (2025)
Kiki or Bouba? Sound Symbolism in Vision-and-Language Models
von: Alper, Morris, et al.
Veröffentlicht: (2023)
von: Alper, Morris, et al.
Veröffentlicht: (2023)
Supercharging Floorplan Localization with Semantic Rays
von: Grader, Yuval, et al.
Veröffentlicht: (2025)
von: Grader, Yuval, et al.
Veröffentlicht: (2025)
Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation
von: Fiebelman, Gal, et al.
Veröffentlicht: (2025)
von: Fiebelman, Gal, et al.
Veröffentlicht: (2025)
A Joint Study of Phrase Grounding and Task Performance in Vision and Language Models
von: Kojima, Noriyuki, et al.
Veröffentlicht: (2023)
von: Kojima, Noriyuki, et al.
Veröffentlicht: (2023)
Lang3D-XL: Language Embedded 3D Gaussians for Large-scale Scenes
von: Krakovsky, Shai, et al.
Veröffentlicht: (2025)
von: Krakovsky, Shai, et al.
Veröffentlicht: (2025)
InstanceGen: Image Generation with Instance-level Instructions
von: Sella, Etai, et al.
Veröffentlicht: (2025)
von: Sella, Etai, et al.
Veröffentlicht: (2025)
Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction Scenes
von: Peng, Wenxuan, et al.
Veröffentlicht: (2026)
von: Peng, Wenxuan, et al.
Veröffentlicht: (2026)
Spice-E : Structural Priors in 3D Diffusion using Cross-Entity Attention
von: Sella, Etai, et al.
Veröffentlicht: (2023)
von: Sella, Etai, et al.
Veröffentlicht: (2023)
Dynamic Scene Understanding from Vision-Language Representations
von: Pruss, Shahaf, et al.
Veröffentlicht: (2025)
von: Pruss, Shahaf, et al.
Veröffentlicht: (2025)
Color Bind: Exploring Color Perception in Text-to-Image Models
von: Shomer-Chai, Shay, et al.
Veröffentlicht: (2025)
von: Shomer-Chai, Shay, et al.
Veröffentlicht: (2025)
Thinking in Blender: Staged Executable Inverse Graphics with Vision-Language Models
von: He, Guangzhao, et al.
Veröffentlicht: (2026)
von: He, Guangzhao, et al.
Veröffentlicht: (2026)
WAFFLE: Multimodal Floorplan Understanding in the Wild
von: Ganon, Keren, et al.
Veröffentlicht: (2024)
von: Ganon, Keren, et al.
Veröffentlicht: (2024)
Blended Point Cloud Diffusion for Localized Text-guided Shape Editing
von: Sella, Etai, et al.
Veröffentlicht: (2025)
von: Sella, Etai, et al.
Veröffentlicht: (2025)
4-LEGS: 4D Language Embedded Gaussian Splatting
von: Fiebelman, Gal, et al.
Veröffentlicht: (2024)
von: Fiebelman, Gal, et al.
Veröffentlicht: (2024)
HaLo-NeRF: Learning Geometry-Guided Semantics for Exploring Unconstrained Photo Collections
von: Dudai, Chen, et al.
Veröffentlicht: (2024)
von: Dudai, Chen, et al.
Veröffentlicht: (2024)
ICC: Quantifying Image Caption Concreteness for Multimodal Dataset Curation
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
von: Yanuka, Moran, et al.
Veröffentlicht: (2024)
MeshOn: Intersection-Free Mesh-to-Mesh Composition
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2026)
von: Kim, Hyunwoo, et al.
Veröffentlicht: (2026)
Scene Grounding In the Wild
von: Cohen, Tamir, et al.
Veröffentlicht: (2026)
von: Cohen, Tamir, et al.
Veröffentlicht: (2026)
Systole-Conditioned Generative Cardiac Motion
von: Zuler, Shahar, et al.
Veröffentlicht: (2025)
von: Zuler, Shahar, et al.
Veröffentlicht: (2025)
Mitigating Open-Vocabulary Caption Hallucinations
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2023)
von: Ben-Kish, Assaf, et al.
Veröffentlicht: (2023)
MegaScenes: Scene-Level View Synthesis at Scale
von: Tung, Joseph, et al.
Veröffentlicht: (2024)
von: Tung, Joseph, et al.
Veröffentlicht: (2024)
ReNoise: Real Image Inversion Through Iterative Noising
von: Garibi, Daniel, et al.
Veröffentlicht: (2024)
von: Garibi, Daniel, et al.
Veröffentlicht: (2024)
ProtoSnap: Prototype Alignment for Cuneiform Signs
von: Mikulinsky, Rachel, et al.
Veröffentlicht: (2025)
von: Mikulinsky, Rachel, et al.
Veröffentlicht: (2025)
3DFIRES: Few Image 3D REconstruction for Scenes with Hidden Surface
von: Jin, Linyi, et al.
Veröffentlicht: (2024)
von: Jin, Linyi, et al.
Veröffentlicht: (2024)
ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
von: Chen, Hanyu, et al.
Veröffentlicht: (2026)
von: Chen, Hanyu, et al.
Veröffentlicht: (2026)
EPIC Fields: Marrying 3D Geometry and Video Understanding
von: Tschernezki, Vadim, et al.
Veröffentlicht: (2023)
von: Tschernezki, Vadim, et al.
Veröffentlicht: (2023)
Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features
von: Xiangli, Yuanbo, et al.
Veröffentlicht: (2024)
von: Xiangli, Yuanbo, et al.
Veröffentlicht: (2024)
Extreme Two-View Geometry From Object Poses with Diffusion Models
von: Sun, Yujing, et al.
Veröffentlicht: (2024)
von: Sun, Yujing, et al.
Veröffentlicht: (2024)
Emergent Outlier View Rejection in Visual Geometry Grounded Transformers
von: Han, Jisang, et al.
Veröffentlicht: (2025)
von: Han, Jisang, et al.
Veröffentlicht: (2025)
Dynamic Camera Poses and Where to Find Them
von: Rockwell, Chris, et al.
Veröffentlicht: (2025)
von: Rockwell, Chris, et al.
Veröffentlicht: (2025)
Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos
von: Jin, Linyi, et al.
Veröffentlicht: (2024)
von: Jin, Linyi, et al.
Veröffentlicht: (2024)
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
von: Qian, Shengyi, et al.
Veröffentlicht: (2024)
HyperCT: Low-Rank Hypernet for Unified Chest CT Analysis
von: Liu, Fengbei, et al.
Veröffentlicht: (2026)
von: Liu, Fengbei, et al.
Veröffentlicht: (2026)
Segmentation by Factorization: Unsupervised Semantic Segmentation for Pathology by Factorizing Foundation Model Features
von: Gildenblat, Jacob, et al.
Veröffentlicht: (2024)
von: Gildenblat, Jacob, et al.
Veröffentlicht: (2024)
Dens3R: A Foundation Model for 3D Geometry Prediction
von: Fang, Xianze, et al.
Veröffentlicht: (2025)
von: Fang, Xianze, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Extreme Rotation Estimation in the Wild
von: Bezalel, Hana, et al.
Veröffentlicht: (2024) -
Emergent Visual-Semantic Hierarchies in Image-Text Representations
von: Alper, Morris, et al.
Veröffentlicht: (2024) -
Long-tail Internet photo reconstruction
von: Li, Yuan, et al.
Veröffentlicht: (2026) -
Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction
von: Phung, Hao, et al.
Veröffentlicht: (2026) -
WildCAT3D: Appearance-Aware Multi-View Diffusion in the Wild
von: Alper, Morris, et al.
Veröffentlicht: (2025)