Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Rui, Fischer, Tobias, Segu, Mattia, Pollefeys, Marc, Van Gool, Luc, Tombari, Federico |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
by: Segu, Mattia, et al.
Published: (2025)
by: Segu, Mattia, et al.
Published: (2025)
Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
by: Mazzucco, Silvio, et al.
Published: (2025)
by: Mazzucco, Silvio, et al.
Published: (2025)
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations
by: Di Lorenzo, Gaia, et al.
Published: (2025)
by: Di Lorenzo, Gaia, et al.
Published: (2025)
Self-supervised Shape Completion via Involution and Implicit Correspondences
by: Liu, Mengya, et al.
Published: (2024)
by: Liu, Mengya, et al.
Published: (2024)
Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Appearance Graphs
by: Segu, Mattia, et al.
Published: (2024)
by: Segu, Mattia, et al.
Published: (2024)
Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking
by: Segu, Mattia, et al.
Published: (2024)
by: Segu, Mattia, et al.
Published: (2024)
Matching Anything by Segmenting Anything
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
SegSplat: Feed-forward Gaussian Splatting and Open-Set Semantic Segmentation
by: Siegel, Peter, et al.
Published: (2025)
by: Siegel, Peter, et al.
Published: (2025)
OpenNeRF: Open Set 3D Neural Scene Segmentation with Pixel-Wise Features and Rendered Novel Views
by: Engelmann, Francis, et al.
Published: (2024)
by: Engelmann, Francis, et al.
Published: (2024)
UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
UniK3D: Universal Camera Monocular 3D Estimation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
UniDepth: Universal Monocular Metric Depth Estimation
by: Piccinelli, Luigi, et al.
Published: (2024)
by: Piccinelli, Luigi, et al.
Published: (2024)
MICDrop: Masking Image and Depth Features via Complementary Dropout for Domain-Adaptive Semantic Segmentation
by: Yang, Linyan, et al.
Published: (2024)
by: Yang, Linyan, et al.
Published: (2024)
UniSDF: Unifying Neural Representations for High-Fidelity 3D Reconstruction of Complex Scenes with Reflections
by: Wang, Fangjinhua, et al.
Published: (2023)
by: Wang, Fangjinhua, et al.
Published: (2023)
One2Any: One-Reference 6D Pose Estimation for Any Object
by: Liu, Mengya, et al.
Published: (2025)
by: Liu, Mengya, et al.
Published: (2025)
SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
OVI-MAP:Open-Vocabulary Instance-Semantic Mapping
by: Deng, Zilong, et al.
Published: (2026)
by: Deng, Zilong, et al.
Published: (2026)
Re-Nerfing: Improving Novel View Synthesis through Novel View Synthesis
by: Tristram, Felix, et al.
Published: (2023)
by: Tristram, Felix, et al.
Published: (2023)
Are Multimodal Large Language Models Ready for Omnidirectional Spatial Reasoning?
by: Dongfang, Zihao, et al.
Published: (2025)
by: Dongfang, Zihao, et al.
Published: (2025)
SG-Tailor: Inter-Object Commonsense Relationship Reasoning for Scene Graph Manipulation
by: Shang, Haoliang, et al.
Published: (2025)
by: Shang, Haoliang, et al.
Published: (2025)
LaRI: Layered Ray Intersections for Single-view 3D Geometric Reasoning
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
3D-MOOD: Lifting 2D to 3D for Monocular Open-Set Object Detection
by: Yang, Yung-Hsu, et al.
Published: (2025)
by: Yang, Yung-Hsu, et al.
Published: (2025)
Splat-SLAM: Globally Optimized RGB-only SLAM with 3D Gaussians
by: Sandström, Erik, et al.
Published: (2024)
by: Sandström, Erik, et al.
Published: (2024)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
by: Shahbazi, Mohamad, et al.
Published: (2024)
by: Shahbazi, Mohamad, et al.
Published: (2024)
Déjà View: Looping Transformers for Multi-View 3D Reconstruction
by: Burzio, Alessandro, et al.
Published: (2026)
by: Burzio, Alessandro, et al.
Published: (2026)
Loc3R-VLM: Language-based Localization and 3D Reasoning with Vision-Language Models
by: Qu, Kevin, et al.
Published: (2026)
by: Qu, Kevin, et al.
Published: (2026)
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
by: Halacheva, Anna-Maria, et al.
Published: (2025)
by: Halacheva, Anna-Maria, et al.
Published: (2025)
Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding
by: Dey, Sombit, et al.
Published: (2024)
by: Dey, Sombit, et al.
Published: (2024)
P2P-Bridge: Diffusion Bridges for 3D Point Cloud Denoising
by: Vogel, Mathias, et al.
Published: (2024)
by: Vogel, Mathias, et al.
Published: (2024)
SceneGraphLoc: Cross-Modal Coarse Visual Localization on 3D Scene Graphs
by: Miao, Yang, et al.
Published: (2024)
by: Miao, Yang, et al.
Published: (2024)
Video Perception Models for 3D Scene Synthesis
by: Huang, Rui, et al.
Published: (2025)
by: Huang, Rui, et al.
Published: (2025)
EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting
by: Zhang, Daiwei, et al.
Published: (2024)
by: Zhang, Daiwei, et al.
Published: (2024)
TrafficBots V1.5: Traffic Simulation via Conditional VAEs and Transformers with Relative Pose Encoding
by: Zhang, Zhejun, et al.
Published: (2024)
by: Zhang, Zhejun, et al.
Published: (2024)
Vision encoders should be image size agnostic and task driven
by: Prisadnikov, Nedyalko, et al.
Published: (2025)
by: Prisadnikov, Nedyalko, et al.
Published: (2025)
GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
by: Fiaz, Mustansar, et al.
Published: (2025)
by: Fiaz, Mustansar, et al.
Published: (2025)
Spatial-Temporal Graph Mamba for Music-Guided Dance Video Synthesis
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding
by: Unal, Ozan, et al.
Published: (2023)
by: Unal, Ozan, et al.
Published: (2023)
Similar Items
-
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
by: Segu, Mattia, et al.
Published: (2025) -
Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
by: Mazzucco, Silvio, et al.
Published: (2025) -
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024) -
Object-X: Learning to Reconstruct Multi-Modal 3D Object Representations
by: Di Lorenzo, Gaia, et al.
Published: (2025) -
Self-supervised Shape Completion via Involution and Implicit Correspondences
by: Liu, Mengya, et al.
Published: (2024)