Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Yufei, Zheng, Jie, Meng, Qianke, Yu, Zhou, Chen, Minghao, Ding, Jiajun, Tan, Min, Xi, Yuling, Chen, Zhiwen, Lv, Chengfei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
von: Yin, Yufei, et al.
Veröffentlicht: (2025)
Z3D: Zero-Shot 3D Visual Grounding from Images
von: Drozdov, Nikita, et al.
Veröffentlicht: (2026)
von: Drozdov, Nikita, et al.
Veröffentlicht: (2026)
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
von: Yin, Yufei, et al.
Veröffentlicht: (2026)
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
von: Xu, Runsen, et al.
Veröffentlicht: (2024)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
von: Li, Rong, et al.
Veröffentlicht: (2024)
von: Li, Rong, et al.
Veröffentlicht: (2024)
Lift3D: Zero-Shot Lifting of Any 2D Vision Model to 3D
von: T, Mukund Varma, et al.
Veröffentlicht: (2024)
von: T, Mukund Varma, et al.
Veröffentlicht: (2024)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
Solving Zero-Shot 3D Visual Grounding as Constraint Satisfaction Problems
von: Yuan, Qihao, et al.
Veröffentlicht: (2024)
von: Yuan, Qihao, et al.
Veröffentlicht: (2024)
Zero-Shot 3D Visual Grounding from Vision-Language Models
von: Li, Rong, et al.
Veröffentlicht: (2025)
von: Li, Rong, et al.
Veröffentlicht: (2025)
FHAvatar: Fast and High-Fidelity Reconstruction of Face-and-Hair Composable 3D Head Avatar from Few Casual Captures
von: Sun, Yujie, et al.
Veröffentlicht: (2026)
von: Sun, Yujie, et al.
Veröffentlicht: (2026)
TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian Splatting
von: Chen, Jianchuan, et al.
Veröffentlicht: (2025)
von: Chen, Jianchuan, et al.
Veröffentlicht: (2025)
ORIGEN: Zero-Shot 3D Orientation Grounding in Text-to-Image Generation
von: Min, Yunhong, et al.
Veröffentlicht: (2025)
von: Min, Yunhong, et al.
Veröffentlicht: (2025)
ConsDreamer: Advancing Multi-View Consistency for Zero-Shot Text-to-3D Generation
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
von: Zhou, Yuan, et al.
Veröffentlicht: (2025)
GeoLoco: Leveraging 3D Geometric Priors from Visual Foundation Model for Robust RGB-Only Humanoid Locomotion
von: Liu, Yufei, et al.
Veröffentlicht: (2026)
von: Liu, Yufei, et al.
Veröffentlicht: (2026)
D$^3$Fields: Dynamic 3D Descriptor Fields for Zero-Shot Generalizable Rearrangement
von: Wang, Yixuan, et al.
Veröffentlicht: (2023)
von: Wang, Yixuan, et al.
Veröffentlicht: (2023)
DGE: Direct Gaussian 3D Editing by Consistent Multi-view Editing
von: Chen, Minghao, et al.
Veröffentlicht: (2024)
von: Chen, Minghao, et al.
Veröffentlicht: (2024)
BeliefMapNav: 3D Voxel-Based Belief Map for Zero-Shot Object Navigation
von: Zhou, Zibo, et al.
Veröffentlicht: (2025)
von: Zhou, Zibo, et al.
Veröffentlicht: (2025)
SeqVLM: Proposal-Guided Multi-View Sequences Reasoning via VLM for Zero-Shot 3D Visual Grounding
von: Lin, Jiawen, et al.
Veröffentlicht: (2025)
von: Lin, Jiawen, et al.
Veröffentlicht: (2025)
AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models
von: Huynh, Cuong, et al.
Veröffentlicht: (2026)
von: Huynh, Cuong, et al.
Veröffentlicht: (2026)
SORT3D: Spatial Object-centric Reasoning Toolbox for Zero-Shot 3D Grounding Using Large Language Models
von: Zantout, Nader, et al.
Veröffentlicht: (2025)
von: Zantout, Nader, et al.
Veröffentlicht: (2025)
Zero-Shot UAV Navigation in Forests via Relightable 3D Gaussian Splatting
von: Lv, Zinan, et al.
Veröffentlicht: (2026)
von: Lv, Zinan, et al.
Veröffentlicht: (2026)
SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation
von: Xu, Mutian, et al.
Veröffentlicht: (2023)
von: Xu, Mutian, et al.
Veröffentlicht: (2023)
Global solutions of the 3D incompressible inhomogeneous viscoelastic system
von: Ai, Chengfei, et al.
Veröffentlicht: (2024)
von: Ai, Chengfei, et al.
Veröffentlicht: (2024)
Think, Act, Build: An Agentic Framework with Vision Language Models for Zero-Shot 3D Visual Grounding
von: Wang, Haibo, et al.
Veröffentlicht: (2026)
von: Wang, Haibo, et al.
Veröffentlicht: (2026)
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
von: Sun, Xuefei, et al.
Veröffentlicht: (2026)
von: Sun, Xuefei, et al.
Veröffentlicht: (2026)
GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
Agent3D-Zero: An Agent for Zero-shot 3D Understanding
von: Zhang, Sha, et al.
Veröffentlicht: (2024)
von: Zhang, Sha, et al.
Veröffentlicht: (2024)
Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding
von: Yuan, Zhihao, et al.
Veröffentlicht: (2023)
von: Yuan, Zhihao, et al.
Veröffentlicht: (2023)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
von: Feng, Chun, et al.
Veröffentlicht: (2024)
von: Feng, Chun, et al.
Veröffentlicht: (2024)
Reasoning3D -- Grounding and Reasoning in 3D: Fine-Grained Zero-Shot Open-Vocabulary 3D Reasoning Part Segmentation via Large Vision-Language Models
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
von: Chen, Tianrun, et al.
Veröffentlicht: (2024)
MCL-AD: Multimodal Collaboration Learning for Zero-Shot 3D Anomaly Detection
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
Sketch3D: Style-Consistent Guidance for Sketch-to-3D Generation
von: Zheng, Wangguandong, et al.
Veröffentlicht: (2024)
von: Zheng, Wangguandong, et al.
Veröffentlicht: (2024)
Reasoning Matters for 3D Visual Grounding
von: Huang, Hsiang-Wei, et al.
Veröffentlicht: (2026)
von: Huang, Hsiang-Wei, et al.
Veröffentlicht: (2026)
InterGSEdit: Interactive 3D Gaussian Splatting Editing with 3D Geometry-Consistent Attention Prior
von: Wen, Minghao, et al.
Veröffentlicht: (2025)
von: Wen, Minghao, et al.
Veröffentlicht: (2025)
Visual-Semantic Graph Matching Net for Zero-Shot Learning
von: Duan, Bowen, et al.
Veröffentlicht: (2024)
von: Duan, Bowen, et al.
Veröffentlicht: (2024)
Zero-P-to-3: Zero-Shot Partial-View Images to 3D Object
von: Lin, Yuxuan, et al.
Veröffentlicht: (2025)
von: Lin, Yuxuan, et al.
Veröffentlicht: (2025)
Unified Representation Space for 3D Visual Grounding
von: Zheng, Yinuo, et al.
Veröffentlicht: (2025)
von: Zheng, Yinuo, et al.
Veröffentlicht: (2025)
Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
Let 2D Diffusion Model Know 3D-Consistency for Robust Text-to-3D Generation
von: Seo, Junyoung, et al.
Veröffentlicht: (2023)
von: Seo, Junyoung, et al.
Veröffentlicht: (2023)
Learning an Efficient Multi-Turn Dialogue Evaluator from Multiple LLM Judges
von: Tang, Yuqi, et al.
Veröffentlicht: (2025)
von: Tang, Yuqi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2025) -
Z3D: Zero-Shot 3D Visual Grounding from Images
von: Drozdov, Nikita, et al.
Veröffentlicht: (2026) -
Progressive Video Condensation with MLLM Agent for Long-form Video Understanding
von: Yin, Yufei, et al.
Veröffentlicht: (2026) -
VLM-Grounder: A VLM Agent for Zero-Shot 3D Visual Grounding
von: Xu, Runsen, et al.
Veröffentlicht: (2024) -
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
von: Li, Rong, et al.
Veröffentlicht: (2024)