Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liao, Liwei, Li, Xufeng, Zheng, Xiaoyun, Liu, Boning, Gao, Feng, Wang, Ronggang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing 3D Gaussian Splatting Compression via Spatial Condition-based Prediction
von: Ma, Jingui, et al.
Veröffentlicht: (2025)
von: Ma, Jingui, et al.
Veröffentlicht: (2025)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes
von: Yan, Jinbo, et al.
Veröffentlicht: (2024)
von: Yan, Jinbo, et al.
Veröffentlicht: (2024)
Disparity-based Stereo Image Compression with Aligned Cross-View Priors
von: Zhai, Yongqi, et al.
Veröffentlicht: (2022)
von: Zhai, Yongqi, et al.
Veröffentlicht: (2022)
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
von: Cheung, Tsun-Hin, et al.
Veröffentlicht: (2024)
von: Cheung, Tsun-Hin, et al.
Veröffentlicht: (2024)
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
von: Duan, Yiqun, et al.
Veröffentlicht: (2025)
von: Duan, Yiqun, et al.
Veröffentlicht: (2025)
PKU-DyMVHumans: A Multi-View Video Benchmark for High-Fidelity Dynamic Human Modeling
von: Zheng, Xiaoyun, et al.
Veröffentlicht: (2024)
von: Zheng, Xiaoyun, et al.
Veröffentlicht: (2024)
Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding
von: Xu, Yue, et al.
Veröffentlicht: (2024)
von: Xu, Yue, et al.
Veröffentlicht: (2024)
Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
SSNVC: Single Stream Neural Video Compression with Implicit Temporal Information
von: Wang, Feng, et al.
Veröffentlicht: (2024)
von: Wang, Feng, et al.
Veröffentlicht: (2024)
Zero-Shot Character Identification and Speaker Prediction in Comics via Iterative Multimodal Fusion
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
von: Li, Yingxuan, et al.
Veröffentlicht: (2024)
Robust Mesh Saliency Ground Truth Acquisition in VR via View Cone Sampling and Manifold Diffusion
von: Zheng, Guoquan, et al.
Veröffentlicht: (2026)
von: Zheng, Guoquan, et al.
Veröffentlicht: (2026)
Adaptive 3D Gaussian Splatting Video Streaming
von: Gong, Han, et al.
Veröffentlicht: (2025)
von: Gong, Han, et al.
Veröffentlicht: (2025)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
Surface-SOS: Self-Supervised Object Segmentation via Neural Surface Representation
von: Zheng, Xiaoyun, et al.
Veröffentlicht: (2025)
von: Zheng, Xiaoyun, et al.
Veröffentlicht: (2025)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
von: Xu, Yifang, et al.
Veröffentlicht: (2025)
Language-Guided Diffusion Model for Visual Grounding
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
GSCodec Studio: A Modular Framework for Gaussian Splat Compression
von: Li, Sicheng, et al.
Veröffentlicht: (2025)
von: Li, Sicheng, et al.
Veröffentlicht: (2025)
GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
von: Yu, Hongyun, et al.
Veröffentlicht: (2024)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
L-LBVC: Long-Term Motion Estimation and Prediction for Learned Bi-Directional Video Compression
von: Zhai, Yongqi, et al.
Veröffentlicht: (2025)
von: Zhai, Yongqi, et al.
Veröffentlicht: (2025)
Boosting Temporal Sentence Grounding via Causal Inference
von: Tang, Kefan, et al.
Veröffentlicht: (2025)
von: Tang, Kefan, et al.
Veröffentlicht: (2025)
3D Gaussian Editing with A Single Image
von: Luo, Guan, et al.
Veröffentlicht: (2024)
von: Luo, Guan, et al.
Veröffentlicht: (2024)
Discrepancy-Aware Attention Network for Enhanced Audio-Visual Zero-Shot Learning
von: Yu, RunLin, et al.
Veröffentlicht: (2024)
von: Yu, RunLin, et al.
Veröffentlicht: (2024)
SizeGS: Size-aware Compression of 3D Gaussian Splatting via Mixed Integer Programming
von: Xie, Shuzhao, et al.
Veröffentlicht: (2024)
von: Xie, Shuzhao, et al.
Veröffentlicht: (2024)
Visual Grounding with Multi-modal Conditional Adaptation
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
Learning Brain Representation with Hierarchical Visual Embeddings
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
von: Zheng, Jiawen, et al.
Veröffentlicht: (2026)
FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
von: Cai, Lingling, et al.
Veröffentlicht: (2024)
3D-LMVIC: Learning-based Multi-View Image Coding with 3D Gaussian Geometric Priors
von: Huang, Yujun, et al.
Veröffentlicht: (2024)
von: Huang, Yujun, et al.
Veröffentlicht: (2024)
DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding
von: Liu, Ting, et al.
Veröffentlicht: (2024)
von: Liu, Ting, et al.
Veröffentlicht: (2024)
Hybrid Local-Global Context Learning for Neural Video Compression
von: Zhai, Yongqi, et al.
Veröffentlicht: (2024)
von: Zhai, Yongqi, et al.
Veröffentlicht: (2024)
GaussianForest: Hierarchical-Hybrid 3D Gaussian Splatting for Compressed Scene Modeling
von: Zhang, Fengyi, et al.
Veröffentlicht: (2024)
von: Zhang, Fengyi, et al.
Veröffentlicht: (2024)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
TimeNeRF: Building Generalizable Neural Radiance Fields across Time from Few-Shot Input Views
von: Hung, Hsiang-Hui, et al.
Veröffentlicht: (2025)
von: Hung, Hsiang-Hui, et al.
Veröffentlicht: (2025)
Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
GS-QA: Comprehensive Quality Assessment Benchmark for Gaussian Splatting View Synthesis
von: Martin, Pedro, et al.
Veröffentlicht: (2025)
von: Martin, Pedro, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
von: Li, Wenrui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Enhancing 3D Gaussian Splatting Compression via Spatial Condition-based Prediction
von: Ma, Jingui, et al.
Veröffentlicht: (2025) -
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026) -
4D Gaussian Splatting with Scale-aware Residual Field and Adaptive Optimization for Real-time Rendering of Temporally Complex Dynamic Scenes
von: Yan, Jinbo, et al.
Veröffentlicht: (2024) -
Disparity-based Stereo Image Compression with Aligned Cross-View Priors
von: Zhai, Yongqi, et al.
Veröffentlicht: (2022) -
Fine-grained Textual Inversion Network for Zero-Shot Composed Image Retrieval
von: Lin, Haoqiang, et al.
Veröffentlicht: (2025)