I Speak and You Find: Robust 3D Visual Grounding with Noisy and Ambiguous Speech Inputs
Fuente:
arXiv
Saved in:
| Main Authors: | Qi, Yu, Gu, Lipeng, Chen, Honghua, Nan, Liangliang, Wei, Mingqiang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified Representation Space for 3D Visual Grounding
by: Zheng, Yinuo, et al.
Published: (2025)
by: Zheng, Yinuo, et al.
Published: (2025)
CrossTracker: Robust Multi-modal 3D Multi-Object Tracking via Cross Correction
by: Gu, Lipeng, et al.
Published: (2024)
by: Gu, Lipeng, et al.
Published: (2024)
RegTrack: Simplicity Beneath Complexity in Robust Multi-Modal 3D Multi-Object Tracking
by: Gu, Lipeng, et al.
Published: (2024)
by: Gu, Lipeng, et al.
Published: (2024)
PointSFDA: Source-free Domain Adaptation for Point Cloud Completion
by: He, Xing, et al.
Published: (2025)
by: He, Xing, et al.
Published: (2025)
CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing
by: Zhu, Zhe, et al.
Published: (2025)
by: Zhu, Zhe, et al.
Published: (2025)
BridgeShape: Latent Diffusion Schrödinger Bridge for 3D Shape Completion
by: Kong, Dequan, et al.
Published: (2025)
by: Kong, Dequan, et al.
Published: (2025)
Hierarchical Error Assessment of CAD Models for Aircraft Manufacturing-and-Measurement
by: Huang, Jin, et al.
Published: (2025)
by: Huang, Jin, et al.
Published: (2025)
PointSea: Point Cloud Completion via Self-structure Augmentation
by: Zhu, Zhe, et al.
Published: (2025)
by: Zhu, Zhe, et al.
Published: (2025)
PointCG: Self-supervised Point Cloud Learning via Joint Completion and Generation
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
RARE: Refine Any Registration of Pairwise Point Clouds via Zero-Shot Learning
by: Zheng, Chengyu, et al.
Published: (2025)
by: Zheng, Chengyu, et al.
Published: (2025)
Improving the Robustness of 3D Human Pose Estimation: A Benchmark and Learning from Noisy Input
by: Hoang, Trung-Hieu, et al.
Published: (2023)
by: Hoang, Trung-Hieu, et al.
Published: (2023)
BSGS: Bi-stage 3D Gaussian Splatting for Camera Motion Deblurring
by: Zhao, An, et al.
Published: (2025)
by: Zhao, An, et al.
Published: (2025)
In-Field 3D Wheat Head Instance Segmentation From TLS Point Clouds Using Deep Learning Without Manual Labels
by: Medic, Tomislav, et al.
Published: (2026)
by: Medic, Tomislav, et al.
Published: (2026)
STAR-Edge: Structure-aware Local Spherical Curve Representation for Thin-walled Edge Extraction from Unstructured Point Clouds
by: Li, Zikuan, et al.
Published: (2025)
by: Li, Zikuan, et al.
Published: (2025)
You Only Speak Once to See
by: Yang, Wenhao, et al.
Published: (2024)
by: Yang, Wenhao, et al.
Published: (2024)
PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
by: Zhu, Zhe, et al.
Published: (2025)
by: Zhu, Zhe, et al.
Published: (2025)
GenPC: Zero-shot Point Cloud Completion via 3D Generative Priors
by: Li, An, et al.
Published: (2025)
by: Li, An, et al.
Published: (2025)
ClickSeg3D: Few-Click Interactive Segmentation via Semantic Embeddings
by: Kang, Xueyang, et al.
Published: (2026)
by: Kang, Xueyang, et al.
Published: (2026)
Multiple Consistent 2D-3D Mappings for Robust Zero-Shot 3D Visual Grounding
by: Yin, Yufei, et al.
Published: (2026)
by: Yin, Yufei, et al.
Published: (2026)
MonoRelief V2: Leveraging Real Data for High-Fidelity Monocular Relief Recovery
by: Zhang, Yu-Wei, et al.
Published: (2025)
by: Zhang, Yu-Wei, et al.
Published: (2025)
VAGUE: Visual Contexts Clarify Ambiguous Expressions
by: Nam, Heejeong, et al.
Published: (2024)
by: Nam, Heejeong, et al.
Published: (2024)
Quantized Visual Geometry Grounded Transformer
by: Feng, Weilun, et al.
Published: (2025)
by: Feng, Weilun, et al.
Published: (2025)
UniBEV: Multi-modal 3D Object Detection with Uniform BEV Encoders for Robustness against Missing Sensor Modalities
by: Wang, Shiming, et al.
Published: (2023)
by: Wang, Shiming, et al.
Published: (2023)
On the Estimation of Image-matching Uncertainty in Visual Place Recognition
by: Zaffar, Mubariz, et al.
Published: (2024)
by: Zaffar, Mubariz, et al.
Published: (2024)
ViSpeak: Visual Instruction Feedback in Streaming Videos
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
Laplace-Mamba: Laplace Frequency Prior-Guided Mamba-CNN Fusion Network for Image Dehazing
by: Wang, Yongzhen, et al.
Published: (2025)
by: Wang, Yongzhen, et al.
Published: (2025)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
Do You See What I Say? Generalizable Deepfake Detection based on Visual Speech Recognition
by: Bora, Maheswar, et al.
Published: (2025)
by: Bora, Maheswar, et al.
Published: (2025)
Data-Efficient 3D Visual Grounding via Order-Aware Referring
by: Wu, Tung-Yu, et al.
Published: (2024)
by: Wu, Tung-Yu, et al.
Published: (2024)
Point What You Mean: Visually Grounded Instruction Policy
by: Yu, Hang, et al.
Published: (2025)
by: Yu, Hang, et al.
Published: (2025)
Holistic Reliability Propagation: Decoupling Annotation and Prediction for Robust Noisy-Label
by: Mao, Jingyang, et al.
Published: (2026)
by: Mao, Jingyang, et al.
Published: (2026)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
AsyncBEV: Cross-modal Flow Alignment in Asynchronous 3D Object Detection
by: Wang, Shiming, et al.
Published: (2026)
by: Wang, Shiming, et al.
Published: (2026)
MVIP-NeRF: Multi-view 3D Inpainting on NeRF Scenes via Diffusion Prior
by: Chen, Honghua, et al.
Published: (2024)
by: Chen, Honghua, et al.
Published: (2024)
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
by: Haliassos, Alexandros, et al.
Published: (2024)
by: Haliassos, Alexandros, et al.
Published: (2024)
Dynamic Visual-semantic Alignment for Zero-shot Learning with Ambiguous Labels
by: Li, Jiangnan, et al.
Published: (2026)
by: Li, Jiangnan, et al.
Published: (2026)
CT-Bound: Robust Boundary Detection From Noisy Images Via Hybrid Convolution and Transformer Neural Networks
by: Xu, Wei, et al.
Published: (2024)
by: Xu, Wei, et al.
Published: (2024)
The Overlooked Value of Test-time Reference Sets in Visual Place Recognition
by: Zaffar, Mubariz, et al.
Published: (2025)
by: Zaffar, Mubariz, et al.
Published: (2025)
Visual Position Prompt for MLLM based Visual Grounding
by: Tang, Wei, et al.
Published: (2025)
by: Tang, Wei, et al.
Published: (2025)
Textured 3D Regenerative Morphing with 3D Diffusion Prior
by: Yang, Songlin, et al.
Published: (2025)
by: Yang, Songlin, et al.
Published: (2025)
Similar Items
-
Unified Representation Space for 3D Visual Grounding
by: Zheng, Yinuo, et al.
Published: (2025) -
CrossTracker: Robust Multi-modal 3D Multi-Object Tracking via Cross Correction
by: Gu, Lipeng, et al.
Published: (2024) -
RegTrack: Simplicity Beneath Complexity in Robust Multi-Modal 3D Multi-Object Tracking
by: Gu, Lipeng, et al.
Published: (2024) -
PointSFDA: Source-free Domain Adaptation for Point Cloud Completion
by: He, Xing, et al.
Published: (2025) -
CoreEditor: Correspondence-constrained Diffusion for Consistent 3D Editing
by: Zhu, Zhe, et al.
Published: (2025)