3D-GRES: Generalized 3D Referring Expression Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Changli, Liu, Yihang, Ji, Jiayi, Ma, Yiwei, Wang, Haowei, Luo, Gen, Ding, Henghui, Sun, Xiaoshuai, Ji, Rongrong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
von: Wu, Changli, et al.
Veröffentlicht: (2024)
von: Wu, Changli, et al.
Veröffentlicht: (2024)
JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues
von: Ji, Jiayi, et al.
Veröffentlicht: (2023)
von: Ji, Jiayi, et al.
Veröffentlicht: (2023)
IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
von: Chen, Qi, et al.
Veröffentlicht: (2025)
von: Chen, Qi, et al.
Veröffentlicht: (2025)
3D-DRES: Detailed 3D Referring Expression Segmentation
von: Chen, Qi, et al.
Veröffentlicht: (2026)
von: Chen, Qi, et al.
Veröffentlicht: (2026)
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
von: Yang, Danni, et al.
Veröffentlicht: (2024)
von: Yang, Danni, et al.
Veröffentlicht: (2024)
Any-to-3D Generation via Hybrid Diffusion Supervision
von: Fan, Yijun, et al.
Veröffentlicht: (2024)
von: Fan, Yijun, et al.
Veröffentlicht: (2024)
Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation
von: Liu, Sihan, et al.
Veröffentlicht: (2023)
von: Liu, Sihan, et al.
Veröffentlicht: (2023)
X-Dreamer: Creating High-quality 3D Content by Bridging the Domain Gap Between Text-to-2D and Text-to-3D Generation
von: Ma, Yiwei, et al.
Veröffentlicht: (2023)
von: Ma, Yiwei, et al.
Veröffentlicht: (2023)
X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar Generation
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
Multi-branch Collaborative Learning Network for 3D Visual Grounding
von: Qian, Zhipeng, et al.
Veröffentlicht: (2024)
von: Qian, Zhipeng, et al.
Veröffentlicht: (2024)
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
von: Luo, Yaxin, et al.
Veröffentlicht: (2024)
von: Luo, Yaxin, et al.
Veröffentlicht: (2024)
MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
von: Wu, Changli, et al.
Veröffentlicht: (2026)
von: Wu, Changli, et al.
Veröffentlicht: (2026)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
von: Yang, Danni, et al.
Veröffentlicht: (2024)
von: Yang, Danni, et al.
Veröffentlicht: (2024)
Omni-Referring Image Segmentation
von: Zheng, Qiancheng, et al.
Veröffentlicht: (2025)
von: Zheng, Qiancheng, et al.
Veröffentlicht: (2025)
RefMask3D: Language-Guided Transformer for 3D Referring Segmentation
von: He, Shuting, et al.
Veröffentlicht: (2024)
von: He, Shuting, et al.
Veröffentlicht: (2024)
Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation
von: Wang, Haowei, et al.
Veröffentlicht: (2023)
von: Wang, Haowei, et al.
Veröffentlicht: (2023)
HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
von: Lin, Weihuang, et al.
Veröffentlicht: (2025)
A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
von: He, Shuting, et al.
Veröffentlicht: (2025)
von: He, Shuting, et al.
Veröffentlicht: (2025)
Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
Image Captioning via Dynamic Path Customization
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
von: Zhang, Jinlu, et al.
Veröffentlicht: (2024)
von: Zhang, Jinlu, et al.
Veröffentlicht: (2024)
MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
von: Ma, Yiwei, et al.
Veröffentlicht: (2025)
von: Ma, Yiwei, et al.
Veröffentlicht: (2025)
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
von: Ding, Henghui, et al.
Veröffentlicht: (2026)
von: Ding, Henghui, et al.
Veröffentlicht: (2026)
Deep Instruction Tuning for Segment Anything Model
von: Huang, Xiaorui, et al.
Veröffentlicht: (2024)
von: Huang, Xiaorui, et al.
Veröffentlicht: (2024)
ReferSplat: Referring Segmentation in 3D Gaussian Splatting
von: He, Shuting, et al.
Veröffentlicht: (2025)
von: He, Shuting, et al.
Veröffentlicht: (2025)
M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
von: Wu, Hang, et al.
Veröffentlicht: (2025)
von: Wu, Hang, et al.
Veröffentlicht: (2025)
PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
von: Ke, Shuyan, et al.
Veröffentlicht: (2026)
von: Ke, Shuyan, et al.
Veröffentlicht: (2026)
INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
von: Ma, Yiwei, et al.
Veröffentlicht: (2024)
Test-Time Computing for Referring Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
von: Zhou, Ziyin, et al.
Veröffentlicht: (2025)
von: Zhou, Ziyin, et al.
Veröffentlicht: (2025)
Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
von: Ying, Kaining, et al.
Veröffentlicht: (2025)
Ref-SAM3D: Bridging SAM3D with Text for Reference 3D Reconstruction
von: Zhou, Yun, et al.
Veröffentlicht: (2025)
von: Zhou, Yun, et al.
Veröffentlicht: (2025)
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
von: Luo, Gen, et al.
Veröffentlicht: (2024)
von: Luo, Gen, et al.
Veröffentlicht: (2024)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
Evading Visual Aphasia: Contrastive Adaptive Semantic Token Pruning for Vision-Language Models
von: Ma, Jie, et al.
Veröffentlicht: (2026)
von: Ma, Jie, et al.
Veröffentlicht: (2026)
Evaluating and Analyzing Relationship Hallucinations in Large Vision-Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
Mixed Degradation Image Restoration via Local Dynamic Optimization and Conditional Embedding
von: Gu, Yubin, et al.
Veröffentlicht: (2024)
von: Gu, Yubin, et al.
Veröffentlicht: (2024)
Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation
von: He, Shuting, et al.
Veröffentlicht: (2024)
von: He, Shuting, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
von: Wu, Changli, et al.
Veröffentlicht: (2024) -
JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues
von: Ji, Jiayi, et al.
Veröffentlicht: (2023) -
IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
von: Chen, Qi, et al.
Veröffentlicht: (2025) -
3D-DRES: Detailed 3D Referring Expression Segmentation
von: Chen, Qi, et al.
Veröffentlicht: (2026) -
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
von: Yang, Danni, et al.
Veröffentlicht: (2024)