RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Wu, Changli, Chen, Qi, Ji, Jiayi, Wang, Haowei, Ma, Yiwei, Huang, You, Luo, Gen, Fei, Hao, Sun, Xiaoshuai, Ji, Rongrong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
3D-GRES: Generalized 3D Referring Expression Segmentation
di: Wu, Changli, et al.
Pubblicazione: (2024)
di: Wu, Changli, et al.
Pubblicazione: (2024)
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
di: Yang, Danni, et al.
Pubblicazione: (2024)
di: Yang, Danni, et al.
Pubblicazione: (2024)
IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
di: Chen, Qi, et al.
Pubblicazione: (2025)
di: Chen, Qi, et al.
Pubblicazione: (2025)
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
di: Lin, Weihuang, et al.
Pubblicazione: (2025)
di: Lin, Weihuang, et al.
Pubblicazione: (2025)
JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues
di: Ji, Jiayi, et al.
Pubblicazione: (2023)
di: Ji, Jiayi, et al.
Pubblicazione: (2023)
Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation
di: Liu, Sihan, et al.
Pubblicazione: (2023)
di: Liu, Sihan, et al.
Pubblicazione: (2023)
3D-DRES: Detailed 3D Referring Expression Segmentation
di: Chen, Qi, et al.
Pubblicazione: (2026)
di: Chen, Qi, et al.
Pubblicazione: (2026)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
di: Yang, Danni, et al.
Pubblicazione: (2024)
di: Yang, Danni, et al.
Pubblicazione: (2024)
Any-to-3D Generation via Hybrid Diffusion Supervision
di: Fan, Yijun, et al.
Pubblicazione: (2024)
di: Fan, Yijun, et al.
Pubblicazione: (2024)
HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
di: Lin, Weihuang, et al.
Pubblicazione: (2025)
di: Lin, Weihuang, et al.
Pubblicazione: (2025)
MMAPS: End-to-End Multi-Grained Multi-Modal Attribute-Aware Product Summarization
di: Chen, Tao, et al.
Pubblicazione: (2023)
di: Chen, Tao, et al.
Pubblicazione: (2023)
Omni-Referring Image Segmentation
di: Zheng, Qiancheng, et al.
Pubblicazione: (2025)
di: Zheng, Qiancheng, et al.
Pubblicazione: (2025)
X-Dreamer: Creating High-quality 3D Content by Bridging the Domain Gap Between Text-to-2D and Text-to-3D Generation
di: Ma, Yiwei, et al.
Pubblicazione: (2023)
di: Ma, Yiwei, et al.
Pubblicazione: (2023)
Multi-branch Collaborative Learning Network for 3D Visual Grounding
di: Qian, Zhipeng, et al.
Pubblicazione: (2024)
di: Qian, Zhipeng, et al.
Pubblicazione: (2024)
Beat: Bi-directional One-to-Many Embedding Alignment for Text-based Person Retrieval
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar Generation
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
di: Luo, Yaxin, et al.
Pubblicazione: (2024)
di: Luo, Yaxin, et al.
Pubblicazione: (2024)
Image Captioning via Dynamic Path Customization
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning
di: Ma, Yiwei, et al.
Pubblicazione: (2025)
di: Ma, Yiwei, et al.
Pubblicazione: (2025)
INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
di: Wu, Changli, et al.
Pubblicazione: (2026)
di: Wu, Changli, et al.
Pubblicazione: (2026)
Deep Instruction Tuning for Segment Anything Model
di: Huang, Xiaorui, et al.
Pubblicazione: (2024)
di: Huang, Xiaorui, et al.
Pubblicazione: (2024)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
di: Wu, Mingrui, et al.
Pubblicazione: (2024)
di: Wu, Mingrui, et al.
Pubblicazione: (2024)
PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation
di: Ke, Shuyan, et al.
Pubblicazione: (2026)
di: Ke, Shuyan, et al.
Pubblicazione: (2026)
Test-Time Computing for Referring Multimodal Large Language Models
di: Wu, Mingrui, et al.
Pubblicazione: (2026)
di: Wu, Mingrui, et al.
Pubblicazione: (2026)
GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
di: Liu, Lin, et al.
Pubblicazione: (2025)
di: Liu, Lin, et al.
Pubblicazione: (2025)
MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
di: Wu, Mingrui, et al.
Pubblicazione: (2026)
di: Wu, Mingrui, et al.
Pubblicazione: (2026)
M4-BLIP: Advancing Multi-Modal Media Manipulation Detection through Face-Enhanced Local Analysis
di: Wu, Hang, et al.
Pubblicazione: (2025)
di: Wu, Hang, et al.
Pubblicazione: (2025)
Referring Expression Instance Retrieval and A Strong End-to-End Baseline
di: Hao, Xiangzhao, et al.
Pubblicazione: (2025)
di: Hao, Xiangzhao, et al.
Pubblicazione: (2025)
Beyond First Impressions: Integrating Joint Multi-modal Cues for Comprehensive 3D Representation
di: Wang, Haowei, et al.
Pubblicazione: (2023)
di: Wang, Haowei, et al.
Pubblicazione: (2023)
Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models
di: Luo, Gen, et al.
Pubblicazione: (2024)
di: Luo, Gen, et al.
Pubblicazione: (2024)
I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
di: Ma, Yiwei, et al.
Pubblicazione: (2024)
HRSAM: Efficient Interactive Segmentation in High-Resolution Images
di: Huang, You, et al.
Pubblicazione: (2024)
di: Huang, You, et al.
Pubblicazione: (2024)
Computational Offloading in Semantic-Aware Cloud-Edge-End Collaborative Networks
di: Ji, Zelin, et al.
Pubblicazione: (2024)
di: Ji, Zelin, et al.
Pubblicazione: (2024)
Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
di: Zhang, Jinlu, et al.
Pubblicazione: (2024)
di: Zhang, Jinlu, et al.
Pubblicazione: (2024)
Mixed Degradation Image Restoration via Local Dynamic Optimization and Conditional Embedding
di: Gu, Yubin, et al.
Pubblicazione: (2024)
di: Gu, Yubin, et al.
Pubblicazione: (2024)
SAGE:State-Aware Guided End-to-End Policy for Multi-Stage Sequential Tasks via Hidden Markov Decision Process
di: Wu, BinXu, et al.
Pubblicazione: (2025)
di: Wu, BinXu, et al.
Pubblicazione: (2025)
Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
di: Ye, Kai, et al.
Pubblicazione: (2025)
di: Ye, Kai, et al.
Pubblicazione: (2025)
MIHBench: Benchmarking and Mitigating Multi-Image Hallucinations in Multimodal Large Language Models
di: Li, Jiale, et al.
Pubblicazione: (2025)
di: Li, Jiale, et al.
Pubblicazione: (2025)
RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning
di: Li, Kun, et al.
Pubblicazione: (2025)
di: Li, Kun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
3D-GRES: Generalized 3D Referring Expression Segmentation
di: Wu, Changli, et al.
Pubblicazione: (2024) -
SAM as the Guide: Mastering Pseudo-Label Refinement in Semi-Supervised Referring Expression Segmentation
di: Yang, Danni, et al.
Pubblicazione: (2024) -
IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation
di: Chen, Qi, et al.
Pubblicazione: (2025) -
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
di: Lin, Weihuang, et al.
Pubblicazione: (2025) -
JM3D & JM3D-LLM: Elevating 3D Understanding with Joint Multi-modal Cues
di: Ji, Jiayi, et al.
Pubblicazione: (2023)