SeCG: Semantic-Enhanced 3D Visual Grounding via Cross-modal Graph Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiao, Feng, Xu, Hongbin, Wu, Qiuxia, Kang, Wenxiong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
von: Xiao, Feng, et al.
Veröffentlicht: (2025)
von: Xiao, Feng, et al.
Veröffentlicht: (2025)
B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding
von: Xiao, Feng, et al.
Veröffentlicht: (2025)
von: Xiao, Feng, et al.
Veröffentlicht: (2025)
PointDC:Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel Clustering
von: Chen, Zisheng, et al.
Veröffentlicht: (2023)
von: Chen, Zisheng, et al.
Veröffentlicht: (2023)
4DStyleGaussian: Zero-shot 4D Style Transfer with Gaussian Splatting
von: Liang, Wanlin, et al.
Veröffentlicht: (2024)
von: Liang, Wanlin, et al.
Veröffentlicht: (2024)
StyleDyRF: Zero-shot 4D Style Transfer for Dynamic Neural Radiance Fields
von: Xu, Hongbin, et al.
Veröffentlicht: (2024)
von: Xu, Hongbin, et al.
Veröffentlicht: (2024)
Pb4U-GNet: Resolution-Adaptive Garment Simulation via Propagation-before-Update Graph Network
von: Liu, Aoran, et al.
Veröffentlicht: (2026)
von: Liu, Aoran, et al.
Veröffentlicht: (2026)
EA-3DGS: Efficient and Adaptive 3D Gaussians with Highly Enhanced Quality for outdoor scenes
von: Guo, Jianlin, et al.
Veröffentlicht: (2025)
von: Guo, Jianlin, et al.
Veröffentlicht: (2025)
Improving 3D Finger Traits Recognition via Generalizable Neural Rendering
von: Xu, Hongbin, et al.
Veröffentlicht: (2024)
von: Xu, Hongbin, et al.
Veröffentlicht: (2024)
ControLRM: Fast and Controllable 3D Generation via Large Reconstruction Model
von: Xu, Hongbin, et al.
Veröffentlicht: (2024)
von: Xu, Hongbin, et al.
Veröffentlicht: (2024)
Resolving Long-Tail Ambiguity in Unsupervised 3D Point Cloud Segmentation with Language Priors
von: Wei, Siqi, et al.
Veröffentlicht: (2026)
von: Wei, Siqi, et al.
Veröffentlicht: (2026)
Gesplat: Robust Pose-Free 3D Reconstruction via Geometry-Guided Gaussian Splatting
von: Lu, Jiahui, et al.
Veröffentlicht: (2025)
von: Lu, Jiahui, et al.
Veröffentlicht: (2025)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
von: Huang, Junming, et al.
Veröffentlicht: (2026)
von: Huang, Junming, et al.
Veröffentlicht: (2026)
Visual Object Tracking on Multi-modal RGB-D Videos: A Review
von: Zhu, Xue-Feng, et al.
Veröffentlicht: (2022)
von: Zhu, Xue-Feng, et al.
Veröffentlicht: (2022)
Enhanced Cross-modal 3D Retrieval via Tri-modal Reconstruction
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
RobustMVS: Single Domain Generalized Deep Multi-view Stereo
von: Xu, Hongbin, et al.
Veröffentlicht: (2024)
von: Xu, Hongbin, et al.
Veröffentlicht: (2024)
Visual Grounding with Attention-Driven Constraint Balancing
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
von: Kang, Weitai, et al.
Veröffentlicht: (2024)
SCA3D: Enhancing Cross-modal 3D Retrieval via 3D Shape and Caption Paired Data Augmentation
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
von: Ren, Junlong, et al.
Veröffentlicht: (2025)
Cross-Stage Attention Propagation for Efficient Semantic Segmentation
von: Kang, Beoungwoo
Veröffentlicht: (2026)
von: Kang, Beoungwoo
Veröffentlicht: (2026)
Semantics-aware Test-time Adaptation for 3D Human Pose Estimation
von: Lin, Qiuxia, et al.
Veröffentlicht: (2025)
von: Lin, Qiuxia, et al.
Veröffentlicht: (2025)
Bridging Modality Gap for Visual Grounding with Effecitve Cross-modal Distillation
von: Wang, Jiaxi, et al.
Veröffentlicht: (2023)
von: Wang, Jiaxi, et al.
Veröffentlicht: (2023)
GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion Transformer
von: Lin, Yihong, et al.
Veröffentlicht: (2024)
von: Lin, Yihong, et al.
Veröffentlicht: (2024)
Enhancing Visual Programming for Visual Reasoning via Probabilistic Graphs
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
von: Wan, Wentao, et al.
Veröffentlicht: (2025)
FusionBERT: Multi-View Image-3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder
von: Li, Wei, et al.
Veröffentlicht: (2026)
von: Li, Wei, et al.
Veröffentlicht: (2026)
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
von: Zhang, Jiaxi, et al.
Veröffentlicht: (2026)
von: Zhang, Jiaxi, et al.
Veröffentlicht: (2026)
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation
von: Kang, Dahyun, et al.
Veröffentlicht: (2024)
von: Kang, Dahyun, et al.
Veröffentlicht: (2024)
Fusion-then-Distillation: Toward Cross-modal Positive Distillation for Domain Adaptive 3D Semantic Segmentation
von: Wu, Yao, et al.
Veröffentlicht: (2024)
von: Wu, Yao, et al.
Veröffentlicht: (2024)
SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
von: Jin, Zhao, et al.
Veröffentlicht: (2025)
von: Jin, Zhao, et al.
Veröffentlicht: (2025)
Cross3DVG: Cross-Dataset 3D Visual Grounding on Different RGB-D Scans
von: Miyanishi, Taiki, et al.
Veröffentlicht: (2023)
von: Miyanishi, Taiki, et al.
Veröffentlicht: (2023)
CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
von: Zhang, Zelin, et al.
Veröffentlicht: (2026)
von: Zhang, Zelin, et al.
Veröffentlicht: (2026)
PhysMamba: State Space Duality Model for Remote Physiological Measurement
von: Yan, Zhixin, et al.
Veröffentlicht: (2024)
von: Yan, Zhixin, et al.
Veröffentlicht: (2024)
Cyc3D: Fine-grained Controllable 3D Generation via Cycle Consistency Regularization
von: Xu, Hongbin, et al.
Veröffentlicht: (2025)
von: Xu, Hongbin, et al.
Veröffentlicht: (2025)
View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
von: Liu, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Liu, Yuanyuan, et al.
Veröffentlicht: (2025)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
von: Zheng, Henry, et al.
Veröffentlicht: (2025)
von: Zheng, Henry, et al.
Veröffentlicht: (2025)
SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering
von: Li, Wenli, et al.
Veröffentlicht: (2026)
von: Li, Wenli, et al.
Veröffentlicht: (2026)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
von: Hu, Miao, et al.
Veröffentlicht: (2025)
von: Hu, Miao, et al.
Veröffentlicht: (2025)
MoniRefer: A Real-world Large-scale Multi-modal Dataset based on Roadside Infrastructure for 3D Visual Grounding
von: Yang, Panquan, et al.
Veröffentlicht: (2025)
von: Yang, Panquan, et al.
Veröffentlicht: (2025)
M3DHMR: Monocular 3D Hand Mesh Recovery
von: Lin, Yihong, et al.
Veröffentlicht: (2025)
von: Lin, Yihong, et al.
Veröffentlicht: (2025)
Data-Efficient 3D Visual Grounding via Order-Aware Referring
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
Multi-modality Affinity Inference for Weakly Supervised 3D Semantic Segmentation
von: Li, Xiawei, et al.
Veröffentlicht: (2023)
von: Li, Xiawei, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
von: Xiao, Feng, et al.
Veröffentlicht: (2025) -
B2N3D: Progressive Learning from Binary to N-ary Relationships for 3D Object Grounding
von: Xiao, Feng, et al.
Veröffentlicht: (2025) -
PointDC:Unsupervised Semantic Segmentation of 3D Point Clouds via Cross-modal Distillation and Super-Voxel Clustering
von: Chen, Zisheng, et al.
Veröffentlicht: (2023) -
4DStyleGaussian: Zero-shot 4D Style Transfer with Gaussian Splatting
von: Liang, Wanlin, et al.
Veröffentlicht: (2024) -
StyleDyRF: Zero-shot 4D Style Transfer for Dynamic Neural Radiance Fields
von: Xu, Hongbin, et al.
Veröffentlicht: (2024)