Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yue, Yang, Kaizhi, Luo, Jiebo, Chen, Xuejin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Holistic Visual-Textual Sentiment Analysis with Prior Models
von: Chen, Junyu, et al.
Veröffentlicht: (2022)
von: Chen, Junyu, et al.
Veröffentlicht: (2022)
DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding
von: Liu, Ting, et al.
Veröffentlicht: (2024)
von: Liu, Ting, et al.
Veröffentlicht: (2024)
Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
von: Chen, Zhikai, et al.
Veröffentlicht: (2024)
von: Chen, Zhikai, et al.
Veröffentlicht: (2024)
Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
DanceCamAnimator: Keyframe-Based Controllable 3D Dance Camera Synthesis
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Language-Guided Diffusion Model for Visual Grounding
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
von: E, Shaojun, et al.
Veröffentlicht: (2025)
von: E, Shaojun, et al.
Veröffentlicht: (2025)
Rendering-Oriented 3D Point Cloud Attribute Compression using Sparse Tensor-based Transformer
von: Huo, Xiao, et al.
Veröffentlicht: (2024)
von: Huo, Xiao, et al.
Veröffentlicht: (2024)
Visual Grounding with Multi-modal Conditional Adaptation
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
von: Yao, Ruilin, et al.
Veröffentlicht: (2024)
Enhancing 3D Gaussian Splatting Compression via Spatial Condition-based Prediction
von: Ma, Jingui, et al.
Veröffentlicht: (2025)
von: Ma, Jingui, et al.
Veröffentlicht: (2025)
Dual Enhancement on 3D Vision-Language Perception for Monocular 3D Visual Grounding
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
von: Zhao, Sihan, et al.
Veröffentlicht: (2025)
von: Zhao, Sihan, et al.
Veröffentlicht: (2025)
RealX3D: A Physically-Degraded 3D Benchmark for Multi-view Visual Restoration and Reconstruction
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
von: Liu, Shuhong, et al.
Veröffentlicht: (2025)
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
Can Multimodal Large Language Models Understand Spatial Relations?
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
von: Liu, Jingping, et al.
Veröffentlicht: (2025)
WaterVG: Waterway Visual Grounding based on Text-Guided Vision and mmWave Radar
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
von: Guan, Runwei, et al.
Veröffentlicht: (2024)
Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
von: Jin, Hyundong, et al.
Veröffentlicht: (2025)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
Grounding is All You Need? Dual Temporal Grounding for Video Dialog
von: Qin, You, et al.
Veröffentlicht: (2024)
von: Qin, You, et al.
Veröffentlicht: (2024)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
PixCLIP: Achieving Fine-grained Visual Language Understanding via Any-granularity Pixel-Text Alignment Learning
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
Generating Attribute-Aware Human Motions from Textual Prompt
von: Wang, Xinghan, et al.
Veröffentlicht: (2025)
von: Wang, Xinghan, et al.
Veröffentlicht: (2025)
3D Gaussian Editing with A Single Image
von: Luo, Guan, et al.
Veröffentlicht: (2024)
von: Luo, Guan, et al.
Veröffentlicht: (2024)
SkyLink: Unifying Street-Satellite Geo-Localization via UAV-Mediated 3D Scene Alignment
von: Zhang, Hongyang, et al.
Veröffentlicht: (2025)
von: Zhang, Hongyang, et al.
Veröffentlicht: (2025)
Dual Mutual Learning Network with Global-local Awareness for RGB-D Salient Object Detection
von: Yi, Kang, et al.
Veröffentlicht: (2025)
von: Yi, Kang, et al.
Veröffentlicht: (2025)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
von: Zhang, Miaosen, et al.
Veröffentlicht: (2025)
CLIP Brings Better Features to Visual Aesthetics Learners
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
Noise-Tolerant Learning for Audio-Visual Action Recognition
von: Han, Haochen, et al.
Veröffentlicht: (2022)
von: Han, Haochen, et al.
Veröffentlicht: (2022)
Attributes-aware Visual Emotion Representation Learning
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
von: Chu, Meng, et al.
Veröffentlicht: (2023)
von: Chu, Meng, et al.
Veröffentlicht: (2023)
A Hierarchical Compression Technique for 3D Gaussian Splatting Compression
von: Huang, He, et al.
Veröffentlicht: (2024)
von: Huang, He, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Holistic Visual-Textual Sentiment Analysis with Prior Models
von: Chen, Junyu, et al.
Veröffentlicht: (2022) -
DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding
von: Liu, Ting, et al.
Veröffentlicht: (2024) -
Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
von: Chen, Zhikai, et al.
Veröffentlicht: (2024) -
Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding
von: Li, Yuzhen, et al.
Veröffentlicht: (2025) -
DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance
von: Wang, Zixuan, et al.
Veröffentlicht: (2024)