G2L: Semantically Aligned and Uniform Video Grounding via Geodesic and Game Theory
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongxiang, Cao, Meng, Cheng, Xuxin, Li, Yaowei, Zhu, Zhihong, Zou, Yuexian |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
by: Li, Hongxiang, et al.
Published: (2024)
by: Li, Hongxiang, et al.
Published: (2024)
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
by: Yang, Bang, et al.
Published: (2024)
by: Yang, Bang, et al.
Published: (2024)
VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
by: Zhuang, Xianwei, et al.
Published: (2025)
by: Zhuang, Xianwei, et al.
Published: (2025)
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
by: Li, Yaowei, et al.
Published: (2025)
by: Li, Yaowei, et al.
Published: (2025)
Uncertainty-aware sign language video retrieval with probability distribution modeling
by: Wu, Xuan, et al.
Published: (2024)
by: Wu, Xuan, et al.
Published: (2024)
ZeroNLG: Aligning and Autoencoding Domains for Zero-Shot Multimodal and Multilingual Natural Language Generation
by: Yang, Bang, et al.
Published: (2023)
by: Yang, Bang, et al.
Published: (2023)
Image Conductor: Precision Control for Interactive Video Synthesis
by: Li, Yaowei, et al.
Published: (2024)
by: Li, Yaowei, et al.
Published: (2024)
Towards Spoken Language Understanding via Multi-level Multi-grained Contrastive Learning
by: Cheng, Xuxin, et al.
Published: (2024)
by: Cheng, Xuxin, et al.
Published: (2024)
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
by: Tang, Lexiang, et al.
Published: (2025)
by: Tang, Lexiang, et al.
Published: (2025)
CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model
by: Yao, Ziyu, et al.
Published: (2025)
by: Yao, Ziyu, et al.
Published: (2025)
AlignGS: Aligning Geometry and Semantics for Robust Indoor Reconstruction from Sparse Views
by: Gao, Yijie, et al.
Published: (2025)
by: Gao, Yijie, et al.
Published: (2025)
Textual Inversion and Self-supervised Refinement for Radiology Report Generation
by: Luo, Yuanjiang, et al.
Published: (2024)
by: Luo, Yuanjiang, et al.
Published: (2024)
SphereVAD: Training-Free Video Anomaly Detection via Geodesic Inference on the Unit Hypersphere
by: Huang, Chao, et al.
Published: (2026)
by: Huang, Chao, et al.
Published: (2026)
VideoGuard: Protecting Video Content from Unauthorized Editing
by: Cao, Junjie, et al.
Published: (2025)
by: Cao, Junjie, et al.
Published: (2025)
IC-Custom: Diverse Image Customization via In-Context Learning
by: Li, Yaowei, et al.
Published: (2025)
by: Li, Yaowei, et al.
Published: (2025)
BrushEdit: All-In-One Image Inpainting and Editing
by: Li, Yaowei, et al.
Published: (2024)
by: Li, Yaowei, et al.
Published: (2024)
Towards Long Video Understanding via Fine-detailed Video Story Generation
by: You, Zeng, et al.
Published: (2024)
by: You, Zeng, et al.
Published: (2024)
Learning Spatial-Semantic Features for Robust Video Object Segmentation
by: Li, Xin, et al.
Published: (2024)
by: Li, Xin, et al.
Published: (2024)
AAformer: Auto-Aligned Transformer for Person Re-Identification
by: Zhu, Kuan, et al.
Published: (2021)
by: Zhu, Kuan, et al.
Published: (2021)
IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Cross-Modal Conditioned Reconstruction for Language-guided Medical Image Segmentation
by: Huang, Xiaoshuang, et al.
Published: (2024)
by: Huang, Xiaoshuang, et al.
Published: (2024)
ActPrompt: In-Domain Feature Adaptation via Action Cues for Video Temporal Grounding
by: Wang, Yubin, et al.
Published: (2024)
by: Wang, Yubin, et al.
Published: (2024)
Align3R: Aligned Monocular Depth Estimation for Dynamic Videos
by: Lu, Jiahao, et al.
Published: (2024)
by: Lu, Jiahao, et al.
Published: (2024)
G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
by: Lim, Jiyoung, et al.
Published: (2026)
by: Lim, Jiyoung, et al.
Published: (2026)
An End-to-End Framework for Video Multi-Person Pose Estimation
by: Wei, Zhihong
Published: (2025)
by: Wei, Zhihong
Published: (2025)
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
by: Zhang, Haoze, et al.
Published: (2025)
by: Zhang, Haoze, et al.
Published: (2025)
Clinically-Grounded Counterfactual Reasoning for Medical Video Diagnosis
by: Gao, Jianzhe, et al.
Published: (2026)
by: Gao, Jianzhe, et al.
Published: (2026)
4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generation
by: Yang, Shuzhou, et al.
Published: (2025)
by: Yang, Shuzhou, et al.
Published: (2025)
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
by: Tang, Jiyang, et al.
Published: (2025)
by: Tang, Jiyang, et al.
Published: (2025)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
by: Li, Zizun, et al.
Published: (2026)
by: Li, Zizun, et al.
Published: (2026)
Towards Visual Grounding: A Survey
by: Xiao, Linhui, et al.
Published: (2024)
by: Xiao, Linhui, et al.
Published: (2024)
SG-LDM: Semantic-Guided LiDAR Generation via Latent-Aligned Diffusion
by: Xiang, Zhengkang, et al.
Published: (2025)
by: Xiang, Zhengkang, et al.
Published: (2025)
ArrowGEV: Grounding Events in Video via Learning the Arrow of Time
by: Yu, Fangxu, et al.
Published: (2026)
by: Yu, Fangxu, et al.
Published: (2026)
FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
by: Yao, Ziyu, et al.
Published: (2024)
by: Yao, Ziyu, et al.
Published: (2024)
Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
by: Zhuang, Weijun, et al.
Published: (2026)
by: Zhuang, Weijun, et al.
Published: (2026)
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
by: Cao, Meng, et al.
Published: (2025)
by: Cao, Meng, et al.
Published: (2025)
VMBench: A Benchmark for Perception-Aligned Video Motion Generation
by: Ling, Xinran, et al.
Published: (2025)
by: Ling, Xinran, et al.
Published: (2025)
How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms
by: Jin, Shengji, et al.
Published: (2026)
by: Jin, Shengji, et al.
Published: (2026)
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
by: Liao, Zhaohe, et al.
Published: (2024)
by: Liao, Zhaohe, et al.
Published: (2024)
Similar Items
-
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023) -
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation
by: Li, Hongxiang, et al.
Published: (2024) -
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
by: Yang, Bang, et al.
Published: (2024) -
VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification
by: Zhuang, Xianwei, et al.
Published: (2025) -
BlobCtrl: Taming Controllable Blob for Element-level Image Editing
by: Li, Yaowei, et al.
Published: (2025)