Video-Language Alignment via Spatio-Temporal Graph Transformer
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Shi-Xue, Wang, Hongfa, Zhu, Xiaobin, Gu, Weibo, Zhang, Tianjin, Yang, Chun, Liu, Wei, Yin, Xu-Cheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Inverse-like Antagonistic Scene Text Spotting via Reading-Order Estimation and Dynamic Sampling
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2024)
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2024)
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2025)
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2025)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
Towards Long-Form Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2026)
von: Gu, Xin, et al.
Veröffentlicht: (2026)
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024)
von: Gu, Xin, et al.
Veröffentlicht: (2024)
STGFormer: Spatio-Temporal GraphFormer for 3D Human Pose Estimation in Video
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
STAF: 3D Human Mesh Recovery from Video with Spatio-Temporal Alignment Fusion
von: Yao, Wei, et al.
Veröffentlicht: (2024)
von: Yao, Wei, et al.
Veröffentlicht: (2024)
Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding
von: Tu, Xuezhen, et al.
Veröffentlicht: (2026)
von: Tu, Xuezhen, et al.
Veröffentlicht: (2026)
Patch Spatio-Temporal Relation Prediction for Video Anomaly Detection
von: Shen, Hao, et al.
Veröffentlicht: (2024)
von: Shen, Hao, et al.
Veröffentlicht: (2024)
Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
von: Gao, Shida, et al.
Veröffentlicht: (2025)
von: Gao, Shida, et al.
Veröffentlicht: (2025)
Autoregressive Queries for Adaptive Tracking with Spatio-TemporalTransformers
von: Xie, Jinxia, et al.
Veröffentlicht: (2024)
von: Xie, Jinxia, et al.
Veröffentlicht: (2024)
Flash-GRPO: Efficient Alignment for Video Diffusion via One-Step Policy Optimization
von: He, Xiaoxuan, et al.
Veröffentlicht: (2026)
von: He, Xiaoxuan, et al.
Veröffentlicht: (2026)
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
von: Xu, Wenhao, et al.
Veröffentlicht: (2025)
von: Xu, Wenhao, et al.
Veröffentlicht: (2025)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
The Spatio-Temporal Poisson Point Process: A Simple Model for the Alignment of Event Camera Data
von: Gu, Cheng, et al.
Veröffentlicht: (2021)
von: Gu, Cheng, et al.
Veröffentlicht: (2021)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
CubeComposer: Spatio-Temporal Autoregressive 4K 360° Video Generation from Perspective Video
von: Li, Lingen, et al.
Veröffentlicht: (2026)
von: Li, Lingen, et al.
Veröffentlicht: (2026)
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
von: Guo, Yanan, et al.
Veröffentlicht: (2025)
Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry
von: Zhang, Zhaoxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhaoxing, et al.
Veröffentlicht: (2024)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
von: Cheng, Zixu, et al.
Veröffentlicht: (2025)
von: Cheng, Zixu, et al.
Veröffentlicht: (2025)
IC-Mapper: Instance-Centric Spatio-Temporal Modeling for Online Vectorized Map Construction
von: Zhu, Jiangtong, et al.
Veröffentlicht: (2025)
von: Zhu, Jiangtong, et al.
Veröffentlicht: (2025)
State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
von: Zhou, Jiahuan, et al.
Veröffentlicht: (2025)
von: Zhou, Jiahuan, et al.
Veröffentlicht: (2025)
VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
von: Fu, Honghao, et al.
Veröffentlicht: (2026)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Spatio-Temporal Distortion Aware Omnidirectional Video Super-Resolution
von: An, Hongyu, et al.
Veröffentlicht: (2024)
von: An, Hongyu, et al.
Veröffentlicht: (2024)
PiTe: Pixel-Temporal Alignment for Large Video-Language Model
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
Static and Dynamic Graph Alignment Network for Temporal Video Grounding
von: Hu, Zhanjie, et al.
Veröffentlicht: (2026)
von: Hu, Zhanjie, et al.
Veröffentlicht: (2026)
STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
von: Chen, Junyang, et al.
Veröffentlicht: (2025)
von: Chen, Junyang, et al.
Veröffentlicht: (2025)
PASTA: Towards Flexible and Efficient HDR Imaging Via Progressively Aggregated Spatio-Temporal Alignment
von: Liu, Xiaoning, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoning, et al.
Veröffentlicht: (2024)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
von: Lee, Jaewoo, et al.
Veröffentlicht: (2023)
von: Lee, Jaewoo, et al.
Veröffentlicht: (2023)
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
von: Gao, Ziqi, et al.
Veröffentlicht: (2026)
von: Gao, Ziqi, et al.
Veröffentlicht: (2026)
Distance-Aware Joint Spatio-Temporal Graph Contrastive Learning for Major Depressive Disorder Diagnosis
von: Hasan, Muhammad Asif, et al.
Veröffentlicht: (2026)
von: Hasan, Muhammad Asif, et al.
Veröffentlicht: (2026)
VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
von: Tang, Linfeng, et al.
Veröffentlicht: (2025)
von: Tang, Linfeng, et al.
Veröffentlicht: (2025)
Towards Spatio-Temporal World Scene Graph Generation from Monocular Videos
von: Peddi, Rohith, et al.
Veröffentlicht: (2026)
von: Peddi, Rohith, et al.
Veröffentlicht: (2026)
UHD-GPGNet: UHD Video Denoising via Gaussian-Process-Guided Local Spatio-Temporal Modeling
von: He, Weiyuan, et al.
Veröffentlicht: (2026)
von: He, Weiyuan, et al.
Veröffentlicht: (2026)
Rethinking the Spatio-Temporal Alignment of End-to-End 3D Perception
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyu, et al.
Veröffentlicht: (2025)
TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
von: Li, Shicheng, et al.
Veröffentlicht: (2025)
von: Li, Shicheng, et al.
Veröffentlicht: (2025)
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
von: Qiu, Tianheng, et al.
Veröffentlicht: (2025)
von: Qiu, Tianheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Inverse-like Antagonistic Scene Text Spotting via Reading-Order Estimation and Dynamic Sampling
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2024) -
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
von: Zhang, Shi-Xue, et al.
Veröffentlicht: (2025) -
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
von: Fei, Hao, et al.
Veröffentlicht: (2024) -
Towards Long-Form Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2026) -
Context-Guided Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2024)