Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
Fuente:
arXiv
Guardado en:
| Autores principales: | Fei, Hao, Wu, Shengqiong, Zhang, Meishan, Zhang, Min, Chua, Tat-Seng, Yan, Shuicheng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
por: Fei, Hao, et al.
Publicado: (2024)
por: Fei, Hao, et al.
Publicado: (2024)
Universal Scene Graph Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
Towards Semantic Equivalence of Tokenization in Multimodal LLM
por: Wu, Shengqiong, et al.
Publicado: (2024)
por: Wu, Shengqiong, et al.
Publicado: (2024)
Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
por: Wu, Shengqiong, et al.
Publicado: (2024)
por: Wu, Shengqiong, et al.
Publicado: (2024)
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
por: Fei, Hao, et al.
Publicado: (2023)
por: Fei, Hao, et al.
Publicado: (2023)
Modeling Cross-vision Synergy for Unified Large Vision Model
por: Wu, Shengqiong, et al.
Publicado: (2026)
por: Wu, Shengqiong, et al.
Publicado: (2026)
Global Commander and Local Operative: A Dual-Agent Framework for Scene Navigation
por: Jin, Kaiming, et al.
Publicado: (2026)
por: Jin, Kaiming, et al.
Publicado: (2026)
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
por: Cui, Chenhang, et al.
Publicado: (2024)
por: Cui, Chenhang, et al.
Publicado: (2024)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
XNLP: An Interactive Demonstration System for Universal Structured NLP
por: Fei, Hao, et al.
Publicado: (2023)
por: Fei, Hao, et al.
Publicado: (2023)
Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
por: Qian, Long, et al.
Publicado: (2024)
por: Qian, Long, et al.
Publicado: (2024)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
por: Liu, Kai, et al.
Publicado: (2025)
por: Liu, Kai, et al.
Publicado: (2025)
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
por: Qi, Ji, et al.
Publicado: (2025)
por: Qi, Ji, et al.
Publicado: (2025)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
por: Wu, Shengqiong, et al.
Publicado: (2026)
por: Wu, Shengqiong, et al.
Publicado: (2026)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
por: Fei, Hao, et al.
Publicado: (2024)
por: Fei, Hao, et al.
Publicado: (2024)
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
por: Liu, Kai, et al.
Publicado: (2025)
por: Liu, Kai, et al.
Publicado: (2025)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
por: Liu, Kai, et al.
Publicado: (2026)
por: Liu, Kai, et al.
Publicado: (2026)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
por: Wang, Yaoting, et al.
Publicado: (2025)
por: Wang, Yaoting, et al.
Publicado: (2025)
Learning 4D Panoptic Scene Graph Generation from Rich 2D Visual Scene
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
STARK: Spatio-Temporal Attention for Representation of Keypoints for Continuous Sign Language Recognition
por: Patra, Suvajit, et al.
Publicado: (2026)
por: Patra, Suvajit, et al.
Publicado: (2026)
Video-Language Alignment via Spatio-Temporal Graph Transformer
por: Zhang, Shi-Xue, et al.
Publicado: (2024)
por: Zhang, Shi-Xue, et al.
Publicado: (2024)
VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
por: Xu, Haidong, et al.
Publicado: (2025)
por: Xu, Haidong, et al.
Publicado: (2025)
STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training
por: Qiu, Haiyi, et al.
Publicado: (2024)
por: Qiu, Haiyi, et al.
Publicado: (2024)
RLHF-V: Towards Trustworthy MLLMs via Behavior Alignment from Fine-grained Correctional Human Feedback
por: Yu, Tianyu, et al.
Publicado: (2023)
por: Yu, Tianyu, et al.
Publicado: (2023)
A Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and Production
por: Hwang, Eui Jun, et al.
Publicado: (2024)
por: Hwang, Eui Jun, et al.
Publicado: (2024)
Thinking with Blueprints: Assisting Vision-Language Models in Spatial Reasoning via Structured Object Representation
por: Ma, Weijian, et al.
Publicado: (2026)
por: Ma, Weijian, et al.
Publicado: (2026)
Spatio-temporal Sign Language Representation and Translation
por: Hamidullah, Yasser, et al.
Publicado: (2025)
por: Hamidullah, Yasser, et al.
Publicado: (2025)
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
por: Jin, Peng, et al.
Publicado: (2024)
por: Jin, Peng, et al.
Publicado: (2024)
Grammar Induction from Visual, Speech and Text
por: Zhao, Yu, et al.
Publicado: (2024)
por: Zhao, Yu, et al.
Publicado: (2024)
Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding
por: Tu, Xuezhen, et al.
Publicado: (2026)
por: Tu, Xuezhen, et al.
Publicado: (2026)
Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector
por: Huang, Youcheng, et al.
Publicado: (2024)
por: Huang, Youcheng, et al.
Publicado: (2024)
OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
por: Zhang, Tao, et al.
Publicado: (2024)
por: Zhang, Tao, et al.
Publicado: (2024)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
por: Fang, Xiang, et al.
Publicado: (2026)
por: Fang, Xiang, et al.
Publicado: (2026)
Cognitive Visual-Language Mapper: Advancing Multimodal Comprehension with Enhanced Visual Knowledge Alignment
por: Li, Yunxin, et al.
Publicado: (2024)
por: Li, Yunxin, et al.
Publicado: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
por: Qu, Leigang, et al.
Publicado: (2024)
por: Qu, Leigang, et al.
Publicado: (2024)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
por: Qu, Leigang, et al.
Publicado: (2025)
por: Qu, Leigang, et al.
Publicado: (2025)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
por: Chu, Meng, et al.
Publicado: (2025)
por: Chu, Meng, et al.
Publicado: (2025)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
por: Zhou, Zhenglin, et al.
Publicado: (2025)
por: Zhou, Zhenglin, et al.
Publicado: (2025)
Extending Visual Dynamics for Video-to-Music Generation
por: Liu, Xiaohao, et al.
Publicado: (2025)
por: Liu, Xiaohao, et al.
Publicado: (2025)
Ejemplares similares
-
Vitron: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
por: Fei, Hao, et al.
Publicado: (2024) -
Universal Scene Graph Generation
por: Wu, Shengqiong, et al.
Publicado: (2025) -
Towards Semantic Equivalence of Tokenization in Multimodal LLM
por: Wu, Shengqiong, et al.
Publicado: (2024) -
Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
por: Wu, Shengqiong, et al.
Publicado: (2024) -
Dysen-VDM: Empowering Dynamics-aware Text-to-Video Diffusion with LLMs
por: Fei, Hao, et al.
Publicado: (2023)