Video sentence grounding with temporally global textual knowledge
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Cai, Zhang, Runzhong, Gao, Jianjun, Wu, Kejun, Yap, Kim-Hui, Wang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
by: Cai, Chen, et al.
Published: (2024)
by: Cai, Chen, et al.
Published: (2024)
PromptSR: Cascade Prompting for Lightweight Image Super-Resolution
by: Liu, Wenyang, et al.
Published: (2025)
by: Liu, Wenyang, et al.
Published: (2025)
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
by: Gao, Jianjun, et al.
Published: (2024)
by: Gao, Jianjun, et al.
Published: (2024)
Adaptive Margin Contrastive Learning for Ambiguity-aware 3D Semantic Segmentation
by: Chen, Yang, et al.
Published: (2025)
by: Chen, Yang, et al.
Published: (2025)
A Structure-aware and Motion-adaptive Framework for 3D Human Pose Estimation with Mamba
by: Lu, Ye, et al.
Published: (2025)
by: Lu, Ye, et al.
Published: (2025)
Multi-sentence Video Grounding for Long Video Generation
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising Timestep
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
Towards Blind Bitstream-corrupted Video Recovery via a Visual Foundation Model-driven Framework
by: Liu, Tianyi, et al.
Published: (2025)
by: Liu, Tianyi, et al.
Published: (2025)
Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges
by: Li, Junlong, et al.
Published: (2025)
by: Li, Junlong, et al.
Published: (2025)
From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition
by: Cai, Chen, et al.
Published: (2025)
by: Cai, Chen, et al.
Published: (2025)
A Video-grounded Dialogue Dataset and Metric for Event-driven Activities
by: Imrattanatrai, Wiradee, et al.
Published: (2025)
by: Imrattanatrai, Wiradee, et al.
Published: (2025)
MultiFuser: Multimodal Fusion Transformer for Enhanced Driver Action Recognition
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
HOI-aware Adaptive Network for Weakly-supervised Action Segmentation
by: Zhang, Runzhong, et al.
Published: (2026)
by: Zhang, Runzhong, et al.
Published: (2026)
OccluTrack: Rethinking Awareness of Occlusion for Enhancing Multiple Pedestrian Tracking
by: Gao, Jianjun, et al.
Published: (2023)
by: Gao, Jianjun, et al.
Published: (2023)
SSH-Net: A Self-Supervised and Hybrid Network for Noisy Image Watermark Removal
by: Liu, Wenyang, et al.
Published: (2025)
by: Liu, Wenyang, et al.
Published: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
LLM-grounded Video Diffusion Models
by: Lian, Long, et al.
Published: (2023)
by: Lian, Long, et al.
Published: (2023)
CM2-Net: Continual Cross-Modal Mapping Network for Driver Action Recognition
by: Wang, Ruoyu, et al.
Published: (2024)
by: Wang, Ruoyu, et al.
Published: (2024)
Top-Down Framework for Weakly-supervised Grounded Image Captioning
by: Cai, Chen, et al.
Published: (2023)
by: Cai, Chen, et al.
Published: (2023)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
by: Waheed, Abdul, et al.
Published: (2025)
by: Waheed, Abdul, et al.
Published: (2025)
VideoChat: Chat-Centric Video Understanding
by: Li, KunChang, et al.
Published: (2023)
by: Li, KunChang, et al.
Published: (2023)
ByteNet: Rethinking Multimedia File Fragment Classification through Visual Perspectives
by: Liu, Wenyang, et al.
Published: (2024)
by: Liu, Wenyang, et al.
Published: (2024)
Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs
by: Zhang, Xuan, et al.
Published: (2025)
by: Zhang, Xuan, et al.
Published: (2025)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
by: Mu, Kuan-Chen, et al.
Published: (2024)
by: Mu, Kuan-Chen, et al.
Published: (2024)
Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos
by: Hu, Kairui, et al.
Published: (2025)
by: Hu, Kairui, et al.
Published: (2025)
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models
by: Wang, Wenhao, et al.
Published: (2024)
by: Wang, Wenhao, et al.
Published: (2024)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
by: Yang, Zongxin, et al.
Published: (2024)
by: Yang, Zongxin, et al.
Published: (2024)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
by: Chou, Shih-Han, et al.
Published: (2023)
by: Chou, Shih-Han, et al.
Published: (2023)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
by: Lin, Jingyang, et al.
Published: (2025)
by: Lin, Jingyang, et al.
Published: (2025)
LLaVA-Video: Video Instruction Tuning With Synthetic Data
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
by: Dou, Zi-Yi, et al.
Published: (2024)
by: Dou, Zi-Yi, et al.
Published: (2024)
VideoVista-CulturalLingo: 360$^\circ$ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
by: Chen, Xinyu, et al.
Published: (2025)
by: Chen, Xinyu, et al.
Published: (2025)
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
by: Ma, Guoqing, et al.
Published: (2025)
by: Ma, Guoqing, et al.
Published: (2025)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
by: Huang, Haoyang, et al.
Published: (2025)
by: Huang, Haoyang, et al.
Published: (2025)
Frame-Voyager: Learning to Query Frames for Video Large Language Models
by: Yu, Sicheng, et al.
Published: (2024)
by: Yu, Sicheng, et al.
Published: (2024)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
ASPO: Adaptive Sentence-Level Preference Optimization for Fine-Grained Multimodal Reasoning
by: Wang, Yeyuan, et al.
Published: (2025)
by: Wang, Yeyuan, et al.
Published: (2025)
Enhancing Video-Language Representations with Structural Spatio-Temporal Alignment
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
Similar Items
-
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
by: Cai, Chen, et al.
Published: (2024) -
PromptSR: Cascade Prompting for Lightweight Image Super-Resolution
by: Liu, Wenyang, et al.
Published: (2025) -
CL-HOI: Cross-Level Human-Object Interaction Distillation from Vision Large Language Models
by: Gao, Jianjun, et al.
Published: (2024) -
Adaptive Margin Contrastive Learning for Ambiguity-aware 3D Semantic Segmentation
by: Chen, Yang, et al.
Published: (2025) -
A Structure-aware and Motion-adaptive Framework for 3D Human Pose Estimation with Mamba
by: Lu, Ye, et al.
Published: (2025)