InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yi, He, Yinan, Li, Yizhuo, Li, Kunchang, Yu, Jiashuo, Ma, Xin, Li, Xinhao, Chen, Guo, Chen, Xinyuan, Wang, Yaohui, He, Conghui, Luo, Ping, Liu, Ziwei, Wang, Yali, Wang, Limin, Qiao, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoMamba: State Space Model for Efficient Video Understanding
by: Li, Kunchang, et al.
Published: (2024)
by: Li, Kunchang, et al.
Published: (2024)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
Harvest Video Foundation Models via Efficient Post-Pretraining
by: Li, Yizhuo, et al.
Published: (2023)
by: Li, Yizhuo, et al.
Published: (2023)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
by: Wang, Yi, et al.
Published: (2024)
by: Wang, Yi, et al.
Published: (2024)
Vlogger: Make Your Dream A Vlog
by: Zhuang, Shaobin, et al.
Published: (2024)
by: Zhuang, Shaobin, et al.
Published: (2024)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
by: Li, Kunchang, et al.
Published: (2023)
by: Li, Kunchang, et al.
Published: (2023)
VideoChat: Chat-Centric Video Understanding
by: Li, KunChang, et al.
Published: (2023)
by: Li, KunChang, et al.
Published: (2023)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
Task Preference Optimization: Improving Multimodal Large Language Models with Vision Task Alignment
by: Yan, Ziang, et al.
Published: (2024)
by: Yan, Ziang, et al.
Published: (2024)
Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
by: Chen, Boyu, et al.
Published: (2025)
by: Chen, Boyu, et al.
Published: (2025)
VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
by: Yan, Ziang, et al.
Published: (2025)
by: Yan, Ziang, et al.
Published: (2025)
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
by: Zeng, Xiangyu, et al.
Published: (2024)
by: Zeng, Xiangyu, et al.
Published: (2024)
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
by: Li, Xinhao, et al.
Published: (2025)
by: Li, Xinhao, et al.
Published: (2025)
TimeStep Master: Asymmetrical Mixture of Timestep LoRA Experts for Versatile and Efficient Diffusion Models in Vision
by: Zhuang, Shaobin, et al.
Published: (2025)
by: Zhuang, Shaobin, et al.
Published: (2025)
Percept, Chat, and then Adapt: Multimodal Knowledge Transfer of Foundation Models for Open-World Video Recognition
by: Chen, Boyu, et al.
Published: (2024)
by: Chen, Boyu, et al.
Published: (2024)
TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration
by: Guo, Yiwei, et al.
Published: (2024)
by: Guo, Yiwei, et al.
Published: (2024)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
by: Peng, Bo, et al.
Published: (2023)
by: Peng, Bo, et al.
Published: (2023)
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
by: Yu, Jiashuo, et al.
Published: (2025)
by: Yu, Jiashuo, et al.
Published: (2025)
AccVideo: Accelerating Video Diffusion Model with Synthetic Dataset
by: Zhang, Haiyu, et al.
Published: (2025)
by: Zhang, Haiyu, et al.
Published: (2025)
Bootstrapping Language-Guided Navigation Learning with Self-Refining Data Flywheel
by: Wang, Zun, et al.
Published: (2024)
by: Wang, Zun, et al.
Published: (2024)
ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
by: Wu, Xiaoxue, et al.
Published: (2025)
by: Wu, Xiaoxue, et al.
Published: (2025)
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
by: Deng, Nianchen, et al.
Published: (2025)
by: Deng, Nianchen, et al.
Published: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
by: Wang, Chenting, et al.
Published: (2025)
by: Wang, Chenting, et al.
Published: (2025)
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
by: Huang, Ziqi, et al.
Published: (2024)
by: Huang, Ziqi, et al.
Published: (2024)
VidLeaks: Membership Inference Attacks Against Text-to-Video Models
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
by: Ding, Yanbo, et al.
Published: (2024)
by: Ding, Yanbo, et al.
Published: (2024)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
by: Yang, Zhoufaran, et al.
Published: (2025)
by: Yang, Zhoufaran, et al.
Published: (2025)
Latte: Latent Diffusion Transformer for Video Generation
by: Ma, Xin, et al.
Published: (2024)
by: Ma, Xin, et al.
Published: (2024)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
CineTrans: Learning to Generate Videos with Cinematic Transitions via Masked Diffusion Models
by: Wu, Xiaoxue, et al.
Published: (2025)
by: Wu, Xiaoxue, et al.
Published: (2025)
Vid-SME: Membership Inference Attacks against Large Video Understanding Models
by: Li, Qi, et al.
Published: (2025)
by: Li, Qi, et al.
Published: (2025)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
by: Li, Kunchang, et al.
Published: (2022)
by: Li, Kunchang, et al.
Published: (2022)
GenHOI: Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects
by: Li, Shujia, et al.
Published: (2025)
by: Li, Shujia, et al.
Published: (2025)
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
4Diffusion: Multi-view Video Diffusion Model for 4D Generation
by: Zhang, Haiyu, et al.
Published: (2024)
by: Zhang, Haiyu, et al.
Published: (2024)
InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
by: Wang, Haomin, et al.
Published: (2025)
by: Wang, Haomin, et al.
Published: (2025)
Similar Items
-
VideoMamba: State Space Model for Efficient Video Understanding
by: Li, Kunchang, et al.
Published: (2024) -
Unmasked Teacher: Towards Training-Efficient Video Foundation Models
by: Li, Kunchang, et al.
Published: (2023) -
Harvest Video Foundation Models via Efficient Post-Pretraining
by: Li, Yizhuo, et al.
Published: (2023) -
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
by: Wang, Yi, et al.
Published: (2024) -
Vlogger: Make Your Dream A Vlog
by: Zhuang, Shaobin, et al.
Published: (2024)