EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Qiu, Zongyang, Wang, Bingyuan, Chen, Xingbei, He, Yingqing, Wang, Zeyu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EmoSpace: Fine-Grained Emotion Prototype Learning for Immersive Affective Content Generation
by: Wang, Bingyuan, et al.
Published: (2026)
by: Wang, Bingyuan, et al.
Published: (2026)
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023)
by: Wang, Yi, et al.
Published: (2023)
EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language
by: Chua, Phoebe, et al.
Published: (2025)
by: Chua, Phoebe, et al.
Published: (2025)
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model
by: Zhang, Bingyuan, et al.
Published: (2024)
by: Zhang, Bingyuan, et al.
Published: (2024)
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
by: Li, Hui, et al.
Published: (2024)
by: Li, Hui, et al.
Published: (2024)
EmoLLM: Multimodal Emotional Understanding Meets Large Language Models
by: Yang, Qu, et al.
Published: (2024)
by: Yang, Qu, et al.
Published: (2024)
EmoTrans: A Benchmark for Understanding, Reasoning, and Predicting Emotion Transitions in Multimodal LLMs
by: Hu, He, et al.
Published: (2026)
by: Hu, He, et al.
Published: (2026)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
EmoArt: A Multidimensional Dataset for Emotion-Aware Artistic Generation
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues
by: Zhang, Liyun
Published: (2024)
by: Zhang, Liyun
Published: (2024)
EmoStory: Emotion-Aware Story Generation
by: Yang, Jingyuan, et al.
Published: (2026)
by: Yang, Jingyuan, et al.
Published: (2026)
UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
by: Zhu, Yijie, et al.
Published: (2025)
by: Zhu, Yijie, et al.
Published: (2025)
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
EmoAttack: Emotion-to-Image Diffusion Models for Emotional Backdoor Generation
by: Wei, Tianyu, et al.
Published: (2024)
by: Wei, Tianyu, et al.
Published: (2024)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
VideoChat: Chat-Centric Video Understanding
by: Li, KunChang, et al.
Published: (2023)
by: Li, KunChang, et al.
Published: (2023)
EmoCtrl: Controllable Emotional Image Content Generation
by: Yang, Jingyuan, et al.
Published: (2025)
by: Yang, Jingyuan, et al.
Published: (2025)
FindingEmo: An Image Dataset for Emotion Recognition in the Wild
by: Mertens, Laurent, et al.
Published: (2024)
by: Mertens, Laurent, et al.
Published: (2024)
EmoMeta: A Multimodal Dataset for Fine-grained Emotion Classification in Chinese Metaphors
by: Lu, Xingyuan, et al.
Published: (2025)
by: Lu, Xingyuan, et al.
Published: (2025)
Multimodal Video Emotion Recognition with Reliable Reasoning Priors
by: Wang, Zhepeng, et al.
Published: (2025)
by: Wang, Zhepeng, et al.
Published: (2025)
Controllable Video Generation: A Survey
by: Ma, Yue, et al.
Published: (2025)
by: Ma, Yue, et al.
Published: (2025)
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025)
by: Luo, Jiabin, et al.
Published: (2025)
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
by: Xu, Shuolin, et al.
Published: (2025)
by: Xu, Shuolin, et al.
Published: (2025)
VidText: Towards Comprehensive Evaluation for Video Text Understanding
by: Yang, Zhoufaran, et al.
Published: (2025)
by: Yang, Zhoufaran, et al.
Published: (2025)
EmoAssist: Emotional Assistant for Visual Impairment Community
by: Qi, Xingyu, et al.
Published: (2025)
by: Qi, Xingyu, et al.
Published: (2025)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
by: Li, Chaoyu, et al.
Published: (2024)
by: Li, Chaoyu, et al.
Published: (2024)
EmoSEM: Segment and Explain Emotion Stimuli in Visual Art
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
AnimationBench: Are Video Models Good at Character-Centric Animation?
by: Wu, Leyi, et al.
Published: (2026)
by: Wu, Leyi, et al.
Published: (2026)
FreeTraj: Tuning-Free Trajectory Control in Video Diffusion Models
by: Qiu, Haonan, et al.
Published: (2024)
by: Qiu, Haonan, et al.
Published: (2024)
VidEvent: A Large Dataset for Understanding Dynamic Evolution of Events in Videos
by: Liang, Baoyu, et al.
Published: (2025)
by: Liang, Baoyu, et al.
Published: (2025)
EmoStyle: Emotion-Driven Image Stylization
by: Yang, Jingyuan, et al.
Published: (2025)
by: Yang, Jingyuan, et al.
Published: (2025)
DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation
by: Di, Donglin, et al.
Published: (2024)
by: Di, Donglin, et al.
Published: (2024)
Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy
by: Huang, Jiahao, et al.
Published: (2026)
by: Huang, Jiahao, et al.
Published: (2026)
EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
by: Seoh, Ronald, et al.
Published: (2025)
by: Seoh, Ronald, et al.
Published: (2025)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
by: Ma, Yue, et al.
Published: (2023)
by: Ma, Yue, et al.
Published: (2023)
Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
by: Hu, Jinpeng, et al.
Published: (2025)
by: Hu, Jinpeng, et al.
Published: (2025)
Player-Centric Multimodal Prompt Generation for Large Language Model Based Identity-Aware Basketball Video Captioning
by: Xi, Zeyu, et al.
Published: (2025)
by: Xi, Zeyu, et al.
Published: (2025)
Similar Items
-
EmoSpace: Fine-Grained Emotion Prototype Learning for Immersive Affective Content Generation
by: Wang, Bingyuan, et al.
Published: (2026) -
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
by: Zhang, Zhicheng, et al.
Published: (2025) -
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
by: Wang, Yi, et al.
Published: (2023) -
EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language
by: Chua, Phoebe, et al.
Published: (2025) -
OmniVid: A Generative Framework for Universal Video Understanding
by: Wang, Junke, et al.
Published: (2024)