VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Zhicheng, Wang, Weicheng, Zhu, Yongjie, Qin, Wenyu, Wan, Pengfei, Zhang, Di, Yang, Jufeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
by: Wang, Weicheng, et al.
Published: (2026)
by: Wang, Weicheng, et al.
Published: (2026)
MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
by: Zhang, Zhicheng, et al.
Published: (2025)
by: Zhang, Zhicheng, et al.
Published: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025)
by: Qiu, Zongyang, et al.
Published: (2025)
Analytic Score Optimization for Multi Dimension Video Quality Assessment
by: Lin, Boda, et al.
Published: (2026)
by: Lin, Boda, et al.
Published: (2026)
Embed-RL: Reinforcement Learning for Reasoning-Driven Multimodal Embeddings
by: Jiang, Haonan, et al.
Published: (2026)
by: Jiang, Haonan, et al.
Published: (2026)
EmoSpace: Fine-Grained Emotion Prototype Learning for Immersive Affective Content Generation
by: Wang, Bingyuan, et al.
Published: (2026)
by: Wang, Bingyuan, et al.
Published: (2026)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
EmoScene: A Dual-space Dataset for Controllable Affective Image Generation
by: He, Li, et al.
Published: (2026)
by: He, Li, et al.
Published: (2026)
EmoKGEdit: Training-free Affective Injection via Visual Cue Transformation
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
VidLeaks: Membership Inference Attacks Against Text-to-Video Models
by: Wang, Li, et al.
Published: (2026)
by: Wang, Li, et al.
Published: (2026)
KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding
by: Lin, Boda, et al.
Published: (2026)
by: Lin, Boda, et al.
Published: (2026)
PhysFlow: Unleashing the Potential of Multi-modal Foundation Models and Video Diffusion for 4D Dynamic Physical Scene Simulation
by: Liu, Zhuoman, et al.
Published: (2024)
by: Liu, Zhuoman, et al.
Published: (2024)
MicroEmo: Time-Sensitive Multimodal Emotion Recognition with Micro-Expression Dynamics in Video Dialogues
by: Zhang, Liyun
Published: (2024)
by: Zhang, Liyun
Published: (2024)
HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment
by: Liao, Zhichao, et al.
Published: (2025)
by: Liao, Zhichao, et al.
Published: (2025)
EmoStyle: Emotion-Driven Image Stylization
by: Yang, Jingyuan, et al.
Published: (2025)
by: Yang, Jingyuan, et al.
Published: (2025)
UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
by: Zhu, Yijie, et al.
Published: (2025)
by: Zhu, Yijie, et al.
Published: (2025)
TinyEmo: Scaling down Emotional Reasoning via Metric Projection
by: Gutierrez, Cristian
Published: (2024)
by: Gutierrez, Cristian
Published: (2024)
EmoStory: Emotion-Aware Story Generation
by: Yang, Jingyuan, et al.
Published: (2026)
by: Yang, Jingyuan, et al.
Published: (2026)
UniVid: The Open-Source Unified Video Model
by: Luo, Jiabin, et al.
Published: (2025)
by: Luo, Jiabin, et al.
Published: (2025)
EmoLat: Text-driven Image Sentiment Transfer via Emotion Latent Space
by: Zhang, Jing, et al.
Published: (2026)
by: Zhang, Jing, et al.
Published: (2026)
EmoAttack: Emotion-to-Image Diffusion Models for Emotional Backdoor Generation
by: Wei, Tianyu, et al.
Published: (2024)
by: Wei, Tianyu, et al.
Published: (2024)
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
by: Chen, Xinlong, et al.
Published: (2025)
by: Chen, Xinlong, et al.
Published: (2025)
EmoLLM: Multimodal Emotional Understanding Meets Large Language Models
by: Yang, Qu, et al.
Published: (2024)
by: Yang, Qu, et al.
Published: (2024)
EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion Models
by: Zhang, Yixuan, et al.
Published: (2025)
by: Zhang, Yixuan, et al.
Published: (2025)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
by: Xu, Yicheng, et al.
Published: (2025)
by: Xu, Yicheng, et al.
Published: (2025)
EmoCtrl: Controllable Emotional Image Content Generation
by: Yang, Jingyuan, et al.
Published: (2025)
by: Yang, Jingyuan, et al.
Published: (2025)
EmoSEM: Segment and Explain Emotion Stimuli in Visual Art
by: Zhang, Jing, et al.
Published: (2025)
by: Zhang, Jing, et al.
Published: (2025)
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
by: Li, Hui, et al.
Published: (2024)
by: Li, Hui, et al.
Published: (2024)
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
by: Qin, Bosheng, et al.
Published: (2023)
by: Qin, Bosheng, et al.
Published: (2023)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
EmoArt: A Multidimensional Dataset for Emotion-Aware Artistic Generation
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models
by: Yang, Jingyuan, et al.
Published: (2024)
by: Yang, Jingyuan, et al.
Published: (2024)
FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
EmoEdit: Evoking Emotions through Image Manipulation
by: Yang, Jingyuan, et al.
Published: (2024)
by: Yang, Jingyuan, et al.
Published: (2024)
CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation
by: Yuan, Kaishen, et al.
Published: (2025)
by: Yuan, Kaishen, et al.
Published: (2025)
IF-VidCap: Can Video Caption Models Follow Instructions?
by: Li, Shihao, et al.
Published: (2025)
by: Li, Shihao, et al.
Published: (2025)
GameFactory: Creating New Games with Generative Interactive Videos
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
by: Fang, Ye, et al.
Published: (2025)
by: Fang, Ye, et al.
Published: (2025)
Similar Items
-
Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning
by: Zhang, Zhicheng, et al.
Published: (2026) -
LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing
by: Wang, Weicheng, et al.
Published: (2026) -
MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
by: Zhang, Zhicheng, et al.
Published: (2025) -
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
by: Qiu, Zongyang, et al.
Published: (2025) -
Analytic Score Optimization for Multi Dimension Video Quality Assessment
by: Lin, Boda, et al.
Published: (2026)