UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhu, Yijie, Zhang, Lingsen, Yu, Zitong, Shao, Rui, Tan, Tao, Nie, Liqiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
di: Shao, Rui, et al.
Pubblicazione: (2025)
di: Shao, Rui, et al.
Pubblicazione: (2025)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
di: Lyu, Yibo, et al.
Pubblicazione: (2025)
di: Lyu, Yibo, et al.
Pubblicazione: (2025)
CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation
di: Yuan, Kaishen, et al.
Pubblicazione: (2025)
di: Yuan, Kaishen, et al.
Pubblicazione: (2025)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
di: Li, Teng, et al.
Pubblicazione: (2025)
di: Li, Teng, et al.
Pubblicazione: (2025)
$Δ$VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation
di: Zhu, Yijie, et al.
Pubblicazione: (2026)
di: Zhu, Yijie, et al.
Pubblicazione: (2026)
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
di: Shen, Leyang, et al.
Pubblicazione: (2024)
di: Shen, Leyang, et al.
Pubblicazione: (2024)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
di: Qiu, Zongyang, et al.
Pubblicazione: (2025)
di: Qiu, Zongyang, et al.
Pubblicazione: (2025)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
di: Wen, Haokun, et al.
Pubblicazione: (2026)
di: Wen, Haokun, et al.
Pubblicazione: (2026)
EmoStory: Emotion-Aware Story Generation
di: Yang, Jingyuan, et al.
Pubblicazione: (2026)
di: Yang, Jingyuan, et al.
Pubblicazione: (2026)
Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
di: Zhang, Renshan, et al.
Pubblicazione: (2024)
di: Zhang, Renshan, et al.
Pubblicazione: (2024)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
di: Wu, Size, et al.
Pubblicazione: (2025)
di: Wu, Size, et al.
Pubblicazione: (2025)
UniMesh: Unifying 3D Mesh Understanding and Generation
di: Huang, Peng, et al.
Pubblicazione: (2026)
di: Huang, Peng, et al.
Pubblicazione: (2026)
UniX: Unifying Autoregression and Diffusion for Chest X-Ray Understanding and Generation
di: Zhang, Ruiheng, et al.
Pubblicazione: (2026)
di: Zhang, Ruiheng, et al.
Pubblicazione: (2026)
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
di: Wang, Duomin, et al.
Pubblicazione: (2025)
di: Wang, Duomin, et al.
Pubblicazione: (2025)
EmoLLM: Multimodal Emotional Understanding Meets Large Language Models
di: Yang, Qu, et al.
Pubblicazione: (2024)
di: Yang, Qu, et al.
Pubblicazione: (2024)
EmoAttack: Emotion-to-Image Diffusion Models for Emotional Backdoor Generation
di: Wei, Tianyu, et al.
Pubblicazione: (2024)
di: Wei, Tianyu, et al.
Pubblicazione: (2024)
UniVideo: Unified Understanding, Generation, and Editing for Videos
di: Wei, Cong, et al.
Pubblicazione: (2025)
di: Wei, Cong, et al.
Pubblicazione: (2025)
EmoCtrl: Controllable Emotional Image Content Generation
di: Yang, Jingyuan, et al.
Pubblicazione: (2025)
di: Yang, Jingyuan, et al.
Pubblicazione: (2025)
EmoArt: A Multidimensional Dataset for Emotion-Aware Artistic Generation
di: Zhang, Cheng, et al.
Pubblicazione: (2025)
di: Zhang, Cheng, et al.
Pubblicazione: (2025)
EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition
di: Schuhmann, Christoph, et al.
Pubblicazione: (2025)
di: Schuhmann, Christoph, et al.
Pubblicazione: (2025)
Nano-EmoX: Unifying Multimodal Emotional Intelligence from Perception to Empathy
di: Huang, Jiahao, et al.
Pubblicazione: (2026)
di: Huang, Jiahao, et al.
Pubblicazione: (2026)
UniPaint: Unified Space-time Video Inpainting via Mixture-of-Experts
di: Wan, Zhen, et al.
Pubblicazione: (2024)
di: Wan, Zhen, et al.
Pubblicazione: (2024)
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
di: Wang, Peiyu, et al.
Pubblicazione: (2025)
di: Wang, Peiyu, et al.
Pubblicazione: (2025)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
di: Liu, Jie, et al.
Pubblicazione: (2026)
di: Liu, Jie, et al.
Pubblicazione: (2026)
EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language
di: Chua, Phoebe, et al.
Pubblicazione: (2025)
di: Chua, Phoebe, et al.
Pubblicazione: (2025)
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
di: Tian, Rui, et al.
Pubblicazione: (2025)
di: Tian, Rui, et al.
Pubblicazione: (2025)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
di: Li, Yi, et al.
Pubblicazione: (2025)
di: Li, Yi, et al.
Pubblicazione: (2025)
Image Aesthetics Assessment via Learnable Queries
di: Xiong, Zhiwei, et al.
Pubblicazione: (2023)
di: Xiong, Zhiwei, et al.
Pubblicazione: (2023)
DeepFake-Adapter: Dual-Level Adapter for DeepFake Detection
di: Shao, Rui, et al.
Pubblicazione: (2023)
di: Shao, Rui, et al.
Pubblicazione: (2023)
UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
di: Xu, Yueming, et al.
Pubblicazione: (2025)
di: Xu, Yueming, et al.
Pubblicazione: (2025)
UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
di: An, Ruichuan, et al.
Pubblicazione: (2025)
di: An, Ruichuan, et al.
Pubblicazione: (2025)
UniVS: Unified and Universal Video Segmentation with Prompts as Queries
di: Li, Minghan, et al.
Pubblicazione: (2024)
di: Li, Minghan, et al.
Pubblicazione: (2024)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
di: Ma, Chuofan, et al.
Pubblicazione: (2025)
UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
di: Lu, Hao, et al.
Pubblicazione: (2025)
di: Lu, Hao, et al.
Pubblicazione: (2025)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
di: Zhang, Chi, et al.
Pubblicazione: (2025)
di: Zhang, Chi, et al.
Pubblicazione: (2025)
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
di: Zhang, Zhicheng, et al.
Pubblicazione: (2025)
di: Zhang, Zhicheng, et al.
Pubblicazione: (2025)
UniShield: Unified Face Attack Detection via KG-Informed Multimodal Reasoning
di: Li, Hongrui, et al.
Pubblicazione: (2026)
di: Li, Hongrui, et al.
Pubblicazione: (2026)
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
di: Xie, Wulin, et al.
Pubblicazione: (2025)
di: Xie, Wulin, et al.
Pubblicazione: (2025)
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
di: Yue, Zhengrong, et al.
Pubblicazione: (2025)
di: Yue, Zhengrong, et al.
Pubblicazione: (2025)
UniCompress: Token Compression for Unified Vision-Language Understanding and Generation
di: Wang, Ziyao, et al.
Pubblicazione: (2026)
di: Wang, Ziyao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
di: Shao, Rui, et al.
Pubblicazione: (2025) -
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
di: Lyu, Yibo, et al.
Pubblicazione: (2025) -
CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation
di: Yuan, Kaishen, et al.
Pubblicazione: (2025) -
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
di: Li, Teng, et al.
Pubblicazione: (2025) -
$Δ$VLA: Prior-Guided Vision-Language-Action Models via World Knowledge Variation
di: Zhu, Yijie, et al.
Pubblicazione: (2026)