Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Juan, Chen, Jiahao, Wang, Cheng, Yu, Zhiwang, Qi, Tangquan, Liu, Can, Wu, Di |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
von: Zhang, Juan, et al.
Veröffentlicht: (2024)
EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot
von: Fei, Hao, et al.
Veröffentlicht: (2024)
von: Fei, Hao, et al.
Veröffentlicht: (2024)
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
von: Zhang, Han, et al.
Veröffentlicht: (2025)
von: Zhang, Han, et al.
Veröffentlicht: (2025)
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025)
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)
MusFlow: Multimodal Music Generation via Conditional Flow Matching
von: Song, Jiahao, et al.
Veröffentlicht: (2025)
von: Song, Jiahao, et al.
Veröffentlicht: (2025)
Multimodal Semantic Communication for Generative Audio-Driven Video Conferencing
von: Tong, Haonan, et al.
Veröffentlicht: (2024)
von: Tong, Haonan, et al.
Veröffentlicht: (2024)
GAIA: Zero-shot Talking Avatar Generation
von: He, Tianyu, et al.
Veröffentlicht: (2023)
von: He, Tianyu, et al.
Veröffentlicht: (2023)
Multimodal Emotion Recognition with Large Language Models
von: Zhang, Hongrui, et al.
Veröffentlicht: (2026)
von: Zhang, Hongrui, et al.
Veröffentlicht: (2026)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
von: Gan, Qijun, et al.
Veröffentlicht: (2025)
Hyperbolic Multimodal Generative Representation Learning for Generalized Zero-Shot Multimodal Information Extraction
von: Zhou, Baohang, et al.
Veröffentlicht: (2026)
von: Zhou, Baohang, et al.
Veröffentlicht: (2026)
Multimodal LLM-based Query Paraphrasing for Video Search
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
von: Wu, Jiaxin, et al.
Veröffentlicht: (2024)
MM-InstructEval: Zero-Shot Evaluation of (Multimodal) Large Language Models on Multimodal Reasoning Tasks
von: Yang, Xiaocui, et al.
Veröffentlicht: (2024)
von: Yang, Xiaocui, et al.
Veröffentlicht: (2024)
SentiAvatar: Towards Expressive and Interactive Digital Humans
von: Jin, Chuhao, et al.
Veröffentlicht: (2026)
von: Jin, Chuhao, et al.
Veröffentlicht: (2026)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
von: Tong, Xinyi, et al.
Veröffentlicht: (2025)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
von: Li, Qingcao, et al.
Veröffentlicht: (2026)
Differential Mental Disorder Detection with Psychology-Inspired Multimodal Stimuli
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Zhou, Zhiyuan, et al.
Veröffentlicht: (2026)
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
von: Yang, Dingyi, et al.
Veröffentlicht: (2024)
MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production
von: Hu, Huanran, et al.
Veröffentlicht: (2026)
von: Hu, Huanran, et al.
Veröffentlicht: (2026)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
MSMF: Multi-Scale Multi-Modal Fusion for Enhanced Stock Market Prediction
von: Qin, Jiahao
Veröffentlicht: (2024)
von: Qin, Jiahao
Veröffentlicht: (2024)
When Top-ranked Recommendations Fail: Modeling Multi-Granular Negative Feedback for Explainable and Robust Video Recommendation
von: Chen, Siran, et al.
Veröffentlicht: (2025)
von: Chen, Siran, et al.
Veröffentlicht: (2025)
Cap2Sum: Learning to Summarize Videos by Generating Captions
von: Zhao, Cairong, et al.
Veröffentlicht: (2024)
von: Zhao, Cairong, et al.
Veröffentlicht: (2024)
VideoZeroBench: Probing the Limits of Video MLLMs with Spatio-Temporal Evidence Verification
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
von: Meng, Jiahao, et al.
Veröffentlicht: (2026)
MInD: Improving Multimodal Sentiment Analysis via Multimodal Information Disentanglement
von: Dai, Weichen, et al.
Veröffentlicht: (2024)
von: Dai, Weichen, et al.
Veröffentlicht: (2024)
SACRED: A Faithful Annotated Multimedia Multimodal Multilingual Dataset for Classifying Connectedness Types in Online Spirituality
von: Guan, Qinghao, et al.
Veröffentlicht: (2026)
von: Guan, Qinghao, et al.
Veröffentlicht: (2026)
Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
von: Lu, Jiacheng, et al.
Veröffentlicht: (2025)
VRAgent-R1: Boosting Video Recommendation with MLLM-based Agents via Reinforcement Learning
von: Chen, Siran, et al.
Veröffentlicht: (2025)
von: Chen, Siran, et al.
Veröffentlicht: (2025)
Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model
von: Chen, Gehui, et al.
Veröffentlicht: (2024)
von: Chen, Gehui, et al.
Veröffentlicht: (2024)
MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
von: Zhang, Lei, et al.
Veröffentlicht: (2025)
Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos
von: Gong, Haisong, et al.
Veröffentlicht: (2025)
von: Gong, Haisong, et al.
Veröffentlicht: (2025)
LoVA: Long-form Video-to-Audio Generation
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
von: Cheng, Xin, et al.
Veröffentlicht: (2024)
Smaller is Better: Generative Models Can Power Short Video Preloading
von: Liu, Liming, et al.
Veröffentlicht: (2026)
von: Liu, Liming, et al.
Veröffentlicht: (2026)
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning
von: Bai, Hayes, et al.
Veröffentlicht: (2026)
von: Bai, Hayes, et al.
Veröffentlicht: (2026)
Hybrid CNN-Mamba Enhancement Network for Robust Multimodal Sentiment Analysis
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
Multimodal Graph-Based Variational Mixture of Experts Network for Zero-Shot Multimodal Information Extraction
von: Zhou, Baohang, et al.
Veröffentlicht: (2025)
von: Zhou, Baohang, et al.
Veröffentlicht: (2025)
Music Grounding by Short Video
von: Xin, Zijie, et al.
Veröffentlicht: (2024)
von: Xin, Zijie, et al.
Veröffentlicht: (2024)
Kubrick: Multimodal Agent Collaborations for Synthetic Video Generation
von: He, Liu, et al.
Veröffentlicht: (2024)
von: He, Liu, et al.
Veröffentlicht: (2024)
LungCURE: Benchmarking Multimodal Real-World Clinical Reasoning for Precision Lung Cancer Diagnosis and Treatment
von: Hao, Fangyu, et al.
Veröffentlicht: (2026)
von: Hao, Fangyu, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
von: Zhang, Juan, et al.
Veröffentlicht: (2024) -
EmpathyEar: An Open-source Avatar Multimodal Empathetic Chatbot
von: Fei, Hao, et al.
Veröffentlicht: (2024) -
Towards Multimodal Empathetic Response Generation: A Rich Text-Speech-Vision Avatar-based Benchmark
von: Zhang, Han, et al.
Veröffentlicht: (2025) -
LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
von: Wei, Jingxuan, et al.
Veröffentlicht: (2025) -
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)