HunyuanVideo: A Systematic Framework For Large Video Generative Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Weijie, Tian, Qi, Zhang, Zijian, Min, Rox, Dai, Zuozhuo, Zhou, Jin, Xiong, Jiangfeng, Li, Xin, Wu, Bo, Zhang, Jianwei, Wu, Kathrina, Lin, Qin, Yuan, Junkun, Long, Yanxin, Wang, Aladdin, Wang, Andong, Li, Changlin, Huang, Duojun, Yang, Fang, Tan, Hao, Wang, Hongmei, Song, Jacob, Bai, Jiawang, Wu, Jianbing, Xue, Jinbao, Wang, Joey, Wang, Kai, Liu, Mengyang, Li, Pengyu, Li, Shuai, Wang, Weiyan, Yu, Wenqing, Deng, Xinchi, Li, Yang, Chen, Yi, Cui, Yutao, Peng, Yuanbo, Yu, Zhentao, He, Zhiyu, Xu, Zhiyong, Zhou, Zixiang, Xu, Zunnan, Tao, Yangyu, Lu, Qinglin, Liu, Songtao, Zhou, Dax, Wang, Hongfa, Yang, Yong, Wang, Di, Liu, Yuhong, Jiang, Jie, Zhong, Caesar |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HunyuanVideo 1.5 Technical Report
by: Wu, Bing, et al.
Published: (2025)
by: Wu, Bing, et al.
Published: (2025)
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
by: Shan, Sizhe, et al.
Published: (2025)
by: Shan, Sizhe, et al.
Published: (2025)
HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
by: Huang, Ziyao, et al.
Published: (2025)
by: Huang, Ziyao, et al.
Published: (2025)
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
by: Chen, Yi, et al.
Published: (2025)
by: Chen, Yi, et al.
Published: (2025)
Hunyuan-Game: Industrial-grade Intelligent Game Creation Model
by: Li, Ruihuang, et al.
Published: (2025)
by: Li, Ruihuang, et al.
Published: (2025)
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
by: Ge, Yuying, et al.
Published: (2025)
by: Ge, Yuying, et al.
Published: (2025)
HunyuanImage 3.0 Technical Report
by: Cao, Siyu, et al.
Published: (2025)
by: Cao, Siyu, et al.
Published: (2025)
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
by: Zhang, Shi-Xue, et al.
Published: (2025)
by: Zhang, Shi-Xue, et al.
Published: (2025)
Perception-Oriented Video Frame Interpolation via Asymmetric Blending
by: Wu, Guangyang, et al.
Published: (2024)
by: Wu, Guangyang, et al.
Published: (2024)
Rethinking Metrics and Benchmarks of Video Anomaly Detection
by: Liu, Zihao, et al.
Published: (2025)
by: Liu, Zihao, et al.
Published: (2025)
HunyuanOCR Technical Report
by: Hunyuan Vision Team, et al.
Published: (2025)
by: Hunyuan Vision Team, et al.
Published: (2025)
PM-VIS+: High-Performance Video Instance Segmentation without Video Annotation
by: Yang, Zhangjing, et al.
Published: (2024)
by: Yang, Zhangjing, et al.
Published: (2024)
Hunyuan-MT Technical Report
by: Zheng, Mao, et al.
Published: (2025)
by: Zheng, Mao, et al.
Published: (2025)
Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
by: Li, Zhimin, et al.
Published: (2024)
by: Li, Zhimin, et al.
Published: (2024)
TSdetector: Temporal-Spatial Self-correction Collaborative Learning for Colonoscopy Video Detection
by: Wang, Kaini, et al.
Published: (2024)
by: Wang, Kaini, et al.
Published: (2024)
Follow-Your-Canvas: Higher-Resolution Video Outpainting with Extensive Content Generation
by: Chen, Qihua, et al.
Published: (2024)
by: Chen, Qihua, et al.
Published: (2024)
EffiVED:Efficient Video Editing via Text-instruction Diffusion Models
by: Zhang, Zhenghao, et al.
Published: (2024)
by: Zhang, Zhenghao, et al.
Published: (2024)
Preacher: Paper-to-Video Agentic System
by: Liu, Jingwei, et al.
Published: (2025)
by: Liu, Jingwei, et al.
Published: (2025)
Online Reasoning Video Object Segmentation
by: Liu, Jinyuan, et al.
Published: (2026)
by: Liu, Jinyuan, et al.
Published: (2026)
HunyuanProver: A Scalable Data Synthesis Framework and Guided Tree Search for Automated Theorem Proving
by: Li, Yang, et al.
Published: (2024)
by: Li, Yang, et al.
Published: (2024)
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
by: Li, Jiaqi, et al.
Published: (2025)
by: Li, Jiaqi, et al.
Published: (2025)
GloTSFormer: Global Video Text Spotting Transformer
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
The Role of Video Generation in Enhancing Data-Limited Action Understanding
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Towards Scalable Video Anomaly Retrieval: A Synthetic Video-Text Benchmark
by: Yang, Shuyu, et al.
Published: (2025)
by: Yang, Shuyu, et al.
Published: (2025)
LaneTCA: Enhancing Video Lane Detection with Temporal Context Aggregation
by: Zhou, Keyi, et al.
Published: (2024)
by: Zhou, Keyi, et al.
Published: (2024)
Temporal-Consistent Video Restoration with Pre-trained Diffusion Models
by: Wang, Hengkang, et al.
Published: (2025)
by: Wang, Hengkang, et al.
Published: (2025)
RefAlign: Representation Alignment for Reference-to-Video Generation
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
SpatioTemporal Difference Network for Video Depth Super-Resolution
by: Wang, Zhengxue, et al.
Published: (2025)
by: Wang, Zhengxue, et al.
Published: (2025)
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
by: Wang, Zanyi, et al.
Published: (2025)
by: Wang, Zanyi, et al.
Published: (2025)
Language-guided Open-world Video Anomaly Detection under Weak Supervision
by: Liu, Zihao, et al.
Published: (2025)
by: Liu, Zihao, et al.
Published: (2025)
LaSe-E2V: Towards Language-guided Semantic-Aware Event-to-Video Reconstruction
by: Chen, Kanghao, et al.
Published: (2024)
by: Chen, Kanghao, et al.
Published: (2024)
UniCUE: Unified Recognition and Generation Framework for Chinese Cued Speech Video-to-Speech Generation
by: Wang, Jinting, et al.
Published: (2025)
by: Wang, Jinting, et al.
Published: (2025)
HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
by: HunyuanWorld Team, et al.
Published: (2025)
by: HunyuanWorld Team, et al.
Published: (2025)
PRVR: Partially Relevant Video Retrieval
by: Chen, Xianke, et al.
Published: (2022)
by: Chen, Xianke, et al.
Published: (2022)
VideoMAP: Toward Scalable Mamba-based Video Autoregressive Pretraining
by: Liu, Yunze, et al.
Published: (2025)
by: Liu, Yunze, et al.
Published: (2025)
MMAD: Multi-label Micro-Action Detection in Videos
by: Li, Kun, et al.
Published: (2024)
by: Li, Kun, et al.
Published: (2024)
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
by: Fan, Yingying, et al.
Published: (2025)
by: Fan, Yingying, et al.
Published: (2025)
Research on Short-Video Platform User Decision-Making via Multimodal Temporal Modeling and Reinforcement Learning
by: Wang, Jinmeiyang, et al.
Published: (2025)
by: Wang, Jinmeiyang, et al.
Published: (2025)
Bi-Directional Deep Contextual Video Compression
by: Sheng, Xihua, et al.
Published: (2024)
by: Sheng, Xihua, et al.
Published: (2024)
Tora: Trajectory-oriented Diffusion Transformer for Video Generation
by: Zhang, Zhenghao, et al.
Published: (2024)
by: Zhang, Zhenghao, et al.
Published: (2024)
Similar Items
-
HunyuanVideo 1.5 Technical Report
by: Wu, Bing, et al.
Published: (2025) -
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
by: Shan, Sizhe, et al.
Published: (2025) -
HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation
by: Huang, Ziyao, et al.
Published: (2025) -
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
by: Chen, Yi, et al.
Published: (2025) -
Hunyuan-Game: Industrial-grade Intelligent Game Creation Model
by: Li, Ruihuang, et al.
Published: (2025)