ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ge, Yuying, Ge, Yixiao, Li, Chen, Wang, Teng, Pu, Junfu, Li, Yizhuo, Qiu, Lu, Ma, Jin, Duan, Lisheng, Zuo, Xinyu, Luo, Jinwen, Gu, Weibo, Li, Zexuan, Zhang, Xiaojing, Tao, Yangyu, Hu, Han, Wang, Di, Shan, Ying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
von: Pu, Junfu, et al.
Veröffentlicht: (2025)
von: Pu, Junfu, et al.
Veröffentlicht: (2025)
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
von: Qiu, Lu, et al.
Veröffentlicht: (2025)
von: Qiu, Lu, et al.
Veröffentlicht: (2025)
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
von: Li, Yizhuo, et al.
Veröffentlicht: (2024)
von: Li, Yizhuo, et al.
Veröffentlicht: (2024)
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
Aligning Latent Spaces with Flow Priors
von: Li, Yizhuo, et al.
Veröffentlicht: (2025)
von: Li, Yizhuo, et al.
Veröffentlicht: (2025)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
von: Chen, Yi, et al.
Veröffentlicht: (2024)
von: Chen, Yi, et al.
Veröffentlicht: (2024)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
Exploring the Effect of Reinforcement Learning on Video Understanding: Insights from SEED-Bench-R1
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
von: Liu, Ruyang, et al.
Veröffentlicht: (2023)
von: Liu, Ruyang, et al.
Veröffentlicht: (2023)
SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension
von: Li, Bohao, et al.
Veröffentlicht: (2024)
von: Li, Bohao, et al.
Veröffentlicht: (2024)
GenHancer: Imperfect Generative Models are Secretly Strong Vision-Centric Enhancers
von: Ma, Shijie, et al.
Veröffentlicht: (2025)
von: Ma, Shijie, et al.
Veröffentlicht: (2025)
SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
HunyuanVideo: A Systematic Framework For Large Video Generative Models
von: Kong, Weijie, et al.
Veröffentlicht: (2024)
von: Kong, Weijie, et al.
Veröffentlicht: (2024)
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
von: Lin, Haokun, et al.
Veröffentlicht: (2025)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
von: Guo, Yuxin, et al.
Veröffentlicht: (2025)
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
von: Qiu, Lu, et al.
Veröffentlicht: (2024)
HunyuanVideo 1.5 Technical Report
von: Wu, Bing, et al.
Veröffentlicht: (2025)
von: Wu, Bing, et al.
Veröffentlicht: (2025)
AnimeGamer: Infinite Anime Life Simulation with Next Game State Prediction
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
von: Ge, Yuying, et al.
Veröffentlicht: (2024)
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
von: Shan, Sizhe, et al.
Veröffentlicht: (2025)
von: Shan, Sizhe, et al.
Veröffentlicht: (2025)
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models
von: Wu, Tao, et al.
Veröffentlicht: (2024)
von: Wu, Tao, et al.
Veröffentlicht: (2024)
SEED-Story: Multimodal Long Story Generation with Large Language Model
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
von: Hu, Teng, et al.
Veröffentlicht: (2025)
von: Hu, Teng, et al.
Veröffentlicht: (2025)
Supervised Fine-tuning in turn Improves Visual Foundation Models
von: Jiang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Jiang, Xiaohu, et al.
Veröffentlicht: (2024)
GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
von: Chen, Yi, et al.
Veröffentlicht: (2025)
von: Chen, Yi, et al.
Veröffentlicht: (2025)
UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling
von: Chen, Boyu, et al.
Veröffentlicht: (2026)
von: Chen, Boyu, et al.
Veröffentlicht: (2026)
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning
von: Chen, Yi, et al.
Veröffentlicht: (2023)
von: Chen, Yi, et al.
Veröffentlicht: (2023)
Empirical models for calculating soil wetting patterns under surface drip irrigation systems: A comprehensive analysis
von: Ge Li, et al.
Veröffentlicht: (2024)
von: Ge Li, et al.
Veröffentlicht: (2024)
ST-LLM: Large Language Models Are Effective Temporal Learners
von: Liu, Ruyang, et al.
Veröffentlicht: (2024)
von: Liu, Ruyang, et al.
Veröffentlicht: (2024)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
von: Li, Kunchang, et al.
Veröffentlicht: (2023)
VideoChat: Chat-Centric Video Understanding
von: Li, KunChang, et al.
Veröffentlicht: (2023)
von: Li, KunChang, et al.
Veröffentlicht: (2023)
Hunyuan-MT Technical Report
von: Zheng, Mao, et al.
Veröffentlicht: (2025)
von: Zheng, Mao, et al.
Veröffentlicht: (2025)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
HunyuanOCR Technical Report
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
von: Hunyuan Vision Team, et al.
Veröffentlicht: (2025)
Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
Optimized Live 4K Video Multicast
von: He, Zhaoyuan, et al.
Veröffentlicht: (2023)
von: He, Zhaoyuan, et al.
Veröffentlicht: (2023)
DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
von: Chen, Yi, et al.
Veröffentlicht: (2026)
von: Chen, Yi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ARC-Chapter: Structuring Hour-Long Videos into Navigable Chapters and Hierarchical Summaries
von: Pu, Junfu, et al.
Veröffentlicht: (2025) -
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
von: Ge, Yuying, et al.
Veröffentlicht: (2024) -
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation
von: Qiu, Lu, et al.
Veröffentlicht: (2025) -
DiCoDe: Diffusion-Compressed Deep Tokens for Autoregressive Video Generation with Language Models
von: Li, Yizhuo, et al.
Veröffentlicht: (2024) -
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)