RepVideo: Rethinking Cross-Layer Representation for Video Generation
Fuente:
arXiv
Guardado en:
| Autores principales: | Si, Chenyang, Fan, Weichen, Lv, Zhengyao, Huang, Ziqi, Qiao, Yu, Liu, Ziwei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
por: Lv, Zhengyao, et al.
Publicado: (2025)
por: Lv, Zhengyao, et al.
Publicado: (2025)
FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
por: Lv, Zhengyao, et al.
Publicado: (2024)
por: Lv, Zhengyao, et al.
Publicado: (2024)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
por: Lv, Zhengyao, et al.
Publicado: (2025)
por: Lv, Zhengyao, et al.
Publicado: (2025)
FreeInit: Bridging Initialization Gap in Video Diffusion Models
por: Wu, Tianxing, et al.
Publicado: (2023)
por: Wu, Tianxing, et al.
Publicado: (2023)
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
por: Yang, Ying, et al.
Publicado: (2026)
por: Yang, Ying, et al.
Publicado: (2026)
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
por: Gao, Jianxiong, et al.
Publicado: (2025)
por: Gao, Jianxiong, et al.
Publicado: (2025)
Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation
por: Chen, Gordon, et al.
Publicado: (2026)
por: Chen, Gordon, et al.
Publicado: (2026)
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
por: Huang, Ziqi, et al.
Publicado: (2024)
por: Huang, Ziqi, et al.
Publicado: (2024)
NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
por: Pan, Tianlin, et al.
Publicado: (2026)
por: Pan, Tianlin, et al.
Publicado: (2026)
VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
por: Zheng, Dian, et al.
Publicado: (2025)
por: Zheng, Dian, et al.
Publicado: (2025)
LongVie 2: Multimodal Controllable Ultra-Long Video World Model
por: Gao, Jianxiong, et al.
Publicado: (2025)
por: Gao, Jianxiong, et al.
Publicado: (2025)
Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT
por: Liu, Dongyang, et al.
Publicado: (2025)
por: Liu, Dongyang, et al.
Publicado: (2025)
Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
por: Fan, Weichen, et al.
Publicado: (2025)
por: Fan, Weichen, et al.
Publicado: (2025)
Evaluation Agent: Efficient and Promptable Evaluation Framework for Visual Generative Models
por: Zhang, Fan, et al.
Publicado: (2024)
por: Zhang, Fan, et al.
Publicado: (2024)
RealDPO: Real or Not Real, that is the Preference
por: Cheng, Guo, et al.
Publicado: (2025)
por: Cheng, Guo, et al.
Publicado: (2025)
The Prism Hypothesis: Harmonizing Semantic and Pixel Representations via Unified Autoencoding
por: Fan, Weichen, et al.
Publicado: (2025)
por: Fan, Weichen, et al.
Publicado: (2025)
DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
por: Lv, Zhengyao, et al.
Publicado: (2026)
por: Lv, Zhengyao, et al.
Publicado: (2026)
FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
por: Cao, Yukang, et al.
Publicado: (2025)
por: Cao, Yukang, et al.
Publicado: (2025)
DiverseAR: Boosting Diversity in Bitwise Autoregressive Image Generation
por: Yang, Ying, et al.
Publicado: (2025)
por: Yang, Ying, et al.
Publicado: (2025)
Latte: Latent Diffusion Transformer for Video Generation
por: Ma, Xin, et al.
Publicado: (2024)
por: Ma, Xin, et al.
Publicado: (2024)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
por: Li, Rui, et al.
Publicado: (2025)
por: Li, Rui, et al.
Publicado: (2025)
VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment
por: Meng, Shibei, et al.
Publicado: (2026)
por: Meng, Shibei, et al.
Publicado: (2026)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
por: Cheng, Zixu, et al.
Publicado: (2025)
por: Cheng, Zixu, et al.
Publicado: (2025)
MR. Video: "MapReduce" is the Principle for Long Video Understanding
por: Pang, Ziqi, et al.
Publicado: (2025)
por: Pang, Ziqi, et al.
Publicado: (2025)
RepNet-VSR: Reparameterizable Architecture for High-Fidelity Video Super-Resolution
por: Wu, Biao, et al.
Publicado: (2025)
por: Wu, Biao, et al.
Publicado: (2025)
HoLa: B-Rep Generation using a Holistic Latent Representation
por: Liu, Yilin, et al.
Publicado: (2025)
por: Liu, Yilin, et al.
Publicado: (2025)
DenoiseRep: Denoising Model for Representation Learning
por: Xu, Zhengrui, et al.
Publicado: (2024)
por: Xu, Zhengrui, et al.
Publicado: (2024)
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
por: Gu, Zekai, et al.
Publicado: (2025)
por: Gu, Zekai, et al.
Publicado: (2025)
Video2BEV: Transforming Drone Videos to BEVs for Video-based Geo-localization
por: Ju, Hao, et al.
Publicado: (2024)
por: Ju, Hao, et al.
Publicado: (2024)
GeoVideo: Introducing Geometric Regularization into Video Generation Model
por: Bai, Yunpeng, et al.
Publicado: (2025)
por: Bai, Yunpeng, et al.
Publicado: (2025)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
por: Fan, Linfeng, et al.
Publicado: (2026)
por: Fan, Linfeng, et al.
Publicado: (2026)
Lighting-grounded Video Generation with Renderer-based Agent Reasoning
por: Cai, Ziqi, et al.
Publicado: (2026)
por: Cai, Ziqi, et al.
Publicado: (2026)
Stencil: Subject-Driven Generation with Context Guidance
por: Chen, Gordon, et al.
Publicado: (2025)
por: Chen, Gordon, et al.
Publicado: (2025)
Demystifying Video Reasoning
por: Wang, Ruisi, et al.
Publicado: (2026)
por: Wang, Ruisi, et al.
Publicado: (2026)
CineScale: Free Lunch in High-Resolution Cinematic Visual Generation
por: Qiu, Haonan, et al.
Publicado: (2025)
por: Qiu, Haonan, et al.
Publicado: (2025)
CoS: Chain-of-Shot Prompting for Long Video Understanding
por: Hu, Jian, et al.
Publicado: (2025)
por: Hu, Jian, et al.
Publicado: (2025)
Rethinking Video with a Universal Event-Based Representation
por: Freeman, Andrew
Publicado: (2024)
por: Freeman, Andrew
Publicado: (2024)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
por: Peng, Bo, et al.
Publicado: (2023)
por: Peng, Bo, et al.
Publicado: (2023)
Generative Omnimatte: Learning to Decompose Video into Layers
por: Lee, Yao-Chih, et al.
Publicado: (2024)
por: Lee, Yao-Chih, et al.
Publicado: (2024)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
por: Lin, Kun-Yu, et al.
Publicado: (2024)
por: Lin, Kun-Yu, et al.
Publicado: (2024)
Ejemplares similares
-
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
por: Lv, Zhengyao, et al.
Publicado: (2025) -
FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
por: Lv, Zhengyao, et al.
Publicado: (2024) -
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
por: Lv, Zhengyao, et al.
Publicado: (2025) -
FreeInit: Bridging Initialization Gap in Video Diffusion Models
por: Wu, Tianxing, et al.
Publicado: (2023) -
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
por: Yang, Ying, et al.
Publicado: (2026)