MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Weinan, Lu, Yuning, Huang, Mengqi, Wang, Hualiang, Huang, Binyuan, Chen, Nan, Liu, Mu, Jiang, Jidong, Mao, Zhendong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
by: Huang, Binyuan, et al.
Published: (2026)
by: Huang, Binyuan, et al.
Published: (2026)
D$^2$iT: Dynamic Diffusion Transformer for Accurate Image Generation
by: Jia, Weinan, et al.
Published: (2025)
by: Jia, Weinan, et al.
Published: (2025)
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction
by: Dong, Zijian, et al.
Published: (2025)
by: Dong, Zijian, et al.
Published: (2025)
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
by: Chen, Nan, et al.
Published: (2025)
by: Chen, Nan, et al.
Published: (2025)
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
by: Wu, Bin, et al.
Published: (2026)
by: Wu, Bin, et al.
Published: (2026)
FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus
by: Jin, Qiaoqiao, et al.
Published: (2025)
by: Jin, Qiaoqiao, et al.
Published: (2025)
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
by: Lin, Yijing, et al.
Published: (2025)
by: Lin, Yijing, et al.
Published: (2025)
AutoMoT: A Unified Vision-Language-Action Model with Asynchronous Mixture-of-Transformers for End-to-End Autonomous Driving
by: Huang, Wenhui, et al.
Published: (2026)
by: Huang, Wenhui, et al.
Published: (2026)
Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation
by: Ma, Enhui, et al.
Published: (2024)
by: Ma, Enhui, et al.
Published: (2024)
DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
by: Wang, Wenchuan, et al.
Published: (2025)
by: Wang, Wenchuan, et al.
Published: (2025)
HDGlyph: A Hierarchical Disentangled Glyph-Based Framework for Long-Tail Text Rendering in Diffusion Models
by: Zhuang, Shuhan, et al.
Published: (2025)
by: Zhuang, Shuhan, et al.
Published: (2025)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
by: Yu, Ting, et al.
Published: (2024)
by: Yu, Ting, et al.
Published: (2024)
GEMINUS: Dual-aware Global and Scene-Adaptive Mixture-of-Experts for End-to-End Autonomous Driving
by: Wan, Chi, et al.
Published: (2025)
by: Wan, Chi, et al.
Published: (2025)
DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving
by: Yang, Zhenjie, et al.
Published: (2025)
by: Yang, Zhenjie, et al.
Published: (2025)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
by: Zhang, Chen-Lin, et al.
Published: (2025)
by: Zhang, Chen-Lin, et al.
Published: (2025)
End-to-End Facial Expression Detection in Long Videos
by: Fang, Yini, et al.
Published: (2025)
by: Fang, Yini, et al.
Published: (2025)
End-to-End Dense Video Grounding via Parallel Regression
by: Shi, Fengyuan, et al.
Published: (2021)
by: Shi, Fengyuan, et al.
Published: (2021)
Stream-T1: Test-Time Scaling for Streaming Video Generation
by: Tu, Yijing, et al.
Published: (2026)
by: Tu, Yijing, et al.
Published: (2026)
DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation
by: Huang, Mengqi, et al.
Published: (2022)
by: Huang, Mengqi, et al.
Published: (2022)
MoCha:End-to-End Video Character Replacement without Structural Guidance
by: Xu, Zhengbo, et al.
Published: (2026)
by: Xu, Zhengbo, et al.
Published: (2026)
Drive&Gen: Co-Evaluating End-to-End Driving and Video Generation Models
by: Wang, Jiahao, et al.
Published: (2025)
by: Wang, Jiahao, et al.
Published: (2025)
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
by: She, Dong, et al.
Published: (2025)
by: She, Dong, et al.
Published: (2025)
End-to-End Shared Attention Estimation via Group Detection with Feedback Refinement
by: Nakatani, Chihiro, et al.
Published: (2026)
by: Nakatani, Chihiro, et al.
Published: (2026)
LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer Learning
by: Fu, Fengyi, et al.
Published: (2025)
by: Fu, Fengyi, et al.
Published: (2025)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
by: Chen, Nan, et al.
Published: (2024)
by: Chen, Nan, et al.
Published: (2024)
Mixture-of-Shape-Experts (MoSE): End-to-End Shape Dictionary Framework to Prompt SAM for Generalizable Medical Segmentation
by: Wei, Jia, et al.
Published: (2025)
by: Wei, Jia, et al.
Published: (2025)
End-to-End Visual Autonomous Parking via Control-Aided Attention
by: Chen, Chao, et al.
Published: (2025)
by: Chen, Chao, et al.
Published: (2025)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
by: Guo, Yuwei, et al.
Published: (2025)
by: Guo, Yuwei, et al.
Published: (2025)
EVE: Towards End-to-End Video Subtitle Extraction with Vision-Language Models
by: Yu, Haiyang, et al.
Published: (2025)
by: Yu, Haiyang, et al.
Published: (2025)
Leverage Cross-Attention for End-to-End Open-Vocabulary Panoptic Reconstruction
by: Yu, Xuan, et al.
Published: (2025)
by: Yu, Xuan, et al.
Published: (2025)
Towards Collaborative Autonomous Driving: Simulation Platform and End-to-End System
by: Liu, Genjia, et al.
Published: (2024)
by: Liu, Genjia, et al.
Published: (2024)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
by: Huang, Mengqi, et al.
Published: (2024)
by: Huang, Mengqi, et al.
Published: (2024)
Guiding Attention in End-to-End Driving Models
by: Porres, Diego, et al.
Published: (2024)
by: Porres, Diego, et al.
Published: (2024)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
by: Zhang, Yaolun, et al.
Published: (2026)
by: Zhang, Yaolun, et al.
Published: (2026)
STORM: End-to-End Referring Multi-Object Tracking in Videos
by: Lu, Zijia, et al.
Published: (2026)
by: Lu, Zijia, et al.
Published: (2026)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
STARFlow-V: End-to-End Video Generative Modeling with Normalizing Flows
by: Gu, Jiatao, et al.
Published: (2025)
by: Gu, Jiatao, et al.
Published: (2025)
WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models
by: Inbasekar, Karthik, et al.
Published: (2026)
by: Inbasekar, Karthik, et al.
Published: (2026)
Mixture of Contexts for Long Video Generation
by: Cai, Shengqu, et al.
Published: (2025)
by: Cai, Shengqu, et al.
Published: (2025)
Similar Items
-
Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation
by: Huang, Binyuan, et al.
Published: (2026) -
D$^2$iT: Dynamic Diffusion Transformer for Accurate Image Generation
by: Jia, Weinan, et al.
Published: (2025) -
NativeTok: Native Visual Tokenization for Improved Image Generation
by: Wu, Bin, et al.
Published: (2026) -
MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction
by: Dong, Zijian, et al.
Published: (2025) -
LongAnimation: Long Animation Generation with Dynamic Global-Local Memory
by: Chen, Nan, et al.
Published: (2025)