DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Zhenhao, Wu, Xiaoshi, Lv, Zhengyao, Shi, Xiaoyu, Wang, Xintao, Wan, Pengfei, Gai, Kun, Wong, Kwan-Yee K. |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
par: Lv, Zhengyao, et autres
Publié: (2026)
par: Lv, Zhengyao, et autres
Publié: (2026)
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
par: Lv, Zhengyao, et autres
Publié: (2025)
par: Lv, Zhengyao, et autres
Publié: (2025)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
par: Ji, Sihui, et autres
Publié: (2025)
par: Ji, Sihui, et autres
Publié: (2025)
WorldMem: Long-term Consistent World Simulation with Memory
par: Xiao, Zeqi, et autres
Publié: (2025)
par: Xiao, Zeqi, et autres
Publié: (2025)
PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis
par: Lv, Zhengyao, et autres
Publié: (2024)
par: Lv, Zhengyao, et autres
Publié: (2024)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
par: Yu, Jiwen, et autres
Publié: (2025)
par: Yu, Jiwen, et autres
Publié: (2025)
SemanticGen: Video Generation in Semantic Space
par: Bai, Jianhong, et autres
Publié: (2025)
par: Bai, Jianhong, et autres
Publié: (2025)
StableWorld: Towards Stable and Consistent Long Interactive Video Generation
par: Yang, Ying, et autres
Publié: (2026)
par: Yang, Ying, et autres
Publié: (2026)
ConceptExpress: Harnessing Diffusion Models for Single-image Unsupervised Concept Extraction
par: Hao, Shaozhe, et autres
Publié: (2024)
par: Hao, Shaozhe, et autres
Publié: (2024)
CamCloneMaster: Enabling Reference-based Camera Control for Video Generation
par: Luo, Yawen, et autres
Publié: (2025)
par: Luo, Yawen, et autres
Publié: (2025)
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
par: Ye, Zixuan, et autres
Publié: (2025)
par: Ye, Zixuan, et autres
Publié: (2025)
LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
par: Wu, Di, et autres
Publié: (2026)
par: Wu, Di, et autres
Publié: (2026)
MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
par: Wang, Qinghe, et autres
Publié: (2025)
par: Wang, Qinghe, et autres
Publié: (2025)
Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
par: Guo, Yanjun, et autres
Publié: (2026)
par: Guo, Yanjun, et autres
Publié: (2026)
CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation
par: Wang, Qinghe, et autres
Publié: (2025)
par: Wang, Qinghe, et autres
Publié: (2025)
Scaling Image and Video Generation via Test-Time Evolutionary Search
par: He, Haoran, et autres
Publié: (2025)
par: He, Haoran, et autres
Publié: (2025)
RelightMaster: Precise Video Relighting with Multi-plane Light Images
par: Bian, Weikang, et autres
Publié: (2025)
par: Bian, Weikang, et autres
Publié: (2025)
UniVideo: Unified Understanding, Generation, and Editing for Videos
par: Wei, Cong, et autres
Publié: (2025)
par: Wei, Cong, et autres
Publié: (2025)
FasterCache: Training-Free Video Diffusion Model Acceleration with High Quality
par: Lv, Zhengyao, et autres
Publié: (2024)
par: Lv, Zhengyao, et autres
Publié: (2024)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
par: Lv, Zhengyao, et autres
Publié: (2025)
par: Lv, Zhengyao, et autres
Publié: (2025)
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
par: Xu, Yiyan, et autres
Publié: (2026)
par: Xu, Yiyan, et autres
Publié: (2026)
MemCam: Memory-Augmented Camera Control for Consistent Video Generation
par: Gao, Xinhang, et autres
Publié: (2026)
par: Gao, Xinhang, et autres
Publié: (2026)
DecLock: A Case of Decoupled Locking for Disaggregated Memory
par: Zhang, Hanze, et autres
Publié: (2025)
par: Zhang, Hanze, et autres
Publié: (2025)
Owl-1: Omni World Model for Consistent Long Video Generation
par: Huang, Yuanhui, et autres
Publié: (2024)
par: Huang, Yuanhui, et autres
Publié: (2024)
CoMem: Context Management with A Decoupled Long-Context Model
par: Zhang, Yuwei, et autres
Publié: (2026)
par: Zhang, Yuwei, et autres
Publié: (2026)
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder
par: Shi, Minglei, et autres
Publié: (2025)
par: Shi, Minglei, et autres
Publié: (2025)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
par: Ju, Xuan, et autres
Publié: (2025)
par: Ju, Xuan, et autres
Publié: (2025)
Latent Diffusion Model without Variational Autoencoder
par: Shi, Minglei, et autres
Publié: (2025)
par: Shi, Minglei, et autres
Publié: (2025)
OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
par: Zhao, Lin, et autres
Publié: (2026)
par: Zhao, Lin, et autres
Publié: (2026)
A Survey of Interactive Generative Video
par: Yu, Jiwen, et autres
Publié: (2025)
par: Yu, Jiwen, et autres
Publié: (2025)
MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
par: Yu, Wei, et autres
Publié: (2026)
par: Yu, Wei, et autres
Publié: (2026)
MemLong: Memory-Augmented Retrieval for Long Text Modeling
par: Liu, Weijie, et autres
Publié: (2024)
par: Liu, Weijie, et autres
Publié: (2024)
VipDiff: Towards Coherent and Diverse Video Inpainting via Training-free Denoising Diffusion Models
par: Xie, Chaohao, et autres
Publié: (2025)
par: Xie, Chaohao, et autres
Publié: (2025)
Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction
par: Cai, Xudong, et autres
Publié: (2025)
par: Cai, Xudong, et autres
Publié: (2025)
Minute-Scale Photonic Quantum Memory
par: Lv, You-Cai, et autres
Publié: (2025)
par: Lv, You-Cai, et autres
Publié: (2025)
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling
par: Luo, Yawen, et autres
Publié: (2026)
par: Luo, Yawen, et autres
Publié: (2026)
HyperMem: Hypergraph Memory for Long-Term Conversations
par: Yue, Juwei, et autres
Publié: (2026)
par: Yue, Juwei, et autres
Publié: (2026)
Mem-T: Densifying Rewards for Long-Horizon Memory Agents
par: Yue, Yanwei, et autres
Publié: (2026)
par: Yue, Yanwei, et autres
Publié: (2026)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
par: Zhang, Zeyu, et autres
Publié: (2025)
par: Zhang, Zeyu, et autres
Publié: (2025)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
par: Wu, Shengqiong, et autres
Publié: (2025)
par: Wu, Shengqiong, et autres
Publié: (2025)
Documents similaires
-
DUO-VSR: Dual-Stream Distillation for One-Step Video Super-Resolution
par: Lv, Zhengyao, et autres
Publié: (2026) -
Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
par: Lv, Zhengyao, et autres
Publié: (2025) -
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
par: Ji, Sihui, et autres
Publié: (2025) -
WorldMem: Long-term Consistent World Simulation with Memory
par: Xiao, Zeqi, et autres
Publié: (2025) -
PLACE: Adaptive Layout-Semantic Fusion for Semantic Image Synthesis
par: Lv, Zhengyao, et autres
Publié: (2024)