OmniMem: Scalable and Adaptive Memory Retrieval for Long Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Lin, Wu, Yushu, Gong, Yifan, Wang, Yanzhi, Zhao, Pu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fast and Memory-Efficient Video Diffusion Using Streamlined Inference
by: Zhan, Zheng, et al.
Published: (2024)
by: Zhan, Zheng, et al.
Published: (2024)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
by: Jin, Hongbo, et al.
Published: (2025)
by: Jin, Hongbo, et al.
Published: (2025)
Lotus: learning-based online thermal and latency variation management for two-stage detectors on edge devices
by: Gong, Yifan, et al.
Published: (2024)
by: Gong, Yifan, et al.
Published: (2024)
EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory
by: Li, Yuyang, et al.
Published: (2026)
by: Li, Yuyang, et al.
Published: (2026)
LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation
by: Xiao, Yang, et al.
Published: (2025)
by: Xiao, Yang, et al.
Published: (2025)
StoryMem: Multi-shot Long Video Storytelling with Memory
by: Zhang, Kaiwen, et al.
Published: (2025)
by: Zhang, Kaiwen, et al.
Published: (2025)
FluxMem: Adaptive Hierarchical Memory for Streaming Video Understanding
by: Xie, Yiweng, et al.
Published: (2026)
by: Xie, Yiweng, et al.
Published: (2026)
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
by: Xue, Zhucun, et al.
Published: (2025)
by: Xue, Zhucun, et al.
Published: (2025)
WorldMem: Long-term Consistent World Simulation with Memory
by: Xiao, Zeqi, et al.
Published: (2025)
by: Xiao, Zeqi, et al.
Published: (2025)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
by: Pu, Junfu, et al.
Published: (2026)
by: Pu, Junfu, et al.
Published: (2026)
DecMem: Towards Minute-Long Consistent World Generation with Decoupled Memory
by: Yang, Zhenhao, et al.
Published: (2026)
by: Yang, Zhenhao, et al.
Published: (2026)
Taming Diffusion for Dataset Distillation with High Representativeness
by: Zhao, Lin, et al.
Published: (2025)
by: Zhao, Lin, et al.
Published: (2025)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
by: Wen, Siwei, et al.
Published: (2026)
by: Wen, Siwei, et al.
Published: (2026)
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning
by: Gao, Xianqiang, et al.
Published: (2026)
by: Gao, Xianqiang, et al.
Published: (2026)
Grid: Omni Visual Generation
by: Wan, Cong, et al.
Published: (2024)
by: Wan, Cong, et al.
Published: (2024)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
by: Wang, Ziyang, et al.
Published: (2026)
by: Wang, Ziyang, et al.
Published: (2026)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
by: Zheng, Xinyi, et al.
Published: (2026)
by: Zheng, Xinyi, et al.
Published: (2026)
OmniRoam: World Wandering via Long-Horizon Panoramic Video Generation
by: Liu, Yuheng, et al.
Published: (2026)
by: Liu, Yuheng, et al.
Published: (2026)
MemCam: Memory-Augmented Camera Control for Consistent Video Generation
by: Gao, Xinhang, et al.
Published: (2026)
by: Gao, Xinhang, et al.
Published: (2026)
DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
by: Shen, Xuan, et al.
Published: (2025)
by: Shen, Xuan, et al.
Published: (2025)
AyE-Edge: Automated Deployment Space Search Empowering Accuracy yet Efficient Real-Time Object Detection on the Edge
by: Wu, Chao, et al.
Published: (2024)
by: Wu, Chao, et al.
Published: (2024)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
MosaicMem: Hybrid Spatial Memory for Controllable Video World Models
by: Yu, Wei, et al.
Published: (2026)
by: Yu, Wei, et al.
Published: (2026)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
FastCar: Cache Attentive Replay for Fast Auto-Regressive Video Generation on the Edge
by: Shen, Xuan, et al.
Published: (2025)
by: Shen, Xuan, et al.
Published: (2025)
Exploring Token Pruning in Vision State Space Models
by: Zhan, Zheng, et al.
Published: (2024)
by: Zhan, Zheng, et al.
Published: (2024)
O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding
by: Wu, Peiran, et al.
Published: (2026)
by: Wu, Peiran, et al.
Published: (2026)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
MemRoPE: Training-Free Infinite Video Generation via Evolving Memory Tokens
by: Kim, Youngrae, et al.
Published: (2026)
by: Kim, Youngrae, et al.
Published: (2026)
Omni-Video 2: Scaling MLLM-Conditioned Diffusion for Unified Video Generation and Editing
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
by: Liu, Yunze, et al.
Published: (2026)
by: Liu, Yunze, et al.
Published: (2026)
Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation
by: Zhang, Haojie, et al.
Published: (2024)
by: Zhang, Haojie, et al.
Published: (2024)
TeleMem: Building Long-Term and Multimodal Memory for Agentic AI
by: Chen, Chunliang, et al.
Published: (2025)
by: Chen, Chunliang, et al.
Published: (2025)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
by: Yang, Yanlai, et al.
Published: (2025)
by: Yang, Yanlai, et al.
Published: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
by: Han, ZhaoYang, et al.
Published: (2025)
by: Han, ZhaoYang, et al.
Published: (2025)
HIERAMP: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
Similar Items
-
Fast and Memory-Efficient Video Diffusion Using Streamlined Inference
by: Zhan, Zheng, et al.
Published: (2024) -
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
by: Ji, Sihui, et al.
Published: (2025) -
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
by: Jin, Hongbo, et al.
Published: (2025) -
Lotus: learning-based online thermal and latency variation management for two-stage detectors on edge devices
by: Gong, Yifan, et al.
Published: (2024) -
EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory
by: Li, Yuyang, et al.
Published: (2026)