SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dou, Weijia, Li, Hui, Cui, Jiahao, Zhou, Lei, Wang, Jingdong, Zhu, Siyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
Slot-VAE: Object-Centric Scene Generation with Slot Attention
von: Wang, Yanbo, et al.
Veröffentlicht: (2023)
von: Wang, Yanbo, et al.
Veröffentlicht: (2023)
Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
von: Zhao, Zengqun, et al.
Veröffentlicht: (2026)
von: Zhao, Zengqun, et al.
Veröffentlicht: (2026)
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
von: Agarwal, Vatsal, et al.
Veröffentlicht: (2026)
von: Agarwal, Vatsal, et al.
Veröffentlicht: (2026)
Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
von: Yang, Yanlai, et al.
Veröffentlicht: (2025)
StreamMeCo: Long-Term Agent Memory Compression for Efficient Streaming Video Understanding
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
von: Wang, Junxi, et al.
Veröffentlicht: (2026)
When Slots Compete: Slot Merging in Object-Centric Learning
von: Chatzisavvas, Christos, et al.
Veröffentlicht: (2026)
von: Chatzisavvas, Christos, et al.
Veröffentlicht: (2026)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
von: Han, Jiwook, et al.
Veröffentlicht: (2026)
von: Han, Jiwook, et al.
Veröffentlicht: (2026)
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
von: Li, Hui, et al.
Veröffentlicht: (2024)
von: Li, Hui, et al.
Veröffentlicht: (2024)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
von: Zhang, Haowei, et al.
Veröffentlicht: (2026)
SlotMatch: Distilling Object-Centric Representations for Unsupervised Video Segmentation
von: Grigore, Diana-Nicoleta, et al.
Veröffentlicht: (2025)
von: Grigore, Diana-Nicoleta, et al.
Veröffentlicht: (2025)
StreamMOS: Streaming Moving Object Segmentation with Multi-View Perception and Dual-Span Memory
von: Li, Zhiheng, et al.
Veröffentlicht: (2024)
von: Li, Zhiheng, et al.
Veröffentlicht: (2024)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams
von: Zhang, Haoji, et al.
Veröffentlicht: (2024)
von: Zhang, Haoji, et al.
Veröffentlicht: (2024)
Learning Global Object-Centric Representations via Disentangled Slot Attention
von: Chen, Tonglin, et al.
Veröffentlicht: (2024)
von: Chen, Tonglin, et al.
Veröffentlicht: (2024)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
von: Yu, Jiwen, et al.
Veröffentlicht: (2025)
von: Yu, Jiwen, et al.
Veröffentlicht: (2025)
VideoLLaMB: Long Streaming Video Understanding with Recurrent Memory Bridges
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
MetaSlot: Break Through the Fixed Number of Slots in Object-Centric Learning
von: Liu, Hongjia, et al.
Veröffentlicht: (2025)
von: Liu, Hongjia, et al.
Veröffentlicht: (2025)
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
von: Moon, WonJun, et al.
Veröffentlicht: (2026)
von: Moon, WonJun, et al.
Veröffentlicht: (2026)
VideoMemory: Toward Consistent Video Generation via Memory Integration
von: Zhou, Jinsong, et al.
Veröffentlicht: (2026)
von: Zhou, Jinsong, et al.
Veröffentlicht: (2026)
CacheFlow: Compressive Streaming Memory for Efficient Long-Form Video Understanding
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
von: Patel, Shrenik, et al.
Veröffentlicht: (2025)
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory
von: Manigrasso, Zaira, et al.
Veröffentlicht: (2024)
von: Manigrasso, Zaira, et al.
Veröffentlicht: (2024)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
Pyramidal Patchification Flow for Visual Generation
von: Li, Hui, et al.
Veröffentlicht: (2025)
von: Li, Hui, et al.
Veröffentlicht: (2025)
PlaySlot: Learning Inverse Latent Dynamics for Controllable Object-Centric Video Prediction and Planning
von: Villar-Corrales, Angel, et al.
Veröffentlicht: (2025)
von: Villar-Corrales, Angel, et al.
Veröffentlicht: (2025)
PRISM: Progressive Reasoning through Iterative Slot Memory for Vision
von: Wang, Ziyu, et al.
Veröffentlicht: (2026)
von: Wang, Ziyu, et al.
Veröffentlicht: (2026)
Learning Object-Centric Representations Based on Slots in Real World Scenarios
von: Akan, Adil Kaan
Veröffentlicht: (2025)
von: Akan, Adil Kaan
Veröffentlicht: (2025)
OpenSlot: Mixed Open-Set Recognition with Object-Centric Learning
von: Yin, Xu, et al.
Veröffentlicht: (2024)
von: Yin, Xu, et al.
Veröffentlicht: (2024)
Joint Modeling of Feature, Correspondence, and a Compressed Memory for Video Object Segmentation
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhang, Jiaming, et al.
Veröffentlicht: (2023)
Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
von: Li, Jiaye, et al.
Veröffentlicht: (2025)
von: Li, Jiaye, et al.
Veröffentlicht: (2025)
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2025)
OASIS: On-Demand Hierarchical Event Memory for Streaming Video Reasoning
von: Liang, Zhijia, et al.
Veröffentlicht: (2026)
von: Liang, Zhijia, et al.
Veröffentlicht: (2026)
Object-Centric Learning with Slot Mixture Module
von: Kirilenko, Daniil, et al.
Veröffentlicht: (2023)
von: Kirilenko, Daniil, et al.
Veröffentlicht: (2023)
Hierarchical Memory for Long Video QA
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
von: Yu, Yifei, et al.
Veröffentlicht: (2025)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
von: Wen, Siwei, et al.
Veröffentlicht: (2026)
von: Wen, Siwei, et al.
Veröffentlicht: (2026)
GLASS: Guided Latent Slot Diffusion for Object-Centric Learning
von: Singh, Krishnakant, et al.
Veröffentlicht: (2024)
von: Singh, Krishnakant, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026) -
Slot-VAE: Object-Centric Scene Generation with Slot Attention
von: Wang, Yanbo, et al.
Veröffentlicht: (2023) -
Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
von: Zhao, Zengqun, et al.
Veröffentlicht: (2026) -
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
von: Agarwal, Vatsal, et al.
Veröffentlicht: (2026) -
Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)