EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vandersanden, Jente, Gadelha, Matheus, Huang, Chun-Hao P., Jeong, Hyeonho, Gryaditskaya, Yulia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TrajectoryMover: Generative Movement of Object Trajectories in Videos
von: Chhatre, Kiran, et al.
Veröffentlicht: (2026)
von: Chhatre, Kiran, et al.
Veröffentlicht: (2026)
SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time
von: Huang, Zhening, et al.
Veröffentlicht: (2025)
von: Huang, Zhening, et al.
Veröffentlicht: (2025)
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
von: Lee, Dohun, et al.
Veröffentlicht: (2026)
von: Lee, Dohun, et al.
Veröffentlicht: (2026)
Edge-preserving noise for diffusion models
von: Vandersanden, Jente, et al.
Veröffentlicht: (2024)
von: Vandersanden, Jente, et al.
Veröffentlicht: (2024)
Edge‐preserving noise for diffusion models
von: Jente Vandersanden, et al.
Veröffentlicht: (2026)
von: Jente Vandersanden, et al.
Veröffentlicht: (2026)
JOG3R: Towards 3D-Consistent Video Generators
von: Huang, Chun-Hao Paul, et al.
Veröffentlicht: (2025)
von: Huang, Chun-Hao Paul, et al.
Veröffentlicht: (2025)
Score-Based Generative Modeling through Anisotropic Stochastic Partial Differential Equations
von: Holl, Sascha, et al.
Veröffentlicht: (2026)
von: Holl, Sascha, et al.
Veröffentlicht: (2026)
BachVid: Training-Free Video Generation with Consistent Background and Character
von: Yan, Han, et al.
Veröffentlicht: (2025)
von: Yan, Han, et al.
Veröffentlicht: (2025)
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
von: He, Ruozhen, et al.
Veröffentlicht: (2026)
von: He, Ruozhen, et al.
Veröffentlicht: (2026)
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2024)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2024)
Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2025)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2025)
SketchingReality: From Freehand Scene Sketches To Photorealistic Images
von: Bourouis, Ahmed, et al.
Veröffentlicht: (2026)
von: Bourouis, Ahmed, et al.
Veröffentlicht: (2026)
Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2023)
Open Vocabulary Semantic Scene Sketch Understanding
von: Bourouis, Ahmed, et al.
Veröffentlicht: (2023)
von: Bourouis, Ahmed, et al.
Veröffentlicht: (2023)
DreamMotion: Space-Time Self-Similar Score Distillation for Zero-Shot Video Editing
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2024)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2024)
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
von: Lee, Dohun, et al.
Veröffentlicht: (2025)
von: Lee, Dohun, et al.
Veröffentlicht: (2025)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
von: Liu, Yexin, et al.
Veröffentlicht: (2025)
von: Liu, Yexin, et al.
Veröffentlicht: (2025)
EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
von: Qiu, Zongyang, et al.
Veröffentlicht: (2025)
EvoVid: Temporal-Centric Self-Evolution for Video Large Language Models
von: Huang, Shiqi, et al.
Veröffentlicht: (2026)
von: Huang, Shiqi, et al.
Veröffentlicht: (2026)
PruneVid: Visual Token Pruning for Efficient Video Large Language Models
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
von: Huang, Xiaohu, et al.
Veröffentlicht: (2024)
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video Synthesis
von: Liang, Feng, et al.
Veröffentlicht: (2023)
von: Liang, Feng, et al.
Veröffentlicht: (2023)
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
von: Wang, Boyang, et al.
Veröffentlicht: (2025)
GroundUp: Rapid Sketch-Based 3D City Massing
von: Unlu, Gizem Esra, et al.
Veröffentlicht: (2024)
von: Unlu, Gizem Esra, et al.
Veröffentlicht: (2024)
LightCache: Memory-Efficient, Training-Free Acceleration for Video Generation
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
von: Xiao, Yang, et al.
Veröffentlicht: (2025)
RelightVid: Temporal-Consistent Diffusion Model for Video Relighting
von: Fang, Ye, et al.
Veröffentlicht: (2025)
von: Fang, Ye, et al.
Veröffentlicht: (2025)
BlockVid: Block Diffusion for High-Quality and Consistent Minute-Long Video Generation
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zeyu, et al.
Veröffentlicht: (2025)
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Zhang, Zhicheng, et al.
Veröffentlicht: (2025)
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
von: Feng, X., et al.
Veröffentlicht: (2026)
von: Feng, X., et al.
Veröffentlicht: (2026)
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
von: Li, Hui, et al.
Veröffentlicht: (2024)
von: Li, Hui, et al.
Veröffentlicht: (2024)
Towards 3D VR-Sketch to 3D Shape Retrieval
von: Luo, Ling, et al.
Veröffentlicht: (2022)
von: Luo, Ling, et al.
Veröffentlicht: (2022)
PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning
von: Yan, Sikuan, et al.
Veröffentlicht: (2026)
von: Yan, Sikuan, et al.
Veröffentlicht: (2026)
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing
von: Couairon, Paul, et al.
Veröffentlicht: (2023)
von: Couairon, Paul, et al.
Veröffentlicht: (2023)
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
Training-free, Perceptually Consistent Low-Resolution Previews with High-Resolution Image for Efficient Workflows of Diffusion Models
von: Jeong, Wongi, et al.
Veröffentlicht: (2026)
von: Jeong, Wongi, et al.
Veröffentlicht: (2026)
QVAD: A Question-Centric Agentic Framework for Efficient and Training-Free Video Anomaly Detection
von: Bekit, Lokman, et al.
Veröffentlicht: (2026)
von: Bekit, Lokman, et al.
Veröffentlicht: (2026)
3D VR Sketch Guided 3D Shape Prototyping and Exploration
von: Luo, Ling, et al.
Veröffentlicht: (2023)
von: Luo, Ling, et al.
Veröffentlicht: (2023)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
VideoMemory: Toward Consistent Video Generation via Memory Integration
von: Zhou, Jinsong, et al.
Veröffentlicht: (2026)
von: Zhou, Jinsong, et al.
Veröffentlicht: (2026)
VidPanos: Generative Panoramic Videos from Casual Panning Videos
von: Ma, Jingwei, et al.
Veröffentlicht: (2024)
von: Ma, Jingwei, et al.
Veröffentlicht: (2024)
GimbalDiffusion: Gravity-Aware Camera Control for Video Generation
von: Fortier-Chouinard, Frédéric, et al.
Veröffentlicht: (2025)
von: Fortier-Chouinard, Frédéric, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TrajectoryMover: Generative Movement of Object Trajectories in Videos
von: Chhatre, Kiran, et al.
Veröffentlicht: (2026) -
SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time
von: Huang, Zhening, et al.
Veröffentlicht: (2025) -
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
von: Lee, Dohun, et al.
Veröffentlicht: (2026) -
Edge-preserving noise for diffusion models
von: Vandersanden, Jente, et al.
Veröffentlicht: (2024) -
Edge‐preserving noise for diffusion models
von: Jente Vandersanden, et al.
Veröffentlicht: (2026)