WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
Fuente:
arXiv
Saved in:
| Main Authors: | Oshima, Yuta, Iwasawa, Yusuke, Suzuki, Masahiro, Matsuo, Yutaka, Furuta, Hiroki |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
by: Oshima, Yuta, et al.
Published: (2025)
by: Oshima, Yuta, et al.
Published: (2025)
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
by: Oshima, Yuta, et al.
Published: (2024)
by: Oshima, Yuta, et al.
Published: (2024)
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
by: Kim, Bum Jun, et al.
Published: (2025)
by: Kim, Bum Jun, et al.
Published: (2025)
Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Enhancing Unimodal Latent Representations in Multimodal VAEs through Iterative Amortized Inference
by: Oshima, Yuta, et al.
Published: (2024)
by: Oshima, Yuta, et al.
Published: (2024)
Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases
by: Meo, Cristian, et al.
Published: (2024)
by: Meo, Cristian, et al.
Published: (2024)
Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
by: Furuta, Hiroki, et al.
Published: (2024)
by: Furuta, Hiroki, et al.
Published: (2024)
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
by: Minegishi, Gouki, et al.
Published: (2025)
by: Minegishi, Gouki, et al.
Published: (2025)
Frequency-Calibrated Membership Inference Attacks on Medical Image Diffusion Models
by: Zhao, Xinkai, et al.
Published: (2025)
by: Zhao, Xinkai, et al.
Published: (2025)
CLIP-like Model as a Foundational Density Ratio Estimator
by: Uchiyama, Fumiya, et al.
Published: (2025)
by: Uchiyama, Fumiya, et al.
Published: (2025)
Owl-1: Omni World Model for Consistent Long Video Generation
by: Huang, Yuanhui, et al.
Published: (2024)
by: Huang, Yuanhui, et al.
Published: (2024)
Vid2World: Crafting Video Diffusion Models to Interactive World Models
by: Huang, Siqiao, et al.
Published: (2025)
by: Huang, Siqiao, et al.
Published: (2025)
Understanding Emergent Misalignment via Feature Superposition Geometry
by: Minegishi, Gouki, et al.
Published: (2026)
by: Minegishi, Gouki, et al.
Published: (2026)
Beyond Spatial Compression: Interface-Centric Generative States for Open-World 3D Structure
by: Chen, Xiang, et al.
Published: (2026)
by: Chen, Xiang, et al.
Published: (2026)
AVID: Adapting Video Diffusion Models to World Models
by: Rigter, Marc, et al.
Published: (2024)
by: Rigter, Marc, et al.
Published: (2024)
Realtime Data-Efficient Portrait Stylization Based On Geometric Alignment
by: Wang, Xinrui, et al.
Published: (2022)
by: Wang, Xinrui, et al.
Published: (2022)
Olaf-World: Orienting Latent Actions for Video World Modeling
by: Jiang, Yuxin, et al.
Published: (2026)
by: Jiang, Yuxin, et al.
Published: (2026)
Pack and Detect: Fast Object Detection in Videos Using Region-of-Interest Packing
by: Kumar, Athindran Ramesh, et al.
Published: (2018)
by: Kumar, Athindran Ramesh, et al.
Published: (2018)
EC-Bench: Enumeration and Counting Benchmark for Ultra-Long Videos
by: Tsuchiya, Fumihiko, et al.
Published: (2026)
by: Tsuchiya, Fumihiko, et al.
Published: (2026)
iVideoGPT: Interactive VideoGPTs are Scalable World Models
by: Wu, Jialong, et al.
Published: (2024)
by: Wu, Jialong, et al.
Published: (2024)
Stable Consistency Tuning: Understanding and Improving Consistency Models
by: Wang, Fu-Yun, et al.
Published: (2024)
by: Wang, Fu-Yun, et al.
Published: (2024)
DMAD: Dual Memory Bank for Real-World Anomaly Detection
by: Hu, Jianlong, et al.
Published: (2024)
by: Hu, Jianlong, et al.
Published: (2024)
Light Interaction: Training-Free Inference Acceleration for Interactive Video World Models
by: Lu, Jiacheng, et al.
Published: (2026)
by: Lu, Jiacheng, et al.
Published: (2026)
WorldEval: World Model as Real-World Robot Policies Evaluator
by: Li, Yaxuan, et al.
Published: (2025)
by: Li, Yaxuan, et al.
Published: (2025)
OwMatch: Conditional Self-Labeling with Consistency for Open-World Semi-Supervised Learning
by: Niu, Shengjie, et al.
Published: (2024)
by: Niu, Shengjie, et al.
Published: (2024)
Video World Models with Long-term Spatial Memory
by: Wu, Tong, et al.
Published: (2025)
by: Wu, Tong, et al.
Published: (2025)
Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing
by: Lee, Dohun, et al.
Published: (2026)
by: Lee, Dohun, et al.
Published: (2026)
Improved Training Technique for Latent Consistency Models
by: Dao, Quan, et al.
Published: (2025)
by: Dao, Quan, et al.
Published: (2025)
Cross-View World Models
by: Sharma, Rishabh, et al.
Published: (2026)
by: Sharma, Rishabh, et al.
Published: (2026)
Slot Structured World Models
by: Collu, Jonathan, et al.
Published: (2024)
by: Collu, Jonathan, et al.
Published: (2024)
Instance-wise Supervision-level Optimization in Active Learning
by: Matsuo, Shinnosuke, et al.
Published: (2025)
by: Matsuo, Shinnosuke, et al.
Published: (2025)
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
by: Zhou, Runjie, et al.
Published: (2026)
by: Zhou, Runjie, et al.
Published: (2026)
Disentangled World Models: Learning to Transfer Semantic Knowledge from Distracting Videos for Reinforcement Learning
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
ColorizeDiffusion v2: Enhancing Reference-based Sketch Colorization Through Separating Utilities
by: Yan, Dingkun, et al.
Published: (2025)
by: Yan, Dingkun, et al.
Published: (2025)
Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models
by: Yada, Yuki, et al.
Published: (2025)
by: Yada, Yuki, et al.
Published: (2025)
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
by: Lillemark, Hansen Jin, et al.
Published: (2026)
by: Lillemark, Hansen Jin, et al.
Published: (2026)
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
by: Li, Xinhao, et al.
Published: (2024)
by: Li, Xinhao, et al.
Published: (2024)
Semantically Consistent Video Inpainting with Conditional Diffusion Models
by: Green, Dylan, et al.
Published: (2024)
by: Green, Dylan, et al.
Published: (2024)
DC-Merge: Improving Model Merging with Directional Consistency
by: Zhang, Han-Chen, et al.
Published: (2026)
by: Zhang, Han-Chen, et al.
Published: (2026)
Similar Items
-
Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
by: Oshima, Yuta, et al.
Published: (2025) -
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
by: Oshima, Yuta, et al.
Published: (2025) -
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
by: Oshima, Yuta, et al.
Published: (2024) -
Unlocking Noise-Resistant Vision: Key Architectural Secrets for Robust Models
by: Kim, Bum Jun, et al.
Published: (2025) -
Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
by: Furuta, Hiroki, et al.
Published: (2024)