OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | An, Zhaochong, Jia, Menglin, Qiu, Haonan, Zhou, Zijian, Huang, Xiaoke, Liu, Zhiheng, Ren, Weiming, Kahatapitiya, Kumara, Liu, Ding, He, Sen, Zhang, Chenyang, Xiang, Tao, Yang, Fanny, Belongie, Serge, Xie, Tian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Caching for Faster Video Generation with Diffusion Transformers
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Scaling Zero-Shot Reference-to-Video Generation
von: Zhou, Zijian, et al.
Veröffentlicht: (2025)
von: Zhou, Zijian, et al.
Veröffentlicht: (2025)
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
von: Qiu, Haonan, et al.
Veröffentlicht: (2025)
von: Qiu, Haonan, et al.
Veröffentlicht: (2025)
VecGlypher: Unified Vector Glyph Generation with Language Models
von: Huang, Xiaoke, et al.
Veröffentlicht: (2026)
von: Huang, Xiaoke, et al.
Veröffentlicht: (2026)
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Understanding Long Videos with Multimodal Language Models
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
von: Ranasinghe, Kanchana, et al.
Veröffentlicht: (2024)
VicTR: Video-conditioned Text Representations for Activity Recognition
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2023)
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
von: Li, Lei, et al.
Veröffentlicht: (2025)
von: Li, Lei, et al.
Veröffentlicht: (2025)
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
Rethinking Few-shot 3D Point Cloud Semantic Segmentation
von: An, Zhaochong, et al.
Veröffentlicht: (2024)
von: An, Zhaochong, et al.
Veröffentlicht: (2024)
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
von: Liu, Zhiheng, et al.
Veröffentlicht: (2025)
von: Liu, Zhiheng, et al.
Veröffentlicht: (2025)
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
von: An, Zhaochong, et al.
Veröffentlicht: (2026)
von: An, Zhaochong, et al.
Veröffentlicht: (2026)
MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
von: An, Zhaochong, et al.
Veröffentlicht: (2024)
von: An, Zhaochong, et al.
Veröffentlicht: (2024)
Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
One-Shot Multilingual Font Generation Via ViT
von: Wang, Zhiheng, et al.
Veröffentlicht: (2024)
von: Wang, Zhiheng, et al.
Veröffentlicht: (2024)
Too Many Frames, Not All Useful: Efficient Strategies for Long-Form Video QA
von: Park, Jongwoo, et al.
Veröffentlicht: (2024)
von: Park, Jongwoo, et al.
Veröffentlicht: (2024)
Object-Centric Diffusion for Efficient Video Editing
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)
Noise-Coded Illumination for Forensic and Photometric Video Analysis
von: Michael, Peter F., et al.
Veröffentlicht: (2025)
von: Michael, Peter F., et al.
Veröffentlicht: (2025)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
Video Understanding: From Geometry and Semantics to Unified Models
von: An, Zhaochong, et al.
Veröffentlicht: (2026)
von: An, Zhaochong, et al.
Veröffentlicht: (2026)
Is Meta-Learning Out? Rethinking Unsupervised Few-Shot Classification with Limited Entropy
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
von: Guan, Yunchuan, et al.
Veröffentlicht: (2025)
Unlearning-based Neural Interpretations
von: Choi, Ching Lam, et al.
Veröffentlicht: (2024)
von: Choi, Ching Lam, et al.
Veröffentlicht: (2024)
Large Vision-Language Models for Knowledge-Grounded Data Annotation of Memes
von: Deng, Shiling, et al.
Veröffentlicht: (2025)
von: Deng, Shiling, et al.
Veröffentlicht: (2025)
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models
von: Ren, Yixuan, et al.
Veröffentlicht: (2024)
von: Ren, Yixuan, et al.
Veröffentlicht: (2024)
Stitched Value Model for Diffusion Alignment
von: Go, Hyojun, et al.
Veröffentlicht: (2026)
von: Go, Hyojun, et al.
Veröffentlicht: (2026)
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
von: Gordon, Lucia, et al.
Veröffentlicht: (2026)
von: Gordon, Lucia, et al.
Veröffentlicht: (2026)
PhysConvex: Physics-Informed 3D Dynamic Convex Radiance Fields for Reconstruction and Simulation
von: Wang, Dan, et al.
Veröffentlicht: (2026)
von: Wang, Dan, et al.
Veröffentlicht: (2026)
Enhancing LLM-Based Text Classification in Political Science: Automatic Prompt Optimization and Dynamic Exemplar Selection for Few-Shot Learning
von: Liu, Menglin, et al.
Veröffentlicht: (2024)
von: Liu, Menglin, et al.
Veröffentlicht: (2024)
One Step is Enough: Multi-Agent Reinforcement Learning based on One-Step Policy Optimization for Order Dispatch on Ride-Sharing Platforms
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
von: Zhao, Zijian, et al.
Veröffentlicht: (2025)
UniPaint: Unified Space-time Video Inpainting via Mixture-of-Experts
von: Wan, Zhen, et al.
Veröffentlicht: (2024)
von: Wan, Zhen, et al.
Veröffentlicht: (2024)
HiddenObjects: Scalable Diffusion-Distilled Spatial Priors for Object Placement
von: Schouten, Marco, et al.
Veröffentlicht: (2026)
von: Schouten, Marco, et al.
Veröffentlicht: (2026)
Labeled Data Selection for Category Discovery
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
von: Zhao, Bingchen, et al.
Veröffentlicht: (2024)
Visual Story-Writing: Writing by Manipulating Visual Representations of Stories
von: Masson, Damien, et al.
Veröffentlicht: (2024)
von: Masson, Damien, et al.
Veröffentlicht: (2024)
Adaptive Super Resolution For One-Shot Talking-Head Generation
von: Song, Luchuan, et al.
Veröffentlicht: (2024)
von: Song, Luchuan, et al.
Veröffentlicht: (2024)
Better Language Models Exhibit Higher Visual Alignment
von: Ruthardt, Jona, et al.
Veröffentlicht: (2024)
von: Ruthardt, Jona, et al.
Veröffentlicht: (2024)
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
von: Yu, Hong-Tao, et al.
Veröffentlicht: (2025)
Familiarity-Based Open-Set Recognition Under Adversarial Attacks
von: Enevoldsen, Philip, et al.
Veröffentlicht: (2023)
von: Enevoldsen, Philip, et al.
Veröffentlicht: (2023)
Multi-Modal Framing Analysis of News
von: Arora, Arnav, et al.
Veröffentlicht: (2025)
von: Arora, Arnav, et al.
Veröffentlicht: (2025)
POEM: Precise Object-level Editing via MLLM control
von: Schouten, Marco, et al.
Veröffentlicht: (2025)
von: Schouten, Marco, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Adaptive Caching for Faster Video Generation with Diffusion Transformers
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024) -
Scaling Zero-Shot Reference-to-Video Generation
von: Zhou, Zijian, et al.
Veröffentlicht: (2025) -
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
von: Qiu, Haonan, et al.
Veröffentlicht: (2025) -
VecGlypher: Unified Vector Glyph Generation with Language Models
von: Huang, Xiaoke, et al.
Veröffentlicht: (2026) -
Language Repository for Long Video Understanding
von: Kahatapitiya, Kumara, et al.
Veröffentlicht: (2024)