Saved in:
| Main Authors: | Tan, Shanwen, Li, Hao, Zhang, Jingtao, Jia, Xiaosong, Yang, Xue, Zhang, Shaofeng, Zhang, Yanyong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.09442 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
by: Lin, Kuanwei, et al.
Published: (2026)
by: Lin, Kuanwei, et al.
Published: (2026)
DreamWorld: Unified World Modeling in Video Generation
by: Tan, Boming, et al.
Published: (2026)
by: Tan, Boming, et al.
Published: (2026)
RISE-Video: Can Video Generators Decode Implicit World Rules?
by: Liu, Mingxin, et al.
Published: (2026)
by: Liu, Mingxin, et al.
Published: (2026)
UAR-NVC: A Unified AutoRegressive Framework for Memory-Efficient Neural Video Compression
by: Wang, Jia, et al.
Published: (2025)
by: Wang, Jia, et al.
Published: (2025)
Pack and Force Your Memory: Long-form and Consistent Video Generation
by: Wu, Xiaofei, et al.
Published: (2025)
by: Wu, Xiaofei, et al.
Published: (2025)
Video-Infinity: Distributed Long Video Generation
by: Tan, Zhenxiong, et al.
Published: (2024)
by: Tan, Zhenxiong, et al.
Published: (2024)
PMG: Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning
by: Xi, Yingjie, et al.
Published: (2025)
by: Xi, Yingjie, et al.
Published: (2025)
AdaCM$^2$: On Understanding Extremely Long-Term Video with Adaptive Cross-Modality Memory Reduction
by: Man, Yuanbin, et al.
Published: (2024)
by: Man, Yuanbin, et al.
Published: (2024)
RoboVIP: Multi-View Video Generation with Visual Identity Prompting Augments Robot Manipulation
by: Wang, Boyang, et al.
Published: (2026)
by: Wang, Boyang, et al.
Published: (2026)
Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning
by: Li, Jiazheng, et al.
Published: (2026)
by: Li, Jiazheng, et al.
Published: (2026)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
by: Zheng, Xinyi, et al.
Published: (2026)
by: Zheng, Xinyi, et al.
Published: (2026)
Streaming Video Understanding and Multi-round Interaction with Memory-enhanced Knowledge
by: Xiong, Haomiao, et al.
Published: (2025)
by: Xiong, Haomiao, et al.
Published: (2025)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
by: Wang, Yun, et al.
Published: (2025)
by: Wang, Yun, et al.
Published: (2025)
Free-Lunch Long Video Generation via Layer-Adaptive O.O.D Correction
by: Tian, Jiahao, et al.
Published: (2026)
by: Tian, Jiahao, et al.
Published: (2026)
Memory-Inspired Temporal Prompt Interaction for Text-Image Classification
by: Yu, Xinyao, et al.
Published: (2024)
by: Yu, Xinyao, et al.
Published: (2024)
Memory-Efficient Continual Learning Object Segmentation for Long Video
by: Nazemi, Amir, et al.
Published: (2023)
by: Nazemi, Amir, et al.
Published: (2023)
Dual-IPO: Dual-Iterative Preference Optimization for Text-to-Video Generation
by: Yang, Xiaomeng, et al.
Published: (2025)
by: Yang, Xiaomeng, et al.
Published: (2025)
Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events
by: Liu, Xiaolin, et al.
Published: (2026)
by: Liu, Xiaolin, et al.
Published: (2026)
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
by: Gan, Qijun, et al.
Published: (2025)
by: Gan, Qijun, et al.
Published: (2025)
Enhancing Long Video Generation Consistency without Tuning
by: Li, Xingyao, et al.
Published: (2024)
by: Li, Xingyao, et al.
Published: (2024)
SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker
by: Su, Junbin, et al.
Published: (2026)
by: Su, Junbin, et al.
Published: (2026)
Spatia: Video Generation with Updatable Spatial Memory
by: Zhao, Jinjing, et al.
Published: (2025)
by: Zhao, Jinjing, et al.
Published: (2025)
GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction
by: Shen, Yedong, et al.
Published: (2026)
by: Shen, Yedong, et al.
Published: (2026)
Capability-aware Prompt Reformulation Learning for Text-to-Image Generation
by: Zhan, Jingtao, et al.
Published: (2024)
by: Zhan, Jingtao, et al.
Published: (2024)
AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video Generation
by: Tan, Haoyue, et al.
Published: (2026)
by: Tan, Haoyue, et al.
Published: (2026)
Shallow Features Matter: Hierarchical Memory with Heterogeneous Interaction for Unsupervised Video Object Segmentation
by: Xiangyu, Zheng, et al.
Published: (2025)
by: Xiangyu, Zheng, et al.
Published: (2025)
Mixture of Contexts for Long Video Generation
by: Cai, Shengqu, et al.
Published: (2025)
by: Cai, Shengqu, et al.
Published: (2025)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025)
by: Ding, Yang, et al.
Published: (2025)
Yan: Foundational Interactive Video Generation
by: Ye, Deheng, et al.
Published: (2025)
by: Ye, Deheng, et al.
Published: (2025)
TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
by: Pan, Junwen, et al.
Published: (2025)
by: Pan, Junwen, et al.
Published: (2025)
EM-Vid: Training-Free Entity-Centric Memory for Efficient and Consistent Multi-Shot Video Generation
by: Vandersanden, Jente, et al.
Published: (2026)
by: Vandersanden, Jente, et al.
Published: (2026)
Semantic-Aware Adaptive Visual Memory for Streaming Video Understanding
by: Wu, Hang, et al.
Published: (2026)
by: Wu, Hang, et al.
Published: (2026)
Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2)
by: Yang, Zhenjie, et al.
Published: (2025)
by: Yang, Zhenjie, et al.
Published: (2025)
HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
by: Zhang, Haowei, et al.
Published: (2026)
by: Zhang, Haowei, et al.
Published: (2026)
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models
by: Zhang, Yue, et al.
Published: (2024)
by: Zhang, Yue, et al.
Published: (2024)
Point2RBox-v2: Rethinking Point-supervised Oriented Object Detection with Spatial Layout Among Instances
by: Yu, Yi, et al.
Published: (2025)
by: Yu, Yi, et al.
Published: (2025)
Adaptive Greedy Frame Selection for Long Video Understanding
by: Huang, Yuning, et al.
Published: (2026)
by: Huang, Yuning, et al.
Published: (2026)
Memory-Efficient Prompt Tuning for Incremental Histopathology Classification
by: Zhu, Yu, et al.
Published: (2024)
by: Zhu, Yu, et al.
Published: (2024)
Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity
by: Tian, Jiahao, et al.
Published: (2026)
by: Tian, Jiahao, et al.
Published: (2026)
Scaling the Long Video Understanding of Multimodal Large Language Models via Visual Memory Mechanism
by: Chen, Tao, et al.
Published: (2026)
by: Chen, Tao, et al.
Published: (2026)
Similar Items
-
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
by: Lin, Kuanwei, et al.
Published: (2026) -
DreamWorld: Unified World Modeling in Video Generation
by: Tan, Boming, et al.
Published: (2026) -
RISE-Video: Can Video Generators Decode Implicit World Rules?
by: Liu, Mingxin, et al.
Published: (2026) -
UAR-NVC: A Unified AutoRegressive Framework for Memory-Efficient Neural Video Compression
by: Wang, Jia, et al.
Published: (2025) -
Pack and Force Your Memory: Long-form and Consistent Video Generation
by: Wu, Xiaofei, et al.
Published: (2025)