Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Zengqun, Lu, Yanzuo, Liu, Ziquan, Song, Jifei, Deng, Jiankang, Patras, Ioannis |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
by: Zhao, Zengqun, et al.
Published: (2026)
by: Zhao, Zengqun, et al.
Published: (2026)
Prompting Visual-Language Models for Dynamic Facial Expression Recognition
by: Zhao, Zengqun, et al.
Published: (2023)
by: Zhao, Zengqun, et al.
Published: (2023)
RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO
by: Lu, Yanzuo, et al.
Published: (2026)
by: Lu, Yanzuo, et al.
Published: (2026)
AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic Data
by: Zhao, Zengqun, et al.
Published: (2025)
by: Zhao, Zengqun, et al.
Published: (2025)
Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts
by: Cao, Yu, et al.
Published: (2025)
by: Cao, Yu, et al.
Published: (2025)
Enhancing Zero-Shot Facial Expression Recognition by LLM Knowledge Transfer
by: Zhao, Zengqun, et al.
Published: (2024)
by: Zhao, Zengqun, et al.
Published: (2024)
ReWind: Understanding Long Videos with Instructed Learnable Memory
by: Diko, Anxhelo, et al.
Published: (2024)
by: Diko, Anxhelo, et al.
Published: (2024)
Pack and Force Your Memory: Long-form and Consistent Video Generation
by: Wu, Xiaofei, et al.
Published: (2025)
by: Wu, Xiaofei, et al.
Published: (2025)
FG-Portrait: 3D Flow Guided Editable Portrait Animation
by: Xu, Yating, et al.
Published: (2026)
by: Xu, Yating, et al.
Published: (2026)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
by: Dou, Weijia, et al.
Published: (2026)
by: Dou, Weijia, et al.
Published: (2026)
Diffusion-Based Makeup Transfer with Facial Region-Aware Makeup Features
by: Gao, Zheng, et al.
Published: (2026)
by: Gao, Zheng, et al.
Published: (2026)
SAGS: Structure-Aware 3D Gaussian Splatting
by: Ververas, Evangelos, et al.
Published: (2024)
by: Ververas, Evangelos, et al.
Published: (2024)
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
by: Chen, Jiayu, et al.
Published: (2026)
by: Chen, Jiayu, et al.
Published: (2026)
Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
by: Yu, Jiwen, et al.
Published: (2025)
by: Yu, Jiwen, et al.
Published: (2025)
Context Forcing: Consistent Autoregressive Video Generation with Long Context
by: Chen, Shuo, et al.
Published: (2026)
by: Chen, Shuo, et al.
Published: (2026)
CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition
by: Sun, Zhonglin, et al.
Published: (2024)
by: Sun, Zhonglin, et al.
Published: (2024)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
by: Foteinopoulou, Niki Maria, et al.
Published: (2023)
FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation
by: Jiang, Chenhan, et al.
Published: (2026)
by: Jiang, Chenhan, et al.
Published: (2026)
Self-Supervised Facial Representation Learning with Facial Region Awareness
by: Gao, Zheng, et al.
Published: (2024)
by: Gao, Zheng, et al.
Published: (2024)
Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation
by: Wu, Mingqiang, et al.
Published: (2026)
by: Wu, Mingqiang, et al.
Published: (2026)
Forcing-KV: Hybrid KV Cache Compression for Efficient Autoregressive Video Diffusion Models
by: Ji, Yicheng, et al.
Published: (2026)
by: Ji, Yicheng, et al.
Published: (2026)
Unlocking the Potential of Diffusion Priors in Blind Face Restoration
by: Miao, Yunqi, et al.
Published: (2025)
by: Miao, Yunqi, et al.
Published: (2025)
ZeroGS: Training 3D Gaussian Splatting from Unposed Images
by: Chen, Yu, et al.
Published: (2024)
by: Chen, Yu, et al.
Published: (2024)
Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
by: Metaxas, Ioannis Maniadis, et al.
Published: (2024)
by: Metaxas, Ioannis Maniadis, et al.
Published: (2024)
Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
by: Huang, Junchao, et al.
Published: (2025)
by: Huang, Junchao, et al.
Published: (2025)
Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
by: Yang, Sicheng, et al.
Published: (2026)
by: Yang, Sicheng, et al.
Published: (2026)
Future Forcing: Future-aware Training-free KV Cache Policy for Autoregressive Video Generation
by: Luo, Jiayi, et al.
Published: (2026)
by: Luo, Jiayi, et al.
Published: (2026)
VidCtx: Context-aware Video Question Answering with Image Models
by: Goulas, Andreas, et al.
Published: (2024)
by: Goulas, Andreas, et al.
Published: (2024)
GlobalPointer: Large-Scale Plane Adjustment with Bi-Convex Relaxation
by: Liao, Bangyan, et al.
Published: (2024)
by: Liao, Bangyan, et al.
Published: (2024)
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
by: Liu, Zijian, et al.
Published: (2026)
by: Liu, Zijian, et al.
Published: (2026)
VideoMemory: Toward Consistent Video Generation via Memory Integration
by: Zhou, Jinsong, et al.
Published: (2026)
by: Zhou, Jinsong, et al.
Published: (2026)
CLIPCleaner: Cleaning Noisy Labels with CLIP
by: Feng, Chen, et al.
Published: (2024)
by: Feng, Chen, et al.
Published: (2024)
FOAA: Flattened Outer Arithmetic Attention For Multimodal Tumor Classification
by: Alwazzan, Omnia, et al.
Published: (2024)
by: Alwazzan, Omnia, et al.
Published: (2024)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning
by: Krestenitis, Marios, et al.
Published: (2026)
by: Krestenitis, Marios, et al.
Published: (2026)
UniMRSeg: Unified Modality-Relax Segmentation via Hierarchical Self-Supervised Compensation
by: Zhao, Xiaoqi, et al.
Published: (2025)
by: Zhao, Xiaoqi, et al.
Published: (2025)
Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
by: Guo, Yanjun, et al.
Published: (2026)
by: Guo, Yanjun, et al.
Published: (2026)
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
by: Singh, Abhishek Kumar, et al.
Published: (2024)
by: Singh, Abhishek Kumar, et al.
Published: (2024)
Color When It Counts: Grayscale-Guided Online Triggering for Always-On Streaming Video Sensing
by: Cai, Weitong, et al.
Published: (2026)
by: Cai, Weitong, et al.
Published: (2026)
Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
by: Liu, Kunhao, et al.
Published: (2025)
by: Liu, Kunhao, et al.
Published: (2025)
Similar Items
-
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
by: Zhao, Zengqun, et al.
Published: (2026) -
Prompting Visual-Language Models for Dynamic Facial Expression Recognition
by: Zhao, Zengqun, et al.
Published: (2023) -
RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO
by: Lu, Yanzuo, et al.
Published: (2026) -
AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic Data
by: Zhao, Zengqun, et al.
Published: (2025) -
Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts
by: Cao, Yu, et al.
Published: (2025)