Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
Fuente:
arXiv
Saved in:
| Main Authors: | Yao, Zhengjian, Li, Yongzhi, Gao, Xinyuan, Chen, Quan, Jiang, Peng, Lu, Yanye |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Bridging Degradation Discrimination and Generation for Universal Image Restoration
by: Hu, JiaKui, et al.
Published: (2026)
by: Hu, JiaKui, et al.
Published: (2026)
Universal Image Restoration Pre-training via Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Universal Image Restoration Pre-training via Masked Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Exploiting Inherent Class Label: Towards Robust Scribble Supervised Semantic Segmentation
by: Zhang, Xinliang, et al.
Published: (2025)
by: Zhang, Xinliang, et al.
Published: (2025)
Chat-CBM: Towards Interactive Concept Bottleneck Models with Frozen Large Language Models
by: He, Hangzhou, et al.
Published: (2025)
by: He, Hangzhou, et al.
Published: (2025)
Auto-Regressively Generating Multi-View Consistent Images
by: Hu, JiaKui, et al.
Published: (2025)
by: Hu, JiaKui, et al.
Published: (2025)
Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration
by: Yao, Zhengjian, et al.
Published: (2026)
by: Yao, Zhengjian, et al.
Published: (2026)
RADA: Region-Aware Dual-encoder Auxiliary learning for Barely-supervised Medical Image Segmentation
by: Zeng, Shuang, et al.
Published: (2026)
by: Zeng, Shuang, et al.
Published: (2026)
Beyond Text: Frozen Large Language Models in Visual Signal Comprehension
by: Zhu, Lei, et al.
Published: (2024)
by: Zhu, Lei, et al.
Published: (2024)
AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs
by: Zhang, Xinliang, et al.
Published: (2025)
by: Zhang, Xinliang, et al.
Published: (2025)
VideoAuteur: Towards Long Narrative Video Generation
by: Xiao, Junfei, et al.
Published: (2025)
by: Xiao, Junfei, et al.
Published: (2025)
Range and Bird's Eye View Fused Cross-Modal Visual Place Recognition
by: Peng, Jianyi, et al.
Published: (2025)
by: Peng, Jianyi, et al.
Published: (2025)
STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives
by: Wang, Bo, et al.
Published: (2025)
by: Wang, Bo, et al.
Published: (2025)
AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation
by: Zhou, Milton, et al.
Published: (2026)
by: Zhou, Milton, et al.
Published: (2026)
Multi-Modal Proxy Learning Towards Personalized Visual Multiple Clustering
by: Yao, Jiawei, et al.
Published: (2024)
by: Yao, Jiawei, et al.
Published: (2024)
FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
Uni-Animator: Towards Unified Visual Colorization
by: Chen, Xinyuan, et al.
Published: (2026)
by: Chen, Xinyuan, et al.
Published: (2026)
Physics-Informed Conditional Diffusion for Motion-Robust Retinal Temporal Laser Speckle Contrast Imaging
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
by: Peng, Bo, et al.
Published: (2023)
by: Peng, Bo, et al.
Published: (2023)
LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models
by: Qiu, Han, et al.
Published: (2024)
by: Qiu, Han, et al.
Published: (2024)
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
by: He, Ruozhen, et al.
Published: (2026)
by: He, Ruozhen, et al.
Published: (2026)
Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs
by: Ghorbani, Saeed
Published: (2025)
by: Ghorbani, Saeed
Published: (2025)
HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes
by: Soroco, Mauricio, et al.
Published: (2026)
by: Soroco, Mauricio, et al.
Published: (2026)
Towards Flexible, Scalable, and Adaptive Multi-Modal Conditioned Face Synthesis
by: Ren, Jingjing, et al.
Published: (2023)
by: Ren, Jingjing, et al.
Published: (2023)
CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
by: Zhang, Zelin, et al.
Published: (2026)
by: Zhang, Zelin, et al.
Published: (2026)
Asymmetric Cross-Modal Knowledge Distillation: Bridging Modalities with Weak Semantic Consistency
by: Wei, Riling, et al.
Published: (2025)
by: Wei, Riling, et al.
Published: (2025)
Unified Sequence-to-Sequence Learning for Single- and Multi-Modal Visual Object Tracking
by: Chen, Xin, et al.
Published: (2023)
by: Chen, Xin, et al.
Published: (2023)
Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
by: Zhu, Lei, et al.
Published: (2024)
by: Zhu, Lei, et al.
Published: (2024)
Weaver: End-to-End Agentic System Training for Video Interleaved Reasoning
by: Shi, Yudi, et al.
Published: (2026)
by: Shi, Yudi, et al.
Published: (2026)
TimeWeaver: Age-Consistent Reference-Based Face Restoration with Identity Preservation
by: Song, Teer, et al.
Published: (2026)
by: Song, Teer, et al.
Published: (2026)
Modality Agnostic Efficient Long Range Encoder
by: Parag, Toufiq, et al.
Published: (2025)
by: Parag, Toufiq, et al.
Published: (2025)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives
by: Ji, Sihui, et al.
Published: (2025)
by: Ji, Sihui, et al.
Published: (2025)
V-CORE: Temporally Consistent Video Understanding for Video-LLM
by: Kang, Zhengjian, et al.
Published: (2026)
by: Kang, Zhengjian, et al.
Published: (2026)
ReWeaver: Towards Simulation-Ready and Topology-Accurate Garment Reconstruction
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
Wan-Weaver: Interleaved Multi-modal Generation via Decoupled Training
by: Xing, Jinbo, et al.
Published: (2026)
by: Xing, Jinbo, et al.
Published: (2026)
ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement
by: Yang, Yufeng, et al.
Published: (2026)
by: Yang, Yufeng, et al.
Published: (2026)
Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers
by: Ma, Xin, et al.
Published: (2025)
by: Ma, Xin, et al.
Published: (2025)
ScaleWeaver: Weaving Efficient Controllable T2I Generation with Multi-Scale Reference Attention
by: Liu, Keli, et al.
Published: (2025)
by: Liu, Keli, et al.
Published: (2025)
Similar Items
-
Bridging Degradation Discrimination and Generation for Universal Image Restoration
by: Hu, JiaKui, et al.
Published: (2026) -
Universal Image Restoration Pre-training via Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025) -
Universal Image Restoration Pre-training via Masked Degradation Classification
by: Hu, JiaKui, et al.
Published: (2025) -
Enhancing Image Restoration Transformer via Adaptive Translation Equivariance
by: Hu, JiaKui, et al.
Published: (2025) -
Exploiting Inherent Class Label: Towards Robust Scribble Supervised Semantic Segmentation
by: Zhang, Xinliang, et al.
Published: (2025)