BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zhaoyang, Qian, Dongjun, Su, Kai, Diao, Qishuai, Xia, Xiangyang, Liu, Chang, Yang, Wenfei, Zhang, Tianzhu, Yuan, Zehuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VC-LLM: Automated Advertisement Video Creation from Raw Footage using Multi-modal LLMs
by: Qian, Dongjun, et al.
Published: (2025)
by: Qian, Dongjun, et al.
Published: (2025)
FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
by: Luo, Xiangyang, et al.
Published: (2025)
by: Luo, Xiangyang, et al.
Published: (2025)
Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose Estimation
by: Lin, Xiao, et al.
Published: (2024)
by: Lin, Xiao, et al.
Published: (2024)
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
by: Hu, Teng, et al.
Published: (2025)
by: Hu, Teng, et al.
Published: (2025)
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
by: Wang, Zun, et al.
Published: (2026)
by: Wang, Zun, et al.
Published: (2026)
UniSOT: A Unified Framework for Multi-Modality Single Object Tracking
by: Ma, Yinchao, et al.
Published: (2025)
by: Ma, Yinchao, et al.
Published: (2025)
CrossWeaver: Cross-modal Weaving for Arbitrary-Modality Semantic Segmentation
by: Zhang, Zelin, et al.
Published: (2026)
by: Zhang, Zelin, et al.
Published: (2026)
VideoMemory: Toward Consistent Video Generation via Memory Integration
by: Zhou, Jinsong, et al.
Published: (2026)
by: Zhou, Jinsong, et al.
Published: (2026)
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
by: Chen, Zhuowei, et al.
Published: (2025)
by: Chen, Zhuowei, et al.
Published: (2025)
Learning Shape-Independent Transformation via Spherical Representations for Category-Level Object Pose Estimation
by: Ren, Huan, et al.
Published: (2025)
by: Ren, Huan, et al.
Published: (2025)
State Space Model Meets Transformer: A New Paradigm for 3D Object Detection
by: Wang, Chuxin, et al.
Published: (2025)
by: Wang, Chuxin, et al.
Published: (2025)
StruMamba3D: Exploring Structural Mamba for Self-supervised Point Cloud Representation Learning
by: Wang, Chuxin, et al.
Published: (2025)
by: Wang, Chuxin, et al.
Published: (2025)
Exploring Semantic Masked Autoencoder for Self-supervised Point Cloud Understanding
by: Zha, Yixin, et al.
Published: (2025)
by: Zha, Yixin, et al.
Published: (2025)
DA-Cal: Towards Cross-Domain Calibration in Semantic Segmentation
by: Li, Wangkai, et al.
Published: (2026)
by: Li, Wangkai, et al.
Published: (2026)
ActionParty: Multi-Subject Action Binding in Generative Video Games
by: Pondaven, Alexander, et al.
Published: (2026)
by: Pondaven, Alexander, et al.
Published: (2026)
Consistent Human Image and Video Generation with Spatially Conditioned Diffusion
by: Cao, Mingdeng, et al.
Published: (2024)
by: Cao, Mingdeng, et al.
Published: (2024)
UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
by: Du, Shian, et al.
Published: (2025)
by: Du, Shian, et al.
Published: (2025)
Cross-view Domain Generalization via Geometric Consistency for LiDAR Semantic Segmentation
by: Zhao, Jindong, et al.
Published: (2026)
by: Zhao, Jindong, et al.
Published: (2026)
Unifying Visual and Vision-Language Tracking via Contrastive Learning
by: Ma, Yinchao, et al.
Published: (2024)
by: Ma, Yinchao, et al.
Published: (2024)
Cross-Modal Bidirectional Interaction Model for Referring Remote Sensing Image Segmentation
by: Dong, Zhe, et al.
Published: (2024)
by: Dong, Zhe, et al.
Published: (2024)
ALIVE: Animate Your World with Lifelike Audio-Video Generation
by: Guo, Ying, et al.
Published: (2026)
by: Guo, Ying, et al.
Published: (2026)
Semantic-Consistent Bidirectional Contrastive Hashing for Noisy Multi-Label Cross-Modal Retrieval
by: Peng, Likang, et al.
Published: (2025)
by: Peng, Likang, et al.
Published: (2025)
Hollywood Town: Long-Video Generation via Cross-Modal Multi-Agent Orchestration
by: Wei, Zheng, et al.
Published: (2025)
by: Wei, Zheng, et al.
Published: (2025)
Dual-End Consistency Model
by: Dong, Linwei, et al.
Published: (2026)
by: Dong, Linwei, et al.
Published: (2026)
Multi-modal Attribute Prompting for Vision-Language Models
by: Liu, Xin, et al.
Published: (2024)
by: Liu, Xin, et al.
Published: (2024)
SMTrack: State-Aware Mamba for Efficient Temporal Modeling in Visual Tracking
by: Ma, Yinchao, et al.
Published: (2026)
by: Ma, Yinchao, et al.
Published: (2026)
MV-S2V: Multi-View Subject-Consistent Video Generation
by: Song, Ziyang, et al.
Published: (2026)
by: Song, Ziyang, et al.
Published: (2026)
Dense Audio-Visual Event Localization under Cross-Modal Consistency and Multi-Temporal Granularity Collaboration
by: Zhou, Ziheng, et al.
Published: (2024)
by: Zhou, Ziheng, et al.
Published: (2024)
Localization and Expansion: A Decoupled Framework for Point Cloud Few-shot Semantic Segmentation
by: Li, Zhaoyang, et al.
Published: (2024)
by: Li, Zhaoyang, et al.
Published: (2024)
Text-Video Retrieval via Variational Multi-Modal Hypergraph Networks
by: Li, Qian, et al.
Published: (2024)
by: Li, Qian, et al.
Published: (2024)
Exposing Cross-Modal Consistency for Fake News Detection in Short-Form Videos
by: Tian, Chong, et al.
Published: (2026)
by: Tian, Chong, et al.
Published: (2026)
Continual Cross-Modal Generalization
by: Xia, Yan, et al.
Published: (2025)
by: Xia, Yan, et al.
Published: (2025)
23‐1: Invited Paper: Effects of Dynamic Spatial Distortion on User Experience in Virtual Environments
by: Zhenping Xia, et al.
Published: (2025)
by: Zhenping Xia, et al.
Published: (2025)
CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation
by: Luo, Xiangyang, et al.
Published: (2026)
by: Luo, Xiangyang, et al.
Published: (2026)
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
by: Sun, Yang-Tian, et al.
Published: (2025)
by: Sun, Yang-Tian, et al.
Published: (2025)
Multiscale Matching Driven by Cross-Modal Similarity Consistency for Audio-Text Retrieval
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
Cross-Modal Clinical Knowledge Integration for Mammography Report Generation
by: Zhu, Jiayi, et al.
Published: (2026)
by: Zhu, Jiayi, et al.
Published: (2026)
Video Generation with Consistency Tuning
by: Wang, Chaoyi, et al.
Published: (2024)
by: Wang, Chaoyi, et al.
Published: (2024)
UniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous Driving
by: Chen, Rui, et al.
Published: (2024)
by: Chen, Rui, et al.
Published: (2024)
MV-Adapter: Multi-view Consistent Image Generation Made Easy
by: Huang, Zehuan, et al.
Published: (2024)
by: Huang, Zehuan, et al.
Published: (2024)
Similar Items
-
VC-LLM: Automated Advertisement Video Creation from Raw Footage using Multi-modal LLMs
by: Qian, Dongjun, et al.
Published: (2025) -
FilmWeaver: Weaving Consistent Multi-Shot Videos with Cache-Guided Autoregressive Diffusion
by: Luo, Xiangyang, et al.
Published: (2025) -
Instance-Adaptive and Geometric-Aware Keypoint Learning for Category-Level 6D Object Pose Estimation
by: Lin, Xiao, et al.
Published: (2024) -
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
by: Hu, Teng, et al.
Published: (2025) -
AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
by: Wang, Zun, et al.
Published: (2026)