Phantom: Subject-consistent video generation via cross-modal alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Lijie, Ma, Tianxiang, Li, Bingchuan, Chen, Zhuowei, Liu, Jiawei, Li, Gen, Zhou, Siyu, He, Qian, Wu, Xinglong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
LibraGen: Playing a Balance Game in Subject-Driven Video Generation
von: Zhu, Jiahao, et al.
Veröffentlicht: (2026)
von: Zhu, Jiahao, et al.
Veröffentlicht: (2026)
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
von: Chen, Jinshu, et al.
Veröffentlicht: (2025)
von: Chen, Jinshu, et al.
Veröffentlicht: (2025)
DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)
VLA-Mark: A cross modal watermark for large vision-language alignment model
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
Customize Your Own Paired Data via Few-shot Way
von: Chen, Jinshu, et al.
Veröffentlicht: (2024)
von: Chen, Jinshu, et al.
Veröffentlicht: (2024)
I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
von: Feng, Wanquan, et al.
Veröffentlicht: (2024)
von: Feng, Wanquan, et al.
Veröffentlicht: (2024)
AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
von: Li, Wei, et al.
Veröffentlicht: (2025)
von: Li, Wei, et al.
Veröffentlicht: (2025)
I2VControl: Disentangled and Unified Video Motion Synthesis Control
von: Feng, Wanquan, et al.
Veröffentlicht: (2024)
von: Feng, Wanquan, et al.
Veröffentlicht: (2024)
SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment
von: Mao, Xinyu, et al.
Veröffentlicht: (2025)
von: Mao, Xinyu, et al.
Veröffentlicht: (2025)
Can video generation replace cinematographers? Research on the cinematic language of generated video
von: Li, Xiaozhe, et al.
Veröffentlicht: (2024)
von: Li, Xiaozhe, et al.
Veröffentlicht: (2024)
XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation
von: Chen, Bowen, et al.
Veröffentlicht: (2025)
von: Chen, Bowen, et al.
Veröffentlicht: (2025)
PuLID: Pure and Lightning ID Customization via Contrastive Alignment
von: Guo, Zinan, et al.
Veröffentlicht: (2024)
von: Guo, Zinan, et al.
Veröffentlicht: (2024)
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
von: Cong, Yuren, et al.
Veröffentlicht: (2023)
von: Cong, Yuren, et al.
Veröffentlicht: (2023)
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
von: Guo, Xu, et al.
Veröffentlicht: (2026)
von: Guo, Xu, et al.
Veröffentlicht: (2026)
3DGS-Enhancer: Enhancing Unbounded 3D Gaussian Splatting with View-consistent 2D Diffusion Priors
von: Liu, Xi, et al.
Veröffentlicht: (2024)
von: Liu, Xi, et al.
Veröffentlicht: (2024)
DartControl: A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control
von: Zhao, Kaifeng, et al.
Veröffentlicht: (2024)
von: Zhao, Kaifeng, et al.
Veröffentlicht: (2024)
HyperLoRA: Parameter-Efficient Adaptive Generation for Portrait Synthesis
von: Li, Mengtian, et al.
Veröffentlicht: (2025)
von: Li, Mengtian, et al.
Veröffentlicht: (2025)
Improving Diffusion-based Inverse Algorithms under Few-Step Constraint via Learnable Linear Extrapolation
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2025)
Show and Segment: Universal Medical Image Segmentation via In-Context Learning
von: Gao, Yunhe, et al.
Veröffentlicht: (2025)
von: Gao, Yunhe, et al.
Veröffentlicht: (2025)
FSMR: A Feature Swapping Multi-modal Reasoning Approach with Joint Textual and Visual Clues
von: Li, Shuang, et al.
Veröffentlicht: (2024)
von: Li, Shuang, et al.
Veröffentlicht: (2024)
MUSAR: Exploring Multi-Subject Customization from Single-Subject Dataset via Attention Routing
von: Guo, Zinan, et al.
Veröffentlicht: (2025)
von: Guo, Zinan, et al.
Veröffentlicht: (2025)
Towards multi-modal forgery representation learning for AI-generated video detection and localization
von: Le, Dat, et al.
Veröffentlicht: (2026)
von: Le, Dat, et al.
Veröffentlicht: (2026)
ECHOPulse: ECG controlled echocardio-grams video generation
von: Li, Yiwei, et al.
Veröffentlicht: (2024)
von: Li, Yiwei, et al.
Veröffentlicht: (2024)
Improving vision-language alignment with graph spiking hybrid Networks
von: Zhang, Siyu, et al.
Veröffentlicht: (2025)
von: Zhang, Siyu, et al.
Veröffentlicht: (2025)
AR-Diffusion: Asynchronous Video Generation with Auto-Regressive Diffusion
von: Sun, Mingzhen, et al.
Veröffentlicht: (2025)
von: Sun, Mingzhen, et al.
Veröffentlicht: (2025)
FFA Sora, video generation as fundus fluorescein angiography simulator
von: Wu, Xinyuan, et al.
Veröffentlicht: (2024)
von: Wu, Xinyuan, et al.
Veröffentlicht: (2024)
KPL: Training-Free Medical Knowledge Mining of Vision-Language Models
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2025)
von: Liu, Jiaxiang, et al.
Veröffentlicht: (2025)
GeoReasoner: Geo-localization with Reasoning in Street Views using a Large Vision-Language Model
von: Li, Ling, et al.
Veröffentlicht: (2024)
von: Li, Ling, et al.
Veröffentlicht: (2024)
Flow caching for autoregressive video generation
von: Ma, Yuexiao, et al.
Veröffentlicht: (2026)
von: Ma, Yuexiao, et al.
Veröffentlicht: (2026)
Model alignment using inter-modal bridges
von: Gholamzadeh, Ali, et al.
Veröffentlicht: (2025)
von: Gholamzadeh, Ali, et al.
Veröffentlicht: (2025)
FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation
von: Li, Gen, et al.
Veröffentlicht: (2026)
von: Li, Gen, et al.
Veröffentlicht: (2026)
MoDA: Multi-modal Diffusion Architecture for Talking Head Generation
von: Li, Xinyang, et al.
Veröffentlicht: (2025)
von: Li, Xinyang, et al.
Veröffentlicht: (2025)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
von: Chen, Nan, et al.
Veröffentlicht: (2024)
von: Chen, Nan, et al.
Veröffentlicht: (2024)
DreamID: High-Fidelity and Fast diffusion-based Face Swapping via Triplet ID Group Learning
von: Ye, Fulong, et al.
Veröffentlicht: (2025)
von: Ye, Fulong, et al.
Veröffentlicht: (2025)
Training Like a Medical Resident: Context-Prior Learning Toward Universal Medical Image Segmentation
von: Gao, Yunhe, et al.
Veröffentlicht: (2023)
von: Gao, Yunhe, et al.
Veröffentlicht: (2023)
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
von: He, Chaoqun, et al.
Veröffentlicht: (2026)
von: He, Chaoqun, et al.
Veröffentlicht: (2026)
Multi-modality transrectal ultrasound video classification for identification of clinically significant prostate cancer
von: Wu, Hong, et al.
Veröffentlicht: (2024)
von: Wu, Hong, et al.
Veröffentlicht: (2024)
A cross-modal network for facial expression recognition
von: Tian, Chunwei, et al.
Veröffentlicht: (2026)
von: Tian, Chunwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
von: Chen, Zhuowei, et al.
Veröffentlicht: (2025) -
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025) -
LibraGen: Playing a Balance Game in Subject-Driven Video Generation
von: Zhu, Jiahao, et al.
Veröffentlicht: (2026) -
OmniInsert: Mask-Free Video Insertion of Any Reference via Diffusion Transformer Models
von: Chen, Jinshu, et al.
Veröffentlicht: (2025) -
DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation
von: Liu, Jiawei, et al.
Veröffentlicht: (2025)