Concat-ID: Towards Universal Identity-Preserving Video Synthesis
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Yong, Yang, Zhuoyi, Teng, Jiayan, Gu, Xiaotao, Li, Chongxuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
by: Zhang, Zhenxing, et al.
Published: (2025)
by: Zhang, Zhenxing, et al.
Published: (2025)
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control
by: Zhong, Yong, et al.
Published: (2024)
by: Zhong, Yong, et al.
Published: (2024)
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
by: Liu, Shizhan, et al.
Published: (2025)
by: Liu, Shizhan, et al.
Published: (2025)
Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding
by: Lai, Yixuan, et al.
Published: (2026)
by: Lai, Yixuan, et al.
Published: (2026)
InstrAct: Towards Action-Centric Understanding in Instructional Videos
by: Yang, Zhuoyi, et al.
Published: (2026)
by: Yang, Zhuoyi, et al.
Published: (2026)
InstantID: Zero-shot Identity-Preserving Generation in Seconds
by: Wang, Qixun, et al.
Published: (2024)
by: Wang, Qixun, et al.
Published: (2024)
ConsistentID: Portrait Generation with Multimodal Fine-Grained Identity Preserving
by: Huang, Jiehui, et al.
Published: (2024)
by: Huang, Jiehui, et al.
Published: (2024)
Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
by: Liu, Zichuan, et al.
Published: (2025)
by: Liu, Zichuan, et al.
Published: (2025)
ID-Aligner: Enhancing Identity-Preserving Text-to-Image Generation with Reward Feedback Learning
by: Chen, Weifeng, et al.
Published: (2024)
by: Chen, Weifeng, et al.
Published: (2024)
ConcatPlexer: Additional Dim1 Batching for Faster ViTs
by: Han, Donghoon, et al.
Published: (2023)
by: Han, Donghoon, et al.
Published: (2023)
SCAIL: Towards Studio-Grade Character Animation via In-Context Learning of 3D-Consistent Pose Representations
by: Yan, Wenhao, et al.
Published: (2025)
by: Yan, Wenhao, et al.
Published: (2025)
FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
by: Wang, Yuanzhi, et al.
Published: (2026)
by: Wang, Yuanzhi, et al.
Published: (2026)
PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories
by: Tarrés, Gemma Canet, et al.
Published: (2026)
by: Tarrés, Gemma Canet, et al.
Published: (2026)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
by: Zhao, Min, et al.
Published: (2024)
by: Zhao, Min, et al.
Published: (2024)
Beyond the Pixels: VLM-based Evaluation of Identity Preservation in Reference-Guided Synthesis
by: Singhania, Aditi, et al.
Published: (2025)
by: Singhania, Aditi, et al.
Published: (2025)
ID-Sim: An Identity-Focused Similarity Metric
by: Chae, Julia, et al.
Published: (2026)
by: Chae, Julia, et al.
Published: (2026)
Finer-Personalization Rank: Fine-Grained Retrieval Examines Identity Preservation for Personalized Generation
by: Kilrain, Connor, et al.
Published: (2025)
by: Kilrain, Connor, et al.
Published: (2025)
FedVideoMAE: Efficient Privacy-Preserving Federated Video Moderation
by: Tao, Ziyuan, et al.
Published: (2025)
by: Tao, Ziyuan, et al.
Published: (2025)
CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion
by: Zheng, Wendi, et al.
Published: (2024)
by: Zheng, Wendi, et al.
Published: (2024)
Magic-Me: Identity-Specific Video Customized Diffusion
by: Ma, Ze, et al.
Published: (2024)
by: Ma, Ze, et al.
Published: (2024)
UniDemoiré: Towards Universal Image Demoiréing with Data Generation and Synthesis
by: Yang, Zemin, et al.
Published: (2025)
by: Yang, Zemin, et al.
Published: (2025)
TruePose: Human-Parsing-guided Attention Diffusion for Full-ID Preserving Pose Transfer
by: Xu, Zhihong, et al.
Published: (2025)
by: Xu, Zhihong, et al.
Published: (2025)
LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
by: Xing, Jiazheng, et al.
Published: (2026)
by: Xing, Jiazheng, et al.
Published: (2026)
UVCG: Leveraging Temporal Consistency for Universal Video Protection
by: Li, KaiZhou, et al.
Published: (2024)
by: Li, KaiZhou, et al.
Published: (2024)
SafeVid: Toward Safety Aligned Video Large Multimodal Models
by: Wang, Yixu, et al.
Published: (2025)
by: Wang, Yixu, et al.
Published: (2025)
StableAnimator: High-Quality Identity-Preserving Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
Video Text Preservation with Synthetic Text-Rich Videos
by: Liu, Ziyang, et al.
Published: (2025)
by: Liu, Ziyang, et al.
Published: (2025)
LVBench: An Extreme Long Video Understanding Benchmark
by: Wang, Weihan, et al.
Published: (2024)
by: Wang, Weihan, et al.
Published: (2024)
Privacy-Preserving SAM Quantization for Efficient Edge Intelligence in Healthcare
by: Li, Zhikai, et al.
Published: (2024)
by: Li, Zhikai, et al.
Published: (2024)
Preserving Forgery Artifacts: AI-Generated Video Detection at Native Scale
by: Li, Zhengcen, et al.
Published: (2026)
by: Li, Zhengcen, et al.
Published: (2026)
When Few Steps Are Enough: Training-Free Acceleration of Identity-Preserved Generation
by: Zheng, Dongqi
Published: (2026)
by: Zheng, Dongqi
Published: (2026)
On Memorization in Diffusion Models
by: Gu, Xiangming, et al.
Published: (2023)
by: Gu, Xiangming, et al.
Published: (2023)
S3-CLIP: Video Super Resolution for Person-ReID
by: Endrei, Tamas, et al.
Published: (2026)
by: Endrei, Tamas, et al.
Published: (2026)
Preserve or Modify? Context-Aware Evaluation for Balancing Preservation and Modification in Text-Guided Image Editing
by: Kim, Yoonjeon, et al.
Published: (2024)
by: Kim, Yoonjeon, et al.
Published: (2024)
SwinIFS: Landmark Guided Swin Transformer For Identity Preserving Face Super Resolution
by: Kausar, Habiba, et al.
Published: (2026)
by: Kausar, Habiba, et al.
Published: (2026)
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
by: Shi, Fengyuan, et al.
Published: (2023)
by: Shi, Fengyuan, et al.
Published: (2023)
WithAnyone: Towards Controllable and ID Consistent Image Generation
by: Xu, Hengyuan, et al.
Published: (2025)
by: Xu, Hengyuan, et al.
Published: (2025)
AerialVLA: A Vision-Language-Action Model for UAV Navigation via Minimalist End-to-End Control
by: Xu, Peng, et al.
Published: (2026)
by: Xu, Peng, et al.
Published: (2026)
LatXGen: Towards Radiation-Free and Accurate Quantitative Analysis of Sagittal Spinal Alignment Via Cross-Modal Radiographic View Synthesis
by: Zhao, Moxin, et al.
Published: (2025)
by: Zhao, Moxin, et al.
Published: (2025)
Similar Items
-
Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
by: Zhang, Zhenxing, et al.
Published: (2025) -
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control
by: Zhong, Yong, et al.
Published: (2024) -
Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability
by: Liu, Shizhan, et al.
Published: (2025) -
Slot-ID: Identity-Preserving Video Generation from Reference Videos via Slot-Based Temporal Identity Encoding
by: Lai, Yixuan, et al.
Published: (2026) -
InstrAct: Towards Action-Centric Understanding in Instructional Videos
by: Yang, Zhuoyi, et al.
Published: (2026)