Disentangling 3D from Large Vision-Language Models for Controlled Portrait Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Nick Yiwen, Caliskan, Akin, Kicanaoglu, Berkay, Tompkin, James, Kim, Hyeongwoo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAV: Personalized Head Avatar from Unstructured Video Collection
by: Caliskan, Akin, et al.
Published: (2024)
by: Caliskan, Akin, et al.
Published: (2024)
The GAN is dead; long live the GAN! A Modern GAN Baseline
by: Huang, Yiwen, et al.
Published: (2025)
by: Huang, Yiwen, et al.
Published: (2025)
Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation
by: Nocentini, Federico, et al.
Published: (2026)
by: Nocentini, Federico, et al.
Published: (2026)
FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
by: Tang, Jiapeng, et al.
Published: (2025)
by: Tang, Jiapeng, et al.
Published: (2025)
DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis
by: Gu, Yuming, et al.
Published: (2023)
by: Gu, Yuming, et al.
Published: (2023)
Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis
by: Ye, Zhenhui, et al.
Published: (2024)
by: Ye, Zhenhui, et al.
Published: (2024)
FILP-3D: Enhancing 3D Few-shot Class-incremental Learning with Pre-trained Vision-Language Models
by: Xu, Wan, et al.
Published: (2023)
by: Xu, Wan, et al.
Published: (2023)
Q-Bench-Portrait: Benchmarking Multimodal Large Language Models on Portrait Image Quality Perception
by: Wu, Sijing, et al.
Published: (2026)
by: Wu, Sijing, et al.
Published: (2026)
MegaPortrait: Revisiting Diffusion Control for High-fidelity Portrait Generation
by: Yang, Han, et al.
Published: (2024)
by: Yang, Han, et al.
Published: (2024)
Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation
by: Tan, Weipeng, et al.
Published: (2025)
by: Tan, Weipeng, et al.
Published: (2025)
Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
by: Kim, Jungeun, et al.
Published: (2024)
by: Kim, Jungeun, et al.
Published: (2024)
PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment
by: Ji, Chaonan, et al.
Published: (2026)
by: Ji, Chaonan, et al.
Published: (2026)
DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion Representations
by: Shi, Yuxiang, et al.
Published: (2025)
by: Shi, Yuxiang, et al.
Published: (2025)
Portrait3D: Text-Guided High-Quality 3D Portrait Generation Using Pyramid Representation and GANs Prior
by: Wu, Yiqian, et al.
Published: (2024)
by: Wu, Yiqian, et al.
Published: (2024)
DiffSLT: Enhancing Diversity in Sign Language Translation via Diffusion Model
by: Moon, JiHwan, et al.
Published: (2024)
by: Moon, JiHwan, et al.
Published: (2024)
OmniSDF: Scene Reconstruction using Omnidirectional Signed Distance Functions and Adaptive Binoctrees
by: Kim, Hakyeong, et al.
Published: (2024)
by: Kim, Hakyeong, et al.
Published: (2024)
FG-Portrait: 3D Flow Guided Editable Portrait Animation
by: Xu, Yating, et al.
Published: (2026)
by: Xu, Yating, et al.
Published: (2026)
Talk3D: High-Fidelity Talking Portrait Synthesis via Personalized 3D Generative Prior
by: Ko, Jaehoon, et al.
Published: (2024)
by: Ko, Jaehoon, et al.
Published: (2024)
Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided Diffusion
by: Ma, Xingpei, et al.
Published: (2025)
by: Ma, Xingpei, et al.
Published: (2025)
PerformRecast: Expression and Head Pose Disentanglement for Portrait Video Editing
by: Liang, Jiadong, et al.
Published: (2026)
by: Liang, Jiadong, et al.
Published: (2026)
On the Adversarial Robustness of 3D Large Vision-Language Models
by: Liu, Chao, et al.
Published: (2026)
by: Liu, Chao, et al.
Published: (2026)
HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation
by: Xu, Zunnan, et al.
Published: (2025)
by: Xu, Zunnan, et al.
Published: (2025)
LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control
by: Guo, Jianzhu, et al.
Published: (2024)
by: Guo, Jianzhu, et al.
Published: (2024)
Zero-Shot Monocular Scene Flow Estimation in the Wild
by: Liang, Yiqing, et al.
Published: (2025)
by: Liang, Yiqing, et al.
Published: (2025)
Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language Models
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
by: An, Na Min, et al.
Published: (2025)
by: An, Na Min, et al.
Published: (2025)
Watch Closely: Mitigating Object Hallucinations in Large Vision-Language Models with Disentangled Decoding
by: Ma, Ruiqi, et al.
Published: (2025)
by: Ma, Ruiqi, et al.
Published: (2025)
P3T: Prototypical Point-level Prompt Tuning with Enhanced Generalization for 3D Vision-Language Models
by: Jung, Geunyoung, et al.
Published: (2026)
by: Jung, Geunyoung, et al.
Published: (2026)
Bringing Your Portrait to 3D Presence
by: Zhang, Jiawei, et al.
Published: (2025)
by: Zhang, Jiawei, et al.
Published: (2025)
Zero-Shot Video Deraining with Video Diffusion Models
by: Varanka, Tuomas, et al.
Published: (2025)
by: Varanka, Tuomas, et al.
Published: (2025)
3DPR: Single Image 3D Portrait Relight using Generative Priors
by: Rao, Pramod, et al.
Published: (2025)
by: Rao, Pramod, et al.
Published: (2025)
PERSE: Personalized 3D Generative Avatars from A Single Portrait
by: Cha, Hyunsoo, et al.
Published: (2024)
by: Cha, Hyunsoo, et al.
Published: (2024)
3D Vision and Language Pretraining with Large-Scale Synthetic Data
by: Yang, Dejie, et al.
Published: (2024)
by: Yang, Dejie, et al.
Published: (2024)
PortraitCraft: A Benchmark for Portrait Composition Understanding and Generation
by: Sha, Yuyang, et al.
Published: (2026)
by: Sha, Yuyang, et al.
Published: (2026)
ARTIST: Improving the Generation of Text-rich Images with Disentangled Diffusion Models and Large Language Models
by: Zhang, Jianyi, et al.
Published: (2024)
by: Zhang, Jianyi, et al.
Published: (2024)
DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait Synthesis
by: Chen, Peiyin, et al.
Published: (2025)
by: Chen, Peiyin, et al.
Published: (2025)
Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement
by: Qiu, Weikang, et al.
Published: (2026)
by: Qiu, Weikang, et al.
Published: (2026)
Know3D: Prompting 3D Generation with Knowledge from Vision-Language Models
by: Chen, Wenyue, et al.
Published: (2026)
by: Chen, Wenyue, et al.
Published: (2026)
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
by: Kim, Eunki, et al.
Published: (2025)
by: Kim, Eunki, et al.
Published: (2025)
Phantom of Latent for Large Language and Vision Models
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
Similar Items
-
PAV: Personalized Head Avatar from Unstructured Video Collection
by: Caliskan, Akin, et al.
Published: (2024) -
The GAN is dead; long live the GAN! A Modern GAN Baseline
by: Huang, Yiwen, et al.
Published: (2025) -
Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation
by: Nocentini, Federico, et al.
Published: (2026) -
FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and Viewpoint
by: Tang, Jiapeng, et al.
Published: (2025) -
DiffPortrait3D: Controllable Diffusion for Zero-Shot Portrait View Synthesis
by: Gu, Yuming, et al.
Published: (2023)