Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Jiayi, Hua, Changcheng, Chen, Qingchao, Peng, Yuxin, Liu, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024)
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
von: Zhang, Beiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Beiyuan, et al.
Veröffentlicht: (2024)
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
von: Mo, Wentao, et al.
Veröffentlicht: (2025)
von: Mo, Wentao, et al.
Veröffentlicht: (2025)
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)
von: Zhou, S. Z., et al.
Veröffentlicht: (2025)
FreeEnhance: Tuning-Free Image Enhancement via Content-Consistent Noising-and-Denoising Process
von: Luo, Yang, et al.
Veröffentlicht: (2024)
von: Luo, Yang, et al.
Veröffentlicht: (2024)
TAVGBench: Benchmarking Text to Audible-Video Generation
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
von: Mao, Yuxin, et al.
Veröffentlicht: (2024)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
von: Xie, Jingjing, et al.
Veröffentlicht: (2024)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
von: Zhang, Deyu, et al.
Veröffentlicht: (2025)
von: Zhang, Deyu, et al.
Veröffentlicht: (2025)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
DeepSPG: Exploring Deep Semantic Prior Guidance for Low-light Image Enhancement with Multimodal Learning
von: Lu, Jialang, et al.
Veröffentlicht: (2025)
von: Lu, Jialang, et al.
Veröffentlicht: (2025)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
von: Ilaslan, Muhammet Furkan, et al.
Veröffentlicht: (2024)
VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Bringing Textual Prompt to AI-Generated Image Quality Assessment
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
von: Qu, Bowen, et al.
Veröffentlicht: (2024)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
von: Qin, Yang, et al.
Veröffentlicht: (2023)
von: Qin, Yang, et al.
Veröffentlicht: (2023)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
von: Yang, Danni, et al.
Veröffentlicht: (2024)
von: Yang, Danni, et al.
Veröffentlicht: (2024)
Reference-Guided Identity Preserving Face Restoration
von: Zhou, Mo, et al.
Veröffentlicht: (2025)
von: Zhou, Mo, et al.
Veröffentlicht: (2025)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
von: E, Shaojun, et al.
Veröffentlicht: (2025)
von: E, Shaojun, et al.
Veröffentlicht: (2025)
Controllable Complex Human Motion Video Generation via Text-to-Skeleton Cascades
von: Taghipour, Ashkan, et al.
Veröffentlicht: (2026)
von: Taghipour, Ashkan, et al.
Veröffentlicht: (2026)
Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
von: Yang, Yuxin, et al.
Veröffentlicht: (2026)
Multimodal Class-aware Semantic Enhancement Network for Audio-Visual Video Parsing
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2024)
Training-and-Prompt-Free General Painterly Harmonization via Zero-Shot Disentenglement on Style and Content References
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
von: Hsiao, Teng-Fang, et al.
Veröffentlicht: (2024)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
von: Dai, Yusheng, et al.
Veröffentlicht: (2026)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
G-Refine: A General Quality Refiner for Text-to-Image Generation
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
Sliced Maximal Information Coefficient: A Training-Free Approach for Image Quality Assessment Enhancement
von: Xiao, Kang, et al.
Veröffentlicht: (2024)
von: Xiao, Kang, et al.
Veröffentlicht: (2024)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
Harnessing the Latent Diffusion Model for Training-Free Image Style Transfer
von: Masui, Kento, et al.
Veröffentlicht: (2024)
von: Masui, Kento, et al.
Veröffentlicht: (2024)
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
von: Ji, Yatai, et al.
Veröffentlicht: (2024)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
von: Mao, Junzhu, et al.
Veröffentlicht: (2025)
von: Mao, Junzhu, et al.
Veröffentlicht: (2025)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
von: Qiao, Xiangshuo, et al.
Veröffentlicht: (2024)
Advancing Unsupervised Low-light Image Enhancement: Noise Estimation, Illumination Interpolation, and Self-Regulation
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2023)
von: Liu, Xiaofeng, et al.
Veröffentlicht: (2023)
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
EntroAD: Structural Entropy-Guided Prompt Adaptation for Zero-Shot Anomaly Detection
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhao, Xinyu, et al.
Veröffentlicht: (2026)
STIV: Scalable Text and Image Conditioned Video Generation
von: Lin, Zongyu, et al.
Veröffentlicht: (2024)
von: Lin, Zongyu, et al.
Veröffentlicht: (2024)
A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
von: Hei, Nailei, et al.
Veröffentlicht: (2024)
von: Hei, Nailei, et al.
Veröffentlicht: (2024)
OpFlowTalker: Realistic and Natural Talking Face Generation via Optical Flow Guidance
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
von: Ge, Shuheng, et al.
Veröffentlicht: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
Towards Realistic Low-Light Image Enhancement via ISP Driven Data Modeling
von: Wang, Zhihua, et al.
Veröffentlicht: (2025)
von: Wang, Zhihua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
von: Yuan, Shenghai, et al.
Veröffentlicht: (2024) -
Follow-Your-MultiPose: Tuning-Free Multi-Character Text-to-Video Generation via Pose Guidance
von: Zhang, Beiyuan, et al.
Veröffentlicht: (2024) -
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
von: Mo, Wentao, et al.
Veröffentlicht: (2025) -
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
von: Zhou, S. Z., et al.
Veröffentlicht: (2025) -
FreeEnhance: Tuning-Free Image Enhancement via Content-Consistent Noising-and-Denoising Process
von: Luo, Yang, et al.
Veröffentlicht: (2024)