PTTA: A Pure Text-to-Animation Framework for High-Quality Creation
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Ruiqi, Cai, Kaitong, Fan, Yijia, Wang, Keze |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
by: Zeng, Qinglin, et al.
Published: (2025)
by: Zeng, Qinglin, et al.
Published: (2025)
MAT-Agent: Adaptive Multi-Agent Training Optimization
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
Top-Down Semantic Refinement for Image Captioning
by: Zhang, Jusheng, et al.
Published: (2025)
by: Zhang, Jusheng, et al.
Published: (2025)
3DAlign-DAER: Dynamic Attention Policy and Efficient Retrieval Strategy for Fine-grained 3D-Text Alignment at Scale
by: Fan, Yijia, et al.
Published: (2025)
by: Fan, Yijia, et al.
Published: (2025)
Process-of-Thought Reasoning for Videos
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
3D-Agent:Tri-Modal Multi-Agent Collaboration for Scalable 3D Object Annotation
by: Zhang, Jusheng, et al.
Published: (2026)
by: Zhang, Jusheng, et al.
Published: (2026)
FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
by: Zhang, Jesen, et al.
Published: (2025)
by: Zhang, Jesen, et al.
Published: (2025)
StableAnimator: High-Quality Identity-Preserving Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
SirenPose: Dynamic Scene Reconstruction via Geometric Supervision
by: Cai, Kaitong, et al.
Published: (2025)
by: Cai, Kaitong, et al.
Published: (2025)
MotiF: Making Text Count in Image Animation with Motion Focal Loss
by: Wang, Shijie, et al.
Published: (2024)
by: Wang, Shijie, et al.
Published: (2024)
High-Quality 3D Creation from A Single Image Using Subject-Specific Knowledge Prior
by: Huang, Nan, et al.
Published: (2023)
by: Huang, Nan, et al.
Published: (2023)
A Stepwise Distillation Learning Strategy for Non-differentiable Visual Programming Frameworks on Visual Reasoning Tasks
by: Wan, Wentao, et al.
Published: (2023)
by: Wan, Wentao, et al.
Published: (2023)
GTMA: Dynamic Representation Optimization for OOD Vision-Language Models
by: Zhang, Jensen, et al.
Published: (2025)
by: Zhang, Jensen, et al.
Published: (2025)
STORM: Search-Guided Generative World Models for Robotic Manipulation
by: Lin, Wenjun, et al.
Published: (2025)
by: Lin, Wenjun, et al.
Published: (2025)
LoopAnimate: Loopable Salient Object Animation
by: Wang, Fanyi, et al.
Published: (2024)
by: Wang, Fanyi, et al.
Published: (2024)
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image Models
by: Zhang, Yiming, et al.
Published: (2023)
by: Zhang, Yiming, et al.
Published: (2023)
Understanding Pure Textual Reasoning for Blind Image Quality Assessment
by: Li, Yuan, et al.
Published: (2026)
by: Li, Yuan, et al.
Published: (2026)
Zero-shot High-fidelity and Pose-controllable Character Animation
by: Zhu, Bingwen, et al.
Published: (2024)
by: Zhu, Bingwen, et al.
Published: (2024)
GenesisTex2: Stable, Consistent and High-Quality Text-to-Texture Generation
by: Lu, Jiawei, et al.
Published: (2024)
by: Lu, Jiawei, et al.
Published: (2024)
Diffusion-Based Visual Art Creation: A Survey and New Perspectives
by: Wang, Bingyuan, et al.
Published: (2024)
by: Wang, Bingyuan, et al.
Published: (2024)
ResAgent: Entropy-based Prior Point Discovery and Visual Reasoning for Referring Expression Segmentation
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
by: Feng, Yukang, et al.
Published: (2025)
by: Feng, Yukang, et al.
Published: (2025)
PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
by: Wang, Baiqin, et al.
Published: (2025)
by: Wang, Baiqin, et al.
Published: (2025)
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding
by: Liu, Tao, et al.
Published: (2024)
by: Liu, Tao, et al.
Published: (2024)
CoF-T2I: Video Models as Pure Visual Reasoners for Text-to-Image Generation
by: Tong, Chengzhuo, et al.
Published: (2026)
by: Tong, Chengzhuo, et al.
Published: (2026)
EvAnimate: Event-conditioned Image-to-Video Generation for Human Animation
by: Qu, Qiang, et al.
Published: (2025)
by: Qu, Qiang, et al.
Published: (2025)
Animate Any Character in Any World
by: Wang, Yitong, et al.
Published: (2025)
by: Wang, Yitong, et al.
Published: (2025)
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration
by: Li, Wuyang, et al.
Published: (2026)
by: Li, Wuyang, et al.
Published: (2026)
Grounding Text-to-Image Diffusion Models for Controlled High-Quality Image Generation
by: Süleyman, Ahmad, et al.
Published: (2025)
by: Süleyman, Ahmad, et al.
Published: (2025)
Is this Generated Person Existed in Real-world? Fine-grained Detecting and Calibrating Abnormal Human-body
by: Wang, Zeqing, et al.
Published: (2024)
by: Wang, Zeqing, et al.
Published: (2024)
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
by: Wang, Ziyi, et al.
Published: (2026)
by: Wang, Ziyi, et al.
Published: (2026)
Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
by: Chen, Weiming, et al.
Published: (2025)
by: Chen, Weiming, et al.
Published: (2025)
Runge-Kutta Approximation and Decoupled Attention for Rectified Flow Inversion and Semantic Editing
by: Chen, Weiming, et al.
Published: (2025)
by: Chen, Weiming, et al.
Published: (2025)
Latent Bias Alignment for High-Fidelity Diffusion Inversion in Real-World Image Reconstruction and Manipulation
by: Chen, Weiming, et al.
Published: (2026)
by: Chen, Weiming, et al.
Published: (2026)
Medical Visual Prompting (MVP): A Unified Framework for Versatile and High-Quality Medical Image Segmentation
by: Chen, Yulin, et al.
Published: (2024)
by: Chen, Yulin, et al.
Published: (2024)
Versatile Multimodal Controls for Expressive Talking Human Animation
by: Qin, Zheng, et al.
Published: (2025)
by: Qin, Zheng, et al.
Published: (2025)
Similar Items
-
CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation
by: Zeng, Qinglin, et al.
Published: (2025) -
MAT-Agent: Adaptive Multi-Agent Training Optimization
by: Zhang, Jusheng, et al.
Published: (2025) -
HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
by: Zhang, Jusheng, et al.
Published: (2025) -
MM-CoT:A Benchmark for Probing Visual Chain-of-Thought Reasoning in Multimodal Models
by: Zhang, Jusheng, et al.
Published: (2025) -
Top-Down Semantic Refinement for Image Captioning
by: Zhang, Jusheng, et al.
Published: (2025)