AToM: Amortized Text-to-Mesh using 2D Diffusion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qian, Guocheng, Cao, Junli, Siarohin, Aliaksandr, Kant, Yash, Wang, Chaoyang, Vasilkovsky, Michael, Lee, Hsin-Ying, Fang, Yuwei, Skorokhodov, Ivan, Zhuang, Peiye, Gilitschenski, Igor, Ren, Jian, Ghanem, Bernard, Aberman, Kfir, Tulyakov, Sergey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diffusion Priors for Dynamic View Synthesis from Monocular Videos
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
SPAD : Spatially Aware Multiview Diffusers
von: Kant, Yash, et al.
Veröffentlicht: (2024)
von: Kant, Yash, et al.
Veröffentlicht: (2024)
Pixel-Aligned Multi-View Generation with Depth Guided Decoder
von: Tang, Zhenggang, et al.
Veröffentlicht: (2024)
von: Tang, Zhenggang, et al.
Veröffentlicht: (2024)
4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024)
Dynamic Concepts Personalization from Single Videos
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
GTR: Improving Large 3D Reconstruction Models through Geometry and Texture Refinement
von: Zhuang, Peiye, et al.
Veröffentlicht: (2024)
von: Zhuang, Peiye, et al.
Veröffentlicht: (2024)
Mind the Time: Temporally-Controlled Multi-Event Video Generation
von: Wu, Ziyi, et al.
Veröffentlicht: (2024)
von: Wu, Ziyi, et al.
Veröffentlicht: (2024)
Hierarchical Patch Diffusion Models for High-Resolution Video Generation
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2024)
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2024)
AlphaFlow: Understanding and Improving MeanFlow Models
von: Zhang, Huijie, et al.
Veröffentlicht: (2025)
von: Zhang, Huijie, et al.
Veröffentlicht: (2025)
4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion Models
von: Yu, Heng, et al.
Veröffentlicht: (2024)
von: Yu, Heng, et al.
Veröffentlicht: (2024)
DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
von: Wu, Ziyi, et al.
Veröffentlicht: (2025)
von: Wu, Ziyi, et al.
Veröffentlicht: (2025)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)
Multi-subject Open-set Personalization in Video Generation
von: Chen, Tsai-Shien, et al.
Veröffentlicht: (2025)
von: Chen, Tsai-Shien, et al.
Veröffentlicht: (2025)
4Real-Video-V2: Fused View-Time Attention and Feedforward Reconstruction for 4D Scene Generation
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2026)
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2026)
ActionParty: Multi-Subject Action Binding in Generative Video Games
von: Pondaven, Alexander, et al.
Veröffentlicht: (2026)
von: Pondaven, Alexander, et al.
Veröffentlicht: (2026)
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation
von: Wang, Kuan-Chieh, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chieh, et al.
Veröffentlicht: (2024)
VIMI: Grounding Video Generation through Multi-modal Instruction
von: Fang, Yuwei, et al.
Veröffentlicht: (2024)
von: Fang, Yuwei, et al.
Veröffentlicht: (2024)
Improving Progressive Generation with Decomposable Flow Matching
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2025)
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2025)
Improving the Diffusability of Autoencoders
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2025)
von: Skorokhodov, Ivan, et al.
Veröffentlicht: (2025)
T2Bs: Text-to-Character Blendshapes via Video Generation
von: Luo, Jiahao, et al.
Veröffentlicht: (2025)
von: Luo, Jiahao, et al.
Veröffentlicht: (2025)
MyVLM: Personalizing VLMs for User-Specific Queries
von: Alaluf, Yuval, et al.
Veröffentlicht: (2024)
von: Alaluf, Yuval, et al.
Veröffentlicht: (2024)
Video Motion Transfer with Diffusion Transformers
von: Pondaven, Alexander, et al.
Veröffentlicht: (2024)
von: Pondaven, Alexander, et al.
Veröffentlicht: (2024)
Omni-ID: Holistic Identity Representation Designed for Generative Tasks
von: Qian, Guocheng, et al.
Veröffentlicht: (2024)
von: Qian, Guocheng, et al.
Veröffentlicht: (2024)
DELTA: Dense Efficient Long-range 3D Tracking for any video
von: Ngo, Tuan Duc, et al.
Veröffentlicht: (2024)
von: Ngo, Tuan Duc, et al.
Veröffentlicht: (2024)
AToM-Bot: Embodied Fulfillment of Unspoken Human Needs with Affective Theory of Mind
von: Ding, Wei, et al.
Veröffentlicht: (2024)
von: Ding, Wei, et al.
Veröffentlicht: (2024)
Nested Attention: Semantic-aware Attention Values for Concept Personalization
von: Patashnik, Or, et al.
Veröffentlicht: (2025)
von: Patashnik, Or, et al.
Veröffentlicht: (2025)
HyperHuman: Hyper-Realistic Human Generation with Latent Structural Diffusion
von: Liu, Xian, et al.
Veröffentlicht: (2023)
von: Liu, Xian, et al.
Veröffentlicht: (2023)
AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2024)
von: Haji-Ali, Moayed, et al.
Veröffentlicht: (2024)
H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models
von: Wu, Yushu, et al.
Veröffentlicht: (2025)
von: Wu, Yushu, et al.
Veröffentlicht: (2025)
AsCAN: Asymmetric Convolution-Attention Networks for Efficient Recognition and Generation
von: Kag, Anil, et al.
Veröffentlicht: (2024)
von: Kag, Anil, et al.
Veröffentlicht: (2024)
I Think, Therefore I Diffuse: Enabling Multimodal In-Context Reasoning in Diffusion Models
von: Mi, Zhenxing, et al.
Veröffentlicht: (2025)
von: Mi, Zhenxing, et al.
Veröffentlicht: (2025)
ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
von: Qian, Guocheng Gordon, et al.
Veröffentlicht: (2025)
von: Qian, Guocheng Gordon, et al.
Veröffentlicht: (2025)
SF-V: Single Forward Video Generation Model
von: Zhang, Zhixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhixing, et al.
Veröffentlicht: (2024)
AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
von: Han, Haonan, et al.
Veröffentlicht: (2024)
von: Han, Haonan, et al.
Veröffentlicht: (2024)
EasyV2V: A High-quality Instruction-based Video Editing Framework
von: Mai, Jinjie, et al.
Veröffentlicht: (2025)
von: Mai, Jinjie, et al.
Veröffentlicht: (2025)
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
von: Menapace, Willi, et al.
Veröffentlicht: (2024)
von: Menapace, Willi, et al.
Veröffentlicht: (2024)
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
von: Li, Runjia, et al.
Veröffentlicht: (2025)
von: Li, Runjia, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Diffusion Priors for Dynamic View Synthesis from Monocular Videos
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024) -
SPAD : Spatially Aware Multiview Diffusers
von: Kant, Yash, et al.
Veröffentlicht: (2024) -
Pixel-Aligned Multi-View Generation with Depth Guided Decoder
von: Tang, Zhenggang, et al.
Veröffentlicht: (2024) -
4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion
von: Wang, Chaoyang, et al.
Veröffentlicht: (2024) -
Dynamic Concepts Personalization from Single Videos
von: Abdal, Rameen, et al.
Veröffentlicht: (2025)