ScaleDreamer: Scalable Text-to-3D Synthesis with Asynchronous Score Distillation
Fuente:
arXiv
Saved in:
| Main Authors: | Ma, Zhiyuan, Wei, Yuxiang, Zhang, Yabin, Zhu, Xiangyu, Lei, Zhen, Zhang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
by: Liu, Xiangyu, et al.
Published: (2026)
by: Liu, Xiangyu, et al.
Published: (2026)
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
by: Zhang, Yabin, et al.
Published: (2024)
by: Zhang, Yabin, et al.
Published: (2024)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
MOC-3D: Manifold-Order Consistency for Text-to-3D Generation
by: Fan, Chenyang, et al.
Published: (2026)
by: Fan, Chenyang, et al.
Published: (2026)
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
by: Wang, Yilin, et al.
Published: (2025)
by: Wang, Yilin, et al.
Published: (2025)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
by: Li, Chunyu, et al.
Published: (2026)
by: Li, Chunyu, et al.
Published: (2026)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
by: Chen, Nan, et al.
Published: (2024)
by: Chen, Nan, et al.
Published: (2024)
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
by: Guo, Zile, et al.
Published: (2026)
by: Guo, Zile, et al.
Published: (2026)
TMFNet: Two-Stream Multi-Channels Fusion Networks for Color Image Operation Chain Detection
by: Niu, Yakun, et al.
Published: (2024)
by: Niu, Yakun, et al.
Published: (2024)
PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
by: Jin, Chuhao, et al.
Published: (2025)
by: Jin, Chuhao, et al.
Published: (2025)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
by: Yang, Shuyu, et al.
Published: (2024)
by: Yang, Shuyu, et al.
Published: (2024)
GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models
by: Sun, Hao, et al.
Published: (2025)
by: Sun, Hao, et al.
Published: (2025)
Magic3DSketch: Create Colorful 3D Models From Sketch-Based 3D Modeling Guided by Text and Language-Image Pre-Training
by: Zang, Ying, et al.
Published: (2024)
by: Zang, Ying, et al.
Published: (2024)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
by: Qiao, Xiangshuo, et al.
Published: (2024)
by: Qiao, Xiangshuo, et al.
Published: (2024)
VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
HDiffTG: A Lightweight Hybrid Diffusion-Transformer-GCN Architecture for 3D Human Pose Estimation
by: Fu, Yajie, et al.
Published: (2025)
by: Fu, Yajie, et al.
Published: (2025)
Riemann-based Multi-scale Attention Reasoning Network for Text-3D Retrieval
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
TIPS: Text-Induced Pose Synthesis
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
G-Refine: A General Quality Refiner for Text-to-Image Generation
by: Li, Chunyi, et al.
Published: (2024)
by: Li, Chunyi, et al.
Published: (2024)
P-GSVC: Layered Progressive 2D Gaussian Splatting for Scalable Image and Video
by: Wang, Longan, et al.
Published: (2026)
by: Wang, Longan, et al.
Published: (2026)
TAVGBench: Benchmarking Text to Audible-Video Generation
by: Mao, Yuxin, et al.
Published: (2024)
by: Mao, Yuxin, et al.
Published: (2024)
Seeing Text in the Dark: Algorithm and Benchmark
by: Xu, Chengpei, et al.
Published: (2024)
by: Xu, Chengpei, et al.
Published: (2024)
Progressive Rendering Distillation: Adapting Stable Diffusion for Instant Text-to-Mesh Generation without 3D Data
by: Ma, Zhiyuan, et al.
Published: (2025)
by: Ma, Zhiyuan, et al.
Published: (2025)
DanceCamera3D: 3D Camera Movement Synthesis with Music and Dance
by: Wang, Zixuan, et al.
Published: (2024)
by: Wang, Zixuan, et al.
Published: (2024)
Hand1000: Generating Realistic Hands from Text with Only 1,000 Images
by: Zhang, Haozhuo, et al.
Published: (2024)
by: Zhang, Haozhuo, et al.
Published: (2024)
CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise
by: Yu, Fuyang, et al.
Published: (2024)
by: Yu, Fuyang, et al.
Published: (2024)
GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
by: Yu, Hongyun, et al.
Published: (2024)
by: Yu, Hongyun, et al.
Published: (2024)
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
by: Li, Zeju, et al.
Published: (2024)
by: Li, Zeju, et al.
Published: (2024)
TEn-CATG:Text-Enriched Audio-Visual Video Parsing with Multi-Scale Category-Aware Temporal Graph
by: Chen, Yaru, et al.
Published: (2025)
by: Chen, Yaru, et al.
Published: (2025)
FeatDistill: A Feature Distillation Enhanced Multi-Expert Ensemble Framework for Robust AI-generated Image Detection
by: Tu, Zhilin, et al.
Published: (2026)
by: Tu, Zhilin, et al.
Published: (2026)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
by: Zhu, Lingsi, et al.
Published: (2026)
by: Zhu, Lingsi, et al.
Published: (2026)
HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision
by: Zhou, Shengli, et al.
Published: (2025)
by: Zhou, Shengli, et al.
Published: (2025)
Tora3: Trajectory-Guided Audio-Video Generation with Physical Coherence
by: Liao, Junchao, et al.
Published: (2026)
by: Liao, Junchao, et al.
Published: (2026)
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction
by: Gong, Zixuan, et al.
Published: (2024)
by: Gong, Zixuan, et al.
Published: (2024)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
by: Zhang, Deyu, et al.
Published: (2025)
by: Zhang, Deyu, et al.
Published: (2025)
Instant3D: Instant Text-to-3D Generation
by: Li, Ming, et al.
Published: (2023)
by: Li, Ming, et al.
Published: (2023)
HDCompression: Hybrid-Diffusion Image Compression for Ultra-Low Bitrates
by: Lu, Lei, et al.
Published: (2025)
by: Lu, Lei, et al.
Published: (2025)
DreamMesh: Jointly Manipulating and Texturing Triangle Meshes for Text-to-3D Generation
by: Yang, Haibo, et al.
Published: (2024)
by: Yang, Haibo, et al.
Published: (2024)
SFFNet: Synergistic Feature Fusion Network With Dual-Domain Edge Enhancement for UAV Image Object Detection
by: Zhang, Wenfeng, et al.
Published: (2026)
by: Zhang, Wenfeng, et al.
Published: (2026)
GaussianForest: Hierarchical-Hybrid 3D Gaussian Splatting for Compressed Scene Modeling
by: Zhang, Fengyi, et al.
Published: (2024)
by: Zhang, Fengyi, et al.
Published: (2024)
Similar Items
-
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation
by: Liu, Xiangyu, et al.
Published: (2026) -
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
by: Zhang, Yabin, et al.
Published: (2024) -
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024) -
MOC-3D: Manifold-Order Consistency for Text-to-3D Generation
by: Fan, Chenyang, et al.
Published: (2026) -
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
by: Wang, Yilin, et al.
Published: (2025)