Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Ziying, Lu, Xuequan, Zhao, Xinkui, Cheng, Guanjie, Deng, Shuiguang, Yin, Jianwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Robust Incomplete Multimodal Low-Rank Adaptation Approach for Emotion Recognition
by: Zhao, Xinkui, et al.
Published: (2025)
by: Zhao, Xinkui, et al.
Published: (2025)
Quality-Aware Robust Multi-View Clustering for Heterogeneous Observation Noise
by: Wu, Peihan, et al.
Published: (2026)
by: Wu, Peihan, et al.
Published: (2026)
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
by: Zhao, Xinkui, et al.
Published: (2025)
by: Zhao, Xinkui, et al.
Published: (2025)
TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References
by: Yu, Jiahong, et al.
Published: (2025)
by: Yu, Jiahong, et al.
Published: (2025)
Bridge 2D-3D: Uncertainty-aware Hierarchical Registration Network with Domain Alignment
by: Cheng, Zhixin, et al.
Published: (2025)
by: Cheng, Zhixin, et al.
Published: (2025)
SatFusion: A Unified Framework for Enhancing Remote Sensing Images via Multi-Frame and Multi-Source Images Fusion
by: Tong, Yufei, et al.
Published: (2025)
by: Tong, Yufei, et al.
Published: (2025)
LightRouter: Towards Efficient LLM Collaboration with Minimal Overhead
by: Zhang, Yifan, et al.
Published: (2025)
by: Zhang, Yifan, et al.
Published: (2025)
Shiva-DiT: Residual-Based Differentiable Top-$k$ Selection for Efficient Diffusion Transformers
by: Zhang, Jiaji, et al.
Published: (2026)
by: Zhang, Jiaji, et al.
Published: (2026)
DiffCrossGait: Trajectory-Level Alignment for 2D-3D Cross-Modal Gait Recognition via Latent Diffusion
by: Lu, Zhiyang, et al.
Published: (2026)
by: Lu, Zhiyang, et al.
Published: (2026)
Integrating Artificial Intelligence into Operating Systems: A Survey on Techniques, Applications, and Future Directions
by: Zhang, Yifan, et al.
Published: (2024)
by: Zhang, Yifan, et al.
Published: (2024)
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
by: Wang, Zitong, et al.
Published: (2025)
by: Wang, Zitong, et al.
Published: (2025)
LTM3D: Bridging Token Spaces for Conditional 3D Generation with Auto-Regressive Diffusion Framework
by: Kang, Xin, et al.
Published: (2025)
by: Kang, Xin, et al.
Published: (2025)
Omni123: Exploring 3D Native Foundation Models with Limited 3D Data by Unifying Text to 2D and 3D Generation
by: Ye, Chongjie, et al.
Published: (2026)
by: Ye, Chongjie, et al.
Published: (2026)
WalkVLM:Aid Visually Impaired People Walking by Vision Language Model
by: Yuan, Zhiqiang, et al.
Published: (2024)
by: Yuan, Zhiqiang, et al.
Published: (2024)
G2LTraj: A Global-to-Local Generation Approach for Trajectory Prediction
by: Zhang, Zhanwei, et al.
Published: (2024)
by: Zhang, Zhanwei, et al.
Published: (2024)
Ambiguous Medical Image Segmentation Using Diffusion Schrödinger Bridge
by: Baru, Lalith Bharadwaj, et al.
Published: (2025)
by: Baru, Lalith Bharadwaj, et al.
Published: (2025)
Bridging Text and Video Generation: A Survey
by: Kumar, Nilay, et al.
Published: (2025)
by: Kumar, Nilay, et al.
Published: (2025)
Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
by: Yang, Xianghui, et al.
Published: (2024)
by: Yang, Xianghui, et al.
Published: (2024)
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
by: Yu, Qiucheng, et al.
Published: (2026)
by: Yu, Qiucheng, et al.
Published: (2026)
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
by: Zhang, Frank, et al.
Published: (2024)
by: Zhang, Frank, et al.
Published: (2024)
An Efficient LiDAR-Camera Fusion Network for Multi-Class 3D Dynamic Object Detection and Trajectory Prediction
by: He, Yushen, et al.
Published: (2025)
by: He, Yushen, et al.
Published: (2025)
RelaxFlow: Text-Driven Amodal 3D Generation
by: Zhu, Jiayin, et al.
Published: (2026)
by: Zhu, Jiayin, et al.
Published: (2026)
Generating Surface for Text-to-3D using 2D Gaussian Splatting
by: Dong, Huanning, et al.
Published: (2025)
by: Dong, Huanning, et al.
Published: (2025)
SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation
by: Agrawal, Vaibhav, et al.
Published: (2026)
by: Agrawal, Vaibhav, et al.
Published: (2026)
ManipDreamer3D : Synthesizing Plausible Robotic Manipulation Video with Occupancy-aware 3D Trajectory
by: Li, Ying, et al.
Published: (2025)
by: Li, Ying, et al.
Published: (2025)
A Comprehensive Survey on Human Video Generation: Challenges, Methods, and Insights
by: Lei, Wentao, et al.
Published: (2024)
by: Lei, Wentao, et al.
Published: (2024)
CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature Leveraging
by: Ling, Zhiwei, et al.
Published: (2025)
by: Ling, Zhiwei, et al.
Published: (2025)
Text-to-3D Generation using Jensen-Shannon Score Distillation
by: Do, Khoi, et al.
Published: (2025)
by: Do, Khoi, et al.
Published: (2025)
Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk?
by: Li, Hongyu, et al.
Published: (2026)
by: Li, Hongyu, et al.
Published: (2026)
Walking Further: Semantic-aware Multimodal Gait Recognition Under Long-Range Conditions
by: Lu, Zhiyang, et al.
Published: (2026)
by: Lu, Zhiyang, et al.
Published: (2026)
Dual-Domain Representation Alignment: Bridging 2D and 3D Vision via Geometry-Aware Architecture Search
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Text2Avatar: Text to 3D Human Avatar Generation with Codebook-Driven Body Controllable Attribute
by: Gong, Chaoqun, et al.
Published: (2024)
by: Gong, Chaoqun, et al.
Published: (2024)
BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving
by: Liu, Shu, et al.
Published: (2025)
by: Liu, Shu, et al.
Published: (2025)
Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
by: Chen, Guanjie, et al.
Published: (2025)
by: Chen, Guanjie, et al.
Published: (2025)
Aes3D: Aesthetic Assessment in 3D Gaussian Splatting
by: Xu, Chuanzhi, et al.
Published: (2026)
by: Xu, Chuanzhi, et al.
Published: (2026)
Object-level Cross-view Geo-localization with Location Enhancement and Multi-Head Cross Attention
by: Huang, Zheyang, et al.
Published: (2025)
by: Huang, Zheyang, et al.
Published: (2025)
Wonder3D++: Cross-domain Diffusion for High-fidelity 3D Generation from a Single Image
by: Yang, Yuxiao, et al.
Published: (2025)
by: Yang, Yuxiao, et al.
Published: (2025)
Native and Compact Structured Latents for 3D Generation
by: Xiang, Jianfeng, et al.
Published: (2025)
by: Xiang, Jianfeng, et al.
Published: (2025)
VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation
by: Li, Mengtian, et al.
Published: (2026)
by: Li, Mengtian, et al.
Published: (2026)
Similar Items
-
A Robust Incomplete Multimodal Low-Rank Adaptation Approach for Emotion Recognition
by: Zhao, Xinkui, et al.
Published: (2025) -
Quality-Aware Robust Multi-View Clustering for Heterogeneous Observation Noise
by: Wu, Peihan, et al.
Published: (2026) -
Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
by: Zhao, Xinkui, et al.
Published: (2025) -
TrackTeller: Temporal Multimodal 3D Grounding for Behavior-Dependent Object References
by: Yu, Jiahong, et al.
Published: (2025) -
Bridge 2D-3D: Uncertainty-aware Hierarchical Registration Network with Domain Alignment
by: Cheng, Zhixin, et al.
Published: (2025)