Saved in:
| Main Authors: | Wang, Luozhou, Chen, Zhifei, Du, Yihua, Yan, Dongyu, Ge, Wenhang, Shen, Guibao, Xu, Xinli, Wu, Leyi, Chen, Man, Xu, Tianshuo, Ren, Peiran, Tao, Xin, Wan, Pengfei, Chen, Ying-Cong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.17067 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion
by: Chen, Zhifei, et al.
Published: (2024)
by: Chen, Zhifei, et al.
Published: (2024)
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
by: Ge, Wenhang, et al.
Published: (2026)
by: Ge, Wenhang, et al.
Published: (2026)
FlexPainter: Flexible and Multi-View Consistent Texture Generation
by: Yan, Dongyu, et al.
Published: (2025)
by: Yan, Dongyu, et al.
Published: (2025)
PRM: Photometric Stereo based Large Reconstruction Model
by: Ge, Wenhang, et al.
Published: (2024)
by: Ge, Wenhang, et al.
Published: (2024)
STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
by: Chen, Zhifei, et al.
Published: (2025)
by: Chen, Zhifei, et al.
Published: (2025)
Text-Anchored Score Composition: Tackling Condition Misalignment in Text-to-Image Diffusion Models
by: Wang, Luozhou, et al.
Published: (2023)
by: Wang, Luozhou, et al.
Published: (2023)
StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative Priors
by: Shen, Guibao, et al.
Published: (2025)
by: Shen, Guibao, et al.
Published: (2025)
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
by: Shen, Guibao, et al.
Published: (2024)
by: Shen, Guibao, et al.
Published: (2024)
Motion Forcing: A Decoupled Framework for Robust Video Generation in Motion Dynamics
by: Xu, Tianshuo, et al.
Published: (2026)
by: Xu, Tianshuo, et al.
Published: (2026)
Motion Inversion for Video Customization
by: Wang, Luozhou, et al.
Published: (2024)
by: Wang, Luozhou, et al.
Published: (2024)
FlexGen: Flexible Multi-View Generation from Text and Image Inputs
by: Xu, Xinli, et al.
Published: (2024)
by: Xu, Xinli, et al.
Published: (2024)
UniCalli: A Unified Diffusion Framework for Column-Level Generation and Recognition of Chinese Calligraphy
by: Xu, Tianshuo, et al.
Published: (2025)
by: Xu, Tianshuo, et al.
Published: (2025)
VideoMemory: Toward Consistent Video Generation via Memory Integration
by: Zhou, Jinsong, et al.
Published: (2026)
by: Zhou, Jinsong, et al.
Published: (2026)
Motion Dreamer: Boundary Conditional Motion Reasoning for Physically Coherent Video Generation
by: Xu, Tianshuo, et al.
Published: (2024)
by: Xu, Tianshuo, et al.
Published: (2024)
RectifiedHR: Enable Efficient High-Resolution Synthesis via Energy Rectification
by: Yang, Zhen, et al.
Published: (2025)
by: Yang, Zhen, et al.
Published: (2025)
GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs
by: Xu, Xinli, et al.
Published: (2024)
by: Xu, Xinli, et al.
Published: (2024)
Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset Generation
by: Lin, Jiantao, et al.
Published: (2025)
by: Lin, Jiantao, et al.
Published: (2025)
RoboEvolve: Co-Evolving Planner-Simulator for Robotic Manipulation with Limited Data
by: Chen, Harold Haodong, et al.
Published: (2026)
by: Chen, Harold Haodong, et al.
Published: (2026)
ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback
by: Guo, Litao, et al.
Published: (2025)
by: Guo, Litao, et al.
Published: (2025)
CalliMaster: Mastering Page-level Chinese Calligraphy via Layout-guided Spatial Planning
by: Xu, Tianshuo, et al.
Published: (2026)
by: Xu, Tianshuo, et al.
Published: (2026)
TransPixeler: Advancing Text-to-Video Generation with Transparency
by: Wang, Luozhou, et al.
Published: (2025)
by: Wang, Luozhou, et al.
Published: (2025)
LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding
by: Zhao, Haoyu, et al.
Published: (2024)
by: Zhao, Haoyu, et al.
Published: (2024)
RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes
by: Wu, Leyi, et al.
Published: (2026)
by: Wu, Leyi, et al.
Published: (2026)
PreGenie: An Agentic Framework for High-quality Visual Presentation Generation
by: Xu, Xiaojie, et al.
Published: (2025)
by: Xu, Xiaojie, et al.
Published: (2025)
From Bird's-Eye to Street View: Crafting Diverse and Condition-Aligned Images with Latent Diffusion Model
by: Xu, Xiaojie, et al.
Published: (2024)
by: Xu, Xiaojie, et al.
Published: (2024)
LucidFusion: Reconstructing 3D Gaussians with Arbitrary Unposed Images
by: He, Hao, et al.
Published: (2024)
by: He, Hao, et al.
Published: (2024)
PresentCoach: Dual-Agent Presentation Coaching through Exemplars and Interactive Feedback
by: Chen, Sirui, et al.
Published: (2025)
by: Chen, Sirui, et al.
Published: (2025)
Long-Video Audio Synthesis with Multi-Agent Collaboration
by: Zhang, Yehang, et al.
Published: (2025)
by: Zhang, Yehang, et al.
Published: (2025)
SCC-YOLO: An Improved Object Detector for Assisting in Brain Tumor Diagnosis
by: Bai, Runci, et al.
Published: (2025)
by: Bai, Runci, et al.
Published: (2025)
DisEnvisioner: Disentangled and Enriched Visual Prompt for Customized Image Generation
by: He, Jing, et al.
Published: (2024)
by: He, Jing, et al.
Published: (2024)
AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
by: Chen, Yutian, et al.
Published: (2026)
by: Chen, Yutian, et al.
Published: (2026)
X-Ray: A Sequential 3D Representation For Generation
by: Hu, Tao, et al.
Published: (2024)
by: Hu, Tao, et al.
Published: (2024)
On the Codegree graphs of finite groups
by: Chen, Jiyong, et al.
Published: (2025)
by: Chen, Jiyong, et al.
Published: (2025)
RHAML: Rendezvous-based Hierarchical Architecture for Mutual Localization
by: Chen, Gaoming, et al.
Published: (2024)
by: Chen, Gaoming, et al.
Published: (2024)
The Right Inference Strategy Is All You Need: Nearly Training-Free Domain-Wise Inference for EgoCross Challenge
by: Wu, Leyi, et al.
Published: (2026)
by: Wu, Leyi, et al.
Published: (2026)
Hawk: Learning to Understand Open-World Video Anomalies
by: Tang, Jiaqi, et al.
Published: (2024)
by: Tang, Jiaqi, et al.
Published: (2024)
Denoising Diffusion Step-aware Models
by: Yang, Shuai, et al.
Published: (2023)
by: Yang, Shuai, et al.
Published: (2023)
Managing Uncertainty in LLM-based Multi-Agent System Operation
by: Zhang, Man, et al.
Published: (2026)
by: Zhang, Man, et al.
Published: (2026)
Training Prompt Matters: State-Adaptive Optimization for Robust Fine-Tuning
by: Shi, Wenhang, et al.
Published: (2026)
by: Shi, Wenhang, et al.
Published: (2026)
Investigating the Impact of Rationales for LLMs on Natural Language Understanding
by: Shi, Wenhang, et al.
Published: (2025)
by: Shi, Wenhang, et al.
Published: (2025)
Similar Items
-
Uni-Renderer: Unifying Rendering and Inverse Rendering Via Dual Stream Diffusion
by: Chen, Zhifei, et al.
Published: (2024) -
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
by: Ge, Wenhang, et al.
Published: (2026) -
FlexPainter: Flexible and Multi-View Consistent Texture Generation
by: Yan, Dongyu, et al.
Published: (2025) -
PRM: Photometric Stereo based Large Reconstruction Model
by: Ge, Wenhang, et al.
Published: (2024) -
STANCE: Motion Coherent Video Generation Via Sparse-to-Dense Anchored Encoding
by: Chen, Zhifei, et al.
Published: (2025)