Saved in:
| Main Authors: | Li, Rui, Hao, Ke, Liang, Yuanzhi, Huang, Haibin, Zhang, Chi, Gu, Yun, Li, XueLong |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.19234 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Reward-Aware Trajectory Shaping for Few-step Visual Generation
by: Li, Rui, et al.
Published: (2026)
by: Li, Rui, et al.
Published: (2026)
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
by: Ni, Ziqi, et al.
Published: (2025)
by: Ni, Ziqi, et al.
Published: (2025)
Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
by: Liu, Ruiying, et al.
Published: (2025)
by: Liu, Ruiying, et al.
Published: (2025)
Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
by: Liang, Yuanzhi, et al.
Published: (2025)
by: Liang, Yuanzhi, et al.
Published: (2025)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
by: Liang, Yuanzhi, et al.
Published: (2026)
by: Liang, Yuanzhi, et al.
Published: (2026)
InterSyn: Interleaved Learning for Dynamic Motion Synthesis in the Wild
by: Ma, Yiyi, et al.
Published: (2025)
by: Ma, Yiyi, et al.
Published: (2025)
VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation
by: Zhang, Chi, et al.
Published: (2024)
by: Zhang, Chi, et al.
Published: (2024)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
Aesthetic Post-Training Diffusion Models from Generic Preferences with Step-by-step Preference Optimization
by: Liang, Zhanhao, et al.
Published: (2024)
by: Liang, Zhanhao, et al.
Published: (2024)
QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
Towards General Multimodal Visual Tracking
by: Lu, Andong, et al.
Published: (2025)
by: Lu, Andong, et al.
Published: (2025)
SlimFlow: Training Smaller One-Step Diffusion Models with Rectified Flow
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
LaxMotion: Rethinking Supervision Granularity for 3D Human Motion Generation
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
by: Zhu, Yuanzhi, et al.
Published: (2025)
by: Zhu, Yuanzhi, et al.
Published: (2025)
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
by: Liang, Sen, et al.
Published: (2026)
by: Liang, Sen, et al.
Published: (2026)
LossAgent: Towards Any Optimization Objectives for Image Processing with LLM Agents
by: Li, Bingchen, et al.
Published: (2024)
by: Li, Bingchen, et al.
Published: (2024)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
SpatialLock: Precise Spatial Control in Text-to-Image Synthesis
by: Liu, Biao, et al.
Published: (2025)
by: Liu, Biao, et al.
Published: (2025)
CAVE: A Structured Credit Assignment Approach for Fragmented Visual Evidence Reasoning
by: Guo, Tengda, et al.
Published: (2026)
by: Guo, Tengda, et al.
Published: (2026)
Beyond the Last Frame: Process-aware Evaluation for Generative Video Reasoning
by: Li, Yifan, et al.
Published: (2025)
by: Li, Yifan, et al.
Published: (2025)
Visual Text Generation in the Wild
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
OSV: One Step is Enough for High-Quality Image to Video Generation
by: Mao, Xiaofeng, et al.
Published: (2024)
by: Mao, Xiaofeng, et al.
Published: (2024)
OFTSR: One-Step Flow for Image Super-Resolution with Tunable Fidelity-Realism Trade-offs
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
One Step Learning, One Step Review
by: Huang, Xiaolong, et al.
Published: (2024)
by: Huang, Xiaolong, et al.
Published: (2024)
StepAL: Step-aware Active Learning for Cataract Surgical Videos
by: Shah, Nisarg A., et al.
Published: (2025)
by: Shah, Nisarg A., et al.
Published: (2025)
Optimizing Prompts for Text-to-Image Generation
by: Hao, Yaru, et al.
Published: (2022)
by: Hao, Yaru, et al.
Published: (2022)
Knowledge-aware Visual Question Generation for Remote Sensing Images
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
Revisiting Disentanglement in Downstream Tasks: A Study on Its Necessity for Abstract Visual Reasoning
by: Nai, Ruiqian, et al.
Published: (2024)
by: Nai, Ruiqian, et al.
Published: (2024)
Object Style Diffusion for Generalized Object Detection in Urban Scene
by: Li, Hao, et al.
Published: (2024)
by: Li, Hao, et al.
Published: (2024)
FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal Attention
by: Lu, Yu, et al.
Published: (2024)
by: Lu, Yu, et al.
Published: (2024)
Let Synthetic Data Shine: Domain Reassembly and Soft-Fusion for Single Domain Generalization
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation
by: Li, Chao, et al.
Published: (2026)
by: Li, Chao, et al.
Published: (2026)
Spatial-Temporal State Propagation Autoregressive Model for 4D Object Generation
by: Yang, Liying, et al.
Published: (2026)
by: Yang, Liying, et al.
Published: (2026)
Hierarchical Instruction-aware Embodied Visual Tracking
by: Wu, Kui, et al.
Published: (2025)
by: Wu, Kui, et al.
Published: (2025)
Segmenting Objectiveness and Task-awareness Unknown Region for Autonomous Driving
by: Zheng, Mi, et al.
Published: (2025)
by: Zheng, Mi, et al.
Published: (2025)
TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
by: Ma, Yukuo, et al.
Published: (2025)
by: Ma, Yukuo, et al.
Published: (2025)
Similar Items
-
Reward-Aware Trajectory Shaping for Few-step Visual Generation
by: Li, Rui, et al.
Published: (2026) -
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
by: Ni, Ziqi, et al.
Published: (2025) -
Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
by: Liu, Ruiying, et al.
Published: (2025) -
Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
by: Liang, Yuanzhi, et al.
Published: (2025) -
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)