Saved in:
| Main Authors: | Ni, Ziqi, Liang, Yuanzhi, Li, Rui, Zhou, Yi, Huang, Haibin, Zhang, Chi, Li, Xuelong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.18719 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
by: Liu, Ruiying, et al.
Published: (2025)
by: Liu, Ruiying, et al.
Published: (2025)
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Reward-Aware Trajectory Shaping for Few-step Visual Generation
by: Li, Rui, et al.
Published: (2026)
by: Li, Rui, et al.
Published: (2026)
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation
by: Li, Rui, et al.
Published: (2026)
by: Li, Rui, et al.
Published: (2026)
Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
by: Liang, Yuanzhi, et al.
Published: (2025)
by: Liang, Yuanzhi, et al.
Published: (2025)
VAST 1.0: A Unified Framework for Controllable and Consistent Video Generation
by: Zhang, Chi, et al.
Published: (2024)
by: Zhang, Chi, et al.
Published: (2024)
InterSyn: Interleaved Learning for Dynamic Motion Synthesis in the Wild
by: Ma, Yiyi, et al.
Published: (2025)
by: Ma, Yiyi, et al.
Published: (2025)
QwenStyle: Content-Preserving Style Transfer with Qwen-Image-Edit
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
CtrlVDiff: Controllable Video Generation via Unified Multimodal Video Diffusion
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
by: Liang, Yuanzhi, et al.
Published: (2026)
by: Liang, Yuanzhi, et al.
Published: (2026)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
What You See Is What Matters: A Novel Visual and Physics-Based Metric for Evaluating Video Generation Quality
by: Wang, Zihan, et al.
Published: (2024)
by: Wang, Zihan, et al.
Published: (2024)
Uni-Inter: Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
TeleStyle: Content-Preserving Style Transfer in Images and Videos
by: Zhang, Shiwen, et al.
Published: (2026)
by: Zhang, Shiwen, et al.
Published: (2026)
Do MLLMs Really See It: Reinforcing Visual Attention in Multimodal LLMs
by: Ou, Siqu, et al.
Published: (2026)
by: Ou, Siqu, et al.
Published: (2026)
ViPO: Visual Preference Optimization at Scale
by: Li, Ming, et al.
Published: (2026)
by: Li, Ming, et al.
Published: (2026)
TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
by: Ma, Yukuo, et al.
Published: (2025)
by: Ma, Yukuo, et al.
Published: (2025)
Visual Implicit Autoregressive Modeling
by: Jiang, Pengfei, et al.
Published: (2026)
by: Jiang, Pengfei, et al.
Published: (2026)
Seeing What Matters: Empowering CLIP with Patch Generation-to-Selection
by: Pei, Gensheng, et al.
Published: (2025)
by: Pei, Gensheng, et al.
Published: (2025)
SpongeBob: Sync-Aware Harmonious Audio-Visual Generative Editing
by: Liang, Sen, et al.
Published: (2026)
by: Liang, Sen, et al.
Published: (2026)
Towards General Multimodal Visual Tracking
by: Lu, Andong, et al.
Published: (2025)
by: Lu, Andong, et al.
Published: (2025)
TelePhysics: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
See the Text: From Tokenization to Visual Reading
by: Xing, Ling, et al.
Published: (2025)
by: Xing, Ling, et al.
Published: (2025)
VERTIGO: Visual Preference Optimization for Cinematic Camera Trajectory Generation
by: Li, Mengtian, et al.
Published: (2026)
by: Li, Mengtian, et al.
Published: (2026)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Visual Text Generation in the Wild
by: Zhu, Yuanzhi, et al.
Published: (2024)
by: Zhu, Yuanzhi, et al.
Published: (2024)
Point2Insert: Video Object Insertion via Sparse Point Guidance
by: Zhou, Yu, et al.
Published: (2026)
by: Zhou, Yu, et al.
Published: (2026)
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation
by: Liu, Jie, et al.
Published: (2026)
by: Liu, Jie, et al.
Published: (2026)
GS-SLAM: Dense Visual SLAM with 3D Gaussian Splatting
by: Yan, Chi, et al.
Published: (2023)
by: Yan, Chi, et al.
Published: (2023)
Visual Preference Optimization with Rubric Rewards
by: Yu, Ya-Qi, et al.
Published: (2026)
by: Yu, Ya-Qi, et al.
Published: (2026)
Focus On What Matters: Separated Models For Visual-Based RL Generalization
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
by: Dai, Ziyun, et al.
Published: (2025)
by: Dai, Ziyun, et al.
Published: (2025)
Conditional Video Generation for High-Efficiency Video Compression
by: Yi, Fangqiu, et al.
Published: (2025)
by: Yi, Fangqiu, et al.
Published: (2025)
FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis
by: Ni, Ziqi, et al.
Published: (2025)
by: Ni, Ziqi, et al.
Published: (2025)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Visual Spatial Tuning
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
by: Liu, Jialun, et al.
Published: (2026)
by: Liu, Jialun, et al.
Published: (2026)
Point What You Mean: Visually Grounded Instruction Policy
by: Yu, Hang, et al.
Published: (2025)
by: Yu, Hang, et al.
Published: (2025)
Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
by: Zhao, Kesen, et al.
Published: (2025)
by: Zhao, Kesen, et al.
Published: (2025)
Similar Items
-
Learning What to Trust: Bayesian Prior-Guided Optimization for Visual Generation
by: Liu, Ruiying, et al.
Published: (2025) -
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
by: Li, Rui, et al.
Published: (2025) -
Reward-Aware Trajectory Shaping for Few-step Visual Generation
by: Li, Rui, et al.
Published: (2026) -
UniModel: A Visual-Only Framework for Unified Multimodal Understanding and Generation
by: Zhang, Chi, et al.
Published: (2025) -
Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation
by: Li, Rui, et al.
Published: (2026)