VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Qunzhong, Liu, Jie, Liang, Jiajun, Jiang, Yilei, Zhang, Yuanxing, Zheng, Yaozhi, Wang, Xintao, Wan, Pengfei, Yue, Xiangyu, Liu, Jiaheng |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents
par: Jiang, Yilei, et autres
Publié: (2025)
par: Jiang, Yilei, et autres
Publié: (2025)
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
par: Wang, Yuan, et autres
Publié: (2026)
par: Wang, Yuan, et autres
Publié: (2026)
OneThinker: All-in-one Reasoning Model for Image and Video
par: Feng, Kaituo, et autres
Publié: (2025)
par: Feng, Kaituo, et autres
Publié: (2025)
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
par: Wang, Chaoyang, et autres
Publié: (2026)
par: Wang, Chaoyang, et autres
Publié: (2026)
Flow-GRPO: Training Flow Matching Models via Online RL
par: Liu, Jie, et autres
Publié: (2025)
par: Liu, Jie, et autres
Publié: (2025)
Scaling Image and Video Generation via Test-Time Evolutionary Search
par: He, Haoran, et autres
Publié: (2025)
par: He, Haoran, et autres
Publié: (2025)
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
par: Wu, Shengqiong, et autres
Publié: (2025)
par: Wu, Shengqiong, et autres
Publié: (2025)
GARDO: Reinforcing Diffusion Models without Reward Hacking
par: He, Haoran, et autres
Publié: (2025)
par: He, Haoran, et autres
Publié: (2025)
V-Thinker: Interactive Thinking with Images
par: Qiao, Runqi, et autres
Publié: (2025)
par: Qiao, Runqi, et autres
Publié: (2025)
ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
par: Ding, Shengyuan, et autres
Publié: (2025)
par: Ding, Shengyuan, et autres
Publié: (2025)
SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking
par: Liu, Junnan, et autres
Publié: (2025)
par: Liu, Junnan, et autres
Publié: (2025)
Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
par: Wang, Shijian, et autres
Publié: (2025)
par: Wang, Shijian, et autres
Publié: (2025)
GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking
par: Cheng, Zixu, et autres
Publié: (2026)
par: Cheng, Zixu, et autres
Publié: (2026)
QuadSentinel: Sequent Safety for Machine-Checkable Control in Multi-agent Systems
par: Yang, Yiliu, et autres
Publié: (2025)
par: Yang, Yiliu, et autres
Publié: (2025)
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
par: Cai, Minghong, et autres
Publié: (2025)
par: Cai, Minghong, et autres
Publié: (2025)
Thinker: Learning to Think Fast and Slow
par: Chung, Stephen, et autres
Publié: (2025)
par: Chung, Stephen, et autres
Publié: (2025)
Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning
par: Li, Xueheng, et autres
Publié: (2026)
par: Li, Xueheng, et autres
Publié: (2026)
TypedThinker: Diversify Large Language Model Reasoning with Typed Thinking
par: Wang, Danqing, et autres
Publié: (2024)
par: Wang, Danqing, et autres
Publié: (2024)
SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
par: Fan, Kaixuan, et autres
Publié: (2025)
par: Fan, Kaixuan, et autres
Publié: (2025)
In-Context Audio Control of Video Diffusion Transformers
par: Liu, Wenze, et autres
Publié: (2025)
par: Liu, Wenze, et autres
Publié: (2025)
Monet: Reasoning in Latent Visual Space Beyond Images and Language
par: Wang, Qixun, et autres
Publié: (2025)
par: Wang, Qixun, et autres
Publié: (2025)
MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
par: Peng, Tianhao, et autres
Publié: (2025)
par: Peng, Tianhao, et autres
Publié: (2025)
KAG-Thinker: Interactive Thinking and Deep Reasoning in LLMs via Knowledge-Augmented Generation
par: Zhang, Dalong, et autres
Publié: (2025)
par: Zhang, Dalong, et autres
Publié: (2025)
Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
par: Ye, Zixuan, et autres
Publié: (2025)
par: Ye, Zixuan, et autres
Publié: (2025)
PreferThinker: Reasoning-based Personalized Image Preference Assessment
par: Xu, Shengqi, et autres
Publié: (2025)
par: Xu, Shengqi, et autres
Publié: (2025)
SemanticGen: Video Generation in Semantic Space
par: Bai, Jianhong, et autres
Publié: (2025)
par: Bai, Jianhong, et autres
Publié: (2025)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
par: Hu, Pengfei, et autres
Publié: (2025)
par: Hu, Pengfei, et autres
Publié: (2025)
EditThinker: Unlocking Iterative Reasoning for Any Image Editor
par: Li, Hongyu, et autres
Publié: (2025)
par: Li, Hongyu, et autres
Publié: (2025)
GameFactory: Creating New Games with Generative Interactive Videos
par: Yu, Jiwen, et autres
Publié: (2025)
par: Yu, Jiwen, et autres
Publié: (2025)
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
par: Wang, Zengzhi, et autres
Publié: (2025)
par: Wang, Zengzhi, et autres
Publié: (2025)
LightThinker++: From Reasoning Compression to Memory Management
par: Zhu, Yuqi, et autres
Publié: (2026)
par: Zhu, Yuqi, et autres
Publié: (2026)
Exploring Reasoning Reward Model for Agents
par: Fan, Kaixuan, et autres
Publié: (2026)
par: Fan, Kaixuan, et autres
Publié: (2026)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
par: Li, Chenglin, et autres
Publié: (2026)
par: Li, Chenglin, et autres
Publié: (2026)
Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning
par: Gan, Siyuan, et autres
Publié: (2026)
par: Gan, Siyuan, et autres
Publié: (2026)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
par: He, Zefeng, et autres
Publié: (2025)
par: He, Zefeng, et autres
Publié: (2025)
Improving Video Generation with Human Feedback
par: Liu, Jie, et autres
Publié: (2025)
par: Liu, Jie, et autres
Publié: (2025)
GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
par: Wang, Jing, et autres
Publié: (2025)
par: Wang, Jing, et autres
Publié: (2025)
Vero: An Open RL Recipe for General Visual Reasoning
par: Sarch, Gabriel, et autres
Publié: (2026)
par: Sarch, Gabriel, et autres
Publié: (2026)
UniVideo: Unified Understanding, Generation, and Editing for Videos
par: Wei, Cong, et autres
Publié: (2025)
par: Wei, Cong, et autres
Publié: (2025)
LightThinker: Thinking Step-by-Step Compression
par: Zhang, Jintian, et autres
Publié: (2025)
par: Zhang, Jintian, et autres
Publié: (2025)
Documents similaires
-
ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents
par: Jiang, Yilei, et autres
Publié: (2025) -
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
par: Wang, Yuan, et autres
Publié: (2026) -
OneThinker: All-in-one Reasoning Model for Image and Video
par: Feng, Kaituo, et autres
Publié: (2025) -
VLA-Thinker: Boosting Vision-Language-Action Models through Thinking-with-Image Reasoning
par: Wang, Chaoyang, et autres
Publié: (2026) -
Flow-GRPO: Training Flow Matching Models via Online RL
par: Liu, Jie, et autres
Publié: (2025)