Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongyu, Han, Songhao, Liao, Yue, Luo, Junfeng, Gao, Jialin, Yan, Shuicheng, Liu, Si |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
by: Han, Songhao, et al.
Published: (2025)
by: Han, Songhao, et al.
Published: (2025)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
by: Kim, Minji, et al.
Published: (2025)
by: Kim, Minji, et al.
Published: (2025)
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026)
by: Guan, Yiran, et al.
Published: (2026)
LLMs as Visual Explainers: Advancing Image Classification with Evolving Visual Descriptions
by: Han, Songhao, et al.
Published: (2023)
by: Han, Songhao, et al.
Published: (2023)
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)
by: Qin, Jialong, et al.
Published: (2025)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
by: Li, Chenglin, et al.
Published: (2026)
by: Li, Chenglin, et al.
Published: (2026)
Lost in Time: A New Temporal Benchmark for VideoLLMs
by: Cores, Daniel, et al.
Published: (2024)
by: Cores, Daniel, et al.
Published: (2024)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
by: Chatterjee, Dibyadip, et al.
Published: (2025)
by: Chatterjee, Dibyadip, et al.
Published: (2025)
Geometry-Guided Camera Motion Understanding in VideoLLMs
by: Feng, Haoan, et al.
Published: (2026)
by: Feng, Haoan, et al.
Published: (2026)
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
by: Li, Jiameng, et al.
Published: (2026)
by: Li, Jiameng, et al.
Published: (2026)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
by: Cai, Jianfeng, et al.
Published: (2025)
by: Cai, Jianfeng, et al.
Published: (2025)
Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
by: Chung, Hyungjin, et al.
Published: (2025)
by: Chung, Hyungjin, et al.
Published: (2025)
VideoLLM-online: Online Video Large Language Model for Streaming Video
by: Chen, Joya, et al.
Published: (2024)
by: Chen, Joya, et al.
Published: (2024)
Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
by: Yan, Weicai, et al.
Published: (2026)
by: Yan, Weicai, et al.
Published: (2026)
WeaveTime: Stream from Earlier Frames into Emergent Memory in VideoLLMs
by: Zhang, Yulin, et al.
Published: (2026)
by: Zhang, Yulin, et al.
Published: (2026)
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
by: Wu, Shiwei, et al.
Published: (2024)
by: Wu, Shiwei, et al.
Published: (2024)
Dynamic-VLM: Simple Dynamic Visual Token Compression for VideoLLM
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
by: Shu, Fangxun, et al.
Published: (2025)
by: Shu, Fangxun, et al.
Published: (2025)
VideoLLM Benchmarks and Evaluation: A Survey
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing
by: Kumar, Yogesh
Published: (2025)
by: Kumar, Yogesh
Published: (2025)
LayoutCoT: Unleashing the Deep Reasoning Potential of Large Language Models for Layout Generation
by: Shi, Hengyu, et al.
Published: (2025)
by: Shi, Hengyu, et al.
Published: (2025)
Data Augmentation in Human-Centric Vision
by: Jiang, Wentao, et al.
Published: (2024)
by: Jiang, Wentao, et al.
Published: (2024)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
by: Fang, Pengcheng, et al.
Published: (2025)
by: Fang, Pengcheng, et al.
Published: (2025)
PosterOmni: Generalized Artistic Poster Creation via Task Distillation and Unified Reward Feedback
by: Chen, Sixiang, et al.
Published: (2026)
by: Chen, Sixiang, et al.
Published: (2026)
MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning
by: Tao, Sicheng, et al.
Published: (2025)
by: Tao, Sicheng, et al.
Published: (2025)
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
IPFormer-VideoLLM: Enhancing Multi-modal Video Understanding for Multi-shot Scenes
by: Liang, Yujia, et al.
Published: (2025)
by: Liang, Yujia, et al.
Published: (2025)
"I Can See Forever!": Evaluating Real-time VideoLLMs for Assisting Individuals with Visual Impairments
by: Zhang, Ziyi, et al.
Published: (2025)
by: Zhang, Ziyi, et al.
Published: (2025)
MAN++: Scaling Momentum Auxiliary Network for Supervised Local Learning in Vision Tasks
by: Su, Junhao, et al.
Published: (2025)
by: Su, Junhao, et al.
Published: (2025)
Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
VideoEspresso: A Large-Scale Chain-of-Thought Dataset for Fine-Grained Video Reasoning via Core Frame Selection
by: Han, Songhao, et al.
Published: (2024)
by: Han, Songhao, et al.
Published: (2024)
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
by: Chen, Jingkun, et al.
Published: (2026)
by: Chen, Jingkun, et al.
Published: (2026)
Reinforcement Learning with Robust Rubric Rewards
by: Yu, Ya-Qi, et al.
Published: (2026)
by: Yu, Ya-Qi, et al.
Published: (2026)
FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
by: Yin, Bo, et al.
Published: (2025)
by: Yin, Bo, et al.
Published: (2025)
Peregrine: One-Shot Fine-Tuning for FHE Inference of General Deep CNNs
by: Ling, Huaming, et al.
Published: (2025)
by: Ling, Huaming, et al.
Published: (2025)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Reinforcing Consistency in Video MLLMs with Structured Rewards
by: Quan, Yihao, et al.
Published: (2026)
by: Quan, Yihao, et al.
Published: (2026)
Learning Temporally Consistent Video Depth from Video Diffusion Priors
by: Shao, Jiahao, et al.
Published: (2024)
by: Shao, Jiahao, et al.
Published: (2024)
Parrot: Pareto-optimal Multi-Reward Reinforcement Learning Framework for Text-to-Image Generation
by: Lee, Seung Hyun, et al.
Published: (2024)
by: Lee, Seung Hyun, et al.
Published: (2024)
MVR: Multi-view Video Reward Shaping for Reinforcement Learning
by: Luo, Lirui, et al.
Published: (2026)
by: Luo, Lirui, et al.
Published: (2026)
Similar Items
-
RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
by: Han, Songhao, et al.
Published: (2025) -
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
by: Kim, Minji, et al.
Published: (2025) -
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
by: Guan, Yiran, et al.
Published: (2026) -
LLMs as Visual Explainers: Advancing Image Classification with Evolving Visual Descriptions
by: Han, Songhao, et al.
Published: (2023) -
Sharp Eyes and Memory for VideoLLMs: Information-Aware Visual Token Pruning for Efficient and Reliable VideoLLM Reasoning
by: Qin, Jialong, et al.
Published: (2025)