VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Xinye, Guo, Hongcan, Qian, Jiawen, Nan, Guoshun, Wang, Chao, Pan, Yuqi, Hou, Tianhao, Wang, Xiaojuan, Gao, Yutong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Disentangling Deception and Hallucination Failures in LLMs
by: Lu, Haolang, et al.
Published: (2026)
by: Lu, Haolang, et al.
Published: (2026)
Advancing Compositional LLM Reasoning with Structured Task Relations in Interactive Multimodal Communications
by: Cao, Xinye, et al.
Published: (2025)
by: Cao, Xinye, et al.
Published: (2025)
Advancing LLM-Based Security Automation with Customized Group Relative Policy Optimization for Zero-Touch Networks
by: Cao, Xinye, et al.
Published: (2025)
by: Cao, Xinye, et al.
Published: (2025)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
by: Zheng, Guangcong, et al.
Published: (2025)
by: Zheng, Guangcong, et al.
Published: (2025)
Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly
by: Du, Hang, et al.
Published: (2024)
by: Du, Hang, et al.
Published: (2024)
NEWTON: Agentic Planning for Physically Grounded Video Generation
by: Feng, Yuxiang, et al.
Published: (2026)
by: Feng, Yuxiang, et al.
Published: (2026)
Event-Anchored Frame Selection for Effective Long-Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
Advancing Expert Specialization for Better MoE
by: Guo, Hongcan, et al.
Published: (2025)
by: Guo, Hongcan, et al.
Published: (2025)
Multi-sentence Video Grounding for Long Video Generation
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
by: Yang, Xuyi, et al.
Published: (2025)
by: Yang, Xuyi, et al.
Published: (2025)
KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
by: Li, Zongyao, et al.
Published: (2025)
by: Li, Zongyao, et al.
Published: (2025)
Think with Grounding: Curriculum Reinforced Reasoning with Video Grounding for Long Video Understanding
by: Chen, Houlun, et al.
Published: (2026)
by: Chen, Houlun, et al.
Published: (2026)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
by: Wang, Ziyang, et al.
Published: (2025)
by: Wang, Ziyang, et al.
Published: (2025)
Supervised Learning Model for Key Frame Identification from Cow Teat Videos
by: Wang, Minghao, et al.
Published: (2024)
by: Wang, Minghao, et al.
Published: (2024)
Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
by: Li, Jialuo, et al.
Published: (2025)
by: Li, Jialuo, et al.
Published: (2025)
Two Is Better Than One: Rotations Scale LoRAs
by: Guo, Hongcan, et al.
Published: (2025)
by: Guo, Hongcan, et al.
Published: (2025)
On the blow-up of harmonic maps from surfaces to homogeneous manifolds
by: Qian, Hongcan, et al.
Published: (2026)
by: Qian, Hongcan, et al.
Published: (2026)
Where, Not What: Compelling Video LLMs to Learn Geometric Causality for 3D-Grounding
by: Zhong, Yutong
Published: (2025)
by: Zhong, Yutong
Published: (2025)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
by: Wang, Chendong, et al.
Published: (2025)
by: Wang, Chendong, et al.
Published: (2025)
LDDR: Linear-DPP-Based Dynamic-Resolution Frame Sampling for Video MLLMs
by: Chen, Jingfeng, et al.
Published: (2026)
by: Chen, Jingfeng, et al.
Published: (2026)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
VerIPO: Cultivating Long Reasoning in Video-LLMs via Verifier-Gudied Iterative Policy Optimization
by: Li, Yunxin, et al.
Published: (2025)
by: Li, Yunxin, et al.
Published: (2025)
Exploring Iterative Refinement with Diffusion Models for Video Grounding
by: Liang, Xiao, et al.
Published: (2023)
by: Liang, Xiao, et al.
Published: (2023)
GRPOformer: Advancing Hyperparameter Optimization via Group Relative Policy Optimization
by: Guo, Haoxin, et al.
Published: (2025)
by: Guo, Haoxin, et al.
Published: (2025)
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
by: Sun, Guangyu, et al.
Published: (2025)
by: Sun, Guangyu, et al.
Published: (2025)
VideoSSM: Autoregressive Long Video Generation with Hybrid State-Space Memory
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
by: Bao, Xiaoyi, et al.
Published: (2025)
by: Bao, Xiaoyi, et al.
Published: (2025)
Multi-Sentence Grounding for Long-term Instructional Video
by: Li, Zeqian, et al.
Published: (2023)
by: Li, Zeqian, et al.
Published: (2023)
Strange Hadron Production at High Baryon Density
by: Li, Hongcan
Published: (2025)
by: Li, Hongcan
Published: (2025)
MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
by: Jia, Weinan, et al.
Published: (2025)
by: Jia, Weinan, et al.
Published: (2025)
Adaptive Greedy Frame Selection for Long Video Understanding
by: Huang, Yuning, et al.
Published: (2026)
by: Huang, Yuning, et al.
Published: (2026)
Hierarchical Indexing with Knowledge Enrichment for Multilingual Video Corpus Retrieval
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Re-Identifying Kākā with AI-Automated Video Key Frame Extraction
by: Maddigan, Paula, et al.
Published: (2025)
by: Maddigan, Paula, et al.
Published: (2025)
Unleashing Hour-Scale Video Training for Long Video-Language Understanding
by: Lin, Jingyang, et al.
Published: (2025)
by: Lin, Jingyang, et al.
Published: (2025)
A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
by: Zou, Yuanhao, et al.
Published: (2025)
by: Zou, Yuanhao, et al.
Published: (2025)
Relative Kinetic Utility for Reasoning-Aware Structural Pruning in Large Language Models
by: Qian, Tianhao
Published: (2026)
by: Qian, Tianhao
Published: (2026)
Velocity Disambiguation for Video Frame Interpolation
by: Zhong, Zhihang, et al.
Published: (2023)
by: Zhong, Zhihang, et al.
Published: (2023)
Multi-Focused Video Group Activities Hashing
by: Qi, Zhongmiao, et al.
Published: (2025)
by: Qi, Zhongmiao, et al.
Published: (2025)
A Rolling Stone Gathers No Moss: Adaptive Policy Optimization for Stable Self-Evaluation in Large Multimodal Models
by: Wang, Wenkai, et al.
Published: (2025)
by: Wang, Wenkai, et al.
Published: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
by: Chen, Qirui, et al.
Published: (2024)
by: Chen, Qirui, et al.
Published: (2024)
Similar Items
-
Disentangling Deception and Hallucination Failures in LLMs
by: Lu, Haolang, et al.
Published: (2026) -
Advancing Compositional LLM Reasoning with Structured Task Relations in Interactive Multimodal Communications
by: Cao, Xinye, et al.
Published: (2025) -
Advancing LLM-Based Security Automation with Customized Group Relative Policy Optimization for Zero-Touch Networks
by: Cao, Xinye, et al.
Published: (2025) -
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
by: Zheng, Guangcong, et al.
Published: (2025) -
Exploring What Why and How: A Multifaceted Benchmark for Causation Understanding of Video Anomaly
by: Du, Hang, et al.
Published: (2024)