From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Lin, Shih-Yao, Paul, Sibendu, Chen, Caren |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Automatic Funny Scene Extraction from Long-form Cinematic Videos
by: Paul, Sibendu, et al.
Published: (2026)
by: Paul, Sibendu, et al.
Published: (2026)
Subtle Motion Blur Detection and Segmentation from Static Image Artworks
by: Samarth, Ganesh, et al.
Published: (2026)
by: Samarth, Ganesh, et al.
Published: (2026)
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
by: Wang, Eileen, et al.
Published: (2024)
by: Wang, Eileen, et al.
Published: (2024)
DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
by: Song, Zhende, et al.
Published: (2024)
by: Song, Zhende, et al.
Published: (2024)
STEC: A Reference-Free Spatio-Temporal Entropy Coverage Metric for Evaluating Sampled Video Frames
by: Lin, Shih-Yao
Published: (2026)
by: Lin, Shih-Yao
Published: (2026)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
by: Wang, Shaoguang, et al.
Published: (2026)
by: Wang, Shaoguang, et al.
Published: (2026)
KFFocus: Highlighting Keyframes for Enhanced Video Understanding
by: Nie, Ming, et al.
Published: (2025)
by: Nie, Ming, et al.
Published: (2025)
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
by: Song, Baiyang, et al.
Published: (2026)
by: Song, Baiyang, et al.
Published: (2026)
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
by: Kan, Shichao, et al.
Published: (2026)
by: Kan, Shichao, et al.
Published: (2026)
Progress-Aware Video Frame Captioning
by: Xue, Zihui, et al.
Published: (2024)
by: Xue, Zihui, et al.
Published: (2024)
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
by: Sun, Guangyu, et al.
Published: (2025)
by: Sun, Guangyu, et al.
Published: (2025)
KeyVID: Keyframe-Aware Video Diffusion for Audio-Synchronized Visual Animation
by: Wang, Xingrui, et al.
Published: (2025)
by: Wang, Xingrui, et al.
Published: (2025)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025)
by: Chen, Xiaofu, et al.
Published: (2025)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
by: Liang, Jianxin, et al.
Published: (2024)
by: Liang, Jianxin, et al.
Published: (2024)
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
by: Chen, Lin, et al.
Published: (2024)
by: Chen, Lin, et al.
Published: (2024)
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
by: Wang, Yuan, et al.
Published: (2026)
by: Wang, Yuan, et al.
Published: (2026)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
by: Ghazanfari, Sara, et al.
Published: (2025)
by: Ghazanfari, Sara, et al.
Published: (2025)
LongCaptioning: Unlocking the Power of Long Video Caption Generation in Large Multimodal Models
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
KeyframeFace: Language-Driven Facial Animation via Semantic Keyframes
by: Wu, Jingchao, et al.
Published: (2025)
by: Wu, Jingchao, et al.
Published: (2025)
Logic-in-Frames: Dynamic Keyframe Search via Visual Semantic-Logical Verification for Long Video Understanding
by: Guo, Weiyu, et al.
Published: (2025)
by: Guo, Weiyu, et al.
Published: (2025)
Adaptive Keyframe Sampling for Long Video Understanding
by: Tang, Xi, et al.
Published: (2025)
by: Tang, Xi, et al.
Published: (2025)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
by: Chou, Shih-Han, et al.
Published: (2023)
by: Chou, Shih-Han, et al.
Published: (2023)
VideoScore2: Think before You Score in Generative Video Evaluation
by: He, Xuan, et al.
Published: (2025)
by: He, Xuan, et al.
Published: (2025)
LFS: Learnable Frame Selector for Event-Aware and Temporally Diverse Video Captioning
by: Chao, Lianying, et al.
Published: (2026)
by: Chao, Lianying, et al.
Published: (2026)
Generative Frame Sampler for Long Video Understanding
by: Yao, Linli, et al.
Published: (2025)
by: Yao, Linli, et al.
Published: (2025)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
by: Zheng, Guangcong, et al.
Published: (2025)
by: Zheng, Guangcong, et al.
Published: (2025)
Agentic Keyframe Search for Video Question Answering
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
Generalized Trajectory Scoring for End-to-end Multimodal Planning
by: Li, Zhenxin, et al.
Published: (2025)
by: Li, Zhenxin, et al.
Published: (2025)
Large Model based Sequential Keyframe Extraction for Video Summarization
by: Tan, Kailong, et al.
Published: (2024)
by: Tan, Kailong, et al.
Published: (2024)
FOCUS: Efficient Keyframe Selection for Long Video Understanding
by: Zhu, Zirui, et al.
Published: (2025)
by: Zhu, Zirui, et al.
Published: (2025)
Occlusion-Aware Physics-Semantic Keyframe Selection for Robust Video Editing
by: Liu, Lin, et al.
Published: (2026)
by: Liu, Lin, et al.
Published: (2026)
Analytic Score Optimization for Multi Dimension Video Quality Assessment
by: Lin, Boda, et al.
Published: (2026)
by: Lin, Boda, et al.
Published: (2026)
Narrating the Video: Boosting Text-Video Retrieval via Comprehensive Utilization of Frame-Level Captions
by: Hur, Chan, et al.
Published: (2025)
by: Hur, Chan, et al.
Published: (2025)
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
by: Bao, Xiaoyi, et al.
Published: (2025)
by: Bao, Xiaoyi, et al.
Published: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
by: Chen, Wang, et al.
Published: (2026)
by: Chen, Wang, et al.
Published: (2026)
Shot2Tactic-Caption: Multi-Scale Captioning of Badminton Videos for Tactical Understanding
by: Ding, Ning, et al.
Published: (2025)
by: Ding, Ning, et al.
Published: (2025)
Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding
by: Wang, Yiheng, et al.
Published: (2026)
by: Wang, Yiheng, et al.
Published: (2026)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
by: Rahman, Aimon, et al.
Published: (2024)
by: Rahman, Aimon, et al.
Published: (2024)
Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring
by: Xie, Qizhi, et al.
Published: (2025)
by: Xie, Qizhi, et al.
Published: (2025)
Similar Items
-
Automatic Funny Scene Extraction from Long-form Cinematic Videos
by: Paul, Sibendu, et al.
Published: (2026) -
Subtle Motion Blur Detection and Segmentation from Static Image Artworks
by: Samarth, Ganesh, et al.
Published: (2026) -
GEM-VPC: A dual Graph-Enhanced Multimodal integration for Video Paragraph Captioning
by: Wang, Eileen, et al.
Published: (2024) -
DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
by: Song, Zhende, et al.
Published: (2024) -
STEC: A Reference-Free Spatio-Temporal Entropy Coverage Metric for Evaluating Sampled Video Frames
by: Lin, Shih-Yao
Published: (2026)