Generative Frame Sampler for Long Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yao, Linli, Wu, Haoning, Ouyang, Kun, Zhang, Yuanxing, Xiong, Caiming, Chen, Bei, Sun, Xu, Li, Junnan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Edit As You Wish: Video Caption Editing with Multi-grained User Control
von: Yao, Linli, et al.
Veröffentlicht: (2023)
von: Yao, Linli, et al.
Veröffentlicht: (2023)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
von: Luo, Ziyang, et al.
Veröffentlicht: (2024)
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
von: Mei, Yuting, et al.
Veröffentlicht: (2024)
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
von: Wu, Haoning, et al.
Veröffentlicht: (2024)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
von: Zhang, Deyu, et al.
Veröffentlicht: (2025)
von: Zhang, Deyu, et al.
Veröffentlicht: (2025)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
von: Fang, Xinyu, et al.
Veröffentlicht: (2024)
Towards Event-oriented Long Video Understanding
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Video-EM: Event-Centric Episodic Memory for Long-Form Video Understanding
von: Wang, Yun, et al.
Veröffentlicht: (2025)
von: Wang, Yun, et al.
Veröffentlicht: (2025)
G-Refine: A General Quality Refiner for Text-to-Image Generation
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
von: Li, Chunyi, et al.
Veröffentlicht: (2024)
AIS 2024 Challenge on Video Quality Assessment of User-Generated Content: Methods and Results
von: Conde, Marcos V., et al.
Veröffentlicht: (2024)
von: Conde, Marcos V., et al.
Veröffentlicht: (2024)
Region-Constraint In-Context Generation for Instructional Video Editing
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
Q-Adapt: Adapting LMM for Visual Quality Assessment with Progressive Instruction Tuning
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
von: Lu, Yiting, et al.
Veröffentlicht: (2025)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
von: Zhang, Zhongwei, et al.
Veröffentlicht: (2025)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
von: Pu, Junfu, et al.
Veröffentlicht: (2026)
DMC$^3$: Dual-Modal Counterfactual Contrastive Construction for Egocentric Video Question Answering
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
Representing Long Volumetric Video with Temporal Gaussian Hierarchy
von: Xu, Zhen, et al.
Veröffentlicht: (2024)
von: Xu, Zhen, et al.
Veröffentlicht: (2024)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
von: Qu, Mengxue, et al.
Veröffentlicht: (2024)
An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval
von: Kandhare, Mahesh, et al.
Veröffentlicht: (2024)
von: Kandhare, Mahesh, et al.
Veröffentlicht: (2024)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
von: Wang, Xiao, et al.
Veröffentlicht: (2024)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
von: Gao, Shibo, et al.
Veröffentlicht: (2025)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
von: Liu, Lingyu, et al.
Veröffentlicht: (2026)
MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation
von: Shi, Haoyuan, et al.
Veröffentlicht: (2026)
von: Shi, Haoyuan, et al.
Veröffentlicht: (2026)
Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
Dual-Branch Network for Portrait Image Quality Assessment
von: Sun, Wei, et al.
Veröffentlicht: (2024)
von: Sun, Wei, et al.
Veröffentlicht: (2024)
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
von: Kong, Fanheng, et al.
Veröffentlicht: (2025)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
von: Zhang, Chen-Lin, et al.
Veröffentlicht: (2025)
von: Zhang, Chen-Lin, et al.
Veröffentlicht: (2025)
Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
von: Yang, Haibo, et al.
Veröffentlicht: (2024)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
von: Zhang, Long, et al.
Veröffentlicht: (2025)
von: Zhang, Long, et al.
Veröffentlicht: (2025)
Q-Bench: A Benchmark for General-Purpose Foundation Models on Low-level Vision
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
von: Wu, Haoning, et al.
Veröffentlicht: (2023)
Human Motion Video Generation: A Survey
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
VideoMem: Constructing, Analyzing, Predicting Short-term and Long-term Video Memorability
von: Cohendet, Romain, et al.
Veröffentlicht: (2018)
von: Cohendet, Romain, et al.
Veröffentlicht: (2018)
Ähnliche Einträge
-
Edit As You Wish: Video Caption Editing with Multi-grained User Control
von: Yao, Linli, et al.
Veröffentlicht: (2023) -
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026) -
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
von: Luo, Ziyang, et al.
Veröffentlicht: (2024) -
UBiSS: A Unified Framework for Bimodal Semantic Summarization of Videos
von: Mei, Yuting, et al.
Veröffentlicht: (2024) -
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2023)