HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Yiqing, Lam, Kin-Man |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
di: Liang, Hao, et al.
Pubblicazione: (2024)
di: Liang, Hao, et al.
Pubblicazione: (2024)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
di: Chow, Wei, et al.
Pubblicazione: (2025)
di: Chow, Wei, et al.
Pubblicazione: (2025)
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
di: Zhang, Dongxu, et al.
Pubblicazione: (2026)
di: Zhang, Dongxu, et al.
Pubblicazione: (2026)
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
di: Zhang, Deyu, et al.
Pubblicazione: (2025)
di: Zhang, Deyu, et al.
Pubblicazione: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
di: Zhang, Xueqiao, et al.
Pubblicazione: (2025)
di: Zhang, Xueqiao, et al.
Pubblicazione: (2025)
MuLTI: Efficient Video-and-Language Understanding with Text-Guided MultiWay-Sampler and Multiple Choice Modeling
di: Xu, Jiaqi, et al.
Pubblicazione: (2023)
di: Xu, Jiaqi, et al.
Pubblicazione: (2023)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
di: Wang, Shaoguang, et al.
Pubblicazione: (2026)
di: Wang, Shaoguang, et al.
Pubblicazione: (2026)
MMBench-Video: A Long-Form Multi-Shot Benchmark for Holistic Video Understanding
di: Fang, Xinyu, et al.
Pubblicazione: (2024)
di: Fang, Xinyu, et al.
Pubblicazione: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
di: Wang, Jiapeng, et al.
Pubblicazione: (2024)
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
di: Cheung, Tsun-Hin, et al.
Pubblicazione: (2024)
di: Cheung, Tsun-Hin, et al.
Pubblicazione: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
di: Ji, Yatai, et al.
Pubblicazione: (2024)
di: Ji, Yatai, et al.
Pubblicazione: (2024)
Video Summarization: Towards Entity-Aware Captions
di: Ayyubi, Hammad A., et al.
Pubblicazione: (2023)
di: Ayyubi, Hammad A., et al.
Pubblicazione: (2023)
Holistic Evaluation of Multimodal LLMs on Spatial Intelligence
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
Recipe Generation from Unsegmented Cooking Videos
di: Nishimura, Taichi, et al.
Pubblicazione: (2022)
di: Nishimura, Taichi, et al.
Pubblicazione: (2022)
Towards Event-oriented Long Video Understanding
di: Du, Yifan, et al.
Pubblicazione: (2024)
di: Du, Yifan, et al.
Pubblicazione: (2024)
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
di: Biswas, Subrata, et al.
Pubblicazione: (2025)
di: Biswas, Subrata, et al.
Pubblicazione: (2025)
Generative Frame Sampler for Long Video Understanding
di: Yao, Linli, et al.
Pubblicazione: (2025)
di: Yao, Linli, et al.
Pubblicazione: (2025)
Beyond Coarse-Grained Matching in Video-Text Retrieval
di: Chen, Aozhu, et al.
Pubblicazione: (2024)
di: Chen, Aozhu, et al.
Pubblicazione: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
di: Doughty, Hazel, et al.
Pubblicazione: (2024)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
di: Chen, Yijing, et al.
Pubblicazione: (2025)
di: Chen, Yijing, et al.
Pubblicazione: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
di: Zhao, Zhixian, et al.
Pubblicazione: (2026)
di: Zhao, Zhixian, et al.
Pubblicazione: (2026)
ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
di: Wang, Xiao, et al.
Pubblicazione: (2024)
di: Wang, Xiao, et al.
Pubblicazione: (2024)
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
di: Sukhani, Siddhant, et al.
Pubblicazione: (2025)
di: Sukhani, Siddhant, et al.
Pubblicazione: (2025)
Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
di: Stamatakis, Markos, et al.
Pubblicazione: (2025)
di: Stamatakis, Markos, et al.
Pubblicazione: (2025)
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
di: Nishimura, Taichi, et al.
Pubblicazione: (2024)
di: Nishimura, Taichi, et al.
Pubblicazione: (2024)
COM Kitchens: An Unedited Overhead-view Video Dataset as a Vision-Language Benchmark
di: Maeda, Koki, et al.
Pubblicazione: (2024)
di: Maeda, Koki, et al.
Pubblicazione: (2024)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
di: Wang, Xiao, et al.
Pubblicazione: (2025)
di: Wang, Xiao, et al.
Pubblicazione: (2025)
EasyAnimate: High-Performance Video Generation Framework with Hybrid Windows Attention and Reward Backpropagation
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation
di: Li, Yunxin, et al.
Pubblicazione: (2024)
di: Li, Yunxin, et al.
Pubblicazione: (2024)
Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
di: Duan, Changxu
Pubblicazione: (2025)
di: Duan, Changxu
Pubblicazione: (2025)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
di: Han, Wei, et al.
Pubblicazione: (2023)
di: Han, Wei, et al.
Pubblicazione: (2023)
An Empirical Comparison of Video Frame Sampling Methods for Multi-Modal RAG Retrieval
di: Kandhare, Mahesh, et al.
Pubblicazione: (2024)
di: Kandhare, Mahesh, et al.
Pubblicazione: (2024)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
di: Liang, Zhengyang, et al.
Pubblicazione: (2024)
di: Liang, Zhengyang, et al.
Pubblicazione: (2024)
LongVALE: Vision-Audio-Language-Event Benchmark Towards Time-Aware Omni-Modal Perception of Long Videos
di: Geng, Tiantian, et al.
Pubblicazione: (2024)
di: Geng, Tiantian, et al.
Pubblicazione: (2024)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
di: Chen, Junyi, et al.
Pubblicazione: (2023)
di: Chen, Junyi, et al.
Pubblicazione: (2023)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
di: Pian, Weiguo, et al.
Pubblicazione: (2026)
di: Pian, Weiguo, et al.
Pubblicazione: (2026)
Can Video Diffusion Models Predict Past Frames? Bidirectional Cycle Consistency for Reversible Interpolation
di: Liu, Lingyu, et al.
Pubblicazione: (2026)
di: Liu, Lingyu, et al.
Pubblicazione: (2026)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
GalleryGPT: Analyzing Paintings with Large Multimodal Models
di: Bin, Yi, et al.
Pubblicazione: (2024)
di: Bin, Yi, et al.
Pubblicazione: (2024)
Holistic Visual-Textual Sentiment Analysis with Prior Models
di: Chen, Junyu, et al.
Pubblicazione: (2022)
di: Chen, Junyu, et al.
Pubblicazione: (2022)
Documenti analoghi
-
KeyVideoLLM: Towards Large-scale Video Keyframe Selection
di: Liang, Hao, et al.
Pubblicazione: (2024) -
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
di: Chow, Wei, et al.
Pubblicazione: (2025) -
Chain-of-Thought Compression Should Not Be Blind: V-Skip for Efficient Multimodal Reasoning via Dual-Path Anchoring
di: Zhang, Dongxu, et al.
Pubblicazione: (2026) -
Prompt-aware of Frame Sampling for Efficient Text-Video Retrieval
di: Zhang, Deyu, et al.
Pubblicazione: (2025) -
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
di: Zhang, Xueqiao, et al.
Pubblicazione: (2025)