HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ben-Ami, Dan, Serussi, Gabriele, Cohen, Kobi, Baskin, Chaim |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
von: Ben-Ami, Dan, et al.
Veröffentlicht: (2025)
von: Ben-Ami, Dan, et al.
Veröffentlicht: (2025)
PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval
von: Serussi, Gabriele, et al.
Veröffentlicht: (2026)
von: Serussi, Gabriele, et al.
Veröffentlicht: (2026)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
von: Yang, Xuyi, et al.
Veröffentlicht: (2025)
von: Yang, Xuyi, et al.
Veröffentlicht: (2025)
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2026)
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2026)
Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges
von: Kosman, Eitan, et al.
Veröffentlicht: (2026)
von: Kosman, Eitan, et al.
Veröffentlicht: (2026)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
von: Kwon, Minchan, et al.
Veröffentlicht: (2026)
von: Kwon, Minchan, et al.
Veröffentlicht: (2026)
Semantic Event Graphs for Long-Form Video Question Answering
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026)
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026)
Encoding and Controlling Global Semantics for Long-form Video Question Answering
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Thong Thanh, et al.
Veröffentlicht: (2024)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
von: Huang, De-An, et al.
Veröffentlicht: (2025)
von: Huang, De-An, et al.
Veröffentlicht: (2025)
Robot Instance Segmentation with Few Annotations for Grasping
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2024)
Foundation Models and Adaptive Feature Selection: A Synergistic Approach to Video Question Answering
von: Rongali, Sai Bhargav, et al.
Veröffentlicht: (2024)
von: Rongali, Sai Bhargav, et al.
Veröffentlicht: (2024)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
von: Oh, Ju-Young, et al.
Veröffentlicht: (2025)
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering
von: Jung, Woojun, et al.
Veröffentlicht: (2026)
von: Jung, Woojun, et al.
Veröffentlicht: (2026)
WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
Question-Answering Dense Video Events
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
Multi-Frame, Lightweight & Efficient Vision-Language Models for Question Answering in Autonomous Driving
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
von: Gopalkrishnan, Akshay, et al.
Veröffentlicht: (2024)
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
Adaptive Greedy Frame Selection for Long Video Understanding
von: Huang, Yuning, et al.
Veröffentlicht: (2026)
von: Huang, Yuning, et al.
Veröffentlicht: (2026)
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
von: Zhang, Junkai, et al.
Veröffentlicht: (2025)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
von: Shen, Ruoyue, et al.
Veröffentlicht: (2024)
von: Shen, Ruoyue, et al.
Veröffentlicht: (2024)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
von: Liu, Huabin, et al.
Veröffentlicht: (2025)
CARES: Context-Aware Resolution Selector for VLMs
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
von: Kimhi, Moshe, et al.
Veröffentlicht: (2025)
Explicit Abstention Knobs for Predictable Reliability in Video Question Answering
von: Ortiz, Jorge
Veröffentlicht: (2025)
von: Ortiz, Jorge
Veröffentlicht: (2025)
Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment
von: Chen, Jin, et al.
Veröffentlicht: (2024)
von: Chen, Jin, et al.
Veröffentlicht: (2024)
CogStream: Context-guided Streaming Video Question Answering
von: Zhao, Zicheng, et al.
Veröffentlicht: (2025)
von: Zhao, Zicheng, et al.
Veröffentlicht: (2025)
VQA$^2$: Visual Question Answering for Video Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
von: Wieczorek, Tobias Jan, et al.
Veröffentlicht: (2025)
von: Wieczorek, Tobias Jan, et al.
Veröffentlicht: (2025)
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering
von: Fu, Xingyu, et al.
Veröffentlicht: (2023)
von: Fu, Xingyu, et al.
Veröffentlicht: (2023)
Hierarchical Question-Answering for Driving Scene Understanding Using Vision-Language Models
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2025)
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2025)
PipeFlow: Pipelined Processing and Motion-Aware Frame Selection for Long-Form Video Editing
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
von: Munir, Mustafa, et al.
Veröffentlicht: (2025)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
von: Patel, Alkesh, et al.
Veröffentlicht: (2025)
von: Patel, Alkesh, et al.
Veröffentlicht: (2025)
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
von: Choi, Changin, et al.
Veröffentlicht: (2025)
von: Choi, Changin, et al.
Veröffentlicht: (2025)
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
von: Wang, Haibo, et al.
Veröffentlicht: (2024)
VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2026)
von: Bhope, Rahul Atul, et al.
Veröffentlicht: (2026)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
von: Zhou, Ziqin, et al.
Veröffentlicht: (2025)
M-LLM Based Video Frame Selection for Efficient Video Understanding
von: Hu, Kai, et al.
Veröffentlicht: (2025)
von: Hu, Kai, et al.
Veröffentlicht: (2025)
InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
von: Xie, Tianchi, et al.
Veröffentlicht: (2025)
von: Xie, Tianchi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
von: Ben-Ami, Dan, et al.
Veröffentlicht: (2025) -
PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval
von: Serussi, Gabriele, et al.
Veröffentlicht: (2026) -
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
von: Yang, Xuyi, et al.
Veröffentlicht: (2025) -
MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2026) -
CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering
von: Bhosale, Mahesh, et al.
Veröffentlicht: (2026)