Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Xuyi, Zhang, Wenhao, Jin, Hongbo, Liu, Lin, Xu, Hongbo, Nie, Yongwei, Yu, Fei, Ma, Fei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
von: Song, Enxin, et al.
Veröffentlicht: (2024)
von: Song, Enxin, et al.
Veröffentlicht: (2024)
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
von: Guo, Chaohong, et al.
Veröffentlicht: (2026)
FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
von: Zemskova, Tatiana, et al.
Veröffentlicht: (2026)
von: Zemskova, Tatiana, et al.
Veröffentlicht: (2026)
TAR-TVG: Enhancing VLMs with Timestamp Anchor-Constrained Reasoning for Temporal Video Grounding
von: Guo, Chaohong, et al.
Veröffentlicht: (2025)
von: Guo, Chaohong, et al.
Veröffentlicht: (2025)
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
von: Ben-Ami, Dan, et al.
Veröffentlicht: (2026)
von: Ben-Ami, Dan, et al.
Veröffentlicht: (2026)
Sketch3DVE: Sketch-based 3D-Aware Scene Video Editing
von: Liu, Feng-Lin, et al.
Veröffentlicht: (2025)
von: Liu, Feng-Lin, et al.
Veröffentlicht: (2025)
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
von: Xue, Junxiao, et al.
Veröffentlicht: (2024)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
3D Question Answering for City Scene Understanding
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
von: Sun, Penglei, et al.
Veröffentlicht: (2024)
Towards Fine-Grained Video Question Answering
von: Dai, Wei, et al.
Veröffentlicht: (2025)
von: Dai, Wei, et al.
Veröffentlicht: (2025)
Frame-Level Captions for Long Video Generation with Complex Multi Scenes
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
von: Zheng, Guangcong, et al.
Veröffentlicht: (2025)
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
von: Zou, Kai, et al.
Veröffentlicht: (2026)
von: Zou, Kai, et al.
Veröffentlicht: (2026)
FunduSAM: A Specialized Deep Learning Model for Enhanced Optic Disc and Cup Segmentation in Fundus Images
von: Yu, Jinchen, et al.
Veröffentlicht: (2025)
von: Yu, Jinchen, et al.
Veröffentlicht: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
von: Chen, Wang, et al.
Veröffentlicht: (2026)
von: Chen, Wang, et al.
Veröffentlicht: (2026)
A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
von: Zou, Yuanhao, et al.
Veröffentlicht: (2025)
von: Zou, Yuanhao, et al.
Veröffentlicht: (2025)
Narrative Aligned Long Form Video Question Answering
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
von: Jain, Rahul, et al.
Veröffentlicht: (2026)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames
von: Chen, Chao, et al.
Veröffentlicht: (2023)
von: Chen, Chao, et al.
Veröffentlicht: (2023)
ViLA: Efficient Video-Language Alignment for Video Question Answering
von: Wang, Xijun, et al.
Veröffentlicht: (2023)
von: Wang, Xijun, et al.
Veröffentlicht: (2023)
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
von: Guo, Diandian, et al.
Veröffentlicht: (2026)
von: Guo, Diandian, et al.
Veröffentlicht: (2026)
CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
von: Chen, Guo, et al.
Veröffentlicht: (2024)
von: Chen, Guo, et al.
Veröffentlicht: (2024)
A Simple LLM Framework for Long-Range Video Question-Answering
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
Bring the Power of Diffusion Model to Defect Detection
von: Yu, Xuyi
Veröffentlicht: (2024)
von: Yu, Xuyi
Veröffentlicht: (2024)
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
von: Luo, Jingzhou, et al.
Veröffentlicht: (2025)
von: Luo, Jingzhou, et al.
Veröffentlicht: (2025)
HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models
von: Bai, Xiangyu, et al.
Veröffentlicht: (2026)
von: Bai, Xiangyu, et al.
Veröffentlicht: (2026)
Ensemble Predicate Decoding for Unbiased Scene Graph Generation
von: Feng, Jiasong, et al.
Veröffentlicht: (2024)
von: Feng, Jiasong, et al.
Veröffentlicht: (2024)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
von: Jin, Hongbo, et al.
Veröffentlicht: (2026)
von: Jin, Hongbo, et al.
Veröffentlicht: (2026)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
Simplifying CLIP: Unleashing the Power of Large-Scale Models on Consumer-level Computers
von: Liu, Hongbo
Veröffentlicht: (2024)
von: Liu, Hongbo
Veröffentlicht: (2024)
Compositional 4D Dynamic Scenes Understanding with Physics Priors for Video Question Answering
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
von: Wang, Xingrui, et al.
Veröffentlicht: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
CL-VISTA: Benchmarking Continual Learning in Video Large Language Models
von: Guo, Haiyang, et al.
Veröffentlicht: (2026)
von: Guo, Haiyang, et al.
Veröffentlicht: (2026)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
Agentic Keyframe Search for Video Question Answering
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
von: Dong, Xinxin, et al.
Veröffentlicht: (2025)
von: Dong, Xinxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
von: Jin, Hongbo, et al.
Veröffentlicht: (2025) -
VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition
von: Jin, Hongbo, et al.
Veröffentlicht: (2025) -
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
von: Song, Enxin, et al.
Veröffentlicht: (2024) -
T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
von: Guo, Chaohong, et al.
Veröffentlicht: (2026) -
FocusGraph: Graph-Structured Frame Selection for Embodied Long Video Question Answering
von: Zemskova, Tatiana, et al.
Veröffentlicht: (2026)