End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liang, Jianxin, Meng, Xiaojun, Wang, Yueqian, Liu, Chang, Liu, Qun, Zhao, Dongyan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
von: Wang, Yueqian, et al.
Veröffentlicht: (2025)
End-to-End Navigation with Vision Language Models: Transforming Spatial Reasoning into Question-Answering
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
von: Goetting, Dylan, et al.
Veröffentlicht: (2024)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Friends-MMC: A Dataset for Multi-modal Multi-party Conversation Understanding
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
von: Han, Wei, et al.
Veröffentlicht: (2023)
von: Han, Wei, et al.
Veröffentlicht: (2023)
FREAK: A Fine-grained Hallucination Evaluation Benchmark for Advanced MLLMs
von: Yin, Zhihan, et al.
Veröffentlicht: (2026)
von: Yin, Zhihan, et al.
Veröffentlicht: (2026)
PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
von: Cai, Chen, et al.
Veröffentlicht: (2024)
von: Cai, Chen, et al.
Veröffentlicht: (2024)
EVA: Efficient Reinforcement Learning for End-to-End Video Agent
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
von: Zhang, Yaolun, et al.
Veröffentlicht: (2026)
ICR-Drive: Instruction Counterfactual Robustness for End-to-End Language-Driven Autonomous Driving
von: Hamid, Kaiser, et al.
Veröffentlicht: (2026)
von: Hamid, Kaiser, et al.
Veröffentlicht: (2026)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
von: Liang, Hao, et al.
Veröffentlicht: (2024)
von: Liang, Hao, et al.
Veröffentlicht: (2024)
E2E-GMNER: End-to-End Generative Grounded Multimodal Named Entity Recognition
von: Zhang, Meng, et al.
Veröffentlicht: (2026)
von: Zhang, Meng, et al.
Veröffentlicht: (2026)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
von: Zhu, Wang, et al.
Veröffentlicht: (2023)
von: Zhu, Wang, et al.
Veröffentlicht: (2023)
Actions and Objects Pathways for Domain Adaptation in Video Question Answering
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2024)
von: Mohamud, Safaa Abdullahi Moallim, et al.
Veröffentlicht: (2024)
TextFormer: A Query-based End-to-End Text Spotter with Mixed Supervision
von: Zhai, Yukun, et al.
Veröffentlicht: (2023)
von: Zhai, Yukun, et al.
Veröffentlicht: (2023)
A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
von: Zou, Yuanhao, et al.
Veröffentlicht: (2025)
von: Zou, Yuanhao, et al.
Veröffentlicht: (2025)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
von: Liu, Shuming, et al.
Veröffentlicht: (2023)
von: Liu, Shuming, et al.
Veröffentlicht: (2023)
Beyond End-to-End VLMs: Leveraging Intermediate Text Representations for Superior Flowchart Understanding
von: Ye, Junyi, et al.
Veröffentlicht: (2024)
von: Ye, Junyi, et al.
Veröffentlicht: (2024)
Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
ANIM-400K: A Large-Scale Dataset for Automated End-To-End Dubbing of Video
von: Cai, Kevin, et al.
Veröffentlicht: (2024)
von: Cai, Kevin, et al.
Veröffentlicht: (2024)
TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
von: You, Ling, et al.
Veröffentlicht: (2025)
von: You, Ling, et al.
Veröffentlicht: (2025)
RIG: Synergizing Reasoning and Imagination in End-to-End Generalist Policy
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
von: Zhao, Zhonghan, et al.
Veröffentlicht: (2025)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
Selectively Answering Visual Questions
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
von: Eisenschlos, Julian Martin, et al.
Veröffentlicht: (2024)
Top-down Activity Representation Learning for Video Question Answering
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts
von: Wang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2025)
Pose-Based Sign Language Spotting via an End-to-End Encoder Architecture
von: Johnny, Samuel Ebimobowei, et al.
Veröffentlicht: (2025)
von: Johnny, Samuel Ebimobowei, et al.
Veröffentlicht: (2025)
Multi-granularity Contrastive Cross-modal Collaborative Generation for End-to-End Long-term Video Question Answering
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
Grounding Chest X-Ray Visual Question Answering with Generated Radiology Reports
von: Serra, Francesco Dalla, et al.
Veröffentlicht: (2025)
von: Serra, Francesco Dalla, et al.
Veröffentlicht: (2025)
LaPA: Latent Prompt Assist Model For Medical Visual Question Answering
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
von: Gu, Tiancheng, et al.
Veröffentlicht: (2024)
End-to-End Agentic RAG System Training for Traceable Diagnostic Reasoning
von: Zheng, Qiaoyu, et al.
Veröffentlicht: (2025)
von: Zheng, Qiaoyu, et al.
Veröffentlicht: (2025)
Multi-object event graph representation learning for Video Question Answering
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
von: Wang, Yanan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HawkEye: Training Video-Text LLMs for Grounding Text in Videos
von: Wang, Yueqian, et al.
Veröffentlicht: (2024) -
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
von: Wang, Yueqian, et al.
Veröffentlicht: (2024) -
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025) -
VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
von: Wang, Yueqian, et al.
Veröffentlicht: (2024) -
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)