QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | He, Zhixian, Zhao, Pengcheng, Zhang, Fuwei, Lin, Shujin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2025)
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2025)
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
VideoQA in the Era of LLMs: An Empirical Study
von: Xiao, Junbin, et al.
Veröffentlicht: (2024)
von: Xiao, Junbin, et al.
Veröffentlicht: (2024)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
von: Liang, Lili, et al.
Veröffentlicht: (2024)
von: Liang, Lili, et al.
Veröffentlicht: (2024)
Reading Between the Lanes: Text VideoQA on the Road
von: Tom, George, et al.
Veröffentlicht: (2023)
von: Tom, George, et al.
Veröffentlicht: (2023)
VideoQA-SC: Adaptive Semantic Communication for Video Question Answering
von: Guo, Jiangyuan, et al.
Veröffentlicht: (2024)
von: Guo, Jiangyuan, et al.
Veröffentlicht: (2024)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
ENTER: Event Based Interpretable Reasoning for VideoQA
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
von: Ayyubi, Hammad, et al.
Veröffentlicht: (2025)
Understanding Complexity in VideoQA via Visual Program Generation
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2025)
CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2026)
TUMTraffic-VideoQA: A Benchmark for Unified Spatio-Temporal Video Understanding in Traffic Scenes
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
von: Zhou, Xingcheng, et al.
Veröffentlicht: (2025)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
von: Song, Zijie, et al.
Veröffentlicht: (2025)
von: Song, Zijie, et al.
Veröffentlicht: (2025)
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
von: Rawal, Ishaan Singh, et al.
Veröffentlicht: (2023)
StreamingCoT: A Dataset for Temporal Dynamics and Multimodal Chain-of-Thought Reasoning in Streaming VideoQA
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
von: Hu, Yuhang, et al.
Veröffentlicht: (2025)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark
von: Heyward, Joseph, et al.
Veröffentlicht: (2024)
von: Heyward, Joseph, et al.
Veröffentlicht: (2024)
Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries
von: Amoroso, Roberto, et al.
Veröffentlicht: (2024)
von: Amoroso, Roberto, et al.
Veröffentlicht: (2024)
QG-VTC: Question-Guided Visual Token Compression in MLLMs for Efficient VQA
von: Li, Shuai, et al.
Veröffentlicht: (2025)
von: Li, Shuai, et al.
Veröffentlicht: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
von: Zhang, Xiaoman, et al.
Veröffentlicht: (2023)
VQA$^2$: Visual Question Answering for Video Quality Assessment
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
von: Jia, Ziheng, et al.
Veröffentlicht: (2024)
ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
von: Wang, Xucheng, et al.
Veröffentlicht: (2026)
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
von: Tateno, Masatoshi, et al.
Veröffentlicht: (2025)
von: Tateno, Masatoshi, et al.
Veröffentlicht: (2025)
GHR-VQA: Graph-guided Hierarchical Relational Reasoning for Video Question Answering
von: Brilli, Dionysia Danai, et al.
Veröffentlicht: (2025)
von: Brilli, Dionysia Danai, et al.
Veröffentlicht: (2025)
YTCommentQA: Video Question Answerability in Instructional Videos
von: Yang, Saelyne, et al.
Veröffentlicht: (2024)
von: Yang, Saelyne, et al.
Veröffentlicht: (2024)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2026)
von: He, Haibin, et al.
Veröffentlicht: (2026)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
Dynamic Try-On: Taming Video Virtual Try-on with Dynamic Attention Mechanism
von: Zheng, Jun, et al.
Veröffentlicht: (2024)
von: Zheng, Jun, et al.
Veröffentlicht: (2024)
Visual Robustness Benchmark for Visual Question Answering (VQA)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
von: Ishmam, Md Farhan, et al.
Veröffentlicht: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
von: Zhou, Sheng, et al.
Veröffentlicht: (2025)
VITON-DiT: Learning In-the-Wild Video Try-On from Human Dance Videos via Diffusion Transformers
von: Zheng, Jun, et al.
Veröffentlicht: (2024)
von: Zheng, Jun, et al.
Veröffentlicht: (2024)
RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
von: Guan, Runwei, et al.
Veröffentlicht: (2025)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
von: Zhang, Yan, et al.
Veröffentlicht: (2025)
Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track
von: Gupta, Deepak, et al.
Veröffentlicht: (2024)
von: Gupta, Deepak, et al.
Veröffentlicht: (2024)
MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
von: Shaar, Shaden, et al.
Veröffentlicht: (2026)
von: Shaar, Shaden, et al.
Veröffentlicht: (2026)
MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering
von: Li, Zhifei, et al.
Veröffentlicht: (2026)
von: Li, Zhifei, et al.
Veröffentlicht: (2026)
SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
von: He, Haibin, et al.
Veröffentlicht: (2025)
von: He, Haibin, et al.
Veröffentlicht: (2025)
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
von: Yu, Suhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
von: Zhao, Pengcheng, et al.
Veröffentlicht: (2025) -
ReasVQA: Advancing VideoQA with Imperfect Reasoning Process
von: Liang, Jianxin, et al.
Veröffentlicht: (2025) -
VideoQA in the Era of LLMs: An Empirical Study
von: Xiao, Junbin, et al.
Veröffentlicht: (2024) -
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
von: Liang, Lili, et al.
Veröffentlicht: (2024) -
Reading Between the Lanes: Text VideoQA on the Road
von: Tom, George, et al.
Veröffentlicht: (2023)