ViLA: Efficient Video-Language Alignment for Video Question Answering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xijun, Liang, Junbang, Wang, Chun-Kai, Deng, Kenan, Lou, Yu, Lin, Ming, Yang, Shan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoDistill: Language-aware Vision Distillation for Video Question Answering
von: Zou, Bo, et al.
Veröffentlicht: (2024)
von: Zou, Bo, et al.
Veröffentlicht: (2024)
Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment
von: Chen, Jin, et al.
Veröffentlicht: (2024)
von: Chen, Jin, et al.
Veröffentlicht: (2024)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025)
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
von: Liao, Zhaohe, et al.
Veröffentlicht: (2024)
von: Liao, Zhaohe, et al.
Veröffentlicht: (2024)
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering
von: Yu, Ting, et al.
Veröffentlicht: (2024)
von: Yu, Ting, et al.
Veröffentlicht: (2024)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
von: Nie, Yuxiang, et al.
Veröffentlicht: (2025)
Agentic Keyframe Search for Video Question Answering
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
von: Fan, Sunqi, et al.
Veröffentlicht: (2025)
ViViD: Video Virtual Try-on using Diffusion Models
von: Fang, Zixun, et al.
Veröffentlicht: (2024)
von: Fang, Zixun, et al.
Veröffentlicht: (2024)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
von: Di, Shangzhe, et al.
Veröffentlicht: (2025)
CogStream: Context-guided Streaming Video Question Answering
von: Zhao, Zicheng, et al.
Veröffentlicht: (2025)
von: Zhao, Zicheng, et al.
Veröffentlicht: (2025)
ViLLa: Video Reasoning Segmentation with Large Language Model
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2023)
von: Chen, Xiuyuan, et al.
Veröffentlicht: (2023)
ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering
von: Guan, Kaisi, et al.
Veröffentlicht: (2025)
von: Guan, Kaisi, et al.
Veröffentlicht: (2025)
A Simple LLM Framework for Long-Range Video Question-Answering
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
von: Zhang, Ce, et al.
Veröffentlicht: (2023)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
von: Meng, Yiran, et al.
Veröffentlicht: (2025)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
von: Song, Enxin, et al.
Veröffentlicht: (2024)
von: Song, Enxin, et al.
Veröffentlicht: (2024)
Semantic Event Graphs for Long-Form Video Question Answering
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026)
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026)
Scene-Text Grounding for Text-Based Video Question Answering
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
von: Zhou, Sheng, et al.
Veröffentlicht: (2024)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
von: Ma, Haodi, et al.
Veröffentlicht: (2025)
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
von: Yang, Xuyi, et al.
Veröffentlicht: (2025)
von: Yang, Xuyi, et al.
Veröffentlicht: (2025)
Progressive Image Restoration via Text-Conditioned Video Generation
von: Kang, Peng, et al.
Veröffentlicht: (2025)
von: Kang, Peng, et al.
Veröffentlicht: (2025)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
Glance and Focus: Memory Prompting for Multi-Event Video Question Answering
von: Bai, Ziyi, et al.
Veröffentlicht: (2024)
von: Bai, Ziyi, et al.
Veröffentlicht: (2024)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
von: Chaybouti, Sofian, et al.
Veröffentlicht: (2025)
MobileViCLIP: An Efficient Video-Text Model for Mobile Devices
von: Yang, Min, et al.
Veröffentlicht: (2025)
von: Yang, Min, et al.
Veröffentlicht: (2025)
Object-centric Video Question Answering with Visual Grounding and Referring
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
von: Wang, Haochen, et al.
Veröffentlicht: (2025)
ImViD: Immersive Volumetric Videos for Enhanced VR Engagement
von: Yang, Zhengxian, et al.
Veröffentlicht: (2025)
von: Yang, Zhengxian, et al.
Veröffentlicht: (2025)
Cross-modal Causal Relation Alignment for Video Question Grounding
von: Chen, Weixing, et al.
Veröffentlicht: (2025)
von: Chen, Weixing, et al.
Veröffentlicht: (2025)
Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels
von: Liang, Tianming, et al.
Veröffentlicht: (2024)
von: Liang, Tianming, et al.
Veröffentlicht: (2024)
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
von: Cai, Chen, et al.
Veröffentlicht: (2024)
von: Cai, Chen, et al.
Veröffentlicht: (2024)
ViSpeak: Visual Instruction Feedback in Streaming Videos
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
von: Liang, Jianxin, et al.
Veröffentlicht: (2024)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
von: Chen, Yilong, et al.
Veröffentlicht: (2025)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
von: Liang, Lili, et al.
Veröffentlicht: (2024)
von: Liang, Lili, et al.
Veröffentlicht: (2024)
A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
von: Zou, Yuanhao, et al.
Veröffentlicht: (2025)
von: Zou, Yuanhao, et al.
Veröffentlicht: (2025)
BIMBA: Selective-Scan Compression for Long-Range Video Question Answering
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
von: Islam, Md Mohaiminul, et al.
Veröffentlicht: (2025)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
von: Fernando, Basura, et al.
Veröffentlicht: (2025)
von: Fernando, Basura, et al.
Veröffentlicht: (2025)
Question-Answering Dense Video Events
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
von: Qin, Hangyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VideoDistill: Language-aware Vision Distillation for Video Question Answering
von: Zou, Bo, et al.
Veröffentlicht: (2024) -
Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment
von: Chen, Jin, et al.
Veröffentlicht: (2024) -
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
von: Wang, Haochen, et al.
Veröffentlicht: (2025) -
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
von: Drago, Mauro Orazio, et al.
Veröffentlicht: (2025) -
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
von: Liao, Zhaohe, et al.
Veröffentlicht: (2024)