REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
Fuente:
arXiv
Saved in:
| Main Authors: | Chaybouti, Sofian, Bousselham, Walid, Wolter, Moritz, Kuehne, Hilde |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
by: Bousselham, Walid, et al.
Published: (2024)
by: Bousselham, Walid, et al.
Published: (2024)
LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity
by: Bousselham, Walid, et al.
Published: (2024)
by: Bousselham, Walid, et al.
Published: (2024)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025)
by: Vogel, Felix, et al.
Published: (2025)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)
by: Jahagirdar, Soumya, et al.
Published: (2026)
DEX-AR: A Dynamic Explainability Method for Autoregressive Vision-Language Models
by: Bousselham, Walid, et al.
Published: (2026)
by: Bousselham, Walid, et al.
Published: (2026)
TimeLogic: A Temporal Logic Benchmark for Video QA
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion
by: Shvetsova, Nina, et al.
Published: (2025)
by: Shvetsova, Nina, et al.
Published: (2025)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
by: Shvetsova, Nina, et al.
Published: (2023)
by: Shvetsova, Nina, et al.
Published: (2023)
Top-down Activity Representation Learning for Video Question Answering
by: Wang, Yanan, et al.
Published: (2024)
by: Wang, Yanan, et al.
Published: (2024)
GHR-VQA: Graph-guided Hierarchical Relational Reasoning for Video Question Answering
by: Brilli, Dionysia Danai, et al.
Published: (2025)
by: Brilli, Dionysia Danai, et al.
Published: (2025)
Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models
by: T V, Sethuraman, et al.
Published: (2026)
by: T V, Sethuraman, et al.
Published: (2026)
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering
by: Li, Haopeng, et al.
Published: (2024)
by: Li, Haopeng, et al.
Published: (2024)
Agentic Keyframe Search for Video Question Answering
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
Grounded Question-Answering in Long Egocentric Videos
by: Di, Shangzhe, et al.
Published: (2023)
by: Di, Shangzhe, et al.
Published: (2023)
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval
by: Di, Shangzhe, et al.
Published: (2025)
by: Di, Shangzhe, et al.
Published: (2025)
ViLA: Efficient Video-Language Alignment for Video Question Answering
by: Wang, Xijun, et al.
Published: (2023)
by: Wang, Xijun, et al.
Published: (2023)
VideoDistill: Language-aware Vision Distillation for Video Question Answering
by: Zou, Bo, et al.
Published: (2024)
by: Zou, Bo, et al.
Published: (2024)
SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
STRIVE: Structured Spatiotemporal Exploration for Reinforcement Learning in Video Question Answering
by: Bahrami, Emad, et al.
Published: (2026)
by: Bahrami, Emad, et al.
Published: (2026)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
Narrative Aligned Long Form Video Question Answering
by: Jain, Rahul, et al.
Published: (2026)
by: Jain, Rahul, et al.
Published: (2026)
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering
by: Liao, Zhaohe, et al.
Published: (2024)
by: Liao, Zhaohe, et al.
Published: (2024)
Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach
by: Oh, Ju-Young
Published: (2025)
by: Oh, Ju-Young
Published: (2025)
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
by: Song, Enxin, et al.
Published: (2024)
by: Song, Enxin, et al.
Published: (2024)
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering
by: Liang, Lili, et al.
Published: (2024)
by: Liang, Lili, et al.
Published: (2024)
Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering
by: Kwon, Minchan, et al.
Published: (2026)
by: Kwon, Minchan, et al.
Published: (2026)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
by: Fernando, Basura, et al.
Published: (2025)
by: Fernando, Basura, et al.
Published: (2025)
Object-centric Video Question Answering with Visual Grounding and Referring
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Admitting Ignorance Helps the Video Question Answering Models to Answer
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
VDMA: Video Question Answering with Dynamically Generated Multi-Agents
by: Kugo, Noriyuki, et al.
Published: (2024)
by: Kugo, Noriyuki, et al.
Published: (2024)
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering
by: Meng, Yiran, et al.
Published: (2025)
by: Meng, Yiran, et al.
Published: (2025)
Fréchet Wavelet Distance: A Domain-Agnostic Metric for Image Generation
by: Veeramacheneni, Lokesh, et al.
Published: (2023)
by: Veeramacheneni, Lokesh, et al.
Published: (2023)
ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
by: Nie, Yuxiang, et al.
Published: (2025)
by: Nie, Yuxiang, et al.
Published: (2025)
DocVideoQA: Towards Comprehensive Understanding of Document-Centric Videos through Question Answering
by: Wang, Haochen, et al.
Published: (2025)
by: Wang, Haochen, et al.
Published: (2025)
Question-Answering Dense Video Events
by: Qin, Hangyu, et al.
Published: (2024)
by: Qin, Hangyu, et al.
Published: (2024)
Improving Video Question Answering through query-based frame selection
by: Patil, Himanshu, et al.
Published: (2026)
by: Patil, Himanshu, et al.
Published: (2026)
StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
by: Chen, Yilong, et al.
Published: (2025)
by: Chen, Yilong, et al.
Published: (2025)
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
by: Fan, Sunqi, et al.
Published: (2025)
by: Fan, Sunqi, et al.
Published: (2025)
Similar Items
-
MaskInversion: Localized Embeddings via Optimization of Explainability Maps
by: Bousselham, Walid, et al.
Published: (2024) -
LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity
by: Bousselham, Walid, et al.
Published: (2024) -
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025) -
VideoGEM: Training-free Action Grounding in Videos
by: Vogel, Felix, et al.
Published: (2025) -
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)