An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Wonkyun, Choi, Changin, Lee, Wonseok, Rhee, Wonjong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025)
by: Choi, Changin, et al.
Published: (2025)
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026)
by: Ko, Jungmin, et al.
Published: (2026)
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
by: Byun, Dongnam, et al.
Published: (2025)
by: Byun, Dongnam, et al.
Published: (2025)
Improving Forward Compatibility in Class Incremental Learning by Increasing Representation Rank and Feature Richness
by: Kim, Jaeill, et al.
Published: (2024)
by: Kim, Jaeill, et al.
Published: (2024)
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2023)
by: Zhang, Jiarui, et al.
Published: (2023)
Task-Specific Preconditioner for Cross-Domain Few-Shot Learning
by: Kang, Suhyun, et al.
Published: (2024)
by: Kang, Suhyun, et al.
Published: (2024)
Empowering Large Language Model for Continual Video Question Answering with Collaborative Prompting
by: Cai, Chen, et al.
Published: (2024)
by: Cai, Chen, et al.
Published: (2024)
Actions and Objects Pathways for Domain Adaptation in Video Question Answering
by: Mohamud, Safaa Abdullahi Moallim, et al.
Published: (2024)
by: Mohamud, Safaa Abdullahi Moallim, et al.
Published: (2024)
Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization
by: Kim, Jimyeong, et al.
Published: (2024)
by: Kim, Jimyeong, et al.
Published: (2024)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling
by: Liang, Jianxin, et al.
Published: (2024)
by: Liang, Jianxin, et al.
Published: (2024)
STAIR: Spatial-Temporal Reasoning with Auditable Intermediate Results for Video Question Answering
by: Wang, Yueqian, et al.
Published: (2024)
by: Wang, Yueqian, et al.
Published: (2024)
Soft Head Selection for Injecting ICL-Derived Task Embeddings
by: Park, Jungwon, et al.
Published: (2025)
by: Park, Jungwon, et al.
Published: (2025)
Flashback: Memory-Driven Zero-shot, Real-time Video Anomaly Detection
by: Lee, Hyogun, et al.
Published: (2025)
by: Lee, Hyogun, et al.
Published: (2025)
Harmonizing Visual and Textual Embeddings for Zero-Shot Text-to-Image Customization
by: Song, Yeji, et al.
Published: (2024)
by: Song, Yeji, et al.
Published: (2024)
EVQAScore: A Fine-grained Metric for Video Question Answering Data Quality Evaluation
by: Liang, Hao, et al.
Published: (2024)
by: Liang, Hao, et al.
Published: (2024)
Can Language Models Laugh at YouTube Short-form Videos?
by: Ko, Dayoon, et al.
Published: (2023)
by: Ko, Dayoon, et al.
Published: (2023)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
by: Han, Wei, et al.
Published: (2023)
by: Han, Wei, et al.
Published: (2023)
Top-down Activity Representation Learning for Video Question Answering
by: Wang, Yanan, et al.
Published: (2024)
by: Wang, Yanan, et al.
Published: (2024)
A Video is Worth 10,000 Words: Training and Benchmarking with Diverse Captions for Better Long Video Retrieval
by: Gwilliam, Matthew, et al.
Published: (2023)
by: Gwilliam, Matthew, et al.
Published: (2023)
ViQAgent: Zero-Shot Video Question Answering via Agent with Open-Vocabulary Grounding Validation
by: Montes, Tony, et al.
Published: (2025)
by: Montes, Tony, et al.
Published: (2025)
Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering
by: Romero, David, et al.
Published: (2024)
by: Romero, David, et al.
Published: (2024)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023)
by: Awal, Rabiul, et al.
Published: (2023)
Multi-object event graph representation learning for Video Question Answering
by: Wang, Yanan, et al.
Published: (2024)
by: Wang, Yanan, et al.
Published: (2024)
VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
by: Meng, Rui, et al.
Published: (2025)
by: Meng, Rui, et al.
Published: (2025)
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
by: Cho, Seunghyuk, et al.
Published: (2025)
by: Cho, Seunghyuk, et al.
Published: (2025)
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
Re:Verse -- Can Your VLM Read a Manga?
by: Baranwal, Aaditya, et al.
Published: (2025)
by: Baranwal, Aaditya, et al.
Published: (2025)
Generalizing Visual Question Answering from Synthetic to Human-Written Questions via a Chain of QA with a Large Language Model
by: Kim, Taehee, et al.
Published: (2024)
by: Kim, Taehee, et al.
Published: (2024)
Selectively Answering Visual Questions
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
by: Eisenschlos, Julian Martin, et al.
Published: (2024)
On-Off Pattern Encoding and Path-Count Encoding as Deep Neural Network Representations
by: Jung, Euna, et al.
Published: (2024)
by: Jung, Euna, et al.
Published: (2024)
ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation
by: Kim, Jimyeong, et al.
Published: (2025)
by: Kim, Jimyeong, et al.
Published: (2025)
FIQ: Fundamental Question Generation with the Integration of Question Embeddings for Video Question Answering
by: Oh, Ju-Young, et al.
Published: (2025)
by: Oh, Ju-Young, et al.
Published: (2025)
SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
by: Guo, Ziyu, et al.
Published: (2024)
by: Guo, Ziyu, et al.
Published: (2024)
Zero-Shot Action Recognition in Surveillance Videos
by: Pereira, Joao, et al.
Published: (2024)
by: Pereira, Joao, et al.
Published: (2024)
IFCap: Image-like Retrieval and Frequency-based Entity Filtering for Zero-shot Captioning
by: Lee, Soeun, et al.
Published: (2024)
by: Lee, Soeun, et al.
Published: (2024)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
by: Kim, Jongha, et al.
Published: (2026)
by: Kim, Jongha, et al.
Published: (2026)
Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
by: Chowdhury, Md Intisar, et al.
Published: (2025)
by: Chowdhury, Md Intisar, et al.
Published: (2025)
Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
by: Yoon, Eunseop, et al.
Published: (2025)
by: Yoon, Eunseop, et al.
Published: (2025)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
by: Oh, Hongseok, et al.
Published: (2025)
by: Oh, Hongseok, et al.
Published: (2025)
Similar Items
-
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
by: Choi, Changin, et al.
Published: (2025) -
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
by: Ko, Jungmin, et al.
Published: (2026) -
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
by: Byun, Dongnam, et al.
Published: (2025) -
Improving Forward Compatibility in Class Incremental Learning by Increasing Representation Rank and Feature Richness
by: Kim, Jaeill, et al.
Published: (2024) -
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2023)