Can VLMs be used on videos for action recognition? LLMs are Visual Reasoning Coordinators
Fuente:
arXiv
Saved in:
| Main Author: | Lunia, Harsh |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024)
by: Lunia, Harsh, et al.
Published: (2024)
Multi-model learning by sequential reading of untrimmed videos for action recognition
by: Kamiya, Kodai, et al.
Published: (2024)
by: Kamiya, Kodai, et al.
Published: (2024)
Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
by: Elmansoury, Sary, et al.
Published: (2025)
by: Elmansoury, Sary, et al.
Published: (2025)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
by: Törtei, Brigitta Malagurski, et al.
Published: (2025)
Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
by: Ballout, Mohamad, et al.
Published: (2025)
by: Ballout, Mohamad, et al.
Published: (2025)
Can masking background and object reduce static bias for zero-shot action recognition?
by: Fukuzawa, Takumi, et al.
Published: (2025)
by: Fukuzawa, Takumi, et al.
Published: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
by: Berman, Shmuel, et al.
Published: (2025)
by: Berman, Shmuel, et al.
Published: (2025)
VisualActBench: Can VLMs See and Act like a Human?
by: Zhang, Daoan, et al.
Published: (2025)
by: Zhang, Daoan, et al.
Published: (2025)
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
by: Park, Simon, et al.
Published: (2025)
by: Park, Simon, et al.
Published: (2025)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
by: Guo, Hao, et al.
Published: (2026)
by: Guo, Hao, et al.
Published: (2026)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
by: Ge, Yuyao, et al.
Published: (2025)
by: Ge, Yuyao, et al.
Published: (2025)
VRIQ: Benchmarking and Analyzing Visual-Reasoning IQ of VLMs
by: Khezresmaeilzadeh, Tina, et al.
Published: (2026)
by: Khezresmaeilzadeh, Tina, et al.
Published: (2026)
Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition
by: Sun, Jian, et al.
Published: (2026)
by: Sun, Jian, et al.
Published: (2026)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
by: Imam, Mohamed Fazli, et al.
Published: (2025)
by: Imam, Mohamed Fazli, et al.
Published: (2025)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
by: Zhang, Xintong, et al.
Published: (2025)
by: Zhang, Xintong, et al.
Published: (2025)
Can VLMs Recall Factual Associations From Visual References?
by: Ashok, Dhananjay, et al.
Published: (2025)
by: Ashok, Dhananjay, et al.
Published: (2025)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
by: Tekin, Selim Furkan, et al.
Published: (2026)
by: Tekin, Selim Furkan, et al.
Published: (2026)
Egocentric zone-aware action recognition across environments
by: Peirone, Simone Alberto, et al.
Published: (2024)
by: Peirone, Simone Alberto, et al.
Published: (2024)
CLGRPO: Reasoning Ability Enhancement for Small VLMs
by: Wang, Fanyi, et al.
Published: (2025)
by: Wang, Fanyi, et al.
Published: (2025)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
by: Tan, Zhangyun, et al.
Published: (2026)
by: Tan, Zhangyun, et al.
Published: (2026)
iVISPAR -- An Interactive Visual-Spatial Reasoning Benchmark for VLMs
by: Mayer, Julius, et al.
Published: (2025)
by: Mayer, Julius, et al.
Published: (2025)
Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
by: Yu, Jiaao, et al.
Published: (2025)
by: Yu, Jiaao, et al.
Published: (2025)
Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
by: Chen, Kaitao, et al.
Published: (2025)
by: Chen, Kaitao, et al.
Published: (2025)
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
by: Chen, Weixin, et al.
Published: (2026)
by: Chen, Weixin, et al.
Published: (2026)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
by: Li, Weiming, et al.
Published: (2025)
by: Li, Weiming, et al.
Published: (2025)
ADHD diagnosis based on action characteristics recorded in videos using machine learning
by: Li, Yichun, et al.
Published: (2024)
by: Li, Yichun, et al.
Published: (2024)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
by: Li, Yiwei, et al.
Published: (2026)
by: Li, Yiwei, et al.
Published: (2026)
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
by: Gambashidze, Alexander, et al.
Published: (2025)
by: Gambashidze, Alexander, et al.
Published: (2025)
Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
by: Izadi, Amirmohammad, et al.
Published: (2025)
by: Izadi, Amirmohammad, et al.
Published: (2025)
[De|Re]constructing VLMs' Reasoning in Counting
by: Alghisi, Simone, et al.
Published: (2025)
by: Alghisi, Simone, et al.
Published: (2025)
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
by: Zhang, Gengyuan, et al.
Published: (2023)
by: Zhang, Gengyuan, et al.
Published: (2023)
Emotion recognition in talking-face videos using persistent entropy and neural networks
by: Paluzo-Hidalgo, Eduardo, et al.
Published: (2021)
by: Paluzo-Hidalgo, Eduardo, et al.
Published: (2021)
CFE-PPAR: Compression-friendly encryption for privacy-preserving action recognition leveraging video transformers
by: Lin, Haiwei, et al.
Published: (2026)
by: Lin, Haiwei, et al.
Published: (2026)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
by: Huang, Irene, et al.
Published: (2024)
by: Huang, Irene, et al.
Published: (2024)
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
Can MLLMs Reason About Visual Persuasion? Evaluating the Efficacy and Faithfulness of Reasoning
by: Lee, Naeun, et al.
Published: (2026)
by: Lee, Naeun, et al.
Published: (2026)
Visual Prompt Based Reasoning for Offroad Mapping using Multimodal LLMs
by: Nasser, Abdelmoamen, et al.
Published: (2026)
by: Nasser, Abdelmoamen, et al.
Published: (2026)
Deep learning for action spotting in association football videos
by: Giancola, Silvio, et al.
Published: (2024)
by: Giancola, Silvio, et al.
Published: (2024)
Deep kernel video approximation for unsupervised action segmentation
by: Pintea, Silvia L., et al.
Published: (2026)
by: Pintea, Silvia L., et al.
Published: (2026)
Caption This, Reason That: VLMs Caught in the Middle
by: Weng, Zihan, et al.
Published: (2025)
by: Weng, Zihan, et al.
Published: (2025)
Similar Items
-
IndicSTR12: A Dataset for Indic Scene Text Recognition
by: Lunia, Harsh, et al.
Published: (2024) -
Multi-model learning by sequential reading of untrimmed videos for action recognition
by: Kamiya, Kodai, et al.
Published: (2024) -
Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
by: Elmansoury, Sary, et al.
Published: (2025) -
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
by: Törtei, Brigitta Malagurski, et al.
Published: (2025) -
Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
by: Ballout, Mohamad, et al.
Published: (2025)