Zero-shot Action Localization via the Confidence of Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Aklilu, Josiah, Wang, Xiaohan, Yeung-Levy, Serena |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting Active Learning in the Era of Vision Foundation Models
by: Gupte, Sanket Rajan, et al.
Published: (2024)
by: Gupte, Sanket Rajan, et al.
Published: (2024)
Depth-guided NeRF Training via Earth Mover's Distance
by: Rau, Anita, et al.
Published: (2024)
by: Rau, Anita, et al.
Published: (2024)
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
by: Endo, Mark, et al.
Published: (2024)
by: Endo, Mark, et al.
Published: (2024)
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024)
by: Sui, Elaine, et al.
Published: (2024)
Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
by: Rau, Anita, et al.
Published: (2025)
by: Rau, Anita, et al.
Published: (2025)
Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
by: Wang, Xiaohan, et al.
Published: (2024)
by: Wang, Xiaohan, et al.
Published: (2024)
BIOMEDICA: An Open Biomedical Image-Caption Archive, Dataset, and Vision-Language Models Derived from Scientific Literature
by: Lozano, Alejandro, et al.
Published: (2025)
by: Lozano, Alejandro, et al.
Published: (2025)
Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
by: Endo, Mark, et al.
Published: (2025)
by: Endo, Mark, et al.
Published: (2025)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
Foundation Models Secretly Understand Neural Network Weights: Enhancing Hypernetwork Architectures with Foundation Models
by: Gu, Jeffrey, et al.
Published: (2025)
by: Gu, Jeffrey, et al.
Published: (2025)
The Impact of Image Resolution on Biomedical Multimodal Large Language Models
by: Chen, Liangyu, et al.
Published: (2025)
by: Chen, Liangyu, et al.
Published: (2025)
Seeing Is Believing? A Benchmark for Multimodal Large Language Models on Visual Illusions and Anomalies
by: Hou, Wenjin, et al.
Published: (2026)
by: Hou, Wenjin, et al.
Published: (2026)
NegVQA: Can Vision Language Models Understand Negation?
by: Zhang, Yuhui, et al.
Published: (2025)
by: Zhang, Yuhui, et al.
Published: (2025)
Multi-Human Mesh Recovery with Transformers
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography
by: Spinaci, Gianmarco, et al.
Published: (2025)
by: Spinaci, Gianmarco, et al.
Published: (2025)
V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
by: Tang, Bingda, et al.
Published: (2026)
by: Tang, Bingda, et al.
Published: (2026)
Video Action Differencing
by: Burgess, James, et al.
Published: (2025)
by: Burgess, James, et al.
Published: (2025)
SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
by: Deng, Andong, et al.
Published: (2025)
by: Deng, Andong, et al.
Published: (2025)
Fine-tuning MLLMs Without Forgetting Is Easier Than You Think
by: Li, He, et al.
Published: (2026)
by: Li, He, et al.
Published: (2026)
SCALE: Semantic- and Confidence-Aware Conditional Variational Autoencoder for Zero-shot Skeleton-based Action Recognition
by: Oraki, Soroush, et al.
Published: (2026)
by: Oraki, Soroush, et al.
Published: (2026)
Towards Zero-shot Human-Object Interaction Detection via Vision-Language Integration
by: Xue, Weiying, et al.
Published: (2024)
by: Xue, Weiying, et al.
Published: (2024)
Why are Visually-Grounded Language Models Bad at Image Classification?
by: Zhang, Yuhui, et al.
Published: (2024)
by: Zhang, Yuhui, et al.
Published: (2024)
Zero-Shot Robustness of Vision Language Models Via Confidence-Aware Weighting
by: Naghavian, Nikoo, et al.
Published: (2025)
by: Naghavian, Nikoo, et al.
Published: (2025)
Boosting Audio-visual Zero-shot Learning with Large Language Models
by: Chen, Haoxing, et al.
Published: (2023)
by: Chen, Haoxing, et al.
Published: (2023)
Improving Zero-shot ADL Recognition with Large Language Models through Event-based Context and Confidence
by: Fiori, Michele, et al.
Published: (2026)
by: Fiori, Michele, et al.
Published: (2026)
TTRV: Test-Time Reinforcement Learning for Vision Language Models
by: Singh, Akshit, et al.
Published: (2025)
by: Singh, Akshit, et al.
Published: (2025)
Ask, Pose, Unite: Scaling Data Acquisition for Close Interactions with Vision Language Models
by: Bravo-Sánchez, Laura, et al.
Published: (2024)
by: Bravo-Sánchez, Laura, et al.
Published: (2024)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
by: Zhou, Jiaming, et al.
Published: (2024)
by: Zhou, Jiaming, et al.
Published: (2024)
Zero-shot Generalizable Incremental Learning for Vision-Language Object Detection
by: Deng, Jieren, et al.
Published: (2024)
by: Deng, Jieren, et al.
Published: (2024)
Diffusion-HPC: Synthetic Data Generation for Human Mesh Recovery in Challenging Domains
by: Weng, Zhenzhen, et al.
Published: (2023)
by: Weng, Zhenzhen, et al.
Published: (2023)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
by: Zhao, Shuai, et al.
Published: (2023)
by: Zhao, Shuai, et al.
Published: (2023)
Masked Diffusion Vision-Language Models for Temporal Action Localization
by: Wang, Fengshun, et al.
Published: (2026)
by: Wang, Fengshun, et al.
Published: (2026)
Test-Time Low Rank Adaptation via Confidence Maximization for Zero-Shot Generalization of Vision-Language Models
by: Imam, Raza, et al.
Published: (2024)
by: Imam, Raza, et al.
Published: (2024)
Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
by: Wu, Shengguang, et al.
Published: (2025)
by: Wu, Shengguang, et al.
Published: (2025)
Training-Free Zero-Shot Temporal Action Detection with Vision-Language Models
by: Han, Chaolei, et al.
Published: (2025)
by: Han, Chaolei, et al.
Published: (2025)
Label Propagation for Zero-shot Classification with Vision-Language Models
by: Stojnić, Vladan, et al.
Published: (2024)
by: Stojnić, Vladan, et al.
Published: (2024)
Part-aware Unified Representation of Language and Skeleton for Zero-shot Action Recognition
by: Zhu, Anqi, et al.
Published: (2024)
by: Zhu, Anqi, et al.
Published: (2024)
Similar Items
-
Revisiting Active Learning in the Era of Vision Foundation Models
by: Gupte, Sanket Rajan, et al.
Published: (2024) -
Depth-guided NeRF Training via Earth Mover's Distance
by: Rau, Anita, et al.
Published: (2024) -
Feather the Throttle: Revisiting Visual Token Pruning for Vision-Language Model Acceleration
by: Endo, Mark, et al.
Published: (2024) -
Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models
by: Sui, Elaine, et al.
Published: (2024) -
Systematic Evaluation of Large Vision-Language Models for Surgical Artificial Intelligence
by: Rau, Anita, et al.
Published: (2025)