Saved in:
| Main Authors: | Beliaev, Mark, Yang, Victor, Raju, Madhura, Sun, Jiachen, Hu, Xinghai |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2502.09573 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
by: Cheng, Jiale, et al.
Published: (2025)
by: Cheng, Jiale, et al.
Published: (2025)
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding
by: Sun, Qi, et al.
Published: (2024)
by: Sun, Qi, et al.
Published: (2024)
Commonsense for Zero-Shot Natural Language Video Localization
by: Holla, Meghana, et al.
Published: (2023)
by: Holla, Meghana, et al.
Published: (2023)
Image-Caption Encoding for Improving Zero-Shot Generalization
by: Yu, Eric Yang, et al.
Published: (2024)
by: Yu, Eric Yang, et al.
Published: (2024)
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
by: Zhang, Yiming, et al.
Published: (2024)
by: Zhang, Yiming, et al.
Published: (2024)
No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance
by: Udandarao, Vishaal, et al.
Published: (2024)
by: Udandarao, Vishaal, et al.
Published: (2024)
Zero-Shot Action Recognition in Surveillance Videos
by: Pereira, Joao, et al.
Published: (2024)
by: Pereira, Joao, et al.
Published: (2024)
IPO: Interpretable Prompt Optimization for Vision-Language Models
by: Du, Yingjun, et al.
Published: (2024)
by: Du, Yingjun, et al.
Published: (2024)
Data Alignment for Zero-Shot Concept Generation in Dermatology AI
by: Gadgil, Soham, et al.
Published: (2024)
by: Gadgil, Soham, et al.
Published: (2024)
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
by: Awal, Rabiul, et al.
Published: (2023)
by: Awal, Rabiul, et al.
Published: (2023)
Spectral Prompt Tuning:Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
by: Xu, Wenhao, et al.
Published: (2023)
by: Xu, Wenhao, et al.
Published: (2023)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
Troika: Multi-Path Cross-Modal Traction for Compositional Zero-Shot Learning
by: Huang, Siteng, et al.
Published: (2023)
by: Huang, Siteng, et al.
Published: (2023)
How Does Diverse Interpretability of Textual Prompts Impact Medical Vision-Language Zero-Shot Tasks?
by: Wang, Sicheng, et al.
Published: (2024)
by: Wang, Sicheng, et al.
Published: (2024)
Zero-Shot Defense Against Toxic Images via Inherent Multimodal Alignment in LVLMs
by: Zhao, Wei, et al.
Published: (2025)
by: Zhao, Wei, et al.
Published: (2025)
CLIP meets DINO for Tuning Zero-Shot Classifier using Unlabeled Image Collections
by: Imam, Mohamed Fazli, et al.
Published: (2024)
by: Imam, Mohamed Fazli, et al.
Published: (2024)
DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
by: Yang, Zongxin, et al.
Published: (2024)
by: Yang, Zongxin, et al.
Published: (2024)
Prompt-Induced Score Variance in Zero-Shot Binary Vision-Language Safety Classification
by: Weng, Charles, et al.
Published: (2026)
by: Weng, Charles, et al.
Published: (2026)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
RadZero: Similarity-Based Cross-Attention for Explainable Vision-Language Alignment in Chest X-ray with Zero-Shot Multi-Task Capability
by: Park, Jonggwon, et al.
Published: (2025)
by: Park, Jonggwon, et al.
Published: (2025)
Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations
by: Rizwan, Naquee, et al.
Published: (2024)
by: Rizwan, Naquee, et al.
Published: (2024)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
by: Luo, Jianjie, et al.
Published: (2024)
by: Luo, Jianjie, et al.
Published: (2024)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
by: Yang, Cheng-Fu, et al.
Published: (2024)
by: Yang, Cheng-Fu, et al.
Published: (2024)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
by: Tanaka, Ryota, et al.
Published: (2024)
by: Tanaka, Ryota, et al.
Published: (2024)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
by: Wu, Haoning, et al.
Published: (2024)
by: Wu, Haoning, et al.
Published: (2024)
Temporal Preference Optimization for Long-Form Video Understanding
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
Zero-Shot Refinement of Buildings' Segmentation Models using SAM
by: Mayladan, Ali, et al.
Published: (2023)
by: Mayladan, Ali, et al.
Published: (2023)
Few-Shot Adversarial Prompt Learning on Vision-Language Models
by: Zhou, Yiwei, et al.
Published: (2024)
by: Zhou, Yiwei, et al.
Published: (2024)
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
by: Jiang, Jiachen, et al.
Published: (2025)
by: Jiang, Jiachen, et al.
Published: (2025)
SimpleGPT: Improving GPT via A Simple Normalization Strategy
by: Chen, Marco, et al.
Published: (2026)
by: Chen, Marco, et al.
Published: (2026)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
by: Nagar, Aishik, et al.
Published: (2024)
by: Nagar, Aishik, et al.
Published: (2024)
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
by: Gupta, Akash, et al.
Published: (2025)
by: Gupta, Akash, et al.
Published: (2025)
Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
by: Dipta, Shubhashis Roy, et al.
Published: (2025)
by: Dipta, Shubhashis Roy, et al.
Published: (2025)
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks
by: Singh, Simranjit, et al.
Published: (2024)
by: Singh, Simranjit, et al.
Published: (2024)
CoS: Chain-of-Shot Prompting for Long Video Understanding
by: Hu, Jian, et al.
Published: (2025)
by: Hu, Jian, et al.
Published: (2025)
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning
by: Lin, Han, et al.
Published: (2023)
by: Lin, Han, et al.
Published: (2023)
GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
by: Yang, Kaichun, et al.
Published: (2025)
by: Yang, Kaichun, et al.
Published: (2025)
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding
by: Wang, Xiao, et al.
Published: (2024)
by: Wang, Xiao, et al.
Published: (2024)
Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
by: Liu, Kaiting, et al.
Published: (2026)
by: Liu, Kaiting, et al.
Published: (2026)
Similar Items
-
VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
by: Cheng, Jiale, et al.
Published: (2025) -
Exploiting GPT-4 Vision for Zero-shot Point Cloud Understanding
by: Sun, Qi, et al.
Published: (2024) -
Commonsense for Zero-Shot Natural Language Video Localization
by: Holla, Meghana, et al.
Published: (2023) -
Image-Caption Encoding for Improving Zero-Shot Generalization
by: Yu, Eric Yang, et al.
Published: (2024) -
Beyond Training: Dynamic Token Merging for Zero-Shot Video Understanding
by: Zhang, Yiming, et al.
Published: (2024)