Meta-Prompting for Automating Zero-shot Visual Recognition with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Mirza, M. Jehanzeb, Karlinsky, Leonid, Lin, Wei, Doveh, Sivan, Micorek, Jakub, Kozinski, Mateusz, Kuehne, Hilde, Possegger, Horst |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MULDE: Multiscale Log-Density Estimation via Denoising Score Matching for Video Anomaly Detection
by: Micorek, Jakub, et al.
Published: (2024)
by: Micorek, Jakub, et al.
Published: (2024)
Comparison Visual Instruction Tuning
by: Lin, Wei, et al.
Published: (2024)
by: Lin, Wei, et al.
Published: (2024)
Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation
by: Spoecklberger, Johannes, et al.
Published: (2025)
by: Spoecklberger, Johannes, et al.
Published: (2025)
Teaching VLMs to Localize Specific Objects from In-context Examples
by: Doveh, Sivan, et al.
Published: (2024)
by: Doveh, Sivan, et al.
Published: (2024)
Towards Multimodal In-Context Learning for Vision & Language Models
by: Doveh, Sivan, et al.
Published: (2024)
by: Doveh, Sivan, et al.
Published: (2024)
VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
by: Gavrikov, Paul, et al.
Published: (2025)
by: Gavrikov, Paul, et al.
Published: (2025)
MAEDAY: MAE for few and zero shot AnomalY-Detection
by: Schwartz, Eli, et al.
Published: (2022)
by: Schwartz, Eli, et al.
Published: (2022)
TTRV: Test-Time Reinforcement Learning for Vision Language Models
by: Singh, Akshit, et al.
Published: (2025)
by: Singh, Akshit, et al.
Published: (2025)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
by: Huang, Irene, et al.
Published: (2024)
by: Huang, Irene, et al.
Published: (2024)
Into the Fog: Evaluating Robustness of Multiple Object Tracking
by: Kirillova, Nadezda, et al.
Published: (2024)
by: Kirillova, Nadezda, et al.
Published: (2024)
GLOV: Guided Large Language Models as Implicit Optimizers for Vision Language Models
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
by: Mirza, M. Jehanzeb, et al.
Published: (2024)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
by: Rouditchenko, Andrew, et al.
Published: (2024)
by: Rouditchenko, Andrew, et al.
Published: (2024)
LiveXiv -- A Multi-Modal Live Benchmark Based on Arxiv Papers Content
by: Shabtay, Nimrod, et al.
Published: (2024)
by: Shabtay, Nimrod, et al.
Published: (2024)
PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
by: Selch, Lukas, et al.
Published: (2025)
by: Selch, Lukas, et al.
Published: (2025)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
by: Hansen, Jacob, et al.
Published: (2025)
by: Hansen, Jacob, et al.
Published: (2025)
NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning
by: Schwartz, Eli, et al.
Published: (2024)
by: Schwartz, Eli, et al.
Published: (2024)
AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers
by: Araujo, Edson, et al.
Published: (2026)
by: Araujo, Edson, et al.
Published: (2026)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
by: Bousselham, Walid, et al.
Published: (2025)
by: Bousselham, Walid, et al.
Published: (2025)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
by: Shvetsova, Nina, et al.
Published: (2023)
by: Shvetsova, Nina, et al.
Published: (2023)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
by: Rouditchenko, Andrew, et al.
Published: (2025)
by: Rouditchenko, Andrew, et al.
Published: (2025)
CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment
by: Araujo, Edson, et al.
Published: (2025)
by: Araujo, Edson, et al.
Published: (2025)
WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition
by: Bock, Marius, et al.
Published: (2023)
by: Bock, Marius, et al.
Published: (2023)
Overflow Prevention Enhances Long-Context Recurrent LLMs
by: Ben-Kish, Assaf, et al.
Published: (2025)
by: Ben-Kish, Assaf, et al.
Published: (2025)
State-Space Large Audio Language Models
by: Bhati, Saurabhchand, et al.
Published: (2024)
by: Bhati, Saurabhchand, et al.
Published: (2024)
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
by: Bhati, Saurabhchand, et al.
Published: (2024)
by: Bhati, Saurabhchand, et al.
Published: (2024)
TTA-Vid: Generalized Test-Time Adaptation for Video Reasoning
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
by: Jahagirdar, Soumya Shamarao, et al.
Published: (2026)
TimeLogic: A Temporal Logic Benchmark for Video QA
by: Swetha, Sirnam, et al.
Published: (2025)
by: Swetha, Sirnam, et al.
Published: (2025)
Efficient Motion Prediction: A Lightweight & Accurate Trajectory Prediction Model With Fast Training and Inference Speed
by: Prutsch, Alexander, et al.
Published: (2024)
by: Prutsch, Alexander, et al.
Published: (2024)
CALM: Class-Conditional Sparse Attention Vectors for Large Audio-Language Models
by: Mehta, Videet, et al.
Published: (2026)
by: Mehta, Videet, et al.
Published: (2026)
Vision-Language Guidance for LiDAR-based Unsupervised 3D Object Detection
by: Fruhwirth-Reisinger, Christian, et al.
Published: (2024)
by: Fruhwirth-Reisinger, Christian, et al.
Published: (2024)
Probing the effectiveness of World Models for Spatial Reasoning through Test-time Scaling
by: Jha, Saurav, et al.
Published: (2025)
by: Jha, Saurav, et al.
Published: (2025)
Zero-shot Prompt-based Video Encoder for Surgical Gesture Recognition
by: Rao, Mingxing, et al.
Published: (2024)
by: Rao, Mingxing, et al.
Published: (2024)
One Model, Many Behaviors: Training-Induced Effects on Out-of-Distribution Detection
by: Krumpl, Gerhard, et al.
Published: (2026)
by: Krumpl, Gerhard, et al.
Published: (2026)
ICONIC-444: A 3.1-Million-Image Dataset for OOD Detection Research
by: Krumpl, Gerhard, et al.
Published: (2026)
by: Krumpl, Gerhard, et al.
Published: (2026)
3VL: Using Trees to Improve Vision-Language Models' Interpretability
by: Yellinek, Nir, et al.
Published: (2023)
by: Yellinek, Nir, et al.
Published: (2023)
Sample- and Parameter-Efficient Auto-Regressive Image Models
by: Amrani, Elad, et al.
Published: (2024)
by: Amrani, Elad, et al.
Published: (2024)
Streaming Real-Time Trajectory Prediction Using Endpoint-Aware Modeling
by: Prutsch, Alexander, et al.
Published: (2026)
by: Prutsch, Alexander, et al.
Published: (2026)
Prompt-based Visual Alignment for Zero-shot Policy Transfer
by: Gao, Haihan, et al.
Published: (2024)
by: Gao, Haihan, et al.
Published: (2024)
When LLaVA Meets Objects: Token Composition for Vision-Language-Models
by: Jahagirdar, Soumya, et al.
Published: (2026)
by: Jahagirdar, Soumya, et al.
Published: (2026)
REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
by: Chaybouti, Sofian, et al.
Published: (2025)
by: Chaybouti, Sofian, et al.
Published: (2025)
Similar Items
-
MULDE: Multiscale Log-Density Estimation via Denoising Score Matching for Video Anomaly Detection
by: Micorek, Jakub, et al.
Published: (2024) -
Comparison Visual Instruction Tuning
by: Lin, Wei, et al.
Published: (2024) -
Exploring Modality Guidance to Enhance VFM-based Feature Fusion for UDA in 3D Semantic Segmentation
by: Spoecklberger, Johannes, et al.
Published: (2025) -
Teaching VLMs to Localize Specific Objects from In-context Examples
by: Doveh, Sivan, et al.
Published: (2024) -
Towards Multimodal In-Context Learning for Vision & Language Models
by: Doveh, Sivan, et al.
Published: (2024)