MAGIC-Enhanced Keyword Prompting for Zero-Shot Audio Captioning with CLIP Models
Fuente:
arXiv
Saved in:
| Main Authors: | Govindarajan, Vijay, Patel, Pratik, Tripathi, Sahil, Hoque, Md Azizul, Kashyap, Gautam Siddharth |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Clinical Question Answering Systems Really Need Specialised Medical Fine Tuning?
by: Ray, Sushant Kumar, et al.
Published: (2026)
by: Ray, Sushant Kumar, et al.
Published: (2026)
They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
by: Tripathi, Sahil, et al.
Published: (2026)
by: Tripathi, Sahil, et al.
Published: (2026)
AlignCultura: Towards Culturally Aligned Large Language Models?
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
by: Kashyap, Gautam Siddharth, et al.
Published: (2026)
FactGenius: Combining Zero-Shot Prompting and Fuzzy Relation Mining to Improve Fact Verification with Knowledge Graphs
by: Gautam, Sushant
Published: (2024)
by: Gautam, Sushant
Published: (2024)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
Too Helpful, Too Harmless, Too Honest or Just Right?
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
ZSE-Cap: A Zero-Shot Ensemble for Image Retrieval and Prompt-Guided Captioning
by: Dinh, Duc-Tai, et al.
Published: (2025)
by: Dinh, Duc-Tai, et al.
Published: (2025)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
by: Yadav, Sumit, et al.
Published: (2025)
by: Yadav, Sumit, et al.
Published: (2025)
From Text to Transformation: A Comprehensive Review of Large Language Models' Versatility
by: Kaur, Pravneet, et al.
Published: (2024)
by: Kaur, Pravneet, et al.
Published: (2024)
Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
by: Li, Ziheng, et al.
Published: (2025)
by: Li, Ziheng, et al.
Published: (2025)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
by: Yen, Hao, et al.
Published: (2024)
by: Yen, Hao, et al.
Published: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
by: Yu, Eric Yang, et al.
Published: (2024)
by: Yu, Eric Yang, et al.
Published: (2024)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
by: Chen, Xiaofu, et al.
Published: (2025)
by: Chen, Xiaofu, et al.
Published: (2025)
Can Large Language Models Make Everyone Happy?
by: Naseem, Usman, et al.
Published: (2026)
by: Naseem, Usman, et al.
Published: (2026)
Do Large Language Models Reflect Demographic Pluralism in Safety?
by: Naseem, Usman, et al.
Published: (2026)
by: Naseem, Usman, et al.
Published: (2026)
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
by: RRV, Aswin, et al.
Published: (2024)
by: RRV, Aswin, et al.
Published: (2024)
Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
by: Liu, Jizhong, et al.
Published: (2024)
by: Liu, Jizhong, et al.
Published: (2024)
Zero Resource Cross-Lingual Part Of Speech Tagging
by: Chopra, Sahil
Published: (2024)
by: Chopra, Sahil
Published: (2024)
Brevity Constraints Reverse Performance Hierarchies in Language Models
by: Hakim, MD Azizul
Published: (2026)
by: Hakim, MD Azizul
Published: (2026)
Which Words Matter Most in Zero-Shot Prompts?
by: Sadr, Nikta Gohari, et al.
Published: (2025)
by: Sadr, Nikta Gohari, et al.
Published: (2025)
Better Zero-Shot Reasoning with Role-Play Prompting
by: Kong, Aobo, et al.
Published: (2023)
by: Kong, Aobo, et al.
Published: (2023)
Are Aligned Large Language Models Still Misaligned?
by: Naseem, Usman, et al.
Published: (2026)
by: Naseem, Usman, et al.
Published: (2026)
Enabling Natural Zero-Shot Prompting on Encoder Models via Statement-Tuning
by: Elshabrawy, Ahmed, et al.
Published: (2024)
by: Elshabrawy, Ahmed, et al.
Published: (2024)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
by: Tang, Changli, et al.
Published: (2025)
by: Tang, Changli, et al.
Published: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
by: Chaffin, Antoine, et al.
Published: (2024)
by: Chaffin, Antoine, et al.
Published: (2024)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
by: Anand, Nishit, et al.
Published: (2024)
by: Anand, Nishit, et al.
Published: (2024)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
by: Huang, Yuchen, et al.
Published: (2025)
by: Huang, Yuchen, et al.
Published: (2025)
Towards Diverse and Efficient Audio Captioning via Diffusion Models
by: Xu, Manjie, et al.
Published: (2024)
by: Xu, Manjie, et al.
Published: (2024)
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
by: Kambhatla, Gauri, et al.
Published: (2025)
by: Kambhatla, Gauri, et al.
Published: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
by: Luo, Jianjie, et al.
Published: (2024)
by: Luo, Jianjie, et al.
Published: (2024)
Enhancing Small Language Models for Cross-Lingual Generalized Zero-Shot Classification with Soft Prompt Tuning
by: Philippy, Fred, et al.
Published: (2025)
by: Philippy, Fred, et al.
Published: (2025)
GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
by: Song, Mingyang, et al.
Published: (2025)
by: Song, Mingyang, et al.
Published: (2025)
Updating CLIP to Prefer Descriptions Over Captions
by: Zur, Amir, et al.
Published: (2024)
by: Zur, Amir, et al.
Published: (2024)
Adapter-state Sharing CLIP for Parameter-efficient Multimodal Sarcasm Detection
by: Jana, Soumyadeep, et al.
Published: (2025)
by: Jana, Soumyadeep, et al.
Published: (2025)
FAQ-Gen: An automated system to generate domain-specific FAQs to aid content comprehension
by: Kale, Sahil, et al.
Published: (2024)
by: Kale, Sahil, et al.
Published: (2024)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
by: Möller, Lucas, et al.
Published: (2024)
by: Möller, Lucas, et al.
Published: (2024)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
by: Phukan, Orchid Chetia, et al.
Published: (2024)
by: Phukan, Orchid Chetia, et al.
Published: (2024)
Adaptive Test-Time Scaling for Zero-Shot Respiratory Audio Classification
by: Wang, Tsai-Ning, et al.
Published: (2026)
by: Wang, Tsai-Ning, et al.
Published: (2026)
Leveraging Zero-Shot Prompting for Efficient Language Model Distillation
by: Vöge, Lukas, et al.
Published: (2024)
by: Vöge, Lukas, et al.
Published: (2024)
Similar Items
-
Do Clinical Question Answering Systems Really Need Specialised Medical Fine Tuning?
by: Ray, Sushant Kumar, et al.
Published: (2026) -
They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
by: Tripathi, Sahil, et al.
Published: (2026) -
AlignCultura: Towards Culturally Aligned Large Language Models?
by: Kashyap, Gautam Siddharth, et al.
Published: (2026) -
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
by: Kashyap, Gautam Siddharth, et al.
Published: (2026) -
FactGenius: Combining Zero-Shot Prompting and Fuzzy Relation Mining to Improve Fact Verification with Knowledge Graphs
by: Gautam, Sushant
Published: (2024)