MAGIC-Enhanced Keyword Prompting for Zero-Shot Audio Captioning with CLIP Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Govindarajan, Vijay, Patel, Pratik, Tripathi, Sahil, Hoque, Md Azizul, Kashyap, Gautam Siddharth |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Do Clinical Question Answering Systems Really Need Specialised Medical Fine Tuning?
par: Ray, Sushant Kumar, et autres
Publié: (2026)
par: Ray, Sushant Kumar, et autres
Publié: (2026)
They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
par: Tripathi, Sahil, et autres
Publié: (2026)
par: Tripathi, Sahil, et autres
Publié: (2026)
AlignCultura: Towards Culturally Aligned Large Language Models?
par: Kashyap, Gautam Siddharth, et autres
Publié: (2026)
par: Kashyap, Gautam Siddharth, et autres
Publié: (2026)
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
par: Kashyap, Gautam Siddharth, et autres
Publié: (2026)
par: Kashyap, Gautam Siddharth, et autres
Publié: (2026)
FactGenius: Combining Zero-Shot Prompting and Fuzzy Relation Mining to Improve Fact Verification with Knowledge Graphs
par: Gautam, Sushant
Publié: (2024)
par: Gautam, Sushant
Publié: (2024)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
par: Kashyap, Gautam Siddharth, et autres
Publié: (2025)
par: Kashyap, Gautam Siddharth, et autres
Publié: (2025)
Too Helpful, Too Harmless, Too Honest or Just Right?
par: Kashyap, Gautam Siddharth, et autres
Publié: (2025)
par: Kashyap, Gautam Siddharth, et autres
Publié: (2025)
ZSE-Cap: A Zero-Shot Ensemble for Image Retrieval and Prompt-Guided Captioning
par: Dinh, Duc-Tai, et autres
Publié: (2025)
par: Dinh, Duc-Tai, et autres
Publié: (2025)
MaiBERT: A Pre-training Corpus and Language Model for Low-Resourced Maithili Language
par: Yadav, Sumit, et autres
Publié: (2025)
par: Yadav, Sumit, et autres
Publié: (2025)
From Text to Transformation: A Comprehensive Review of Large Language Models' Versatility
par: Kaur, Pravneet, et autres
Publié: (2024)
par: Kaur, Pravneet, et autres
Publié: (2024)
Keyword-Centric Prompting for One-Shot Event Detection with Self-Generated Rationale Enhancements
par: Li, Ziheng, et autres
Publié: (2025)
par: Li, Ziheng, et autres
Publié: (2025)
Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
par: Yen, Hao, et autres
Publié: (2024)
par: Yen, Hao, et autres
Publié: (2024)
Image-Caption Encoding for Improving Zero-Shot Generalization
par: Yu, Eric Yang, et autres
Publié: (2024)
par: Yu, Eric Yang, et autres
Publié: (2024)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
par: Chen, Xiaofu, et autres
Publié: (2025)
par: Chen, Xiaofu, et autres
Publié: (2025)
Can Large Language Models Make Everyone Happy?
par: Naseem, Usman, et autres
Publié: (2026)
par: Naseem, Usman, et autres
Publié: (2026)
Do Large Language Models Reflect Demographic Pluralism in Safety?
par: Naseem, Usman, et autres
Publié: (2026)
par: Naseem, Usman, et autres
Publié: (2026)
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
par: RRV, Aswin, et autres
Publié: (2024)
par: RRV, Aswin, et autres
Publié: (2024)
Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
par: Liu, Jizhong, et autres
Publié: (2024)
par: Liu, Jizhong, et autres
Publié: (2024)
Zero Resource Cross-Lingual Part Of Speech Tagging
par: Chopra, Sahil
Publié: (2024)
par: Chopra, Sahil
Publié: (2024)
Brevity Constraints Reverse Performance Hierarchies in Language Models
par: Hakim, MD Azizul
Publié: (2026)
par: Hakim, MD Azizul
Publié: (2026)
Which Words Matter Most in Zero-Shot Prompts?
par: Sadr, Nikta Gohari, et autres
Publié: (2025)
par: Sadr, Nikta Gohari, et autres
Publié: (2025)
Better Zero-Shot Reasoning with Role-Play Prompting
par: Kong, Aobo, et autres
Publié: (2023)
par: Kong, Aobo, et autres
Publié: (2023)
Are Aligned Large Language Models Still Misaligned?
par: Naseem, Usman, et autres
Publié: (2026)
par: Naseem, Usman, et autres
Publié: (2026)
Enabling Natural Zero-Shot Prompting on Encoder Models via Statement-Tuning
par: Elshabrawy, Ahmed, et autres
Publié: (2024)
par: Elshabrawy, Ahmed, et autres
Publié: (2024)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
par: Tang, Changli, et autres
Publié: (2025)
par: Tang, Changli, et autres
Publié: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
par: Chaffin, Antoine, et autres
Publié: (2024)
par: Chaffin, Antoine, et autres
Publié: (2024)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
par: Anand, Nishit, et autres
Publié: (2024)
par: Anand, Nishit, et autres
Publié: (2024)
CultureCLIP: Empowering CLIP with Cultural Awareness through Synthetic Images and Contextualized Captions
par: Huang, Yuchen, et autres
Publié: (2025)
par: Huang, Yuchen, et autres
Publié: (2025)
Towards Diverse and Efficient Audio Captioning via Diffusion Models
par: Xu, Manjie, et autres
Publié: (2024)
par: Xu, Manjie, et autres
Publié: (2024)
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
par: Kambhatla, Gauri, et autres
Publié: (2025)
par: Kambhatla, Gauri, et autres
Publié: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
par: Luo, Jianjie, et autres
Publié: (2024)
par: Luo, Jianjie, et autres
Publié: (2024)
Enhancing Small Language Models for Cross-Lingual Generalized Zero-Shot Classification with Soft Prompt Tuning
par: Philippy, Fred, et autres
Publié: (2025)
par: Philippy, Fred, et autres
Publié: (2025)
GRP: Goal-Reversed Prompting for Zero-Shot Evaluation with LLMs
par: Song, Mingyang, et autres
Publié: (2025)
par: Song, Mingyang, et autres
Publié: (2025)
Updating CLIP to Prefer Descriptions Over Captions
par: Zur, Amir, et autres
Publié: (2024)
par: Zur, Amir, et autres
Publié: (2024)
Adapter-state Sharing CLIP for Parameter-efficient Multimodal Sarcasm Detection
par: Jana, Soumyadeep, et autres
Publié: (2025)
par: Jana, Soumyadeep, et autres
Publié: (2025)
FAQ-Gen: An automated system to generate domain-specific FAQs to aid content comprehension
par: Kale, Sahil, et autres
Publié: (2024)
par: Kale, Sahil, et autres
Publié: (2024)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
par: Möller, Lucas, et autres
Publié: (2024)
par: Möller, Lucas, et autres
Publié: (2024)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
par: Phukan, Orchid Chetia, et autres
Publié: (2024)
Adaptive Test-Time Scaling for Zero-Shot Respiratory Audio Classification
par: Wang, Tsai-Ning, et autres
Publié: (2026)
par: Wang, Tsai-Ning, et autres
Publié: (2026)
Leveraging Zero-Shot Prompting for Efficient Language Model Distillation
par: Vöge, Lukas, et autres
Publié: (2024)
par: Vöge, Lukas, et autres
Publié: (2024)
Documents similaires
-
Do Clinical Question Answering Systems Really Need Specialised Medical Fine Tuning?
par: Ray, Sushant Kumar, et autres
Publié: (2026) -
They Said Memes Were Harmless-We Found the Ones That Hurt: Decoding Jokes, Symbols, and Cultural References
par: Tripathi, Sahil, et autres
Publié: (2026) -
AlignCultura: Towards Culturally Aligned Large Language Models?
par: Kashyap, Gautam Siddharth, et autres
Publié: (2026) -
When the Model Said 'No Comment', We Knew Helpfulness Was Dead, Honesty Was Alive, and Safety Was Terrified
par: Kashyap, Gautam Siddharth, et autres
Publié: (2026) -
FactGenius: Combining Zero-Shot Prompting and Fuzzy Relation Mining to Improve Fact Verification with Knowledge Graphs
par: Gautam, Sushant
Publié: (2024)