A Closer Look at the Limitations of Instruction Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ghosh, Sreyan, Evuru, Chandra Kiran Reddy, Kumar, Sonal, S, Ramaneswaran, Aneja, Deepali, Jin, Zeyu, Duraiswami, Ramani, Manocha, Dinesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
CoDa: Constrained Generation based Data Augmentation for Low-Resource NLP
von: Evuru, Chandra Kiran Reddy, et al.
Veröffentlicht: (2024)
von: Evuru, Chandra Kiran Reddy, et al.
Veröffentlicht: (2024)
Visual Description Grounding Reduces Hallucinations and Boosts Reasoning in LVLMs
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
von: Sakshi, S, et al.
Veröffentlicht: (2024)
von: Sakshi, S, et al.
Veröffentlicht: (2024)
ASPIRE: Language-Guided Data Augmentation for Improving Robustness Against Spurious Correlations
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
von: Seth, Ashish, et al.
Veröffentlicht: (2025)
von: Seth, Ashish, et al.
Veröffentlicht: (2025)
ProSE: Diffusion Priors for Speech Enhancement
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
von: Sakshi, S, et al.
Veröffentlicht: (2025)
von: Sakshi, S, et al.
Veröffentlicht: (2025)
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
von: Goel, Arushi, et al.
Veröffentlicht: (2025)
von: Goel, Arushi, et al.
Veröffentlicht: (2025)
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
von: Lokegaonkar, Vaibhavi, et al.
Veröffentlicht: (2026)
von: Lokegaonkar, Vaibhavi, et al.
Veröffentlicht: (2026)
Do Audio-Visual Large Language Models Really See and Hear?
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2026)
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2026)
MultiVox: A Benchmark for Evaluating Voice Assistants for Multimodal Interactions
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2025)
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2025)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
Do Vision-Language Models Understand Compound Nouns?
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
Quantifying Document Impact in RAG-LLMs
von: Gerami, Armin, et al.
Veröffentlicht: (2025)
von: Gerami, Armin, et al.
Veröffentlicht: (2025)
Learning Illumination Control in Diffusion Models
von: Anand, Nishit, et al.
Veröffentlicht: (2026)
von: Anand, Nishit, et al.
Veröffentlicht: (2026)
A Closer Look into LLMs for Table Understanding
von: Wang, Jia, et al.
Veröffentlicht: (2026)
von: Wang, Jia, et al.
Veröffentlicht: (2026)
A Closer Look at System Prompt Robustness
von: Mu, Norman, et al.
Veröffentlicht: (2025)
von: Mu, Norman, et al.
Veröffentlicht: (2025)
Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training
von: Du, Wenyu, et al.
Veröffentlicht: (2024)
von: Du, Wenyu, et al.
Veröffentlicht: (2024)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
Do Audio-Language Models Understand Linguistic Variations?
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
A Closer Look at Machine Unlearning for Large Language Models
von: Yuan, Xiaojian, et al.
Veröffentlicht: (2024)
von: Yuan, Xiaojian, et al.
Veröffentlicht: (2024)
A Closer Look at the Self-Verification Abilities of Large Language Models in Logical Reasoning
von: Hong, Ruixin, et al.
Veröffentlicht: (2023)
von: Hong, Ruixin, et al.
Veröffentlicht: (2023)
What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
von: Cheng, Stephen, et al.
Veröffentlicht: (2026)
Closer Look at Efficient Inference Methods: A Survey of Speculative Decoding
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
von: Ryu, Hyun, et al.
Veröffentlicht: (2024)
Interpretable Question Answering with Knowledge Graphs
von: Aneja, Kartikeya, et al.
Veröffentlicht: (2025)
von: Aneja, Kartikeya, et al.
Veröffentlicht: (2025)
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
von: Chehade, Mohamad, et al.
Veröffentlicht: (2025)
von: Chehade, Mohamad, et al.
Veröffentlicht: (2025)
Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2025)
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2025)
Disperse-Then-Merge: Pushing the Limits of Instruction Tuning via Alignment Tax Reduction
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
AV-RIR: Audio-Visual Room Impulse Response Estimation
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023) -
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023) -
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024) -
ABEX: Data Augmentation for Low-Resource NLU via Expanding Abstract Descriptions
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024) -
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)