Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Goel, Arushi, Ghosh, Sreyan, Kim, Jaehyeon, Kumar, Sonal, Kong, Zhifeng, Lee, Sang-gil, Yang, Chao-Han Huck, Duraiswami, Ramani, Manocha, Dinesh, Valle, Rafael, Catanzaro, Bryan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Music Flamingo: Scaling Music Understanding in Audio Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025)
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026)
Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
RECAP: Retrieval-Augmented Audio Captioning
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
ETTA: Elucidating the Design Space of Text-to-Audio Models
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
Audio Dialogues: Dialogues dataset for audio and music understanding
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
von: Sakshi, S, et al.
Veröffentlicht: (2025)
von: Sakshi, S, et al.
Veröffentlicht: (2025)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
PAT: Parameter-Free Audio-Text Aligner to Boost Zero-Shot Audio Classification
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
von: Sakshi, S, et al.
Veröffentlicht: (2024)
von: Sakshi, S, et al.
Veröffentlicht: (2024)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2023)
A2SB: Audio-to-Audio Schrodinger Bridges
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2025)
Improving Text-To-Audio Models with Synthetic Captions
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
von: Anand, Nishit, et al.
Veröffentlicht: (2024)
Do Audio-Language Models Understand Linguistic Variations?
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
von: Selvakumar, Ramaneswaran, et al.
Veröffentlicht: (2024)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
ProSE: Diffusion Priors for Speech Enhancement
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
Room Impulse Response Synthesis via Differentiable Feedback Delay Networks for Efficient Spatial Audio Rendering
von: Gerami, Armin, et al.
Veröffentlicht: (2025)
von: Gerami, Armin, et al.
Veröffentlicht: (2025)
SILA: Signal-to-Language Augmentation for Enhanced Control in Text-to-Audio Generation
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
von: Kumar, Sonal, et al.
Veröffentlicht: (2024)
UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
von: Tian, Jinchuan, et al.
Veröffentlicht: (2025)
Biomimetic Frontend for Differentiable Audio Processing
von: Famularo, Ruolan Leslie, et al.
Veröffentlicht: (2024)
von: Famularo, Ruolan Leslie, et al.
Veröffentlicht: (2024)
MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
von: Kumar, Sonal, et al.
Veröffentlicht: (2025)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
von: Seth, Ashish, et al.
Veröffentlicht: (2026)
AV-RIR: Audio-Visual Room Impulse Response Estimation
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
von: Ratnarajah, Anton, et al.
Veröffentlicht: (2023)
EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation Learning
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
von: Seth, Ashish, et al.
Veröffentlicht: (2024)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
von: Hung, Chia-Yu, et al.
Veröffentlicht: (2024)
mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2025)
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
von: Bharadwaj, Shikhar, et al.
Veröffentlicht: (2025)
Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
von: Wan, Zhen, et al.
Veröffentlicht: (2026)
von: Wan, Zhen, et al.
Veröffentlicht: (2026)
Whisper-Flamingo: Integrating Visual Features into Whisper for Audio-Visual Speech Recognition and Translation
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
von: Rouditchenko, Andrew, et al.
Veröffentlicht: (2024)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Stable Audio Open
von: Evans, Zach, et al.
Veröffentlicht: (2024)
von: Evans, Zach, et al.
Veröffentlicht: (2024)
Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models
von: He, Haolin, et al.
Veröffentlicht: (2025)
von: He, Haolin, et al.
Veröffentlicht: (2025)
ODAQ: Open Dataset of Audio Quality
von: Torcoli, Matteo, et al.
Veröffentlicht: (2023)
von: Torcoli, Matteo, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Music Flamingo: Scaling Music Understanding in Audio Language Models
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025) -
Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2025) -
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2026) -
Synthio: Augmenting Small-Scale Audio Classification Datasets with Synthetic Data
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024) -
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)