Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghosh, Sreyan, Goel, Arushi, Jayakumar, Kaousheik, Koroshinadze, Lasha, Anand, Nishit, Kong, Zhifeng, Gururani, Siddharth, Lee, Sang-gil, Kim, Jaehyeon, Aljafari, Aya, Yang, Chao-Han Huck, Kim, Sungwon, Duraiswami, Ramani, Manocha, Dinesh, Shoeybi, Mohammad, Catanzaro, Bryan, Liu, Ming-Yu, Ping, Wei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910123114364928
author Ghosh, Sreyan
Goel, Arushi
Jayakumar, Kaousheik
Koroshinadze, Lasha
Anand, Nishit
Kong, Zhifeng
Gururani, Siddharth
Lee, Sang-gil
Kim, Jaehyeon
Aljafari, Aya
Yang, Chao-Han Huck
Kim, Sungwon
Duraiswami, Ramani
Manocha, Dinesh
Shoeybi, Mohammad
Catanzaro, Bryan
Liu, Ming-Yu
Ping, Wei
author_facet Ghosh, Sreyan
Goel, Arushi
Jayakumar, Kaousheik
Koroshinadze, Lasha
Anand, Nishit
Kong, Zhifeng
Gururani, Siddharth
Lee, Sang-gil
Kim, Jaehyeon
Aljafari, Aya
Yang, Chao-Han Huck
Kim, Sungwon
Duraiswami, Ramani
Manocha, Dinesh
Shoeybi, Mohammad
Catanzaro, Bryan
Liu, Ming-Yu
Ping, Wei
contents We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to advance understanding and reasoning over speech, environmental sounds and music. Compared to Audio Flamingo 3, AF-Next introduces: (i) a stronger foundational audio-language model that significantly improves accuracy across diverse audio understanding tasks; (ii) scalable strategies for constructing large-scale audio understanding and reasoning data beyond existing academic benchmarks; (iii) support for long and complex audio inputs up to 30 minutes; and (iv) Temporal Audio Chain-of-Thought, a new reasoning paradigm that explicitly grounds intermediate reasoning steps to timestamps in long audio, enabling fine-grained temporal alignment and improved interpretability. To enable these capabilities, we first conduct a systematic analysis of Audio Flamingo 3 to identify key gaps in audio understanding and reasoning. We then curate and scale new large-scale datasets totaling over 1 million hours to address these limitations and expand the existing AudioSkills-XL, LongAudio-XL, AF-Think and AF-Chat datasets. AF-Next is trained using a curriculum-based strategy spanning pre-training, mid-training and post-training stages. Extensive experiments across 20 audio understanding and reasoning benchmarks, including challenging long-audio tasks, show that AF-Next outperforms similarly sized open models by large margins and remains highly competitive with and sometimes surpasses, much larger open-weight and closed models. Beyond benchmark performance, AF-Next exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. In addition to all data, code and methods, we open-source 3 variants of AF-Next, including AF-Next-Instruct, AF-Next-Think and AF-Next-Captioner.
format Preprint
id arxiv_https___arxiv_org_abs_2604_10905
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
Ghosh, Sreyan
Goel, Arushi
Jayakumar, Kaousheik
Koroshinadze, Lasha
Anand, Nishit
Kong, Zhifeng
Gururani, Siddharth
Lee, Sang-gil
Kim, Jaehyeon
Aljafari, Aya
Yang, Chao-Han Huck
Kim, Sungwon
Duraiswami, Ramani
Manocha, Dinesh
Shoeybi, Mohammad
Catanzaro, Bryan
Liu, Ming-Yu
Ping, Wei
Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
We present Audio Flamingo Next (AF-Next), the next-generation and most capable large audio-language model in the Audio Flamingo series, designed to advance understanding and reasoning over speech, environmental sounds and music. Compared to Audio Flamingo 3, AF-Next introduces: (i) a stronger foundational audio-language model that significantly improves accuracy across diverse audio understanding tasks; (ii) scalable strategies for constructing large-scale audio understanding and reasoning data beyond existing academic benchmarks; (iii) support for long and complex audio inputs up to 30 minutes; and (iv) Temporal Audio Chain-of-Thought, a new reasoning paradigm that explicitly grounds intermediate reasoning steps to timestamps in long audio, enabling fine-grained temporal alignment and improved interpretability. To enable these capabilities, we first conduct a systematic analysis of Audio Flamingo 3 to identify key gaps in audio understanding and reasoning. We then curate and scale new large-scale datasets totaling over 1 million hours to address these limitations and expand the existing AudioSkills-XL, LongAudio-XL, AF-Think and AF-Chat datasets. AF-Next is trained using a curriculum-based strategy spanning pre-training, mid-training and post-training stages. Extensive experiments across 20 audio understanding and reasoning benchmarks, including challenging long-audio tasks, show that AF-Next outperforms similarly sized open models by large margins and remains highly competitive with and sometimes surpasses, much larger open-weight and closed models. Beyond benchmark performance, AF-Next exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability. In addition to all data, code and methods, we open-source 3 variants of AF-Next, including AF-Next-Instruct, AF-Next-Think and AF-Next-Captioner.
title Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
topic Sound
Artificial Intelligence
Computation and Language
Audio and Speech Processing
url https://arxiv.org/abs/2604.10905