Eureka-Audio: Triggering Audio Intelligence in Compact Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Dan, Lei, Yishu, Hu, Jing, He, Shuwei, Deng, Songhe, Luo, Xianlong, Zhu, Danxiang, Feng, Shikun, Liu, Rui, He, Jingzhou, Sun, Yu, Wu, Hua, Wang, Haifeng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MoE Adapter for Large Audio Language Models: Sparsity, Disentanglement, and Gradient-Conflict-Free
by: Lei, Yishu, et al.
Published: (2026)
by: Lei, Yishu, et al.
Published: (2026)
CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
by: Hu, Jing, et al.
Published: (2026)
by: Hu, Jing, et al.
Published: (2026)
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
by: Lin, Zihan, et al.
Published: (2026)
by: Lin, Zihan, et al.
Published: (2026)
AudioFab: Building A General and Intelligent Audio Factory through Tool Learning
by: Zhu, Cheng, et al.
Published: (2025)
by: Zhu, Cheng, et al.
Published: (2025)
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
by: Zhu, Xinfa, et al.
Published: (2025)
by: Zhu, Xinfa, et al.
Published: (2025)
AudioKV: KV Cache Eviction in Efficient Large Audio Language Models
by: Wang, Yuxuan, et al.
Published: (2026)
by: Wang, Yuxuan, et al.
Published: (2026)
Audio-Visual Speech Enhancement: Architectural Design and Deployment Strategies
by: Hamadouche, Anis, et al.
Published: (2025)
by: Hamadouche, Anis, et al.
Published: (2025)
AudioRWKV: Efficient and Stable Bidirectional RWKV for Audio Pattern Recognition
by: Xiong, Jiayu, et al.
Published: (2025)
by: Xiong, Jiayu, et al.
Published: (2025)
Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
by: He, Shuwei, et al.
Published: (2024)
by: He, Shuwei, et al.
Published: (2024)
AudioSpa: Spatializing Sound Events with Text
by: Feng, Linfeng, et al.
Published: (2025)
by: Feng, Linfeng, et al.
Published: (2025)
HeadRouter: Dynamic Head-Weight Routing for Task-Adaptive Audio Token Pruning in Large Audio Language Models
by: He, Peize, et al.
Published: (2026)
by: He, Peize, et al.
Published: (2026)
DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners
by: Luo, Xiaoxue, et al.
Published: (2025)
by: Luo, Xiaoxue, et al.
Published: (2025)
ChronosAudio: A Comprehensive Long-Audio Benchmark for Evaluating Audio-Large Language Models
by: Luo, Kaiwen, et al.
Published: (2026)
by: Luo, Kaiwen, et al.
Published: (2026)
TinyMU: A Compact Audio-Language Model for Music Understanding
by: Li, Xiquan, et al.
Published: (2026)
by: Li, Xiquan, et al.
Published: (2026)
SemanticAudio: Audio Generation and Editing in Semantic Space
by: Dai, Zheqi, et al.
Published: (2026)
by: Dai, Zheqi, et al.
Published: (2026)
AudioGAN: A Compact and Efficient Framework for Real-Time High-Fidelity Text-to-Audio Generation
by: Chung, HaeChun
Published: (2025)
by: Chung, HaeChun
Published: (2025)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
by: Huang, Zhiqi, et al.
Published: (2024)
by: Huang, Zhiqi, et al.
Published: (2024)
Discrete Audio Representations for Automated Audio Captioning
by: Tian, Jingguang, et al.
Published: (2025)
by: Tian, Jingguang, et al.
Published: (2025)
Audio Mamba: Pretrained Audio State Space Model For Audio Tagging
by: Lin, Jiaju, et al.
Published: (2024)
by: Lin, Jiaju, et al.
Published: (2024)
Precise and Simple Audio-to-Score Alignment
by: Peter, Silvan, et al.
Published: (2026)
by: Peter, Silvan, et al.
Published: (2026)
AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models
by: Zhao, Junqi, et al.
Published: (2025)
by: Zhao, Junqi, et al.
Published: (2025)
A Streamable Neural Audio Codec with Residual Scalar-Vector Quantization for Real-Time Communication
by: Jiang, Xiao-Hang, et al.
Published: (2025)
by: Jiang, Xiao-Hang, et al.
Published: (2025)
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation
by: Xie, Tianxin, et al.
Published: (2025)
by: Xie, Tianxin, et al.
Published: (2025)
Whisper-AuT: Domain-Adapted Audio Encoder for Efficient Audio-LLM Training
by: Qiu, Jielin, et al.
Published: (2026)
by: Qiu, Jielin, et al.
Published: (2026)
AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
by: He, Peize, et al.
Published: (2025)
by: He, Peize, et al.
Published: (2025)
AudioSet-R: A Refined AudioSet with Multi-Stage LLM Label Reannotation
by: Sun, Yulin, et al.
Published: (2025)
by: Sun, Yulin, et al.
Published: (2025)
TAME: Temporal Audio-based Mamba for Enhanced Drone Trajectory Estimation and Classification
by: Xiao, Zhenyuan, et al.
Published: (2024)
by: Xiao, Zhenyuan, et al.
Published: (2024)
Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
by: Yadav, Sarthak, et al.
Published: (2024)
by: Yadav, Sarthak, et al.
Published: (2024)
Resonate: Reinforcing Text-to-Audio Generation via Online Feedback from Large Audio Language Models
by: Li, Xiquan, et al.
Published: (2026)
by: Li, Xiquan, et al.
Published: (2026)
AudioToolAgent: An Agentic Framework for Audio-Language Models
by: Wijngaard, Gijs, et al.
Published: (2025)
by: Wijngaard, Gijs, et al.
Published: (2025)
MUKA: Multi Kernel Audio Adaptation Of Audio-Language Models
by: Bensaid, Reda, et al.
Published: (2026)
by: Bensaid, Reda, et al.
Published: (2026)
UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
by: Yang, Dongchao, et al.
Published: (2026)
by: Yang, Dongchao, et al.
Published: (2026)
AudioMoG: Guiding Audio Generation with Mixture-of-Guidance
by: Wang, Junyou, et al.
Published: (2025)
by: Wang, Junyou, et al.
Published: (2025)
Retrieval-Augmented Audio Deepfake Detection
by: Kang, Zuheng, et al.
Published: (2024)
by: Kang, Zuheng, et al.
Published: (2024)
The Sonar Moment: Benchmarking Audio-Language Models in Audio Geo-Localization
by: Zhang, Ruixing, et al.
Published: (2026)
by: Zhang, Ruixing, et al.
Published: (2026)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
by: Wang, Yucheng, et al.
Published: (2026)
by: Wang, Yucheng, et al.
Published: (2026)
FoleySpace: Vision-Aligned Binaural Spatial Audio Generation
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
Manipulated Regions Localization For Partially Deepfake Audio: A Survey
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing
by: Chen, William, et al.
Published: (2026)
by: Chen, William, et al.
Published: (2026)
Similar Items
-
MoE Adapter for Large Audio Language Models: Sparsity, Disentanglement, and Gradient-Conflict-Free
by: Lei, Yishu, et al.
Published: (2026) -
CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation
by: Hu, Jing, et al.
Published: (2026) -
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
by: Lin, Zihan, et al.
Published: (2026) -
AudioFab: Building A General and Intelligent Audio Factory through Tool Learning
by: Zhu, Cheng, et al.
Published: (2025) -
CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions
by: Zhu, Xinfa, et al.
Published: (2025)