ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Li, Bohan, Huang, Wenbin, Qiu, Yuhang, Guo, Yiwei, Wang, Hankun, Li, Zhihan, Peng, Jing, Ma, Ziyang, Chen, Xie, Yu, Kai |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
par: Guo, Yiwei, et autres
Publié: (2025)
par: Guo, Yiwei, et autres
Publié: (2025)
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
par: Wang, Haoran, et autres
Publié: (2025)
par: Wang, Haoran, et autres
Publié: (2025)
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
par: Wang, Hankun, et autres
Publié: (2025)
par: Wang, Hankun, et autres
Publié: (2025)
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
par: Li, Bohan, et autres
Publié: (2025)
par: Li, Bohan, et autres
Publié: (2025)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
par: Wang, Hankun, et autres
Publié: (2024)
par: Wang, Hankun, et autres
Publié: (2024)
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
par: Li, Bohan, et autres
Publié: (2026)
par: Li, Bohan, et autres
Publié: (2026)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech
par: Chen, Huakang, et autres
Publié: (2026)
par: Chen, Huakang, et autres
Publié: (2026)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
par: Li, Bohan, et autres
Publié: (2024)
par: Li, Bohan, et autres
Publié: (2024)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
par: Wang, Hankun, et autres
Publié: (2024)
par: Wang, Hankun, et autres
Publié: (2024)
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
par: Wang, Hui, et autres
Publié: (2025)
par: Wang, Hui, et autres
Publié: (2025)
Towards Weakly Supervised Text-to-Audio Grounding
par: Xu, Xuenan, et autres
Publié: (2024)
par: Xu, Xuenan, et autres
Publié: (2024)
Recent Advances in Discrete Speech Tokens: A Review
par: Guo, Yiwei, et autres
Publié: (2025)
par: Guo, Yiwei, et autres
Publié: (2025)
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
par: Li, Baihan, et autres
Publié: (2024)
par: Li, Baihan, et autres
Publié: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
par: Du, Chenpeng, et autres
Publié: (2024)
par: Du, Chenpeng, et autres
Publié: (2024)
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs
par: Chen, Wenxi, et autres
Publié: (2024)
par: Chen, Wenxi, et autres
Publié: (2024)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
par: Zhang, Hanglei, et autres
Publié: (2025)
par: Zhang, Hanglei, et autres
Publié: (2025)
PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio
par: Chen, Yuanjian, et autres
Publié: (2026)
par: Chen, Yuanjian, et autres
Publié: (2026)
AudioBench: A Universal Benchmark for Audio Large Language Models
par: Wang, Bin, et autres
Publié: (2024)
par: Wang, Bin, et autres
Publié: (2024)
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
par: Li, Bohan, et autres
Publié: (2024)
par: Li, Bohan, et autres
Publié: (2024)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
par: Wang, Yucheng, et autres
Publié: (2026)
par: Wang, Yucheng, et autres
Publié: (2026)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
par: Du, Chenpeng, et autres
Publié: (2022)
par: Du, Chenpeng, et autres
Publié: (2022)
NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization
par: Niu, Zhikang, et autres
Publié: (2024)
par: Niu, Zhikang, et autres
Publié: (2024)
LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models
par: Zhao, Zhiyuan, et autres
Publié: (2026)
par: Zhao, Zhiyuan, et autres
Publié: (2026)
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
par: Deng, Ruifan, et autres
Publié: (2025)
par: Deng, Ruifan, et autres
Publié: (2025)
From Contrast to Commonality: Audio Commonality Captioning for Enhanced Audio-Text Cross-modal Understanding in Multimodal LLMs
par: Jia, Yuhang, et autres
Publié: (2025)
par: Jia, Yuhang, et autres
Publié: (2025)
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
par: Lin, Jingru, et autres
Publié: (2026)
par: Lin, Jingru, et autres
Publié: (2026)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
par: Hou, Yixuan, et autres
Publié: (2025)
par: Hou, Yixuan, et autres
Publié: (2025)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
par: Li, Yangze, et autres
Publié: (2024)
par: Li, Yangze, et autres
Publié: (2024)
AeroGPT: Leveraging Large-Scale Audio Model for Aero-Engine Bearing Fault Diagnosis
par: Liu, Jiale, et autres
Publié: (2025)
par: Liu, Jiale, et autres
Publié: (2025)
Audio-CoT: Exploring Chain-of-Thought Reasoning in Large Audio Language Model
par: Ma, Ziyang, et autres
Publié: (2025)
par: Ma, Ziyang, et autres
Publié: (2025)
SpeechEditBench: A Bilingual Multi-Attribute Benchmark for Instruction-Guided Speech Editing
par: Zhang, Hanlin, et autres
Publié: (2026)
par: Zhang, Hanlin, et autres
Publié: (2026)
Acoustic BPE for Speech Generation with Discrete Tokens
par: Shen, Feiyu, et autres
Publié: (2023)
par: Shen, Feiyu, et autres
Publié: (2023)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
par: Yang, Dongchao, et autres
Publié: (2024)
par: Yang, Dongchao, et autres
Publié: (2024)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
par: Li, Junjie, et autres
Publié: (2023)
par: Li, Junjie, et autres
Publié: (2023)
AudioMarkBench: Benchmarking Robustness of Audio Watermarking
par: Liu, Hongbin, et autres
Publié: (2024)
par: Liu, Hongbin, et autres
Publié: (2024)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
par: Xi, Yu, et autres
Publié: (2024)
par: Xi, Yu, et autres
Publié: (2024)
FoleyBench: A Benchmark For Video-to-Audio Models
par: Dixit, Satvik, et autres
Publié: (2025)
par: Dixit, Satvik, et autres
Publié: (2025)
STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
par: Liu, Zihan, et autres
Publié: (2025)
par: Liu, Zihan, et autres
Publié: (2025)
Documents similaires
-
AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
par: Guo, Yiwei, et autres
Publié: (2025) -
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
par: Wang, Haoran, et autres
Publié: (2025) -
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
par: Wang, Hankun, et autres
Publié: (2025) -
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
par: Li, Bohan, et autres
Publié: (2025) -
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
par: Guo, Yiwei, et autres
Publié: (2024)