AHAMask: Reliable Task Specification for Large Audio Language Models without Instructions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Guo, Yiwei, Li, Bohan, Wang, Hankun, Li, Zhihan, Wang, Shuai, Chen, Xie, Yu, Kai |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
par: Li, Bohan, et autres
Publié: (2025)
par: Li, Bohan, et autres
Publié: (2025)
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
par: Wang, Haoran, et autres
Publié: (2025)
par: Wang, Haoran, et autres
Publié: (2025)
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
par: Wang, Hankun, et autres
Publié: (2025)
par: Wang, Hankun, et autres
Publié: (2025)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
par: Wang, Hankun, et autres
Publié: (2024)
par: Wang, Hankun, et autres
Publié: (2024)
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
par: Li, Bohan, et autres
Publié: (2025)
par: Li, Bohan, et autres
Publié: (2025)
Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
par: Wang, Hankun, et autres
Publié: (2024)
par: Wang, Hankun, et autres
Publié: (2024)
LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
par: Guo, Yiwei, et autres
Publié: (2024)
par: Guo, Yiwei, et autres
Publié: (2024)
HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
par: Li, Bohan, et autres
Publié: (2026)
par: Li, Bohan, et autres
Publié: (2026)
Fast and High-Quality Auto-Regressive Speech Synthesis via Speculative Decoding
par: Li, Bohan, et autres
Publié: (2024)
par: Li, Bohan, et autres
Publié: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
par: Du, Chenpeng, et autres
Publié: (2024)
par: Du, Chenpeng, et autres
Publié: (2024)
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
par: Li, Bohan, et autres
Publié: (2024)
par: Li, Bohan, et autres
Publié: (2024)
Can Audio Large Language Models Verify Speaker Identity?
par: Ren, Yiming, et autres
Publié: (2025)
par: Ren, Yiming, et autres
Publié: (2025)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
par: Yang, Dongchao, et autres
Publié: (2024)
par: Yang, Dongchao, et autres
Publié: (2024)
Recent Advances in Discrete Speech Tokens: A Review
par: Guo, Yiwei, et autres
Publié: (2025)
par: Guo, Yiwei, et autres
Publié: (2025)
Pengi: An Audio Language Model for Audio Tasks
par: Deshmukh, Soham, et autres
Publié: (2023)
par: Deshmukh, Soham, et autres
Publié: (2023)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
par: Wang, Yucheng, et autres
Publié: (2026)
par: Wang, Yucheng, et autres
Publié: (2026)
DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
par: Li, Baihan, et autres
Publié: (2024)
par: Li, Baihan, et autres
Publié: (2024)
LEMAS: Large A 150K-Hour Large-scale Extensible Multilingual Audio Suite with Generative Speech Models
par: Zhao, Zhiyuan, et autres
Publié: (2026)
par: Zhao, Zhiyuan, et autres
Publié: (2026)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
par: Dinkel, Heinrich, et autres
Publié: (2026)
par: Dinkel, Heinrich, et autres
Publié: (2026)
A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
par: Li, Yangze, et autres
Publié: (2024)
par: Li, Yangze, et autres
Publié: (2024)
Can Large Language Models Understand Spatial Audio?
par: Tang, Changli, et autres
Publié: (2024)
par: Tang, Changli, et autres
Publié: (2024)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
par: Zhang, Hanglei, et autres
Publié: (2025)
par: Zhang, Hanglei, et autres
Publié: (2025)
VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature
par: Du, Chenpeng, et autres
Publié: (2022)
par: Du, Chenpeng, et autres
Publié: (2022)
Unlocking Large Audio-Language Models for Interactive Language Learning
par: Liu, Hongfu, et autres
Publié: (2026)
par: Liu, Hongfu, et autres
Publié: (2026)
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models
par: Wang, Shuiyuan, et autres
Publié: (2026)
par: Wang, Shuiyuan, et autres
Publié: (2026)
Summary on The Multilingual Conversational Speech Language Model Challenge: Datasets, Tasks, Baselines, and Methods
par: Mu, Bingshen, et autres
Publié: (2025)
par: Mu, Bingshen, et autres
Publié: (2025)
UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding
par: Du, Chenpeng, et autres
Publié: (2023)
par: Du, Chenpeng, et autres
Publié: (2023)
Contrastive Learning With Audio Discrimination For Customizable Keyword Spotting In Continuous Speech
par: Xi, Yu, et autres
Publié: (2024)
par: Xi, Yu, et autres
Publié: (2024)
DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
par: Liu, Shengqiang, et autres
Publié: (2024)
par: Liu, Shengqiang, et autres
Publié: (2024)
AeroGPT: Leveraging Large-Scale Audio Model for Aero-Engine Bearing Fault Diagnosis
par: Liu, Jiale, et autres
Publié: (2025)
par: Liu, Jiale, et autres
Publié: (2025)
Acoustic BPE for Speech Generation with Discrete Tokens
par: Shen, Feiyu, et autres
Publié: (2023)
par: Shen, Feiyu, et autres
Publié: (2023)
MINT: Boosting Audio-Language Model via Multi-Target Pre-Training and Instruction Tuning
par: Zhao, Hang, et autres
Publié: (2024)
par: Zhao, Hang, et autres
Publié: (2024)
HearFit+: Personalized Fitness Monitoring via Audio Signals on Smart Speakers
par: Xie, Yadong, et autres
Publié: (2025)
par: Xie, Yadong, et autres
Publié: (2025)
A Survey on Speech Large Language Models for Understanding
par: Peng, Jing, et autres
Publié: (2024)
par: Peng, Jing, et autres
Publié: (2024)
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
par: Zhao, Xiaohan, et autres
Publié: (2025)
par: Zhao, Xiaohan, et autres
Publié: (2025)
Jointly Recognizing Speech and Singing Voices Based on Multi-Task Audio Source Separation
par: Bai, Ye, et autres
Publié: (2024)
par: Bai, Ye, et autres
Publié: (2024)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
par: Li, Junjie, et autres
Publié: (2023)
par: Li, Junjie, et autres
Publié: (2023)
SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor
par: Yang, Chenyu, et autres
Publié: (2024)
par: Yang, Chenyu, et autres
Publié: (2024)
Multi-Speaker Multi-Lingual VQTTS System for LIMMITS 2023 Challenge
par: Du, Chenpeng, et autres
Publié: (2023)
par: Du, Chenpeng, et autres
Publié: (2023)
Documents similaires
-
ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
par: Li, Bohan, et autres
Publié: (2025) -
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
par: Wang, Haoran, et autres
Publié: (2025) -
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
par: Wang, Hankun, et autres
Publié: (2025) -
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
par: Wang, Hankun, et autres
Publié: (2024) -
Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy
par: Li, Bohan, et autres
Publié: (2025)