MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Niu, Yadong, Wang, Tianzi, Dinkel, Heinrich, Sun, Xingwei, Zhou, Jiahao, Li, Gang, Liu, Jizhong, Liu, Xunying, Zhang, Junbo, Luan, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding
von: Niu, Yadong, et al.
Veröffentlicht: (2026)
von: Niu, Yadong, et al.
Veröffentlicht: (2026)
Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
von: Li, Gang, et al.
Veröffentlicht: (2025)
von: Li, Gang, et al.
Veröffentlicht: (2025)
GLAP: General contrastive audio-text pretraining across domains and languages
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
DashengTokenizer: One layer is enough for unified audio understanding and generation
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders
von: Sun, Xingwei, et al.
Veröffentlicht: (2025)
von: Sun, Xingwei, et al.
Veröffentlicht: (2025)
X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding
von: Liu, Jizhong, et al.
Veröffentlicht: (2024)
von: Liu, Jizhong, et al.
Veröffentlicht: (2024)
Bridging Language Gaps in Audio-Text Retrieval
von: Yan, Zhiyong, et al.
Veröffentlicht: (2024)
von: Yan, Zhiyong, et al.
Veröffentlicht: (2024)
The ICME 2025 Audio Encoder Capability Challenge
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
von: Zhang, Junbo, et al.
Veröffentlicht: (2025)
Streaming Audio Transformers for Online Audio Tagging
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2023)
Enhancing Pre-trained ASR System Fine-tuning for Dysarthric Speech Recognition using Adversarial Data Augmentation
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
von: Wang, Huimeng, et al.
Veröffentlicht: (2024)
Unfolding A Few Structures for The Many: Memory-Efficient Compression of Conformer and Speech Foundation Models
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
Joint Speaker Features Learning for Audio-visual Multichannel Speech Separation and Recognition
von: Li, Guinan, et al.
Veröffentlicht: (2024)
von: Li, Guinan, et al.
Veröffentlicht: (2024)
Scaling up masked audio encoder learning for general audio classification
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
Towards One-bit ASR: Extremely Low-bit Conformer Quantization Using Co-training and Stochastic Precision
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2025)
UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion
von: Li, Zhaoqing, et al.
Veröffentlicht: (2026)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2026)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
Variational Auto-Encoder Based Variability Encoding for Dysarthric Speech Recognition
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
von: Xie, Xurong, et al.
Veröffentlicht: (2022)
On-the-fly Routing for Zero-shot MoE Speaker Adaptation of Speech Foundation Models for Dysarthric Speech Recognition
von: HU, Shujie, et al.
Veröffentlicht: (2025)
von: HU, Shujie, et al.
Veröffentlicht: (2025)
Structured Speaker-Deficiency Adaptation of Foundation Models for Dysarthric and Elderly Speech Recognition
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
AudioGenie-Reasoner: A Training-Free Multi-Agent Framework for Coarse-to-Fine Audio Deep Reasoning
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
HearFit+: Personalized Fitness Monitoring via Audio Signals on Smart Speakers
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
Pengi: An Audio Language Model for Audio Tasks
von: Deshmukh, Soham, et al.
Veröffentlicht: (2023)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2023)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
von: Chen, Youjun, et al.
Veröffentlicht: (2025)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
Audio Entailment: Assessing Deductive Reasoning for Audio Understanding
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2024)
PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio
von: Chen, Yuanjian, et al.
Veröffentlicht: (2026)
von: Chen, Yuanjian, et al.
Veröffentlicht: (2026)
TTA-Bench: A Comprehensive Benchmark for Evaluating Text-to-Audio Models
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
QuarkAudio Technical Report
von: Liu, Chengwei, et al.
Veröffentlicht: (2025)
von: Liu, Chengwei, et al.
Veröffentlicht: (2025)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2024)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
von: Cui, Mingyu, et al.
Veröffentlicht: (2024)
One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
Efficient Adapter Tuning for Joint Singing Voice Beat and Downbeat Tracking with Self-supervised Learning Features
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
von: Deng, Jiajun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025) -
ACAVCaps: Enabling large-scale training for fine-grained and diverse audio understanding
von: Niu, Yadong, et al.
Veröffentlicht: (2026) -
Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
von: Li, Gang, et al.
Veröffentlicht: (2025) -
GLAP: General contrastive audio-text pretraining across domains and languages
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025) -
DashengTokenizer: One layer is enough for unified audio understanding and generation
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)