A Large-Scale Evaluation of Speech Foundation Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Shu-wen, Chang, Heng-Jui, Huang, Zili, Liu, Andy T., Lai, Cheng-I, Wu, Haibin, Shi, Jiatong, Chang, Xuankai, Tsai, Hsiang-Sheng, Huang, Wen-Chin, Feng, Tzu-hsun, Chi, Po-Han, Lin, Yist Y., Chuang, Yung-Sung, Huang, Tzu-Hsien, Tseng, Wei-Cheng, Lakhotia, Kushal, Li, Shang-Wen, Mohamed, Abdelrahman, Watanabe, Shinji, Lee, Hung-yi |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
par: Hsu, Tzu-wen, et autres
Publié: (2025)
par: Hsu, Tzu-wen, et autres
Publié: (2025)
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
par: Shi, Jiatong, et autres
Publié: (2023)
par: Shi, Jiatong, et autres
Publié: (2023)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
par: Lin, Tzu-Quan, et autres
Publié: (2022)
par: Lin, Tzu-Quan, et autres
Publié: (2022)
Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
par: Shi, Jiatong, et autres
Publié: (2023)
par: Shi, Jiatong, et autres
Publié: (2023)
Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
par: Lin, Yi-Cheng, et autres
Publié: (2025)
par: Lin, Yi-Cheng, et autres
Publié: (2025)
LV-CTC: Non-autoregressive ASR with CTC and latent variable models
par: Fujita, Yuya, et autres
Publié: (2024)
par: Fujita, Yuya, et autres
Publié: (2024)
Real-Time Minimum-Energy Operating-Point Tracking for Battery-Powered Micro DC Motors Under Dynamically Variable Loading
par: Huang, Tzu-Hsiang, et autres
Publié: (2026)
par: Huang, Tzu-Hsiang, et autres
Publié: (2026)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
par: Chou, Huang-Cheng, et autres
Publié: (2024)
par: Chou, Huang-Cheng, et autres
Publié: (2024)
How Contrastive Decoding Enhances Large Audio Language Models?
par: Lin, Tzu-Quan, et autres
Publié: (2026)
par: Lin, Tzu-Quan, et autres
Publié: (2026)
A Dataset and Baselines for Measuring and Predicting the Music Piece Memorability
par: Tseng, Li-Yang, et autres
Publié: (2024)
par: Tseng, Li-Yang, et autres
Publié: (2024)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
par: Wu, Haibin, et autres
Publié: (2024)
par: Wu, Haibin, et autres
Publié: (2024)
Analytical Model of Groundwater Flow in a Rectangular Domain for Spatiotemporally Distributed Recharge
par: Ping‐Cheng Hsieh, et autres
Publié: (2025)
par: Ping‐Cheng Hsieh, et autres
Publié: (2025)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
par: Wang, Shih-heng, et autres
Publié: (2024)
par: Wang, Shih-heng, et autres
Publié: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
par: Guo, Pengcheng, et autres
Publié: (2024)
par: Guo, Pengcheng, et autres
Publié: (2024)
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
par: Lin, Yi-Cheng, et autres
Publié: (2024)
par: Lin, Yi-Cheng, et autres
Publié: (2024)
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
par: Tsai, Yun-Shao, et autres
Publié: (2026)
par: Tsai, Yun-Shao, et autres
Publié: (2026)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
par: Lin, Tzu-Quan, et autres
Publié: (2025)
par: Lin, Tzu-Quan, et autres
Publié: (2025)
AI-based CSI Feedback with Digital Twins: Real-World Validation and Insights
par: Huang, Tzu-Hao, et autres
Publié: (2025)
par: Huang, Tzu-Hao, et autres
Publié: (2025)
The Interspeech 2024 Challenge on Speech Processing Using Discrete Units
par: Chang, Xuankai, et autres
Publié: (2024)
par: Chang, Xuankai, et autres
Publié: (2024)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
par: Shi, Jiatong, et autres
Publié: (2024)
par: Shi, Jiatong, et autres
Publié: (2024)
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
par: Cornell, Samuele, et autres
Publié: (2024)
par: Cornell, Samuele, et autres
Publié: (2024)
How to Learn a New Language? An Efficient Solution for Self-Supervised Learning Models Unseen Languages Adaption in Low-Resource Scenario
par: Wang, Shih-Heng, et autres
Publié: (2024)
par: Wang, Shih-Heng, et autres
Publié: (2024)
Learning Displacement-Aware WiFi Representations for Weakly Supervised Relative Localization
par: Wei, Tzu-Ti, et autres
Publié: (2026)
par: Wei, Tzu-Ti, et autres
Publié: (2026)
ART: Artifact Removal Transformer for Reconstructing Noise-Free Multichannel Electroencephalographic Signals
par: Chuang, Chun-Hsiang, et autres
Publié: (2024)
par: Chuang, Chun-Hsiang, et autres
Publié: (2024)
Comparing the Effects of Different Dielectric Materials on an Atmospheric Pressure Plasma Jet by Experiments and Simulations
par: Po‐Chun Huang, et autres
Publié: (2024)
par: Po‐Chun Huang, et autres
Publié: (2024)
On the social bias of speech self-supervised models
par: Lin, Yi-Cheng, et autres
Publié: (2024)
par: Lin, Yi-Cheng, et autres
Publié: (2024)
Full-Duplex-Bench-v2: A Multi-Turn Evaluation Framework for Duplex Dialogue Systems with an Automated Examiner
par: Lin, Guan-Ting, et autres
Publié: (2025)
par: Lin, Guan-Ting, et autres
Publié: (2025)
Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
par: Tseng, Liang-Hsuan, et autres
Publié: (2024)
par: Tseng, Liang-Hsuan, et autres
Publié: (2024)
Enhancing Multilingual ASR for Unseen Languages via Language Embedding Modeling
par: Huang, Shao-Syuan, et autres
Publié: (2024)
par: Huang, Shao-Syuan, et autres
Publié: (2024)
MelHuBERT: A simplified HuBERT on Mel spectrograms
par: Lin, Tzu-Quan, et autres
Publié: (2022)
par: Lin, Tzu-Quan, et autres
Publié: (2022)
DAISY: Data Adaptive Self-Supervised Early Exit for Speech Representation Models
par: Lin, Tzu-Quan, et autres
Publié: (2024)
par: Lin, Tzu-Quan, et autres
Publié: (2024)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
par: Lin, Yi-Cheng, et autres
Publié: (2025)
par: Lin, Yi-Cheng, et autres
Publié: (2025)
Voxtlm: unified decoder-only models for consolidating speech recognition/synthesis and speech/text continuation tasks
par: Maiti, Soumi, et autres
Publié: (2023)
par: Maiti, Soumi, et autres
Publié: (2023)
Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks
par: Wang, Shih-Heng, et autres
Publié: (2026)
par: Wang, Shih-Heng, et autres
Publié: (2026)
How Does Instrumental Music Help SingFake Detection?
par: Chen, Xuanjun, et autres
Publié: (2025)
par: Chen, Xuanjun, et autres
Publié: (2025)
Towards Robust Speech Representation Learning for Thousands of Languages
par: Chen, William, et autres
Publié: (2024)
par: Chen, William, et autres
Publié: (2024)
Joint Fullband-Subband Modeling for High-Resolution SingFake Detection
par: Chen, Xuanjun, et autres
Publié: (2026)
par: Chen, Xuanjun, et autres
Publié: (2026)
End-User-Centric Collaborative MIMO: Performance Analysis and Proof of Concept
par: Wen, Chao-Kai, et autres
Publié: (2024)
par: Wen, Chao-Kai, et autres
Publié: (2024)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
par: Yang, Dongchao, et autres
Publié: (2023)
par: Yang, Dongchao, et autres
Publié: (2023)
Quantum Information-Empowered Graph Neural Network for Hyperspectral Change Detection
par: Lin, Chia-Hsiang, et autres
Publié: (2024)
par: Lin, Chia-Hsiang, et autres
Publié: (2024)
Documents similaires
-
Reducing Object Hallucination in Large Audio-Language Models via Audio-Aware Decoding
par: Hsu, Tzu-wen, et autres
Publié: (2025) -
ML-SUPERB: Multilingual Speech Universal PERformance Benchmark
par: Shi, Jiatong, et autres
Publié: (2023) -
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
par: Lin, Tzu-Quan, et autres
Publié: (2022) -
Findings of the 2023 ML-SUPERB Challenge: Pre-Training and Evaluation over More Languages and Beyond
par: Shi, Jiatong, et autres
Publié: (2023) -
Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
par: Lin, Yi-Cheng, et autres
Publié: (2025)