Towards audio language modeling -- an overview
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Haibin, Chen, Xuanjun, Lin, Yi-Cheng, Chang, Kai-wei, Chung, Ho-Lam, Liu, Alexander H., Lee, Hung-yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Singing Voice Graph Modeling for SingFake Detection
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
Neural Codec-based Adversarial Sample Detection for Speaker Verification
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
CodecFake: Enhancing Anti-Spoofing Models Against Deepfake Audios from Codec-Based Speech Synthesis Systems
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Towards Generalized Source Tracing for Codec-Based Deepfake Speech
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
Joint Fullband-Subband Modeling for High-Resolution SingFake Detection
von: Chen, Xuanjun, et al.
Veröffentlicht: (2026)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2026)
EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
DFADD: The Diffusion and Flow-Matching Based Audio Deepfake Dataset
von: Du, Jiawei, et al.
Veröffentlicht: (2024)
von: Du, Jiawei, et al.
Veröffentlicht: (2024)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
How Does Instrumental Music Help SingFake Detection?
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
Speaker anonymization using neural audio codec language models
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
von: Panariello, Michele, et al.
Veröffentlicht: (2023)
Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
von: Lin, Tsung-En, et al.
Veröffentlicht: (2025)
von: Lin, Tsung-En, et al.
Veröffentlicht: (2025)
Multimodal Transformer Distillation for Audio-Visual Synchronization
von: Chen, Xuanjun, et al.
Veröffentlicht: (2022)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2022)
Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
von: Ren, Wenze, et al.
Veröffentlicht: (2024)
Mitigating Subgroup Disparities in Multi-Label Speech Emotion Recognition: A Pseudo-Labeling and Unsupervised Learning Approach
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2025)
Towards General-Purpose Text-Instruction-Guided Voice Conversion
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2023)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2023)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
EMO-SUPERB: An In-depth Look at Speech Emotion Recognition
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
von: Wu, Haibin, et al.
Veröffentlicht: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Guan-Ting, et al.
Veröffentlicht: (2024)
MI-Fuse: Label Fusion for Unsupervised Domain Adaptation with Closed-Source Large-Audio Language Model
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2025)
von: Huang, Hsiao-Ying, et al.
Veröffentlicht: (2025)
Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
A Preliminary Exploration with GPT-4o Voice Mode
von: Lin, Yu-Xiang, et al.
Veröffentlicht: (2025)
von: Lin, Yu-Xiang, et al.
Veröffentlicht: (2025)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
von: Lin, Hsi-Che, et al.
Veröffentlicht: (2024)
Improving the Adversarial Robustness for Speaker Verification by Self-Supervised Learning
von: Wu, Haibin, et al.
Veröffentlicht: (2021)
von: Wu, Haibin, et al.
Veröffentlicht: (2021)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
von: Takano, Taisei, et al.
Veröffentlicht: (2025)
Can Large Audio-Language Models Truly Hear? Tackling Hallucinations with Multi-Task Assessment and Stepwise Audio Reasoning
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
Teaching Audio-Aware Large Language Models What Does Not Hear: Mitigating Hallucinations through Synthesized Negative Samples
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
von: Lin, Yi-Cheng, et al.
Veröffentlicht: (2026)
How Contrastive Decoding Enhances Large Audio Language Models?
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2026)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2026)
SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
ASTAR-NTU solution to AudioMOS Challenge 2025 Track1
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2025)
Mellow: a small audio language model for reasoning
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
Are audio DeepFake detection models polyglots?
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
von: Marek, Bartłomiej, et al.
Veröffentlicht: (2024)
Parallel Synthesis for Autoregressive Speech Generation
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2022)
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
von: Wu, Haibin, et al.
Veröffentlicht: (2024) -
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
von: Wu, Haibin, et al.
Veröffentlicht: (2024) -
Singing Voice Graph Modeling for SingFake Detection
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024) -
Neural Codec-based Adversarial Sample Detection for Speaker Verification
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024) -
Codec-Based Deepfake Source Tracing via Neural Audio Codec Taxonomy
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)