Gespeichert in:
| Hauptverfasser: | Wang, Xiangbo, Jiang, Wenbin, Wang, Jin, You, Yubo, Fang, Sheng, Wen, Fei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.20362 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
von: Wang, Jin, et al.
Veröffentlicht: (2025)
von: Wang, Jin, et al.
Veröffentlicht: (2025)
TQCodec: Towards neural audio codec for high-fidelity music streaming
von: He, Lixing, et al.
Veröffentlicht: (2026)
von: He, Lixing, et al.
Veröffentlicht: (2026)
Stage-adaptive audio diffusion modeling
von: Zhang, Xuanhao, et al.
Veröffentlicht: (2026)
von: Zhang, Xuanhao, et al.
Veröffentlicht: (2026)
Making deep neural networks work for medical audio: representation, compression and domain adaptation
von: Onu, Charles C
Veröffentlicht: (2025)
von: Onu, Charles C
Veröffentlicht: (2025)
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
Multi-layer attentive probing improves transfer of audio representations for bioacoustics
von: Miron, Marius, et al.
Veröffentlicht: (2026)
von: Miron, Marius, et al.
Veröffentlicht: (2026)
Learning to reconstruct from saturated data: audio declipping and high-dynamic range imaging
von: Sechaud, Victor, et al.
Veröffentlicht: (2026)
von: Sechaud, Victor, et al.
Veröffentlicht: (2026)
USAT: A Universal Speaker-Adaptive Text-to-Speech Approach
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
Keep what you need : extracting efficient subnetworks from large audio representation models
von: Genova, David, et al.
Veröffentlicht: (2025)
von: Genova, David, et al.
Veröffentlicht: (2025)
SS-DPPN: A self-supervised dual-path foundation model for the generalizable cardiac audio representation
von: Muna, Ummy Maria, et al.
Veröffentlicht: (2025)
von: Muna, Ummy Maria, et al.
Veröffentlicht: (2025)
ADIFF: Explaining audio difference using natural language
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
Mellow: a small audio language model for reasoning
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
von: Deshmukh, Soham, et al.
Veröffentlicht: (2025)
End-to-end audio-visual learning for cochlear implant sound coding simulations in noisy environments
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
Omni-CLST: Error-aware Curriculum Learning with guided Selective chain-of-Thought for audio question answering
von: Zhao, Jinghua, et al.
Veröffentlicht: (2025)
von: Zhao, Jinghua, et al.
Veröffentlicht: (2025)
MBCodec:Thorough disentangle for high-fidelity audio compression
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruonan, et al.
Veröffentlicht: (2025)
SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture
von: Sui, Kehan, et al.
Veröffentlicht: (2025)
von: Sui, Kehan, et al.
Veröffentlicht: (2025)
Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on SoC FPGA
von: Jeziorek, Kamil, et al.
Veröffentlicht: (2026)
von: Jeziorek, Kamil, et al.
Veröffentlicht: (2026)
RAS: a Reliability Oriented Metric for Automatic Speech Recognition
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
AudioMAE++: learning better masked audio representations with SwiGLU FFNs
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)
GRAM: Spatial general-purpose audio representation models for real-world applications
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
von: Yuksel, Goksenin, et al.
Veröffentlicht: (2025)
Beyond saliency: enhancing explanation of speech emotion recognition with expert-referenced acoustic cues
von: Nasr, Seham, et al.
Veröffentlicht: (2025)
von: Nasr, Seham, et al.
Veröffentlicht: (2025)
AISTAT lab system for DCASE2025 Task6: Language-based audio retrieval
von: Kim, Hyun Jun, et al.
Veröffentlicht: (2025)
von: Kim, Hyun Jun, et al.
Veröffentlicht: (2025)
Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes
von: Roman, Adrian S., et al.
Veröffentlicht: (2024)
von: Roman, Adrian S., et al.
Veröffentlicht: (2024)
Recomposer: Event-roll-guided generative audio editing
von: Ellis, Daniel P. W., et al.
Veröffentlicht: (2025)
von: Ellis, Daniel P. W., et al.
Veröffentlicht: (2025)
AST: Adaptive, Seamless, and Training-Free Precise Speech Editing
von: Lv, Sihan, et al.
Veröffentlicht: (2026)
von: Lv, Sihan, et al.
Veröffentlicht: (2026)
Where are we in audio deepfake detection? A systematic analysis over generative and detection models
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
DualSpec: Text-to-spatial-audio Generation via Dual-Spectrogram Guided Diffusion Model
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
von: Zhao, Lei, et al.
Veröffentlicht: (2025)
Forensic deepfake audio detection using segmental speech features
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
von: Yang, Tianle, et al.
Veröffentlicht: (2025)
NatureLM-audio: an Audio-Language Foundation Model for Bioacoustics
von: Robinson, David, et al.
Veröffentlicht: (2024)
von: Robinson, David, et al.
Veröffentlicht: (2024)
A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
von: Olvera, Michel, et al.
Veröffentlicht: (2024)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
von: Chen, Jinming, et al.
Veröffentlicht: (2024)
von: Chen, Jinming, et al.
Veröffentlicht: (2024)
Keyword spotting using convolutional neural network for speech recognition in Hindi
von: Bharti, Saru, et al.
Veröffentlicht: (2026)
von: Bharti, Saru, et al.
Veröffentlicht: (2026)
Exploring bat song syllable representations in self-supervised audio encoders
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
TimberAgent: Gram-Guided Retrieval for Executable Music Effect Control
von: He, Shihao, et al.
Veröffentlicht: (2026)
von: He, Shihao, et al.
Veröffentlicht: (2026)
Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
A unified multichannel far-field speech recognition system: combining neural beamforming with attention based end-to-end model
von: Zhao, Dongdi, et al.
Veröffentlicht: (2024)
von: Zhao, Dongdi, et al.
Veröffentlicht: (2024)
Joint sentiment analysis of lyrics and audio in music
von: Schaab, Lea, et al.
Veröffentlicht: (2024)
von: Schaab, Lea, et al.
Veröffentlicht: (2024)
Adaptive Accompaniment with ReaLchords
von: Wu, Yusong, et al.
Veröffentlicht: (2025)
von: Wu, Yusong, et al.
Veröffentlicht: (2025)
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2026)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
von: Wang, Jin, et al.
Veröffentlicht: (2025) -
TQCodec: Towards neural audio codec for high-fidelity music streaming
von: He, Lixing, et al.
Veröffentlicht: (2026) -
Stage-adaptive audio diffusion modeling
von: Zhang, Xuanhao, et al.
Veröffentlicht: (2026) -
Making deep neural networks work for medical audio: representation, compression and domain adaptation
von: Onu, Charles C
Veröffentlicht: (2025) -
An overview of neural architectures for self-supervised audio representation learning from masked spectrograms
von: Yadav, Sarthak, et al.
Veröffentlicht: (2025)