Gespeichert in:
| Hauptverfasser: | Li, Yin, Liu, Bo, Nanadakumar, Rajalakshmi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2403.10796 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Past, Present, and Future of Spatial Audio and Room Acoustics
von: Koyama, Shoichi, et al.
Veröffentlicht: (2025)
von: Koyama, Shoichi, et al.
Veröffentlicht: (2025)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024)
Structural and Statistical Audio Texture Knowledge Distillation for Acoustic Classification
von: Ritu, Jarin, et al.
Veröffentlicht: (2025)
von: Ritu, Jarin, et al.
Veröffentlicht: (2025)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
IMPACT: Industrial Machine Perception via Acoustic Cognitive Transformer
von: Han, Changheon, et al.
Veröffentlicht: (2025)
von: Han, Changheon, et al.
Veröffentlicht: (2025)
HearSmoking: Smoking Detection in Driving Environment via Acoustic Sensing on Smartphones
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
von: Xie, Yadong, et al.
Veröffentlicht: (2025)
Beyond-Voice: Towards Continuous 3D Hand Pose Tracking on Commercial Home Assistant Devices
von: Li, Yin, et al.
Veröffentlicht: (2023)
von: Li, Yin, et al.
Veröffentlicht: (2023)
DARAS: Dynamic Audio-Room Acoustic Synthesis for Blind Room Impulse Response Estimation
von: Wang, Chunxi, et al.
Veröffentlicht: (2025)
von: Wang, Chunxi, et al.
Veröffentlicht: (2025)
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
Acoustic Disturbance Sensing Level Detection for ASD Diagnosis and Intelligibility Enhancement
von: Pillonetto, Marcelo, et al.
Veröffentlicht: (2024)
von: Pillonetto, Marcelo, et al.
Veröffentlicht: (2024)
AeroGPT: Leveraging Large-Scale Audio Model for Aero-Engine Bearing Fault Diagnosis
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
Disentangled Representation Learning for Environment-agnostic Speaker Recognition
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
von: Nam, KiHyun, et al.
Veröffentlicht: (2024)
UniSep: Universal Target Audio Separation with Language Models at Scale
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
Neural Speech Tracking in a Virtual Acoustic Environment: Audio-Visual Benefit for Unscripted Continuous Speech
von: Daeglau, Mareike, et al.
Veröffentlicht: (2025)
von: Daeglau, Mareike, et al.
Veröffentlicht: (2025)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
von: Manakul, Potsawee, et al.
Veröffentlicht: (2026)
von: Manakul, Potsawee, et al.
Veröffentlicht: (2026)
PolyBench: A Benchmark for Compositional Reasoning in Polyphonic Audio
von: Chen, Yuanjian, et al.
Veröffentlicht: (2026)
von: Chen, Yuanjian, et al.
Veröffentlicht: (2026)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
Semantic and Semiotic Interplays in Text-to-Audio AI: Exploring Cognitive Dynamics and Musical Interactions
von: Coelho, Guilherme
Veröffentlicht: (2025)
von: Coelho, Guilherme
Veröffentlicht: (2025)
AudioGenie-Reasoner: A Training-Free Multi-Agent Framework for Coarse-to-Fine Audio Deep Reasoning
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
von: Dang, Ting, et al.
Veröffentlicht: (2025)
von: Dang, Ting, et al.
Veröffentlicht: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
von: Yang, Dongchao, et al.
Veröffentlicht: (2023)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
AudioLCM: Text-to-Audio Generation with Latent Consistency Models
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
Prior-agnostic Multi-scale Contrastive Text-Audio Pre-training for Parallelized TTS Frontend Modeling
von: Wang, Quanxiu, et al.
Veröffentlicht: (2024)
von: Wang, Quanxiu, et al.
Veröffentlicht: (2024)
MiDashengLM: Efficient Audio Understanding with General Audio Captions
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2025)
SRC-gAudio: Sampling-Rate-Controlled Audio Generation
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
von: Li, Chenxing, et al.
Veröffentlicht: (2024)
FlashAudio: Rectified Flows for Fast and High-Fidelity Text-to-Audio Generation
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
von: Liu, Huadai, et al.
Veröffentlicht: (2024)
The Interspeech 2026 Audio Encoder Capability Challenge for Large Audio Language Models
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
von: Dinkel, Heinrich, et al.
Veröffentlicht: (2026)
SonicSense: Object Perception from In-Hand Acoustic Vibration
von: Liu, Jiaxun, et al.
Veröffentlicht: (2024)
von: Liu, Jiaxun, et al.
Veröffentlicht: (2024)
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
AND: Audio Network Dissection for Interpreting Deep Acoustic Models
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
Audio-Mind: An Auditable Agentic Framework for Audio Understanding
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
von: Wang, Yucheng, et al.
Veröffentlicht: (2026)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
AudioRAG: A Challenging Benchmark for Audio Reasoning and Information Retrieval
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
von: Lin, Jingru, et al.
Veröffentlicht: (2026)
ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing
von: Hu, Jinbo, et al.
Veröffentlicht: (2025)
von: Hu, Jinbo, et al.
Veröffentlicht: (2025)
Evaluation of Virtual Acoustic Environments with Different Acoustic Level of Detail
von: Fichna, Stefan, et al.
Veröffentlicht: (2023)
von: Fichna, Stefan, et al.
Veröffentlicht: (2023)
Polyphonia: Zero-Shot Timbre Transfer in Polyphonic Music with Acoustic-Informed Attention Calibration
von: Li, Haowen, et al.
Veröffentlicht: (2026)
von: Li, Haowen, et al.
Veröffentlicht: (2026)
EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
von: Hai, Jiarui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Past, Present, and Future of Spatial Audio and Room Acoustics
von: Koyama, Shoichi, et al.
Veröffentlicht: (2025) -
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026) -
Leveraging Self-supervised Audio Representations for Data-Efficient Acoustic Scene Classification
von: Cai, Yiqiang, et al.
Veröffentlicht: (2024) -
Structural and Statistical Audio Texture Knowledge Distillation for Acoustic Classification
von: Ritu, Jarin, et al.
Veröffentlicht: (2025) -
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
von: Guo, Hongming, et al.
Veröffentlicht: (2024)