Guardado en:
| Autores principales: | Galougah, Siminfar Samakoush, Pulijala, Pranav, Duraiswami, Ramani |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.02205 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Applying Automatic Differentiation to Optimize Differential Microphone Array Designs
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2024)
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2024)
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2025)
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2025)
Spectrum Coexistence, Network Dimensioning, and Cell-Free Architectures in 5G and 5G-Advanced Wireless Networks
por: Galougah, Siminfar Samakoush
Publicado: (2026)
por: Galougah, Siminfar Samakoush
Publicado: (2026)
Room Impulse Response Synthesis via Differentiable Feedback Delay Networks for Efficient Spatial Audio Rendering
por: Gerami, Armin, et al.
Publicado: (2025)
por: Gerami, Armin, et al.
Publicado: (2025)
Biomimetic Frontend for Differentiable Audio Processing
por: Famularo, Ruolan Leslie, et al.
Publicado: (2024)
por: Famularo, Ruolan Leslie, et al.
Publicado: (2024)
Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models
por: Seth, Ashish, et al.
Publicado: (2026)
por: Seth, Ashish, et al.
Publicado: (2026)
TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification
por: Anand, Nishit, et al.
Publicado: (2024)
por: Anand, Nishit, et al.
Publicado: (2024)
RECAP: Retrieval-Augmented Audio Captioning
por: Ghosh, Sreyan, et al.
Publicado: (2023)
por: Ghosh, Sreyan, et al.
Publicado: (2023)
MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
por: Sakshi, S, et al.
Publicado: (2024)
por: Sakshi, S, et al.
Publicado: (2024)
ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models
por: Ghosh, Sreyan, et al.
Publicado: (2023)
por: Ghosh, Sreyan, et al.
Publicado: (2023)
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
por: Goel, Arushi, et al.
Publicado: (2025)
por: Goel, Arushi, et al.
Publicado: (2025)
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
por: Sedláček, Šimon, et al.
Publicado: (2025)
por: Sedláček, Šimon, et al.
Publicado: (2025)
Audio Flamingo Next: Next-Generation Open Audio-Language Models for Speech, Sound, and Music
por: Ghosh, Sreyan, et al.
Publicado: (2026)
por: Ghosh, Sreyan, et al.
Publicado: (2026)
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
por: Yang, Chao-Han Huck, et al.
Publicado: (2025)
por: Yang, Chao-Han Huck, et al.
Publicado: (2025)
Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
por: Wang, Ruoyu, et al.
Publicado: (2024)
por: Wang, Ruoyu, et al.
Publicado: (2024)
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
por: Chen, Tianxiang, et al.
Publicado: (2024)
por: Chen, Tianxiang, et al.
Publicado: (2024)
Cinematic Audio Source Separation Using Visual Cues
por: Zhang, Kang, et al.
Publicado: (2026)
por: Zhang, Kang, et al.
Publicado: (2026)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
por: Pan, Tianrui, et al.
Publicado: (2024)
por: Pan, Tianrui, et al.
Publicado: (2024)
MUKA: Multi Kernel Audio Adaptation Of Audio-Language Models
por: Bensaid, Reda, et al.
Publicado: (2026)
por: Bensaid, Reda, et al.
Publicado: (2026)
Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
por: Cho, Hyunsung, et al.
Publicado: (2024)
por: Cho, Hyunsung, et al.
Publicado: (2024)
SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video Generation
por: Pham, Kien T., et al.
Publicado: (2025)
por: Pham, Kien T., et al.
Publicado: (2025)
Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration
por: Xie, Siyi, et al.
Publicado: (2025)
por: Xie, Siyi, et al.
Publicado: (2025)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
por: Zhou, Dingkun, et al.
Publicado: (2025)
por: Zhou, Dingkun, et al.
Publicado: (2025)
A Comprehensive Corpus of Biomechanically Constrained Piano Chords: Generation, Analysis, and Implications for Voicing and Psychoacoustics
por: Ramani, Mahesh
Publicado: (2026)
por: Ramani, Mahesh
Publicado: (2026)
MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model
por: Tao, Ye, et al.
Publicado: (2025)
por: Tao, Ye, et al.
Publicado: (2025)
Audio-Visual Speech Enhancement in Noisy Environments via Emotion-Based Contextual Cues
por: Hussain, Tassadaq, et al.
Publicado: (2024)
por: Hussain, Tassadaq, et al.
Publicado: (2024)
NablAFx: A Framework for Differentiable Black-box and Gray-box Modeling of Audio Effects
por: Comunità, Marco, et al.
Publicado: (2025)
por: Comunità, Marco, et al.
Publicado: (2025)
A Lightweight and Real-Time Binaural Speech Enhancement Model with Spatial Cues Preservation
por: Wang, Jingyuan, et al.
Publicado: (2024)
por: Wang, Jingyuan, et al.
Publicado: (2024)
A Sensitivity Analysis of Multi-Event Audio Grounding in Audio LLMs
por: Lee, Taehan, et al.
Publicado: (2026)
por: Lee, Taehan, et al.
Publicado: (2026)
SIREN: Spatially-Informed Reconstruction of Binaural Audio with Vision
por: Song, Mingyeong, et al.
Publicado: (2026)
por: Song, Mingyeong, et al.
Publicado: (2026)
Fundamental Survey on Neuromorphic Based Audio Classification
por: Basu, Amlan, et al.
Publicado: (2025)
por: Basu, Amlan, et al.
Publicado: (2025)
Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation
por: Lokegaonkar, Vaibhavi, et al.
Publicado: (2026)
por: Lokegaonkar, Vaibhavi, et al.
Publicado: (2026)
Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation
por: Zeng, Runhao, et al.
Publicado: (2025)
por: Zeng, Runhao, et al.
Publicado: (2025)
BreathNet: Generalizable Audio Deepfake Detection via Breath-Cue-Guided Feature Refinement
por: Ye, Zhe, et al.
Publicado: (2026)
por: Ye, Zhe, et al.
Publicado: (2026)
A Lightweight Fourier-based Network for Binaural Speech Enhancement with Spatial Cue Preservation
por: Lu, Xikun, et al.
Publicado: (2025)
por: Lu, Xikun, et al.
Publicado: (2025)
WTFormer: A Wavelet Conformer Network for MIMO Speech Enhancement with Spatial Cues Peservation
por: Han, Lu, et al.
Publicado: (2025)
por: Han, Lu, et al.
Publicado: (2025)
Can Large Language Models Understand Spatial Audio?
por: Tang, Changli, et al.
Publicado: (2024)
por: Tang, Changli, et al.
Publicado: (2024)
Universal Spatial Audio Transcoder
por: Sagasti, Amaia, et al.
Publicado: (2024)
por: Sagasti, Amaia, et al.
Publicado: (2024)
Ejemplares similares
-
Applying Automatic Differentiation to Optimize Differential Microphone Array Designs
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2024) -
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
por: Galougah, Siminfar Samakoush, et al.
Publicado: (2025) -
Spectrum Coexistence, Network Dimensioning, and Cell-Free Architectures in 5G and 5G-Advanced Wireless Networks
por: Galougah, Siminfar Samakoush
Publicado: (2026) -
Room Impulse Response Synthesis via Differentiable Feedback Delay Networks for Efficient Spatial Audio Rendering
por: Gerami, Armin, et al.
Publicado: (2025) -
Biomimetic Frontend for Differentiable Audio Processing
por: Famularo, Ruolan Leslie, et al.
Publicado: (2024)