Spatial-Aware Conditioned Fusion for Audio-Visual Navigation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Shaohang, Yu, Yinfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Audio Spatially-Guided Fusion for Audio-Visual Navigation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Reliability-Aware Geometric Fusion for Robust Audio-Visual Navigation
von: Liu, Teng, et al.
Veröffentlicht: (2026)
von: Liu, Teng, et al.
Veröffentlicht: (2026)
Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction
von: Li, Jia, et al.
Veröffentlicht: (2026)
von: Li, Jia, et al.
Veröffentlicht: (2026)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)
TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion
von: Pegg, Samuel, et al.
Veröffentlicht: (2024)
von: Pegg, Samuel, et al.
Veröffentlicht: (2024)
ML-SAN: Multi-Level Speaker-Adaptive Network for Emotion Recognition in Conversations
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
von: Wang, Kexue, et al.
Veröffentlicht: (2026)
Leveraging Label Potential for Enhanced Multimodal Emotion Recognition
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
von: Shao, Xuechun, et al.
Veröffentlicht: (2025)
HRTFformer: A Spatially-Aware Transformer for Individual HRTF Upsampling in Immersive Audio Rendering
von: Hu, Xuyi, et al.
Veröffentlicht: (2025)
von: Hu, Xuyi, et al.
Veröffentlicht: (2025)
FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
MLCA-AVSR: Multi-Layer Cross Attention Fusion based Audio-Visual Speech Recognition
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
In-the-wild Audio Spatialization with Flexible Text-guided Localization
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
von: Pan, Tianrui, et al.
Veröffentlicht: (2025)
Audio Atlas: Visualizing and Exploring Audio Datasets
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2024)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2024)
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
von: Park, Hansol, et al.
Veröffentlicht: (2025)
von: Park, Hansol, et al.
Veröffentlicht: (2025)
Magnitude-Phase Dual-Path Speech Enhancement Network based on Self-Supervised Embedding and Perceptual Contrast Stretch Boosting
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2025)
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
von: Cao, Yubing, et al.
Veröffentlicht: (2025)
Audio-Guided Fusion Techniques for Multimodal Emotion Analysis
von: Shi, Pujin, et al.
Veröffentlicht: (2024)
von: Shi, Pujin, et al.
Veröffentlicht: (2024)
ViSAGe: Video-to-Spatial Audio Generation
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2025)
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
von: Chen, Chih-Ning, et al.
Veröffentlicht: (2026)
von: Chen, Chih-Ning, et al.
Veröffentlicht: (2026)
Audio-Conditioned Diffusion LLMs for ASR and Deliberation Processing
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
von: Wang, Mengqi, et al.
Veröffentlicht: (2025)
ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
AND: Audio Network Dissection for Interpreting Deep Acoustic Models
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
A Multi-Stream Fusion Approach with One-Class Learning for Audio-Visual Deepfake Detection
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2024)
PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
von: Dementyev, Artem, et al.
Veröffentlicht: (2026)
von: Dementyev, Artem, et al.
Veröffentlicht: (2026)
DSpAST: Disentangled Representations for Spatial Audio Reasoning with Large Language Models
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
von: Wilkinghoff, Kevin, et al.
Veröffentlicht: (2025)
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
von: Gao, Ming, et al.
Veröffentlicht: (2025)
von: Gao, Ming, et al.
Veröffentlicht: (2025)
BrewCLIP: A Bifurcated Representation Learning Framework for Audio-Visual Retrieval
von: Lu, Zhenyu, et al.
Veröffentlicht: (2024)
von: Lu, Zhenyu, et al.
Veröffentlicht: (2024)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
Plug-and-Play Co-Occurring Face Attention for Robust Audio-Visual Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
Unveiling Visual Biases in Audio-Visual Localization Benchmarks
von: Chen, Liangyu, et al.
Veröffentlicht: (2024)
von: Chen, Liangyu, et al.
Veröffentlicht: (2024)
RBA-FE: A Robust Brain-Inspired Audio Feature Extractor for Depression Diagnosis
von: Wu, Yu-Xuan, et al.
Veröffentlicht: (2025)
von: Wu, Yu-Xuan, et al.
Veröffentlicht: (2025)
Audio Mamba: Pretrained Audio State Space Model For Audio Tagging
von: Lin, Jiaju, et al.
Veröffentlicht: (2024)
von: Lin, Jiaju, et al.
Veröffentlicht: (2024)
WavShape: Information-Theoretic Speech Representation Learning for Fair and Privacy-Aware Audio Processing
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
von: Baser, Oguzhan, et al.
Veröffentlicht: (2025)
Do Models Hear Like Us? Probing the Representational Alignment of Audio LLMs and Naturalistic EEG
von: Yang, Haoyun, et al.
Veröffentlicht: (2026)
von: Yang, Haoyun, et al.
Veröffentlicht: (2026)
Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech
von: He, Shuwei, et al.
Veröffentlicht: (2024)
von: He, Shuwei, et al.
Veröffentlicht: (2024)
Decoding Ambiguous Emotions with Test-Time Scaling in Audio-Language Models
von: Jia, Hong, et al.
Veröffentlicht: (2026)
von: Jia, Hong, et al.
Veröffentlicht: (2026)
RFWave: Multi-band Rectified Flow for Audio Waveform Reconstruction
von: Liu, Peng, et al.
Veröffentlicht: (2024)
von: Liu, Peng, et al.
Veröffentlicht: (2024)
Representation-Regularized Convolutional Audio Transformer for Audio Understanding
von: Han, Bing, et al.
Veröffentlicht: (2026)
von: Han, Bing, et al.
Veröffentlicht: (2026)
When Scaling Fails: Mitigating Audio Perception Decay of LALMs via Multi-Step Perception-Aware Reasoning
von: Mao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Mao, Ruixiang, et al.
Veröffentlicht: (2026)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaeyeon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Audio Spatially-Guided Fusion for Audio-Visual Navigation
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026) -
Reliability-Aware Geometric Fusion for Robust Audio-Visual Navigation
von: Liu, Teng, et al.
Veröffentlicht: (2026) -
Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction
von: Li, Jia, et al.
Veröffentlicht: (2026) -
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024) -
EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian Deformation
von: Zhu, Tianheng, et al.
Veröffentlicht: (2025)