Visual-based spatial audio generation system for multi-speaker environments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Xiaojing, Gurelli, Ogulcan, Wang, Yan, Reiss, Joshua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024)
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
von: Sun, Chang, et al.
Veröffentlicht: (2024)
von: Sun, Chang, et al.
Veröffentlicht: (2024)
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
Dance2MIDI: Dance-driven multi-instruments music generation
von: Han, Bo, et al.
Veröffentlicht: (2023)
von: Han, Bo, et al.
Veröffentlicht: (2023)
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
von: Wang, Anna, et al.
Veröffentlicht: (2024)
von: Wang, Anna, et al.
Veröffentlicht: (2024)
Robust LLM-based Audio-Visual Speech Recognition with Sparse Modality Alignment and Visual Unit-Guided Refinement
von: Su, Fei, et al.
Veröffentlicht: (2026)
von: Su, Fei, et al.
Veröffentlicht: (2026)
RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
von: Pan, Tianrui, et al.
Veröffentlicht: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
von: Yuan, Yi, et al.
Veröffentlicht: (2024)
Human-Inspired Computing for Robust and Efficient Audio-Visual Speech Recognition
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
von: Liu, Qianhui, et al.
Veröffentlicht: (2024)
Zero-Shot Fake Video Detection by Audio-Visual Consistency
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
von: Li, Xiaolou, et al.
Veröffentlicht: (2024)
Singer separation for karaoke content generation
von: Lin, Hsuan-Yu, et al.
Veröffentlicht: (2021)
von: Lin, Hsuan-Yu, et al.
Veröffentlicht: (2021)
Building Audio-Visual Digital Twins with Smartphones
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
von: Lan, Zitong, et al.
Veröffentlicht: (2025)
LCB-net: Long-Context Biasing for Audio-Visual Speech Recognition
von: Yu, Fan, et al.
Veröffentlicht: (2024)
von: Yu, Fan, et al.
Veröffentlicht: (2024)
TEAdapter: Supply abundant guidance for controllable text-to-music generation
von: Zou, Jialing, et al.
Veröffentlicht: (2024)
von: Zou, Jialing, et al.
Veröffentlicht: (2024)
Quality-Aware End-to-End Audio-Visual Neural Speaker Diarization
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
von: He, Mao-Kui, et al.
Veröffentlicht: (2024)
Versatile audio-visual learning for emotion recognition
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2023)
Rhythmic Foley: A Framework For Seamless Audio-Visual Alignment In Video-to-Audio Synthesis
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
von: Huang, Zhiqi, et al.
Veröffentlicht: (2024)
Robust Wake Word Spotting With Frame-Level Cross-Modal Attention Based Audio-Visual Conformer
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
von: Wang, Haoxu, et al.
Veröffentlicht: (2024)
Efficient Video to Audio Mapper with Visual Scene Detection
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
von: Yi, Mingjing, et al.
Veröffentlicht: (2024)
Cinematic Audio Source Separation Using Visual Cues
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
von: Zhang, Kang, et al.
Veröffentlicht: (2026)
Audio-Visual Speech Separation via Bottleneck Iterative Network
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
von: Zhang, Sidong, et al.
Veröffentlicht: (2025)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
SONIQUE: Video Background Music Generation Using Unpaired Audio-Visual Data
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
von: Zhang, Liqian, et al.
Veröffentlicht: (2024)
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
von: Han, Runduo, et al.
Veröffentlicht: (2024)
von: Han, Runduo, et al.
Veröffentlicht: (2024)
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
von: Rong, Yan, et al.
Veröffentlicht: (2025)
von: Rong, Yan, et al.
Veröffentlicht: (2025)
M6: Multi-generator, Multi-domain, Multi-lingual and cultural, Multi-genres, Multi-instrument Machine-Generated Music Detection Databases
von: Li, Yupei, et al.
Veröffentlicht: (2024)
von: Li, Yupei, et al.
Veröffentlicht: (2024)
StereoFoley: Object-Aware Stereo Audio Generation from Video
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
von: Karchkhadze, Tornike, et al.
Veröffentlicht: (2025)
Conformer-based Ultrasound-to-Speech Conversion
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
von: Ibrahimov, Ibrahim, et al.
Veröffentlicht: (2025)
FGAS: Fixed Decoder Network-Based Audio Steganography with Adversarial Perturbation Generation
von: Yan, Jialin, et al.
Veröffentlicht: (2025)
von: Yan, Jialin, et al.
Veröffentlicht: (2025)
Low-latency Speech Enhancement via Speech Token Generation
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
von: Xue, Huaying, et al.
Veröffentlicht: (2023)
DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
von: Liu, Shengqiang, et al.
Veröffentlicht: (2024)
von: Liu, Shengqiang, et al.
Veröffentlicht: (2024)
Dance-to-Music Generation with Encoder-based Textual Inversion
von: Li, Sifei, et al.
Veröffentlicht: (2024)
von: Li, Sifei, et al.
Veröffentlicht: (2024)
Exploring compressibility of transformer based text-to-music (TTM) models
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
von: Moschopoulos, Vasileios, et al.
Veröffentlicht: (2024)
Attentive-based Multi-level Feature Fusion for Voice Disorder Diagnosis
von: Shen, Lipeng, et al.
Veröffentlicht: (2024)
von: Shen, Lipeng, et al.
Veröffentlicht: (2024)
Improved symbolic drum style classification with grammar-based hierarchical representations
von: Géré, Léo, et al.
Veröffentlicht: (2024)
von: Géré, Léo, et al.
Veröffentlicht: (2024)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
von: Cui, Meng, et al.
Veröffentlicht: (2023)
von: Cui, Meng, et al.
Veröffentlicht: (2023)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
von: Guo, Hongming, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
An automatic mixing speech enhancement system for multi-track audio
von: Liu, Xiaojing, et al.
Veröffentlicht: (2024) -
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
von: Sun, Chang, et al.
Veröffentlicht: (2024) -
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
von: Kwak, Doyeop, et al.
Veröffentlicht: (2026) -
Dance2MIDI: Dance-driven multi-instruments music generation
von: Han, Bo, et al.
Veröffentlicht: (2023) -
M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
von: Wang, Anna, et al.
Veröffentlicht: (2024)