DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Luo, Ziyu, Chen, Lin, Qu, Qiang, Chen, Xiaoming, Shen, Yiran |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
di: Luo, Ziyu, et al.
Pubblicazione: (2026)
di: Luo, Ziyu, et al.
Pubblicazione: (2026)
FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss
di: Sudarsanam, Parthasaarathy, et al.
Pubblicazione: (2025)
di: Sudarsanam, Parthasaarathy, et al.
Pubblicazione: (2025)
OmniAudio: Generating Spatial Audio from 360-Degree Video
di: Liu, Huadai, et al.
Pubblicazione: (2025)
di: Liu, Huadai, et al.
Pubblicazione: (2025)
SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering
di: Gayer, Yhonatan
Pubblicazione: (2026)
di: Gayer, Yhonatan
Pubblicazione: (2026)
Compression of Higher Order Ambisonics with Multichannel RVQGAN
di: Hirvonen, Toni, et al.
Pubblicazione: (2024)
di: Hirvonen, Toni, et al.
Pubblicazione: (2024)
DiffAU: Diffusion-Based Ambisonics Upscaling
di: Milstein, Amit, et al.
Pubblicazione: (2025)
di: Milstein, Amit, et al.
Pubblicazione: (2025)
Ambisonizer: Neural Upmixing as Spherical Harmonics Generation
di: Zang, Yongyi, et al.
Pubblicazione: (2024)
di: Zang, Yongyi, et al.
Pubblicazione: (2024)
HiFi-HARP: A High-Fidelity 7th-Order Ambisonic Room Impulse Response Dataset
di: Saini, Shivam, et al.
Pubblicazione: (2025)
di: Saini, Shivam, et al.
Pubblicazione: (2025)
Residual Learning for Neural Ambisonics Encoders
di: Deppisch, Thomas, et al.
Pubblicazione: (2026)
di: Deppisch, Thomas, et al.
Pubblicazione: (2026)
Improving Acoustic Scene Classification in Low-Resource Conditions
di: Chen, Zhi, et al.
Pubblicazione: (2024)
di: Chen, Zhi, et al.
Pubblicazione: (2024)
Ambisonics Networks -- The Effect Of Radial Functions Regularization
di: Shaybet, Bar, et al.
Pubblicazione: (2024)
di: Shaybet, Bar, et al.
Pubblicazione: (2024)
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
di: Chen, Tuochao, et al.
Pubblicazione: (2025)
di: Chen, Tuochao, et al.
Pubblicazione: (2025)
Neural Ambisonics encoding for compact irregular microphone arrays
di: Heikkinen, Mikko, et al.
Pubblicazione: (2024)
di: Heikkinen, Mikko, et al.
Pubblicazione: (2024)
Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays
di: Heikkinen, Mikko, et al.
Pubblicazione: (2025)
di: Heikkinen, Mikko, et al.
Pubblicazione: (2025)
Acoustic BPE for Speech Generation with Discrete Tokens
di: Shen, Feiyu, et al.
Pubblicazione: (2023)
di: Shen, Feiyu, et al.
Pubblicazione: (2023)
Ambisonics Binaural Rendering via Masked Magnitude Least Squares
di: Berebi, Or, et al.
Pubblicazione: (2025)
di: Berebi, Or, et al.
Pubblicazione: (2025)
Diffusion Reconstruction towards Generalizable Audio Deepfake Detection
di: Cheng, Bo, et al.
Pubblicazione: (2026)
di: Cheng, Bo, et al.
Pubblicazione: (2026)
Blind Estimation of Sub-band Acoustic Parameters from Ambisonics Recordings using Spectro-Spatial Covariance Features
di: Meng, Hanyu, et al.
Pubblicazione: (2024)
di: Meng, Hanyu, et al.
Pubblicazione: (2024)
HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset
di: Saini, Shivam, et al.
Pubblicazione: (2024)
di: Saini, Shivam, et al.
Pubblicazione: (2024)
Direction-Aware Neural Acoustic Fields for Few-Shot Interpolation of Ambisonic Impulse Responses
di: Ick, Christopher, et al.
Pubblicazione: (2025)
di: Ick, Christopher, et al.
Pubblicazione: (2025)
GSound-SIR: A Spatial Impulse Response Ray-Tracing and High-order Ambisonic Auralization Python Toolkit
di: Zang, Yongyi, et al.
Pubblicazione: (2025)
di: Zang, Yongyi, et al.
Pubblicazione: (2025)
Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
di: Nawfal, Ismael, et al.
Pubblicazione: (2025)
di: Nawfal, Ismael, et al.
Pubblicazione: (2025)
Room Impulse Response Generation Conditioned on Acoustic Parameters
di: Arellano, Silvia, et al.
Pubblicazione: (2025)
di: Arellano, Silvia, et al.
Pubblicazione: (2025)
Tri-Ergon: Fine-grained Video-to-Audio Generation with Multi-modal Conditions and LUFS Control
di: Li, Bingliang, et al.
Pubblicazione: (2024)
di: Li, Bingliang, et al.
Pubblicazione: (2024)
Can We Hear from Events? Generating Speech from Event Camera
di: Fang, Jingping, et al.
Pubblicazione: (2026)
di: Fang, Jingping, et al.
Pubblicazione: (2026)
Video Echoed in Music: Semantic, Temporal, and Rhythmic Alignment for Video-to-Music Generation
di: Tong, Xinyi, et al.
Pubblicazione: (2025)
di: Tong, Xinyi, et al.
Pubblicazione: (2025)
Towards General Discrete Speech Codec for Complex Acoustic Environments: A Study of Reconstruction and Downstream Task Consistency
di: Wang, Haoran, et al.
Pubblicazione: (2025)
di: Wang, Haoran, et al.
Pubblicazione: (2025)
Data-Efficient Low-Complexity Acoustic Scene Classification via Distilling and Progressive Pruning
di: Han, Bing, et al.
Pubblicazione: (2024)
di: Han, Bing, et al.
Pubblicazione: (2024)
RSA-Bench: Benchmarking Audio Large Models in Real-World Acoustic Scenarios
di: Zhang, Yibo, et al.
Pubblicazione: (2026)
di: Zhang, Yibo, et al.
Pubblicazione: (2026)
Array-Aware Ambisonics and HRTF Encoding for Binaural Reproduction With Wearable Arrays
di: Gayer, Yhonatan, et al.
Pubblicazione: (2025)
di: Gayer, Yhonatan, et al.
Pubblicazione: (2025)
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2026)
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2026)
LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation
di: Wang, Qi, et al.
Pubblicazione: (2026)
di: Wang, Qi, et al.
Pubblicazione: (2026)
Apollo: Unified Multi-Task Audio-Video Joint Generation
di: Wang, Jun, et al.
Pubblicazione: (2026)
di: Wang, Jun, et al.
Pubblicazione: (2026)
Low-Complexity Acoustic Scene Classification Using Parallel Attention-Convolution Network
di: Li, Yanxiong, et al.
Pubblicazione: (2024)
di: Li, Yanxiong, et al.
Pubblicazione: (2024)
ViTex: Visual Texture Control for Multi-Track Symbolic Music Generation via Discrete Diffusion Models
di: Yi, Xiaoyu, et al.
Pubblicazione: (2026)
di: Yi, Xiaoyu, et al.
Pubblicazione: (2026)
Before the Mic: Physical-Layer Voiceprint Anonymization with Acoustic Metamaterials
di: Ning, Zhiyuan, et al.
Pubblicazione: (2026)
di: Ning, Zhiyuan, et al.
Pubblicazione: (2026)
MSF-SER: Enriching Acoustic Modeling with Multi-Granularity Semantics for Speech Emotion Recognition
di: Li, Haoxun, et al.
Pubblicazione: (2025)
di: Li, Haoxun, et al.
Pubblicazione: (2025)
LipDiffuser: Lip-to-Speech Generation with Conditional Diffusion Models
di: Richter, Julius, et al.
Pubblicazione: (2025)
di: Richter, Julius, et al.
Pubblicazione: (2025)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
di: Lin, Wan, et al.
Pubblicazione: (2024)
di: Lin, Wan, et al.
Pubblicazione: (2024)
Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
di: Wang, Jun, et al.
Pubblicazione: (2025)
di: Wang, Jun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
di: Luo, Ziyu, et al.
Pubblicazione: (2026) -
FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss
di: Sudarsanam, Parthasaarathy, et al.
Pubblicazione: (2025) -
OmniAudio: Generating Spatial Audio from 360-Degree Video
di: Liu, Huadai, et al.
Pubblicazione: (2025) -
SHroom: A Python Framework for Ambisonics Room Acoustics Simulation and Binaural Rendering
di: Gayer, Yhonatan
Pubblicazione: (2026) -
Compression of Higher Order Ambisonics with Multichannel RVQGAN
di: Hirvonen, Toni, et al.
Pubblicazione: (2024)