FOA Tokenizer: Low-bitrate Neural Codec for First Order Ambisonics with Spatial Consistency Loss
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sudarsanam, Parthasaarathy, Braun, Sebastian, Gamper, Hannes |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Spatial Audio Understanding via Question Answering
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
von: Luo, Ziyu, et al.
Veröffentlicht: (2026)
von: Luo, Ziyu, et al.
Veröffentlicht: (2026)
DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
von: Luo, Ziyu, et al.
Veröffentlicht: (2026)
von: Luo, Ziyu, et al.
Veröffentlicht: (2026)
ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
von: Yang, Dongchao, et al.
Veröffentlicht: (2025)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)
Sci-Phi: A Large Language Model Spatial Audio Descriptor
von: Jiang, Xilin, et al.
Veröffentlicht: (2025)
von: Jiang, Xilin, et al.
Veröffentlicht: (2025)
PSCodec: A Series of High-Fidelity Low-bitrate Neural Speech Codecs Leveraging Prompt Encoders
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
von: Moliner, Eloi, et al.
Veröffentlicht: (2024)
CMMD: Contrastive Multi-Modal Diffusion for Video-Audio Conditional Modeling
von: Yang, Ruihan, et al.
Veröffentlicht: (2023)
von: Yang, Ruihan, et al.
Veröffentlicht: (2023)
FlexiCodec: A Dynamic Neural Audio Codec for Low Frame Rates
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
Distillation and Pruning for Scalable Self-Supervised Representation-Based Speech Quality Assessment
von: Stahl, Benjamin, et al.
Veröffentlicht: (2025)
von: Stahl, Benjamin, et al.
Veröffentlicht: (2025)
SpatialCodec: Neural Spatial Speech Coding
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
Residual Learning for Neural Ambisonics Encoders
von: Deppisch, Thomas, et al.
Veröffentlicht: (2026)
von: Deppisch, Thomas, et al.
Veröffentlicht: (2026)
Spatial Audio Rendering for Real-Time Speech Translation in Virtual Meetings
von: Geleta, Margarita, et al.
Veröffentlicht: (2025)
von: Geleta, Margarita, et al.
Veröffentlicht: (2025)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
Ambisonizer: Neural Upmixing as Spherical Harmonics Generation
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
von: Zang, Yongyi, et al.
Veröffentlicht: (2024)
Compression of Higher Order Ambisonics with Multichannel RVQGAN
von: Hirvonen, Toni, et al.
Veröffentlicht: (2024)
von: Hirvonen, Toni, et al.
Veröffentlicht: (2024)
Universal Speech Token Learning via Low-Bitrate Neural Codec and Pretrained Representations
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
von: Jiang, Xue, et al.
Veröffentlicht: (2025)
Neural Ambisonics encoding for compact irregular microphone arrays
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2024)
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2024)
LLM-Codec: Neural Audio Codec Meets Language Model Objectives
von: Chung, Ho-Lam, et al.
Veröffentlicht: (2026)
von: Chung, Ho-Lam, et al.
Veröffentlicht: (2026)
DAC-JAX: A JAX Implementation of the Descript Audio Codec
von: Braun, David
Veröffentlicht: (2024)
von: Braun, David
Veröffentlicht: (2024)
HiFi-HARP: A High-Fidelity 7th-Order Ambisonic Room Impulse Response Dataset
von: Saini, Shivam, et al.
Veröffentlicht: (2025)
von: Saini, Shivam, et al.
Veröffentlicht: (2025)
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling
von: Shi, Jiacheng, et al.
Veröffentlicht: (2026)
von: Shi, Jiacheng, et al.
Veröffentlicht: (2026)
GSound-SIR: A Spatial Impulse Response Ray-Tracing and High-order Ambisonic Auralization Python Toolkit
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
von: Zang, Yongyi, et al.
Veröffentlicht: (2025)
AffectCodec: Emotion-Preserving Neural Speech Codec with Block-Diagonal Residual FSQ
von: Meng, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Meng, Zhaoyang, et al.
Veröffentlicht: (2026)
Analyzing and Mitigating Inconsistency in Discrete Audio Tokens for Neural Codec Language Models
von: Liu, Wenrui, et al.
Veröffentlicht: (2024)
von: Liu, Wenrui, et al.
Veröffentlicht: (2024)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
RepCodec: A Speech Representation Codec for Speech Tokenization
von: Huang, Zhichao, et al.
Veröffentlicht: (2023)
von: Huang, Zhichao, et al.
Veröffentlicht: (2023)
MuCodec: Ultra Low-Bitrate Music Codec
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
von: Xu, Yaoxun, et al.
Veröffentlicht: (2024)
Ambisonics Super-Resolution Using A Waveform-Domain Neural Network
von: Nawfal, Ismael, et al.
Veröffentlicht: (2025)
von: Nawfal, Ismael, et al.
Veröffentlicht: (2025)
SACodec: Asymmetric Quantization with Semantic Anchoring for Low-Bitrate High-Fidelity Neural Speech Codecs
von: Dong, Zhongren, et al.
Veröffentlicht: (2025)
von: Dong, Zhongren, et al.
Veröffentlicht: (2025)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
Soft Disentanglement in Frequency Bands for Neural Audio Codecs
von: Ginies, Benoit, et al.
Veröffentlicht: (2025)
von: Ginies, Benoit, et al.
Veröffentlicht: (2025)
U-Codec: Ultra Low Frame-rate Neural Speech Codec for Fast High-fidelity Speech Generation
von: Yang, Xusheng, et al.
Veröffentlicht: (2025)
von: Yang, Xusheng, et al.
Veröffentlicht: (2025)
Adapting Neural Audio Codecs to EEG
von: Kastrati, Ard, et al.
Veröffentlicht: (2025)
von: Kastrati, Ard, et al.
Veröffentlicht: (2025)
Personalized Neural Speech Codec
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents
von: Banerjee, Adhiraj, et al.
Veröffentlicht: (2025)
von: Banerjee, Adhiraj, et al.
Veröffentlicht: (2025)
Evaluating Objective Speech Quality Metrics for Neural Audio Codecs
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
von: Lanzendörfer, Luca A., et al.
Veröffentlicht: (2025)
Harmonic-Percussive Disentangled Neural Audio Codec for Bandwidth Extension
von: Giniès, Benoît, et al.
Veröffentlicht: (2025)
von: Giniès, Benoît, et al.
Veröffentlicht: (2025)
Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone Arrays
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2025)
von: Heikkinen, Mikko, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Spatial Audio Understanding via Question Answering
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025) -
DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
von: Luo, Ziyu, et al.
Veröffentlicht: (2026) -
DynFOA: Generating First-Order Ambisonics with Conditional Diffusion for Dynamic and Acoustically Complex 360-Degree Videos
von: Luo, Ziyu, et al.
Veröffentlicht: (2026) -
ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
von: Yang, Dongchao, et al.
Veröffentlicht: (2025) -
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
von: Sudarsanam, Parthasaarathy, et al.
Veröffentlicht: (2025)