Compression of Higher Order Ambisonics with Multichannel RVQGAN
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hirvonen, Toni, Namazi, Mahmoud |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset
von: Saini, Shivam, et al.
Veröffentlicht: (2024)
von: Saini, Shivam, et al.
Veröffentlicht: (2024)
Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
von: Bandyopadhyay, Tathagata
Veröffentlicht: (2024)
von: Bandyopadhyay, Tathagata
Veröffentlicht: (2024)
Siamese Residual Neural Network for Musical Shape Evaluation in Piano Performance Assessment
von: Li, Xiaoquan, et al.
Veröffentlicht: (2024)
von: Li, Xiaoquan, et al.
Veröffentlicht: (2024)
Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation
von: Yu, Jun, et al.
Veröffentlicht: (2024)
von: Yu, Jun, et al.
Veröffentlicht: (2024)
A Recurrent Neural Network Approach to the Answering Machine Detection Problem
von: Altwlkany, Kemal, et al.
Veröffentlicht: (2024)
von: Altwlkany, Kemal, et al.
Veröffentlicht: (2024)
Source Separation of Multi-source Raw Music using a Residual Quantized Variational Autoencoder
von: Berti, Leonardo
Veröffentlicht: (2024)
von: Berti, Leonardo
Veröffentlicht: (2024)
CHORDONOMICON: A Dataset of 666,000 Songs and their Chord Progressions
von: Kantarelis, Spyridon, et al.
Veröffentlicht: (2024)
von: Kantarelis, Spyridon, et al.
Veröffentlicht: (2024)
Audiopedia: Audio QA with Knowledge
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2024)
von: Penamakuri, Abhirama Subramanyam, et al.
Veröffentlicht: (2024)
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
von: Liu, Renhang, et al.
Veröffentlicht: (2024)
von: Liu, Renhang, et al.
Veröffentlicht: (2024)
LSTMSE-Net: Long Short Term Speech Enhancement Network for Audio-visual Speech Enhancement
von: Jain, Arnav, et al.
Veröffentlicht: (2024)
von: Jain, Arnav, et al.
Veröffentlicht: (2024)
Unified Microphone Conversion: Many-to-Many Device Mapping via Feature-wise Linear Modulation
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2024)
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2024)
A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation
von: Ishii, Masato, et al.
Veröffentlicht: (2024)
von: Ishii, Masato, et al.
Veröffentlicht: (2024)
MidiCaps: A large-scale MIDI dataset with text captions
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2024)
Microphone Conversion: Mitigating Device Variability in Sound Event Classification
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2024)
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2024)
Just Label the Repeats for In-The-Wild Audio-to-Score Alignment
von: Bukey, Irmak, et al.
Veröffentlicht: (2024)
von: Bukey, Irmak, et al.
Veröffentlicht: (2024)
Music102: An $D_{12}$-equivariant transformer for chord progression accompaniment
von: Luo, Weiliang
Veröffentlicht: (2024)
von: Luo, Weiliang
Veröffentlicht: (2024)
Network Bending of Diffusion Models for Audio-Visual Generation
von: Dzwonczyk, Luke, et al.
Veröffentlicht: (2024)
von: Dzwonczyk, Luke, et al.
Veröffentlicht: (2024)
Speech Separation with Pretrained Frontend to Minimize Domain Mismatch
von: Wang, Wupeng, et al.
Veröffentlicht: (2024)
von: Wang, Wupeng, et al.
Veröffentlicht: (2024)
AWARE: Audio Watermarking with Adversarial Resistance to Edits
von: Pavlović, Kosta, et al.
Veröffentlicht: (2025)
von: Pavlović, Kosta, et al.
Veröffentlicht: (2025)
Bridging The Multi-Modality Gaps of Audio, Visual and Linguistic for Speech Enhancement
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
von: Lin, Meng-Ping, et al.
Veröffentlicht: (2025)
Multimodal Emotion Coupling via Speech-to-Facial and Bodily Gestures in Dyadic Interaction
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
von: Herbuela, Von Ralph Dane Marquez, et al.
Veröffentlicht: (2025)
Revisit Modality Imbalance at the Decision Layer
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Ma, Xiaoyu, et al.
Veröffentlicht: (2025)
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
von: Bai, Yatong, et al.
Veröffentlicht: (2023)
Multimodal Speech Enhancement Using Burst Propagation
von: Raza, Mohsin, et al.
Veröffentlicht: (2022)
von: Raza, Mohsin, et al.
Veröffentlicht: (2022)
$C^2$AV-TSE: Context and Confidence-aware Audio Visual Target Speaker Extraction
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2025)
BERT-like Pre-training for Symbolic Piano Music Classification Tasks
von: Chou, Yi-Hui, et al.
Veröffentlicht: (2021)
von: Chou, Yi-Hui, et al.
Veröffentlicht: (2021)
MIDI-GPT: A Controllable Generative Model for Computer-Assisted Multitrack Music Composition
von: Pasquier, Philippe, et al.
Veröffentlicht: (2025)
von: Pasquier, Philippe, et al.
Veröffentlicht: (2025)
Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
von: Wang, Juncheng, et al.
Veröffentlicht: (2025)
A Traditional Approach to Symbolic Piano Continuation
von: Zhou-Zheng, Christian, et al.
Veröffentlicht: (2025)
von: Zhou-Zheng, Christian, et al.
Veröffentlicht: (2025)
Versatile audio-visual learning for emotion recognition
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2023)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
IML-Spikeformer: Input-aware Multi-Level Spiking Transformer for Speech Processing
von: Song, Zeyang, et al.
Veröffentlicht: (2025)
von: Song, Zeyang, et al.
Veröffentlicht: (2025)
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
A vector quantized masked autoencoder for audiovisual speech emotion recognition
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
Leveraging Pre-Trained Models for Multimodal Class-Incremental Learning under Adaptive Fusion
von: Chen, Yukun, et al.
Veröffentlicht: (2025)
von: Chen, Yukun, et al.
Veröffentlicht: (2025)
Improving BERT for Symbolic Music Understanding Using Token Denoising and Pianoroll Prediction
von: Wang, Jun-You, et al.
Veröffentlicht: (2025)
von: Wang, Jun-You, et al.
Veröffentlicht: (2025)
Multimodal Dataset Normalization and Perceptual Validation for Music-Taste Correspondences
von: Spanio, Matteo, et al.
Veröffentlicht: (2026)
von: Spanio, Matteo, et al.
Veröffentlicht: (2026)
Tiny is not small enough: High-quality, low-resource facial animation models through hybrid knowledge distillation
von: Han, Zhen, et al.
Veröffentlicht: (2025)
von: Han, Zhen, et al.
Veröffentlicht: (2025)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset
von: Saini, Shivam, et al.
Veröffentlicht: (2024) -
Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
von: Bandyopadhyay, Tathagata
Veröffentlicht: (2024) -
Siamese Residual Neural Network for Musical Shape Evaluation in Piano Performance Assessment
von: Li, Xiaoquan, et al.
Veröffentlicht: (2024) -
Efficient Feature Extraction and Late Fusion Strategy for Audiovisual Emotional Mimicry Intensity Estimation
von: Yu, Jun, et al.
Veröffentlicht: (2024) -
A Recurrent Neural Network Approach to the Answering Machine Detection Problem
von: Altwlkany, Kemal, et al.
Veröffentlicht: (2024)