SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Haohe, Xu, Xuenan, Yuan, Yi, Wu, Mengyue, Wang, Wenwu, Plumbley, Mark D. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning Temporal Resolution in Spectrogram for Audio Classification
por: Liu, Haohe, et al.
Publicado: (2022)
por: Liu, Haohe, et al.
Publicado: (2022)
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
por: Liu, Haohe, et al.
Publicado: (2023)
por: Liu, Haohe, et al.
Publicado: (2023)
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
por: Yuan, Yi, et al.
Publicado: (2023)
por: Yuan, Yi, et al.
Publicado: (2023)
Retrieval-Augmented Text-to-Audio Generation
por: Yuan, Yi, et al.
Publicado: (2023)
por: Yuan, Yi, et al.
Publicado: (2023)
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
por: Xu, Xuenan, et al.
Publicado: (2024)
por: Xu, Xuenan, et al.
Publicado: (2024)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
por: Ji, Shengpeng, et al.
Publicado: (2024)
por: Ji, Shengpeng, et al.
Publicado: (2024)
Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions
por: Yuan, Yi, et al.
Publicado: (2024)
por: Yuan, Yi, et al.
Publicado: (2024)
MuCodec: Ultra Low-Bitrate Music Codec
por: Xu, Yaoxun, et al.
Publicado: (2024)
por: Xu, Yaoxun, et al.
Publicado: (2024)
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
por: Liu, Haohe, et al.
Publicado: (2024)
por: Liu, Haohe, et al.
Publicado: (2024)
WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
por: Mei, Xinhao, et al.
Publicado: (2023)
por: Mei, Xinhao, et al.
Publicado: (2023)
Towards Generating Diverse Audio Captions via Adversarial Training
por: Mei, Xinhao, et al.
Publicado: (2022)
por: Mei, Xinhao, et al.
Publicado: (2022)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
por: Xin, Detai, et al.
Publicado: (2024)
por: Xin, Detai, et al.
Publicado: (2024)
FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
por: Yuan, Yi, et al.
Publicado: (2024)
por: Yuan, Yi, et al.
Publicado: (2024)
Separate Anything You Describe
por: Liu, Xubo, et al.
Publicado: (2023)
por: Liu, Xubo, et al.
Publicado: (2023)
FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
por: Xu, Zhongweiyang, et al.
Publicado: (2024)
por: Xu, Zhongweiyang, et al.
Publicado: (2024)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
por: Li, Jiaqi, et al.
Publicado: (2025)
por: Li, Jiaqi, et al.
Publicado: (2025)
Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec
por: Ren, Yanzhou, et al.
Publicado: (2026)
por: Ren, Yanzhou, et al.
Publicado: (2026)
DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
por: Tian, Wenjie, et al.
Publicado: (2025)
por: Tian, Wenjie, et al.
Publicado: (2025)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
por: Chou, Huang-Cheng, et al.
Publicado: (2024)
por: Chou, Huang-Cheng, et al.
Publicado: (2024)
Gull: A Generative Multifunctional Audio Codec
por: Luo, Yi, et al.
Publicado: (2024)
por: Luo, Yi, et al.
Publicado: (2024)
On the Design of Diffusion-based Neural Speech Codecs
por: Foti, Pietro, et al.
Publicado: (2025)
por: Foti, Pietro, et al.
Publicado: (2025)
Codec-SUPERB: An In-Depth Analysis of Sound Codec Models
por: Wu, Haibin, et al.
Publicado: (2024)
por: Wu, Haibin, et al.
Publicado: (2024)
Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning
por: Sun, Luoyi, et al.
Publicado: (2023)
por: Sun, Luoyi, et al.
Publicado: (2023)
Multimodal Fish Feeding Intensity Assessment in Aquaculture
por: Cui, Meng, et al.
Publicado: (2023)
por: Cui, Meng, et al.
Publicado: (2023)
Latent Granular Resynthesis using Neural Audio Codecs
por: Tokui, Nao, et al.
Publicado: (2025)
por: Tokui, Nao, et al.
Publicado: (2025)
Exploring the User Experience of AI-Assisted Sound Searching Systems for Creative Workflows
por: Liu, Haohe, et al.
Publicado: (2025)
por: Liu, Haohe, et al.
Publicado: (2025)
Smooth-Foley: Creating Continuous Sound for Video-to-Audio Generation Under Semantic Guidance
por: Zhang, Yaoyun, et al.
Publicado: (2024)
por: Zhang, Yaoyun, et al.
Publicado: (2024)
Integrating IP Broadcasting with Audio Tags: Workflow and Challenges
por: Burchett-Vass, Rhys, et al.
Publicado: (2024)
por: Burchett-Vass, Rhys, et al.
Publicado: (2024)
Region-Specific Audio Tagging for Spatial Sound
por: Zhao, Jinzheng, et al.
Publicado: (2025)
por: Zhao, Jinzheng, et al.
Publicado: (2025)
Audio-Visual Speaker Tracking: Progress, Challenges, and Future Directions
por: Zhao, Jinzheng, et al.
Publicado: (2023)
por: Zhao, Jinzheng, et al.
Publicado: (2023)
STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
por: Ren, Yong, et al.
Publicado: (2024)
por: Ren, Yong, et al.
Publicado: (2024)
Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model
por: Chen, Gehui, et al.
Publicado: (2024)
por: Chen, Gehui, et al.
Publicado: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
por: Xie, Zeyu, et al.
Publicado: (2023)
por: Xie, Zeyu, et al.
Publicado: (2023)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
por: Kim, Minje, et al.
Publicado: (2024)
por: Kim, Minje, et al.
Publicado: (2024)
MDCTCodec: A Lightweight MDCT-based Neural Audio Codec towards High Sampling Rate and Low Bitrate Scenarios
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
por: Jiang, Xiao-Hang, et al.
Publicado: (2024)
Representation Learning for Semantic Alignment of Language, Audio, and Visual Modalities
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
por: Sudarsanam, Parthasaarathy, et al.
Publicado: (2025)
LoVA: Long-form Video-to-Audio Generation
por: Cheng, Xin, et al.
Publicado: (2024)
por: Cheng, Xin, et al.
Publicado: (2024)
Optimizing Neural Speech Codec for Low-Bitrate Compression via Multi-Scale Encoding
por: Yang, Peiji, et al.
Publicado: (2024)
por: Yang, Peiji, et al.
Publicado: (2024)
Audio-Language Models for Audio-Centric Tasks: A Systematic Survey
por: Su, Yi, et al.
Publicado: (2025)
por: Su, Yi, et al.
Publicado: (2025)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
por: Gong, Yitian, et al.
Publicado: (2025)
por: Gong, Yitian, et al.
Publicado: (2025)
Ejemplares similares
-
Learning Temporal Resolution in Spectrogram for Audio Classification
por: Liu, Haohe, et al.
Publicado: (2022) -
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining
por: Liu, Haohe, et al.
Publicado: (2023) -
Leveraging Pre-trained AudioLDM for Sound Generation: A Benchmark Study
por: Yuan, Yi, et al.
Publicado: (2023) -
Retrieval-Augmented Text-to-Audio Generation
por: Yuan, Yi, et al.
Publicado: (2023) -
Efficient Audio Captioning with Encoder-Level Knowledge Distillation
por: Xu, Xuenan, et al.
Publicado: (2024)