STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Kaiyuan, Shi, Mohan, Eren, Eray, Shankar, Natarajan Balaji, Wang, Zilai, Alwan, Abeer |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
von: Shi, Mohan, et al.
Veröffentlicht: (2025)
Selective Attention Merging for low resource tasks: A case study of Child ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025)
Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR
von: Wang, Zilai, et al.
Veröffentlicht: (2026)
von: Wang, Zilai, et al.
Veröffentlicht: (2026)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
SOA: Reducing Domain Mismatch in SSL Pipeline by Speech Only Adaptation for Low Resource ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2024)
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2024)
ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs
von: Eren, Eray, et al.
Veröffentlicht: (2025)
von: Eren, Eray, et al.
Veröffentlicht: (2025)
UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
An Age-Agnostic System for Robust Speaker Verification
von: Zheng, Jiusi, et al.
Veröffentlicht: (2025)
von: Zheng, Jiusi, et al.
Veröffentlicht: (2025)
OmniCodec: Low Frame Rate Universal Audio Codec with Semantic-Acoustic Disentanglement
von: Hu, Jingbin, et al.
Veröffentlicht: (2026)
von: Hu, Jingbin, et al.
Veröffentlicht: (2026)
Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
von: Manakul, Potsawee, et al.
Veröffentlicht: (2026)
von: Manakul, Potsawee, et al.
Veröffentlicht: (2026)
Speech Codec Probing from Semantic and Phonetic Perspectives
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
von: Shi, Xuan, et al.
Veröffentlicht: (2026)
Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
Analysing the Language of Neural Audio Codecs
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
von: Park, Joonyong, et al.
Veröffentlicht: (2025)
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
von: Deng, Ruifan, et al.
Veröffentlicht: (2025)
von: Deng, Ruifan, et al.
Veröffentlicht: (2025)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
Acoustic Teleportation via Disentangled Neural Audio Codec Representations
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
von: Grundhuber, Philipp, et al.
Veröffentlicht: (2025)
Indic-CodecFake meets SATYAM: Towards Detecting Neural Audio Codec Synthesized Speech Deepfakes in Indic Languages
von: Girish, et al.
Veröffentlicht: (2026)
von: Girish, et al.
Veröffentlicht: (2026)
SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization
von: Wang, Jin, et al.
Veröffentlicht: (2025)
von: Wang, Jin, et al.
Veröffentlicht: (2025)
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025)
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
HILCodec: High-Fidelity and Lightweight Neural Audio Codec
von: Ahn, Sunghwan, et al.
Veröffentlicht: (2024)
von: Ahn, Sunghwan, et al.
Veröffentlicht: (2024)
AudioLog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive Learning
von: Bai, Jisheng, et al.
Veröffentlicht: (2023)
von: Bai, Jisheng, et al.
Veröffentlicht: (2023)
Towards Neural Audio Codec Source Parsing
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2025)
SAGA-SR: Semantically and Acoustically Guided Audio Super-Resolution
von: Im, Jaekwon, et al.
Veröffentlicht: (2025)
von: Im, Jaekwon, et al.
Veröffentlicht: (2025)
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
von: Liu, Haohe, et al.
Veröffentlicht: (2024)
MSR-Codec: A Low-Bitrate Multi-Stream Residual Codec for High-Fidelity Speech Generation with Information Disentanglement
von: Li, Jingyu, et al.
Veröffentlicht: (2025)
von: Li, Jingyu, et al.
Veröffentlicht: (2025)
UniCodec: Unified Audio Codec with Single Domain-Adaptive Codebook
von: Jiang, Yidi, et al.
Veröffentlicht: (2025)
von: Jiang, Yidi, et al.
Veröffentlicht: (2025)
Bringing Interpretability to Neural Audio Codecs
von: Sadok, Samir, et al.
Veröffentlicht: (2025)
von: Sadok, Samir, et al.
Veröffentlicht: (2025)
SemanticAudio: Audio Generation and Editing in Semantic Space
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
von: Dai, Zheqi, et al.
Veröffentlicht: (2026)
Analyzing and Mitigating Inconsistency in Discrete Audio Tokens for Neural Codec Language Models
von: Liu, Wenrui, et al.
Veröffentlicht: (2024)
von: Liu, Wenrui, et al.
Veröffentlicht: (2024)
LongCat-Audio-Codec: An Audio Tokenizer and Detokenizer Solution Designed for Speech Large Language Models
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2025)
von: Zhao, Xiaohan, et al.
Veröffentlicht: (2025)
G-IFT: A Gated Linear Unit adapter with Iterative Fine-Tuning for Low-Resource Children's Speaker Verification
von: Shetty, Vishwas M., et al.
Veröffentlicht: (2025)
von: Shetty, Vishwas M., et al.
Veröffentlicht: (2025)
LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
von: Zhang, Zhisheng, et al.
Veröffentlicht: (2026)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs
von: Langman, Ryan, et al.
Veröffentlicht: (2024)
von: Langman, Ryan, et al.
Veröffentlicht: (2024)
Exploring the Benefits of Tokenization of Discrete Acoustic Units
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
von: Dekel, Avihu, et al.
Veröffentlicht: (2024)
ScoreDec: A Phase-preserving High-Fidelity Audio Codec with A Generalized Score-based Diffusion Post-filter
von: Wu, Yi-Chiao, et al.
Veröffentlicht: (2024)
von: Wu, Yi-Chiao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Comparing Unsupervised and Supervised Semantic Speech Tokens: A Case Study of Child ASR
von: Shi, Mohan, et al.
Veröffentlicht: (2025) -
Selective Attention Merging for low resource tasks: A case study of Child ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025) -
CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
von: Shankar, Natarajan Balaji, et al.
Veröffentlicht: (2025) -
Mind the Shift: Using Delta SSL Embeddings to Enhance Child ASR
von: Wang, Zilai, et al.
Veröffentlicht: (2026) -
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)