Gespeichert in:
| Hauptverfasser: | Shi, Xuan, Zeng, Chang, Feng, Tiantian, Wang, Shih-Heng, Ma, Jianbo, Narayanan, Shrikanth |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2603.10371 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
von: Feng, Tiantian, et al.
Veröffentlicht: (2023)
Examining Test-Time Adaptation for Personalized Child Speech Recognition
von: Shi, Zhonghao, et al.
Veröffentlicht: (2024)
von: Shi, Zhonghao, et al.
Veröffentlicht: (2024)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
Large Language Models based ASR Error Correction for Child Conversations
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
von: Nguyen, Hong, et al.
Veröffentlicht: (2024)
von: Nguyen, Hong, et al.
Veröffentlicht: (2024)
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
von: Prescott, Jordan, et al.
Veröffentlicht: (2026)
von: Prescott, Jordan, et al.
Veröffentlicht: (2026)
Can Synthetic Audio From Generative Foundation Models Assist Audio Recognition and Speech Modeling?
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
Self-Supervised Speech Representations are More Phonetic than Semantic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2024)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
Joint ASR and Speaker Role Tagging with Serialized Output Training
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
STACodec: Semantic Token Assignment for Balancing Acoustic Fidelity and Semantic Information in Audio Codecs
von: Zhang, Kaiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Kaiyuan, et al.
Veröffentlicht: (2026)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
von: Zhang, Hezhao, et al.
Veröffentlicht: (2026)
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
von: Feng, Tiantian, et al.
Veröffentlicht: (2024)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
von: Lee, Jihwan, et al.
Veröffentlicht: (2024)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
von: Feng, Tiantian, et al.
Veröffentlicht: (2026)
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
von: Tsai, Yun-Shao, et al.
Veröffentlicht: (2026)
von: Tsai, Yun-Shao, et al.
Veröffentlicht: (2026)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
von: Baas, Matthew, et al.
Veröffentlicht: (2025)
von: Baas, Matthew, et al.
Veröffentlicht: (2025)
An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production
von: Lee, Jihwan, et al.
Veröffentlicht: (2026)
von: Lee, Jihwan, et al.
Veröffentlicht: (2026)
Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
von: Yang, Yifan, et al.
Veröffentlicht: (2025)
Emotion-Aligned Contrastive Learning Between Images and Music
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
von: Stewart, Shanti, et al.
Veröffentlicht: (2023)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
Can a Machine Distinguish High and Low Amount of Social Creak in Speech?
von: Laukkanen, Anne-Maria, et al.
Veröffentlicht: (2024)
von: Laukkanen, Anne-Maria, et al.
Veröffentlicht: (2024)
Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test
von: Bleeck, Stefan
Veröffentlicht: (2025)
von: Bleeck, Stefan
Veröffentlicht: (2025)
Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
von: Feng, Tiantian, et al.
Veröffentlicht: (2025)
FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025)
von: Ahasan, Md Mubtasim, et al.
Veröffentlicht: (2025)
End-to-End Joint ASR and Speaker Role Diarization with Child-Adult Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)
NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2025)
von: Casanova, Edresson, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TI-ASU: Toward Robust Automatic Speech Understanding through Text-to-speech Imputation Against Missing Speech Modality
von: Feng, Tiantian, et al.
Veröffentlicht: (2024) -
PEFT-SER: On the Use of Parameter Efficient Transfer Learning Approaches For Speech Emotion Recognition Using Pre-trained Speech Models
von: Feng, Tiantian, et al.
Veröffentlicht: (2023) -
Examining Test-Time Adaptation for Personalized Child Speech Recognition
von: Shi, Zhonghao, et al.
Veröffentlicht: (2024) -
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
von: Feng, Tiantian, et al.
Veröffentlicht: (2025) -
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
von: Xu, Anfeng, et al.
Veröffentlicht: (2026)