Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits
Fuente:
arXiv
Salvato in:
| Autori principali: | Feng, Tiantian, Lee, Jihwan, Xu, Anfeng, Lee, Yoonjeong, Lertpetchpun, Thanathai, Shi, Xuan, Wang, Helin, Thebaud, Thomas, Moro-Velazquez, Laureano, Byrd, Dani, Dehak, Najim, Narayanan, Shrikanth |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2026)
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2026)
Learning-free L2-Accented Speech Generation using Phonological Rules
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2026)
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2026)
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
di: Feng, Tiantian, et al.
Pubblicazione: (2025)
Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2025)
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2025)
ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
di: Lee, Jihwan, et al.
Pubblicazione: (2025)
di: Lee, Jihwan, et al.
Pubblicazione: (2025)
Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2026)
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2026)
VoxGuard: Evaluating User and Attribute Privacy in Speech via Membership Inference Attacks
di: Tsaprazlis, Efthymios, et al.
Pubblicazione: (2025)
di: Tsaprazlis, Efthymios, et al.
Pubblicazione: (2025)
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
di: Thebaud, Thomas, et al.
Pubblicazione: (2026)
di: Thebaud, Thomas, et al.
Pubblicazione: (2026)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
di: Xu, Anfeng, et al.
Pubblicazione: (2026)
Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
di: Lee, Junhyeok, et al.
Pubblicazione: (2026)
di: Lee, Junhyeok, et al.
Pubblicazione: (2026)
Noise-robust Speech Separation with Fast Generative Correction
di: Wang, Helin, et al.
Pubblicazione: (2024)
di: Wang, Helin, et al.
Pubblicazione: (2024)
Trade-offs Between Capacity and Robustness in Neural Audio Codecs for Adversarially Robust Speech Recognition
di: Prescott, Jordan, et al.
Pubblicazione: (2026)
di: Prescott, Jordan, et al.
Pubblicazione: (2026)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
di: Wang, Helin, et al.
Pubblicazione: (2025)
di: Wang, Helin, et al.
Pubblicazione: (2025)
On the Relationship between Accent Strength and Articulatory Features
di: Huang, Kevin, et al.
Pubblicazione: (2025)
di: Huang, Kevin, et al.
Pubblicazione: (2025)
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
di: Lee, Junhyeok, et al.
Pubblicazione: (2025)
di: Lee, Junhyeok, et al.
Pubblicazione: (2025)
VoxCog: Towards End-to-End Multilingual Cognitive Impairment Classification through Dialectal Knowledge
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
di: Feng, Tiantian, et al.
Pubblicazione: (2026)
Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech
di: Chaparala, Kaavya, et al.
Pubblicazione: (2026)
di: Chaparala, Kaavya, et al.
Pubblicazione: (2026)
SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline
di: Wang, Helin, et al.
Pubblicazione: (2025)
di: Wang, Helin, et al.
Pubblicazione: (2025)
Interpretable Modeling of Articulatory Temporal Dynamics from real-time MRI for Phoneme Recognition
di: Park, Jay, et al.
Pubblicazione: (2025)
di: Park, Jay, et al.
Pubblicazione: (2025)
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
di: Xu, Anfeng, et al.
Pubblicazione: (2024)
Demographic Attributes Prediction from Speech Using WavLM Embeddings
di: Yang, Yuchen, et al.
Pubblicazione: (2025)
di: Yang, Yuchen, et al.
Pubblicazione: (2025)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
di: Lu, Yen-Ju, et al.
Pubblicazione: (2024)
di: Lu, Yen-Ju, et al.
Pubblicazione: (2024)
Cross-Corpus and Cross-domain Handwriting Assessment of NeuroDegenerative Diseases via Time-Series-to-Image Conversion
di: Chavez, Gabrielle, et al.
Pubblicazione: (2025)
di: Chavez, Gabrielle, et al.
Pubblicazione: (2025)
Joint ASR and Speaker Role Tagging with Serialized Output Training
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
DiT-Flow: Speech Enhancement Robust to Multiple Distortions based on Flow Matching in Latent Space and Diffusion Transformers
di: Cao, Tianyu, et al.
Pubblicazione: (2026)
di: Cao, Tianyu, et al.
Pubblicazione: (2026)
Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
Study of Pre-processing Defenses against Adversarial Attacks on State-of-the-art Speaker Recognition Systems
di: Joshi, Sonal, et al.
Pubblicazione: (2021)
di: Joshi, Sonal, et al.
Pubblicazione: (2021)
Rhythm Features for Speaker Identification
di: Mehlman, Nick, et al.
Pubblicazione: (2025)
di: Mehlman, Nick, et al.
Pubblicazione: (2025)
Detecting Neurodegenerative Diseases using Frame-Level Handwriting Embeddings
di: Laouedj, Sarah, et al.
Pubblicazione: (2025)
di: Laouedj, Sarah, et al.
Pubblicazione: (2025)
Dynamics of Handwriting for Cognitive Assessment
di: Gabrielle Chavez, et al.
Pubblicazione: (2024)
di: Gabrielle Chavez, et al.
Pubblicazione: (2024)
Analyzing Attention Focus in the Cookie TheftPicture Description Task Using Word Alignment
di: Anna Favaro, et al.
Pubblicazione: (2024)
di: Anna Favaro, et al.
Pubblicazione: (2024)
Cognitive Assessment through Writing Tasks
di: Casey Chen, et al.
Pubblicazione: (2024)
di: Casey Chen, et al.
Pubblicazione: (2024)
Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
di: Trachu, Thanapat, et al.
Pubblicazione: (2025)
di: Trachu, Thanapat, et al.
Pubblicazione: (2025)
Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
di: Feng, Tiantian, et al.
Pubblicazione: (2024)
Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
di: Lu, Yen-Ju, et al.
Pubblicazione: (2025)
Enhancing Listened Speech Decoding from EEG via Parallel Phoneme Sequence Prediction
di: Lee, Jihwan, et al.
Pubblicazione: (2025)
di: Lee, Jihwan, et al.
Pubblicazione: (2025)
Towards Interpretable Framework for Neural Audio Codecs via Sparse Autoencoders: A Case Study on Accent Information
di: Wang, Shih-Heng, et al.
Pubblicazione: (2026)
di: Wang, Shih-Heng, et al.
Pubblicazione: (2026)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2026) -
Learning-free L2-Accented Speech Generation using Phonological Rules
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2026) -
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe
di: Feng, Tiantian, et al.
Pubblicazione: (2025) -
Developing a Top-tier Framework in Naturalistic Conditions Challenge for Categorized Emotion Prediction: From Speech Foundation Models and Learning Objective to Data Augmentation and Engineering Choices
di: Feng, Tiantian, et al.
Pubblicazione: (2025) -
Developing a High-performance Framework for Speech Emotion Recognition in Naturalistic Conditions Challenge for Emotional Attribute Prediction
di: Lertpetchpun, Thanathai, et al.
Pubblicazione: (2025)