Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kloots, Marianne de Heer, Zuidema, Willem |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Linguists should learn to love speech-based deep learning models
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024)
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)
SpecWav-Attack: Leveraging Spectrogram Resizing and Wav2Vec 2.0 for Attacking Anonymized Speech
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
Exploring bat song syllable representations in self-supervised audio encoders
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2024)
Over-the-air White-box Attack on the Wav2Vec Speech Recognition Neural Network
von: Alexey, Protopopov
Veröffentlicht: (2026)
von: Alexey, Protopopov
Veröffentlicht: (2026)
Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
von: Pietroń, Marcin, et al.
Veröffentlicht: (2026)
Disentangling Textual and Acoustic Features of Neural Speech Representations
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
von: Mohebbi, Hosein, et al.
Veröffentlicht: (2024)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
von: Stuhlmann, Linus, et al.
Veröffentlicht: (2025)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Phonetic Segmentation of the UCLA Phonetics Lab Archive
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
von: Chodroff, Eleanor, et al.
Veröffentlicht: (2024)
Tracking the emergence of linguistic structure in self-supervised models learning from speech
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2026)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2026)
Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
von: Ridoy, Md Sazzadul Islam, et al.
Veröffentlicht: (2025)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
von: Chou, Ju-Chieh, et al.
Veröffentlicht: (2023)
von: Chou, Ju-Chieh, et al.
Veröffentlicht: (2023)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
WavMark: Watermarking for Audio Generation
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
von: Chen, Guangyu, et al.
Veröffentlicht: (2023)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
(SimPhon Speech Test): A Data-Driven Method for In Silico Design and Validation of a Phonetically Balanced Speech Test
von: Bleeck, Stefan
Veröffentlicht: (2025)
von: Bleeck, Stefan
Veröffentlicht: (2025)
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
ManWav: The First Manchu ASR Model
von: Seo, Jean, et al.
Veröffentlicht: (2024)
von: Seo, Jean, et al.
Veröffentlicht: (2024)
High-Fidelity Neural Phonetic Posteriorgrams
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
von: Churchwell, Cameron, et al.
Veröffentlicht: (2024)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
von: Deng, Keqi, et al.
Veröffentlicht: (2024)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
von: Han, Seungu, et al.
Veröffentlicht: (2026)
von: Han, Seungu, et al.
Veröffentlicht: (2026)
WavLLM: Towards Robust and Adaptive Speech Large Language Model
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
von: Hu, Shujie, et al.
Veröffentlicht: (2024)
Bailing-TTS: Chinese Dialectal Speech Synthesis Towards Human-like Spontaneous Representation
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
von: Di, Xinhan, et al.
Veröffentlicht: (2024)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
Decoding Linguistic Representations of Human Brain
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
von: Chen, Sanyuan, et al.
Veröffentlicht: (2024)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
von: Sudo, Yui, et al.
Veröffentlicht: (2025)
Multimodal Input Aids a Bayesian Model of Phonetic Learning
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
von: Zhi, Sophia, et al.
Veröffentlicht: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
In-Context Learning Boosts Speech Recognition via Human-like Adaptation to Speakers and Language Varieties
von: Roll, Nathan, et al.
Veröffentlicht: (2025)
von: Roll, Nathan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Linguists should learn to love speech-based deep learning models
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025) -
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025) -
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024) -
Wav2Small: Distilling Wav2Vec2 to 72K parameters for Low-Resource Speech emotion recognition
von: Kounadis-Bastian, Dionyssos, et al.
Veröffentlicht: (2024) -
Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
von: Nguyen, Tuan, et al.
Veröffentlicht: (2024)