Languages in Whisper-Style Speech Encoders Align Both Phonetically and Semantically
Fuente:
arXiv
Saved in:
| Main Authors: | Shim, Ryan Soh-Eun, De Cristofaro, Domenico, Hu, Chengzhi Martin, Vietti, Alessandro, Plank, Barbara |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
by: De Cristofaro, Domenico, et al.
Published: (2025)
by: De Cristofaro, Domenico, et al.
Published: (2025)
When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding
by: De Cristofaro, Domenico, et al.
Published: (2026)
by: De Cristofaro, Domenico, et al.
Published: (2026)
Rashid: A Cipher-Based Framework for Exploring In-Context Language Learning
by: Bafna, Niyati, et al.
Published: (2026)
by: Bafna, Niyati, et al.
Published: (2026)
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
by: Li, Chin-Jou, et al.
Published: (2025)
by: Li, Chin-Jou, et al.
Published: (2025)
SteerEval: Inference-time Interventions Strengthen Multilingual Generalization in Neural Summarization Metrics
by: Casola, Silvia, et al.
Published: (2026)
by: Casola, Silvia, et al.
Published: (2026)
Surgical, Cheap, and Flexible: Mitigating False Refusal in Language Models via Single Vector Ablation
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
Phonotactic Complexity across Dialects
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
by: Shim, Ryan Soh-Eun, et al.
Published: (2024)
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration
by: Shim, Ryan Soh-Eun, et al.
Published: (2026)
by: Shim, Ryan Soh-Eun, et al.
Published: (2026)
Look at the Text: Instruction-Tuned Language Models are More Robust Multiple Choice Selectors than You Think
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
Speech Codec Probing from Semantic and Phonetic Perspectives
by: Shi, Xuan, et al.
Published: (2026)
by: Shi, Xuan, et al.
Published: (2026)
Noise-Robust AV-ASR Using Visual Features Both in the Whisper Encoder and Decoder
by: Li, Zhengyang, et al.
Published: (2026)
by: Li, Zhengyang, et al.
Published: (2026)
Refusal Direction is Universal Across Safety-Aligned Languages
by: Wang, Xinpeng, et al.
Published: (2025)
by: Wang, Xinpeng, et al.
Published: (2025)
Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context
by: Trinh, Viet Anh, et al.
Published: (2025)
by: Trinh, Viet Anh, et al.
Published: (2025)
Self-Supervised Speech Representations are More Phonetic than Semantic
by: Choi, Kwanghee, et al.
Published: (2024)
by: Choi, Kwanghee, et al.
Published: (2024)
PHISH in MESH: Korean Adversarial Phonetic Substitution and Phonetic-Semantic Feature Integration Defense
by: Kim, Byungjun, et al.
Published: (2025)
by: Kim, Byungjun, et al.
Published: (2025)
Phonetic Modeling of Dialectal Variation in Vietnamese Speech
by: Hoang, Quan Ngoc, et al.
Published: (2026)
by: Hoang, Quan Ngoc, et al.
Published: (2026)
Whispering Context: Distilling Syntax and Semantics for Long Speech Transcripts
by: Altinok, Duygu
Published: (2025)
by: Altinok, Duygu
Published: (2025)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
by: Zhou, Kun, et al.
Published: (2024)
by: Zhou, Kun, et al.
Published: (2024)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
by: Sudo, Yui, et al.
Published: (2025)
by: Sudo, Yui, et al.
Published: (2025)
Liar, Liar, Logical Mire: A Benchmark for Suppositional Reasoning in Large Language Models
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
"My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
by: Wang, Xinpeng, et al.
Published: (2024)
by: Wang, Xinpeng, et al.
Published: (2024)
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
by: Blaschke, Verena, et al.
Published: (2025)
by: Blaschke, Verena, et al.
Published: (2025)
Calm-Whisper: Reduce Whisper Hallucination On Non-Speech By Calming Crazy Heads Down
by: Wang, Yingzhi, et al.
Published: (2025)
by: Wang, Yingzhi, et al.
Published: (2025)
Careless Whisper: Speech-to-Text Hallucination Harms
by: Koenecke, Allison, et al.
Published: (2024)
by: Koenecke, Allison, et al.
Published: (2024)
WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning
by: Rao, Rajath, et al.
Published: (2025)
by: Rao, Rajath, et al.
Published: (2025)
Comparing Inferential Strategies of Humans and Large Language Models in Deductive Reasoning
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
Aligning NLP Models with Target Population Perspectives using PAIR: Population-Aligned Instance Replication
by: Eckman, Stephanie, et al.
Published: (2025)
by: Eckman, Stephanie, et al.
Published: (2025)
Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
by: Mondorf, Philipp, et al.
Published: (2024)
by: Mondorf, Philipp, et al.
Published: (2024)
PAST: Phonetic-Acoustic Speech Tokenizer
by: Har-Tuv, Nadav, et al.
Published: (2025)
by: Har-Tuv, Nadav, et al.
Published: (2025)
Approaching Dialogue State Tracking via Aligning Speech Encoders and LLMs
by: Sedláček, Šimon, et al.
Published: (2025)
by: Sedláček, Šimon, et al.
Published: (2025)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
by: Peng, Yifan, et al.
Published: (2025)
by: Peng, Yifan, et al.
Published: (2025)
OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer
by: Peng, Yifan, et al.
Published: (2024)
by: Peng, Yifan, et al.
Published: (2024)
If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models
by: Orth, Jasmin, et al.
Published: (2025)
by: Orth, Jasmin, et al.
Published: (2025)
Unlocking Fine-Grained and Within-Utterance Speaking Style Control in Prompt-Based Text-to-Speech Models
by: Kang, Jaehoon, et al.
Published: (2026)
by: Kang, Jaehoon, et al.
Published: (2026)
Whispering in Amharic: Fine-tuning Whisper for Low-resource Language
by: Gete, Dawit Ketema, et al.
Published: (2025)
by: Gete, Dawit Ketema, et al.
Published: (2025)
Disagreeing Rationales: Rethinking Classification and Explainability Evaluation in Hate Speech Detection
by: Muscato, Benedetta, et al.
Published: (2026)
by: Muscato, Benedetta, et al.
Published: (2026)
StyleBench: Evaluating Speech Language Models on Conversational Speaking Style Control
by: Zhao, Haishu, et al.
Published: (2026)
by: Zhao, Haishu, et al.
Published: (2026)
Polynomial Mixing for Efficient Self-supervised Speech Encoders
by: Feillet, Eva, et al.
Published: (2026)
by: Feillet, Eva, et al.
Published: (2026)
Indirect Question Answering in English, German and Bavarian: A Challenging Task for High- and Low-Resource Languages Alike
by: Winkler, Miriam, et al.
Published: (2026)
by: Winkler, Miriam, et al.
Published: (2026)
Similar Items
-
Dialetto, ma Quanto Dialetto? Transcribing and Evaluating Dialects on a Continuum
by: Shim, Ryan Soh-Eun, et al.
Published: (2024) -
Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
by: De Cristofaro, Domenico, et al.
Published: (2025) -
When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding
by: De Cristofaro, Domenico, et al.
Published: (2026) -
Rashid: A Cipher-Based Framework for Exploring In-Context Language Learning
by: Bafna, Niyati, et al.
Published: (2026) -
POWSM: A Phonetic Open Whisper-Style Speech Foundation Model
by: Li, Chin-Jou, et al.
Published: (2025)