Probing neural audio codecs for distinctions among English nuclear tunes
Fuente:
arXiv
Guardado en:
| Autores principales: | Vigneaux, Juan Pablo, Cole, Jennifer |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Sociolinguistic Analysis of Automatic Speech Recognition Bias in Newcastle English
por: Serditova, Dana, et al.
Publicado: (2026)
por: Serditova, Dana, et al.
Publicado: (2026)
Mixer Metaphors: audio interfaces for non-musical applications
por: McNamara, Tace, et al.
Publicado: (2025)
por: McNamara, Tace, et al.
Publicado: (2025)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
por: Kim, Minu, et al.
Publicado: (2025)
por: Kim, Minu, et al.
Publicado: (2025)
BMdataset: A Musicologically Curated LilyPond Dataset
por: Spanio, Matteo, et al.
Publicado: (2026)
por: Spanio, Matteo, et al.
Publicado: (2026)
How BERT Speaks Shakespearean English? Evaluating Historical Bias in Contextual Language Models
por: Cuscito, Miriam, et al.
Publicado: (2024)
por: Cuscito, Miriam, et al.
Publicado: (2024)
LLMs Generate Kitsch
por: Klinge, Xenia, et al.
Publicado: (2026)
por: Klinge, Xenia, et al.
Publicado: (2026)
Comparative Study of Large Language Models on Chinese Film Script Continuation: An Empirical Analysis Based on GPT-5.2 and Qwen-Max
por: Cao, Yuxuan, et al.
Publicado: (2026)
por: Cao, Yuxuan, et al.
Publicado: (2026)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
por: Viswanathan, Janaki, et al.
Publicado: (2025)
por: Viswanathan, Janaki, et al.
Publicado: (2025)
Self-Supervised Borrowing Detection on Multilingual Wordlists
por: Wientzek, Tim
Publicado: (2025)
por: Wientzek, Tim
Publicado: (2025)
Improving French Synthetic Speech Quality via SSML Prosody Control
por: Ouali, Nassima Ould, et al.
Publicado: (2025)
por: Ouali, Nassima Ould, et al.
Publicado: (2025)
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
por: Chien, Sheng-You, et al.
Publicado: (2026)
por: Chien, Sheng-You, et al.
Publicado: (2026)
An accurate and revised version of optical character recognition-based speech synthesis using LabVIEW
por: Mehta, Prateek, et al.
Publicado: (2025)
por: Mehta, Prateek, et al.
Publicado: (2025)
BlasBench: An Open Benchmark for Irish Speech Recognition
por: Raj, Jyoutir, et al.
Publicado: (2026)
por: Raj, Jyoutir, et al.
Publicado: (2026)
Step-Audio-R1 Technical Report
por: Tian, Fei, et al.
Publicado: (2025)
por: Tian, Fei, et al.
Publicado: (2025)
Adversarially Probing Cross-Family Sound Symbolism in 27 Languages
por: Sharma, Anika, et al.
Publicado: (2025)
por: Sharma, Anika, et al.
Publicado: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
por: Alamr, Meshal, et al.
Publicado: (2026)
por: Alamr, Meshal, et al.
Publicado: (2026)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
por: Wang, Hsuan-Yu, et al.
Publicado: (2025)
por: Wang, Hsuan-Yu, et al.
Publicado: (2025)
Language Predicts Identity Fusion Across Cultures and Reveals Divergent Pathways to Violence
por: Wright, Devin R., et al.
Publicado: (2026)
por: Wright, Devin R., et al.
Publicado: (2026)
Where is my Glass Slipper? AI, Poetry and Art
por: Pagiaslis, Anastasios P.
Publicado: (2025)
por: Pagiaslis, Anastasios P.
Publicado: (2025)
Chronic pain patient narratives allow for the estimation of current pain intensity
por: Nunes, Diogo A. P., et al.
Publicado: (2022)
por: Nunes, Diogo A. P., et al.
Publicado: (2022)
Reasoning Over the Glyphs: Evaluation of LLM's Decipherment of Rare Scripts
por: Shih, Yu-Fei, et al.
Publicado: (2025)
por: Shih, Yu-Fei, et al.
Publicado: (2025)
An audio-to-analysis pipeline with certified transcription for information-theoretic profiling of the piano repertoire
por: Jalbert-Desforges, Fred
Publicado: (2026)
por: Jalbert-Desforges, Fred
Publicado: (2026)
Next Token Prediction Is a Dead End for Creativity
por: Olatunji, Ibukun, et al.
Publicado: (2025)
por: Olatunji, Ibukun, et al.
Publicado: (2025)
Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech
por: Floresca, Rez Samantha Z., et al.
Publicado: (2026)
por: Floresca, Rez Samantha Z., et al.
Publicado: (2026)
Modeling Changing Scientific Concepts with Complex Networks: A Case Study on the Chemical Revolution
por: Aguilar-Valdez, Sofía, et al.
Publicado: (2026)
por: Aguilar-Valdez, Sofía, et al.
Publicado: (2026)
Domain Adaptation of the Pyannote Diarization Pipeline for Conversational Indonesian Audio
por: Prasetyo, Muhammad Daffa'i Rafi, et al.
Publicado: (2026)
por: Prasetyo, Muhammad Daffa'i Rafi, et al.
Publicado: (2026)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
por: Hori, Takaaki, et al.
Publicado: (2025)
por: Hori, Takaaki, et al.
Publicado: (2025)
Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset
por: Marie, Ambre, et al.
Publicado: (2025)
por: Marie, Ambre, et al.
Publicado: (2025)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
por: Chowdhury, MD. Sagor, et al.
Publicado: (2026)
por: Chowdhury, MD. Sagor, et al.
Publicado: (2026)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
por: Zaragozá, Lucía Gómez, et al.
Publicado: (2024)
por: Zaragozá, Lucía Gómez, et al.
Publicado: (2024)
Assessing Latency in ASR Systems: A Methodological Perspective for Real-Time Use
por: Arriaga, Carlos, et al.
Publicado: (2024)
por: Arriaga, Carlos, et al.
Publicado: (2024)
Window Size Versus Accuracy Experiments in Voice Activity Detectors
por: McKinnon, Max, et al.
Publicado: (2026)
por: McKinnon, Max, et al.
Publicado: (2026)
The Table of Media Bias Elements: A sentence-level taxonomy of media bias types and propaganda techniques
por: Menzner, Tim, et al.
Publicado: (2026)
por: Menzner, Tim, et al.
Publicado: (2026)
ASR Error Correction in Low-Resource Burmese with Alignment-Enhanced Transformers using Phonetic Features
por: Lin, Ye Bhone, et al.
Publicado: (2025)
por: Lin, Ye Bhone, et al.
Publicado: (2025)
NAAQA: A Neural Architecture for Acoustic Question Answering
por: Abdelnour, Jerome, et al.
Publicado: (2021)
por: Abdelnour, Jerome, et al.
Publicado: (2021)
Cross-lingual Transfer in Programming Languages: An Extensive Empirical Study
por: Baltaji, Razan, et al.
Publicado: (2023)
por: Baltaji, Razan, et al.
Publicado: (2023)
Analysis of LLM as a grammatical feature tagger for African American English
por: Porwal, Rahul, et al.
Publicado: (2025)
por: Porwal, Rahul, et al.
Publicado: (2025)
Leveraging large multimodal models for audio-video deepfake detection: a pilot study
por: Cao, Songjun, et al.
Publicado: (2026)
por: Cao, Songjun, et al.
Publicado: (2026)
BEAT: Tokenizing and Generating Symbolic Music by Uniform Temporal Steps
por: Qian, Lekai, et al.
Publicado: (2026)
por: Qian, Lekai, et al.
Publicado: (2026)
The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS
por: Chen, Kuan-Yu, et al.
Publicado: (2026)
por: Chen, Kuan-Yu, et al.
Publicado: (2026)
Ejemplares similares
-
A Sociolinguistic Analysis of Automatic Speech Recognition Bias in Newcastle English
por: Serditova, Dana, et al.
Publicado: (2026) -
Mixer Metaphors: audio interfaces for non-musical applications
por: McNamara, Tace, et al.
Publicado: (2025) -
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
por: Kim, Minu, et al.
Publicado: (2025) -
BMdataset: A Musicologically Curated LilyPond Dataset
por: Spanio, Matteo, et al.
Publicado: (2026) -
How BERT Speaks Shakespearean English? Evaluating Historical Bias in Contextual Language Models
por: Cuscito, Miriam, et al.
Publicado: (2024)