Segment Boundary Detection via Class Entropy Measurements in Connectionist Phoneme Recognition
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Salvi, Giampiero |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints
von: Salvi, Giampiero
Veröffentlicht: (2024)
von: Salvi, Giampiero
Veröffentlicht: (2024)
Developing Acoustic Models for Automatic Speech Recognition in Swedish
von: Salvi, Giampiero
Veröffentlicht: (2024)
von: Salvi, Giampiero
Veröffentlicht: (2024)
NAAQA: A Neural Architecture for Acoustic Question Answering
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021)
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021)
Graph Connectionist Temporal Classification for Phoneme Recognition
von: Grafé, Henry, et al.
Veröffentlicht: (2025)
von: Grafé, Henry, et al.
Veröffentlicht: (2025)
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2025)
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2025)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device
von: Kozak, Nazar
Veröffentlicht: (2026)
von: Kozak, Nazar
Veröffentlicht: (2026)
Beyond Deep Learning: Speech Segmentation and Phone Classification with Neural Assemblies
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
von: Adelson, Trevor, et al.
Veröffentlicht: (2026)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
Less Stress, More Privacy: Stress Detection on Anonymized Speech of Air Traffic Controllers
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
von: Viswanathan, Janaki, et al.
Veröffentlicht: (2025)
Enhancing Speech Emotion Recognition Leveraging Aligning Timestamps of ASR Transcripts and Speaker Diarization
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
von: Wang, Hsuan-Yu, et al.
Veröffentlicht: (2025)
Impact of Phonetics on Speaker Identity in Adversarial Voice Attack
von: Dar, Daniyal Kabir, et al.
Veröffentlicht: (2025)
von: Dar, Daniyal Kabir, et al.
Veröffentlicht: (2025)
An End-to-End Approach for Korean Wakeword Systems with Speaker Authentication
von: Seo, Geonwoo
Veröffentlicht: (2025)
von: Seo, Geonwoo
Veröffentlicht: (2025)
Generation of Musical Timbres using a Text-Guided Diffusion Model
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
von: Yuan, Weixuan, et al.
Veröffentlicht: (2025)
Self-Improvement for Audio Large Language Model using Unlabeled Speech
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
von: Wang, Shaowen, et al.
Veröffentlicht: (2025)
MAIN-VC: Lightweight Speech Representation Disentanglement for One-shot Voice Conversion
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
von: Li, Pengcheng, et al.
Veröffentlicht: (2024)
Measuring the Accuracy of Automatic Speech Recognition Solutions
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
von: Kuhn, Korbinian, et al.
Veröffentlicht: (2024)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
SeQuiFi: Mitigating Catastrophic Forgetting in Speech Emotion Recognition with Sequential Class-Finetuning
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
von: Jain, Sarthak, et al.
Veröffentlicht: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
Quantifying the effect of speech pathology on automatic and human speaker verification
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
von: Halpern, Bence Mark, et al.
Veröffentlicht: (2024)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Quantization for OpenAI's Whisper Models: A Comparative Analysis
von: Andreyev, Allison
Veröffentlicht: (2025)
von: Andreyev, Allison
Veröffentlicht: (2025)
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
von: Bhadra, Dipayan, et al.
Veröffentlicht: (2025)
von: Bhadra, Dipayan, et al.
Veröffentlicht: (2025)
Taming Audio VAEs via Target-KL Regularization
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
von: Seetharaman, Prem, et al.
Veröffentlicht: (2026)
A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction
von: Cheripally, Sowmya
Veröffentlicht: (2024)
von: Cheripally, Sowmya
Veröffentlicht: (2024)
SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset
von: Gautam, Sushant, et al.
Veröffentlicht: (2024)
von: Gautam, Sushant, et al.
Veröffentlicht: (2024)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
von: Ahn, Taekyung, et al.
Veröffentlicht: (2024)
Improving Speech Recognition Accuracy Using Custom Language Models with the Vosk Toolkit
von: Soni, Aniket Abhishek
Veröffentlicht: (2025)
von: Soni, Aniket Abhishek
Veröffentlicht: (2025)
STRUM: A Spectral Transcription and Rhythm Understanding Model for End-to-End Generation of Playable Rhythm-Game Charts
von: Opria, Joshua
Veröffentlicht: (2026)
von: Opria, Joshua
Veröffentlicht: (2026)
SW-ASR: A Context-Aware Hybrid ASR Pipeline for Robust Single Word Speech Recognition
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
von: Sharma, Manali, et al.
Veröffentlicht: (2026)
Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
PerceiverS: A Multi-Scale Perceiver with Effective Segmentation for Long-Term Expressive Symbolic Music Generation
von: Yi, Yungang, et al.
Veröffentlicht: (2024)
von: Yi, Yungang, et al.
Veröffentlicht: (2024)
SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
von: Chen, Kuan-Yu, et al.
Veröffentlicht: (2025)
SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
von: Donepudi, Dharma Teja
Veröffentlicht: (2025)
Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
AI-based Drone Assisted Human Rescue in Disaster Environments: Challenges and Opportunities
von: Papyan, Narek, et al.
Veröffentlicht: (2024)
von: Papyan, Narek, et al.
Veröffentlicht: (2024)
Pictures Of MIDI: Controlled Music Generation via Graphical Prompts for Image-Based Diffusion Inpainting
von: Hawley, Scott H.
Veröffentlicht: (2024)
von: Hawley, Scott H.
Veröffentlicht: (2024)
Representation Loss Minimization with Randomized Selection Strategy for Efficient Environmental Fake Audio Detection
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
von: Phukan, Orchid Chetia, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints
von: Salvi, Giampiero
Veröffentlicht: (2024) -
Developing Acoustic Models for Automatic Speech Recognition in Swedish
von: Salvi, Giampiero
Veröffentlicht: (2024) -
NAAQA: A Neural Architecture for Acoustic Question Answering
von: Abdelnour, Jerome, et al.
Veröffentlicht: (2021) -
Graph Connectionist Temporal Classification for Phoneme Recognition
von: Grafé, Henry, et al.
Veröffentlicht: (2025) -
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2025)