The Conformer Encoder May Reverse the Time Dimension
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schmitt, Robin, Zeyer, Albert, Zeineldeen, Mohammad, Schlüter, Ralf, Ney, Hermann |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023)
LLMs and Speech: Integration vs. Combination
von: Schmitt, Robin, et al.
Veröffentlicht: (2026)
von: Schmitt, Robin, et al.
Veröffentlicht: (2026)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
von: Yang, Zijian, et al.
Veröffentlicht: (2026)
von: Yang, Zijian, et al.
Veröffentlicht: (2026)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
von: Raissi, Tina, et al.
Veröffentlicht: (2025)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
Unified Learnable 2D Convolutional Feature Extraction for ASR
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
Label-Context-Dependent Internal Language Model Estimation for CTC
von: Yang, Zijian, et al.
Veröffentlicht: (2025)
von: Yang, Zijian, et al.
Veröffentlicht: (2025)
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
von: Raissi, Tina, et al.
Veröffentlicht: (2024)
von: Raissi, Tina, et al.
Veröffentlicht: (2024)
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2025)
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2025)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
von: Vieting, Peter, et al.
Veröffentlicht: (2023)
von: Vieting, Peter, et al.
Veröffentlicht: (2023)
Exploring System Adaptations For Minimum Latency Real-Time Piano Transcription
von: Hu, Patricia, et al.
Veröffentlicht: (2025)
von: Hu, Patricia, et al.
Veröffentlicht: (2025)
Beat this! Accurate beat tracking without DBN postprocessing
von: Foscarin, Francesco, et al.
Veröffentlicht: (2024)
von: Foscarin, Francesco, et al.
Veröffentlicht: (2024)
Music Boomerang: Reusing Diffusion Models for Data Augmentation and Audio Manipulation
von: Fichtinger, Alexander, et al.
Veröffentlicht: (2025)
von: Fichtinger, Alexander, et al.
Veröffentlicht: (2025)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2024)
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2024)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
Do Foundational Audio Encoders Understand Music Structure?
von: Toyama, Keisuke, et al.
Veröffentlicht: (2025)
von: Toyama, Keisuke, et al.
Veröffentlicht: (2025)
Hold Me Tight: Stable Encoder-Decoder Design for Speech Enhancement
von: Haider, Daniel, et al.
Veröffentlicht: (2024)
von: Haider, Daniel, et al.
Veröffentlicht: (2024)
Voice Conversion with Diverse Intonation using Conditional Variational Auto-Encoder
von: Suh, Soobin, et al.
Veröffentlicht: (2025)
von: Suh, Soobin, et al.
Veröffentlicht: (2025)
WhiSQA: Non-Intrusive Speech Quality Prediction Using Whisper Encoder Features
von: Close, George, et al.
Veröffentlicht: (2025)
von: Close, George, et al.
Veröffentlicht: (2025)
Alternating Weak Triphone/BPE Alignment Supervision from Hybrid Model Improves End-to-End ASR
von: Jiang, Jintao, et al.
Veröffentlicht: (2024)
von: Jiang, Jintao, et al.
Veröffentlicht: (2024)
Optimizing Multi-Stuttered Speech Classification: Leveraging Whisper's Encoder for Efficient Parameter Reduction in Automated Assessment
von: Ameer, Huma, et al.
Veröffentlicht: (2024)
von: Ameer, Huma, et al.
Veröffentlicht: (2024)
Analysis-Driven Procedural Generation of an Engine Sound Dataset with Embedded Control Annotations
von: Doerfler, Robin, et al.
Veröffentlicht: (2026)
von: Doerfler, Robin, et al.
Veröffentlicht: (2026)
Voice Quality Dimensions as Interpretable Primitives for Speaking Style for Atypical Speech and Affect
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
SC-MoE: Switch Conformer Mixture of Experts for Unified Streaming and Non-streaming Code-Switching ASR
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
von: Ye, Shuaishuai, et al.
Veröffentlicht: (2024)
CleanCTG: A Deep Learning Model for Multi-Artefact Detection and Reconstruction in Cardiotocography
von: Wong, Sheng, et al.
Veröffentlicht: (2025)
von: Wong, Sheng, et al.
Veröffentlicht: (2025)
Reverse-Speech-Finder: A Neural Network Backtracking Architecture for Generating Alzheimer's Disease Speech Samples and Improving Diagnosis Performance
von: Li, Victor OK, et al.
Veröffentlicht: (2025)
von: Li, Victor OK, et al.
Veröffentlicht: (2025)
Speech After Gender: A Trans-Feminine Perspective on Next Steps for Speech Science and Technology
von: Netzorg, Robin, et al.
Veröffentlicht: (2024)
von: Netzorg, Robin, et al.
Veröffentlicht: (2024)
Investigation of Time-Frequency Feature Combinations with Histogram Layer Time Delay Neural Networks
von: Mohammadi, Amirmohammad, et al.
Veröffentlicht: (2024)
von: Mohammadi, Amirmohammad, et al.
Veröffentlicht: (2024)
Error Analysis in a Modular Meeting Transcription System
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
von: Vieting, Peter, et al.
Veröffentlicht: (2025)
Test-Time Training for Depression Detection
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Test-Time Training for Speech Enhancement
von: Behera, Avishkar, et al.
Veröffentlicht: (2025)
von: Behera, Avishkar, et al.
Veröffentlicht: (2025)
Fast Timing-Conditioned Latent Audio Diffusion
von: Evans, Zach, et al.
Veröffentlicht: (2024)
von: Evans, Zach, et al.
Veröffentlicht: (2024)
Test-Time Adaptation for Speech Emotion Recognition
von: Dong, Jiaheng, et al.
Veröffentlicht: (2026)
von: Dong, Jiaheng, et al.
Veröffentlicht: (2026)
Differentiable All-pole Filters for Time-varying Audio Systems
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
von: Yu, Chin-Yun, et al.
Veröffentlicht: (2024)
Improving Real-Time Music Accompaniment Separation with MMDenseNet
von: Wang, Chun-Hsiang, et al.
Veröffentlicht: (2024)
von: Wang, Chun-Hsiang, et al.
Veröffentlicht: (2024)
StreamVC: Real-Time Low-Latency Voice Conversion
von: Yang, Yang, et al.
Veröffentlicht: (2024)
von: Yang, Yang, et al.
Veröffentlicht: (2024)
Whispy: Adapting STT Whisper Models to Real-Time Environments
von: Bevilacqua, Antonio, et al.
Veröffentlicht: (2024)
von: Bevilacqua, Antonio, et al.
Veröffentlicht: (2024)
TF-MLPNet: Tiny Real-Time Neural Speech Separation
von: Itani, Malek, et al.
Veröffentlicht: (2025)
von: Itani, Malek, et al.
Veröffentlicht: (2025)
Distribution Preserving Source Separation With Time Frequency Predictive Models
von: T., Pedro J. Villasana, et al.
Veröffentlicht: (2023)
von: T., Pedro J. Villasana, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
von: Zeineldeen, Mohammad, et al.
Veröffentlicht: (2023) -
LLMs and Speech: Integration vs. Combination
von: Schmitt, Robin, et al.
Veröffentlicht: (2026) -
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
von: Yang, Zijian, et al.
Veröffentlicht: (2026) -
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
von: Raissi, Tina, et al.
Veröffentlicht: (2025) -
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
von: Yang, Zijian, et al.
Veröffentlicht: (2023)