On the Relation of State Space Models and Hidden Markov Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ghojogh, Aydin, Sepanj, M. Hadi, Ghojogh, Benyamin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Self-Supervised Learning Using Nonlinear Dependence
von: Sepanj, M. Hadi, et al.
Veröffentlicht: (2025)
von: Sepanj, M. Hadi, et al.
Veröffentlicht: (2025)
Self-Supervised Learning by Curvature Alignment
von: Ghojogh, Benyamin, et al.
Veröffentlicht: (2025)
von: Ghojogh, Benyamin, et al.
Veröffentlicht: (2025)
Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative
von: Xuan, Xi, et al.
Veröffentlicht: (2025)
von: Xuan, Xi, et al.
Veröffentlicht: (2025)
Kernel VICReg for Self-Supervised Learning in Reproducing Kernel Hilbert Space
von: Sepanj, M. Hadi, et al.
Veröffentlicht: (2025)
von: Sepanj, M. Hadi, et al.
Veröffentlicht: (2025)
ETTA: Elucidating the Design Space of Text-to-Audio Models
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
von: Varshavsky-Hassid, Miri, et al.
Veröffentlicht: (2024)
von: Varshavsky-Hassid, Miri, et al.
Veröffentlicht: (2024)
A New Class of Efficient Adaptive Filters for Online Nonlinear Modeling
von: Comminiello, Danilo, et al.
Veröffentlicht: (2021)
von: Comminiello, Danilo, et al.
Veröffentlicht: (2021)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
von: Yang, Zijian, et al.
Veröffentlicht: (2023)
A Joint Spectro-Temporal Relational Thinking Based Acoustic Modeling Framework
von: Nan, Zheng, et al.
Veröffentlicht: (2024)
von: Nan, Zheng, et al.
Veröffentlicht: (2024)
Integrating Self-supervised Speech Model with Pseudo Word-level Targets from Visually-grounded Speech Model
von: Fang, Hung-Chieh, et al.
Veröffentlicht: (2024)
von: Fang, Hung-Chieh, et al.
Veröffentlicht: (2024)
Gender Bias in Instruction-Guided Speech Synthesis Models
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2025)
A Contrastive Learning Approach to Mitigate Bias in Speech Models
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
von: Koudounas, Alkis, et al.
Veröffentlicht: (2024)
SDiaReward: Modeling and Benchmarking Spoken Dialogue Rewards with Modality and Colloquialness
von: Lu, Jingyu, et al.
Veröffentlicht: (2026)
von: Lu, Jingyu, et al.
Veröffentlicht: (2026)
What Do Self-Supervised Speech Models Know About Words?
von: Pasad, Ankita, et al.
Veröffentlicht: (2023)
von: Pasad, Ankita, et al.
Veröffentlicht: (2023)
Memory-Efficient Training for Text-Dependent SV with Independent Pre-trained Models
von: Farokh, Seyed Ali, et al.
Veröffentlicht: (2024)
von: Farokh, Seyed Ali, et al.
Veröffentlicht: (2024)
Exploring Speech Foundation Models for Speaker Diarization in Child-Adult Dyadic Interactions
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
von: Xu, Anfeng, et al.
Veröffentlicht: (2024)
A Multimodal Approach to Device-Directed Speech Detection with Large Language Models
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
von: Wagner, Dominik, et al.
Veröffentlicht: (2024)
Initial Decoding with Minimally Augmented Language Model for Improved Lattice Rescoring in Low Resource ASR
von: Murthy, Savitha, et al.
Veröffentlicht: (2024)
von: Murthy, Savitha, et al.
Veröffentlicht: (2024)
Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing
von: Trinh, Viet Anh, et al.
Veröffentlicht: (2024)
von: Trinh, Viet Anh, et al.
Veröffentlicht: (2024)
NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
von: Lin, Yen-Ting, et al.
Veröffentlicht: (2024)
von: Lin, Yen-Ting, et al.
Veröffentlicht: (2024)
The State Of TTS: A Case Study with Human Fooling Rates
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
von: Varadhan, Praveen Srinivasa, et al.
Veröffentlicht: (2025)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
von: Ngo, Huong, et al.
Veröffentlicht: (2025)
Modeling Overlapped Speech with Shuffles
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
von: Wiesner, Matthew, et al.
Veröffentlicht: (2026)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
von: Qian, Mengjie, et al.
Veröffentlicht: (2024)
Textually Pretrained Speech Language Models
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
Objective Soups: Multilingual Multi-Task Modeling for Speech Processing
von: Saif, A F M, et al.
Veröffentlicht: (2025)
von: Saif, A F M, et al.
Veröffentlicht: (2025)
Coupling Speech Encoders with Downstream Text Models
von: Chelba, Ciprian, et al.
Veröffentlicht: (2024)
von: Chelba, Ciprian, et al.
Veröffentlicht: (2024)
Improving Text-To-Audio Models with Synthetic Captions
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
von: Kong, Zhifeng, et al.
Veröffentlicht: (2024)
Revisiting ASR Error Correction with Specialized Models
von: Gu, Zijin, et al.
Veröffentlicht: (2024)
von: Gu, Zijin, et al.
Veröffentlicht: (2024)
Bayesian Low-Rank Factorization for Robust Model Adaptation
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
Energy-Based Models with Applications to Speech and Language Processing
von: Ou, Zhijian
Veröffentlicht: (2024)
von: Ou, Zhijian
Veröffentlicht: (2024)
How Redundant Is the Transformer Stack in Speech Representation Models?
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
von: Dorszewski, Teresa, et al.
Veröffentlicht: (2024)
Federated Learning of Large ASR Models in the Real World
von: Xiao, Yonghui, et al.
Veröffentlicht: (2024)
von: Xiao, Yonghui, et al.
Veröffentlicht: (2024)
Aligning Spoken Dialogue Models from User Interactions
von: Wu, Anne, et al.
Veröffentlicht: (2025)
von: Wu, Anne, et al.
Veröffentlicht: (2025)
Audio-to-Score Conversion Model Based on Whisper methodology
von: Zhang, Hongyao, et al.
Veröffentlicht: (2024)
von: Zhang, Hongyao, et al.
Veröffentlicht: (2024)
Foundations of Riemannian Geometry for Riemannian Optimization: A Monograph with Detailed Derivations
von: Ghojogh, Benyamin
Veröffentlicht: (2026)
von: Ghojogh, Benyamin
Veröffentlicht: (2026)
WhisperRT -- Turning Whisper into a Causal Streaming Model
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
von: Krichli, Tomer, et al.
Veröffentlicht: (2025)
Label-Context-Dependent Internal Language Model Estimation for CTC
von: Yang, Zijian, et al.
Veröffentlicht: (2025)
von: Yang, Zijian, et al.
Veröffentlicht: (2025)
tinyCLAP: Distilling Constrastive Language-Audio Pretrained Models
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
von: Paissan, Francesco, et al.
Veröffentlicht: (2023)
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Self-Supervised Learning Using Nonlinear Dependence
von: Sepanj, M. Hadi, et al.
Veröffentlicht: (2025) -
Self-Supervised Learning by Curvature Alignment
von: Ghojogh, Benyamin, et al.
Veröffentlicht: (2025) -
Fake-Mamba: Real-Time Speech Deepfake Detection Using Bidirectional Mamba as Self-Attention's Alternative
von: Xuan, Xi, et al.
Veröffentlicht: (2025) -
Kernel VICReg for Self-Supervised Learning in Reproducing Kernel Hilbert Space
von: Sepanj, M. Hadi, et al.
Veröffentlicht: (2025) -
ETTA: Elucidating the Design Space of Text-to-Audio Models
von: Lee, Sang-gil, et al.
Veröffentlicht: (2024)