Digits micro-model for accurate and secure transactions
Fuente:
arXiv
Salvato in:
| Autori principali: | Chhablani, Chirag, Sharma, Nikhita, Hosier, Jordan, Gurbani, Vijay K. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
di: Ngo, Huong, et al.
Pubblicazione: (2025)
di: Ngo, Huong, et al.
Pubblicazione: (2025)
Gradient boundaries through confidence intervals for forced alignment estimates using model ensembles
di: Kelley, Matthew C.
Pubblicazione: (2025)
di: Kelley, Matthew C.
Pubblicazione: (2025)
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
di: You, Jian, et al.
Pubblicazione: (2024)
di: You, Jian, et al.
Pubblicazione: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech
di: R, Aron, et al.
Pubblicazione: (2024)
di: R, Aron, et al.
Pubblicazione: (2024)
Text-only adaptation in LLM-based ASR through text denoising
di: Carofilis, Andrés, et al.
Pubblicazione: (2026)
di: Carofilis, Andrés, et al.
Pubblicazione: (2026)
ETTA: Elucidating the Design Space of Text-to-Audio Models
di: Lee, Sang-gil, et al.
Pubblicazione: (2024)
di: Lee, Sang-gil, et al.
Pubblicazione: (2024)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
di: Park, Taejin, et al.
Pubblicazione: (2024)
di: Park, Taejin, et al.
Pubblicazione: (2024)
Introduction to speech recognition
di: Dauphin, Gabriel
Pubblicazione: (2024)
di: Dauphin, Gabriel
Pubblicazione: (2024)
AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
di: Kim, Jongsuk, et al.
Pubblicazione: (2024)
di: Kim, Jongsuk, et al.
Pubblicazione: (2024)
An Analysis of Linear Complexity Attention Substitutes with BEST-RQ
di: Whetten, Ryan, et al.
Pubblicazione: (2024)
di: Whetten, Ryan, et al.
Pubblicazione: (2024)
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
di: Liu, Alexander H., et al.
Pubblicazione: (2024)
di: Liu, Alexander H., et al.
Pubblicazione: (2024)
On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models
di: Varshavsky-Hassid, Miri, et al.
Pubblicazione: (2024)
di: Varshavsky-Hassid, Miri, et al.
Pubblicazione: (2024)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
di: Shah, Muhammad A., et al.
Pubblicazione: (2024)
di: Shah, Muhammad A., et al.
Pubblicazione: (2024)
Property Neurons in Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
di: Veluri, Bandhav, et al.
Pubblicazione: (2024)
di: Veluri, Bandhav, et al.
Pubblicazione: (2024)
Rethinking MUSHRA: Addressing Modern Challenges in Text-to-Speech Evaluation
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2024)
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2024)
TelME: Teacher-leading Multimodal Fusion Network for Emotion Recognition in Conversation
di: Yun, Taeyang, et al.
Pubblicazione: (2024)
di: Yun, Taeyang, et al.
Pubblicazione: (2024)
Africa-Centric Self-Supervised Pre-Training for Multilingual Speech Representation in a Sub-Saharan Context
di: Caubrière, Antoine, et al.
Pubblicazione: (2024)
di: Caubrière, Antoine, et al.
Pubblicazione: (2024)
Usefulness of Emotional Prosody in Neural Machine Translation
di: Brazier, Charles, et al.
Pubblicazione: (2024)
di: Brazier, Charles, et al.
Pubblicazione: (2024)
Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping
di: Zhang, Kevin, et al.
Pubblicazione: (2024)
di: Zhang, Kevin, et al.
Pubblicazione: (2024)
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese
di: Arisaputra, Panji, et al.
Pubblicazione: (2024)
di: Arisaputra, Panji, et al.
Pubblicazione: (2024)
Exploring Pathological Speech Quality Assessment with ASR-Powered Wav2Vec2 in Data-Scarce Context
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
di: Nguyen, Tuan, et al.
Pubblicazione: (2024)
Energy-Based Models with Applications to Speech and Language Processing
di: Ou, Zhijian
Pubblicazione: (2024)
di: Ou, Zhijian
Pubblicazione: (2024)
Disentangling Textual and Acoustic Features of Neural Speech Representations
di: Mohebbi, Hosein, et al.
Pubblicazione: (2024)
di: Mohebbi, Hosein, et al.
Pubblicazione: (2024)
Moonshine: Speech Recognition for Live Transcription and Voice Commands
di: Jeffries, Nat, et al.
Pubblicazione: (2024)
di: Jeffries, Nat, et al.
Pubblicazione: (2024)
Investigating Disentanglement in a Phoneme-level Speech Codec for Prosody Modeling
di: Karapiperis, Sotirios, et al.
Pubblicazione: (2024)
di: Karapiperis, Sotirios, et al.
Pubblicazione: (2024)
Robust and Explainable Depression Identification from Speech Using Vowel-Based Ensemble Learning Approaches
di: Feng, Kexin, et al.
Pubblicazione: (2024)
di: Feng, Kexin, et al.
Pubblicazione: (2024)
How Redundant Is the Transformer Stack in Speech Representation Models?
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
Papez: Resource-Efficient Speech Separation with Auditory Working Memory
di: Oh, Hyunseok, et al.
Pubblicazione: (2024)
di: Oh, Hyunseok, et al.
Pubblicazione: (2024)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
di: Anand, Srija, et al.
Pubblicazione: (2024)
di: Anand, Srija, et al.
Pubblicazione: (2024)
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
di: Li, Weiqin, et al.
Pubblicazione: (2024)
di: Li, Weiqin, et al.
Pubblicazione: (2024)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
di: Lux, Florian, et al.
Pubblicazione: (2024)
di: Lux, Florian, et al.
Pubblicazione: (2024)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
di: Lu, Yen-Ju, et al.
Pubblicazione: (2024)
di: Lu, Yen-Ju, et al.
Pubblicazione: (2024)
A multilingual training strategy for low resource Text to Speech
di: Amalas, Asma, et al.
Pubblicazione: (2024)
di: Amalas, Asma, et al.
Pubblicazione: (2024)
Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2024)
Self-Train Before You Transcribe
di: Flynn, Robert, et al.
Pubblicazione: (2024)
di: Flynn, Robert, et al.
Pubblicazione: (2024)
Remastering Divide and Remaster: A Cinematic Audio Source Separation Dataset with Multilingual Support
di: Watcharasupat, Karn N., et al.
Pubblicazione: (2024)
di: Watcharasupat, Karn N., et al.
Pubblicazione: (2024)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
di: Duret, Jarod, et al.
Pubblicazione: (2024)
di: Duret, Jarod, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
di: Nagpal, Chirag, et al.
Pubblicazione: (2024) -
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
di: Ngo, Huong, et al.
Pubblicazione: (2025) -
Gradient boundaries through confidence intervals for forced alignment estimates using model ensembles
di: Kelley, Matthew C.
Pubblicazione: (2025) -
A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
di: You, Jian, et al.
Pubblicazione: (2024) -
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)