Morse Code-Enabled Speech Recognition for Individuals with Visual and Hearing Impairments
Fuente:
arXiv
Guardado en:
| Autor principal: | Choudhury, Ritabrata Roy |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Arabic Little STT: Arabic Children Speech Recognition Dataset
por: Alkadri, Mouhand, et al.
Publicado: (2025)
por: Alkadri, Mouhand, et al.
Publicado: (2025)
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
por: Xie, Zhifei, et al.
Publicado: (2024)
por: Xie, Zhifei, et al.
Publicado: (2024)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
por: Wang, Xihuai, et al.
Publicado: (2025)
por: Wang, Xihuai, et al.
Publicado: (2025)
Do AI Voices Learn Social Nuances? A Case of Politeness and Speech Rate
por: Rabin, Eyal, et al.
Publicado: (2025)
por: Rabin, Eyal, et al.
Publicado: (2025)
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
por: Liu, Tianyun
Publicado: (2025)
por: Liu, Tianyun
Publicado: (2025)
Adaptation and Optimization of Automatic Speech Recognition (ASR) for the Maritime Domain in the Field of VHF Communication
por: Nakilcioglu, Emin Cagatay, et al.
Publicado: (2023)
por: Nakilcioglu, Emin Cagatay, et al.
Publicado: (2023)
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
por: Tiwari, Upasana, et al.
Publicado: (2025)
por: Tiwari, Upasana, et al.
Publicado: (2025)
Cross-Lingual Speech Emotion Recognition: Humans vs. Self-Supervised Models
por: Han, Zhichen, et al.
Publicado: (2024)
por: Han, Zhichen, et al.
Publicado: (2024)
What Does it Take to Generalize SER Model Across Datasets? A Comprehensive Benchmark
por: Ibrahim, Adham, et al.
Publicado: (2024)
por: Ibrahim, Adham, et al.
Publicado: (2024)
Voice "Cloning" is Style Transfer
por: Zhou, Kaitlyn, et al.
Publicado: (2026)
por: Zhou, Kaitlyn, et al.
Publicado: (2026)
Evaluating Human-AI Interaction via Usability, User Experience and Acceptance Measures for MMM-C: A Creative AI System for Music Composition
por: Tchemeube, Renaud Bougueng, et al.
Publicado: (2025)
por: Tchemeube, Renaud Bougueng, et al.
Publicado: (2025)
Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection
por: Chandra, Joydeep
Publicado: (2026)
por: Chandra, Joydeep
Publicado: (2026)
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
por: Darwish, Hazem, et al.
Publicado: (2024)
por: Darwish, Hazem, et al.
Publicado: (2024)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
por: Kawamura, Kazuki, et al.
Publicado: (2024)
por: Kawamura, Kazuki, et al.
Publicado: (2024)
Beyond Acoustic Emotion Recognition: Multimodal Pathos Analysis in Political Speech Using LLM-Based and Acoustic Emotion Models
por: Dietrich, Juergen
Publicado: (2026)
por: Dietrich, Juergen
Publicado: (2026)
Call2Instruct: Automated Pipeline for Generating Q&A Datasets from Call Center Recordings for LLM Fine-Tuning
por: Echeverria, Alex, et al.
Publicado: (2025)
por: Echeverria, Alex, et al.
Publicado: (2025)
EmoHeal: An End-to-End System for Personalized Therapeutic Music Retrieval from Fine-grained Emotions
por: Wan, Xinchen, et al.
Publicado: (2025)
por: Wan, Xinchen, et al.
Publicado: (2025)
EmoAugNet: A Signal-Augmented Hybrid CNN-LSTM Framework for Speech Emotion Recognition
por: Paul, Durjoy Chandra, et al.
Publicado: (2025)
por: Paul, Durjoy Chandra, et al.
Publicado: (2025)
Evaluating Large Language Models for Health-related Queries with Presuppositions
por: Kaur, Navreet, et al.
Publicado: (2023)
por: Kaur, Navreet, et al.
Publicado: (2023)
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
por: Sinha, Abhijit, et al.
Publicado: (2025)
por: Sinha, Abhijit, et al.
Publicado: (2025)
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
por: Huang, Ailin, et al.
Publicado: (2025)
por: Huang, Ailin, et al.
Publicado: (2025)
Spontaneous Informal Speech Dataset for Punctuation Restoration
por: Liu, Xing Yi, et al.
Publicado: (2024)
por: Liu, Xing Yi, et al.
Publicado: (2024)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
por: Snoubara, Abdul Aziz, et al.
Publicado: (2026)
por: Snoubara, Abdul Aziz, et al.
Publicado: (2026)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
por: Torgashov, Nikita, et al.
Publicado: (2025)
por: Torgashov, Nikita, et al.
Publicado: (2025)
Observing Dialogue in Therapy: Categorizing and Forecasting Behavioral Codes
por: Cao, Jie, et al.
Publicado: (2019)
por: Cao, Jie, et al.
Publicado: (2019)
Deep Learning Models in Speech Recognition: Measuring GPU Energy Consumption, Impact of Noise and Model Quantization for Edge Deployment
por: Chakravarty, Aditya
Publicado: (2024)
por: Chakravarty, Aditya
Publicado: (2024)
A Computational Method for Measuring "Open Codes" in Qualitative Analysis
por: Chen, John, et al.
Publicado: (2024)
por: Chen, John, et al.
Publicado: (2024)
LLM Attributor: Interactive Visual Attribution for LLM Generation
por: Lee, Seongmin, et al.
Publicado: (2024)
por: Lee, Seongmin, et al.
Publicado: (2024)
Stress Detection on Code-Mixed Texts in Dravidian Languages using Machine Learning
por: Ramos, L., et al.
Publicado: (2024)
por: Ramos, L., et al.
Publicado: (2024)
LingVarBench: Benchmarking LLMs on Entity Recognitions and Linguistic Verbalization Patterns in Phone-Call Transcripts
por: Mohammadi, Seyedali, et al.
Publicado: (2025)
por: Mohammadi, Seyedali, et al.
Publicado: (2025)
Acoustic and perceptual differences between standard and accented speech and their voice clones
por: Yang, Tianle, et al.
Publicado: (2026)
por: Yang, Tianle, et al.
Publicado: (2026)
Diffusion Explainer: Visual Explanation for Text-to-image Stable Diffusion
por: Lee, Seongmin, et al.
Publicado: (2023)
por: Lee, Seongmin, et al.
Publicado: (2023)
LLM Comparator: Visual Analytics for Side-by-Side Evaluation of Large Language Models
por: Kahng, Minsuk, et al.
Publicado: (2024)
por: Kahng, Minsuk, et al.
Publicado: (2024)
A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
por: Benster, Tyler, et al.
Publicado: (2024)
por: Benster, Tyler, et al.
Publicado: (2024)
Human Feedback Driven Dynamic Speech Emotion Recognition
por: Fedorov, Ilya, et al.
Publicado: (2025)
por: Fedorov, Ilya, et al.
Publicado: (2025)
MIST: Multimodal Interactive Speech-based Tool-calling Conversational Assistants for Smart Homes
por: Chen, Maximillian, et al.
Publicado: (2026)
por: Chen, Maximillian, et al.
Publicado: (2026)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
por: Chang, Yi, et al.
Publicado: (2024)
por: Chang, Yi, et al.
Publicado: (2024)
"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
por: Dammu, Preetam Prabhu Srikar, et al.
Publicado: (2024)
por: Dammu, Preetam Prabhu Srikar, et al.
Publicado: (2024)
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
por: Park, Joon Sung, et al.
Publicado: (2024)
por: Park, Joon Sung, et al.
Publicado: (2024)
Personalized Benchmarking: Evaluating LLMs by Individual Preferences
por: Garbacea, Cristina, et al.
Publicado: (2026)
por: Garbacea, Cristina, et al.
Publicado: (2026)
Ejemplares similares
-
Arabic Little STT: Arabic Children Speech Recognition Dataset
por: Alkadri, Mouhand, et al.
Publicado: (2025) -
Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
por: Xie, Zhifei, et al.
Publicado: (2024) -
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
por: Wang, Xihuai, et al.
Publicado: (2025) -
Do AI Voices Learn Social Nuances? A Case of Politeness and Speech Rate
por: Rabin, Eyal, et al.
Publicado: (2025) -
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
por: Liu, Tianyun
Publicado: (2025)