SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance
Fuente:
arXiv
Guardado en:
| Autores principales: | Lou, Haowei, Huang, Chengkai, Paik, Hye-young, Hu, Yongquan, Quigley, Aaron, Hu, Wen, Yao, Lina |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
por: Lou, Haowei, et al.
Publicado: (2026)
por: Lou, Haowei, et al.
Publicado: (2026)
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
por: Lou, Haowei, et al.
Publicado: (2025)
por: Lou, Haowei, et al.
Publicado: (2025)
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
por: Lou, Haowei, et al.
Publicado: (2025)
por: Lou, Haowei, et al.
Publicado: (2025)
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
por: Lou, Haowei, et al.
Publicado: (2024)
por: Lou, Haowei, et al.
Publicado: (2024)
Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration
por: Lou, Haowei, et al.
Publicado: (2024)
por: Lou, Haowei, et al.
Publicado: (2024)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
por: Lou, Haowei, et al.
Publicado: (2024)
por: Lou, Haowei, et al.
Publicado: (2024)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
por: Huang, Wuwei, et al.
Publicado: (2025)
por: Huang, Wuwei, et al.
Publicado: (2025)
Recent Advances in End-to-End Simultaneous Speech Translation
por: Liu, Xiaoqian, et al.
Publicado: (2024)
por: Liu, Xiaoqian, et al.
Publicado: (2024)
Speech-to-See: End-to-End Speech-Driven Open-Set Object Detection
por: Lu, Wenhuan, et al.
Publicado: (2025)
por: Lu, Wenhuan, et al.
Publicado: (2025)
Ti-Audio: The First Multi-Dialectal End-to-End Speech LLM for Tibetan
por: Wang, Jialing, et al.
Publicado: (2026)
por: Wang, Jialing, et al.
Publicado: (2026)
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
por: Min, Anna, et al.
Publicado: (2025)
por: Min, Anna, et al.
Publicado: (2025)
Representation Purification for End-to-End Speech Translation
por: Zhang, Chengwei, et al.
Publicado: (2024)
por: Zhang, Chengwei, et al.
Publicado: (2024)
Continual Test-time Adaptation for End-to-end Speech Recognition on Noisy Speech
por: Lin, Guan-Ting, et al.
Publicado: (2024)
por: Lin, Guan-Ting, et al.
Publicado: (2024)
Probing Human Articulatory Constraints in End-to-End TTS with Reverse and Mismatched Speech-Text Directions
por: Khadse, Parth, et al.
Publicado: (2026)
por: Khadse, Parth, et al.
Publicado: (2026)
End-to-End Speech-to-Text Translation: A Survey
por: Sethiya, Nivedita, et al.
Publicado: (2023)
por: Sethiya, Nivedita, et al.
Publicado: (2023)
A Parallel Ultra-Low Power Silent Speech Interface based on a Wearable, Fully-dry EMG Neckband
por: Meier, Fiona, et al.
Publicado: (2025)
por: Meier, Fiona, et al.
Publicado: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
por: Chen, Junyang, et al.
Publicado: (2026)
por: Chen, Junyang, et al.
Publicado: (2026)
An End-to-End Speech Summarization Using Large Language Model
por: Shang, Hengchao, et al.
Publicado: (2024)
por: Shang, Hengchao, et al.
Publicado: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
por: Gupta, Kishan, et al.
Publicado: (2024)
por: Gupta, Kishan, et al.
Publicado: (2024)
Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
por: Zhou, Xuanru, et al.
Publicado: (2024)
por: Zhou, Xuanru, et al.
Publicado: (2024)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
por: Ma, Zhengrui, et al.
Publicado: (2024)
por: Ma, Zhengrui, et al.
Publicado: (2024)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
por: Zhou, Junzuo, et al.
Publicado: (2024)
por: Zhou, Junzuo, et al.
Publicado: (2024)
Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training
por: Ahmad, Hawraz A., et al.
Publicado: (2024)
por: Ahmad, Hawraz A., et al.
Publicado: (2024)
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
por: Hu, Jiliang, et al.
Publicado: (2025)
por: Hu, Jiliang, et al.
Publicado: (2025)
Frequency-Specific Neural Response and Cross-Correlation Analysis of Envelope Following Responses to Native Speech and Music Using Multichannel EEG Signals: A Case Study
por: Hasan, Md. Mahbub, et al.
Publicado: (2025)
por: Hasan, Md. Mahbub, et al.
Publicado: (2025)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
por: Huang, Wuwei, et al.
Publicado: (2025)
por: Huang, Wuwei, et al.
Publicado: (2025)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
por: Farhadipour, Aref, et al.
Publicado: (2023)
por: Farhadipour, Aref, et al.
Publicado: (2023)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
por: Li, Tianpeng, et al.
Publicado: (2025)
por: Li, Tianpeng, et al.
Publicado: (2025)
Meta-Learning in Audio and Speech Processing: An End to End Comprehensive Review
por: Raimon, Athul, et al.
Publicado: (2024)
por: Raimon, Athul, et al.
Publicado: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
por: Guo, Yinlin, et al.
Publicado: (2024)
por: Guo, Yinlin, et al.
Publicado: (2024)
ASCEND: Accurate yet Efficient End-to-End Stochastic Computing Acceleration of Vision Transformer
por: Xie, Tong, et al.
Publicado: (2024)
por: Xie, Tong, et al.
Publicado: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
por: Bataev, Vladimir, et al.
Publicado: (2025)
por: Bataev, Vladimir, et al.
Publicado: (2025)
TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving
por: Wang, Jiawei, et al.
Publicado: (2025)
por: Wang, Jiawei, et al.
Publicado: (2025)
An Efficient End-to-End Approach to Noise Invariant Speech Features via Multi-Task Learning
por: Guimarães, Heitor R., et al.
Publicado: (2024)
por: Guimarães, Heitor R., et al.
Publicado: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
por: Kocour, Martin, et al.
Publicado: (2025)
por: Kocour, Martin, et al.
Publicado: (2025)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
por: Moslem, Yasmin
Publicado: (2024)
por: Moslem, Yasmin
Publicado: (2024)
ML-ARIS: Multilayer Underwater Acoustic Reconfigurable Intelligent Surface with High-Resolution Reflection Control
por: Pu, Lina, et al.
Publicado: (2025)
por: Pu, Lina, et al.
Publicado: (2025)
End-to-End Simultaneous Dysarthric Speech Reconstruction with Frame-Level Adaptor and Multiple Wait-k Knowledge Distillation
por: Wu, Minghui, et al.
Publicado: (2026)
por: Wu, Minghui, et al.
Publicado: (2026)
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
por: Cheng, Shanbo, et al.
Publicado: (2024)
por: Cheng, Shanbo, et al.
Publicado: (2024)
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
por: Zhuang, Zhuoran, et al.
Publicado: (2026)
por: Zhuang, Zhuoran, et al.
Publicado: (2026)
Ejemplares similares
-
ParaMETA: Towards Learning Disentangled Paralinguistic Speaking Styles Representations from Speech
por: Lou, Haowei, et al.
Publicado: (2026) -
Generalized Multilingual Text-to-Speech Generation with Language-Aware Style Adaptation
por: Lou, Haowei, et al.
Publicado: (2025) -
ParaStyleTTS: Toward Efficient and Robust Paralinguistic Style Control for Expressive Text-to-Speech Generation
por: Lou, Haowei, et al.
Publicado: (2025) -
StyleSpeech: Parameter-efficient Fine Tuning for Pre-trained Controllable Text-to-Speech
por: Lou, Haowei, et al.
Publicado: (2024) -
Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration
por: Lou, Haowei, et al.
Publicado: (2024)