Wireless Hearables With Programmable Speech AI Accelerators
Fuente:
arXiv
Guardado en:
| Autores principales: | Itani, Malek, Chen, Tuochao, Raghavan, Arun, Kohlberg, Gavriel, Gollakota, Shyamnath |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TF-MLPNet: Tiny Real-Time Neural Speech Separation
por: Itani, Malek, et al.
Publicado: (2025)
por: Itani, Malek, et al.
Publicado: (2025)
Proactive Hearing Assistants that Isolate Egocentric Conversations
por: Hu, Guilin, et al.
Publicado: (2025)
por: Hu, Guilin, et al.
Publicado: (2025)
Look Once to Hear: Target Speech Hearing with Noisy Examples
por: Veluri, Bandhav, et al.
Publicado: (2024)
por: Veluri, Bandhav, et al.
Publicado: (2024)
Spatial Speech Translation: Translating Across Space With Binaural Hearables
por: Chen, Tuochao, et al.
Publicado: (2025)
por: Chen, Tuochao, et al.
Publicado: (2025)
Neural Speech Extraction with Human Feedback
por: Itani, Malek, et al.
Publicado: (2025)
por: Itani, Malek, et al.
Publicado: (2025)
Knowledge boosting during low-latency inference
por: Srinivas, Vidya, et al.
Publicado: (2024)
por: Srinivas, Vidya, et al.
Publicado: (2024)
Fine-grained Soundscape Control for Augmented Hearing
por: Oh, Seunghyun, et al.
Publicado: (2026)
por: Oh, Seunghyun, et al.
Publicado: (2026)
LLAMAPIE: Proactive In-Ear Conversation Assistants
por: Chen, Tuochao, et al.
Publicado: (2025)
por: Chen, Tuochao, et al.
Publicado: (2025)
Target conversation extraction: Source separation using turn-taking dynamics
por: Chen, Tuochao, et al.
Publicado: (2024)
por: Chen, Tuochao, et al.
Publicado: (2024)
Modeling of Speech-dependent Own Voice Transfer Characteristics for Hearables with In-ear Microphones
por: Ohlenbusch, Mattes, et al.
Publicado: (2023)
por: Ohlenbusch, Mattes, et al.
Publicado: (2023)
Speech-dependent Modeling of Own Voice Transfer Characteristics for In-ear Microphones in Hearables
por: Ohlenbusch, Mattes, et al.
Publicado: (2023)
por: Ohlenbusch, Mattes, et al.
Publicado: (2023)
Towards Sub-millisecond Latency Real-Time Speech Enhancement Models on Hearables
por: Dementyev, Artem, et al.
Publicado: (2024)
por: Dementyev, Artem, et al.
Publicado: (2024)
Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents
por: Veluri, Bandhav, et al.
Publicado: (2024)
por: Veluri, Bandhav, et al.
Publicado: (2024)
SoundSculpt: Direction and Semantics Driven Ambisonic Target Sound Extraction
por: Chen, Tuochao, et al.
Publicado: (2025)
por: Chen, Tuochao, et al.
Publicado: (2025)
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
por: Zheng, Qixi, et al.
Publicado: (2025)
por: Zheng, Qixi, et al.
Publicado: (2025)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
por: Choi, Jeongsoo, et al.
Publicado: (2025)
por: Choi, Jeongsoo, et al.
Publicado: (2025)
Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
por: Lin, Yi-Cheng, et al.
Publicado: (2026)
por: Lin, Yi-Cheng, et al.
Publicado: (2026)
Llasa+: Free Lunch for Accelerated and Streaming Llama-Based Speech Synthesis
por: Tian, Wenjie, et al.
Publicado: (2025)
por: Tian, Wenjie, et al.
Publicado: (2025)
Accelerating Diffusion Transformer-Based Text-to-Speech with Transformer Layer Caching
por: Sakpiboonchit, Siratish
Publicado: (2025)
por: Sakpiboonchit, Siratish
Publicado: (2025)
IKFST: IOO and KOO Algorithms for Accelerated and Precise WFST-based End-to-End Automatic Speech Recognition
por: Zhuang, Zhuoran, et al.
Publicado: (2026)
por: Zhuang, Zhuoran, et al.
Publicado: (2026)
Towards Machine Unlearning for Paralinguistic Speech Processing
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis
por: R, Vinotha, et al.
Publicado: (2024)
por: R, Vinotha, et al.
Publicado: (2024)
SUBARU: A Practical Approach to Power Saving in Hearables Using SUB-Nyquist Audio Resolution Upsampling
por: Tamiti, Tarikul Islam, et al.
Publicado: (2025)
por: Tamiti, Tarikul Islam, et al.
Publicado: (2025)
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
por: Wei, Linye, et al.
Publicado: (2025)
por: Wei, Linye, et al.
Publicado: (2025)
Spatially Selective Active Noise Control for Open-fitting Hearables with Acausal Optimization
por: Xiao, Tong, et al.
Publicado: (2025)
por: Xiao, Tong, et al.
Publicado: (2025)
AntiDeepFake: AI for Deep Fake Speech Recognition
por: Togootogtokh, Enkhtogtokh, et al.
Publicado: (2024)
por: Togootogtokh, Enkhtogtokh, et al.
Publicado: (2024)
Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations
por: Girish, et al.
Publicado: (2025)
por: Girish, et al.
Publicado: (2025)
NeckCare: Preventing Tech Neck using Hearable-based Multimodal Sensing
por: Chhaglani, Bhawana, et al.
Publicado: (2024)
por: Chhaglani, Bhawana, et al.
Publicado: (2024)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
por: Yang, Qian, et al.
Publicado: (2024)
por: Yang, Qian, et al.
Publicado: (2024)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
por: Gao, Xiaoxue, et al.
Publicado: (2024)
por: Gao, Xiaoxue, et al.
Publicado: (2024)
Investigating Polyglot Speech Foundation Models for Learning Collective Emotion from Crowds
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
por: Phukan, Orchid Chetia, et al.
Publicado: (2025)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
por: Chen, Weidong, et al.
Publicado: (2025)
por: Chen, Weidong, et al.
Publicado: (2025)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
por: Wang, Helin, et al.
Publicado: (2024)
por: Wang, Helin, et al.
Publicado: (2024)
ToneUnit: A Speech Discretization Approach for Tonal Language Speech Synthesis
por: Tao, Dehua, et al.
Publicado: (2024)
por: Tao, Dehua, et al.
Publicado: (2024)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
por: Zhang, Leying, et al.
Publicado: (2025)
por: Zhang, Leying, et al.
Publicado: (2025)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
por: Wang, Yuanyuan, et al.
Publicado: (2025)
por: Wang, Yuanyuan, et al.
Publicado: (2025)
Noise-Aware Speech Separation with Contrastive Learning
por: Zhang, Zizheng, et al.
Publicado: (2023)
por: Zhang, Zizheng, et al.
Publicado: (2023)
Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
por: Chen, Peikun, et al.
Publicado: (2024)
por: Chen, Peikun, et al.
Publicado: (2024)
From Human Speech to Ocean Signals: Transferring Speech Large Models for Underwater Acoustic Target Recognition
por: Huang, Mengcheng, et al.
Publicado: (2026)
por: Huang, Mengcheng, et al.
Publicado: (2026)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
por: Wang, Hui, et al.
Publicado: (2025)
por: Wang, Hui, et al.
Publicado: (2025)
Ejemplares similares
-
TF-MLPNet: Tiny Real-Time Neural Speech Separation
por: Itani, Malek, et al.
Publicado: (2025) -
Proactive Hearing Assistants that Isolate Egocentric Conversations
por: Hu, Guilin, et al.
Publicado: (2025) -
Look Once to Hear: Target Speech Hearing with Noisy Examples
por: Veluri, Bandhav, et al.
Publicado: (2024) -
Spatial Speech Translation: Translating Across Space With Binaural Hearables
por: Chen, Tuochao, et al.
Publicado: (2025) -
Neural Speech Extraction with Human Feedback
por: Itani, Malek, et al.
Publicado: (2025)