An Intelligent AI glasses System with Multi-Agent Architecture for Real-Time Voice Processing and Task Execution
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Sheng-Kai, Wu, Jyh-Horng, Lin, Ching-Yao, Lin, Yen-Ting |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
por: Chen, Junjie, et al.
Publicado: (2025)
por: Chen, Junjie, et al.
Publicado: (2025)
Sona: Real-Time Multi-Target Sound Attenuation for Noise Sensitivity
por: Huang, Jeremy Zhengqi, et al.
Publicado: (2026)
por: Huang, Jeremy Zhengqi, et al.
Publicado: (2026)
AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic Features
por: Xu, Weiye, et al.
Publicado: (2025)
por: Xu, Weiye, et al.
Publicado: (2025)
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
por: Shi, Weiyan, et al.
Publicado: (2025)
por: Shi, Weiyan, et al.
Publicado: (2025)
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
por: Martin, Charles Patrick
Publicado: (2026)
por: Martin, Charles Patrick
Publicado: (2026)
Beyond Voice Assistants: Exploring Advantages and Risks of an In-Car Social Robot in Real Driving Scenarios
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
From Generation to Attribution: Music AI Agent Architectures for the Post-Streaming Era
por: Kim, Wonil, et al.
Publicado: (2025)
por: Kim, Wonil, et al.
Publicado: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
por: Mertes, Silvan, et al.
Publicado: (2024)
por: Mertes, Silvan, et al.
Publicado: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
por: Li, Yue, et al.
Publicado: (2024)
por: Li, Yue, et al.
Publicado: (2024)
Transhuman Ansambl - Voice Beyond Language
por: Ivsic, Lucija, et al.
Publicado: (2024)
por: Ivsic, Lucija, et al.
Publicado: (2024)
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
por: Ma, Yong, et al.
Publicado: (2025)
por: Ma, Yong, et al.
Publicado: (2025)
SpeechCompass: Enhancing Mobile Captioning with Diarization and Directional Guidance via Multi-Microphone Localization
por: Dementyev, Artem, et al.
Publicado: (2025)
por: Dementyev, Artem, et al.
Publicado: (2025)
Do AI Voices Learn Social Nuances? A Case of Politeness and Speech Rate
por: Rabin, Eyal, et al.
Publicado: (2025)
por: Rabin, Eyal, et al.
Publicado: (2025)
A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation
por: Li, Kai, et al.
Publicado: (2026)
por: Li, Kai, et al.
Publicado: (2026)
Opening Musical Creativity? Embedded Ideologies in Generative-AI Music Systems
por: Pram, Liam, et al.
Publicado: (2025)
por: Pram, Liam, et al.
Publicado: (2025)
Real-time and Continuous Turn-taking Prediction Using Voice Activity Projection
por: Inoue, Koji, et al.
Publicado: (2024)
por: Inoue, Koji, et al.
Publicado: (2024)
Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
por: Doukhan, David, et al.
Publicado: (2024)
por: Doukhan, David, et al.
Publicado: (2024)
Attribution-by-design: Ensuring Inference-Time Provenance in Generative Music Systems
por: Morreale, Fabio, et al.
Publicado: (2025)
por: Morreale, Fabio, et al.
Publicado: (2025)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
por: Guo, Yiwei, et al.
Publicado: (2023)
por: Guo, Yiwei, et al.
Publicado: (2023)
Calliope: An Online Generative Music System for Symbolic Multi-Track Composition
por: Tchemeube, Renaud Bougueng, et al.
Publicado: (2025)
por: Tchemeube, Renaud Bougueng, et al.
Publicado: (2025)
Voice "Cloning" is Style Transfer
por: Zhou, Kaitlyn, et al.
Publicado: (2026)
por: Zhou, Kaitlyn, et al.
Publicado: (2026)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
por: Xue, Liumeng, et al.
Publicado: (2024)
por: Xue, Liumeng, et al.
Publicado: (2024)
Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
por: Kato, Kazushi, et al.
Publicado: (2025)
por: Kato, Kazushi, et al.
Publicado: (2025)
Yeah, Un, Oh: Continuous and Real-time Backchannel Prediction with Fine-tuning of Voice Activity Projection
por: Inoue, Koji, et al.
Publicado: (2024)
por: Inoue, Koji, et al.
Publicado: (2024)
Chord Colourizer: A Near Real-Time System for Visualizing Musical Key
por: Haimes, Paul
Publicado: (2025)
por: Haimes, Paul
Publicado: (2025)
LoopLens: Supporting Search as Creation in Loop-Based Music Composition
por: Long, Sheng, et al.
Publicado: (2026)
por: Long, Sheng, et al.
Publicado: (2026)
Beyond-Voice: Towards Continuous 3D Hand Pose Tracking on Commercial Home Assistant Devices
por: Li, Yin, et al.
Publicado: (2023)
por: Li, Yin, et al.
Publicado: (2023)
AI TrackMate: Finally, Someone Who Will Give Your Music More Than Just "Sounds Great!"
por: Jiang, Yi-Lin, et al.
Publicado: (2024)
por: Jiang, Yi-Lin, et al.
Publicado: (2024)
Springboard, Roadblock or "Crutch"?: How Transgender Users Leverage Voice Changers for Gender Presentation in Social Virtual Reality
por: Povinelli, Kassie, et al.
Publicado: (2024)
por: Povinelli, Kassie, et al.
Publicado: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
por: Nishida, Naoto, et al.
Publicado: (2025)
por: Nishida, Naoto, et al.
Publicado: (2025)
Talking Spell: A Wearable System Enabling Real-Time Anthropomorphic Voice Interaction with Everyday Objects
por: Wang, Xuetong, et al.
Publicado: (2025)
por: Wang, Xuetong, et al.
Publicado: (2025)
I Know You're Listening: Adaptive Voice for HRI
por: Tuttösí, Paige
Publicado: (2025)
por: Tuttösí, Paige
Publicado: (2025)
Improved Dysarthric Speech to Text Conversion via TTS Personalization
por: Mihajlik, Péter, et al.
Publicado: (2025)
por: Mihajlik, Péter, et al.
Publicado: (2025)
Bridging the Gap between Micro-scale Traffic Simulation and 4D Digital Cityscapes
por: Jiao, Longxiang, et al.
Publicado: (2026)
por: Jiao, Longxiang, et al.
Publicado: (2026)
Erie: A Declarative Grammar for Data Sonification
por: Kim, Hyeok, et al.
Publicado: (2024)
por: Kim, Hyeok, et al.
Publicado: (2024)
Effect of Avatar Head Movement on Communication Behaviour, Experience of Presence and Conversation Success in Triadic Conversations
por: Kothe, Angelika, et al.
Publicado: (2025)
por: Kothe, Angelika, et al.
Publicado: (2025)
Three-Class Emotion Classification for Audiovisual Scenes Based on Ensemble Learning Scheme
por: Xiong, Xiangrui, et al.
Publicado: (2025)
por: Xiong, Xiangrui, et al.
Publicado: (2025)
Accessible Fine-grained Data Representation via Spatial Audio
por: Liu, Can, et al.
Publicado: (2026)
por: Liu, Can, et al.
Publicado: (2026)
LightBeam: An Accurate and Memory-Efficient CTC Decoder for Speech Neuroprostheses
por: Feghhi, Ebrahim, et al.
Publicado: (2026)
por: Feghhi, Ebrahim, et al.
Publicado: (2026)
FlueBricks: A Construction Kit of Flute-like Instruments for Acoustic Reasoning
por: Chen, Bo-Yu, et al.
Publicado: (2026)
por: Chen, Bo-Yu, et al.
Publicado: (2026)
Ejemplares similares
-
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations
por: Chen, Junjie, et al.
Publicado: (2025) -
Sona: Real-Time Multi-Target Sound Attenuation for Noise Sensitivity
por: Huang, Jeremy Zhengqi, et al.
Publicado: (2026) -
AuthGlass: Benchmarking Voice Liveness Detection and Authentication on Smart Glasses via Comprehensive Acoustic Features
por: Xu, Weiye, et al.
Publicado: (2025) -
TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
por: Shi, Weiyan, et al.
Publicado: (2025) -
Opening the Design Space: Two Years of Performance with Intelligent Musical Instruments
por: Martin, Charles Patrick
Publicado: (2026)