IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zhang, Xin, Lyu, Xiang, Du, Zhihao, Chen, Qian, Zhang, Dong, Hu, Hangrui, Tan, Chaohong, Zhao, Tianyu, Wang, Yuxuan, Zhang, Bin, Lu, Heng, Zhou, Yaqian, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
par: An, Keyu, et autres
Publié: (2024)
par: An, Keyu, et autres
Publié: (2024)
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
par: Zhang, Qinglin, et autres
Publié: (2024)
par: Zhang, Qinglin, et autres
Publié: (2024)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
par: Du, Zhihao, et autres
Publié: (2024)
par: Du, Zhihao, et autres
Publié: (2024)
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
par: Huang, Kexin, et autres
Publié: (2026)
par: Huang, Kexin, et autres
Publié: (2026)
VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures
par: Qiu, Jielin, et autres
Publié: (2026)
par: Qiu, Jielin, et autres
Publié: (2026)
Enkidu: Universal Frequential Perturbation for Real-Time Audio Privacy Protection against Voice Deepfakes
par: Feng, Zhou, et autres
Publié: (2025)
par: Feng, Zhou, et autres
Publié: (2025)
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
par: Shi, Yemin, et autres
Publié: (2025)
par: Shi, Yemin, et autres
Publié: (2025)
Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
par: Lin, Yueqian, et autres
Publié: (2025)
par: Lin, Yueqian, et autres
Publié: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
par: Zhang, Xin, et autres
Publié: (2023)
par: Zhang, Xin, et autres
Publié: (2023)
VoiceBench: Benchmarking LLM-Based Voice Assistants
par: Chen, Yiming, et autres
Publié: (2024)
par: Chen, Yiming, et autres
Publié: (2024)
$τ$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
par: Ray, Soham, et autres
Publié: (2026)
par: Ray, Soham, et autres
Publié: (2026)
Marco-Voice Technical Report
par: Tian, Fengping, et autres
Publié: (2025)
par: Tian, Fengping, et autres
Publié: (2025)
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
par: Du, Zhihao, et autres
Publié: (2024)
par: Du, Zhihao, et autres
Publié: (2024)
SceneGuard: Training-Time Voice Protection with Scene-Consistent Audible Background Noise
par: Sang, Rui, et autres
Publié: (2025)
par: Sang, Rui, et autres
Publié: (2025)
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
par: Chen, Qian, et autres
Publié: (2025)
par: Chen, Qian, et autres
Publié: (2025)
VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
par: Zhan, Jun, et autres
Publié: (2025)
par: Zhan, Jun, et autres
Publié: (2025)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
par: Du, Zhihao, et autres
Publié: (2025)
par: Du, Zhihao, et autres
Publié: (2025)
SingVERSE: A Diverse, Real-World Benchmark for Singing Voice Enhancement
par: Jiang, Shaohan, et autres
Publié: (2025)
par: Jiang, Shaohan, et autres
Publié: (2025)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
par: Huo, Mingyue, et autres
Publié: (2025)
par: Huo, Mingyue, et autres
Publié: (2025)
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
par: Xu, Rixi, et autres
Publié: (2026)
par: Xu, Rixi, et autres
Publié: (2026)
Controlling your Attributes in Voice
par: Li, Xuyuan, et autres
Publié: (2025)
par: Li, Xuyuan, et autres
Publié: (2025)
Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection
par: Xue, Jun, et autres
Publié: (2026)
par: Xue, Jun, et autres
Publié: (2026)
SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
par: Zhang, Dong, et autres
Publié: (2024)
par: Zhang, Dong, et autres
Publié: (2024)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
par: Hou, Yixuan, et autres
Publié: (2025)
par: Hou, Yixuan, et autres
Publié: (2025)
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
par: Ma, Yong, et autres
Publié: (2025)
par: Ma, Yong, et autres
Publié: (2025)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
par: Qian, Zhiwen, et autres
Publié: (2025)
par: Qian, Zhiwen, et autres
Publié: (2025)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2026)
par: Wang, Zhichao, et autres
Publié: (2026)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
par: Du, Zongyang, et autres
Publié: (2024)
par: Du, Zongyang, et autres
Publié: (2024)
Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models
par: Kim, Heeseung, et autres
Publié: (2025)
par: Kim, Heeseung, et autres
Publié: (2025)
Building Enterprise Realtime Voice Agents from Scratch: A Technical Tutorial
par: Qiu, Jielin, et autres
Publié: (2026)
par: Qiu, Jielin, et autres
Publié: (2026)
Perpetual Dialogues: A Computational Analysis of Voice-Guitar Interaction in Carlos Paredes's Discography
par: Bernardes, Gilberto, et autres
Publié: (2026)
par: Bernardes, Gilberto, et autres
Publié: (2026)
SpeechAlign: Aligning Speech Generation to Human Preferences
par: Zhang, Dong, et autres
Publié: (2024)
par: Zhang, Dong, et autres
Publié: (2024)
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children
par: Xu, Haiying, et autres
Publié: (2025)
par: Xu, Haiying, et autres
Publié: (2025)
Robust Singing Voice Transcription Serves Synthesis
par: Li, Ruiqi, et autres
Publié: (2024)
par: Li, Ruiqi, et autres
Publié: (2024)
i-LAVA: Insights on Low Latency Voice-2-Voice Architecture for Agents
par: Purwar, Anupam, et autres
Publié: (2025)
par: Purwar, Anupam, et autres
Publié: (2025)
StyleStream: Real-Time Zero-Shot Voice Style Conversion
par: Liu, Yisi, et autres
Publié: (2026)
par: Liu, Yisi, et autres
Publié: (2026)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
par: Anastassiou, Philip, et autres
Publié: (2024)
par: Anastassiou, Philip, et autres
Publié: (2024)
Self Voice Conversion as an Attack against Neural Audio Watermarking
par: Özer, Yigitcan, et autres
Publié: (2026)
par: Özer, Yigitcan, et autres
Publié: (2026)
OpenVoice: Versatile Instant Voice Cloning
par: Qin, Zengyi, et autres
Publié: (2023)
par: Qin, Zengyi, et autres
Publié: (2023)
Documents similaires
-
FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
par: An, Keyu, et autres
Publié: (2024) -
OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
par: Zhang, Qinglin, et autres
Publié: (2024) -
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
par: Du, Zhihao, et autres
Publié: (2024) -
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
par: Huang, Kexin, et autres
Publié: (2026) -
VoiceAgentRAG: Solving the RAG Latency Bottleneck in Real-Time Voice Agents Using Dual-Agent Architectures
par: Qiu, Jielin, et autres
Publié: (2026)