ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Xi, Xue, Wei, Guo, Yike |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
von: Cheng, Sitong, et al.
Veröffentlicht: (2025)
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors
von: Bao, Guangyin, et al.
Veröffentlicht: (2026)
von: Bao, Guangyin, et al.
Veröffentlicht: (2026)
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
von: Wu, Weihao, et al.
Veröffentlicht: (2025)
von: Wu, Weihao, et al.
Veröffentlicht: (2025)
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
von: Dvirniak, Artem, et al.
Veröffentlicht: (2026)
von: Dvirniak, Artem, et al.
Veröffentlicht: (2026)
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
von: Park, Kyeongman, et al.
Veröffentlicht: (2025)
von: Park, Kyeongman, et al.
Veröffentlicht: (2025)
Human or Machine? A Preliminary Turing Test for Speech-to-Speech Interaction
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
SLM-SS: Speech Language Model for Generative Speech Separation
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueyao, et al.
Veröffentlicht: (2025)
AST: Adaptive, Seamless, and Training-Free Precise Speech Editing
von: Lv, Sihan, et al.
Veröffentlicht: (2026)
von: Lv, Sihan, et al.
Veröffentlicht: (2026)
UniSE: A Unified Framework for Decoder-only Autoregressive LM-based Speech Enhancement
von: Yan, Haoyin, et al.
Veröffentlicht: (2025)
von: Yan, Haoyin, et al.
Veröffentlicht: (2025)
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
von: Li, Wenyu, et al.
Veröffentlicht: (2025)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
Inference-time Scaling for Diffusion-based Audio Super-resolution
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
von: Jin, Yizhu, et al.
Veröffentlicht: (2025)
RAS: a Reliability Oriented Metric for Automatic Speech Recognition
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
von: Huang, Wenbin, et al.
Veröffentlicht: (2026)
FlashSpeech: Efficient Zero-Shot Speech Synthesis
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
von: Ye, Zhen, et al.
Veröffentlicht: (2024)
Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
von: Xue, Jun, et al.
Veröffentlicht: (2026)
von: Xue, Jun, et al.
Veröffentlicht: (2026)
SyncSpeech: Efficient and Low-Latency Text-to-Speech based on Temporal Masked Transformer
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
DDSP-QbE++: Improving Speech Quality for Speech Anonymisation for Atypical Speech
von: Ghosh, Suhita, et al.
Veröffentlicht: (2026)
von: Ghosh, Suhita, et al.
Veröffentlicht: (2026)
FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
von: Chen, Jianyi, et al.
Veröffentlicht: (2024)
CoMoSVC: Consistency Model-based Singing Voice Conversion
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
von: Lu, Yiwen, et al.
Veröffentlicht: (2024)
AutoStyle-TTS: Retrieval-Augmented Generation based Automatic Style Matching Text-to-Speech Synthesis
von: Luo, Dan, et al.
Veröffentlicht: (2025)
von: Luo, Dan, et al.
Veröffentlicht: (2025)
GSRM: Generative Speech Reward Model for Speech RLHF
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
von: Shen, Maohao, et al.
Veröffentlicht: (2026)
Speech Emotion Recognition via Entropy-Aware Score Selection
von: Chua, ChenYi, et al.
Veröffentlicht: (2025)
von: Chua, ChenYi, et al.
Veröffentlicht: (2025)
Understanding Frechet Speech Distance for Synthetic Speech Quality Evaluation
von: Kim, June-Woo, et al.
Veröffentlicht: (2026)
von: Kim, June-Woo, et al.
Veröffentlicht: (2026)
Dynamic Fusion Multimodal Network for SpeechWellness Detection
von: Sun, Wenqiang, et al.
Veröffentlicht: (2025)
von: Sun, Wenqiang, et al.
Veröffentlicht: (2025)
Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play
von: Shi, Yemin, et al.
Veröffentlicht: (2025)
von: Shi, Yemin, et al.
Veröffentlicht: (2025)
Spatial Audio Question Answering and Reasoning on Dynamic Source Movements
von: Sridhar, Arvind Krishna, et al.
Veröffentlicht: (2026)
von: Sridhar, Arvind Krishna, et al.
Veröffentlicht: (2026)
Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
von: Xie, Yuankun, et al.
Veröffentlicht: (2025)
SpeechQualityLLM: LLM-Based Multimodal Assessment of Speech Quality
von: Monjur, Mahathir, et al.
Veröffentlicht: (2025)
von: Monjur, Mahathir, et al.
Veröffentlicht: (2025)
SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
ROSE: A Recognition-Oriented Speech Enhancement Framework in Air Traffic Control Using Multi-Objective Learning
von: Yu, Xincheng, et al.
Veröffentlicht: (2023)
von: Yu, Xincheng, et al.
Veröffentlicht: (2023)
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations
von: Han, Yichen, et al.
Veröffentlicht: (2025)
von: Han, Yichen, et al.
Veröffentlicht: (2025)
Towards Explicit Acoustic Evidence Perception in Audio LLMs for Speech Deepfake Detection
von: Guo, Xiaoxuan, et al.
Veröffentlicht: (2026)
von: Guo, Xiaoxuan, et al.
Veröffentlicht: (2026)
Enabling Automatic Disordered Speech Recognition: An Impaired Speech Dataset in the Akan Language
von: Wiafe, Isaac, et al.
Veröffentlicht: (2026)
von: Wiafe, Isaac, et al.
Veröffentlicht: (2026)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
von: Wang, Peng, et al.
Veröffentlicht: (2026)
von: Wang, Peng, et al.
Veröffentlicht: (2026)
MF-Speech: Achieving Fine-Grained and Compositional Control in Speech Generation via Factor Disentanglement
von: Yu, Xinyue, et al.
Veröffentlicht: (2025)
von: Yu, Xinyue, et al.
Veröffentlicht: (2025)
MindMelody: A Closed-Loop EEG-Driven System for Personalized Music Intervention
von: Zhang, Yimeng, et al.
Veröffentlicht: (2026)
von: Zhang, Yimeng, et al.
Veröffentlicht: (2026)
The NPU-ASLP-LiAuto System Description for Visual Speech Recognition in CNVSRC 2023
von: Wang, He, et al.
Veröffentlicht: (2024)
von: Wang, He, et al.
Veröffentlicht: (2024)
LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
von: Kang, Boyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice
von: Cheng, Sitong, et al.
Veröffentlicht: (2025) -
Speech-DRAME: A Framework for Human-Aligned Benchmarks in Speech Role-Play
von: Shi, Jiatong, et al.
Veröffentlicht: (2025) -
MindVoice: Reconstructing Intelligible Speech from Non-invasive Neural Signals with Pretrained Priors
von: Bao, Guangyin, et al.
Veröffentlicht: (2026) -
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
von: Wu, Weihao, et al.
Veröffentlicht: (2025) -
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
von: Dvirniak, Artem, et al.
Veröffentlicht: (2026)