MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xingjian, Xu, Zhe, Cheng, Qinyuan, Fei, Zhaoye, Jin, Luozhijie, Wang, Yang, Chen, Hanfu, Jiang, Yaozhou, Gao, Qinghui, Chen, Ke, Li, Ruixiao, Chen, Mingshu, Wang, Ruiming, Zhang, Wenbo, Zhang, Yiyang, Yu, Donghua, Gao, Yang, Yang, Xiaogui, Gong, Yitian, Xu, Yuanfan, Zhou, Yaqian, Huang, Xuanjing, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MOSS-TTSD: Text to Spoken Dialogue Generation
by: Zhang, Yuqian, et al.
Published: (2026)
by: Zhang, Yuqian, et al.
Published: (2026)
MOSS-TTS Technical Report
by: Gong, Yitian, et al.
Published: (2026)
by: Gong, Yitian, et al.
Published: (2026)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
by: Gong, Yitian, et al.
Published: (2026)
by: Gong, Yitian, et al.
Published: (2026)
MOSS-Audio Technical Report
by: Yang, Chen, et al.
Published: (2026)
by: Yang, Chen, et al.
Published: (2026)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
by: Gong, Yitian, et al.
Published: (2025)
by: Gong, Yitian, et al.
Published: (2025)
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
by: Deng, Ruifan, et al.
Published: (2025)
by: Deng, Ruifan, et al.
Published: (2025)
MOSS Transcribe Diarize Technical Report
by: AI, MOSI., et al.
Published: (2026)
by: AI, MOSI., et al.
Published: (2026)
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
by: Huang, Kexin, et al.
Published: (2026)
by: Huang, Kexin, et al.
Published: (2026)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
by: Zhang, Xin, et al.
Published: (2023)
by: Zhang, Xin, et al.
Published: (2023)
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
by: Huang, Kexin, et al.
Published: (2025)
by: Huang, Kexin, et al.
Published: (2025)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
by: Yang, Chenchen, et al.
Published: (2026)
by: Yang, Chenchen, et al.
Published: (2026)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
by: Wang, Peng, et al.
Published: (2026)
by: Wang, Peng, et al.
Published: (2026)
SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation
by: Zhang, Dong, et al.
Published: (2024)
by: Zhang, Dong, et al.
Published: (2024)
SpeechAlign: Aligning Speech Generation to Human Preferences
by: Zhang, Dong, et al.
Published: (2024)
by: Zhang, Dong, et al.
Published: (2024)
Speech-Mamba: Long-Context Speech Recognition with Selective State Spaces Models
by: Gao, Xiaoxue, et al.
Published: (2024)
by: Gao, Xiaoxue, et al.
Published: (2024)
The Disintegration of Free Speech
by: Mei, Yiyang
Published: (2026)
by: Mei, Yiyang
Published: (2026)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
by: Peng, Runyu, et al.
Published: (2026)
by: Peng, Runyu, et al.
Published: (2026)
Inner Speech in Motor Cortex: Laying the Foundation for Privacy‐Preserving Speech Neuroprostheses
by: Chenzhitong Chen, et al.
Published: (2026)
by: Chenzhitong Chen, et al.
Published: (2026)
DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
by: Chen, Weidong, et al.
Published: (2025)
by: Chen, Weidong, et al.
Published: (2025)
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
by: Liu, Wenrui, et al.
Published: (2025)
by: Liu, Wenrui, et al.
Published: (2025)
AS-Speech: Adaptive Style For Speech Synthesis
by: Li, Zhipeng, et al.
Published: (2024)
by: Li, Zhipeng, et al.
Published: (2024)
SpeechAgents: Human-Communication Simulation with Multi-Modal Multi-Agent Systems
by: Zhang, Dong, et al.
Published: (2024)
by: Zhang, Dong, et al.
Published: (2024)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
by: Xu, Jing, et al.
Published: (2026)
by: Xu, Jing, et al.
Published: (2026)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
by: Yang, Qian, et al.
Published: (2024)
by: Yang, Qian, et al.
Published: (2024)
Textless Streaming Speech-to-Speech Translation using Semantic Speech Tokens
by: Zhao, Jinzheng, et al.
Published: (2024)
by: Zhao, Jinzheng, et al.
Published: (2024)
WEST: LLM based Speech Toolkit for Speech Understanding, Generation, and Interaction
by: Zhang, Binbin, et al.
Published: (2025)
by: Zhang, Binbin, et al.
Published: (2025)
SLM-SS: Speech Language Model for Generative Speech Separation
by: Li, Tianhua, et al.
Published: (2026)
by: Li, Tianhua, et al.
Published: (2026)
SpeechR: A Benchmark for Speech Reasoning in Large Audio-Language Models
by: Yang, Wanqi, et al.
Published: (2025)
by: Yang, Wanqi, et al.
Published: (2025)
GSRM: Generative Speech Reward Model for Speech RLHF
by: Shen, Maohao, et al.
Published: (2026)
by: Shen, Maohao, et al.
Published: (2026)
LoASR-Bench: Evaluating Large Speech Language Models on Low-Resource Automatic Speech Recognition Across Language Families
by: Chen, Jianan, et al.
Published: (2026)
by: Chen, Jianan, et al.
Published: (2026)
NeuSpeech: Decode Neural signal as Speech
by: Yang, Yiqian, et al.
Published: (2024)
by: Yang, Yiqian, et al.
Published: (2024)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
by: Tian, Jingguang, et al.
Published: (2024)
by: Tian, Jingguang, et al.
Published: (2024)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
by: Deng, Keqi, et al.
Published: (2025)
by: Deng, Keqi, et al.
Published: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
by: Jiang, Feng, et al.
Published: (2025)
by: Jiang, Feng, et al.
Published: (2025)
PolySpeech: Exploring Unified Multitask Speech Models for Competitiveness with Single-task Models
by: Yang, Runyan, et al.
Published: (2024)
by: Yang, Runyan, et al.
Published: (2024)
Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models
by: Tao, Dehua, et al.
Published: (2026)
by: Tao, Dehua, et al.
Published: (2026)
Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression Manipulation
by: Lu, Zhenxuan, et al.
Published: (2026)
by: Lu, Zhenxuan, et al.
Published: (2026)
SynParaSpeech: Automated Synthesis of Paralinguistic Datasets for Speech Generation and Understanding
by: Bai, Bingsong, et al.
Published: (2025)
by: Bai, Bingsong, et al.
Published: (2025)
MORE: Multi-Objective Adversarial Attacks on Speech Recognition
by: Gao, Xiaoxue, et al.
Published: (2026)
by: Gao, Xiaoxue, et al.
Published: (2026)
Too Good to Be True: A Study on Modern Automatic Speech Recognition for the Evaluation of Speech Enhancement
by: de Oliveira, Danilo, et al.
Published: (2026)
by: de Oliveira, Danilo, et al.
Published: (2026)
Similar Items
-
MOSS-TTSD: Text to Spoken Dialogue Generation
by: Zhang, Yuqian, et al.
Published: (2026) -
MOSS-TTS Technical Report
by: Gong, Yitian, et al.
Published: (2026) -
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
by: Gong, Yitian, et al.
Published: (2026) -
MOSS-Audio Technical Report
by: Yang, Chen, et al.
Published: (2026) -
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
by: Gong, Yitian, et al.
Published: (2025)