MOSS-Audio Technical Report
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Chen, Yu, Chufan, Chen, Hanfu, Zhu, Jie, Chen, Jingqi, Chen, Ke, Wang, Wenxuan, Wang, Yang, Jiang, Yaozhou, Jiang, Yi, Lin, Zhengyuan, Chen, Ziqi, Fei, Zhaoye, Liu, Chenghao, Zhan, Jun, Yu, Kang, Huang, Kexin, Chen, Mingshu, Cheng, Qinyuan, Li, Ruixiao, Li, Shimin, Wang, Songlin, Gao, Yang, Zhang, Yiyang, Qiu, Xipeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MOSS-TTS Technical Report
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
MOSS Transcribe Diarize Technical Report
von: AI, MOSI., et al.
Veröffentlicht: (2026)
von: AI, MOSI., et al.
Veröffentlicht: (2026)
MOSS-TTSD: Text to Spoken Dialogue Generation
von: Zhang, Yuqian, et al.
Veröffentlicht: (2026)
von: Zhang, Yuqian, et al.
Veröffentlicht: (2026)
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
von: Gong, Yitian, et al.
Veröffentlicht: (2026)
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
von: Zhao, Xingjian, et al.
Veröffentlicht: (2025)
von: Zhao, Xingjian, et al.
Veröffentlicht: (2025)
MOSS-VoiceGenerator: Create Realistic Voices with Natural Language Descriptions
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
von: Huang, Kexin, et al.
Veröffentlicht: (2026)
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
How Attention Sinks Emerge in Large Language Models: An Interpretability Perspective
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
von: Peng, Runyu, et al.
Veröffentlicht: (2026)
MOVA: Towards Scalable and Synchronized Video-Audio Generation
von: OpenMOSS Team, et al.
Veröffentlicht: (2026)
von: OpenMOSS Team, et al.
Veröffentlicht: (2026)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
Agent Alignment in Evolving Social Norms
von: Li, Shimin, et al.
Veröffentlicht: (2024)
von: Li, Shimin, et al.
Veröffentlicht: (2024)
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation
von: Deng, Ruifan, et al.
Veröffentlicht: (2025)
von: Deng, Ruifan, et al.
Veröffentlicht: (2025)
Technical Report for ActivityNet Challenge 2022 -- Temporal Action Localization
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
Technical Report for Soccernet 2023 -- Dense Video Captioning
von: Ruan, Zheng, et al.
Veröffentlicht: (2024)
von: Ruan, Zheng, et al.
Veröffentlicht: (2024)
Technical Report for SoccerNet Challenge 2022 -- Replay Grounding Task
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
RAIRS: Optimizing Redundant Assignment and List Layout for IVF-Based ANN Search
von: Yang, Zehai, et al.
Veröffentlicht: (2026)
von: Yang, Zehai, et al.
Veröffentlicht: (2026)
LITS: An Optimized Learned Index for Strings (An Extended Version)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
von: Yang, Yifan, et al.
Veröffentlicht: (2024)
Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition
von: Wang, Peng, et al.
Veröffentlicht: (2026)
von: Wang, Peng, et al.
Veröffentlicht: (2026)
Floquet Nonadiabatic Mixed Quantum-Classical Dynamics in Laser-Dressed Solid Systems
von: Chen, Jingqi, et al.
Veröffentlicht: (2024)
von: Chen, Jingqi, et al.
Veröffentlicht: (2024)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
X-VC: Zero-shot Streaming Voice Conversion in Codec Space
von: Zheng, Qixi, et al.
Veröffentlicht: (2026)
von: Zheng, Qixi, et al.
Veröffentlicht: (2026)
Improving the Computational Efficiency and Explainability of GeoAggregator
von: Deng, Rui, et al.
Veröffentlicht: (2025)
von: Deng, Rui, et al.
Veröffentlicht: (2025)
GeoAggregator: An Efficient Transformer Model for Geo-Spatial Tabular Data
von: Deng, Rui, et al.
Veröffentlicht: (2025)
von: Deng, Rui, et al.
Veröffentlicht: (2025)
Glycoside Hydrolase Family 16 Enzyme RsEG146 From Rhizoctonia solani AG1 IA Induces Cell Death and Triggers Defence Response in Nicotiana tabacum
von: Chen Chen, et al.
Veröffentlicht: (2025)
von: Chen Chen, et al.
Veröffentlicht: (2025)
InsCL: A Data-efficient Continual Learning Paradigm for Fine-tuning Large Language Models with Instructions
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
von: Wang, Yifan, et al.
Veröffentlicht: (2024)
Computer‐Aided Technology for Bioactive Protein Design and Clinical Application
von: Chufan Wang, et al.
Veröffentlicht: (2025)
von: Chufan Wang, et al.
Veröffentlicht: (2025)
Reflective Unit Test Generation for Precise Type Error Detection with Large Language Models
von: Yang, Chen, et al.
Veröffentlicht: (2025)
von: Yang, Chen, et al.
Veröffentlicht: (2025)
World Modeling Makes a Better Planner: Dual Preference Optimization for Embodied Task Planning
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
von: Wang, Siyin, et al.
Veröffentlicht: (2025)
Largest $2$-regular Subgraphs in complete $S$-partite Graphs
von: Jiang, Yiyang, et al.
Veröffentlicht: (2026)
von: Jiang, Yiyang, et al.
Veröffentlicht: (2026)
CLIPDrag: Combining Text-based and Drag-based Instructions for Image Editing
von: Jiang, Ziqi, et al.
Veröffentlicht: (2024)
von: Jiang, Ziqi, et al.
Veröffentlicht: (2024)
Exploring Cross-Modal Flows for Few-Shot Learning
von: Jiang, Ziqi, et al.
Veröffentlicht: (2025)
von: Jiang, Ziqi, et al.
Veröffentlicht: (2025)
How to Mitigate Overfitting in Weak-to-strong Generalization?
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
von: Shi, Junhao, et al.
Veröffentlicht: (2025)
VisuoThink: Empowering LVLM Reasoning with Multimodal Tree Search
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
von: Wang, Yikun, et al.
Veröffentlicht: (2025)
Clarifying Semantics of In-Context Examples for Unit Test Generation
von: Yang, Chen, et al.
Veröffentlicht: (2025)
von: Yang, Chen, et al.
Veröffentlicht: (2025)
Vibration Suppression and Trajectory Tracking Control of Flexible Joint Manipulator Based on PSO Algorithm and Fixed-Time Control
von: Yan Guan, et al.
Veröffentlicht: (2024)
von: Yan Guan, et al.
Veröffentlicht: (2024)
Sampling from Constrained Gibbs Measures: with Applications to High-Dimensional Bayesian Inference
von: Wang, Ruixiao, et al.
Veröffentlicht: (2026)
von: Wang, Ruixiao, et al.
Veröffentlicht: (2026)
Authors' Reply: “Deep Learning for Staging Periodontitis Using Panoramic Radiographs”
von: Xin Li, et al.
Veröffentlicht: (2025)
von: Xin Li, et al.
Veröffentlicht: (2025)
On bicanonical maps of threefolds of general type with large volumes
von: Jiang, Chen, et al.
Veröffentlicht: (2025)
von: Jiang, Chen, et al.
Veröffentlicht: (2025)
Semi-Clairvoyant Scheduling of Speculative Decoding Requests to Minimize LLM Inference Latency
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
von: Li, Ruixiao, et al.
Veröffentlicht: (2025)
The Effect of Income Polarization on Crime: Evidence From Court Judicial Documents in China
von: Jingqi Liu, et al.
Veröffentlicht: (2025)
von: Jingqi Liu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MOSS-TTS Technical Report
von: Gong, Yitian, et al.
Veröffentlicht: (2026) -
MOSS Transcribe Diarize Technical Report
von: AI, MOSI., et al.
Veröffentlicht: (2026) -
MOSS-TTSD: Text to Spoken Dialogue Generation
von: Zhang, Yuqian, et al.
Veröffentlicht: (2026) -
MOSS-Audio-Tokenizer: Scaling Audio Tokenizers for Future Audio Foundation Models
von: Gong, Yitian, et al.
Veröffentlicht: (2026) -
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
von: Zhao, Xingjian, et al.
Veröffentlicht: (2025)