Kimi-Audio Technical Report
Fuente:
arXiv
Saved in:
| Main Authors: | KimiTeam, Ding, Ding, Ju, Zeqian, Leng, Yichong, Liu, Songxiang, Liu, Tong, Shang, Zeyu, Shen, Kai, Song, Wei, Tan, Xu, Tang, Heyi, Wang, Zhengtao, Wei, Chu, Xin, Yifei, Xu, Xinran, Yu, Jianwei, Zhang, Yutao, Zhou, Xinyu, Charles, Y., Chen, Jun, Chen, Yanru, Du, Yulun, He, Weiran, Hu, Zhenxing, Lai, Guokun, Li, Qingcheng, Liu, Yangyang, Sun, Weidong, Wang, Jianzhou, Wang, Yuzhi, Wu, Yuefeng, Wu, Yuxin, Yang, Dongchao, Yang, Hao, Yang, Ying, Yang, Zhilin, Yin, Aoxiong, Yuan, Ruibin, Zhang, Yutong, Zhou, Zaida |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Kimi-VL Technical Report
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
Kimi Linear: An Expressive, Efficient Attention Architecture
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
MoonCast: High-Quality Zero-Shot Podcast Generation
by: Ju, Zeqian, et al.
Published: (2025)
by: Ju, Zeqian, et al.
Published: (2025)
Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
by: Yang, Dongchao, et al.
Published: (2025)
by: Yang, Dongchao, et al.
Published: (2025)
UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
by: Yang, Dongchao, et al.
Published: (2026)
by: Yang, Dongchao, et al.
Published: (2026)
Kimi k1.5: Scaling Reinforcement Learning with LLMs
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
Kimi K2: Open Agentic Intelligence
by: Kimi Team, et al.
Published: (2025)
by: Kimi Team, et al.
Published: (2025)
ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
by: Yang, Dongchao, et al.
Published: (2025)
by: Yang, Dongchao, et al.
Published: (2025)
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
by: Yang, Zonghan, et al.
Published: (2025)
by: Yang, Zonghan, et al.
Published: (2025)
Kimi K2.5: Visual Agentic Intelligence
by: Kimi Team, et al.
Published: (2026)
by: Kimi Team, et al.
Published: (2026)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
by: Yang, Dongchao, et al.
Published: (2024)
by: Yang, Dongchao, et al.
Published: (2024)
Towards Specialized Generalists: A Multi-Task MoE-LoRA Framework for Domain-Specific LLM Adaptation
by: Yang, Yuxin, et al.
Published: (2026)
by: Yang, Yuxin, et al.
Published: (2026)
The Long-Term Effects of Data Selection in LLM Fine-Tuning
by: Yang, Yuxin, et al.
Published: (2026)
by: Yang, Yuxin, et al.
Published: (2026)
Coarse-to-Fine Lightweight Meta-Embedding for ID-Based Recommendation
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Influence of the COVID‐19 Pandemic on the Prevalence of Depression Among US Adults
by: Jixiang Liu, et al.
Published: (2026)
by: Jixiang Liu, et al.
Published: (2026)
The Schur complements for $SDD_{1}$ matrices and their application to linear complementarity problems
by: Hu, Yang, et al.
Published: (2025)
by: Hu, Yang, et al.
Published: (2025)
Disentangling Task Conflicts in Multi-Task LoRA via Orthogonal Gradient Projection
by: Yang, Ziyu, et al.
Published: (2026)
by: Yang, Ziyu, et al.
Published: (2026)
Cascade-KDE: Robust Time-Series Restoration under Out-of-Distribution Impulse Corruptions
by: Liu, Yuefeng, et al.
Published: (2026)
by: Liu, Yuefeng, et al.
Published: (2026)
V-FAT: Benchmarking Visual Fidelity Against Text-bias
by: Wang, Ziteng, et al.
Published: (2026)
by: Wang, Ziteng, et al.
Published: (2026)
Weakly distance-regular digraphs of diameter 2
by: Wang, Xiangli, et al.
Published: (2025)
by: Wang, Xiangli, et al.
Published: (2025)
The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
by: Yin, Aoxiong, et al.
Published: (2025)
by: Yin, Aoxiong, et al.
Published: (2025)
Time-Series Learning for Proactive Fault Prediction in Distributed Systems with Deep Neural Structures
by: Wang, Yang, et al.
Published: (2025)
by: Wang, Yang, et al.
Published: (2025)
Realizing Tunable Long Persistent Luminescent in Novel Cu2+‐Doped NaGaO2 for Multi‐Level Information Storage and Encryption
by: Liang Liang, et al.
Published: (2024)
by: Liang Liang, et al.
Published: (2024)
Rewritable Optical Information Storage and Dual‐Channel Encryption Based on Near‐Infrared Enhancement of Photostimulated Luminescence Phosphors
by: Yulong Ye, et al.
Published: (2025)
by: Yulong Ye, et al.
Published: (2025)
Multiple Linear Regression‐Enhanced Optical Thermometry via Phonon‐Assisted Back Energy Transfer in Tm3+‐Eu3+ Co‐Doped Phosphors
by: Yu Xue, et al.
Published: (2025)
by: Yu Xue, et al.
Published: (2025)
MVOC: a training-free multiple video object composition method with diffusion models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
A general thermodynamically consistent phase-field-micromechanics model of sintering with coupled diffusion and grain motion
by: Yang, Qingcheng, et al.
Published: (2024)
by: Yang, Qingcheng, et al.
Published: (2024)
A thermodynamically consistent phase‐field‐micromechanics model of sintering with coupled diffusion and grain motion
by: Qingcheng Yang, et al.
Published: (2024)
by: Qingcheng Yang, et al.
Published: (2024)
Every nonsymmetric $4$-class association scheme can be generated by a digraph
by: Yang, Yuefeng
Published: (2024)
by: Yang, Yuefeng
Published: (2024)
Light Up Your Face: A Physically Consistent Dataset and Diffusion Model for Face Fill-Light Enhancement
by: Gong, Jue, et al.
Published: (2026)
by: Gong, Jue, et al.
Published: (2026)
A Generalized Non-local Quasicontinuum Approach for Efficient Modeling of Architected Truss-based Lattice Structures
by: Li, Zi, et al.
Published: (2025)
by: Li, Zi, et al.
Published: (2025)
A Generalized Summation Rule‐Based Nonlocal Quasicontinuum Approach (GSR‐QC) for Efficient Modeling of Architected Lattice Structures
by: Zi Li, et al.
Published: (2025)
by: Zi Li, et al.
Published: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
by: Yang, Dongchao, et al.
Published: (2023)
by: Yang, Dongchao, et al.
Published: (2023)
Muon is Scalable for LLM Training
by: Liu, Jingyuan, et al.
Published: (2025)
by: Liu, Jingyuan, et al.
Published: (2025)
Two-disjoint-cycle-cover vertex pancyclicity of split-star networks
by: Liu, Chang, et al.
Published: (2026)
by: Liu, Chang, et al.
Published: (2026)
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
by: Xin, Detai, et al.
Published: (2024)
by: Xin, Detai, et al.
Published: (2024)
SCMM: Calibrating Cross-modal Representations for Text-Based Person Search
by: Liu, Jing, et al.
Published: (2023)
by: Liu, Jing, et al.
Published: (2023)
NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
by: Ju, Zeqian, et al.
Published: (2024)
by: Ju, Zeqian, et al.
Published: (2024)
Inconsistent performance of multi‐type genomic data in phylogenomics of neuropteridan insects, with solutions toward conflicting results
by: Ruyue Zhang, et al.
Published: (2025)
by: Ruyue Zhang, et al.
Published: (2025)
Kinodynamic Model Predictive Control for Energy Efficient Locomotion of Legged Robots with Parallel Elasticity
by: Zhuang, Yulun, et al.
Published: (2025)
by: Zhuang, Yulun, et al.
Published: (2025)
Similar Items
-
Kimi-VL Technical Report
by: Kimi Team, et al.
Published: (2025) -
Kimi Linear: An Expressive, Efficient Attention Architecture
by: Kimi Team, et al.
Published: (2025) -
MoonCast: High-Quality Zero-Shot Podcast Generation
by: Ju, Zeqian, et al.
Published: (2025) -
Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
by: Yang, Dongchao, et al.
Published: (2025) -
UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
by: Yang, Dongchao, et al.
Published: (2026)