Saved in:
| Main Authors: | Yang, Dongchao, Xie, Yuxin, Yin, Yuguo, Wang, Zheyu, Yi, Xiaoyu, Zhu, Gongxi, Weng, Xiaolong, Xiong, Zihan, Ma, Yingzhe, Cong, Dading, Liu, Jingliang, Huang, Zihang, Ru, Jinghan, Huang, Rongjie, Wan, Haoran, Wang, Peixu, Yu, Kuoxi, Wang, Helin, Liang, Liming, Zhuang, Xianwei, Wang, Yuanyuan, Dingdong, Wang, Guo, Haohan, Cao, Junjie, Ju, Zeqian, Liu, Songxiang, Cao, Yuewen, Weng, Heming, Zou, Yuexian |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2601.10547 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
by: Yang, Dongchao, et al.
Published: (2024)
by: Yang, Dongchao, et al.
Published: (2024)
UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
by: Yang, Dongchao, et al.
Published: (2026)
by: Yang, Dongchao, et al.
Published: (2026)
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
by: Zhuang, Xianwei, et al.
Published: (2025)
by: Zhuang, Xianwei, et al.
Published: (2025)
ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
by: Yin, Yuguo, et al.
Published: (2025)
by: Yin, Yuguo, et al.
Published: (2025)
ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
by: Yang, Dongchao, et al.
Published: (2025)
by: Yang, Dongchao, et al.
Published: (2025)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
by: Yang, Dongchao, et al.
Published: (2024)
by: Yang, Dongchao, et al.
Published: (2024)
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
by: Zhuang, Xianwei, et al.
Published: (2025)
by: Zhuang, Xianwei, et al.
Published: (2025)
Do we really have to filter out random noise in pre-training data for language models?
by: Ru, Jinghan, et al.
Published: (2025)
by: Ru, Jinghan, et al.
Published: (2025)
Omni-AutoThink: Adaptive Multimodal Reasoning via Reinforcement Learning
by: Yang, Dongchao, et al.
Published: (2025)
by: Yang, Dongchao, et al.
Published: (2025)
UniAudio 1.5: Large Language Model-driven Audio Codec is A Few-shot Audio Task Learner
by: Yang, Dongchao, et al.
Published: (2024)
by: Yang, Dongchao, et al.
Published: (2024)
Artificial Geographically Weighted Neural Network: A Novel Framework for Spatial Analysis with Geographically Weighted Layers
by: Cao, Jianfei, et al.
Published: (2025)
by: Cao, Jianfei, et al.
Published: (2025)
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
by: Fu, Siyuan, et al.
Published: (2025)
by: Fu, Siyuan, et al.
Published: (2025)
DermoGPT: Open Weights and Open Data for Morphology-Grounded Dermatological Reasoning MLLMs
by: Ru, Jinghan, et al.
Published: (2026)
by: Ru, Jinghan, et al.
Published: (2026)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
by: Wang, Dingdong, et al.
Published: (2024)
by: Wang, Dingdong, et al.
Published: (2024)
SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization
by: Luo, Jiehui, et al.
Published: (2025)
by: Luo, Jiehui, et al.
Published: (2025)
Anatomy-Slot: Unsupervised Anatomical Factorization for Homologous Bilateral Reasoning in Retinal Diagnosis
by: Ma, Yingzhe, et al.
Published: (2026)
by: Ma, Yingzhe, et al.
Published: (2026)
MoMu-Diffusion: On Learning Long-Term Motion-Music Synchronization and Correspondence
by: You, Fuming, et al.
Published: (2024)
by: You, Fuming, et al.
Published: (2024)
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
by: Chen, Xueyuan, et al.
Published: (2024)
by: Chen, Xueyuan, et al.
Published: (2024)
MIGA: Mixture-of-Experts with Group Aggregation for Stock Market Prediction
by: Yu, Zhaojian, et al.
Published: (2024)
by: Yu, Zhaojian, et al.
Published: (2024)
Detecting Mislabeled and Corrupted Data via Pointwise Mutual Information
by: Yang, Jinghan, et al.
Published: (2025)
by: Yang, Jinghan, et al.
Published: (2025)
Composite phenotypes and multiple organ systems aging clocks predict cognitive decline
by: Yingzhe Wang
Published: (2025)
by: Yingzhe Wang
Published: (2025)
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
UniAudio: An Audio Foundation Model Toward Universal Audio Generation
by: Yang, Dongchao, et al.
Published: (2023)
by: Yang, Dongchao, et al.
Published: (2023)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
by: Wang, Dingdong, et al.
Published: (2025)
by: Wang, Dingdong, et al.
Published: (2025)
Unveiling the dynamic evolution of innovation networks in emerging economies: A complex network approach
by: Zeqian Wang, et al.
Published: (2024)
by: Zeqian Wang, et al.
Published: (2024)
Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
by: Tang, Lexiang, et al.
Published: (2025)
by: Tang, Lexiang, et al.
Published: (2025)
Fortunellin ameliorates LPS‐induced acute lung injury, inflammation, and collagen deposition by restraining the TLR4/NF‐κB/NLRP3 pathway
by: Danjuan Liu, et al.
Published: (2024)
by: Danjuan Liu, et al.
Published: (2024)
Learning What Matters Now: Dynamic Preference Inference under Contextual Shifts
by: Cao, Xianwei, et al.
Published: (2026)
by: Cao, Xianwei, et al.
Published: (2026)
PivotMesh: Generic 3D Mesh Generation via Pivot Vertices Guidance
by: Weng, Haohan, et al.
Published: (2024)
by: Weng, Haohan, et al.
Published: (2024)
Default Contagion, Matrix Approximation, and Control in Sparse Financial Networks
by: Zhang, Aoxin, et al.
Published: (2026)
by: Zhang, Aoxin, et al.
Published: (2026)
Deep Reinforcement Learning-aided Transmission Design for Energy-efficient Link Optimization in Vehicular Communications
by: Wang, Zhengpeng, et al.
Published: (2024)
by: Wang, Zhengpeng, et al.
Published: (2024)
BeamCKM: A Framework of Channel Knowledge Map Construction for Multi-Antenna Systems
by: Wang, Haohan, et al.
Published: (2025)
by: Wang, Haohan, et al.
Published: (2025)
MeshFIM: Local Low-Poly Mesh Editing via Fill-in-the-Middle Autoregressive Generation
by: Yang, Dingdong, et al.
Published: (2026)
by: Yang, Dingdong, et al.
Published: (2026)
Beamforming-Codebook-Aware Channel Knowledge Map Construction for Multi-Antenna Systems
by: Wang, Haohan, et al.
Published: (2025)
by: Wang, Haohan, et al.
Published: (2025)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
by: Wang, Helin, et al.
Published: (2025)
by: Wang, Helin, et al.
Published: (2025)
Homogeneous fractional integral operators on weighted Lebesgue, Morrey and Campanato spaces
by: Du, Jingliang, et al.
Published: (2025)
by: Du, Jingliang, et al.
Published: (2025)
Influence of Interstitial Segregation on Grain Boundary Cohesion in Ferritic Steels
by: Jingliang Wang, et al.
Published: (2026)
by: Jingliang Wang, et al.
Published: (2026)
Bound states and atomic interaction in giant atom waveguide QED with dispersive coupling
by: Weng, Mingzhu, et al.
Published: (2024)
by: Weng, Mingzhu, et al.
Published: (2024)
A mean curvature type flow with capillary boundary in a unit ball
by: Wang, Guofang, et al.
Published: (2020)
by: Wang, Guofang, et al.
Published: (2020)
Similar Items
-
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
by: Yang, Dongchao, et al.
Published: (2024) -
UniAudio 2.0: A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization
by: Yang, Dongchao, et al.
Published: (2026) -
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
by: Zhuang, Xianwei, et al.
Published: (2025) -
ATRI: Mitigating Multilingual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors
by: Yin, Yuguo, et al.
Published: (2025) -
ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling
by: Yang, Dongchao, et al.
Published: (2025)