InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Wang, Dingdong, Xu, Jin, Chu, Ruihang, Guo, Zhifang, Wang, Xiong, Wu, Jincenzi, Yang, Dongchao, Ji, Shengpeng, Lin, Junyang |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
par: Wang, Dingdong, et autres
Publié: (2025)
par: Wang, Dingdong, et autres
Publié: (2025)
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
par: Wang, Dingdong, et autres
Publié: (2024)
par: Wang, Dingdong, et autres
Publié: (2024)
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
par: Zeng, Aohan, et autres
Publié: (2024)
par: Zeng, Aohan, et autres
Publié: (2024)
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
par: Yang, Dongchao, et autres
Publié: (2024)
par: Yang, Dongchao, et autres
Publié: (2024)
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
par: Chen, Xueyuan, et autres
Publié: (2024)
par: Chen, Xueyuan, et autres
Publié: (2024)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
par: Ji, Shengpeng, et autres
Publié: (2024)
par: Ji, Shengpeng, et autres
Publié: (2024)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
par: Wang, Hongbin, et autres
Publié: (2025)
par: Wang, Hongbin, et autres
Publié: (2025)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
par: Yu, Luca Jiang-Tao, et autres
Publié: (2024)
par: Yu, Luca Jiang-Tao, et autres
Publié: (2024)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
par: Khanday, Owais Mujtaba, et autres
Publié: (2025)
par: Khanday, Owais Mujtaba, et autres
Publié: (2025)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
par: Chen, Youjun, et autres
Publié: (2025)
par: Chen, Youjun, et autres
Publié: (2025)
S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
par: Jiang, Feng, et autres
Publié: (2025)
par: Jiang, Feng, et autres
Publié: (2025)
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
par: Cheng, Xize, et autres
Publié: (2025)
par: Cheng, Xize, et autres
Publié: (2025)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
par: Li, Yue, et autres
Publié: (2024)
par: Li, Yue, et autres
Publié: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
par: Park, Seohyun, et autres
Publié: (2025)
par: Park, Seohyun, et autres
Publié: (2025)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
par: Feng, Tiantian, et autres
Publié: (2023)
par: Feng, Tiantian, et autres
Publié: (2023)
VoiceX: A Text-To-Speech Framework for Custom Voices
par: Mertes, Silvan, et autres
Publié: (2024)
par: Mertes, Silvan, et autres
Publié: (2024)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
par: Xiao, Yi, et autres
Publié: (2022)
par: Xiao, Yi, et autres
Publié: (2022)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
par: Khanday, Owais Mujtaba, et autres
Publié: (2025)
par: Khanday, Owais Mujtaba, et autres
Publié: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
par: Nowrin, Sadia, et autres
Publié: (2024)
par: Nowrin, Sadia, et autres
Publié: (2024)
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
par: Lu, Ke-Han, et autres
Publié: (2024)
par: Lu, Ke-Han, et autres
Publié: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
par: Dutta, Satwik, et autres
Publié: (2025)
par: Dutta, Satwik, et autres
Publié: (2025)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
par: Liu, Ailin, et autres
Publié: (2024)
par: Liu, Ailin, et autres
Publié: (2024)
Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
par: Sharma, Roshan, et autres
Publié: (2024)
par: Sharma, Roshan, et autres
Publié: (2024)
A cross-talk robust multichannel VAD model for multiparty agent interactions trained using synthetic re-recordings
par: Han, Hyewon, et autres
Publié: (2024)
par: Han, Hyewon, et autres
Publié: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
par: Futami, Hayato, et autres
Publié: (2025)
par: Futami, Hayato, et autres
Publié: (2025)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
par: Wang, Chen, et autres
Publié: (2024)
par: Wang, Chen, et autres
Publié: (2024)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
par: Wang, He, et autres
Publié: (2025)
par: Wang, He, et autres
Publié: (2025)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
par: Fu, Yu-Kuan, et autres
Publié: (2024)
par: Fu, Yu-Kuan, et autres
Publié: (2024)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
par: Zhou, Dongliang, et autres
Publié: (2025)
par: Zhou, Dongliang, et autres
Publié: (2025)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
par: Xie, Jingran, et autres
Publié: (2025)
par: Xie, Jingran, et autres
Publié: (2025)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
par: Xue, Liumeng, et autres
Publié: (2024)
par: Xue, Liumeng, et autres
Publié: (2024)
Scaling Analysis of Interleaved Speech-Text Language Models
par: Maimon, Gallil, et autres
Publié: (2025)
par: Maimon, Gallil, et autres
Publié: (2025)
Efficient Interleaved Speech Modeling through Knowledge Distillation
par: Nouriborji, Mohammadmahdi, et autres
Publié: (2025)
par: Nouriborji, Mohammadmahdi, et autres
Publié: (2025)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
par: Liu, Alexander H., et autres
Publié: (2025)
par: Liu, Alexander H., et autres
Publié: (2025)
Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
par: Pan, Zihan, et autres
Publié: (2024)
par: Pan, Zihan, et autres
Publié: (2024)
Robust Dual-Modal Speech Keyword Spotting for XR Headsets
par: Cai, Zhuojiang, et autres
Publié: (2024)
par: Cai, Zhuojiang, et autres
Publié: (2024)
BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
par: Wang, Chen, et autres
Publié: (2023)
par: Wang, Chen, et autres
Publié: (2023)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
par: Ji, Shengpeng, et autres
Publié: (2024)
par: Ji, Shengpeng, et autres
Publié: (2024)
Real-Time Word-Level Temporal Segmentation in Streaming Speech Recognition
par: Nishida, Naoto, et autres
Publié: (2025)
par: Nishida, Naoto, et autres
Publié: (2025)
SimpleSpeech 2: Towards Simple and Efficient Text-to-Speech with Flow-based Scalar Latent Transformer Diffusion Models
par: Yang, Dongchao, et autres
Publié: (2024)
par: Yang, Dongchao, et autres
Publié: (2024)
Documents similaires
-
MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
par: Wang, Dingdong, et autres
Publié: (2025) -
A Comparative Study of Discrete Speech Tokens for Semantic-Related Tasks with Large Language Models
par: Wang, Dingdong, et autres
Publié: (2024) -
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
par: Zeng, Aohan, et autres
Publié: (2024) -
SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
par: Yang, Dongchao, et autres
Publié: (2024) -
CoLM-DSR: Leveraging Neural Codec Language Modeling for Multi-Modal Dysarthric Speech Reconstruction
par: Chen, Xueyuan, et autres
Publié: (2024)