Speech-XL: Towards Long-Form Speech Understanding in Large Speech Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Haoqin, Lyu, Chenyang, Zhao, Shiwan, Ni, Xuanfan, Kong, Xiangyu, Wang, Longyue, Luo, Weihua, Qin, Yong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
von: Yang, Fei, et al.
Veröffentlicht: (2026)
von: Yang, Fei, et al.
Veröffentlicht: (2026)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
The Affective Bridge: Preserving Speech Representations while Enhancing Deepfake Detection vian emotional Constraints
von: Li, Yupei, et al.
Veröffentlicht: (2025)
von: Li, Yupei, et al.
Veröffentlicht: (2025)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
Marco-ASR: A Principled and Metric-Driven Framework for Fine-Tuning Large-Scale ASR Models for Domain Adaptation
von: Ni, Xuanfan, et al.
Veröffentlicht: (2025)
von: Ni, Xuanfan, et al.
Veröffentlicht: (2025)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
Zero- and One-Shot Data Augmentation for Sentence-Level Dysarthric Speech Recognition in Constrained Scenarios
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
von: Wang, Shiyao, et al.
Veröffentlicht: (2025)
Marco-Voice Technical Report
von: Tian, Fengping, et al.
Veröffentlicht: (2025)
von: Tian, Fengping, et al.
Veröffentlicht: (2025)
Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
von: Wang, Shiyao, et al.
Veröffentlicht: (2024)
von: Wang, Shiyao, et al.
Veröffentlicht: (2024)
DualSpeechLM: Towards Unified Speech Understanding and Generation via Dual Speech Token Modeling with Large Language Models
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
von: Wang, Yuanyuan, et al.
Veröffentlicht: (2025)
ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
DIFFA: Large Language Diffusion Models Can Listen and Understand
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
Long-Form Speech Generation with Spoken Language Models
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
von: Park, Se Jin, et al.
Veröffentlicht: (2024)
FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2025)
WildElder: A Chinese Elderly Speech Dataset from the Wild with Fine-Grained Manual Annotations
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
von: Shao, Mingchen, et al.
Veröffentlicht: (2025)
SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
von: Zhang, Xin, et al.
Veröffentlicht: (2023)
DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Emotion Recognition through Multi-Granularity Cross-Modal Alignment
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
von: Wang, Xuechen, et al.
Veröffentlicht: (2024)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Enhancing Emotion Recognition in Incomplete Data: A Novel Cross-Modal Alignment, Reconstruction, and Refinement Framework
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
Can Large Audio Language Models Understand Audio Well? Speech, Scene and Events Understanding Benchmark for LALMs
von: Yin, Han, et al.
Veröffentlicht: (2025)
von: Yin, Han, et al.
Veröffentlicht: (2025)
DIFFA-2: A Practical Diffusion Large Language Model for General Audio Understanding
von: Zhou, Jiaming, et al.
Veröffentlicht: (2026)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2026)
DisCo-Speech: Controllable Zero-Shot Speech Generation with A Disentangled Speech Codec
von: Li, Tao, et al.
Veröffentlicht: (2025)
von: Li, Tao, et al.
Veröffentlicht: (2025)
Large Speech Model Enabled Semantic Communication
von: Tian, Yun, et al.
Veröffentlicht: (2025)
von: Tian, Yun, et al.
Veröffentlicht: (2025)
SLM-SS: Speech Language Model for Generative Speech Separation
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
von: Li, Tianhua, et al.
Veröffentlicht: (2026)
Understanding Frechet Speech Distance for Synthetic Speech Quality Evaluation
von: Kim, June-Woo, et al.
Veröffentlicht: (2026)
von: Kim, June-Woo, et al.
Veröffentlicht: (2026)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
EmoSURA: Towards Accurate Evaluation of Detailed and Long-Context Emotional Speech Captions
von: Jing, Xin, et al.
Veröffentlicht: (2026)
von: Jing, Xin, et al.
Veröffentlicht: (2026)
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
Robust Long-Form Bangla Speech Processing: Automatic Speech Recognition and Speaker Diarization
von: Chowdhury, MD. Sagor, et al.
Veröffentlicht: (2026)
von: Chowdhury, MD. Sagor, et al.
Veröffentlicht: (2026)
BLSP-Emo: Towards Empathetic Large Speech-Language Models
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
von: Fang, Yuanbo, et al.
Veröffentlicht: (2025)
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LongSpeech: A Scalable Benchmark for Transcription, Translation and Understanding in Long Speech
von: Yang, Fei, et al.
Veröffentlicht: (2026) -
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
von: Sun, Haoqin, et al.
Veröffentlicht: (2024) -
MECap-R1: Emotion-aware Policy with Reinforcement Learning for Multimodal Emotion Captioning
von: Sun, Haoqin, et al.
Veröffentlicht: (2025) -
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation
von: Wang, Hui, et al.
Veröffentlicht: (2025) -
The Affective Bridge: Preserving Speech Representations while Enhancing Deepfake Detection vian emotional Constraints
von: Li, Yupei, et al.
Veröffentlicht: (2025)