FastSLM: Hierarchical Temporal Abstraction for Efficient Long-Form Speech Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Junseok, Lee, Sangyong, Chun, Chang-Jae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention
von: Lee, HyeYoung, et al.
Veröffentlicht: (2025)
von: Lee, HyeYoung, et al.
Veröffentlicht: (2025)
Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
von: Lee, Yongjoon, et al.
Veröffentlicht: (2024)
von: Lee, Yongjoon, et al.
Veröffentlicht: (2024)
Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
von: Ahn, Hoseong, et al.
Veröffentlicht: (2026)
von: Ahn, Hoseong, et al.
Veröffentlicht: (2026)
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
von: Lee, Seo-Hyun, et al.
Veröffentlicht: (2023)
von: Lee, Seo-Hyun, et al.
Veröffentlicht: (2023)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
von: Qi, Xin, et al.
Veröffentlicht: (2024)
von: Qi, Xin, et al.
Veröffentlicht: (2024)
Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
von: Truong, Duc-Tuan, et al.
Veröffentlicht: (2024)
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech
von: Bak, Taejun, et al.
Veröffentlicht: (2024)
von: Bak, Taejun, et al.
Veröffentlicht: (2024)
Bangla-WhisperDiar: Fine-Tuning Whisper and PyAnnote for Bangla Long-Form Speech Recognition and Speaker Diarization
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
von: Bhuiyan, Mohammed Aman, et al.
Veröffentlicht: (2026)
Speech Recognition-based Feature Extraction for Enhanced Automatic Severity Classification in Dysarthric Speech
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
von: Choi, Yerin, et al.
Veröffentlicht: (2024)
Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech from EEG Signals
von: Lee, Jung-Sun, et al.
Veröffentlicht: (2024)
von: Lee, Jung-Sun, et al.
Veröffentlicht: (2024)
Seewo's Submission to MLC-SLM: Lessons learned from Speech Reasoning Language Models
von: Li, Bo, et al.
Veröffentlicht: (2025)
von: Li, Bo, et al.
Veröffentlicht: (2025)
Removing Speaker Information from Speech Representation using Variable-Length Soft Pooling
von: Hwang, Injune, et al.
Veröffentlicht: (2024)
von: Hwang, Injune, et al.
Veröffentlicht: (2024)
Temporal-Aware Iterative Speech Model for Dementia Detection
von: Ugwu, Chukwuemeka, et al.
Veröffentlicht: (2025)
von: Ugwu, Chukwuemeka, et al.
Veröffentlicht: (2025)
Continuous Modeling of the Denoising Process for Speech Enhancement Based on Deep Learning
von: Guo, Zilu, et al.
Veröffentlicht: (2023)
von: Guo, Zilu, et al.
Veröffentlicht: (2023)
Active Learning with Task Adaptation Pre-training for Speech Emotion Recognition
von: Li, Dongyuan, et al.
Veröffentlicht: (2024)
von: Li, Dongyuan, et al.
Veröffentlicht: (2024)
HASS: Hierarchical Simulation of Logopenic Aphasic Speech for Scalable PPA Detection
von: Li, Harrison, et al.
Veröffentlicht: (2026)
von: Li, Harrison, et al.
Veröffentlicht: (2026)
Articulatory Feature Prediction from Surface EMG during Speech Production
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
von: Lee, Jihwan, et al.
Veröffentlicht: (2025)
Incremental FastPitch: Chunk-based High Quality Text to Speech
von: Du, Muyang, et al.
Veröffentlicht: (2024)
von: Du, Muyang, et al.
Veröffentlicht: (2024)
Perceiver-Prompt: Flexible Speaker Adaptation in Whisper for Chinese Disordered Speech Recognition
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
von: Jiang, Yicong, et al.
Veröffentlicht: (2024)
Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanglei, et al.
Veröffentlicht: (2025)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
von: Yen, Hao, et al.
Veröffentlicht: (2024)
von: Yen, Hao, et al.
Veröffentlicht: (2024)
A Two-Stage Hierarchical Deep Filtering Framework for Real-Time Speech Enhancement
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
von: Lu, Shenghui, et al.
Veröffentlicht: (2025)
CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
von: Wang, Helin, et al.
Veröffentlicht: (2025)
von: Wang, Helin, et al.
Veröffentlicht: (2025)
MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and Hierarchical Cooperative Attention
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
von: Jiao, Xinxin, et al.
Veröffentlicht: (2024)
Temporal Information Reconstruction and Non-Aligned Residual in Spiking Neural Networks for Speech Classification
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
von: Polák, Peter, et al.
Veröffentlicht: (2023)
von: Polák, Peter, et al.
Veröffentlicht: (2023)
An Effective Context-Balanced Adaptation Approach for Long-Tailed Speech Recognition
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
von: Wang, Yi-Cheng, et al.
Veröffentlicht: (2024)
EmoSphere-SER: Enhancing Speech Emotion Recognition Through Spherical Representation with Auxiliary Classification
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2025)
Efficient Finetuning for Dimensional Speech Emotion Recognition in the Age of Transformers
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
von: Sampath, Aneesha, et al.
Veröffentlicht: (2025)
SepPrune: Structured Pruning for Efficient Deep Speech Separation
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
FreeAudio: Training-Free Timing Planning for Controllable Long-Form Text-to-Audio Generation
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxuan, et al.
Veröffentlicht: (2025)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-spoofing
von: Liu, Tianchi, et al.
Veröffentlicht: (2025)
von: Liu, Tianchi, et al.
Veröffentlicht: (2025)
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
von: Nguyen, Tan Dat, et al.
Veröffentlicht: (2024)
ECTSpeech: Enhancing Efficient Speech Synthesis via Easy Consistency Tuning
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
von: Zhu, Tao, et al.
Veröffentlicht: (2025)
Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
von: Kim, Nam-Gyu, et al.
Veröffentlicht: (2025)
Token Pruning in Audio Transformers: Optimizing Performance and Decoding Patch Importance
von: Lee, Taehan, et al.
Veröffentlicht: (2025)
von: Lee, Taehan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Toward Efficient Speech Emotion Recognition via Spectral Learning and Attention
von: Lee, HyeYoung, et al.
Veröffentlicht: (2025) -
Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
von: Lee, Yongjoon, et al.
Veröffentlicht: (2024) -
Whisper-CD: Accurate Long-Form Speech Recognition using Multi-Negative Contrastive Decoding
von: Ahn, Hoseong, et al.
Veröffentlicht: (2026) -
Neural Speech Embeddings for Speech Synthesis Based on Deep Generative Networks
von: Lee, Seo-Hyun, et al.
Veröffentlicht: (2023) -
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
von: Qi, Xin, et al.
Veröffentlicht: (2024)