Gespeichert in:
| Hauptverfasser: | Sankar, Sanjana, Lenglet, Martin, Bailly, Gerard, Beautemps, Denis, Hueber, Thomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2501.04799 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
Is Self-Supervised Learning Enough to Fill in the Gap? A Study on Speech Inpainting
von: Asaad, Ihab, et al.
Veröffentlicht: (2024)
von: Asaad, Ihab, et al.
Veröffentlicht: (2024)
Generative Pre-training for Speech with Flow Matching
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
von: Liu, Alexander H., et al.
Veröffentlicht: (2023)
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2025)
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2025)
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
von: Zhu, Yongxin, et al.
Veröffentlicht: (2024)
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
von: Yuhang, Yang, et al.
Veröffentlicht: (2024)
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
von: Romero-Díaz, Jacobo, et al.
Veröffentlicht: (2025)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
von: Pareras, Oriol, et al.
Veröffentlicht: (2025)
von: Pareras, Oriol, et al.
Veröffentlicht: (2025)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
von: Liu, Alexander H., et al.
Veröffentlicht: (2025)
A Simple Method to Enhance Pre-trained Language Models with Speech Tokens for Classification
von: Calbucura, Nicolas, et al.
Veröffentlicht: (2025)
von: Calbucura, Nicolas, et al.
Veröffentlicht: (2025)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
von: Ma, Ziyang, et al.
Veröffentlicht: (2023)
MOSS-Speech: Towards True Speech-to-Speech Models Without Text Guidance
von: Zhao, Xingjian, et al.
Veröffentlicht: (2025)
von: Zhao, Xingjian, et al.
Veröffentlicht: (2025)
RASMALAI: Resources for Adaptive Speech Modeling in Indian Languages with Accents and Intonations
von: Sankar, Ashwin, et al.
Veröffentlicht: (2025)
von: Sankar, Ashwin, et al.
Veröffentlicht: (2025)
SpeechAlign: a Framework for Speech Translation Alignment Evaluation
von: Alastruey, Belen, et al.
Veröffentlicht: (2023)
von: Alastruey, Belen, et al.
Veröffentlicht: (2023)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
von: Labrak, Yanis, et al.
Veröffentlicht: (2025)
On Leveraging Encoder-only Pre-trained Language Models for Effective Keyphrase Generation
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
von: Wang, Dingdong, et al.
Veröffentlicht: (2025)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2024)
von: Toyin, Hawau Olamide, et al.
Veröffentlicht: (2024)
Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2025)
Text-to-Code Generation with Modality-relative Pre-training
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
von: Christopoulou, Fenia, et al.
Veröffentlicht: (2024)
Continuous Speech Tokenizer in Text To Speech
von: Li, Yixing, et al.
Veröffentlicht: (2024)
von: Li, Yixing, et al.
Veröffentlicht: (2024)
LLM-to-Speech: A Synthetic Data Pipeline for Training Dialectal Text-to-Speech Models
von: Khamis, Ahmed Khaled, et al.
Veröffentlicht: (2026)
von: Khamis, Ahmed Khaled, et al.
Veröffentlicht: (2026)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Attentive Merging of Hidden Embeddings from Pre-trained Speech Model for Anti-spoofing Detection
von: Pan, Zihan, et al.
Veröffentlicht: (2024)
von: Pan, Zihan, et al.
Veröffentlicht: (2024)
Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization
von: Tang, Yun, et al.
Veröffentlicht: (2025)
von: Tang, Yun, et al.
Veröffentlicht: (2025)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
von: Yadav, Hemant, et al.
Veröffentlicht: (2024)
Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
von: Kuan, Chun-Yi, et al.
Veröffentlicht: (2024)
A multilingual training strategy for low resource Text to Speech
von: Amalas, Asma, et al.
Veröffentlicht: (2024)
von: Amalas, Asma, et al.
Veröffentlicht: (2024)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
von: Rossenbach, Nick, et al.
Veröffentlicht: (2024)
LatentSpeech: Latent Diffusion for Text-To-Speech Generation
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
von: Lou, Haowei, et al.
Veröffentlicht: (2024)
Streaming Speech-to-Text Translation with a SpeechLLM
von: Parcollet, Titouan, et al.
Veröffentlicht: (2026)
von: Parcollet, Titouan, et al.
Veröffentlicht: (2026)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Streaming Speech-to-Confusion Network Speech Recognition
von: Filimonov, Denis, et al.
Veröffentlicht: (2023)
von: Filimonov, Denis, et al.
Veröffentlicht: (2023)
Revisiting Interpolation Augmentation for Speech-to-Text Generation
von: Xu, Chen, et al.
Veröffentlicht: (2024)
von: Xu, Chen, et al.
Veröffentlicht: (2024)
Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
von: Xie, Jingran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Scaling Speech-Text Pre-training with Synthetic Interleaved Data
von: Zeng, Aohan, et al.
Veröffentlicht: (2024) -
Is Self-Supervised Learning Enough to Fill in the Gap? A Study on Speech Inpainting
von: Asaad, Ihab, et al.
Veröffentlicht: (2024) -
Generative Pre-training for Speech with Flow Matching
von: Liu, Alexander H., et al.
Veröffentlicht: (2023) -
Speech-FT: Merging Pre-trained And Fine-Tuned Speech Representation Models For Cross-Task Generalization
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025) -
Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2025)