Textually Pretrained Speech Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hassid, Michael, Remez, Tal, Nguyen, Tu Anh, Gat, Itai, Conneau, Alexis, Kreuk, Felix, Copet, Jade, Defossez, Alexandre, Synnaeve, Gabriel, Dupoux, Emmanuel, Schwartz, Roy, Adi, Yossi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Simple and Controllable Music Generation
von: Copet, Jade, et al.
Veröffentlicht: (2023)
von: Copet, Jade, et al.
Veröffentlicht: (2023)
Masked Audio Generation using a Single Non-Autoregressive Transformer
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
Discrete Flow Matching
von: Gat, Itai, et al.
Veröffentlicht: (2024)
von: Gat, Itai, et al.
Veröffentlicht: (2024)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024)
von: Tal, Or, et al.
Veröffentlicht: (2024)
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
von: Hassid, Michael, et al.
Veröffentlicht: (2024)
von: Hassid, Michael, et al.
Veröffentlicht: (2024)
Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
von: Hassid, Michael, et al.
Veröffentlicht: (2025)
von: Hassid, Michael, et al.
Veröffentlicht: (2025)
Audio Conditioning for Music Generation via Discrete Bottleneck Features
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
von: Rouard, Simon, et al.
Veröffentlicht: (2024)
Auto-Regressive vs Flow-Matching: a Comparative Study of Modeling Paradigms for Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2025)
von: Tal, Or, et al.
Veröffentlicht: (2025)
An Independence-promoting Loss for Music Generation with Language Models
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
von: Lemercier, Jean-Marie, et al.
Veröffentlicht: (2024)
Low-Resource Self-Supervised Learning with SSL-Enhanced TTS
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
von: Hsu, Po-chun, et al.
Veröffentlicht: (2023)
Transformers are Multi-State RNNs
von: Oren, Matanel, et al.
Veröffentlicht: (2024)
von: Oren, Matanel, et al.
Veröffentlicht: (2024)
Self-Execution Simulation Improves Coding Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2026)
von: Maimon, Gallil, et al.
Veröffentlicht: (2026)
Scaling Analysis of Interleaved Speech-Text Language Models
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
PAST: Phonetic-Acoustic Speech Tokenizer
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
von: Har-Tuv, Nadav, et al.
Veröffentlicht: (2025)
Code Llama: Open Foundation Models for Code
von: Rozière, Baptiste, et al.
Veröffentlicht: (2023)
von: Rozière, Baptiste, et al.
Veröffentlicht: (2023)
On Pruning State-Space LLMs
von: Ghattas, Tamer, et al.
Veröffentlicht: (2025)
von: Ghattas, Tamer, et al.
Veröffentlicht: (2025)
Enhancing TTS Stability in Hebrew using Discrete Semantic Units
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
von: Zeldes, Ella, et al.
Veröffentlicht: (2024)
NAST: Noise Aware Speech Tokenization for Speech Language Models
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
von: Messica, Shoval, et al.
Veröffentlicht: (2024)
RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
von: Gehring, Jonas, et al.
Veröffentlicht: (2024)
von: Gehring, Jonas, et al.
Veröffentlicht: (2024)
LAST: Language Model Aware Speech Tokenization
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
von: Turetzky, Arnon, et al.
Veröffentlicht: (2024)
Short window attention enables long-term memorization
von: Cabannes, Loïc, et al.
Veröffentlicht: (2025)
von: Cabannes, Loïc, et al.
Veröffentlicht: (2025)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
von: Elmakies, Avishai, et al.
Veröffentlicht: (2025)
Spirit LM: Interleaved Spoken and Written Language Model
von: Nguyen, Tu Anh, et al.
Veröffentlicht: (2024)
von: Nguyen, Tu Anh, et al.
Veröffentlicht: (2024)
StressTest: Can YOUR Speech LM Handle the Stress?
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
von: Yosha, Iddo, et al.
Veröffentlicht: (2025)
Stochastic activations
von: Lomeli, Maria, et al.
Veröffentlicht: (2025)
von: Lomeli, Maria, et al.
Veröffentlicht: (2025)
Slamming: Training a Speech Language Model on One GPU in a Day
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
von: Maimon, Gallil, et al.
Veröffentlicht: (2025)
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
von: Yariv, Guy, et al.
Veröffentlicht: (2024)
Simultaneous Speech-to-Speech Translation Without Aligned Data
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
von: de Seyssel, Maureen, et al.
Veröffentlicht: (2023)
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
von: Wei, Yuxiang, et al.
Veröffentlicht: (2025)
Reseñas
von: Mora Hassid
Veröffentlicht: (2025)
von: Mora Hassid
Veröffentlicht: (2025)
Lowering the Horizon on Dark Energy: A Late-Time Response to Early Solutions for the Hubble Tension
von: Adi, Tal
Veröffentlicht: (2025)
von: Adi, Tal
Veröffentlicht: (2025)
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
von: Poli, Maxime, et al.
Veröffentlicht: (2024)
von: Poli, Maxime, et al.
Veröffentlicht: (2024)
fastabx: A library for efficient computation of ABX discriminability
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
Trends in adolescent unions and childbearing in four Central American countries
von: Lisa Remez
Veröffentlicht: (2009)
von: Lisa Remez
Veröffentlicht: (2009)
Formal Language Knowledge Corpus for Retrieval Augmented Generation
von: Zayyad, Majd, et al.
Veröffentlicht: (2024)
von: Zayyad, Majd, et al.
Veröffentlicht: (2024)
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
Corrector Sampling in Language Models
von: Gat, Itai, et al.
Veröffentlicht: (2025)
von: Gat, Itai, et al.
Veröffentlicht: (2025)
Transition Matching: Scalable and Flexible Generative Modeling
von: Shaul, Neta, et al.
Veröffentlicht: (2025)
von: Shaul, Neta, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Simple and Controllable Music Generation
von: Copet, Jade, et al.
Veröffentlicht: (2023) -
Masked Audio Generation using a Single Non-Autoregressive Transformer
von: Ziv, Alon, et al.
Veröffentlicht: (2024) -
Discrete Flow Matching
von: Gat, Itai, et al.
Veröffentlicht: (2024) -
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024) -
The Larger the Better? Improved LLM Code-Generation via Budget Reallocation
von: Hassid, Michael, et al.
Veröffentlicht: (2024)