Spirit LM: Interleaved Spoken and Written Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nguyen, Tu Anh, Muller, Benjamin, Yu, Bokai, Costa-jussa, Marta R., Elbayad, Maha, Popuri, Sravya, Ropers, Christophe, Duquenne, Paul-Ambroise, Algayres, Robin, Mavlyutov, Ruslan, Gat, Itai, Williamson, Mary, Synnaeve, Gabriel, Pino, Juan, Sagot, Benoit, Dupoux, Emmanuel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Investigating Decoder-only Large Language Models for Speech-to-text Translation
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024)
Merging Text Transformer Models from Different Initializations
von: Verma, Neha, et al.
Veröffentlicht: (2024)
von: Verma, Neha, et al.
Veröffentlicht: (2024)
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
von: Chambon, Pierre, et al.
Veröffentlicht: (2025)
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
von: Poli, Maxime, et al.
Veröffentlicht: (2025)
Unified Vision-Language Modeling via Concept Space Alignment
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)
Interference Matrix: Quantifying Cross-Lingual Interference in Transformer Encoders
von: Alastruey, Belen, et al.
Veröffentlicht: (2025)
von: Alastruey, Belen, et al.
Veröffentlicht: (2025)
Towards Massive Multilingual Holistic Bias
von: Tan, Xiaoqing Ellen, et al.
Veröffentlicht: (2024)
von: Tan, Xiaoqing Ellen, et al.
Veröffentlicht: (2024)
Large Concept Models: Language Modeling in a Sentence Representation Space
von: LCM team, et al.
Veröffentlicht: (2024)
von: LCM team, et al.
Veröffentlicht: (2024)
Textually Pretrained Speech Language Models
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
von: Hassid, Michael, et al.
Veröffentlicht: (2023)
Improving Spoken Language Modeling with Phoneme Classification: A Simple Fine-tuning Approach
von: Poli, Maxime, et al.
Veröffentlicht: (2024)
von: Poli, Maxime, et al.
Veröffentlicht: (2024)
Spoken to Spoken vs. Spoken to Written: Corpus Approach to Exploring Interpreting and Subtitling
von: Mikhail Mikhailov
Veröffentlicht: (2010)
von: Mikhail Mikhailov
Veröffentlicht: (2010)
Simple and Controllable Music Generation
von: Copet, Jade, et al.
Veröffentlicht: (2023)
von: Copet, Jade, et al.
Veröffentlicht: (2023)
Discrete Flow Matching
von: Gat, Itai, et al.
Veröffentlicht: (2024)
von: Gat, Itai, et al.
Veröffentlicht: (2024)
VUGEN: Visual Understanding priors for GENeration
von: Chen, Xiangyi, et al.
Veröffentlicht: (2025)
von: Chen, Xiangyi, et al.
Veröffentlicht: (2025)
Text-Guided Semantic Image Encoder
von: Thirukovalluru, Raghuveer, et al.
Veröffentlicht: (2025)
von: Thirukovalluru, Raghuveer, et al.
Veröffentlicht: (2025)
Chapter C7 Spoken and Written Performatives
von: Durant, Alan, et al.
Veröffentlicht: (2021)
von: Durant, Alan, et al.
Veröffentlicht: (2021)
LongTail-Swap: benchmarking language models' abilities on rare words
von: Algayres, Robin, et al.
Veröffentlicht: (2025)
von: Algayres, Robin, et al.
Veröffentlicht: (2025)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
CHECK-MAT: Checking Hand-Written Mathematical Answers for the Russian Unified State Exam
von: Khrulev, Ruslan
Veröffentlicht: (2025)
von: Khrulev, Ruslan
Veröffentlicht: (2025)
Written Term Detection Improves Spoken Term Detection
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
von: Yusuf, Bolaji, et al.
Veröffentlicht: (2024)
Combining the flipped classroom and simulation games in engineering education: a methodological survey
von: Algayres, Muriel, et al.
Veröffentlicht: (2019)
von: Algayres, Muriel, et al.
Veröffentlicht: (2019)
Computational Modeling of the Segmentation of Sentence Stimuli From an Infant Word‐Finding Study
von: Daniel Swingley, et al.
Veröffentlicht: (2024)
von: Daniel Swingley, et al.
Veröffentlicht: (2024)
Masked Audio Generation using a Single Non-Autoregressive Transformer
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
von: Ziv, Alon, et al.
Veröffentlicht: (2024)
Corrector Sampling in Language Models
von: Gat, Itai, et al.
Veröffentlicht: (2025)
von: Gat, Itai, et al.
Veröffentlicht: (2025)
Transition Matching: Scalable and Flexible Generative Modeling
von: Shaul, Neta, et al.
Veröffentlicht: (2025)
von: Shaul, Neta, et al.
Veröffentlicht: (2025)
Set Block Decoding is a Language Model Inference Accelerator
von: Gat, Itai, et al.
Veröffentlicht: (2025)
von: Gat, Itai, et al.
Veröffentlicht: (2025)
2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
Lead Zirconate Titanate Reservoir Computing for Classification of Written and Spoken Digits
von: Buckley, Thomas, et al.
Veröffentlicht: (2026)
von: Buckley, Thomas, et al.
Veröffentlicht: (2026)
Joint Audio and Symbolic Conditioning for Temporally Controlled Text-to-Music Generation
von: Tal, Or, et al.
Veröffentlicht: (2024)
von: Tal, Or, et al.
Veröffentlicht: (2024)
Edit Flows: Flow Matching with Edit Operations
von: Havasi, Marton, et al.
Veröffentlicht: (2025)
von: Havasi, Marton, et al.
Veröffentlicht: (2025)
Y-NQ: English-Yorùbá Evaluation dataset for Open-Book Reading Comprehension and Text Generation
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
von: Costa-jussà, Marta R., et al.
Veröffentlicht: (2024)
Linguini: A benchmark for language-agnostic linguistic reasoning
von: Sánchez, Eduardo, et al.
Veröffentlicht: (2024)
von: Sánchez, Eduardo, et al.
Veröffentlicht: (2024)
A French Version of the OLDI Seed Corpus
von: Marmonier, Malik, et al.
Veröffentlicht: (2025)
von: Marmonier, Malik, et al.
Veröffentlicht: (2025)
Tree of Problems: Improving structured problem solving with compositionality
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
Testing the Deliteralization Hypothesis in Human and Machine Translation
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
von: Marmonier, Malik, et al.
Veröffentlicht: (2026)
In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
von: Zebaze, Armel, et al.
Veröffentlicht: (2024)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
von: Antoun, Wissam, et al.
Veröffentlicht: (2025)
von: Antoun, Wissam, et al.
Veröffentlicht: (2025)
LLM Reasoning for Machine Translation: Synthetic Data Generation over Thinking Tokens
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
von: Zebaze, Armel, et al.
Veröffentlicht: (2025)
Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?
von: Riabi, Arij, et al.
Veröffentlicht: (2021)
von: Riabi, Arij, et al.
Veröffentlicht: (2021)
Ähnliche Einträge
-
Investigating Decoder-only Large Language Models for Speech-to-text Translation
von: Huang, Chao-Wei, et al.
Veröffentlicht: (2024) -
Merging Text Transformer Models from Different Initializations
von: Verma, Neha, et al.
Veröffentlicht: (2024) -
BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity?
von: Chambon, Pierre, et al.
Veröffentlicht: (2025) -
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
von: Poli, Maxime, et al.
Veröffentlicht: (2025) -
Unified Vision-Language Modeling via Concept Space Alignment
von: Qiu, Yifu, et al.
Veröffentlicht: (2026)