Adapting Text LLMs to Speech via Multimodal Depth Up-Scaling
Fuente:
arXiv
Guardado en:
| Autores principales: | Yano, Kazuki, Suzuki, Jun, Watanabe, Shinji |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
STEP: Staged Parameter-Efficient Pre-training for Large Language Models
por: Yano, Kazuki, et al.
Publicado: (2025)
por: Yano, Kazuki, et al.
Publicado: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
por: Futami, Hayato, et al.
Publicado: (2025)
por: Futami, Hayato, et al.
Publicado: (2025)
TimeMachine-bench: A Benchmark for Evaluating Model Capabilities in Repository-Level Migration Tasks
por: Fujii, Ryo, et al.
Publicado: (2026)
por: Fujii, Ryo, et al.
Publicado: (2026)
Efficient Construction of Model Family through Progressive Training Using Model Expansion
por: Yano, Kazuki, et al.
Publicado: (2025)
por: Yano, Kazuki, et al.
Publicado: (2025)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
por: Cornell, Samuele, et al.
Publicado: (2024)
por: Cornell, Samuele, et al.
Publicado: (2024)
Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning
por: Yano, Kazuki, et al.
Publicado: (2026)
por: Yano, Kazuki, et al.
Publicado: (2026)
FastAdaSP: Multitask-Adapted Efficient Inference for Large Speech Language Model
por: Lu, Yichen, et al.
Publicado: (2024)
por: Lu, Yichen, et al.
Publicado: (2024)
Suppressing Final Layer Hidden State Jumps in Transformer Pretraining
por: Shibata, Keigo, et al.
Publicado: (2026)
por: Shibata, Keigo, et al.
Publicado: (2026)
Layerwise Importance Analysis of Feed-Forward Networks in Transformer-based Language Models
por: Ikeda, Wataru, et al.
Publicado: (2025)
por: Ikeda, Wataru, et al.
Publicado: (2025)
On the Effects of Heterogeneous Data Sources on Speech-to-Text Foundation Models
por: Tian, Jinchuan, et al.
Publicado: (2024)
por: Tian, Jinchuan, et al.
Publicado: (2024)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
Wav2Gloss: Generating Interlinear Glossed Text from Speech
por: He, Taiqi, et al.
Publicado: (2024)
por: He, Taiqi, et al.
Publicado: (2024)
Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting?
por: Pareras, Oriol, et al.
Publicado: (2025)
por: Pareras, Oriol, et al.
Publicado: (2025)
OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning
por: Peng, Yifan, et al.
Publicado: (2025)
por: Peng, Yifan, et al.
Publicado: (2025)
Chitranuvad: Adapting Multi-Lingual LLMs for Multimodal Translation
por: Khan, Shaharukh, et al.
Publicado: (2025)
por: Khan, Shaharukh, et al.
Publicado: (2025)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
IterKey: Iterative Keyword Generation with LLMs for Enhanced Retrieval Augmented Generation
por: Hayashi, Kazuki, et al.
Publicado: (2025)
por: Hayashi, Kazuki, et al.
Publicado: (2025)
Progressive Depth Up-scaling via Optimal Transport
por: Cao, Mingzi, et al.
Publicado: (2025)
por: Cao, Mingzi, et al.
Publicado: (2025)
Building Corpora for Single-Channel Speech Separation Across Multiple Domains
por: Maciejewski, Matthew, et al.
Publicado: (2018)
por: Maciejewski, Matthew, et al.
Publicado: (2018)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
por: Nagpal, Chirag, et al.
Publicado: (2024)
por: Nagpal, Chirag, et al.
Publicado: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
Can LLMs Understand Unvoiced Speech? Exploring EMG-to-Text Conversion with LLMs
por: Mohapatra, Payal, et al.
Publicado: (2025)
por: Mohapatra, Payal, et al.
Publicado: (2025)
OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models
por: Chen, William, et al.
Publicado: (2025)
por: Chen, William, et al.
Publicado: (2025)
Do LLMs Implicitly Determine the Suitable Text Difficulty for Users?
por: Gobara, Seiji, et al.
Publicado: (2024)
por: Gobara, Seiji, et al.
Publicado: (2024)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
por: Tsunoo, Emiru, et al.
Publicado: (2024)
por: Tsunoo, Emiru, et al.
Publicado: (2024)
SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
por: Wang, Pu, et al.
Publicado: (2026)
por: Wang, Pu, et al.
Publicado: (2026)
UMETTS: A Unified Framework for Emotional Text-to-Speech Synthesis with Multimodal Prompts
por: Cheng, Zhi-Qi, et al.
Publicado: (2024)
por: Cheng, Zhi-Qi, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
ALAS: Measuring Latent Speech-Text Alignment For Spoken Language Understanding In Multimodal LLMs
por: Mousavi, Pooneh, et al.
Publicado: (2025)
por: Mousavi, Pooneh, et al.
Publicado: (2025)
Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion
por: Du, Yexing, et al.
Publicado: (2026)
por: Du, Yexing, et al.
Publicado: (2026)
Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
por: Yoo, HaeJun, et al.
Publicado: (2026)
por: Yoo, HaeJun, et al.
Publicado: (2026)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
por: Wang, Qingzheng, et al.
Publicado: (2025)
por: Wang, Qingzheng, et al.
Publicado: (2025)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
por: Kawamura, Kazuki, et al.
Publicado: (2024)
por: Kawamura, Kazuki, et al.
Publicado: (2024)
Multimodal LLMs are not all you need for Pediatric Speech Language Pathology
por: Fürst, Darren, et al.
Publicado: (2026)
por: Fürst, Darren, et al.
Publicado: (2026)
SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data
por: Lu, Yichen, et al.
Publicado: (2024)
por: Lu, Yichen, et al.
Publicado: (2024)
Scaling Up RL: Unlocking Diverse Reasoning in LLMs via Prolonged Training
por: Liu, Mingjie, et al.
Publicado: (2025)
por: Liu, Mingjie, et al.
Publicado: (2025)
Explainable Depression Detection using Masked Hard Instance Mining
por: Prakrankamanant, Patawee, et al.
Publicado: (2025)
por: Prakrankamanant, Patawee, et al.
Publicado: (2025)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
Adapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource Languages
por: Xiao, Yang, et al.
Publicado: (2026)
por: Xiao, Yang, et al.
Publicado: (2026)
Ejemplares similares
-
STEP: Staged Parameter-Efficient Pre-training for Large Language Models
por: Yano, Kazuki, et al.
Publicado: (2025) -
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
por: Futami, Hayato, et al.
Publicado: (2025) -
TimeMachine-bench: A Benchmark for Evaluating Model Capabilities in Repository-Level Migration Tasks
por: Fujii, Ryo, et al.
Publicado: (2026) -
Efficient Construction of Model Family through Progressive Training Using Model Expansion
por: Yano, Kazuki, et al.
Publicado: (2025) -
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
por: Cornell, Samuele, et al.
Publicado: (2024)