Spoken Language Intelligence of Large Language Models for Language Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Peng, Linkai, Nuchged, Baorian, Gao, Yingming |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
di: Chen, Li-Wei, et al.
Pubblicazione: (2025)
di: Chen, Li-Wei, et al.
Pubblicazione: (2025)
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
di: Xue, Jinlong, et al.
Pubblicazione: (2024)
di: Xue, Jinlong, et al.
Pubblicazione: (2024)
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
di: Hsiao, Chi-Yuan, et al.
Pubblicazione: (2025)
di: Hsiao, Chi-Yuan, et al.
Pubblicazione: (2025)
PhonologyBench: Evaluating Phonological Skills of Large Language Models
di: Suvarna, Ashima, et al.
Pubblicazione: (2024)
di: Suvarna, Ashima, et al.
Pubblicazione: (2024)
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
di: Sakai, Yusuke, et al.
Pubblicazione: (2024)
di: Sakai, Yusuke, et al.
Pubblicazione: (2024)
Evaluating and Improving Continual Learning in Spoken Language Understanding
di: Yang, Muqiao, et al.
Pubblicazione: (2024)
di: Yang, Muqiao, et al.
Pubblicazione: (2024)
Towards Signal Processing In Large Language Models
di: Verma, Prateek, et al.
Pubblicazione: (2024)
di: Verma, Prateek, et al.
Pubblicazione: (2024)
Adaptive Large Language Models By Layerwise Attention Shortcuts
di: Verma, Prateek, et al.
Pubblicazione: (2024)
di: Verma, Prateek, et al.
Pubblicazione: (2024)
Large Language Models' Internal Perception of Symbolic Music
di: Shin, Andrew, et al.
Pubblicazione: (2025)
di: Shin, Andrew, et al.
Pubblicazione: (2025)
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
di: Choi, Youngwon, et al.
Pubblicazione: (2025)
di: Choi, Youngwon, et al.
Pubblicazione: (2025)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
Whisper-GPT: A Hybrid Representation Audio Large Language Model
di: Verma, Prateek
Pubblicazione: (2024)
di: Verma, Prateek
Pubblicazione: (2024)
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
di: Wang, Jinhan, et al.
Pubblicazione: (2024)
di: Wang, Jinhan, et al.
Pubblicazione: (2024)
What Do Language Models Hear? Probing for Auditory Representations in Language Models
di: Ngo, Jerry, et al.
Pubblicazione: (2024)
di: Ngo, Jerry, et al.
Pubblicazione: (2024)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
SEAL: Speaker Error Correction using Acoustic-conditioned Large Language Models
di: Kumar, Anurag, et al.
Pubblicazione: (2025)
di: Kumar, Anurag, et al.
Pubblicazione: (2025)
Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2023)
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2023)
C3LLM: Conditional Multimodal Content Generation Using Large Language Models
di: Wang, Zixuan, et al.
Pubblicazione: (2024)
di: Wang, Zixuan, et al.
Pubblicazione: (2024)
GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
di: Chen, Hongjie, et al.
Pubblicazione: (2025)
di: Chen, Hongjie, et al.
Pubblicazione: (2025)
Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2026)
di: Kuan, Chun-Yi, et al.
Pubblicazione: (2026)
TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
di: Li, Zehan, et al.
Pubblicazione: (2025)
di: Li, Zehan, et al.
Pubblicazione: (2025)
Enhancing Neural Spoken Language Recognition: An Exploration with Multilingual Datasets
di: Anidjar, Or Haim, et al.
Pubblicazione: (2025)
di: Anidjar, Or Haim, et al.
Pubblicazione: (2025)
Wavelet GPT: Wavelet Inspired Large Language Models
di: Verma, Prateek
Pubblicazione: (2024)
di: Verma, Prateek
Pubblicazione: (2024)
Listen Again and Choose the Right Answer: A New Paradigm for Automatic Speech Recognition with Large Language Models
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
di: Chang, Heng-Jui, et al.
Pubblicazione: (2024)
di: Chang, Heng-Jui, et al.
Pubblicazione: (2024)
Advancing Speech Understanding in Speech-Aware Language Models with GRPO
di: Elmakies, Avishai, et al.
Pubblicazione: (2025)
di: Elmakies, Avishai, et al.
Pubblicazione: (2025)
BAT: Learning to Reason about Spatial Sounds with Large Language Models
di: Zheng, Zhisheng, et al.
Pubblicazione: (2024)
di: Zheng, Zhisheng, et al.
Pubblicazione: (2024)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2024)
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2024)
Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
di: Xie, Zhifei, et al.
Pubblicazione: (2025)
di: Xie, Zhifei, et al.
Pubblicazione: (2025)
Imagine to Hear: Auditory Knowledge Generation can be an Effective Assistant for Language Models
di: Yoo, Suho, et al.
Pubblicazione: (2025)
di: Yoo, Suho, et al.
Pubblicazione: (2025)
Unsupervised Speech Segmentation: A General Approach Using Speech Language Models
di: Elmakies, Avishai, et al.
Pubblicazione: (2025)
di: Elmakies, Avishai, et al.
Pubblicazione: (2025)
Slamming: Training a Speech Language Model on One GPU in a Day
di: Maimon, Gallil, et al.
Pubblicazione: (2025)
di: Maimon, Gallil, et al.
Pubblicazione: (2025)
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction
di: Wang, Qichao, et al.
Pubblicazione: (2025)
di: Wang, Qichao, et al.
Pubblicazione: (2025)
Decoding Poultry Vocalizations -- Natural Language Processing and Transformer Models for Semantic and Emotional Analysis
di: Manikandan, Venkatraman, et al.
Pubblicazione: (2024)
di: Manikandan, Venkatraman, et al.
Pubblicazione: (2024)
Joint Fine-tuning and Conversion of Pretrained Speech and Language Models towards Linear Complexity
di: He, Mutian, et al.
Pubblicazione: (2024)
di: He, Mutian, et al.
Pubblicazione: (2024)
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
di: Mitsui, Kentaro, et al.
Pubblicazione: (2024)
Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
di: Xie, Tianxin, et al.
Pubblicazione: (2024)
di: Xie, Tianxin, et al.
Pubblicazione: (2024)
RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
di: Xin, Detai, et al.
Pubblicazione: (2024)
di: Xin, Detai, et al.
Pubblicazione: (2024)
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
di: Goel, Arushi, et al.
Pubblicazione: (2025)
di: Goel, Arushi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
di: Chen, Li-Wei, et al.
Pubblicazione: (2025) -
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation
di: Xue, Jinlong, et al.
Pubblicazione: (2024) -
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models
di: Hsiao, Chi-Yuan, et al.
Pubblicazione: (2025) -
PhonologyBench: Evaluating Phonological Skills of Large Language Models
di: Suvarna, Ashima, et al.
Pubblicazione: (2024) -
Simultaneous Interpretation Corpus Construction by Large Language Models in Distant Language Pair
di: Sakai, Yusuke, et al.
Pubblicazione: (2024)