Chain-of-Thought Prompting for Speech Translation
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Ke, Chen, Zhehuai, Yang, Chao-Han Huck, Żelasko, Piotr, Hrinchuk, Oleksii, Lavrukhin, Vitaly, Balam, Jagadeesh, Ginsburg, Boris |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Anticipating Future with Large Language Model for Simultaneous Machine Translation
di: Ouyang, Siqi, et al.
Pubblicazione: (2024)
di: Ouyang, Siqi, et al.
Pubblicazione: (2024)
EMMeTT: Efficient Multimodal Machine Translation Training
di: Żelasko, Piotr, et al.
Pubblicazione: (2024)
di: Żelasko, Piotr, et al.
Pubblicazione: (2024)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
di: Puvvada, Krishna C., et al.
Pubblicazione: (2024)
di: Puvvada, Krishna C., et al.
Pubblicazione: (2024)
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
di: Chen, Zhehuai, et al.
Pubblicazione: (2024)
di: Chen, Zhehuai, et al.
Pubblicazione: (2024)
Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech
di: Ouyang, Siqi, et al.
Pubblicazione: (2026)
di: Ouyang, Siqi, et al.
Pubblicazione: (2026)
Training and Inference Efficiency of Encoder-Decoder Speech Models
di: Żelasko, Piotr, et al.
Pubblicazione: (2025)
di: Żelasko, Piotr, et al.
Pubblicazione: (2025)
Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2024)
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2024)
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
di: Hu, Ke, et al.
Pubblicazione: (2025)
di: Hu, Ke, et al.
Pubblicazione: (2025)
VoiceTextBlender: Augmenting Large Language Models with Speech Capabilities via Single-Stage Joint Speech-Text Supervised Fine-Tuning
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
di: Hu, Ke, et al.
Pubblicazione: (2025)
di: Hu, Ke, et al.
Pubblicazione: (2025)
Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
di: Noroozi, Vahid, et al.
Pubblicazione: (2024)
di: Noroozi, Vahid, et al.
Pubblicazione: (2024)
Extending Automatic Machine Translation Evaluation to Book-Length Documents
di: Wang, Kuang-Da, et al.
Pubblicazione: (2025)
di: Wang, Kuang-Da, et al.
Pubblicazione: (2025)
Canary-1B-v2 & Parakeet-TDT-0.6B-v3: Efficient and High-Performance Models for Multilingual ASR and AST
di: Sekoyan, Monica, et al.
Pubblicazione: (2025)
di: Sekoyan, Monica, et al.
Pubblicazione: (2025)
Stateful Conformer with Cache-based Inference for Streaming Automatic Speech Recognition
di: Noroozi, Vahid, et al.
Pubblicazione: (2023)
di: Noroozi, Vahid, et al.
Pubblicazione: (2023)
Granary: Speech Recognition and Translation Dataset in 25 European Languages
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2025)
di: Koluguri, Nithin Rao, et al.
Pubblicazione: (2025)
Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
NeKo: Cross-Modality Post-Recognition Error Correction with Tasks-Guided Mixture-of-Experts Language Model
di: Lin, Yen-Ting, et al.
Pubblicazione: (2024)
di: Lin, Yen-Ting, et al.
Pubblicazione: (2024)
GenTranslate: Large Language Models are Generative Multilingual Speech and Machine Translators
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2024)
di: Xu, Hainan, et al.
Pubblicazione: (2024)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
Large Language Model Based Generative Error Correction: A Challenge and Baselines for Speech Recognition, Speaker Tagging, and Emotion Recognition
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2024)
di: Yang, Chao-Han Huck, et al.
Pubblicazione: (2024)
TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
di: Andrusenko, Andrei, et al.
Pubblicazione: (2025)
di: Andrusenko, Andrei, et al.
Pubblicazione: (2025)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
di: Bataev, Vladimir, et al.
Pubblicazione: (2023)
di: Bataev, Vladimir, et al.
Pubblicazione: (2023)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
di: Andrusenko, Andrei, et al.
Pubblicazione: (2024)
di: Andrusenko, Andrei, et al.
Pubblicazione: (2024)
Label-Looping: Highly Efficient Decoding for Transducers
di: Bataev, Vladimir, et al.
Pubblicazione: (2024)
di: Bataev, Vladimir, et al.
Pubblicazione: (2024)
Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages
di: Ayrapetyan, Alexan, et al.
Pubblicazione: (2025)
di: Ayrapetyan, Alexan, et al.
Pubblicazione: (2025)
Flexible Multichannel Speech Enhancement for Noise-Robust Frontend
di: Jukić, Ante, et al.
Pubblicazione: (2024)
di: Jukić, Ante, et al.
Pubblicazione: (2024)
A Chat About Boring Problems: Studying GPT-based text normalization
di: Zhang, Yang, et al.
Pubblicazione: (2023)
di: Zhang, Yang, et al.
Pubblicazione: (2023)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
di: Lu, Ke-Han, et al.
Pubblicazione: (2024)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
di: Andrusenko, Andrei, et al.
Pubblicazione: (2026)
di: Andrusenko, Andrei, et al.
Pubblicazione: (2026)
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
di: Chen, Chen, et al.
Pubblicazione: (2025)
di: Chen, Chen, et al.
Pubblicazione: (2025)
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
di: Grigoryan, Lilit, et al.
Pubblicazione: (2025)
SpeechIQ: Speech-Agentic Intelligence Quotient Across Cognitive Levels in Voice Understanding by Large Language Models
di: Wan, Zhen, et al.
Pubblicazione: (2025)
di: Wan, Zhen, et al.
Pubblicazione: (2025)
Schrödinger Bridge for Generative Speech Enhancement
di: Jukić, Ante, et al.
Pubblicazione: (2024)
di: Jukić, Ante, et al.
Pubblicazione: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
di: Bataev, Vladimir, et al.
Pubblicazione: (2025)
Sortformer: A Novel Approach for Permutation-Resolved Speaker Supervision in Speech-to-Text Systems
di: Park, Taejin, et al.
Pubblicazione: (2024)
di: Park, Taejin, et al.
Pubblicazione: (2024)
Listening or Reading? Evaluating Speech Awareness in Chain-of-Thought Speech-to-Text Translation
di: Romero-Díaz, Jacobo, et al.
Pubblicazione: (2025)
di: Romero-Díaz, Jacobo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Anticipating Future with Large Language Model for Simultaneous Machine Translation
di: Ouyang, Siqi, et al.
Pubblicazione: (2024) -
EMMeTT: Efficient Multimodal Machine Translation Training
di: Żelasko, Piotr, et al.
Pubblicazione: (2024) -
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
di: Puvvada, Krishna C., et al.
Pubblicazione: (2024) -
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
di: Chen, Zhehuai, et al.
Pubblicazione: (2024) -
Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech
di: Ouyang, Siqi, et al.
Pubblicazione: (2026)