Hierarchical Policy Optimization for Simultaneous Translation of Unbounded Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ouyang, Siqi, Ding, Shuoyang, Hrinchuk, Oleksii, Lavrukhin, Vitaly, Yan, Brian, Ginsburg, Boris, Li, Lei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Anticipating Future with Large Language Model for Simultaneous Machine Translation
von: Ouyang, Siqi, et al.
Veröffentlicht: (2024)
von: Ouyang, Siqi, et al.
Veröffentlicht: (2024)
Chain-of-Thought Prompting for Speech Translation
von: Hu, Ke, et al.
Veröffentlicht: (2024)
von: Hu, Ke, et al.
Veröffentlicht: (2024)
EMMeTT: Efficient Multimodal Machine Translation Training
von: Żelasko, Piotr, et al.
Veröffentlicht: (2024)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2024)
InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model
von: Ouyang, Siqi, et al.
Veröffentlicht: (2025)
von: Ouyang, Siqi, et al.
Veröffentlicht: (2025)
Extending Automatic Machine Translation Evaluation to Book-Length Documents
von: Wang, Kuang-Da, et al.
Veröffentlicht: (2025)
von: Wang, Kuang-Da, et al.
Veröffentlicht: (2025)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
von: Puvvada, Krishna C., et al.
Veröffentlicht: (2024)
Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2024)
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2024)
CMU's IWSLT 2025 Simultaneous Speech Translation System
von: Ouyang, Siqi, et al.
Veröffentlicht: (2025)
von: Ouyang, Siqi, et al.
Veröffentlicht: (2025)
RASST: Fast Cross-modal Retrieval-Augmented Simultaneous Speech Translation
von: Luo, Jiaxuan, et al.
Veröffentlicht: (2026)
von: Luo, Jiaxuan, et al.
Veröffentlicht: (2026)
Open Automatic Speech Recognition Models for Classical and Modern Standard Arabic
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
FASST: Fast LLM-based Simultaneous Speech Translation
von: Ouyang, Siqi, et al.
Veröffentlicht: (2024)
von: Ouyang, Siqi, et al.
Veröffentlicht: (2024)
CMU's IWSLT 2024 Simultaneous Speech Translation System
von: Xu, Xi, et al.
Veröffentlicht: (2024)
von: Xu, Xi, et al.
Veröffentlicht: (2024)
Granary: Speech Recognition and Translation Dataset in 25 European Languages
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2025)
von: Koluguri, Nithin Rao, et al.
Veröffentlicht: (2025)
CA*: Addressing Evaluation Pitfalls in Computation-Aware Latency for Simultaneous Speech Translation
von: Xu, Xi, et al.
Veröffentlicht: (2024)
von: Xu, Xi, et al.
Veröffentlicht: (2024)
TurboBias: Universal ASR Context-Biasing powered by GPU-accelerated Phrase-Boosting Tree
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2025)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2025)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
von: Bataev, Vladimir, et al.
Veröffentlicht: (2023)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2023)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
Label-Looping: Highly Efficient Decoding for Transducers
von: Bataev, Vladimir, et al.
Veröffentlicht: (2024)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2024)
Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages
von: Ayrapetyan, Alexan, et al.
Veröffentlicht: (2025)
von: Ayrapetyan, Alexan, et al.
Veröffentlicht: (2025)
Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2026)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2026)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5
von: Chen, Zhehuai, et al.
Veröffentlicht: (2024)
von: Chen, Zhehuai, et al.
Veröffentlicht: (2024)
A Chat About Boring Problems: Studying GPT-based text normalization
von: Zhang, Yang, et al.
Veröffentlicht: (2023)
von: Zhang, Yang, et al.
Veröffentlicht: (2023)
NGPU-LM: GPU-Accelerated N-Gram Language Model for Context-Biasing in Greedy ASR Decoding
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
Pushing the Limits of Beam Search Decoding for Transducer-based ASR models
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
Training and Inference Efficiency of Encoder-Decoder Speech Models
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
von: Żelasko, Piotr, et al.
Veröffentlicht: (2025)
Translation Canvas: An Explainable Interface to Pinpoint and Analyze Translation Systems
von: Dandekar, Chinmay, et al.
Veröffentlicht: (2024)
von: Dandekar, Chinmay, et al.
Veröffentlicht: (2024)
Optimizing Rare Word Accuracy in Direct Speech Translation with a Retrieval-and-Demonstration Approach
von: Li, Siqi, et al.
Veröffentlicht: (2024)
von: Li, Siqi, et al.
Veröffentlicht: (2024)
Mending the Holes: Mitigating Reward Hacking in Reinforcement Learning for Multilingual Translation
von: Liu, Yifeng, et al.
Veröffentlicht: (2026)
von: Liu, Yifeng, et al.
Veröffentlicht: (2026)
Fine-Tuned Machine Translation Metrics Struggle in Unseen Domains
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
Contrastive Feedback Mechanism for Simultaneous Speech Translation
von: Tan, Haotian, et al.
Veröffentlicht: (2024)
von: Tan, Haotian, et al.
Veröffentlicht: (2024)
SimulTron: On-Device Simultaneous Speech to Speech Translation
von: Agranovich, Alex, et al.
Veröffentlicht: (2024)
von: Agranovich, Alex, et al.
Veröffentlicht: (2024)
SimulU: Training-free Policy for Long-form Simultaneous Speech-to-Speech Translation
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
von: Djanibekov, Amirbek, et al.
Veröffentlicht: (2026)
Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture
von: Fu, Biao, et al.
Veröffentlicht: (2025)
von: Fu, Biao, et al.
Veröffentlicht: (2025)
SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
von: Le, Chenyang, et al.
Veröffentlicht: (2025)
von: Le, Chenyang, et al.
Veröffentlicht: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
Simultaneous Speech-to-Speech Translation Without Aligned Data
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
Continuous Rating as Reliable Human Evaluation of Simultaneous Speech Translation
von: Javorský, Dávid, et al.
Veröffentlicht: (2022)
von: Javorský, Dávid, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
Anticipating Future with Large Language Model for Simultaneous Machine Translation
von: Ouyang, Siqi, et al.
Veröffentlicht: (2024) -
Chain-of-Thought Prompting for Speech Translation
von: Hu, Ke, et al.
Veröffentlicht: (2024) -
EMMeTT: Efficient Multimodal Machine Translation Training
von: Żelasko, Piotr, et al.
Veröffentlicht: (2024) -
InfiniSST: Simultaneous Translation of Unbounded Speech with Large Language Model
von: Ouyang, Siqi, et al.
Veröffentlicht: (2025) -
Extending Automatic Machine Translation Evaluation to Book-Length Documents
von: Wang, Kuang-Da, et al.
Veröffentlicht: (2025)