Saved in:
| Main Authors: | Mullov, Carlos, Pham, Ngoc-Quan, Waibel, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2408.02290 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weight Factorization and Centralization for Continual Learning in Speech Recognition
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Adapting Language Balance in Code-Switching Speech
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Towards continually learning new languages
by: Pham, Ngoc-Quan, et al.
Published: (2022)
by: Pham, Ngoc-Quan, et al.
Published: (2022)
Cocktail-Party Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)
by: Nguyen, Tuan Nam, et al.
Published: (2024)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
by: Zink, Oswald, et al.
Published: (2024)
by: Zink, Oswald, et al.
Published: (2024)
Zero-Shot Strategies for Length-Controllable Summarization
by: Retkowski, Fabian, et al.
Published: (2024)
by: Retkowski, Fabian, et al.
Published: (2024)
End-to-End Evaluation for Low-Latency Simultaneous Speech Translation
by: Huber, Christian, et al.
Published: (2023)
by: Huber, Christian, et al.
Published: (2023)
Bayesian Low-Rank Factorization for Robust Model Adaptation
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Blending LLMs into Cascaded Speech Translation: KIT's Offline Speech Translation System for IWSLT 2024
by: Koneru, Sai, et al.
Published: (2024)
by: Koneru, Sai, et al.
Published: (2024)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
PIER: A Novel Metric for Evaluating What Matters in Code-Switching
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
KIT's Low-resource Speech Translation Systems for IWSLT2025: System Enhancement with Synthetic Data and Model Regularization
by: Li, Zhaolin, et al.
Published: (2025)
by: Li, Zhaolin, et al.
Published: (2025)
A Cocktail-Party Benchmark: Multi-Modal dataset and Comparative Evaluation Results
by: Nguyen, Thai-Binh, et al.
Published: (2025)
by: Nguyen, Thai-Binh, et al.
Published: (2025)
Continuously Learning New Words in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
From Text Segmentation to Smart Chaptering: A Novel Benchmark for Structuring Video Transcriptions
by: Retkowski, Fabian, et al.
Published: (2024)
by: Retkowski, Fabian, et al.
Published: (2024)
Paragraph Segmentation Revisited: Towards a Standard Task for Structuring Speech
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Towards Zero-Shot, Controllable Dialog Planning with LLMs
by: Väth, Dirk, et al.
Published: (2024)
by: Väth, Dirk, et al.
Published: (2024)
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
by: Akti, Seymanur, et al.
Published: (2026)
by: Akti, Seymanur, et al.
Published: (2026)
CharSpan: Utilizing Lexical Similarity to Enable Zero-Shot Machine Translation for Extremely Low-resource Languages
by: Maurya, Kaushal Kumar, et al.
Published: (2023)
by: Maurya, Kaushal Kumar, et al.
Published: (2023)
Spectral Prompt Tuning:Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
by: Xu, Wenhao, et al.
Published: (2023)
by: Xu, Wenhao, et al.
Published: (2023)
Context Biasing for Pronunciation-Orthography Mismatch in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2025)
by: Huber, Christian, et al.
Published: (2025)
Nsanku: Evaluating Zero-Shot Translation Performance of LLMs for Ghanaian Languages
by: Moore, Stephen E., et al.
Published: (2026)
by: Moore, Stephen E., et al.
Published: (2026)
Machine Translation Models are Zero-Shot Detectors of Translation Direction
by: Wastl, Michelle, et al.
Published: (2024)
by: Wastl, Michelle, et al.
Published: (2024)
A Zero-Shot Open-Vocabulary Pipeline for Dialogue Understanding
by: Safa, Abdulfattah, et al.
Published: (2024)
by: Safa, Abdulfattah, et al.
Published: (2024)
Zero-Shot Hierarchical Classification on the Common Procurement Vocabulary Taxonomy
by: Moiraghi, Federico, et al.
Published: (2024)
by: Moiraghi, Federico, et al.
Published: (2024)
Towards Zero-Shot Multimodal Machine Translation
by: Futeral, Matthieu, et al.
Published: (2024)
by: Futeral, Matthieu, et al.
Published: (2024)
Languages Transferred Within the Encoder: On Representation Transfer in Zero-Shot Multilingual Translation
by: Qu, Zhi, et al.
Published: (2024)
by: Qu, Zhi, et al.
Published: (2024)
Handling Numeric Expressions in Automatic Speech Recognition
by: Huber, Christian, et al.
Published: (2024)
by: Huber, Christian, et al.
Published: (2024)
LLMs Are Zero-Shot Context-Aware Simultaneous Translators
by: Koshkin, Roman, et al.
Published: (2024)
by: Koshkin, Roman, et al.
Published: (2024)
KIT's Offline Speech Translation and Instruction Following Submission for IWSLT 2025
by: Koneru, Sai, et al.
Published: (2025)
by: Koneru, Sai, et al.
Published: (2025)
Zero-Shot Detection of LLM-Generated Text using Token Cohesiveness
by: Ma, Shixuan, et al.
Published: (2024)
by: Ma, Shixuan, et al.
Published: (2024)
LCS: A Language Converter Strategy for Zero-Shot Neural Machine Translation
by: Sun, Zengkui, et al.
Published: (2024)
by: Sun, Zengkui, et al.
Published: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024)
by: Nguyen, Thai-Binh, et al.
Published: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
by: Nguyen, Thai-Binh, et al.
Published: (2023)
by: Nguyen, Thai-Binh, et al.
Published: (2023)
The AI Co-Ethnographer: How Far Can Automation Take Qualitative Research?
by: Retkowski, Fabian, et al.
Published: (2025)
by: Retkowski, Fabian, et al.
Published: (2025)
Understanding and Mitigating the Uncertainty in Zero-Shot Translation
by: Wang, Wenxuan, et al.
Published: (2022)
by: Wang, Wenxuan, et al.
Published: (2022)
SciEx: Benchmarking Large Language Models on Scientific Exams with Human Expert Grading and Automatic Grading
by: Dinh, Tu Anh, et al.
Published: (2024)
by: Dinh, Tu Anh, et al.
Published: (2024)
Leveraging Sentence-oriented Augmentation and Transformer-Based Architecture for Vietnamese-Bahnaric Translation
by: Nguyen, Tan Sang, et al.
Published: (2026)
by: Nguyen, Tan Sang, et al.
Published: (2026)
Tuning LLMs with Contrastive Alignment Instructions for Machine Translation in Unseen, Low-resource Languages
by: Mao, Zhuoyuan, et al.
Published: (2024)
by: Mao, Zhuoyuan, et al.
Published: (2024)
Similar Items
-
Weight Factorization and Centralization for Continual Learning in Speech Recognition
by: Ugan, Enes Yavuz, et al.
Published: (2025) -
Adapting Language Balance in Code-Switching Speech
by: Ugan, Enes Yavuz, et al.
Published: (2025) -
Towards continually learning new languages
by: Pham, Ngoc-Quan, et al.
Published: (2022) -
Cocktail-Party Audio-Visual Speech Recognition
by: Nguyen, Thai-Binh, et al.
Published: (2025) -
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)