Typhoon ASR Real-time: FastConformer-Transducer for Thai Automatic Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Sirichotedumrong, Warit, Na-Thalang, Adisai, Manakul, Potsawee, Taveekitworachai, Pittawat, Sripaisarnmongkol, Sittipong, Pipatanakul, Kunat |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models
di: Pipatanakul, Kunat, et al.
Pubblicazione: (2024)
di: Pipatanakul, Kunat, et al.
Pubblicazione: (2024)
Typhoon T1: An Open Thai Reasoning Model
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
di: Chaichana, Yuatyong, et al.
Pubblicazione: (2025)
di: Chaichana, Yuatyong, et al.
Pubblicazione: (2025)
FinCoT: Grounding Chain-of-Thought in Expert Financial Reasoning
di: Nitarach, Natapong, et al.
Pubblicazione: (2025)
di: Nitarach, Natapong, et al.
Pubblicazione: (2025)
Typhoon-S: Minimal Open Post-Training for Sovereign Large Language Models
di: Pipatanakul, Kunat, et al.
Pubblicazione: (2026)
di: Pipatanakul, Kunat, et al.
Pubblicazione: (2026)
Prior Prompt Engineering for Reinforcement Fine-Tuning
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging -- An Open Recipe
di: Pipatanakul, Kunat, et al.
Pubblicazione: (2025)
di: Pipatanakul, Kunat, et al.
Pubblicazione: (2025)
ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
di: Nonesung, Surapon, et al.
Pubblicazione: (2025)
di: Nonesung, Surapon, et al.
Pubblicazione: (2025)
Enhancing Low-Resource Language and Instruction Following Capabilities of Audio Language Models
di: Manakul, Potsawee, et al.
Pubblicazione: (2024)
di: Manakul, Potsawee, et al.
Pubblicazione: (2024)
Typhoon OCR: Open Vision-Language Model For Thai Document Extraction
di: Nonesung, Surapon, et al.
Pubblicazione: (2026)
di: Nonesung, Surapon, et al.
Pubblicazione: (2026)
Talk Less, Call Right: Enhancing Role-Play LLM Agents with Automatic Prompt Optimization and Role Prompting
di: Ruangtanusak, Saksorn, et al.
Pubblicazione: (2025)
di: Ruangtanusak, Saksorn, et al.
Pubblicazione: (2025)
Formula-One Prompting: A Composable Equation-First Prefix for Applied Mathematics
di: Nitarach, Natapong, et al.
Pubblicazione: (2026)
di: Nitarach, Natapong, et al.
Pubblicazione: (2026)
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation
di: Manakul, Potsawee, et al.
Pubblicazione: (2025)
di: Manakul, Potsawee, et al.
Pubblicazione: (2025)
On the Robustness of Answer Formats in Medical Reasoning Models
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
CrossCheckGPT: Universal Hallucination Ranking for Multimodal Foundation Models
di: Sun, Guangzhi, et al.
Pubblicazione: (2024)
di: Sun, Guangzhi, et al.
Pubblicazione: (2024)
Developing an Open Conversational Speech Corpus for the Isan Language
di: Na-Thalang, Adisai, et al.
Pubblicazione: (2025)
di: Na-Thalang, Adisai, et al.
Pubblicazione: (2025)
ThaiSafetyBench: Assessing Language Model Safety in Thai Cultural Contexts
di: Ukarapol, Trapoom, et al.
Pubblicazione: (2026)
di: Ukarapol, Trapoom, et al.
Pubblicazione: (2026)
Mind the Gap! Static and Interactive Evaluations of Large Audio Models
di: Li, Minzhi, et al.
Pubblicazione: (2025)
di: Li, Minzhi, et al.
Pubblicazione: (2025)
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
Large Language Models are Null-Shot Learners
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2024)
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2024)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2025)
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2025)
Promptformer: Prompted Conformer Transducer for ASR
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
LLM Comparative Assessment: Zero-shot NLG Evaluation through Pairwise Comparisons using Large Language Models
di: Liusie, Adian, et al.
Pubblicazione: (2023)
di: Liusie, Adian, et al.
Pubblicazione: (2023)
Unlearning vs. Obfuscation: Are We Truly Removing Knowledge?
di: Sun, Guangzhi, et al.
Pubblicazione: (2025)
di: Sun, Guangzhi, et al.
Pubblicazione: (2025)
WST: Weakly Supervised Transducer for Automatic Speech Recognition
di: Gao, Dongji, et al.
Pubblicazione: (2025)
di: Gao, Dongji, et al.
Pubblicazione: (2025)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2024)
di: Xu, Hainan, et al.
Pubblicazione: (2024)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
di: Cui, Mingyu, et al.
Pubblicazione: (2025)
di: Cui, Mingyu, et al.
Pubblicazione: (2025)
HAINAN: Fast and Accurate Transducer for Hybrid-Autoregressive ASR
di: Xu, Hainan, et al.
Pubblicazione: (2024)
di: Xu, Hainan, et al.
Pubblicazione: (2024)
SkillAggregation: Reference-free LLM-Dependent Aggregation
di: Sun, Guangzhi, et al.
Pubblicazione: (2024)
di: Sun, Guangzhi, et al.
Pubblicazione: (2024)
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
di: Zhang, Tian-Hao, et al.
Pubblicazione: (2023)
di: Zhang, Tian-Hao, et al.
Pubblicazione: (2023)
Granular feedback merits sophisticated aggregation
di: Kagrecha, Anmol, et al.
Pubblicazione: (2025)
di: Kagrecha, Anmol, et al.
Pubblicazione: (2025)
Multi-blank Transducers for Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2022)
di: Xu, Hainan, et al.
Pubblicazione: (2022)
Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
di: Deng, Keqi, et al.
Pubblicazione: (2024)
di: Deng, Keqi, et al.
Pubblicazione: (2024)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
LiteASR: Efficient Automatic Speech Recognition with Low-Rank Approximation
di: Kamahori, Keisuke, et al.
Pubblicazione: (2025)
di: Kamahori, Keisuke, et al.
Pubblicazione: (2025)
FairASR: Fair Audio Contrastive Learning for Automatic Speech Recognition
di: Kim, Jongsuk, et al.
Pubblicazione: (2025)
di: Kim, Jongsuk, et al.
Pubblicazione: (2025)
AS-ASR: A Lightweight Framework for Aphasia-Specific Automatic Speech Recognition
di: Bao, Chen, et al.
Pubblicazione: (2025)
di: Bao, Chen, et al.
Pubblicazione: (2025)
Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
di: Dhawan, Kunal, et al.
Pubblicazione: (2024)
Automatic Speech Recognition (ASR) for the Diagnosis of pronunciation of Speech Sound Disorders in Korean children
di: Ahn, Taekyung, et al.
Pubblicazione: (2024)
di: Ahn, Taekyung, et al.
Pubblicazione: (2024)
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
di: Feng, Chen, et al.
Pubblicazione: (2025)
di: Feng, Chen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models
di: Pipatanakul, Kunat, et al.
Pubblicazione: (2024) -
Typhoon T1: An Open Thai Reasoning Model
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025) -
Extending Audio Context for Long-Form Understanding in Large Audio-Language Models
di: Chaichana, Yuatyong, et al.
Pubblicazione: (2025) -
FinCoT: Grounding Chain-of-Thought in Expert Financial Reasoning
di: Nitarach, Natapong, et al.
Pubblicazione: (2025) -
Typhoon-S: Minimal Open Post-Training for Sovereign Large Language Models
di: Pipatanakul, Kunat, et al.
Pubblicazione: (2026)