Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Chih-Kai, Fu, Yu-Kuan, Li, Chen-An, Lin, Yi-Cheng, Lin, Yu-Xiang, Chen, Wei-Chih, Chung, Ho Lam, Kuan, Chun-Yi, Huang, Wei-Ping, Lu, Ke-Han, Lin, Tzu-Quan, Wang, Hsiu-Hsuan, Hu, En-Pei, Hsu, Chan-Jan, Tseng, Liang-Hsuan, Chiu, I-Hsiang, Sanga, Ulin, Chen, Xuanjun, Hsu, Po-chun, Yang, Shu-wen, Lee, Hung-yi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Age Aware Scheduling for Differentially-Private Federated Learning
by: Lin, Kuan-Yu, et al.
Published: (2024)
by: Lin, Kuan-Yu, et al.
Published: (2024)
Impact of violence on work morale on Taiwanese nurses: The moderator of perceived organizational support
by: Kuan‐Yang Chen, et al.
Published: (2025)
by: Kuan‐Yang Chen, et al.
Published: (2025)
BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights
by: Hsu, Chan-Jan, et al.
Published: (2025)
by: Hsu, Chan-Jan, et al.
Published: (2025)
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
by: Tseng, Liang-Hsuan, et al.
Published: (2025)
by: Tseng, Liang-Hsuan, et al.
Published: (2025)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
by: Lin, Yi-Cheng, et al.
Published: (2024)
by: Lin, Yi-Cheng, et al.
Published: (2024)
Weakly Supervised 3D Object Detection via Multi-Level Visual Guidance
by: Huang, Kuan-Chih, et al.
Published: (2023)
by: Huang, Kuan-Chih, et al.
Published: (2023)
FROAV: A Framework for RAG Observation and Agent Verification -- Lowering the Barrier to LLM Agent Research
by: Lin, Tzu-Hsuan, et al.
Published: (2026)
by: Lin, Tzu-Hsuan, et al.
Published: (2026)
Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
by: Lin, Yi-Cheng, et al.
Published: (2024)
by: Lin, Yi-Cheng, et al.
Published: (2024)
Measuring Taiwanese Mandarin Language Understanding
by: Chen, Po-Heng, et al.
Published: (2024)
by: Chen, Po-Heng, et al.
Published: (2024)
Investigating the Effects of Large-Scale Pseudo-Stereo Data and Different Speech Foundation Model on Dialogue Generative Spoken Language Model
by: Fu, Yu-Kuan, et al.
Published: (2024)
by: Fu, Yu-Kuan, et al.
Published: (2024)
Timely Information Updating for Mobile Devices Without and With ML Advice
by: Hsu, Yu-Pin, et al.
Published: (2025)
by: Hsu, Yu-Pin, et al.
Published: (2025)
PTT: Point-Trajectory Transformer for Efficient Temporal 3D Object Detection
by: Huang, Kuan-Chih, et al.
Published: (2023)
by: Huang, Kuan-Chih, et al.
Published: (2023)
AnchorSteer: Self-Discovered Concept Injection for Structure-Preserving Music Editing
by: Chang, Chih-Heng, et al.
Published: (2026)
by: Chang, Chih-Heng, et al.
Published: (2026)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
by: Huang, Kuan-Po, et al.
Published: (2023)
by: Huang, Kuan-Po, et al.
Published: (2023)
Game-Time: Evaluating Temporal Dynamics in Spoken Language Models
by: Chang, Kai-Wei, et al.
Published: (2025)
by: Chang, Kai-Wei, et al.
Published: (2025)
Randomized Scheduling for Periodic Multi-Source Systems with PAoI Violation Guarantees
by: Lin, Kuan-Yu, et al.
Published: (2025)
by: Lin, Kuan-Yu, et al.
Published: (2025)
A Preliminary Exploration with GPT-4o Voice Mode
by: Lin, Yu-Xiang, et al.
Published: (2025)
by: Lin, Yu-Xiang, et al.
Published: (2025)
Ranking-aware adapter for text-driven image ordering with CLIP
by: Yu, Wei-Hsiang, et al.
Published: (2024)
by: Yu, Wei-Hsiang, et al.
Published: (2024)
Self‐Transforming Bioadhesive Patch with Topological and Ionic Crosslinking for Adaptive Mucoadhesion and Enhanced Wound Healing
by: Shih‐Yung Liao, et al.
Published: (2026)
by: Shih‐Yung Liao, et al.
Published: (2026)
Training-Efficient Text-to-Music Generation with State-Space Modeling
by: Lee, Wei-Jaw, et al.
Published: (2026)
by: Lee, Wei-Jaw, et al.
Published: (2026)
Exploring State-Space-Model based Language Model in Music Generation
by: Lee, Wei-Jaw, et al.
Published: (2025)
by: Lee, Wei-Jaw, et al.
Published: (2025)
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
by: Hsu, Chan-Jan, et al.
Published: (2026)
by: Hsu, Chan-Jan, et al.
Published: (2026)
Recent Advances in Soil Stabilization and Reinforcement: A Comprehensive Review of Emerging Technologies
by: Ching Hung, et al.
Published: (2026)
by: Ching Hung, et al.
Published: (2026)
Factors influencing students' listening learning performance in mobile vocabulary‐assisted listening learning: An extended technology acceptance model
by: Hui‐Tzu Hsu, et al.
Published: (2024)
by: Hui‐Tzu Hsu, et al.
Published: (2024)
TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
by: Tseng, Liang-Hsuan, et al.
Published: (2026)
by: Tseng, Liang-Hsuan, et al.
Published: (2026)
Investigating Zero-Shot Generalizability on Mandarin-English Code-Switched ASR and Speech-to-text Translation of Recent Foundation Models with Self-Supervision and Weak Supervision
by: Yang, Chih-Kai, et al.
Published: (2023)
by: Yang, Chih-Kai, et al.
Published: (2023)
Conditional Semi-Supervised Data Augmentation for Spam Message Detection with Low Resource Data
by: Nuha, Ulin, et al.
Published: (2024)
by: Nuha, Ulin, et al.
Published: (2024)
Curriculum leadership in a rural indigenous high school in Taiwan implementing the 108 Curriculum Guidelines
by: Kuan‐Pei Lin, et al.
Published: (2024)
by: Kuan‐Pei Lin, et al.
Published: (2024)
Néel Tensor Torque in Polycrystalline Antiferromagnets (Adv. Mater. 9/2026)
by: Chao‐Yao Yang, et al.
Published: (2026)
by: Chao‐Yao Yang, et al.
Published: (2026)
Néel Tensor Torque in Polycrystalline Antiferromagnets
by: Chao‐Yao Yang, et al.
Published: (2025)
by: Chao‐Yao Yang, et al.
Published: (2025)
TiCo: Time-Controllable Spoken Dialogue Model
by: Chang, Kai-Wei, et al.
Published: (2026)
by: Chang, Kai-Wei, et al.
Published: (2026)
Higher tumor mutational burden is associated with inferior outcomes among pediatric patients with neuroblastoma
by: Ya‐Hsuan Chang, et al.
Published: (2024)
by: Ya‐Hsuan Chang, et al.
Published: (2024)
Precision‐Targeted Injection Laryngoplasty and Longitudinal Biomaterial Effects Evaluation Using High‐resolution Ultrasonography in a Rat Model
by: Wen‐Hsuan Tseng, et al.
Published: (2024)
by: Wen‐Hsuan Tseng, et al.
Published: (2024)
Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-To-Speech Systems
by: Lin, Yi-Cheng, et al.
Published: (2025)
by: Lin, Yi-Cheng, et al.
Published: (2025)
Singer separation for karaoke content generation
by: Lin, Hsuan-Yu, et al.
Published: (2021)
by: Lin, Hsuan-Yu, et al.
Published: (2021)
Pharmacogenomic Calling From Whole‐Exome Sequencing in the Taiwanese Population—A Real‐World Experience
by: Hsu‐Heng Lin, et al.
Published: (2026)
by: Hsu‐Heng Lin, et al.
Published: (2026)
Quantum-Train with Tensor Network Mapping Model and Distributed Circuit Ansatz
by: Liu, Chen-Yu, et al.
Published: (2024)
by: Liu, Chen-Yu, et al.
Published: (2024)
Quantum-Train Long Short-Term Memory: Application on Flood Prediction Problem
by: Lin, Chu-Hsuan Abraham, et al.
Published: (2024)
by: Lin, Chu-Hsuan Abraham, et al.
Published: (2024)
TTC7B Activates the AKT–JKAMP Signaling Axis to Promote Tumor Progression in Head and Neck Cancer
by: Yu‐Hsuan Lin, et al.
Published: (2025)
by: Yu‐Hsuan Lin, et al.
Published: (2025)
Creativity in LLM-based Multi-Agent Systems: A Survey
by: Lin, Yi-Cheng, et al.
Published: (2025)
by: Lin, Yi-Cheng, et al.
Published: (2025)
Similar Items
-
Age Aware Scheduling for Differentially-Private Federated Learning
by: Lin, Kuan-Yu, et al.
Published: (2024) -
Impact of violence on work morale on Taiwanese nurses: The moderator of perceived organizational support
by: Kuan‐Yang Chen, et al.
Published: (2025) -
BreezyVoice: Adapting TTS for Taiwanese Mandarin with Enhanced Polyphone Disambiguation -- Challenges and Insights
by: Hsu, Chan-Jan, et al.
Published: (2025) -
TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
by: Tseng, Liang-Hsuan, et al.
Published: (2025) -
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
by: Lin, Yi-Cheng, et al.
Published: (2024)