Mining Word Boundaries from Speech-Text Parallel Data for Cross-domain Chinese Word Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Xuebin, Zhang, Lei, Li, Zhenghua, Zhou, Shilin, Gong, Chen, Hou, Yang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Character-Level Chinese Dependency Parsing via Modeling Latent Intra-Word Structure
von: Hou, Yang, et al.
Veröffentlicht: (2024)
von: Hou, Yang, et al.
Veröffentlicht: (2024)
Parsing Through Boundaries in Chinese Word Segmentation
von: Chen, Yige, et al.
Veröffentlicht: (2025)
von: Chen, Yige, et al.
Veröffentlicht: (2025)
JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR
von: Zhou, Shilin, et al.
Veröffentlicht: (2026)
von: Zhou, Shilin, et al.
Veröffentlicht: (2026)
Chinese Word Boundary Recovery through Character Alignment Projection
von: Wang, Lusha, et al.
Veröffentlicht: (2026)
von: Wang, Lusha, et al.
Veröffentlicht: (2026)
Word Segmentation for Asian Languages: Chinese, Korean, and Japanese
von: Rho, Matthew, et al.
Veröffentlicht: (2024)
von: Rho, Matthew, et al.
Veröffentlicht: (2024)
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
von: Sheng, Zhichao, et al.
Veröffentlicht: (2025)
von: Sheng, Zhichao, et al.
Veröffentlicht: (2025)
Improving Contextual ASR via Multi-grained Fusion with Large Language Models
von: Zhou, Shilin, et al.
Veröffentlicht: (2025)
von: Zhou, Shilin, et al.
Veröffentlicht: (2025)
Self-Correction Makes LLMs Better Parsers
von: Zhang, Ziyan, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyan, et al.
Veröffentlicht: (2025)
Mandarin Chinese Words and Parts of Speech
von: Huang, Chu-Ren, et al.
Veröffentlicht: (2021)
von: Huang, Chu-Ren, et al.
Veröffentlicht: (2021)
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
von: Goriely, Zébulon, et al.
Veröffentlicht: (2025)
Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination
von: He, Xiaoqi, et al.
Veröffentlicht: (2026)
von: He, Xiaoqi, et al.
Veröffentlicht: (2026)
Chinese ModernBERT with Whole-Word Masking
von: Zhao, Zeyu, et al.
Veröffentlicht: (2025)
von: Zhao, Zeyu, et al.
Veröffentlicht: (2025)
Using Context to Improve Word Segmentation
von: Hu, Stephanie, et al.
Veröffentlicht: (2025)
von: Hu, Stephanie, et al.
Veröffentlicht: (2025)
Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2025)
von: Pulipaka, Sidharth, et al.
Veröffentlicht: (2025)
Span-Aggregatable, Contextualized Word Embeddings for Effective Phrase Mining
von: Orbach, Eyal, et al.
Veröffentlicht: (2024)
von: Orbach, Eyal, et al.
Veröffentlicht: (2024)
A Comparative Analysis of Word Segmentation, Part-of-Speech Tagging, and Named Entity Recognition for Historical Chinese Sources, 1900-1950
von: Fang, Zhao, et al.
Veröffentlicht: (2025)
von: Fang, Zhao, et al.
Veröffentlicht: (2025)
Cross-Lingual Word Alignment for ASEAN Languages with Contrastive Learning
von: Zhang, Jingshen, et al.
Veröffentlicht: (2024)
von: Zhang, Jingshen, et al.
Veröffentlicht: (2024)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
von: Dai, Xunlian, et al.
Veröffentlicht: (2025)
von: Dai, Xunlian, et al.
Veröffentlicht: (2025)
Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency Parsing
von: Li, Zhuoran, et al.
Veröffentlicht: (2025)
von: Li, Zhuoran, et al.
Veröffentlicht: (2025)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
von: Park, Chanho, et al.
Veröffentlicht: (2023)
von: Park, Chanho, et al.
Veröffentlicht: (2023)
Mixture of Small and Large Models for Chinese Spelling Check
von: Qiao, Ziheng, et al.
Veröffentlicht: (2025)
von: Qiao, Ziheng, et al.
Veröffentlicht: (2025)
From Words to Worlds: Benchmarking Cross-Cultural Cultural Understanding in Machine Translation
von: Han, Bangju, et al.
Veröffentlicht: (2026)
von: Han, Bangju, et al.
Veröffentlicht: (2026)
RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis
von: Wang, Enzhi, et al.
Veröffentlicht: (2025)
von: Wang, Enzhi, et al.
Veröffentlicht: (2025)
From Word2Vec to Transformers: Text-Derived Composition Embeddings for Filtering Combinatorial Electrocatalysts
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
von: Zhang, Lei, et al.
Veröffentlicht: (2026)
SWE2: SubWord Enriched and Significant Word Emphasized Framework for Hate Speech Detection
von: Mou, Guanyi, et al.
Veröffentlicht: (2024)
von: Mou, Guanyi, et al.
Veröffentlicht: (2024)
Word Boundary Information Isn't Useful for Encoder Language Models
von: Gow-Smith, Edward, et al.
Veröffentlicht: (2024)
von: Gow-Smith, Edward, et al.
Veröffentlicht: (2024)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
von: Yang, Chenchen, et al.
Veröffentlicht: (2026)
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
von: Greenfeld, Refael Shaked, et al.
Veröffentlicht: (2026)
von: Greenfeld, Refael Shaked, et al.
Veröffentlicht: (2026)
Constructing Cross-lingual Consumer Health Vocabulary with Word-Embedding from Comparable User Generated Content
von: Chang, Chia-Hsuan, et al.
Veröffentlicht: (2022)
von: Chang, Chia-Hsuan, et al.
Veröffentlicht: (2022)
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
von: Darwish, Hazem, et al.
Veröffentlicht: (2024)
von: Darwish, Hazem, et al.
Veröffentlicht: (2024)
From Word to World: Can Large Language Models be Implicit Text-based World Models?
von: Li, Yixia, et al.
Veröffentlicht: (2025)
von: Li, Yixia, et al.
Veröffentlicht: (2025)
CopyNE: Better Contextual ASR by Copying Named Entities
von: Zhou, Shilin, et al.
Veröffentlicht: (2023)
von: Zhou, Shilin, et al.
Veröffentlicht: (2023)
SAMGPT: Text-free Graph Foundation Model for Multi-domain Pre-training and Cross-domain Adaptation
von: Yu, Xingtong, et al.
Veröffentlicht: (2025)
von: Yu, Xingtong, et al.
Veröffentlicht: (2025)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2022)
von: Lopardo, Gianluigi, et al.
Veröffentlicht: (2022)
Word Order's Impacts: Insights from Reordering and Generation Analysis
von: Zhao, Qinghua, et al.
Veröffentlicht: (2024)
von: Zhao, Qinghua, et al.
Veröffentlicht: (2024)
Word Chain Generators for Prefix Normal Words
von: Adamson, Duncan, et al.
Veröffentlicht: (2025)
von: Adamson, Duncan, et al.
Veröffentlicht: (2025)
Cross-domain Chinese Sentence Pattern Parsing
von: Yu, Jingsi, et al.
Veröffentlicht: (2024)
von: Yu, Jingsi, et al.
Veröffentlicht: (2024)
Continuously Learning New Words in Automatic Speech Recognition
von: Huber, Christian, et al.
Veröffentlicht: (2024)
von: Huber, Christian, et al.
Veröffentlicht: (2024)
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
von: Zhang, Zihong, et al.
Veröffentlicht: (2025)
von: Zhang, Zihong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Character-Level Chinese Dependency Parsing via Modeling Latent Intra-Word Structure
von: Hou, Yang, et al.
Veröffentlicht: (2024) -
Parsing Through Boundaries in Chinese Word Segmentation
von: Chen, Yige, et al.
Veröffentlicht: (2025) -
JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR
von: Zhou, Shilin, et al.
Veröffentlicht: (2026) -
Chinese Word Boundary Recovery through Character Alignment Projection
von: Wang, Lusha, et al.
Veröffentlicht: (2026) -
Word Segmentation for Asian Languages: Chinese, Korean, and Japanese
von: Rho, Matthew, et al.
Veröffentlicht: (2024)