Mining Word Boundaries from Speech-Text Parallel Data for Cross-domain Chinese Word Segmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Xuebin, Zhang, Lei, Li, Zhenghua, Zhou, Shilin, Gong, Chen, Hou, Yang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Character-Level Chinese Dependency Parsing via Modeling Latent Intra-Word Structure
por: Hou, Yang, et al.
Publicado: (2024)
por: Hou, Yang, et al.
Publicado: (2024)
Parsing Through Boundaries in Chinese Word Segmentation
por: Chen, Yige, et al.
Publicado: (2025)
por: Chen, Yige, et al.
Publicado: (2025)
JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR
por: Zhou, Shilin, et al.
Publicado: (2026)
por: Zhou, Shilin, et al.
Publicado: (2026)
Chinese Word Boundary Recovery through Character Alignment Projection
por: Wang, Lusha, et al.
Publicado: (2026)
por: Wang, Lusha, et al.
Publicado: (2026)
Word Segmentation for Asian Languages: Chinese, Korean, and Japanese
por: Rho, Matthew, et al.
Publicado: (2024)
por: Rho, Matthew, et al.
Publicado: (2024)
UniSLU: Unified Spoken Language Understanding from Heterogeneous Cross-Task Datasets
por: Sheng, Zhichao, et al.
Publicado: (2025)
por: Sheng, Zhichao, et al.
Publicado: (2025)
Improving Contextual ASR via Multi-grained Fusion with Large Language Models
por: Zhou, Shilin, et al.
Publicado: (2025)
por: Zhou, Shilin, et al.
Publicado: (2025)
Self-Correction Makes LLMs Better Parsers
por: Zhang, Ziyan, et al.
Publicado: (2025)
por: Zhang, Ziyan, et al.
Publicado: (2025)
Mandarin Chinese Words and Parts of Speech
por: Huang, Chu-Ren, et al.
Publicado: (2021)
por: Huang, Chu-Ren, et al.
Publicado: (2021)
BabyLM's First Words: Word Segmentation as a Phonological Probing Task
por: Goriely, Zébulon, et al.
Publicado: (2025)
por: Goriely, Zébulon, et al.
Publicado: (2025)
Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination
por: He, Xiaoqi, et al.
Publicado: (2026)
por: He, Xiaoqi, et al.
Publicado: (2026)
Chinese ModernBERT with Whole-Word Masking
por: Zhao, Zeyu, et al.
Publicado: (2025)
por: Zhao, Zeyu, et al.
Publicado: (2025)
Using Context to Improve Word Segmentation
por: Hu, Stephanie, et al.
Publicado: (2025)
por: Hu, Stephanie, et al.
Publicado: (2025)
Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts
por: Pulipaka, Sidharth, et al.
Publicado: (2025)
por: Pulipaka, Sidharth, et al.
Publicado: (2025)
Span-Aggregatable, Contextualized Word Embeddings for Effective Phrase Mining
por: Orbach, Eyal, et al.
Publicado: (2024)
por: Orbach, Eyal, et al.
Publicado: (2024)
A Comparative Analysis of Word Segmentation, Part-of-Speech Tagging, and Named Entity Recognition for Historical Chinese Sources, 1900-1950
por: Fang, Zhao, et al.
Publicado: (2025)
por: Fang, Zhao, et al.
Publicado: (2025)
Cross-Lingual Word Alignment for ASEAN Languages with Contrastive Learning
por: Zhang, Jingshen, et al.
Publicado: (2024)
por: Zhang, Jingshen, et al.
Publicado: (2024)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association Test
por: Dai, Xunlian, et al.
Publicado: (2025)
por: Dai, Xunlian, et al.
Publicado: (2025)
Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency Parsing
por: Li, Zhuoran, et al.
Publicado: (2025)
por: Li, Zhuoran, et al.
Publicado: (2025)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
por: Park, Chanho, et al.
Publicado: (2023)
por: Park, Chanho, et al.
Publicado: (2023)
Mixture of Small and Large Models for Chinese Spelling Check
por: Qiao, Ziheng, et al.
Publicado: (2025)
por: Qiao, Ziheng, et al.
Publicado: (2025)
From Words to Worlds: Benchmarking Cross-Cultural Cultural Understanding in Machine Translation
por: Han, Bangju, et al.
Publicado: (2026)
por: Han, Bangju, et al.
Publicado: (2026)
RealTalk-CN: A Realistic Chinese Speech-Text Dialogue Benchmark With Cross-Modal Interaction Analysis
por: Wang, Enzhi, et al.
Publicado: (2025)
por: Wang, Enzhi, et al.
Publicado: (2025)
From Word2Vec to Transformers: Text-Derived Composition Embeddings for Filtering Combinatorial Electrocatalysts
por: Zhang, Lei, et al.
Publicado: (2026)
por: Zhang, Lei, et al.
Publicado: (2026)
SWE2: SubWord Enriched and Significant Word Emphasized Framework for Hate Speech Detection
por: Mou, Guanyi, et al.
Publicado: (2024)
por: Mou, Guanyi, et al.
Publicado: (2024)
Word Boundary Information Isn't Useful for Encoder Language Models
por: Gow-Smith, Edward, et al.
Publicado: (2024)
por: Gow-Smith, Edward, et al.
Publicado: (2024)
WESR: Scaling and Evaluating Word-level Event-Speech Recognition
por: Yang, Chenchen, et al.
Publicado: (2026)
por: Yang, Chenchen, et al.
Publicado: (2026)
Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text
por: Greenfeld, Refael Shaked, et al.
Publicado: (2026)
por: Greenfeld, Refael Shaked, et al.
Publicado: (2026)
Constructing Cross-lingual Consumer Health Vocabulary with Word-Embedding from Comparable User Generated Content
por: Chang, Chia-Hsuan, et al.
Publicado: (2022)
por: Chang, Chia-Hsuan, et al.
Publicado: (2022)
ArEEG_Words: Dataset for Envisioned Speech Recognition using EEG for Arabic Words
por: Darwish, Hazem, et al.
Publicado: (2024)
por: Darwish, Hazem, et al.
Publicado: (2024)
From Word to World: Can Large Language Models be Implicit Text-based World Models?
por: Li, Yixia, et al.
Publicado: (2025)
por: Li, Yixia, et al.
Publicado: (2025)
CopyNE: Better Contextual ASR by Copying Named Entities
por: Zhou, Shilin, et al.
Publicado: (2023)
por: Zhou, Shilin, et al.
Publicado: (2023)
SAMGPT: Text-free Graph Foundation Model for Multi-domain Pre-training and Cross-domain Adaptation
por: Yu, Xingtong, et al.
Publicado: (2025)
por: Yu, Xingtong, et al.
Publicado: (2025)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
por: Lopardo, Gianluigi, et al.
Publicado: (2022)
por: Lopardo, Gianluigi, et al.
Publicado: (2022)
Word Order's Impacts: Insights from Reordering and Generation Analysis
por: Zhao, Qinghua, et al.
Publicado: (2024)
por: Zhao, Qinghua, et al.
Publicado: (2024)
Word Chain Generators for Prefix Normal Words
por: Adamson, Duncan, et al.
Publicado: (2025)
por: Adamson, Duncan, et al.
Publicado: (2025)
Cross-domain Chinese Sentence Pattern Parsing
por: Yu, Jingsi, et al.
Publicado: (2024)
por: Yu, Jingsi, et al.
Publicado: (2024)
Continuously Learning New Words in Automatic Speech Recognition
por: Huber, Christian, et al.
Publicado: (2024)
por: Huber, Christian, et al.
Publicado: (2024)
Segment First or Comprehend First? Explore the Limit of Unsupervised Word Segmentation with Large Language Models
por: Zhang, Zihong, et al.
Publicado: (2025)
por: Zhang, Zihong, et al.
Publicado: (2025)
Ejemplares similares
-
Character-Level Chinese Dependency Parsing via Modeling Latent Intra-Word Structure
por: Hou, Yang, et al.
Publicado: (2024) -
Parsing Through Boundaries in Chinese Word Segmentation
por: Chen, Yige, et al.
Publicado: (2025) -
JSPG: Dynamic Dictionary Filtering via Joint Semantic-Pinyin-Glyph Retrieval for Chinese Contextual ASR
por: Zhou, Shilin, et al.
Publicado: (2026) -
Chinese Word Boundary Recovery through Character Alignment Projection
por: Wang, Lusha, et al.
Publicado: (2026) -
Word Segmentation for Asian Languages: Chinese, Korean, and Japanese
por: Rho, Matthew, et al.
Publicado: (2024)