CSCD-NS: a Chinese Spelling Check Dataset for Native Speakers
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Yong, Meng, Fandong, Zhou, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
C-LLM: Learn to Check Chinese Spelling Errors Character by Character
by: Li, Kunting, et al.
Published: (2024)
by: Li, Kunting, et al.
Published: (2024)
DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling Check
by: Qiao, Ziheng, et al.
Published: (2024)
by: Qiao, Ziheng, et al.
Published: (2024)
An Empirical Investigation of Domain Adaptation Ability for Chinese Spelling Check Models
by: Wang, Xi, et al.
Published: (2024)
by: Wang, Xi, et al.
Published: (2024)
DeepTrans: Deep Reasoning Translation via Reinforcement Learning
by: Wang, Jiaan, et al.
Published: (2025)
by: Wang, Jiaan, et al.
Published: (2025)
ExTrans: Multilingual Deep Reasoning Translation via Exemplar-Enhanced Reinforcement Learning
by: Wang, Jiaan, et al.
Published: (2025)
by: Wang, Jiaan, et al.
Published: (2025)
ChineseErrorCorrector3-4B: State-of-the-Art Chinese Spelling and Grammar Corrector
by: Tian, Wei, et al.
Published: (2025)
by: Tian, Wei, et al.
Published: (2025)
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
by: Shao, Chenze, et al.
Published: (2024)
by: Shao, Chenze, et al.
Published: (2024)
EAG: Extract and Generate Multi-way Aligned Corpus for Complete Multi-lingual Neural Machine Translation
by: Xu, Yulin, et al.
Published: (2022)
by: Xu, Yulin, et al.
Published: (2022)
Retrieval-Augmented Machine Translation with Unstructured Knowledge
by: Wang, Jiaan, et al.
Published: (2024)
by: Wang, Jiaan, et al.
Published: (2024)
TEAL: Tokenize and Embed ALL for Multi-modal Large Language Models
by: Yang, Zhen, et al.
Published: (2023)
by: Yang, Zhen, et al.
Published: (2023)
DRT: Deep Reasoning Translation via Long Chain-of-Thought
by: Wang, Jiaan, et al.
Published: (2024)
by: Wang, Jiaan, et al.
Published: (2024)
CANDY: Benchmarking LLMs' Limitations and Assistive Potential in Chinese Misinformation Fact-Checking
by: Guo, Ruiling, et al.
Published: (2025)
by: Guo, Ruiling, et al.
Published: (2025)
General learned delegation by clones
by: Li, Darren, et al.
Published: (2026)
by: Li, Darren, et al.
Published: (2026)
Chinese Spelling Correction: A Comprehensive Survey of Progress, Challenges, and Opportunities
by: Liu, Changchun, et al.
Published: (2025)
by: Liu, Changchun, et al.
Published: (2025)
Continuous Autoregressive Language Models
by: Shao, Chenze, et al.
Published: (2025)
by: Shao, Chenze, et al.
Published: (2025)
EdaCSC: Two Easy Data Augmentation Methods for Chinese Spelling Correction
by: Sheng, Lei, et al.
Published: (2024)
by: Sheng, Lei, et al.
Published: (2024)
LCS: A Language Converter Strategy for Zero-Shot Neural Machine Translation
by: Sun, Zengkui, et al.
Published: (2024)
by: Sun, Zengkui, et al.
Published: (2024)
Translatotron-V(ison): An End-to-End Model for In-Image Machine Translation
by: Lan, Zhibin, et al.
Published: (2024)
by: Lan, Zhibin, et al.
Published: (2024)
TasTe: Teaching Large Language Models to Translate through Self-Reflection
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
AraSpell: A Deep Learning Approach for Arabic Spelling Correction
by: Salhab, Mahmoud, et al.
Published: (2024)
by: Salhab, Mahmoud, et al.
Published: (2024)
Error-Robust Retrieval for Chinese Spelling Check
by: Yin, Xunjian, et al.
Published: (2022)
by: Yin, Xunjian, et al.
Published: (2022)
On the token distance modeling ability of higher RoPE attention dimension
by: Hong, Xiangyu, et al.
Published: (2024)
by: Hong, Xiangyu, et al.
Published: (2024)
MedFact: Benchmarking the Fact-Checking Capabilities of Large Language Models on Chinese Medical Texts
by: He, Jiayi, et al.
Published: (2025)
by: He, Jiayi, et al.
Published: (2025)
SpellForger: Prompting Custom Spell Properties In-Game using BERT supervised-trained model
by: Silva, Emanuel C., et al.
Published: (2025)
by: Silva, Emanuel C., et al.
Published: (2025)
Outdated Issue Aware Decoding for Reasoning Questions on Edited Knowledge
by: Sun, Zengkui, et al.
Published: (2024)
by: Sun, Zengkui, et al.
Published: (2024)
Mixture of Small and Large Models for Chinese Spelling Check
by: Qiao, Ziheng, et al.
Published: (2025)
by: Qiao, Ziheng, et al.
Published: (2025)
DelTA: An Online Document-Level Translation Agent Based on Multi-Level Memory
by: Wang, Yutong, et al.
Published: (2024)
by: Wang, Yutong, et al.
Published: (2024)
An Empirical Study of Many-to-Many Summarization with Large Language Models
by: Wang, Jiaan, et al.
Published: (2025)
by: Wang, Jiaan, et al.
Published: (2025)
HGMEM: Hypergraph-based Working Memory to Improve Multi-step RAG for Long-Context Complex Relational Modeling
by: Zhou, Chulun, et al.
Published: (2025)
by: Zhou, Chulun, et al.
Published: (2025)
AVG-LLaVA: An Efficient Large Multimodal Model with Adaptive Visual Granularity
by: Lan, Zhibin, et al.
Published: (2024)
by: Lan, Zhibin, et al.
Published: (2024)
ClaimGen-CN: A Large-scale Chinese Dataset for Legal Claim Generation
by: Zhou, Siying, et al.
Published: (2025)
by: Zhou, Siying, et al.
Published: (2025)
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning
by: Lan, Zhibin, et al.
Published: (2025)
by: Lan, Zhibin, et al.
Published: (2025)
Imperfect Language, Artificial Intelligence, and the Human Mind: An Interdisciplinary Approach to Linguistic Errors in Native Spanish Speakers
by: López, Francisco Portillo
Published: (2025)
by: López, Francisco Portillo
Published: (2025)
Retrieval Augmented Spelling Correction for E-Commerce Applications
by: Guo, Xuan, et al.
Published: (2024)
by: Guo, Xuan, et al.
Published: (2024)
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations
by: Xing, Rui, et al.
Published: (2025)
by: Xing, Rui, et al.
Published: (2025)
Towards Threshold-Free KV Cache Pruning
by: Ni, Xuanfan, et al.
Published: (2025)
by: Ni, Xuanfan, et al.
Published: (2025)
A Dual-Space Framework for General Knowledge Distillation of Large Language Models
by: Zhang, Xue, et al.
Published: (2025)
by: Zhang, Xue, et al.
Published: (2025)
Unsupervised Information Refinement Training of Large Language Models for Retrieval-Augmented Generation
by: Xu, Shicheng, et al.
Published: (2024)
by: Xu, Shicheng, et al.
Published: (2024)
Survey of Pseudonymization, Abstractive Summarization & Spell Checker for Hindi and Marathi
by: Ransing, Rasika, et al.
Published: (2024)
by: Ransing, Rasika, et al.
Published: (2024)
CFEVER: A Chinese Fact Extraction and VERification Dataset
by: Lin, Ying-Jia, et al.
Published: (2024)
by: Lin, Ying-Jia, et al.
Published: (2024)
Similar Items
-
C-LLM: Learn to Check Chinese Spelling Errors Character by Character
by: Li, Kunting, et al.
Published: (2024) -
DISC: Plug-and-Play Decoding Intervention with Similarity of Characters for Chinese Spelling Check
by: Qiao, Ziheng, et al.
Published: (2024) -
An Empirical Investigation of Domain Adaptation Ability for Chinese Spelling Check Models
by: Wang, Xi, et al.
Published: (2024) -
DeepTrans: Deep Reasoning Translation via Reinforcement Learning
by: Wang, Jiaan, et al.
Published: (2025) -
ExTrans: Multilingual Deep Reasoning Translation via Exemplar-Enhanced Reinforcement Learning
by: Wang, Jiaan, et al.
Published: (2025)