Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations
Fuente:
arXiv
Salvato in:
| Autori principali: | Pan, Linrong, Jiang, Chenglong, Hou, Gaoze, Gao, Ying |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Language Choice of Teochew Chinese Students in Vinh Chau – Soc Trang Today
di: Vu Nhat Tan
Pubblicazione: (2025)
di: Vu Nhat Tan
Pubblicazione: (2025)
Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin
di: Lin, Pin-Jie, et al.
Pubblicazione: (2024)
di: Lin, Pin-Jie, et al.
Pubblicazione: (2024)
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
di: Khouja, Jude, et al.
Pubblicazione: (2025)
di: Khouja, Jude, et al.
Pubblicazione: (2025)
Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't
di: Taguchi, Chihiro, et al.
Pubblicazione: (2024)
di: Taguchi, Chihiro, et al.
Pubblicazione: (2024)
Enhancing Robustness of Autoregressive Language Models against Orthographic Attacks via Pixel-based Approach
di: Yang, Han, et al.
Pubblicazione: (2025)
di: Yang, Han, et al.
Pubblicazione: (2025)
WildReward: Learning Reward Models from In-the-Wild Human Interactions
di: Peng, Hao, et al.
Pubblicazione: (2026)
di: Peng, Hao, et al.
Pubblicazione: (2026)
Herald: A Natural Language Annotated Lean 4 Dataset
di: Gao, Guoxiong, et al.
Pubblicazione: (2024)
di: Gao, Guoxiong, et al.
Pubblicazione: (2024)
TruthStance: An Annotated Dataset of Conversations on Truth Social
di: Ameen, Fathima, et al.
Pubblicazione: (2026)
di: Ameen, Fathima, et al.
Pubblicazione: (2026)
AKEW: Assessing Knowledge Editing in the Wild
di: Wu, Xiaobao, et al.
Pubblicazione: (2024)
di: Wu, Xiaobao, et al.
Pubblicazione: (2024)
WildIFEval: Instruction Following in the Wild
di: Lior, Gili, et al.
Pubblicazione: (2025)
di: Lior, Gili, et al.
Pubblicazione: (2025)
ShareChat: A Dataset of Chatbot Conversations in the Wild
di: Yan, Yueru, et al.
Pubblicazione: (2025)
di: Yan, Yueru, et al.
Pubblicazione: (2025)
Legend: Leveraging Representation Engineering to Annotate Safety Margin for Preference Datasets
di: Feng, Duanyu, et al.
Pubblicazione: (2024)
di: Feng, Duanyu, et al.
Pubblicazione: (2024)
Sina at FigNews 2024: Multilingual Datasets Annotated with Bias and Propaganda
di: Duaibes, Lina, et al.
Pubblicazione: (2024)
di: Duaibes, Lina, et al.
Pubblicazione: (2024)
ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
di: Jung, Jeesu, et al.
Pubblicazione: (2025)
BoundingDocs: a Unified Dataset for Document Question Answering with Spatial Annotations
di: Giovannini, Simone, et al.
Pubblicazione: (2025)
di: Giovannini, Simone, et al.
Pubblicazione: (2025)
MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotations
di: Scott, Aaron, et al.
Pubblicazione: (2025)
di: Scott, Aaron, et al.
Pubblicazione: (2025)
Reasoning with OmniThought: A Large CoT Dataset with Verbosity and Cognitive Difficulty Annotations
di: Cai, Wenrui, et al.
Pubblicazione: (2025)
di: Cai, Wenrui, et al.
Pubblicazione: (2025)
CMDAG: A Chinese Metaphor Dataset with Annotated Grounds as CoT for Boosting Metaphor Generation
di: Shao, Yujie, et al.
Pubblicazione: (2024)
di: Shao, Yujie, et al.
Pubblicazione: (2024)
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
di: Choi, Juhwan, et al.
Pubblicazione: (2024)
WildSci: Advancing Scientific Reasoning from In-the-Wild Literature
di: Liu, Tengxiao, et al.
Pubblicazione: (2026)
di: Liu, Tengxiao, et al.
Pubblicazione: (2026)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
AnnoCaseLaw: A Richly-Annotated Dataset For Benchmarking Explainable Legal Judgment Prediction
di: Sesodia, Magnus, et al.
Pubblicazione: (2025)
di: Sesodia, Magnus, et al.
Pubblicazione: (2025)
DeFine: A Decomposed and Fine-Grained Annotated Dataset for Long-form Article Generation
di: Wang, Ming, et al.
Pubblicazione: (2025)
di: Wang, Ming, et al.
Pubblicazione: (2025)
CasiMedicos-Arg: A Medical Question Answering Dataset Annotated with Explanatory Argumentative Structures
di: Sviridova, Ekaterina, et al.
Pubblicazione: (2024)
di: Sviridova, Ekaterina, et al.
Pubblicazione: (2024)
Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis
di: Lee, Nayeon, et al.
Pubblicazione: (2023)
di: Lee, Nayeon, et al.
Pubblicazione: (2023)
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
di: Kunilovskaya, Maria, et al.
Pubblicazione: (2026)
di: Kunilovskaya, Maria, et al.
Pubblicazione: (2026)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
di: Sheth, Rajvee, et al.
Pubblicazione: (2025)
di: Sheth, Rajvee, et al.
Pubblicazione: (2025)
NLP Security and Ethics, in the Wild
di: Lent, Heather, et al.
Pubblicazione: (2025)
di: Lent, Heather, et al.
Pubblicazione: (2025)
Vero: An Open RL Recipe for General Visual Reasoning
di: Sarch, Gabriel, et al.
Pubblicazione: (2026)
di: Sarch, Gabriel, et al.
Pubblicazione: (2026)
A Highly Clean Recipe Dataset with Ingredient States Annotation for State Probing Task
di: Toyooka, Mashiro, et al.
Pubblicazione: (2025)
di: Toyooka, Mashiro, et al.
Pubblicazione: (2025)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
di: Lu, Yujie, et al.
Pubblicazione: (2024)
di: Lu, Yujie, et al.
Pubblicazione: (2024)
CocoaBench: Evaluating Unified Digital Agents in the Wild
di: CocoaBench Team, et al.
Pubblicazione: (2026)
di: CocoaBench Team, et al.
Pubblicazione: (2026)
Knots: A Large-Scale Multi-Agent Enhanced Expert-Annotated Dataset and LLM Prompt Optimization for NOTAM Semantic Parsing
di: Liu, Maoqi, et al.
Pubblicazione: (2025)
di: Liu, Maoqi, et al.
Pubblicazione: (2025)
QACP: An Annotated Question Answering Dataset for Assisting Chinese Python Programming Learners
di: Xiao, Rui, et al.
Pubblicazione: (2024)
di: Xiao, Rui, et al.
Pubblicazione: (2024)
Repurposing Annotation Guidelines to Instruct LLM Annotators: A Case Study
di: Kim, Kon Woo, et al.
Pubblicazione: (2025)
di: Kim, Kon Woo, et al.
Pubblicazione: (2025)
How Annotation Trains Annotators: Competence Development in Social Influence Recognition
di: Markiewicz, Maciej, et al.
Pubblicazione: (2026)
di: Markiewicz, Maciej, et al.
Pubblicazione: (2026)
VLQA: The First Comprehensive, Large, and High-Quality Vietnamese Dataset for Legal Question Answering
di: Nguyen, Tan-Minh, et al.
Pubblicazione: (2025)
di: Nguyen, Tan-Minh, et al.
Pubblicazione: (2025)
SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching Dataset
di: Xie, Peng, et al.
Pubblicazione: (2025)
di: Xie, Peng, et al.
Pubblicazione: (2025)
Mapping Overlaps in Benchmarks through Perplexity in the Wild
di: Wu, Siyang, et al.
Pubblicazione: (2025)
di: Wu, Siyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Language Choice of Teochew Chinese Students in Vinh Chau – Soc Trang Today
di: Vu Nhat Tan
Pubblicazione: (2025) -
Modeling Orthographic Variation Improves NLP Performance for Nigerian Pidgin
di: Lin, Pin-Jie, et al.
Pubblicazione: (2024) -
LINGOLY-TOO: Disentangling Reasoning from Knowledge with Templatised Orthographic Obfuscation
di: Khouja, Jude, et al.
Pubblicazione: (2025) -
Language Complexity and Speech Recognition Accuracy: Orthographic Complexity Hurts, Phonological Complexity Doesn't
di: Taguchi, Chihiro, et al.
Pubblicazione: (2024) -
Enhancing Robustness of Autoregressive Language Models against Orthographic Attacks via Pixel-based Approach
di: Yang, Han, et al.
Pubblicazione: (2025)