Comprehensive Evaluation on Lexical Normalization: Boundary-Aware Approaches for Unsegmented Languages
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Higashiyama, Shohei, Utiyama, Masao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CADEL: A Corpus of Administrative Web Documents for Japanese Entity Linking
von: Higashiyama, Shohei, et al.
Veröffentlicht: (2026)
von: Higashiyama, Shohei, et al.
Veröffentlicht: (2026)
ATD-Trans: A Geographically Grounded Japanese-English Travelogue Translation Dataset
von: Higashiyama, Shohei, et al.
Veröffentlicht: (2026)
von: Higashiyama, Shohei, et al.
Veröffentlicht: (2026)
Language-free Experience at Expo 2025 Osaka
von: Paul, Michael, et al.
Veröffentlicht: (2026)
von: Paul, Michael, et al.
Veröffentlicht: (2026)
On Eliciting Syntax from Language Models via Hashing
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
To be Continuous, or to be Discrete, Those are Bits of Questions
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
von: Wang, Yiran, et al.
Veröffentlicht: (2024)
OptiMer: Optimal Distribution Vector Merging Is Better than Data Mixing for Continual Pre-Training
von: Song, Haiyue, et al.
Veröffentlicht: (2026)
von: Song, Haiyue, et al.
Veröffentlicht: (2026)
Improving Language Transfer Capability of Decoder-only Architecture in Multilingual Neural Machine Translation
von: Qu, Zhi, et al.
Veröffentlicht: (2024)
von: Qu, Zhi, et al.
Veröffentlicht: (2024)
Recipe Generation from Unsegmented Cooking Videos
von: Nishimura, Taichi, et al.
Veröffentlicht: (2022)
von: Nishimura, Taichi, et al.
Veröffentlicht: (2022)
PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation
von: Kaing, Hour, et al.
Veröffentlicht: (2025)
von: Kaing, Hour, et al.
Veröffentlicht: (2025)
Registering Source Tokens to Target Language Spaces in Multilingual Neural Machine Translation
von: Qu, Zhi, et al.
Veröffentlicht: (2025)
von: Qu, Zhi, et al.
Veröffentlicht: (2025)
Centroid-Based Efficient Minimum Bayes Risk Decoding
von: Deguchi, Hiroyuki, et al.
Veröffentlicht: (2024)
von: Deguchi, Hiroyuki, et al.
Veröffentlicht: (2024)
ChiKhaPo: A Large-Scale Multilingual Benchmark for Evaluating Lexical Comprehension and Generation in Large Language Models
von: Chang, Emily, et al.
Veröffentlicht: (2025)
von: Chang, Emily, et al.
Veröffentlicht: (2025)
MultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian Languages
von: Buaphet, Weerayut, et al.
Veröffentlicht: (2026)
von: Buaphet, Weerayut, et al.
Veröffentlicht: (2026)
Word Sense Disambiguation in Native Spanish: A Comprehensive Lexical Evaluation Resource
von: Ortega, Pablo, et al.
Veröffentlicht: (2024)
von: Ortega, Pablo, et al.
Veröffentlicht: (2024)
When Alignment Hurts: Decoupling Representational Spaces in Multilingual Models
von: Elshabrawy, Ahmed, et al.
Veröffentlicht: (2025)
von: Elshabrawy, Ahmed, et al.
Veröffentlicht: (2025)
New Evaluation Paradigm for Lexical Simplification
von: Qiang, Jipeng, et al.
Veröffentlicht: (2025)
von: Qiang, Jipeng, et al.
Veröffentlicht: (2025)
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
von: Lyu, Boxuan, et al.
Veröffentlicht: (2025)
von: Lyu, Boxuan, et al.
Veröffentlicht: (2025)
Graph-Structured Trajectory Extraction from Travelogues
von: Yamamoto, Aitaro, et al.
Veröffentlicht: (2024)
von: Yamamoto, Aitaro, et al.
Veröffentlicht: (2024)
ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
von: Nguyen, Thanh-Nhi, et al.
Veröffentlicht: (2024)
von: Nguyen, Thanh-Nhi, et al.
Veröffentlicht: (2024)
IteRABRe: Iterative Recovery-Aided Block Reduction
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2025)
von: Wibowo, Haryo Akbarianto, et al.
Veröffentlicht: (2025)
Latent Lexical Projection in Large Language Models: A Novel Approach to Implicit Representation Refinement
von: Shaker, Ziad, et al.
Veröffentlicht: (2025)
von: Shaker, Ziad, et al.
Veröffentlicht: (2025)
Lexical Manifold Reconfiguration in Large Language Models: A Novel Architectural Approach for Contextual Modulation
von: Vassilis, Koinis, et al.
Veröffentlicht: (2025)
von: Vassilis, Koinis, et al.
Veröffentlicht: (2025)
Generative Psycho-Lexical Approach for Constructing Value Systems in Large Language Models
von: Ye, Haoran, et al.
Veröffentlicht: (2025)
von: Ye, Haoran, et al.
Veröffentlicht: (2025)
LexInstructEval: Lexical Instruction Following Evaluation for Large Language Models
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
von: Ren, Huimin, et al.
Veröffentlicht: (2025)
ViSoLex: An Open-Source Repository for Vietnamese Social Media Lexical Normalization
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2025)
von: Nguyen, Anh Thi-Hoang, et al.
Veröffentlicht: (2025)
SetLexSem Challenge: Using Set Operations to Evaluate the Lexical and Semantic Robustness of Language Models
von: Akhbari, Bardiya, et al.
Veröffentlicht: (2024)
von: Akhbari, Bardiya, et al.
Veröffentlicht: (2024)
TikZero: Zero-Shot Text-Guided Graphics Program Synthesis
von: Belouadi, Jonas, et al.
Veröffentlicht: (2025)
von: Belouadi, Jonas, et al.
Veröffentlicht: (2025)
The Development of a Comprehensive Spanish Dictionary for Phonetic and Lexical Tagging in Socio-phonetic Research (ESPADA)
von: Gonzalez, Simon
Veröffentlicht: (2024)
von: Gonzalez, Simon
Veröffentlicht: (2024)
Structured Document Translation via Format Reinforcement Learning
von: Song, Haiyue, et al.
Veröffentlicht: (2025)
von: Song, Haiyue, et al.
Veröffentlicht: (2025)
A Weakly Supervised Data Labeling Framework for Machine Lexical Normalization in Vietnamese Social Media
von: Nguyen, Dung Ha, et al.
Veröffentlicht: (2024)
von: Nguyen, Dung Ha, et al.
Veröffentlicht: (2024)
Crowdsourcing Lexical Diversity
von: Khalilia, Hadi, et al.
Veröffentlicht: (2024)
von: Khalilia, Hadi, et al.
Veröffentlicht: (2024)
Multi-Level Narrative Evaluation Outperforms Lexical Features for Mental Health
von: Ma, Yuxi, et al.
Veröffentlicht: (2026)
von: Ma, Yuxi, et al.
Veröffentlicht: (2026)
Dispersion Measures as Predictors of Lexical Decision Time, Word Familiarity, and Lexical Complexity
von: Nohejl, Adam, et al.
Veröffentlicht: (2025)
von: Nohejl, Adam, et al.
Veröffentlicht: (2025)
Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection
von: Phan-Tat, Bach, et al.
Veröffentlicht: (2026)
von: Phan-Tat, Bach, et al.
Veröffentlicht: (2026)
Towards Controllable Natural Language Inference through Lexical Inference Types
von: Zhang, Yingji, et al.
Veröffentlicht: (2023)
von: Zhang, Yingji, et al.
Veröffentlicht: (2023)
ProLex: A Benchmark for Language Proficiency-oriented Lexical Substitution
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanming, et al.
Veröffentlicht: (2024)
NAIST Academic Travelogue Dataset
von: Ouchi, Hiroki, et al.
Veröffentlicht: (2023)
von: Ouchi, Hiroki, et al.
Veröffentlicht: (2023)
A Multidimensional Framework for Evaluating Lexical Semantic Change with Social Science Applications
von: Baes, Naomi, et al.
Veröffentlicht: (2024)
von: Baes, Naomi, et al.
Veröffentlicht: (2024)
Evaluating Distributed Representations for Multi-Level Lexical Semantics: A Research Proposal
von: Liu, Zhu
Veröffentlicht: (2024)
von: Liu, Zhu
Veröffentlicht: (2024)
Two Spelling Normalization Approaches Based on Large Language Models
von: Domingo, Miguel, et al.
Veröffentlicht: (2025)
von: Domingo, Miguel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
CADEL: A Corpus of Administrative Web Documents for Japanese Entity Linking
von: Higashiyama, Shohei, et al.
Veröffentlicht: (2026) -
ATD-Trans: A Geographically Grounded Japanese-English Travelogue Translation Dataset
von: Higashiyama, Shohei, et al.
Veröffentlicht: (2026) -
Language-free Experience at Expo 2025 Osaka
von: Paul, Michael, et al.
Veröffentlicht: (2026) -
On Eliciting Syntax from Language Models via Hashing
von: Wang, Yiran, et al.
Veröffentlicht: (2024) -
To be Continuous, or to be Discrete, Those are Bits of Questions
von: Wang, Yiran, et al.
Veröffentlicht: (2024)