MultiLexNorm++: A Unified Benchmark and a Generative Model for Lexical Normalization for Asian Languages
Fuente:
arXiv
Salvato in:
| Autori principali: | Buaphet, Weerayut, Nguyen, Thanh-Nhi, Kondo, Risa, Kajiwara, Tomoyuki, Kim, Yumin, Lee, Jimin, Lee, Hwanhee, Lovenia, Holy, Limkonchotiwat, Peerat, Nutanong, Sarana, Van der Goot, Rob |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai
di: Pengpun, Parinthapat, et al.
Pubblicazione: (2024)
di: Pengpun, Parinthapat, et al.
Pubblicazione: (2024)
Space Decomposition for Sentence Embedding
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2024)
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2024)
When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
di: Tuchinda, Pume, et al.
Pubblicazione: (2025)
di: Tuchinda, Pume, et al.
Pubblicazione: (2025)
Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
di: Sriratanawilai, Sukrit, et al.
Pubblicazione: (2025)
di: Sriratanawilai, Sukrit, et al.
Pubblicazione: (2025)
ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
di: Nguyen, Thanh-Nhi, et al.
Pubblicazione: (2024)
di: Nguyen, Thanh-Nhi, et al.
Pubblicazione: (2024)
Assessing Thai Dialect Performance in LLMs with Automatic Benchmarks and Human Evaluation
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
Mangosteen: An Open Thai Corpus for Language Model Pretraining
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2025)
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2025)
WangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
di: Limkonchotiwat, Peerat, et al.
Pubblicazione: (2025)
Towards Better Understanding of Program-of-Thought Reasoning in Cross-Lingual and Multilingual Environments
di: Payoungkhamdee, Patomporn, et al.
Pubblicazione: (2025)
di: Payoungkhamdee, Patomporn, et al.
Pubblicazione: (2025)
WangchanLion and WangchanX MRC Eval
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2024)
di: Phatthiyaphaibun, Wannaphong, et al.
Pubblicazione: (2024)
SEA-BED: How Do Embedding Models Represent Southeast Asian Languages?
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2025)
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2025)
Selective Demonstration Retrieval for Improved Implicit Hate Speech Detection
di: Kim, Yumin, et al.
Pubblicazione: (2025)
di: Kim, Yumin, et al.
Pubblicazione: (2025)
Addressing Topic Leakage in Cross-Topic Evaluation for Authorship Verification
di: Sawatphol, Jitkapat, et al.
Pubblicazione: (2024)
di: Sawatphol, Jitkapat, et al.
Pubblicazione: (2024)
LLMs Are Few-Shot In-Context Low-Resource Language Learners
di: Cahyawijaya, Samuel, et al.
Pubblicazione: (2024)
di: Cahyawijaya, Samuel, et al.
Pubblicazione: (2024)
Prior Prompt Engineering for Reinforcement Fine-Tuning
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
di: Taveekitworachai, Pittawat, et al.
Pubblicazione: (2025)
A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2026)
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2026)
Crafting the Path: Robust Query Rewriting for Information Retrieval
di: Baek, Ingeol, et al.
Pubblicazione: (2024)
di: Baek, Ingeol, et al.
Pubblicazione: (2024)
Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering
di: Chang, Hwan, et al.
Pubblicazione: (2025)
di: Chang, Hwan, et al.
Pubblicazione: (2025)
Exploring Cross-Client Memorization of Training Data in Large Language Models for Federated Learning
di: Udsa, Tinnakit, et al.
Pubblicazione: (2025)
di: Udsa, Tinnakit, et al.
Pubblicazione: (2025)
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models
di: Lovenia, Holy, et al.
Pubblicazione: (2023)
di: Lovenia, Holy, et al.
Pubblicazione: (2023)
Probing-RAG: Self-Probing to Guide Language Models in Selective Document Retrieval
di: Baek, Ingeol, et al.
Pubblicazione: (2024)
di: Baek, Ingeol, et al.
Pubblicazione: (2024)
SAFE-SQL: Self-Augmented In-Context Learning with Fine-grained Example Selection for Text-to-SQL
di: Lee, Jimin, et al.
Pubblicazione: (2025)
di: Lee, Jimin, et al.
Pubblicazione: (2025)
Distilling Monolingual and Crosslingual Word-in-Context Representations
di: Arase, Yuki, et al.
Pubblicazione: (2024)
di: Arase, Yuki, et al.
Pubblicazione: (2024)
BURMESE-SAN: Burmese NLP Benchmark for Evaluating Large Language Models
di: Aung, Thura, et al.
Pubblicazione: (2026)
di: Aung, Thura, et al.
Pubblicazione: (2026)
KoCoSa: Korean Context-aware Sarcasm Detection Dataset
di: Kim, Yumin, et al.
Pubblicazione: (2024)
di: Kim, Yumin, et al.
Pubblicazione: (2024)
Personality Editing for Language Models through Adjusting Self-Referential Queries
di: Hwang, Seojin, et al.
Pubblicazione: (2025)
di: Hwang, Seojin, et al.
Pubblicazione: (2025)
LexBoost: Improving Lexical Document Retrieval with Nearest Neighbors
di: Kulkarni, Hrishikesh, et al.
Pubblicazione: (2024)
di: Kulkarni, Hrishikesh, et al.
Pubblicazione: (2024)
Big City Bias: Evaluating the Impact of Metropolitan Size on Computational Job Market Abilities of Language Models
di: Campanella, Charlie, et al.
Pubblicazione: (2024)
di: Campanella, Charlie, et al.
Pubblicazione: (2024)
NLPnorth @ TalentCLEF 2025: Comparing Discriminative, Contrastive, and Prompt-Based Methods for Job Title and Skill Matching
di: Zhang, Mike, et al.
Pubblicazione: (2025)
di: Zhang, Mike, et al.
Pubblicazione: (2025)
Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking
di: Chaichana, Yuatyong, et al.
Pubblicazione: (2025)
di: Chaichana, Yuatyong, et al.
Pubblicazione: (2025)
SEA-SafeguardBench: Evaluating AI Safety in SEA Languages and Cultures
di: Tasawong, Panuthep, et al.
Pubblicazione: (2025)
di: Tasawong, Panuthep, et al.
Pubblicazione: (2025)
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
di: Tasawong, Panuthep, et al.
Pubblicazione: (2026)
di: Tasawong, Panuthep, et al.
Pubblicazione: (2026)
Can Group Relative Policy Optimization Improve Thai Legal Reasoning and Question Answering?
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2025)
di: Akarajaradwong, Pawitsapak, et al.
Pubblicazione: (2025)
ProLex: A Benchmark for Language Proficiency-oriented Lexical Substitution
di: Zhang, Xuanming, et al.
Pubblicazione: (2024)
di: Zhang, Xuanming, et al.
Pubblicazione: (2024)
Exploring Persona Sentiment Sensitivity in Personalized Dialogue Generation
di: Jun, Yonghyun, et al.
Pubblicazione: (2025)
di: Jun, Yonghyun, et al.
Pubblicazione: (2025)
Dynamic Order Template Prediction for Generative Aspect-Based Sentiment Analysis
di: Jun, Yonghyun, et al.
Pubblicazione: (2024)
di: Jun, Yonghyun, et al.
Pubblicazione: (2024)
Conversational Query Reformulation with the Guidance of Retrieved Documents
di: Park, Jeonghyun, et al.
Pubblicazione: (2024)
di: Park, Jeonghyun, et al.
Pubblicazione: (2024)
Which Retain Set Matters for LLM Unlearning? A Case Study on Entity Unlearning
di: Chang, Hwan, et al.
Pubblicazione: (2025)
di: Chang, Hwan, et al.
Pubblicazione: (2025)
Investigating Language Preference of Multilingual RAG Systems
di: Park, Jeonghyun, et al.
Pubblicazione: (2025)
di: Park, Jeonghyun, et al.
Pubblicazione: (2025)
Edit-Constrained Decoding for Sentence Simplification
di: Zetsu, Tatsuya, et al.
Pubblicazione: (2024)
di: Zetsu, Tatsuya, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Seed-Free Synthetic Data Generation Framework for Instruction-Tuning LLMs: A Case Study in Thai
di: Pengpun, Parinthapat, et al.
Pubblicazione: (2024) -
Space Decomposition for Sentence Embedding
di: Ponwitayarat, Wuttikorn, et al.
Pubblicazione: (2024) -
When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
di: Tuchinda, Pume, et al.
Pubblicazione: (2025) -
Distilling Multilingual Vision-Language Models: When Smaller Models Stay Multilingual
di: Sriratanawilai, Sukrit, et al.
Pubblicazione: (2025) -
ViLexNorm: A Lexical Normalization Corpus for Vietnamese Social Media Text
di: Nguyen, Thanh-Nhi, et al.
Pubblicazione: (2024)