Tokenization Matters: Improving Zero-Shot NER for Indic Languages
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pattnayak, Priyaranjan, Patel, Hitesh Laxmichand, Agarwal, Amit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
von: Patel, Hitesh Laxmichand, et al.
Veröffentlicht: (2024)
von: Patel, Hitesh Laxmichand, et al.
Veröffentlicht: (2024)
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2025)
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2025)
IndicJR: A Judge-Free Benchmark of Jailbreak Robustness in South Asian Languages
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
LLM-Guided Lifecycle-Aware Clustering of Multi-Turn Customer Support Conversations
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)
Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems
von: Meghwani, Hansa, et al.
Veröffentlicht: (2025)
von: Meghwani, Hansa, et al.
Veröffentlicht: (2025)
Hybrid AI for Responsive Multi-Turn Online Conversations with Novel Dynamic Routing and Feedback Adaptation
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2025)
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2025)
AccessEval: Benchmarking Disability Bias in Large Language Models
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
SN-WER: Script-Normalized WER for Multi-Script Indic ASR Evaluation
von: Pattnayak, Priyaranjan
Veröffentlicht: (2026)
von: Pattnayak, Priyaranjan
Veröffentlicht: (2026)
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2024)
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2024)
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
von: Patel, Hitesh Laxmichand, et al.
Veröffentlicht: (2025)
von: Patel, Hitesh Laxmichand, et al.
Veröffentlicht: (2025)
DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
Cross-Lingual IPA Contrastive Learning for Zero-Shot NER
von: Sohn, Jimin, et al.
Veröffentlicht: (2025)
von: Sohn, Jimin, et al.
Veröffentlicht: (2025)
Zero-Shot Cross-Lingual NER Using Phonemic Representations for Low-Resource Languages
von: Sohn, Jimin, et al.
Veröffentlicht: (2024)
von: Sohn, Jimin, et al.
Veröffentlicht: (2024)
SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks
von: Safarzadeh, Mohammadtaher, et al.
Veröffentlicht: (2026)
von: Safarzadeh, Mohammadtaher, et al.
Veröffentlicht: (2026)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
NADIR: Differential Attention Flow for Non-Autoregressive Transliteration in Indic Languages
von: Tomar, Lakshya, et al.
Veröffentlicht: (2026)
von: Tomar, Lakshya, et al.
Veröffentlicht: (2026)
SAM-NER: Semantic Archetype Mediation for Zero-Shot Named Entity Recognition
von: Cai, Ruichu, et al.
Veröffentlicht: (2026)
von: Cai, Ruichu, et al.
Veröffentlicht: (2026)
ReverseNER: A Self-Generated Example-Driven Framework for Zero-Shot Named Entity Recognition with Large Language Models
von: Wang, Anbang, et al.
Veröffentlicht: (2024)
von: Wang, Anbang, et al.
Veröffentlicht: (2024)
Think Twice Before You Write -- an Entropy-based Decoding Strategy to Enhance LLM Reasoning
von: He, Jiashu, et al.
Veröffentlicht: (2026)
von: He, Jiashu, et al.
Veröffentlicht: (2026)
NER Retriever: Zero-Shot Named Entity Retrieval with Type-Aware Embeddings
von: Shachar, Or, et al.
Veröffentlicht: (2025)
von: Shachar, Or, et al.
Veröffentlicht: (2025)
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding
von: KJ, Sankalp, et al.
Veröffentlicht: (2025)
von: KJ, Sankalp, et al.
Veröffentlicht: (2025)
Relic: Enhancing Reward Model Generalization for Low-Resource Indic Languages with Few-Shot Examples
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2025)
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2025)
Language-Independent Representations Improve Zero-Shot Summarization
von: Solovyev, Vladimir, et al.
Veröffentlicht: (2024)
von: Solovyev, Vladimir, et al.
Veröffentlicht: (2024)
BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic Languages
von: Manoj, Guduru, et al.
Veröffentlicht: (2025)
von: Manoj, Guduru, et al.
Veröffentlicht: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
von: Kim, Eunsu, et al.
Veröffentlicht: (2025)
Aligning LLMs for Multilingual Consistency in Enterprise Applications
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
Zero- and Few-Shot Named-Entity Recognition: Case Study and Dataset in the Crime Domain (CrimeNER)
von: Lopez-Duran, Miguel, et al.
Veröffentlicht: (2026)
von: Lopez-Duran, Miguel, et al.
Veröffentlicht: (2026)
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations
von: Singh, Jyotika, et al.
Veröffentlicht: (2026)
von: Singh, Jyotika, et al.
Veröffentlicht: (2026)
IndicMedDialog: A Parallel Multi-Turn Medical Dialogue Dataset for Accessible Healthcare in Indic Languages
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2026)
von: Nigam, Shubham Kumar, et al.
Veröffentlicht: (2026)
2M-NER: Contrastive Learning for Multilingual and Multimodal NER with Language and Modal Fusion
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
von: Dua, Karan, et al.
Veröffentlicht: (2025)
von: Dua, Karan, et al.
Veröffentlicht: (2025)
Indic-TunedLens: Interpreting Multilingual Models in Indian Languages
von: Panchal, Mihir, et al.
Veröffentlicht: (2026)
von: Panchal, Mihir, et al.
Veröffentlicht: (2026)
Tokenization Matters! Degrading Large Language Models through Challenging Their Tokenization
von: Wang, Dixuan, et al.
Veröffentlicht: (2024)
von: Wang, Dixuan, et al.
Veröffentlicht: (2024)
Multilingual State Space Models for Structured Question Answering in Indic Languages
von: Vats, Arpita, et al.
Veröffentlicht: (2025)
von: Vats, Arpita, et al.
Veröffentlicht: (2025)
IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?
von: Aravapalli, Akhilesh, et al.
Veröffentlicht: (2024)
von: Aravapalli, Akhilesh, et al.
Veröffentlicht: (2024)
WikiNER-fr-gold: A Gold-Standard NER Corpus
von: Cao, Danrun, et al.
Veröffentlicht: (2024)
von: Cao, Danrun, et al.
Veröffentlicht: (2024)
Using Large Language Model for End-to-End Chinese ASR and NER
von: Li, Yuang, et al.
Veröffentlicht: (2024)
von: Li, Yuang, et al.
Veröffentlicht: (2024)
Tokenizer Fertility and Zero-Shot Performance of Foundation Models on Ukrainian Legal Text: A Comparative Study
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
von: Ovcharov, Volodymyr
Veröffentlicht: (2026)
Ähnliche Einträge
-
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
von: Patel, Hitesh Laxmichand, et al.
Veröffentlicht: (2024) -
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2025) -
IndicJR: A Judge-Free Benchmark of Jailbreak Robustness in South Asian Languages
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026) -
LLM-Guided Lifecycle-Aware Clustering of Multi-Turn Customer Support Conversations
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026) -
IndicSafe: A Benchmark for Evaluating Multilingual LLM Safety in South Asia
von: Pattnayak, Priyaranjan, et al.
Veröffentlicht: (2026)