A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
Fuente:
arXiv
Guardado en:
| Autores principales: | Anikina, Tatiana, Cegin, Jan, Simko, Jakub, Ostermann, Simon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
por: Cegin, Jan, et al.
Publicado: (2025)
por: Cegin, Jan, et al.
Publicado: (2025)
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
por: Cegin, Jan, et al.
Publicado: (2024)
por: Cegin, Jan, et al.
Publicado: (2024)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
por: Pecher, Branislav, et al.
Publicado: (2026)
por: Pecher, Branislav, et al.
Publicado: (2026)
Large Language Models for Multilingual Previously Fact-Checked Claim Detection
por: Vykopal, Ivan, et al.
Publicado: (2025)
por: Vykopal, Ivan, et al.
Publicado: (2025)
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
por: Cegin, Jan, et al.
Publicado: (2024)
por: Cegin, Jan, et al.
Publicado: (2024)
Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
por: Cegin, Jan, et al.
Publicado: (2024)
por: Cegin, Jan, et al.
Publicado: (2024)
Soft Language Prompts for Language Transfer
por: Vykopal, Ivan, et al.
Publicado: (2024)
por: Vykopal, Ivan, et al.
Publicado: (2024)
MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models
por: Hyben, Martin, et al.
Publicado: (2026)
por: Hyben, Martin, et al.
Publicado: (2026)
Generative Large Language Models in Automated Fact-Checking: A Survey
por: Vykopal, Ivan, et al.
Publicado: (2024)
por: Vykopal, Ivan, et al.
Publicado: (2024)
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem
por: Wang, Qianli, et al.
Publicado: (2024)
por: Wang, Qianli, et al.
Publicado: (2024)
Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation
por: Pecher, Branislav, et al.
Publicado: (2024)
por: Pecher, Branislav, et al.
Publicado: (2024)
CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems
por: Wang, Qianli, et al.
Publicado: (2024)
por: Wang, Qianli, et al.
Publicado: (2024)
Assessing Web Search Credibility and Response Groundedness in Chat Assistants
por: Vykopal, Ivan, et al.
Publicado: (2025)
por: Vykopal, Ivan, et al.
Publicado: (2025)
On Multilingual Encoder Language Model Compression for Low-Resource Languages
por: Gurgurov, Daniil, et al.
Publicado: (2025)
por: Gurgurov, Daniil, et al.
Publicado: (2025)
Reverse Probing: Evaluating Knowledge Transfer via Finetuned Task Embeddings for Coreference Resolution
por: Anikina, Tatiana, et al.
Publicado: (2025)
por: Anikina, Tatiana, et al.
Publicado: (2025)
Multilingual Large Language Models and Curse of Multilinguality
por: Gurgurov, Daniil, et al.
Publicado: (2024)
por: Gurgurov, Daniil, et al.
Publicado: (2024)
Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters
por: Gurgurov, Daniil, et al.
Publicado: (2024)
por: Gurgurov, Daniil, et al.
Publicado: (2024)
GrEmLIn: A Repository of Green Baseline Embeddings for 87 Low-Resource Languages Injected with Multilingual Graph Knowledge
por: Gurgurov, Daniil, et al.
Publicado: (2024)
por: Gurgurov, Daniil, et al.
Publicado: (2024)
When Scale Meets Diversity: Evaluating Language Models on Fine-Grained Multilingual Claim Verification
por: Shcharbakova, Hanna, et al.
Publicado: (2025)
por: Shcharbakova, Hanna, et al.
Publicado: (2025)
What Language(s) Does Aya-23 Think In? How Multilinguality Affects Internal Language Representations
por: Trinley, Katharina, et al.
Publicado: (2025)
por: Trinley, Katharina, et al.
Publicado: (2025)
Small Models, Big Impact: Efficient Corpus and Graph-Based Adaptation of Small Multilingual Language Models for Low-Resource Languages
por: Gurgurov, Daniil, et al.
Publicado: (2025)
por: Gurgurov, Daniil, et al.
Publicado: (2025)
Political Leaning and Politicalness Classification of Texts
por: Volf, Matous, et al.
Publicado: (2025)
por: Volf, Matous, et al.
Publicado: (2025)
Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems
por: Wang, Qianli, et al.
Publicado: (2025)
por: Wang, Qianli, et al.
Publicado: (2025)
A Generative-AI-Driven Claim Retrieval System Capable of Detecting and Retrieving Claims from Social Media Platforms in Multiple Languages
por: Vykopal, Ivan, et al.
Publicado: (2025)
por: Vykopal, Ivan, et al.
Publicado: (2025)
Cross-Prompt Encoder for Low-Performing Languages
por: Mikaberidze, Beso, et al.
Publicado: (2025)
por: Mikaberidze, Beso, et al.
Publicado: (2025)
OpenFActScore: Open-Source Atomic Evaluation of Factuality in Text Generation
por: Lage, Lucas Fonseca, et al.
Publicado: (2025)
por: Lage, Lucas Fonseca, et al.
Publicado: (2025)
LLM Probe: Evaluating LLMs for Low-Resource Languages
por: Teklehaymanot, Hailay Kidu, et al.
Publicado: (2026)
por: Teklehaymanot, Hailay Kidu, et al.
Publicado: (2026)
Multilingual and Multi-topical Benchmark of Fine-tuned Language models and Large Language Models for Check-Worthy Claim Detection
por: Hyben, Martin, et al.
Publicado: (2023)
por: Hyben, Martin, et al.
Publicado: (2023)
Revisiting Prompt Sensitivity in Large Language Models for Text Classification: The Role of Prompt Underspecification
por: Pecher, Branislav, et al.
Publicado: (2026)
por: Pecher, Branislav, et al.
Publicado: (2026)
mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code
por: Skurla, Adam, et al.
Publicado: (2026)
por: Skurla, Adam, et al.
Publicado: (2026)
LLM-Based Data Generation and Clinical Skills Evaluation for Low-Resource French OSCEs
por: Huang, Tian, et al.
Publicado: (2026)
por: Huang, Tian, et al.
Publicado: (2026)
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
por: Macko, Dominik, et al.
Publicado: (2026)
por: Macko, Dominik, et al.
Publicado: (2026)
Interpretable Predictability-Based AI Text Detection: A Replication Study
por: Skurla, Adam, et al.
Publicado: (2026)
por: Skurla, Adam, et al.
Publicado: (2026)
TRepLiNa: Layer-wise CKA+REPINA Alignment Improves Low-Resource Machine Translation in Aya-23 8B
por: Nakai, Toshiki, et al.
Publicado: (2025)
por: Nakai, Toshiki, et al.
Publicado: (2025)
The AI Language Proficiency Monitor -- Tracking the Progress of LLMs on Multilingual Benchmarks
por: Pomerenke, David, et al.
Publicado: (2025)
por: Pomerenke, David, et al.
Publicado: (2025)
Probing Context Localization of Polysemous Words in Pre-trained Language Model Sub-Layers
por: Vijayakumar, Soniya, et al.
Publicado: (2024)
por: Vijayakumar, Soniya, et al.
Publicado: (2024)
ReasonXL: Shifting LLM Reasoning Language Without Sacrificing Performance
por: Gurgurov, Daniil, et al.
Publicado: (2026)
por: Gurgurov, Daniil, et al.
Publicado: (2026)
Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models
por: Gurgurov, Daniil, et al.
Publicado: (2025)
por: Gurgurov, Daniil, et al.
Publicado: (2025)
The Emperor's New Clothes in Benchmarking? A Rigorous Examination of Mitigation Strategies for LLM Benchmark Data Contamination
por: Sun, Yifan, et al.
Publicado: (2025)
por: Sun, Yifan, et al.
Publicado: (2025)
Unlocking LLM Safeguards for Low-Resource Languages via Reasoning and Alignment with Minimal Training Data
por: Chen, Zhuowei, et al.
Publicado: (2025)
por: Chen, Zhuowei, et al.
Publicado: (2025)
Ejemplares similares
-
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
por: Cegin, Jan, et al.
Publicado: (2025) -
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
por: Cegin, Jan, et al.
Publicado: (2024) -
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
por: Pecher, Branislav, et al.
Publicado: (2026) -
Large Language Models for Multilingual Previously Fact-Checked Claim Detection
por: Vykopal, Ivan, et al.
Publicado: (2025) -
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
por: Cegin, Jan, et al.
Publicado: (2024)