Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Pecher, Branislav, Cegin, Jan, Belanec, Robert, Srba, Ivan, Simko, Jakub, Bielikova, Maria |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
by: Cegin, Jan, et al.
Published: (2025)
by: Cegin, Jan, et al.
Published: (2025)
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
by: Cegin, Jan, et al.
Published: (2024)
by: Cegin, Jan, et al.
Published: (2024)
Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
by: Cegin, Jan, et al.
Published: (2024)
by: Cegin, Jan, et al.
Published: (2024)
PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
by: Belanec, Robert, et al.
Published: (2025)
by: Belanec, Robert, et al.
Published: (2025)
Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
Revisiting Prompt Sensitivity in Large Language Models for Text Classification: The Role of Prompt Underspecification
by: Pecher, Branislav, et al.
Published: (2026)
by: Pecher, Branislav, et al.
Published: (2026)
A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness
by: Pecher, Branislav, et al.
Published: (2023)
by: Pecher, Branislav, et al.
Published: (2023)
On Sensitivity of Learning with Limited Labelled Data to the Effects of Randomness: Impact of Interactions and Systematic Choices
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
by: Belanec, Robert, et al.
Published: (2025)
by: Belanec, Robert, et al.
Published: (2025)
Automatic Combination of Sample Selection Strategies for Few-Shot Learning
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer
by: Belanec, Robert, et al.
Published: (2024)
by: Belanec, Robert, et al.
Published: (2024)
A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
by: Anikina, Tatiana, et al.
Published: (2025)
by: Anikina, Tatiana, et al.
Published: (2025)
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
by: Cegin, Jan, et al.
Published: (2024)
by: Cegin, Jan, et al.
Published: (2024)
MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models
by: Hyben, Martin, et al.
Published: (2026)
by: Hyben, Martin, et al.
Published: (2026)
Multilingual Previously Fact-Checked Claim Retrieval
by: Pikuliak, Matúš, et al.
Published: (2023)
by: Pikuliak, Matúš, et al.
Published: (2023)
Authorship Obfuscation in Multilingual Machine-Generated Text Detection
by: Macko, Dominik, et al.
Published: (2024)
by: Macko, Dominik, et al.
Published: (2024)
KInITVeraAI at SemEval-2023 Task 3: Simple yet Powerful Multilingual Fine-Tuning for Persuasion Techniques Detection
by: Hromadka, Timo, et al.
Published: (2023)
by: Hromadka, Timo, et al.
Published: (2023)
Multilingual and Multi-topical Benchmark of Fine-tuned Language models and Large Language Models for Check-Worthy Claim Detection
by: Hyben, Martin, et al.
Published: (2023)
by: Hyben, Martin, et al.
Published: (2023)
MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark
by: Macko, Dominik, et al.
Published: (2023)
by: Macko, Dominik, et al.
Published: (2023)
Autonomation, Not Automation: Activities and Needs of European Fact-checkers as a Basis for Designing Human-Centered AI Systems
by: Hrckova, Andrea, et al.
Published: (2022)
by: Hrckova, Andrea, et al.
Published: (2022)
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
by: Macko, Dominik, et al.
Published: (2024)
by: Macko, Dominik, et al.
Published: (2024)
Disinformation Capabilities of Large Language Models
by: Vykopal, Ivan, et al.
Published: (2023)
by: Vykopal, Ivan, et al.
Published: (2023)
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
by: Macko, Dominik, et al.
Published: (2026)
by: Macko, Dominik, et al.
Published: (2026)
Increasing the Robustness of the Fine-tuned Multilingual Machine-Generated Text Detectors
by: Macko, Dominik, et al.
Published: (2025)
by: Macko, Dominik, et al.
Published: (2025)
Political Leaning and Politicalness Classification of Texts
by: Volf, Matous, et al.
Published: (2025)
by: Volf, Matous, et al.
Published: (2025)
Beyond the Checkbox: Strengthening DSA Compliance Through Social Media Algorithmic Auditing
by: Solarova, Sara, et al.
Published: (2026)
by: Solarova, Sara, et al.
Published: (2026)
Enhancing Low-Resource LLMs Classification with PEFT and Synthetic Data
by: Patwa, Parth, et al.
Published: (2024)
by: Patwa, Parth, et al.
Published: (2024)
BLUFF: Benchmarking the Detection of False and Synthetic Content across 58 Low-Resource Languages
by: Lucas, Jason, et al.
Published: (2026)
by: Lucas, Jason, et al.
Published: (2026)
mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code
by: Skurla, Adam, et al.
Published: (2026)
by: Skurla, Adam, et al.
Published: (2026)
SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
Scaling Low-Resource MT via Synthetic Data Generation with LLMs
by: de Gibert, Ona, et al.
Published: (2025)
by: de Gibert, Ona, et al.
Published: (2025)
A Generative-AI-Driven Claim Retrieval System Capable of Detecting and Retrieving Claims from Social Media Platforms in Multiple Languages
by: Vykopal, Ivan, et al.
Published: (2025)
by: Vykopal, Ivan, et al.
Published: (2025)
Authorship Attribution in Multilingual Machine-Generated Texts
by: La Cava, Lucio, et al.
Published: (2025)
by: La Cava, Lucio, et al.
Published: (2025)
Better To Ask in English? Evaluating Factual Accuracy of Multilingual LLMs in English and Low-Resource Languages
by: Rohera, Pritika, et al.
Published: (2025)
by: Rohera, Pritika, et al.
Published: (2025)
Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain
by: Kaur, Rishemjit, et al.
Published: (2025)
by: Kaur, Rishemjit, et al.
Published: (2025)
Multiple Sources are Better Than One: Incorporating External Knowledge in Low-Resource Glossing
by: Yang, Changbing, et al.
Published: (2024)
by: Yang, Changbing, et al.
Published: (2024)
Adapting Multilingual LLMs to Low-Resource Languages using Continued Pre-training and Synthetic Corpus
by: Joshi, Raviraj, et al.
Published: (2024)
by: Joshi, Raviraj, et al.
Published: (2024)
Leveraging LLMs for Translating and Classifying Mental Health Data
by: Skianis, Konstantinos, et al.
Published: (2024)
by: Skianis, Konstantinos, et al.
Published: (2024)
YouTube Comments Decoded: Leveraging LLMs for Low Resource Language Classification
by: Deroy, Aniket, et al.
Published: (2024)
by: Deroy, Aniket, et al.
Published: (2024)
Similar Items
-
Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation
by: Pecher, Branislav, et al.
Published: (2024) -
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
by: Cegin, Jan, et al.
Published: (2025) -
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
by: Cegin, Jan, et al.
Published: (2024) -
Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
by: Cegin, Jan, et al.
Published: (2024) -
PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
by: Belanec, Robert, et al.
Published: (2025)