RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
Fuente:
arXiv
Saved in:
| Main Authors: | Cegin, Jan, Pecher, Branislav, Srba, Ivan, Simko, Jakub |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
by: Cegin, Jan, et al.
Published: (2024)
by: Cegin, Jan, et al.
Published: (2024)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
by: Pecher, Branislav, et al.
Published: (2026)
by: Pecher, Branislav, et al.
Published: (2026)
Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
by: Cegin, Jan, et al.
Published: (2024)
by: Cegin, Jan, et al.
Published: (2024)
Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
by: Anikina, Tatiana, et al.
Published: (2025)
by: Anikina, Tatiana, et al.
Published: (2025)
MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models
by: Hyben, Martin, et al.
Published: (2026)
by: Hyben, Martin, et al.
Published: (2026)
A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness
by: Pecher, Branislav, et al.
Published: (2023)
by: Pecher, Branislav, et al.
Published: (2023)
On Sensitivity of Learning with Limited Labelled Data to the Effects of Randomness: Impact of Interactions and Systematic Choices
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
by: Cegin, Jan, et al.
Published: (2024)
by: Cegin, Jan, et al.
Published: (2024)
Automatic Combination of Sample Selection Strategies for Few-Shot Learning
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
Revisiting Prompt Sensitivity in Large Language Models for Text Classification: The Role of Prompt Underspecification
by: Pecher, Branislav, et al.
Published: (2026)
by: Pecher, Branislav, et al.
Published: (2026)
PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
by: Belanec, Robert, et al.
Published: (2025)
by: Belanec, Robert, et al.
Published: (2025)
KInITVeraAI at SemEval-2023 Task 3: Simple yet Powerful Multilingual Fine-Tuning for Persuasion Techniques Detection
by: Hromadka, Timo, et al.
Published: (2023)
by: Hromadka, Timo, et al.
Published: (2023)
Multilingual and Multi-topical Benchmark of Fine-tuned Language models and Large Language Models for Check-Worthy Claim Detection
by: Hyben, Martin, et al.
Published: (2023)
by: Hyben, Martin, et al.
Published: (2023)
Beyond the Checkbox: Strengthening DSA Compliance Through Social Media Algorithmic Auditing
by: Solarova, Sara, et al.
Published: (2026)
by: Solarova, Sara, et al.
Published: (2026)
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
by: Zugecova, Aneta, et al.
Published: (2024)
by: Zugecova, Aneta, et al.
Published: (2024)
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
by: Macko, Dominik, et al.
Published: (2024)
by: Macko, Dominik, et al.
Published: (2024)
Autonomation, Not Automation: Activities and Needs of European Fact-checkers as a Basis for Designing Human-Centered AI Systems
by: Hrckova, Andrea, et al.
Published: (2022)
by: Hrckova, Andrea, et al.
Published: (2022)
Authorship Obfuscation in Multilingual Machine-Generated Text Detection
by: Macko, Dominik, et al.
Published: (2024)
by: Macko, Dominik, et al.
Published: (2024)
Multilingual Previously Fact-Checked Claim Retrieval
by: Pikuliak, Matúš, et al.
Published: (2023)
by: Pikuliak, Matúš, et al.
Published: (2023)
The Effect of Human v/s Synthetic Test Data and Round-tripping on Assessment of Sentiment Analysis Systems for Bias
by: Lakkaraju, Kausik, et al.
Published: (2024)
by: Lakkaraju, Kausik, et al.
Published: (2024)
Political Leaning and Politicalness Classification of Texts
by: Volf, Matous, et al.
Published: (2025)
by: Volf, Matous, et al.
Published: (2025)
A Generative-AI-Driven Claim Retrieval System Capable of Detecting and Retrieving Claims from Social Media Platforms in Multiple Languages
by: Vykopal, Ivan, et al.
Published: (2025)
by: Vykopal, Ivan, et al.
Published: (2025)
MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark
by: Macko, Dominik, et al.
Published: (2023)
by: Macko, Dominik, et al.
Published: (2023)
Increasing the Robustness of the Fine-tuned Multilingual Machine-Generated Text Detectors
by: Macko, Dominik, et al.
Published: (2025)
by: Macko, Dominik, et al.
Published: (2025)
mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code
by: Skurla, Adam, et al.
Published: (2026)
by: Skurla, Adam, et al.
Published: (2026)
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
by: Macko, Dominik, et al.
Published: (2026)
by: Macko, Dominik, et al.
Published: (2026)
Evaluating LLM-Generated Q&A Test: a Student-Centered Study
by: Wróblewska, Anna, et al.
Published: (2025)
by: Wróblewska, Anna, et al.
Published: (2025)
Test Set Quality in Multilingual LLM Evaluation
by: Kranti, Chalamalasetti, et al.
Published: (2025)
by: Kranti, Chalamalasetti, et al.
Published: (2025)
PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
by: Belanec, Robert, et al.
Published: (2025)
by: Belanec, Robert, et al.
Published: (2025)
Generative Large Language Models in Automated Fact-Checking: A Survey
by: Vykopal, Ivan, et al.
Published: (2024)
by: Vykopal, Ivan, et al.
Published: (2024)
Interpretable Predictability-Based AI Text Detection: A Replication Study
by: Skurla, Adam, et al.
Published: (2026)
by: Skurla, Adam, et al.
Published: (2026)
Soft Language Prompts for Language Transfer
by: Vykopal, Ivan, et al.
Published: (2024)
by: Vykopal, Ivan, et al.
Published: (2024)
Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation
by: Miranda, Lester James V., et al.
Published: (2026)
by: Miranda, Lester James V., et al.
Published: (2026)
A Ship of Theseus: Curious Cases of Paraphrasing in LLM-Generated Texts
by: Tripto, Nafis Irtiza, et al.
Published: (2023)
by: Tripto, Nafis Irtiza, et al.
Published: (2023)
CrowdSelect: Synthetic Instruction Data Selection with Multi-LLM Wisdom
by: Li, Yisen, et al.
Published: (2025)
by: Li, Yisen, et al.
Published: (2025)
Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer
by: Belanec, Robert, et al.
Published: (2024)
by: Belanec, Robert, et al.
Published: (2024)
Data Swarms: Optimizable Generation of Synthetic Evaluation Data
by: Feng, Shangbin, et al.
Published: (2025)
by: Feng, Shangbin, et al.
Published: (2025)
SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval
by: Peng, Qiwei, et al.
Published: (2025)
by: Peng, Qiwei, et al.
Published: (2025)
Similar Items
-
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
by: Cegin, Jan, et al.
Published: (2024) -
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
by: Pecher, Branislav, et al.
Published: (2026) -
Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
by: Cegin, Jan, et al.
Published: (2024) -
Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation
by: Pecher, Branislav, et al.
Published: (2024) -
A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
by: Anikina, Tatiana, et al.
Published: (2025)