Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cegin, Jan, Pecher, Branislav, Simko, Jakub, Srba, Ivan, Bielikova, Maria, Brusilovsky, Peter |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
von: Cegin, Jan, et al.
Veröffentlicht: (2024)
von: Cegin, Jan, et al.
Veröffentlicht: (2024)
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
von: Cegin, Jan, et al.
Veröffentlicht: (2025)
von: Cegin, Jan, et al.
Veröffentlicht: (2025)
Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation
von: Pecher, Branislav, et al.
Veröffentlicht: (2024)
von: Pecher, Branislav, et al.
Veröffentlicht: (2024)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness
von: Pecher, Branislav, et al.
Veröffentlicht: (2023)
von: Pecher, Branislav, et al.
Veröffentlicht: (2023)
On Sensitivity of Learning with Limited Labelled Data to the Effects of Randomness: Impact of Interactions and Systematic Choices
von: Pecher, Branislav, et al.
Veröffentlicht: (2024)
von: Pecher, Branislav, et al.
Veröffentlicht: (2024)
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
von: Cegin, Jan, et al.
Veröffentlicht: (2024)
von: Cegin, Jan, et al.
Veröffentlicht: (2024)
Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance
von: Pecher, Branislav, et al.
Veröffentlicht: (2024)
von: Pecher, Branislav, et al.
Veröffentlicht: (2024)
PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
von: Belanec, Robert, et al.
Veröffentlicht: (2025)
von: Belanec, Robert, et al.
Veröffentlicht: (2025)
Automatic Combination of Sample Selection Strategies for Few-Shot Learning
von: Pecher, Branislav, et al.
Veröffentlicht: (2024)
von: Pecher, Branislav, et al.
Veröffentlicht: (2024)
MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models
von: Hyben, Martin, et al.
Veröffentlicht: (2026)
von: Hyben, Martin, et al.
Veröffentlicht: (2026)
A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
von: Anikina, Tatiana, et al.
Veröffentlicht: (2025)
von: Anikina, Tatiana, et al.
Veröffentlicht: (2025)
Revisiting Prompt Sensitivity in Large Language Models for Text Classification: The Role of Prompt Underspecification
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
von: Pecher, Branislav, et al.
Veröffentlicht: (2026)
PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
von: Belanec, Robert, et al.
Veröffentlicht: (2025)
von: Belanec, Robert, et al.
Veröffentlicht: (2025)
Autonomation, Not Automation: Activities and Needs of European Fact-checkers as a Basis for Designing Human-Centered AI Systems
von: Hrckova, Andrea, et al.
Veröffentlicht: (2022)
von: Hrckova, Andrea, et al.
Veröffentlicht: (2022)
KInITVeraAI at SemEval-2023 Task 3: Simple yet Powerful Multilingual Fine-Tuning for Persuasion Techniques Detection
von: Hromadka, Timo, et al.
Veröffentlicht: (2023)
von: Hromadka, Timo, et al.
Veröffentlicht: (2023)
Multilingual and Multi-topical Benchmark of Fine-tuned Language models and Large Language Models for Check-Worthy Claim Detection
von: Hyben, Martin, et al.
Veröffentlicht: (2023)
von: Hyben, Martin, et al.
Veröffentlicht: (2023)
Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer
von: Belanec, Robert, et al.
Veröffentlicht: (2024)
von: Belanec, Robert, et al.
Veröffentlicht: (2024)
Multilingual Previously Fact-Checked Claim Retrieval
von: Pikuliak, Matúš, et al.
Veröffentlicht: (2023)
von: Pikuliak, Matúš, et al.
Veröffentlicht: (2023)
Authorship Obfuscation in Multilingual Machine-Generated Text Detection
von: Macko, Dominik, et al.
Veröffentlicht: (2024)
von: Macko, Dominik, et al.
Veröffentlicht: (2024)
Disinformation Capabilities of Large Language Models
von: Vykopal, Ivan, et al.
Veröffentlicht: (2023)
von: Vykopal, Ivan, et al.
Veröffentlicht: (2023)
Beyond the Checkbox: Strengthening DSA Compliance Through Social Media Algorithmic Auditing
von: Solarova, Sara, et al.
Veröffentlicht: (2026)
von: Solarova, Sara, et al.
Veröffentlicht: (2026)
MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark
von: Macko, Dominik, et al.
Veröffentlicht: (2023)
von: Macko, Dominik, et al.
Veröffentlicht: (2023)
Political Leaning and Politicalness Classification of Texts
von: Volf, Matous, et al.
Veröffentlicht: (2025)
von: Volf, Matous, et al.
Veröffentlicht: (2025)
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
von: Macko, Dominik, et al.
Veröffentlicht: (2024)
von: Macko, Dominik, et al.
Veröffentlicht: (2024)
Beyond speculation: Measuring the growing presence of LLM-generated texts in multilingual disinformation
von: Macko, Dominik, et al.
Veröffentlicht: (2025)
von: Macko, Dominik, et al.
Veröffentlicht: (2025)
Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation
von: Zugecova, Aneta, et al.
Veröffentlicht: (2024)
von: Zugecova, Aneta, et al.
Veröffentlicht: (2024)
A Generative-AI-Driven Claim Retrieval System Capable of Detecting and Retrieving Claims from Social Media Platforms in Multiple Languages
von: Vykopal, Ivan, et al.
Veröffentlicht: (2025)
von: Vykopal, Ivan, et al.
Veröffentlicht: (2025)
Do LLMs produce texts with "human-like" lexical diversity?
von: Kendro, Kelly, et al.
Veröffentlicht: (2025)
von: Kendro, Kelly, et al.
Veröffentlicht: (2025)
mdok-style at SemEval-2026 Task 9: Finetuning LLMs for Multilingual Polarization Detection
von: Macko, Dominik, et al.
Veröffentlicht: (2026)
von: Macko, Dominik, et al.
Veröffentlicht: (2026)
Performance of diverse evaluation metrics in NLP-based assessment and text generation of consumer complaints
von: Gao, Peiheng, et al.
Veröffentlicht: (2025)
von: Gao, Peiheng, et al.
Veröffentlicht: (2025)
Interpretable Predictability-Based AI Text Detection: A Replication Study
von: Skurla, Adam, et al.
Veröffentlicht: (2026)
von: Skurla, Adam, et al.
Veröffentlicht: (2026)
Soft Language Prompts for Language Transfer
von: Vykopal, Ivan, et al.
Veröffentlicht: (2024)
von: Vykopal, Ivan, et al.
Veröffentlicht: (2024)
Evaluating the capability of large language models to personalize science texts for diverse middle-school-age learners
von: Vaccaro Jr, Michael, et al.
Veröffentlicht: (2024)
von: Vaccaro Jr, Michael, et al.
Veröffentlicht: (2024)
mcdok at SemEval-2026 Task 13: Finetuning LLMs for Detection of Machine-Generated Code
von: Skurla, Adam, et al.
Veröffentlicht: (2026)
von: Skurla, Adam, et al.
Veröffentlicht: (2026)
Increasing the Robustness of the Fine-tuned Multilingual Machine-Generated Text Detectors
von: Macko, Dominik, et al.
Veröffentlicht: (2025)
von: Macko, Dominik, et al.
Veröffentlicht: (2025)
SemEval-2025 Task 7: Multilingual and Crosslingual Fact-Checked Claim Retrieval
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
von: Peng, Qiwei, et al.
Veröffentlicht: (2025)
Social Reasoning in Machines: Investigating Collective Truth-Seeking Dynamics in Large Language Model Debate
von: Pecher, Tom
Veröffentlicht: (2026)
von: Pecher, Tom
Veröffentlicht: (2026)
Assessing Web Search Credibility and Response Groundedness in Chat Assistants
von: Vykopal, Ivan, et al.
Veröffentlicht: (2025)
von: Vykopal, Ivan, et al.
Veröffentlicht: (2025)
Generative Large Language Models in Automated Fact-Checking: A Survey
von: Vykopal, Ivan, et al.
Veröffentlicht: (2024)
von: Vykopal, Ivan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
von: Cegin, Jan, et al.
Veröffentlicht: (2024) -
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
von: Cegin, Jan, et al.
Veröffentlicht: (2025) -
Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation
von: Pecher, Branislav, et al.
Veröffentlicht: (2024) -
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
von: Pecher, Branislav, et al.
Veröffentlicht: (2026) -
A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness
von: Pecher, Branislav, et al.
Veröffentlicht: (2023)