Use Random Selection for Now: Investigation of Few-Shot Selection Strategies in LLM-based Text Augmentation for Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Cegin, Jan, Pecher, Branislav, Simko, Jakub, Srba, Ivan, Bielikova, Maria, Brusilovsky, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
by: Cegin, Jan, et al.
Published: (2024)
by: Cegin, Jan, et al.
Published: (2024)
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
by: Cegin, Jan, et al.
Published: (2025)
by: Cegin, Jan, et al.
Published: (2025)
Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
by: Pecher, Branislav, et al.
Published: (2026)
by: Pecher, Branislav, et al.
Published: (2026)
Automatic Combination of Sample Selection Strategies for Few-Shot Learning
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
LLMs vs Established Text Augmentation Techniques for Classification: When do the Benefits Outweight the Costs?
by: Cegin, Jan, et al.
Published: (2024)
by: Cegin, Jan, et al.
Published: (2024)
Comparing Specialised Small and General Large Language Models on Text Classification: 100 Labelled Samples to Achieve Break-Even Performance
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness
by: Pecher, Branislav, et al.
Published: (2023)
by: Pecher, Branislav, et al.
Published: (2023)
On Sensitivity of Learning with Limited Labelled Data to the Effects of Randomness: Impact of Interactions and Systematic Choices
by: Pecher, Branislav, et al.
Published: (2024)
by: Pecher, Branislav, et al.
Published: (2024)
PEFT-Bench: A Parameter-Efficient Fine-Tuning Methods Benchmark
by: Belanec, Robert, et al.
Published: (2025)
by: Belanec, Robert, et al.
Published: (2025)
Revisiting Prompt Sensitivity in Large Language Models for Text Classification: The Role of Prompt Underspecification
by: Pecher, Branislav, et al.
Published: (2026)
by: Pecher, Branislav, et al.
Published: (2026)
A Rigorous Evaluation of LLM Data Generation Strategies for Low-Resource Languages
by: Anikina, Tatiana, et al.
Published: (2025)
by: Anikina, Tatiana, et al.
Published: (2025)
MultiCW: A Large-Scale Balanced Benchmark Dataset for Training Robust Check-Worthiness Detection Models
by: Hyben, Martin, et al.
Published: (2026)
by: Hyben, Martin, et al.
Published: (2026)
PEFT-Factory: Unified Parameter-Efficient Fine-Tuning of Autoregressive Large Language Models
by: Belanec, Robert, et al.
Published: (2025)
by: Belanec, Robert, et al.
Published: (2025)
KInITVeraAI at SemEval-2023 Task 3: Simple yet Powerful Multilingual Fine-Tuning for Persuasion Techniques Detection
by: Hromadka, Timo, et al.
Published: (2023)
by: Hromadka, Timo, et al.
Published: (2023)
Political Leaning and Politicalness Classification of Texts
by: Volf, Matous, et al.
Published: (2025)
by: Volf, Matous, et al.
Published: (2025)
Authorship Obfuscation in Multilingual Machine-Generated Text Detection
by: Macko, Dominik, et al.
Published: (2024)
by: Macko, Dominik, et al.
Published: (2024)
Autonomation, Not Automation: Activities and Needs of European Fact-checkers as a Basis for Designing Human-Centered AI Systems
by: Hrckova, Andrea, et al.
Published: (2022)
by: Hrckova, Andrea, et al.
Published: (2022)
Multilingual and Multi-topical Benchmark of Fine-tuned Language models and Large Language Models for Check-Worthy Claim Detection
by: Hyben, Martin, et al.
Published: (2023)
by: Hyben, Martin, et al.
Published: (2023)
Multilingual Previously Fact-Checked Claim Retrieval
by: Pikuliak, Matúš, et al.
Published: (2023)
by: Pikuliak, Matúš, et al.
Published: (2023)
Task Prompt Vectors: Effective Initialization through Multi-Task Soft-Prompt Transfer
by: Belanec, Robert, et al.
Published: (2024)
by: Belanec, Robert, et al.
Published: (2024)
MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark
by: Macko, Dominik, et al.
Published: (2023)
by: Macko, Dominik, et al.
Published: (2023)
MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts
by: Macko, Dominik, et al.
Published: (2024)
by: Macko, Dominik, et al.
Published: (2024)
Beyond the Checkbox: Strengthening DSA Compliance Through Social Media Algorithmic Auditing
by: Solarova, Sara, et al.
Published: (2026)
by: Solarova, Sara, et al.
Published: (2026)
Disinformation Capabilities of Large Language Models
by: Vykopal, Ivan, et al.
Published: (2023)
by: Vykopal, Ivan, et al.
Published: (2023)
Enhancing LLM-Based Text Classification in Political Science: Automatic Prompt Optimization and Dynamic Exemplar Selection for Few-Shot Learning
by: Liu, Menglin, et al.
Published: (2024)
by: Liu, Menglin, et al.
Published: (2024)
Interpretable Predictability-Based AI Text Detection: A Replication Study
by: Skurla, Adam, et al.
Published: (2026)
by: Skurla, Adam, et al.
Published: (2026)
Label-template based Few-Shot Text Classification with Contrastive Learning
by: Hou, Guanghua, et al.
Published: (2024)
by: Hou, Guanghua, et al.
Published: (2024)
Active Few-Shot Learning for Text Classification
by: Ahmadnia, Saeed, et al.
Published: (2025)
by: Ahmadnia, Saeed, et al.
Published: (2025)
Mask-guided BERT for Few Shot Text Classification
by: Liao, Wenxiong, et al.
Published: (2023)
by: Liao, Wenxiong, et al.
Published: (2023)
Increasing the Robustness of the Fine-tuned Multilingual Machine-Generated Text Detectors
by: Macko, Dominik, et al.
Published: (2025)
by: Macko, Dominik, et al.
Published: (2025)
Towards Robust Few-Shot Text Classification Using Transformer Architectures and Dual Loss Strategies
by: Han, Xu, et al.
Published: (2025)
by: Han, Xu, et al.
Published: (2025)
Boosting Reinforcement Learning with Verifiable Rewards via Randomly Selected Few-Shot Guidance
by: Yan, Kai, et al.
Published: (2026)
by: Yan, Kai, et al.
Published: (2026)
Manual Verbalizer Enrichment for Few-Shot Text Classification
by: Nguyen, Quang Anh, et al.
Published: (2024)
by: Nguyen, Quang Anh, et al.
Published: (2024)
Designing Informative Metrics for Few-Shot Example Selection
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
Exploring Selective Retrieval-Augmentation for Long-Tail Legal Text Classification
by: Mao, Boheng
Published: (2025)
by: Mao, Boheng
Published: (2025)
Social Reasoning in Machines: Investigating Collective Truth-Seeking Dynamics in Large Language Model Debate
by: Pecher, Tom
Published: (2026)
by: Pecher, Tom
Published: (2026)
Few-Shot Fairness: Unveiling LLM's Potential for Fairness-Aware Classification
by: Chhikara, Garima, et al.
Published: (2024)
by: Chhikara, Garima, et al.
Published: (2024)
PolyNorm: Few-Shot LLM-Based Text Normalization for Text-to-Speech
by: Wong, Michel, et al.
Published: (2025)
by: Wong, Michel, et al.
Published: (2025)
Exploiting Text Semantics for Few and Zero Shot Node Classification on Text-attributed Graph
by: Wang, Yuxiang, et al.
Published: (2025)
by: Wang, Yuxiang, et al.
Published: (2025)
Similar Items
-
Effects of diversity incentives on sample diversity and downstream model performance in LLM-based text augmentation
by: Cegin, Jan, et al.
Published: (2024) -
RoSE: Round-robin Synthetic Data Evaluation for Selecting LLM Generators without Human Test Sets
by: Cegin, Jan, et al.
Published: (2025) -
Fighting Randomness with Randomness: Mitigating Optimisation Instability of Fine-Tuning using Delayed Ensemble and Noisy Interpolation
by: Pecher, Branislav, et al.
Published: (2024) -
Better as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
by: Pecher, Branislav, et al.
Published: (2026) -
Automatic Combination of Sample Selection Strategies for Few-Shot Learning
by: Pecher, Branislav, et al.
Published: (2024)