On Evaluation Protocols for Data Augmentation in a Limited Data Scenario
Fuente:
arXiv
Guardado en:
| Autores principales: | Piedboeuf, Frédéric, Langlais, Philippe |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
EUROPA: A Legal Multilingual Keyphrase Generation Dataset
por: Salaün, Olivier, et al.
Publicado: (2024)
por: Salaün, Olivier, et al.
Publicado: (2024)
Evaluating the Effectiveness of Data Augmentation for Emotion Classification in Low-Resource Settings
por: Arora, Aashish, et al.
Publicado: (2024)
por: Arora, Aashish, et al.
Publicado: (2024)
Retrieval-Augmented Data Augmentation for Low-Resource Domain Tasks
por: Seo, Minju, et al.
Publicado: (2024)
por: Seo, Minju, et al.
Publicado: (2024)
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
por: Lopardo, Gianluigi, et al.
Publicado: (2022)
por: Lopardo, Gianluigi, et al.
Publicado: (2022)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
por: Jadon, Aryan, et al.
Publicado: (2025)
por: Jadon, Aryan, et al.
Publicado: (2025)
Diversity-oriented Data Augmentation with Large Language Models
por: Wang, Zaitian, et al.
Publicado: (2025)
por: Wang, Zaitian, et al.
Publicado: (2025)
Labrador: Exploring the Limits of Masked Language Modeling for Laboratory Data
por: Bellamy, David R., et al.
Publicado: (2023)
por: Bellamy, David R., et al.
Publicado: (2023)
A Statistical Framework for Data-dependent Retrieval-Augmented Models
por: Basu, Soumya, et al.
Publicado: (2024)
por: Basu, Soumya, et al.
Publicado: (2024)
Bolster Hallucination Detection via Prompt-Guided Data Augmentation
por: Li, Wenyun, et al.
Publicado: (2025)
por: Li, Wenyun, et al.
Publicado: (2025)
ATLAS: Autoformalizing Theorems through Lifting, Augmentation, and Synthesis of Data
por: Liu, Xiaoyang, et al.
Publicado: (2025)
por: Liu, Xiaoyang, et al.
Publicado: (2025)
UniTabE: A Universal Pretraining Protocol for Tabular Foundation Model in Data Science
por: Yang, Yazheng, et al.
Publicado: (2023)
por: Yang, Yazheng, et al.
Publicado: (2023)
The ALCHEmist: Automated Labeling 500x CHEaper Than LLM Data Annotators
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
por: Huang, Tzu-Heng, et al.
Publicado: (2024)
Are Data Augmentation Methods in Named Entity Recognition Applicable for Uncertainty Estimation?
por: Hashimoto, Wataru, et al.
Publicado: (2024)
por: Hashimoto, Wataru, et al.
Publicado: (2024)
Forging the Forger: An Attempt to Improve Authorship Verification via Data Augmentation
por: Corbara, Silvia, et al.
Publicado: (2024)
por: Corbara, Silvia, et al.
Publicado: (2024)
Self-Supervised Time-Series Anomaly Detection Using Learnable Data Augmentation
por: Choi, Kukjin, et al.
Publicado: (2024)
por: Choi, Kukjin, et al.
Publicado: (2024)
Understanding and Mitigating Bias Inheritance in LLM-based Data Augmentation on Downstream Tasks
por: Li, Miaomiao, et al.
Publicado: (2025)
por: Li, Miaomiao, et al.
Publicado: (2025)
Prompting Test-Time Scaling Is A Strong LLM Reasoning Data Augmentation
por: Bsharat, Sondos Mahmoud, et al.
Publicado: (2025)
por: Bsharat, Sondos Mahmoud, et al.
Publicado: (2025)
Reasoning-Driven Synthetic Data Generation and Evaluation
por: Davidson, Tim R., et al.
Publicado: (2026)
por: Davidson, Tim R., et al.
Publicado: (2026)
BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data
por: Zou, Bo, et al.
Publicado: (2026)
por: Zou, Bo, et al.
Publicado: (2026)
Database Entity Recognition with Data Augmentation and Deep Learning
por: Fu, Zikun, et al.
Publicado: (2025)
por: Fu, Zikun, et al.
Publicado: (2025)
On Sensitivity of Learning with Limited Labelled Data to the Effects of Randomness: Impact of Interactions and Systematic Choices
por: Pecher, Branislav, et al.
Publicado: (2024)
por: Pecher, Branislav, et al.
Publicado: (2024)
Fine-tuning Large Language Models with Limited Data: A Survey and Practical Guide
por: Szep, Marton, et al.
Publicado: (2024)
por: Szep, Marton, et al.
Publicado: (2024)
A Survey on Stability of Learning with Limited Labelled Data and its Sensitivity to the Effects of Randomness
por: Pecher, Branislav, et al.
Publicado: (2023)
por: Pecher, Branislav, et al.
Publicado: (2023)
LLLMs: A Data-Driven Survey of Evolving Research on Limitations of Large Language Models
por: Kostikova, Aida, et al.
Publicado: (2025)
por: Kostikova, Aida, et al.
Publicado: (2025)
Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
por: Jin, Kyohoon, et al.
Publicado: (2024)
por: Jin, Kyohoon, et al.
Publicado: (2024)
Exploring Response Uncertainty in MLLMs: An Empirical Evaluation under Misleading Scenarios
por: Dang, Yunkai, et al.
Publicado: (2024)
por: Dang, Yunkai, et al.
Publicado: (2024)
ToxiGAN: Toxic Data Augmentation via LLM-Guided Directional Adversarial Generation
por: Li, Peiran, et al.
Publicado: (2026)
por: Li, Peiran, et al.
Publicado: (2026)
LEIA: Facilitating Cross-lingual Knowledge Transfer in Language Models with Entity-based Data Augmentation
por: Yamada, Ikuya, et al.
Publicado: (2024)
por: Yamada, Ikuya, et al.
Publicado: (2024)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
por: Yang, Shiping, et al.
Publicado: (2025)
por: Yang, Shiping, et al.
Publicado: (2025)
Training and Evaluating Language Models with Template-based Data Generation
por: Zhang, Yifan
Publicado: (2024)
por: Zhang, Yifan
Publicado: (2024)
Retrieval-Augmented Generation Meets Data-Driven Tabula Rasa Approach for Temporal Knowledge Graph Forecasting
por: Sannidhi, Geethan, et al.
Publicado: (2024)
por: Sannidhi, Geethan, et al.
Publicado: (2024)
MLB: A Scenario-Driven Benchmark for Evaluating Large Language Models in Clinical Applications
por: He, Qing, et al.
Publicado: (2026)
por: He, Qing, et al.
Publicado: (2026)
Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs
por: Jan, Essa, et al.
Publicado: (2025)
por: Jan, Essa, et al.
Publicado: (2025)
Evaluating Tool-Augmented Agents in Remote Sensing Platforms
por: Singh, Simranjit, et al.
Publicado: (2024)
por: Singh, Simranjit, et al.
Publicado: (2024)
The Limits of Preference Data for Post-Training
por: Zhao, Eric, et al.
Publicado: (2025)
por: Zhao, Eric, et al.
Publicado: (2025)
Evaluating the Limits of Large Language Models in Multilingual Legal Reasoning
por: Ioannou, Antreas, et al.
Publicado: (2025)
por: Ioannou, Antreas, et al.
Publicado: (2025)
Evaluating GPT's Capability in Identifying Stages of Cognitive Impairment from Electronic Health Data
por: Leng, Yu, et al.
Publicado: (2025)
por: Leng, Yu, et al.
Publicado: (2025)
XL-Suite: Cross-Lingual Synthetic Training and Evaluation Data for Open-Ended Generation
por: Iyer, Vivek, et al.
Publicado: (2025)
por: Iyer, Vivek, et al.
Publicado: (2025)
R&B: Domain Regrouping and Data Mixture Balancing for Efficient Foundation Model Training
por: Ge, Albert, et al.
Publicado: (2025)
por: Ge, Albert, et al.
Publicado: (2025)
From Static Benchmarks to Dynamic Protocol: Agent-Centric Text Anomaly Detection for Evaluating LLM Reasoning
por: Yoa, Seungdong, et al.
Publicado: (2026)
por: Yoa, Seungdong, et al.
Publicado: (2026)
Ejemplares similares
-
EUROPA: A Legal Multilingual Keyphrase Generation Dataset
por: Salaün, Olivier, et al.
Publicado: (2024) -
Evaluating the Effectiveness of Data Augmentation for Emotion Classification in Low-Resource Settings
por: Arora, Aashish, et al.
Publicado: (2024) -
Retrieval-Augmented Data Augmentation for Low-Resource Domain Tasks
por: Seo, Minju, et al.
Publicado: (2024) -
A Sea of Words: An In-Depth Analysis of Anchors for Text Data
por: Lopardo, Gianluigi, et al.
Publicado: (2022) -
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
por: Jadon, Aryan, et al.
Publicado: (2025)