Data Augmentations for Improved (Large) Language Model Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Feder, Amir, Wald, Yoav, Shi, Claudia, Saria, Suchi, Blei, David |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Open-Set Domain Adaptation Under Background Distribution Shift: Challenges and A Provably Efficient Solution
by: Chaudhari, Shravan, et al.
Published: (2025)
by: Chaudhari, Shravan, et al.
Published: (2025)
MILM: Large Language Models for Multimodal Irregular Time Series with Informative Sampling
by: Chung, Hsing-Huan, et al.
Published: (2026)
by: Chung, Hsing-Huan, et al.
Published: (2026)
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
by: Zheng, Carolina, et al.
Published: (2025)
by: Zheng, Carolina, et al.
Published: (2025)
Multi-environment Topic Models
by: Sobhani, Dominic, et al.
Published: (2024)
by: Sobhani, Dominic, et al.
Published: (2024)
Between Linear and Sinusoidal: Rethinking the Time Encoder in Dynamic Graph Learning
by: Chung, Hsing-Huan, et al.
Published: (2025)
by: Chung, Hsing-Huan, et al.
Published: (2025)
Understanding Understanding: A Pragmatic Framework Motivated by Large Language Models
by: Leyton-Brown, Kevin, et al.
Published: (2024)
by: Leyton-Brown, Kevin, et al.
Published: (2024)
Can Generative AI Solve Your In-Context Learning Problem? A Martingale Perspective
by: Jesson, Andrew, et al.
Published: (2024)
by: Jesson, Andrew, et al.
Published: (2024)
Integrating Large Language Models and Reinforcement Learning for Non-Linear Reasoning
by: Alon, Yoav, et al.
Published: (2024)
by: Alon, Yoav, et al.
Published: (2024)
Confidence-Aware Sub-Structure Beam Search (CABS): Mitigating Hallucination in Structured Data Generation with Large Language Models
by: Wei, Chengwei, et al.
Published: (2024)
by: Wei, Chengwei, et al.
Published: (2024)
NANOGPT: A Query-Driven Large Language Model Retrieval-Augmented Generation System for Nanotechnology Research
by: Chandrasekhar, Achuth, et al.
Published: (2025)
by: Chandrasekhar, Achuth, et al.
Published: (2025)
Augmenting Human-Annotated Training Data with Large Language Model Generation and Distillation in Open-Response Assessment
by: Borchers, Conrad, et al.
Published: (2025)
by: Borchers, Conrad, et al.
Published: (2025)
Diversity-oriented Data Augmentation with Large Language Models
by: Wang, Zaitian, et al.
Published: (2025)
by: Wang, Zaitian, et al.
Published: (2025)
On the Invariance and Generality of Neural Scaling Laws
by: Han, Xing, et al.
Published: (2026)
by: Han, Xing, et al.
Published: (2026)
Conformal Validity Guarantees Exist for Any Data Distribution (and How to Find Them)
by: Prinster, Drew, et al.
Published: (2024)
by: Prinster, Drew, et al.
Published: (2024)
Large Language Models for Document-Level Event-Argument Data Augmentation for Challenging Role Types
by: Gatto, Joseph, et al.
Published: (2024)
by: Gatto, Joseph, et al.
Published: (2024)
Causal Graph Discovery with Retrieval-Augmented Generation based Large Language Models
by: Zhang, Yuzhe, et al.
Published: (2024)
by: Zhang, Yuzhe, et al.
Published: (2024)
Improving Socratic Question Generation using Data Augmentation and Preference Optimization
by: Kumar, Nischal Ashok, et al.
Published: (2024)
by: Kumar, Nischal Ashok, et al.
Published: (2024)
Evaluating the Performance of Large Language Models for SDG Mapping (Technical Report)
by: Yin, Hui, et al.
Published: (2024)
by: Yin, Hui, et al.
Published: (2024)
Listen to the Context: Towards Faithful Large Language Models for Retrieval Augmented Generation on Climate Questions
by: Thulke, David, et al.
Published: (2025)
by: Thulke, David, et al.
Published: (2025)
Reasoning in Trees: Improving Retrieval-Augmented Generation for Multi-Hop Question Answering
by: Shi, Yuling, et al.
Published: (2026)
by: Shi, Yuling, et al.
Published: (2026)
Interactive and Expressive Code-Augmented Planning with Large Language Models
by: Liu, Anthony Z., et al.
Published: (2024)
by: Liu, Anthony Z., et al.
Published: (2024)
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
by: Cohen, Amir DN, et al.
Published: (2024)
by: Cohen, Amir DN, et al.
Published: (2024)
Min-K%++: Improved Baseline for Detecting Pre-Training Data from Large Language Models
by: Zhang, Jingyang, et al.
Published: (2024)
by: Zhang, Jingyang, et al.
Published: (2024)
Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
by: Cheng, Zicong, et al.
Published: (2026)
by: Cheng, Zicong, et al.
Published: (2026)
Nostra Domina at EvaLatin 2024: Improving Latin Polarity Detection through Data Augmentation
by: Bothwell, Stephen, et al.
Published: (2024)
by: Bothwell, Stephen, et al.
Published: (2024)
Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models
by: Li, Zhaoyi, et al.
Published: (2026)
by: Li, Zhaoyi, et al.
Published: (2026)
Lightweight Retrieval-Augmented Generation and Large Language Model-Based Modeling for Scalable Patient-Trial Matching
by: Li, Xiaodi, et al.
Published: (2026)
by: Li, Xiaodi, et al.
Published: (2026)
Language Models Improve When Pretraining Data Matches Target Tasks
by: Mizrahi, David, et al.
Published: (2025)
by: Mizrahi, David, et al.
Published: (2025)
Improving Coverage in Combined Prediction Sets with Weighted p-values
by: Wong, Gina, et al.
Published: (2025)
by: Wong, Gina, et al.
Published: (2025)
Unlocking the Potentials of Retrieval-Augmented Generation for Diffusion Language Models
by: Yu, Chuanyue, et al.
Published: (2026)
by: Yu, Chuanyue, et al.
Published: (2026)
Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
by: Yang, Shiping, et al.
Published: (2025)
by: Yang, Shiping, et al.
Published: (2025)
Personalizing Large Language Models using Retrieval Augmented Generation and Knowledge Graph
by: Prahlad, Deeksha, et al.
Published: (2025)
by: Prahlad, Deeksha, et al.
Published: (2025)
SimRAG: Self-Improving Retrieval-Augmented Generation for Adapting Large Language Models to Specialized Domains
by: Xu, Ran, et al.
Published: (2024)
by: Xu, Ran, et al.
Published: (2024)
Nuisances via Negativa: Adjusting for Spurious Correlations via Data Augmentation
by: Puli, Aahlad, et al.
Published: (2022)
by: Puli, Aahlad, et al.
Published: (2022)
Progressive-Hint Prompting Improves Reasoning in Large Language Models
by: Zheng, Chuanyang, et al.
Published: (2023)
by: Zheng, Chuanyang, et al.
Published: (2023)
Negation Blindness in Large Language Models: Unveiling the NO Syndrome in Image Generation
by: Nadeem, Mohammad, et al.
Published: (2024)
by: Nadeem, Mohammad, et al.
Published: (2024)
Large Language Model Agent as a Mechanical Designer
by: Jadhav, Yayati, et al.
Published: (2024)
by: Jadhav, Yayati, et al.
Published: (2024)
AMGPT: a Large Language Model for Contextual Querying in Additive Manufacturing
by: Chandrasekhar, Achuth, et al.
Published: (2024)
by: Chandrasekhar, Achuth, et al.
Published: (2024)
Branch-Solve-Merge Improves Large Language Model Evaluation and Generation
by: Saha, Swarnadeep, et al.
Published: (2023)
by: Saha, Swarnadeep, et al.
Published: (2023)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
by: Rubashevskii, Aleksandr, et al.
Published: (2026)
by: Rubashevskii, Aleksandr, et al.
Published: (2026)
Similar Items
-
Open-Set Domain Adaptation Under Background Distribution Shift: Challenges and A Provably Efficient Solution
by: Chaudhari, Shravan, et al.
Published: (2025) -
MILM: Large Language Models for Multimodal Irregular Time Series with Informative Sampling
by: Chung, Hsing-Huan, et al.
Published: (2026) -
Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders
by: Zheng, Carolina, et al.
Published: (2025) -
Multi-environment Topic Models
by: Sobhani, Dominic, et al.
Published: (2024) -
Between Linear and Sinusoidal: Rethinking the Time Encoder in Dynamic Graph Learning
by: Chung, Hsing-Huan, et al.
Published: (2025)