Saved in:
| Main Authors: | Ignashina, Mariia, Ive, Julia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2404.19486 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Assistance for Pediatric Depression
by: Ignashina, Mariia, et al.
Published: (2025)
by: Ignashina, Mariia, et al.
Published: (2025)
Privacy-Preserving Behaviour of Chatbot Users: Steering Through Trust Dynamics
by: Ive, Julia, et al.
Published: (2024)
by: Ive, Julia, et al.
Published: (2024)
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
by: Kern, Christoph, et al.
Published: (2023)
by: Kern, Christoph, et al.
Published: (2023)
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
by: Nakka, Krishna Kanth, et al.
Published: (2024)
by: Nakka, Krishna Kanth, et al.
Published: (2024)
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
by: Chen, Howard, et al.
Published: (2025)
by: Chen, Howard, et al.
Published: (2025)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
by: Shetty, Pranav, et al.
Published: (2025)
by: Shetty, Pranav, et al.
Published: (2025)
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training
by: Tran, Toan, et al.
Published: (2025)
by: Tran, Toan, et al.
Published: (2025)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
by: Gadhikar, Advait, et al.
Published: (2025)
by: Gadhikar, Advait, et al.
Published: (2025)
Leveraging Big Data Frameworks for Spam Detection in Amazon Reviews
by: Khatun, Mst Eshita, et al.
Published: (2025)
by: Khatun, Mst Eshita, et al.
Published: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
by: Belenki, Lior, et al.
Published: (2025)
by: Belenki, Lior, et al.
Published: (2025)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
by: Guo, Wenkai, et al.
Published: (2025)
by: Guo, Wenkai, et al.
Published: (2025)
Data Shapley in One Training Run
by: Wang, Jiachen T., et al.
Published: (2024)
by: Wang, Jiachen T., et al.
Published: (2024)
On Training Data Influence of GPT Models
by: Chai, Yekun, et al.
Published: (2024)
by: Chai, Yekun, et al.
Published: (2024)
Synthetic Data Generation for Intersectional Fairness by Leveraging Hierarchical Group Structure
by: Maheshwari, Gaurav, et al.
Published: (2024)
by: Maheshwari, Gaurav, et al.
Published: (2024)
Prescriptive Scaling Laws for Data Constrained Training
by: Lovelace, Justin, et al.
Published: (2026)
by: Lovelace, Justin, et al.
Published: (2026)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
by: Hayase, Jonathan, et al.
Published: (2024)
by: Hayase, Jonathan, et al.
Published: (2024)
Financial Sentiment Analysis: Leveraging Actual and Synthetic Data for Supervised Fine-tuning
by: Atsiwo, Abraham
Published: (2024)
by: Atsiwo, Abraham
Published: (2024)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
by: Soto, Rafael Rivera, et al.
Published: (2024)
by: Soto, Rafael Rivera, et al.
Published: (2024)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
by: He, Luxi, et al.
Published: (2024)
by: He, Luxi, et al.
Published: (2024)
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
by: Iskander, Shadi, et al.
Published: (2024)
by: Iskander, Shadi, et al.
Published: (2024)
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
Unsupervised Data Validation Methods for Efficient Model Training
by: Paniv, Yurii
Published: (2024)
by: Paniv, Yurii
Published: (2024)
Training Bilingual LMs with Data Constraints in the Targeted Language
by: Seto, Skyler, et al.
Published: (2024)
by: Seto, Skyler, et al.
Published: (2024)
ZClip: Adaptive Spike Mitigation for LLM Pre-Training
by: Kumar, Abhay, et al.
Published: (2025)
by: Kumar, Abhay, et al.
Published: (2025)
Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
by: Eiras, Francisco, et al.
Published: (2024)
by: Eiras, Francisco, et al.
Published: (2024)
MALTO at SemEval-2024 Task 6: Leveraging Synthetic Data for LLM Hallucination Detection
by: Borra, Federico, et al.
Published: (2024)
by: Borra, Federico, et al.
Published: (2024)
ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
by: Zhang, Guangwei, et al.
Published: (2025)
by: Zhang, Guangwei, et al.
Published: (2025)
Mitigating Data Imbalance in Automated Speaking Assessment
by: Tsai, Fong-Chun, et al.
Published: (2025)
by: Tsai, Fong-Chun, et al.
Published: (2025)
Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
by: Bai, Yang, et al.
Published: (2024)
by: Bai, Yang, et al.
Published: (2024)
Leveraging Variation Theory in Counterfactual Data Augmentation for Optimized Active Learning
by: Gebreegziabher, Simret Araya, et al.
Published: (2024)
by: Gebreegziabher, Simret Araya, et al.
Published: (2024)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
by: Zhu, Banghua, et al.
Published: (2024)
by: Zhu, Banghua, et al.
Published: (2024)
How to Train Data-Efficient LLMs
by: Sachdeva, Noveen, et al.
Published: (2024)
by: Sachdeva, Noveen, et al.
Published: (2024)
Reinforcement Learning on Pre-Training Data
by: Li, Siheng, et al.
Published: (2025)
by: Li, Siheng, et al.
Published: (2025)
How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
by: Xiao, Teng, et al.
Published: (2024)
by: Xiao, Teng, et al.
Published: (2024)
A Natural Language Processing Approach to Support Biomedical Data Harmonization: Leveraging Large Language Models
by: Li, Zexu, et al.
Published: (2024)
by: Li, Zexu, et al.
Published: (2024)
QuRating: Selecting High-Quality Data for Training Language Models
by: Wettig, Alexander, et al.
Published: (2024)
by: Wettig, Alexander, et al.
Published: (2024)
Sequence-Level Leakage Risk of Training Data in Large Language Models
by: Tiwari, Trishita, et al.
Published: (2024)
by: Tiwari, Trishita, et al.
Published: (2024)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
by: Aerni, Michael, et al.
Published: (2024)
by: Aerni, Michael, et al.
Published: (2024)
Efficient RLVR Training via Weighted Mutual Information Data Selection
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
by: Liang, Hao, et al.
Published: (2026)
by: Liang, Hao, et al.
Published: (2026)
Similar Items
-
LLM Assistance for Pediatric Depression
by: Ignashina, Mariia, et al.
Published: (2025) -
Privacy-Preserving Behaviour of Chatbot Users: Steering Through Trust Dynamics
by: Ive, Julia, et al.
Published: (2024) -
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
by: Kern, Christoph, et al.
Published: (2023) -
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
by: Nakka, Krishna Kanth, et al.
Published: (2024) -
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
by: Chen, Howard, et al.
Published: (2025)