Gespeichert in:
| Hauptverfasser: | Ignashina, Mariia, Ive, Julia |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.19486 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM Assistance for Pediatric Depression
von: Ignashina, Mariia, et al.
Veröffentlicht: (2025)
von: Ignashina, Mariia, et al.
Veröffentlicht: (2025)
Privacy-Preserving Behaviour of Chatbot Users: Steering Through Trust Dynamics
von: Ive, Julia, et al.
Veröffentlicht: (2024)
von: Ive, Julia, et al.
Veröffentlicht: (2024)
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
von: Kern, Christoph, et al.
Veröffentlicht: (2023)
von: Kern, Christoph, et al.
Veröffentlicht: (2023)
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024)
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
von: Chen, Howard, et al.
Veröffentlicht: (2025)
von: Chen, Howard, et al.
Veröffentlicht: (2025)
Perturb Your Data: Paraphrase-Guided Training Data Watermarking
von: Shetty, Pranav, et al.
Veröffentlicht: (2025)
von: Shetty, Pranav, et al.
Veröffentlicht: (2025)
Tokens for Learning, Tokens for Unlearning: Mitigating Membership Inference Attacks in Large Language Models via Dual-Purpose Training
von: Tran, Toan, et al.
Veröffentlicht: (2025)
von: Tran, Toan, et al.
Veröffentlicht: (2025)
OptRot: Mitigating Weight Outliers via Data-Free Rotations for Post-Training Quantization
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
von: Gadhikar, Advait, et al.
Veröffentlicht: (2025)
Leveraging Big Data Frameworks for Spam Detection in Amazon Reviews
von: Khatun, Mst Eshita, et al.
Veröffentlicht: (2025)
von: Khatun, Mst Eshita, et al.
Veröffentlicht: (2025)
Optimizing Pre-Training Data Mixtures with Mixtures of Data Expert Models
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
von: Belenki, Lior, et al.
Veröffentlicht: (2025)
Can Federated Learning Safeguard Private Data in LLM Training? Vulnerabilities, Attacks, and Defense Evaluation
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
von: Guo, Wenkai, et al.
Veröffentlicht: (2025)
Data Shapley in One Training Run
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
von: Wang, Jiachen T., et al.
Veröffentlicht: (2024)
On Training Data Influence of GPT Models
von: Chai, Yekun, et al.
Veröffentlicht: (2024)
von: Chai, Yekun, et al.
Veröffentlicht: (2024)
Synthetic Data Generation for Intersectional Fairness by Leveraging Hierarchical Group Structure
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
von: Maheshwari, Gaurav, et al.
Veröffentlicht: (2024)
Prescriptive Scaling Laws for Data Constrained Training
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
von: Lovelace, Justin, et al.
Veröffentlicht: (2026)
Data Mixture Inference: What do BPE Tokenizers Reveal about their Training Data?
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
von: Hayase, Jonathan, et al.
Veröffentlicht: (2024)
Financial Sentiment Analysis: Leveraging Actual and Synthetic Data for Supervised Fine-tuning
von: Atsiwo, Abraham
Veröffentlicht: (2024)
von: Atsiwo, Abraham
Veröffentlicht: (2024)
Mitigating Paraphrase Attacks on Machine-Text Detectors via Paraphrase Inversion
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2024)
von: Soto, Rafael Rivera, et al.
Veröffentlicht: (2024)
What is in Your Safe Data? Identifying Benign Data that Breaks Safety
von: He, Luxi, et al.
Veröffentlicht: (2024)
von: He, Luxi, et al.
Veröffentlicht: (2024)
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
von: Iskander, Shadi, et al.
Veröffentlicht: (2024)
DataFlex: A Unified Framework for Data-Centric Dynamic Training of Large Language Models
von: Liang, Hao, et al.
Veröffentlicht: (2026)
von: Liang, Hao, et al.
Veröffentlicht: (2026)
Unsupervised Data Validation Methods for Efficient Model Training
von: Paniv, Yurii
Veröffentlicht: (2024)
von: Paniv, Yurii
Veröffentlicht: (2024)
Training Bilingual LMs with Data Constraints in the Targeted Language
von: Seto, Skyler, et al.
Veröffentlicht: (2024)
von: Seto, Skyler, et al.
Veröffentlicht: (2024)
ZClip: Adaptive Spike Mitigation for LLM Pre-Training
von: Kumar, Abhay, et al.
Veröffentlicht: (2025)
von: Kumar, Abhay, et al.
Veröffentlicht: (2025)
Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
MALTO at SemEval-2024 Task 6: Leveraging Synthetic Data for LLM Hallucination Detection
von: Borra, Federico, et al.
Veröffentlicht: (2024)
von: Borra, Federico, et al.
Veröffentlicht: (2024)
ISACL: Internal State Analyzer for Copyrighted Training Data Leakage
von: Zhang, Guangwei, et al.
Veröffentlicht: (2025)
von: Zhang, Guangwei, et al.
Veröffentlicht: (2025)
Mitigating Data Imbalance in Automated Speaking Assessment
von: Tsai, Fong-Chun, et al.
Veröffentlicht: (2025)
von: Tsai, Fong-Chun, et al.
Veröffentlicht: (2025)
Special Characters Attack: Toward Scalable Training Data Extraction From Large Language Models
von: Bai, Yang, et al.
Veröffentlicht: (2024)
von: Bai, Yang, et al.
Veröffentlicht: (2024)
Leveraging Variation Theory in Counterfactual Data Augmentation for Optimized Active Learning
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2024)
von: Gebreegziabher, Simret Araya, et al.
Veröffentlicht: (2024)
Iterative Data Smoothing: Mitigating Reward Overfitting and Overoptimization in RLHF
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
von: Zhu, Banghua, et al.
Veröffentlicht: (2024)
How to Train Data-Efficient LLMs
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
von: Sachdeva, Noveen, et al.
Veröffentlicht: (2024)
Reinforcement Learning on Pre-Training Data
von: Li, Siheng, et al.
Veröffentlicht: (2025)
von: Li, Siheng, et al.
Veröffentlicht: (2025)
How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
von: Xiao, Teng, et al.
Veröffentlicht: (2024)
A Natural Language Processing Approach to Support Biomedical Data Harmonization: Leveraging Large Language Models
von: Li, Zexu, et al.
Veröffentlicht: (2024)
von: Li, Zexu, et al.
Veröffentlicht: (2024)
QuRating: Selecting High-Quality Data for Training Language Models
von: Wettig, Alexander, et al.
Veröffentlicht: (2024)
von: Wettig, Alexander, et al.
Veröffentlicht: (2024)
Sequence-Level Leakage Risk of Training Data in Large Language Models
von: Tiwari, Trishita, et al.
Veröffentlicht: (2024)
von: Tiwari, Trishita, et al.
Veröffentlicht: (2024)
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
von: Aerni, Michael, et al.
Veröffentlicht: (2024)
Efficient RLVR Training via Weighted Mutual Information Data Selection
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
von: Zhou, Xinyu, et al.
Veröffentlicht: (2026)
Towards Next-Generation LLM Training: From the Data-Centric Perspective
von: Liang, Hao, et al.
Veröffentlicht: (2026)
von: Liang, Hao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
LLM Assistance for Pediatric Depression
von: Ignashina, Mariia, et al.
Veröffentlicht: (2025) -
Privacy-Preserving Behaviour of Chatbot Users: Steering Through Trust Dynamics
von: Ive, Julia, et al.
Veröffentlicht: (2024) -
Annotation Sensitivity: Training Data Collection Methods Affect Model Performance
von: Kern, Christoph, et al.
Veröffentlicht: (2023) -
PII-Scope: A Comprehensive Study on Training Data PII Extraction Attacks in LLMs
von: Nakka, Krishna Kanth, et al.
Veröffentlicht: (2024) -
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
von: Chen, Howard, et al.
Veröffentlicht: (2025)