Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Mendu, Sai Krishna, Yenala, Harish, Gulati, Aditi, Kumar, Shanu, Agrawal, Parag |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SCULPT: Systematic Tuning of Long Prompts
di: Kumar, Shanu, et al.
Pubblicazione: (2024)
di: Kumar, Shanu, et al.
Pubblicazione: (2024)
Safer Policy Compliance with Dynamic Epistemic Fallback
di: Imperial, Joseph Marvin, et al.
Pubblicazione: (2026)
di: Imperial, Joseph Marvin, et al.
Pubblicazione: (2026)
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
di: Sam, Dylan, et al.
Pubblicazione: (2025)
di: Sam, Dylan, et al.
Pubblicazione: (2025)
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
di: Kumar, Shanu, et al.
Pubblicazione: (2024)
di: Kumar, Shanu, et al.
Pubblicazione: (2024)
Keep Calm and Avoid Harmful Content: Concept Alignment and Latent Manipulation Towards Safer Answers
di: Belo, Ruben, et al.
Pubblicazione: (2025)
di: Belo, Ruben, et al.
Pubblicazione: (2025)
SMLT-MUGC: Small, Medium, and Large Texts -- Machine versus User-Generated Content Detection and Comparison
di: Rawal, Anjali, et al.
Pubblicazione: (2024)
di: Rawal, Anjali, et al.
Pubblicazione: (2024)
Tracing the Representation Geometry of Language Models from Pretraining to Post-training
di: Li, Melody Zixuan, et al.
Pubblicazione: (2025)
di: Li, Melody Zixuan, et al.
Pubblicazione: (2025)
Biomed-Enriched: A Biomedical Dataset Enriched with LLMs for Pretraining and Extracting Rare and Hidden Content
di: Touchent, Rian, et al.
Pubblicazione: (2025)
di: Touchent, Rian, et al.
Pubblicazione: (2025)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
di: Watts, Ishaan, et al.
Pubblicazione: (2026)
Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs
di: Si, Wai Man, et al.
Pubblicazione: (2026)
di: Si, Wai Man, et al.
Pubblicazione: (2026)
MUGC: Machine Generated versus User Generated Content Detection
di: Xie, Yaqi, et al.
Pubblicazione: (2024)
di: Xie, Yaqi, et al.
Pubblicazione: (2024)
FCoReBench: Can Large Language Models Solve Challenging First-Order Combinatorial Reasoning Problems?
di: Mittal, Chinmay, et al.
Pubblicazione: (2024)
di: Mittal, Chinmay, et al.
Pubblicazione: (2024)
Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
Toward a Safer Web: Multilingual Multi-Agent LLMs for Mitigating Adversarial Misinformation Attacks
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
di: Aldahoul, Nouar, et al.
Pubblicazione: (2025)
GRAID: Synthetic Data Generation with Geometric Constraints and Multi-Agentic Reflection for Harmful Content Detection
di: Rad, Melissa Kazemi, et al.
Pubblicazione: (2025)
di: Rad, Melissa Kazemi, et al.
Pubblicazione: (2025)
Leveraging User-Generated Reviews for Recommender Systems with Dynamic Headers
di: Vashishtha, Shanu, et al.
Pubblicazione: (2024)
di: Vashishtha, Shanu, et al.
Pubblicazione: (2024)
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
di: Zhong, Ziqian, et al.
Pubblicazione: (2025)
Semantic Anchors in In-Context Learning: Why Small LLMs Cannot Flip Their Labels
di: Kumar, Anantha Padmanaban Krishna
Pubblicazione: (2025)
di: Kumar, Anantha Padmanaban Krishna
Pubblicazione: (2025)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
di: Saada, Thiziri Nait, et al.
Pubblicazione: (2025)
di: Saada, Thiziri Nait, et al.
Pubblicazione: (2025)
Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
di: Cen, Zhepeng, et al.
Pubblicazione: (2025)
di: Cen, Zhepeng, et al.
Pubblicazione: (2025)
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning
di: Feng, Weitao, et al.
Pubblicazione: (2025)
di: Feng, Weitao, et al.
Pubblicazione: (2025)
ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark
di: Liu, Kangwei, et al.
Pubblicazione: (2025)
di: Liu, Kangwei, et al.
Pubblicazione: (2025)
SOI Matters: Analyzing Multi-Setting Training Dynamics in Pretrained Language Models via Subsets of Interest
di: Vassef, Shayan, et al.
Pubblicazione: (2025)
di: Vassef, Shayan, et al.
Pubblicazione: (2025)
Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
di: Zhang, Haoyu, et al.
Pubblicazione: (2026)
Towards Universal and Black-Box Query-Response Only Attack on LLMs with QROA
di: Jawad, Hussein, et al.
Pubblicazione: (2024)
di: Jawad, Hussein, et al.
Pubblicazione: (2024)
Towards Efficient Active Learning in NLP via Pretrained Representations
di: Vysogorets, Artem, et al.
Pubblicazione: (2024)
di: Vysogorets, Artem, et al.
Pubblicazione: (2024)
ProxSparse: Regularized Learning of Semi-Structured Sparsity Masks for Pretrained LLMs
di: Liu, Hongyi, et al.
Pubblicazione: (2025)
di: Liu, Hongyi, et al.
Pubblicazione: (2025)
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
di: Fan, Dongyang, et al.
Pubblicazione: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
di: Sheshadri, Abhay, et al.
Pubblicazione: (2024)
di: Sheshadri, Abhay, et al.
Pubblicazione: (2024)
Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey
di: Agrawal, Garima, et al.
Pubblicazione: (2023)
di: Agrawal, Garima, et al.
Pubblicazione: (2023)
Format as a Prior: Quantifying and Analyzing Bias in LLMs for Heterogeneous Data
di: Liu, Jiacheng, et al.
Pubblicazione: (2025)
di: Liu, Jiacheng, et al.
Pubblicazione: (2025)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
di: Farashah, Alireza Dehghanpour, et al.
Pubblicazione: (2026)
di: Farashah, Alireza Dehghanpour, et al.
Pubblicazione: (2026)
Pretrained Hybrids with MAD Skills
di: Roberts, Nicholas, et al.
Pubblicazione: (2024)
di: Roberts, Nicholas, et al.
Pubblicazione: (2024)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
di: Ni, Kangqi, et al.
Pubblicazione: (2025)
di: Ni, Kangqi, et al.
Pubblicazione: (2025)
Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content
di: Stepanov, Ihor, et al.
Pubblicazione: (2026)
di: Stepanov, Ihor, et al.
Pubblicazione: (2026)
GPT-4o as the Gold Standard: A Scalable and General Purpose Approach to Filter Language Model Pretraining Data
di: Zhang, Jifan, et al.
Pubblicazione: (2024)
di: Zhang, Jifan, et al.
Pubblicazione: (2024)
Search Arena: Analyzing Search-Augmented LLMs
di: Miroyan, Mihran, et al.
Pubblicazione: (2025)
di: Miroyan, Mihran, et al.
Pubblicazione: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
di: Chaudhary, Maheep, et al.
Pubblicazione: (2025)
Towards Analyzing and Understanding the Limitations of VAPO: A Theoretical Perspective
di: Shao, Jintian, et al.
Pubblicazione: (2025)
di: Shao, Jintian, et al.
Pubblicazione: (2025)
Emergent Response Planning in LLMs
di: Dong, Zhichen, et al.
Pubblicazione: (2025)
di: Dong, Zhichen, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SCULPT: Systematic Tuning of Long Prompts
di: Kumar, Shanu, et al.
Pubblicazione: (2024) -
Safer Policy Compliance with Dynamic Epistemic Fallback
di: Imperial, Joseph Marvin, et al.
Pubblicazione: (2026) -
Analyzing Similarity Metrics for Data Selection for Language Model Pretraining
di: Sam, Dylan, et al.
Pubblicazione: (2025) -
Socio-Culturally Aware Evaluation Framework for LLM-Based Content Moderation
di: Kumar, Shanu, et al.
Pubblicazione: (2024) -
Keep Calm and Avoid Harmful Content: Concept Alignment and Latent Manipulation Towards Safer Answers
di: Belo, Ruben, et al.
Pubblicazione: (2025)