What Are They Filtering Out? An Experimental Benchmark of Filtering Strategies for Harm Reduction in Pretraining Datasets
Fuente:
arXiv
Salvato in:
| Autori principali: | Stranisci, Marco Antonio, Hardmeier, Christian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A Dataset for the Detection of Dehumanizing Language
di: Engelmann, Paul, et al.
Pubblicazione: (2024)
di: Engelmann, Paul, et al.
Pubblicazione: (2024)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
di: Mendu, Sai Krishna, et al.
Pubblicazione: (2025)
di: Mendu, Sai Krishna, et al.
Pubblicazione: (2025)
With Good MT There is No Need For End-to-End: A Case for Translate-then-Summarize Cross-lingual Summarization
di: Varab, Daniel, et al.
Pubblicazione: (2024)
di: Varab, Daniel, et al.
Pubblicazione: (2024)
Mention Attention for Pronoun Translation
di: Tang, Gongbo, et al.
Pubblicazione: (2024)
di: Tang, Gongbo, et al.
Pubblicazione: (2024)
Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
di: Ulmer, Dennis, et al.
Pubblicazione: (2025)
di: Ulmer, Dennis, et al.
Pubblicazione: (2025)
Dealing with Controversy: An Emotion and Coping Strategy Corpus Based on Role Playing
di: Troiano, Enrica, et al.
Pubblicazione: (2024)
di: Troiano, Enrica, et al.
Pubblicazione: (2024)
Are you sure? Measuring models bias in content moderation through uncertainty
di: Urbinati, Alessandra, et al.
Pubblicazione: (2025)
di: Urbinati, Alessandra, et al.
Pubblicazione: (2025)
Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text
di: Ruby, Ahmed, et al.
Pubblicazione: (2026)
di: Ruby, Ahmed, et al.
Pubblicazione: (2026)
The Data-Quality Illusion: Rethinking Classifier-Based Quality Filtering for LLM Pretraining
di: Saada, Thiziri Nait, et al.
Pubblicazione: (2025)
di: Saada, Thiziri Nait, et al.
Pubblicazione: (2025)
Beyond Line-Level Filtering for the Pretraining Corpora of LLMs
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
di: Park, Chanwoo, et al.
Pubblicazione: (2025)
That is Unacceptable: the Moral Foundations of Canceling
di: Lo, Soda Marem, et al.
Pubblicazione: (2025)
di: Lo, Soda Marem, et al.
Pubblicazione: (2025)
Improving Model Evaluation using SMART Filtering of Benchmark Datasets
di: Gupta, Vipul, et al.
Pubblicazione: (2024)
di: Gupta, Vipul, et al.
Pubblicazione: (2024)
Improving Romanian LLM Pretraining Data using Diversity and Quality Filtering
di: Negoita, Vlad, et al.
Pubblicazione: (2025)
di: Negoita, Vlad, et al.
Pubblicazione: (2025)
Beyond Filtering: Adaptive Image-Text Quality Enhancement for MLLM Pretraining
di: Huang, Han, et al.
Pubblicazione: (2024)
di: Huang, Han, et al.
Pubblicazione: (2024)
An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms
di: Grünefeld, Nils, et al.
Pubblicazione: (2026)
di: Grünefeld, Nils, et al.
Pubblicazione: (2026)
AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters
di: Lucy, Li, et al.
Pubblicazione: (2024)
di: Lucy, Li, et al.
Pubblicazione: (2024)
What to Keep and What to Drop: Adaptive Table Filtering Framework
di: Jang, WonJune
Pubblicazione: (2025)
di: Jang, WonJune
Pubblicazione: (2025)
Wikibio: a Semantic Resource for the Intersectional Analysis of Biographical Events
di: Stranisci, Marco Antonio, et al.
Pubblicazione: (2023)
di: Stranisci, Marco Antonio, et al.
Pubblicazione: (2023)
CoLoR-Filter: Conditional Loss Reduction Filtering for Targeted Language Model Pre-training
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
di: Brandfonbrener, David, et al.
Pubblicazione: (2024)
TailNLG: A Multilingual Benchmark Addressing Verbalization of Long-Tail Entities
di: Draetta, Lia, et al.
Pubblicazione: (2026)
di: Draetta, Lia, et al.
Pubblicazione: (2026)
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
di: Wang, Kaixuan, et al.
Pubblicazione: (2025)
di: Wang, Kaixuan, et al.
Pubblicazione: (2025)
ReFilter: Improving Robustness of Retrieval-Augmented Generation via Gated Filter
di: Chen, Yixin, et al.
Pubblicazione: (2026)
di: Chen, Yixin, et al.
Pubblicazione: (2026)
COUNTDOWN: Contextually Sparse Activation Filtering Out Unnecessary Weights in Down Projection
di: Cheon, Jaewon, et al.
Pubblicazione: (2025)
di: Cheon, Jaewon, et al.
Pubblicazione: (2025)
Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
di: Yang, Yuchen, et al.
Pubblicazione: (2026)
CONGRAD:Conflicting Gradient Filtering for Multilingual Preference Alignment
di: Li, Jiangnan, et al.
Pubblicazione: (2025)
di: Li, Jiangnan, et al.
Pubblicazione: (2025)
Garbage In, Reasoning Out? Why Benchmark Scores are Unreliable and What to Do About It
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2025)
di: Mousavi, Seyed Mahed, et al.
Pubblicazione: (2025)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
di: Ali, Mehdi, et al.
Pubblicazione: (2025)
di: Ali, Mehdi, et al.
Pubblicazione: (2025)
Conspiracy Frame: a Semiotically-Driven Approach for Conspiracy Theories Detection
di: Piva, Heidi Campana, et al.
Pubblicazione: (2026)
di: Piva, Heidi Campana, et al.
Pubblicazione: (2026)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
HarmMetric Eval: Benchmarking Metrics and Judges for LLM Harmfulness Assessment
di: Yang, Langqi, et al.
Pubblicazione: (2025)
di: Yang, Langqi, et al.
Pubblicazione: (2025)
GPT-4o as the Gold Standard: A Scalable and General Purpose Approach to Filter Language Model Pretraining Data
di: Zhang, Jifan, et al.
Pubblicazione: (2024)
di: Zhang, Jifan, et al.
Pubblicazione: (2024)
BlendFilter: Advancing Retrieval-Augmented Large Language Models via Query Generation Blending and Knowledge Filtering
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
di: Sharshar, Ahmed, et al.
Pubblicazione: (2026)
di: Sharshar, Ahmed, et al.
Pubblicazione: (2026)
ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis
di: Tu, Zeao, et al.
Pubblicazione: (2024)
di: Tu, Zeao, et al.
Pubblicazione: (2024)
CBRS: Cognitive Blood Request System with Bilingual Dataset and Dual-Layer Filtering for Multi-Platform Social Streams
di: Saha, Anik, et al.
Pubblicazione: (2026)
di: Saha, Anik, et al.
Pubblicazione: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
Fair Play in the Newsroom: Actor-Based Filtering Gender Discrimination in Text Corpora
di: Urchs, Stefanie, et al.
Pubblicazione: (2025)
di: Urchs, Stefanie, et al.
Pubblicazione: (2025)
Pretraining and Benchmarking Modern Encoders for Latvian
di: Znotins, Arturs
Pubblicazione: (2026)
di: Znotins, Arturs
Pubblicazione: (2026)
Context Filtering with Reward Modeling in Question Answering
di: Kim, Sangryul, et al.
Pubblicazione: (2024)
di: Kim, Sangryul, et al.
Pubblicazione: (2024)
A comparison of data filtering techniques for English-Polish LLM-based machine translation in the biomedical domain
di: Lérida, Jorge del Pozo, et al.
Pubblicazione: (2025)
di: Lérida, Jorge del Pozo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
A Dataset for the Detection of Dehumanizing Language
di: Engelmann, Paul, et al.
Pubblicazione: (2024) -
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
di: Mendu, Sai Krishna, et al.
Pubblicazione: (2025) -
With Good MT There is No Need For End-to-End: A Case for Translate-then-Summarize Cross-lingual Summarization
di: Varab, Daniel, et al.
Pubblicazione: (2024) -
Mention Attention for Pronoun Translation
di: Tang, Gongbo, et al.
Pubblicazione: (2024) -
Anthropomimetic Uncertainty: What Verbalized Uncertainty in Language Models is Missing
di: Ulmer, Dennis, et al.
Pubblicazione: (2025)