Balancing Quality and Variation: Spam Filtering Distorts Data Label Distributions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fleisig, Eve, Orlikowski, Matthias, Cimiano, Philipp, Klein, Dan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It's Best to Relate Perspectives!
von: Heinisch, Philipp, et al.
Veröffentlicht: (2023)
von: Heinisch, Philipp, et al.
Veröffentlicht: (2023)
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
von: Fleisig, Eve, et al.
Veröffentlicht: (2023)
von: Fleisig, Eve, et al.
Veröffentlicht: (2023)
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
von: Verma, Vivek, et al.
Veröffentlicht: (2023)
von: Verma, Vivek, et al.
Veröffentlicht: (2023)
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
von: Orlikowski, Matthias, et al.
Veröffentlicht: (2023)
von: Orlikowski, Matthias, et al.
Veröffentlicht: (2023)
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
von: Fleisig, Eve, et al.
Veröffentlicht: (2024)
von: Fleisig, Eve, et al.
Veröffentlicht: (2024)
CompoST: A Benchmark for Analyzing the Ability of LLMs To Compositionally Interpret Questions in a QALD Setting
von: Schmidt, David Maria, et al.
Veröffentlicht: (2025)
von: Schmidt, David Maria, et al.
Veröffentlicht: (2025)
Mapping Social Choice Theory to RLHF
von: Dai, Jessica, et al.
Veröffentlicht: (2024)
von: Dai, Jessica, et al.
Veröffentlicht: (2024)
From Argumentation to Deliberation: Perspectivized Stance Vectors for Fine-grained (Dis)agreement Analysis
von: Plenz, Moritz, et al.
Veröffentlicht: (2025)
von: Plenz, Moritz, et al.
Veröffentlicht: (2025)
Conditional Semi-Supervised Data Augmentation for Spam Message Detection with Low Resource Data
von: Nuha, Ulin, et al.
Veröffentlicht: (2024)
von: Nuha, Ulin, et al.
Veröffentlicht: (2024)
Lexicalization Is All You Need: Examining the Impact of Lexical Knowledge in a Compositional QALD System
von: Schmidt, David Maria, et al.
Veröffentlicht: (2024)
von: Schmidt, David Maria, et al.
Veröffentlicht: (2024)
GCC-Spam: Spam Detection via GAN, Contrastive Learning, and Character Similarity Networks
von: Wang, Zhijie, et al.
Veröffentlicht: (2025)
von: Wang, Zhijie, et al.
Veröffentlicht: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
von: Li, Jing-Jing, et al.
Veröffentlicht: (2026)
AI, Take the Wheel: What Drives Delegation and Trust in Human-Computer Cooperative Question Answering?
von: Gor, Maharshi, et al.
Veröffentlicht: (2026)
von: Gor, Maharshi, et al.
Veröffentlicht: (2026)
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
von: Orlikowski, Matthias, et al.
Veröffentlicht: (2025)
von: Orlikowski, Matthias, et al.
Veröffentlicht: (2025)
SmolKalam: Ensemble Quality-Filtered Translation at Scale for High Quality Arabic Post-Training Data
von: Alrashed, Sultan, et al.
Veröffentlicht: (2025)
von: Alrashed, Sultan, et al.
Veröffentlicht: (2025)
Evaluating the Performance of ChatGPT for Spam Email Detection
von: Si, Shijing, et al.
Veröffentlicht: (2024)
von: Si, Shijing, et al.
Veröffentlicht: (2024)
Linguistic Bias in ChatGPT: Language Models Reinforce Dialect Discrimination
von: Fleisig, Eve, et al.
Veröffentlicht: (2024)
von: Fleisig, Eve, et al.
Veröffentlicht: (2024)
American Sign Language Handshapes Reflect Pressures for Communicative Efficiency
von: Yin, Kayo, et al.
Veröffentlicht: (2024)
von: Yin, Kayo, et al.
Veröffentlicht: (2024)
THOUGHTSCULPT: Reasoning with Intermediate Revision and Search
von: Chi, Yizhou, et al.
Veröffentlicht: (2024)
von: Chi, Yizhou, et al.
Veröffentlicht: (2024)
High-Quality Data Augmentation for Low-Resource NMT: Combining a Translation Memory, a GAN Generator, and Filtering
von: Liu, Hengjie, et al.
Veröffentlicht: (2024)
von: Liu, Hengjie, et al.
Veröffentlicht: (2024)
SMS Spam Detection and Classification to Combat Abuse in Telephone Networks Using Natural Language Processing
von: Oyeyemi, Dare Azeez, et al.
Veröffentlicht: (2024)
von: Oyeyemi, Dare Azeez, et al.
Veröffentlicht: (2024)
Accurate and Data-Efficient Toxicity Prediction when Annotators Disagree
von: Jaggi, Harbani, et al.
Veröffentlicht: (2024)
von: Jaggi, Harbani, et al.
Veröffentlicht: (2024)
Zero-Shot Spam Email Classification Using Pre-trained Large Language Models
von: Rojas-Galeano, Sergio
Veröffentlicht: (2024)
von: Rojas-Galeano, Sergio
Veröffentlicht: (2024)
Label Distribution Learning-Enhanced Dual-KNN for Text Classification
von: Yuan, Bo, et al.
Veröffentlicht: (2025)
von: Yuan, Bo, et al.
Veröffentlicht: (2025)
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
von: Pathak, Manas, et al.
Veröffentlicht: (2026)
von: Pathak, Manas, et al.
Veröffentlicht: (2026)
Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
von: Ali, Mehdi, et al.
Veröffentlicht: (2025)
Improving Neural Topic Modeling with Semantically-Grounded Soft Label Distributions
von: Li, Raymond, et al.
Veröffentlicht: (2026)
von: Li, Raymond, et al.
Veröffentlicht: (2026)
Balanced Data Sampling for Language Model Training with Clustering
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
von: Shao, Yunfan, et al.
Veröffentlicht: (2024)
How LLMs Distort Our Written Language
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2026)
von: Abdulhai, Marwa, et al.
Veröffentlicht: (2026)
HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on Text
von: Liu, Han, et al.
Veröffentlicht: (2024)
von: Liu, Han, et al.
Veröffentlicht: (2024)
From Noise to Signal to Selbstzweck: Reframing Human Label Variation in the Era of Post-training in NLP
von: Xu, Shanshan, et al.
Veröffentlicht: (2025)
von: Xu, Shanshan, et al.
Veröffentlicht: (2025)
Signature vs. Substance: Evaluating the Balance of Adversarial Resistance and Linguistic Quality in Watermarking Large Language Models
von: Guo, William, et al.
Veröffentlicht: (2025)
von: Guo, William, et al.
Veröffentlicht: (2025)
RLCD: Reinforcement Learning from Contrastive Distillation for Language Model Alignment
von: Yang, Kevin, et al.
Veröffentlicht: (2023)
von: Yang, Kevin, et al.
Veröffentlicht: (2023)
Rethinking KenLM: Good and Bad Model Ensembles for Efficient Text Quality Filtering in Large Web Corpora
von: Kim, Yungi, et al.
Veröffentlicht: (2024)
von: Kim, Yungi, et al.
Veröffentlicht: (2024)
SoftEDA: Rethinking Rule-Based Data Augmentation with Soft Labels
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
von: Choi, Juhwan, et al.
Veröffentlicht: (2024)
LLM as a Broken Telephone: Iterative Generation Distorts Information
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
von: Mohamed, Amr, et al.
Veröffentlicht: (2025)
Signs of Struggle: Spotting Cognitive Distortions across Language and Register
von: Kuber, Abhishek, et al.
Veröffentlicht: (2025)
von: Kuber, Abhishek, et al.
Veröffentlicht: (2025)
DECT: Harnessing LLM-assisted Fine-Grained Linguistic Knowledge and Label-Switched and Label-Preserved Data Generation for Diagnosis of Alzheimer's Disease
von: Mo, Tingyu, et al.
Veröffentlicht: (2025)
von: Mo, Tingyu, et al.
Veröffentlicht: (2025)
Distortion Instead of Hallucination: The Effect of Reasoning Under Strict Constraints
von: Niimi, Junichiro
Veröffentlicht: (2026)
von: Niimi, Junichiro
Veröffentlicht: (2026)
TrustDataFilter:Leveraging Trusted Knowledge Base Data for More Effective Filtering of Unknown Information
von: Zhang, Jinghong, et al.
Veröffentlicht: (2025)
von: Zhang, Jinghong, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Architectural Sweet Spots for Modeling Human Label Variation by the Example of Argument Quality: It's Best to Relate Perspectives!
von: Heinisch, Philipp, et al.
Veröffentlicht: (2023) -
When the Majority is Wrong: Modeling Annotator Disagreement for Subjective Tasks
von: Fleisig, Eve, et al.
Veröffentlicht: (2023) -
Ghostbuster: Detecting Text Ghostwritten by Large Language Models
von: Verma, Vivek, et al.
Veröffentlicht: (2023) -
The Ecological Fallacy in Annotation: Modelling Human Label Variation goes beyond Sociodemographics
von: Orlikowski, Matthias, et al.
Veröffentlicht: (2023) -
The Perspectivist Paradigm Shift: Assumptions and Challenges of Capturing Human Labels
von: Fleisig, Eve, et al.
Veröffentlicht: (2024)