Unified Game Moderation: Soft-Prompting and LLM-Assisted Label Transfer for Resource-Efficient Toxicity Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Zachary, Tullo, Domenico, Rabbany, Reihaneh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BenCSSmark: Making the Social Sciences Count in LLM Research
by: Chatelain, Arnault, et al.
Published: (2026)
by: Chatelain, Arnault, et al.
Published: (2026)
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
by: Petrova, Tatiana, et al.
Published: (2026)
by: Petrova, Tatiana, et al.
Published: (2026)
Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
by: Ge, Zhuohan, et al.
Published: (2025)
by: Ge, Zhuohan, et al.
Published: (2025)
What distinguishes conspiracy from critical narratives? A computational analysis of oppositional discourse
by: Korenčić, Damir, et al.
Published: (2024)
by: Korenčić, Damir, et al.
Published: (2024)
Using Letter Positional Probabilities to Assess Word Complexity
by: Dalvean, Michael
Published: (2024)
by: Dalvean, Michael
Published: (2024)
The Representational Alignment between Humans and Language Models is implicitly driven by a Concreteness Effect
by: Iaia, Cosimo, et al.
Published: (2025)
by: Iaia, Cosimo, et al.
Published: (2025)
Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
by: Castro-Gonzalez, Leonardo, et al.
Published: (2024)
by: Castro-Gonzalez, Leonardo, et al.
Published: (2024)
Collective Memory and Narrative Cohesion: A Computational Study of Palestinian Refugee Oral Histories in Lebanon
by: Awwad, Ghadeer, et al.
Published: (2025)
by: Awwad, Ghadeer, et al.
Published: (2025)
Language Predicts Identity Fusion Across Cultures and Reveals Divergent Pathways to Violence
by: Wright, Devin R., et al.
Published: (2026)
by: Wright, Devin R., et al.
Published: (2026)
Chronic pain patient narratives allow for the estimation of current pain intensity
by: Nunes, Diogo A. P., et al.
Published: (2022)
by: Nunes, Diogo A. P., et al.
Published: (2022)
The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games
by: Mozikov, Mikhail, et al.
Published: (2024)
by: Mozikov, Mikhail, et al.
Published: (2024)
Do Political Opinions Transfer Between Western Languages? An Analysis of Unaligned and Aligned Multilingual LLMs
by: Weeber, Franziska, et al.
Published: (2025)
by: Weeber, Franziska, et al.
Published: (2025)
Scaling In, Not Up? Testing Thick Citation Context Analysis with GPT-5 and Fragile Prompts
by: Simons, Arno
Published: (2026)
by: Simons, Arno
Published: (2026)
Evaluating LLM Prompts for Data Augmentation in Multi-label Classification of Ecological Texts
by: Glazkova, Anna, et al.
Published: (2024)
by: Glazkova, Anna, et al.
Published: (2024)
Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization
by: Elganayni, Mohamed Hesham, et al.
Published: (2026)
by: Elganayni, Mohamed Hesham, et al.
Published: (2026)
Forgotten Words: Benchmarking NeoBERT for Dementia Detection in Low-Resource Conversational Filipino and English Speech
by: Floresca, Rez Samantha Z., et al.
Published: (2026)
by: Floresca, Rez Samantha Z., et al.
Published: (2026)
A new mapping of technological interdependence
by: Colladon, A. Fronzetti, et al.
Published: (2023)
by: Colladon, A. Fronzetti, et al.
Published: (2023)
Aligning Multilingual News for Stock Return Prediction
by: Wu, Yuntao, et al.
Published: (2025)
by: Wu, Yuntao, et al.
Published: (2025)
Discursive objection strategies in online comments: Developing a classification schema and validating its training
by: Shea, Ashley L., et al.
Published: (2024)
by: Shea, Ashley L., et al.
Published: (2024)
NLP Occupational Emergence Analysis: How Occupations Form and Evolve in Real Time -- A Zero-Assumption Method Demonstrated on AI in the US Technology Workforce, 2022-2026
by: Nordfors, David
Published: (2026)
by: Nordfors, David
Published: (2026)
ARGUS: Seeing the Influence of Narrative Features on Persuasion in Argumentative Texts
by: Nabhani, Sara, et al.
Published: (2026)
by: Nabhani, Sara, et al.
Published: (2026)
NTLRAG: Narrative Topic Labels derived with Retrieval Augmented Generation
by: Grobelscheg, Lisa, et al.
Published: (2026)
by: Grobelscheg, Lisa, et al.
Published: (2026)
Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from Text
by: Wright, Devin R., et al.
Published: (2025)
by: Wright, Devin R., et al.
Published: (2025)
Modeling Changing Scientific Concepts with Complex Networks: A Case Study on the Chemical Revolution
by: Aguilar-Valdez, Sofía, et al.
Published: (2026)
by: Aguilar-Valdez, Sofía, et al.
Published: (2026)
On Fact and Frequency: LLM Responses to Misinformation Expressed with Uncertainty
by: van de Sande, Yana, et al.
Published: (2025)
by: van de Sande, Yana, et al.
Published: (2025)
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
by: Zhang, Li, et al.
Published: (2026)
by: Zhang, Li, et al.
Published: (2026)
Prompting from the bench: Large-scale pretraining is not sufficient to prepare LLMs for ordinary meaning analysis
by: Purushothama, Abhishek, et al.
Published: (2025)
by: Purushothama, Abhishek, et al.
Published: (2025)
PromptAug: Fine-grained Conflict Classification Using Data Augmentation
by: Warke, Oliver, et al.
Published: (2025)
by: Warke, Oliver, et al.
Published: (2025)
Detecting Effects of AI-Mediated Communication on Language Complexity and Sentiment
by: Sussman, Kristen, et al.
Published: (2025)
by: Sussman, Kristen, et al.
Published: (2025)
LLMs Simulate Big Five Personality Traits: Further Evidence
by: Sorokovikova, Aleksandra, et al.
Published: (2024)
by: Sorokovikova, Aleksandra, et al.
Published: (2024)
Repetition Without Exclusivity: Scale Sensitivity of Referential Mechanisms in Child-Scale Language Models
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
by: Dhole, Kaustubh D.
Published: (2026)
by: Dhole, Kaustubh D.
Published: (2026)
Categorical Perception in Large Language Model Hidden States: Structural Warping at Digit-Count Boundaries
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Therapy as an NLP Task: Psychologists' Comparison of LLMs and Human Peers in CBT
by: Iftikhar, Zainab, et al.
Published: (2024)
by: Iftikhar, Zainab, et al.
Published: (2024)
Ornithologist: Towards Trustworthy "Reasoning" about Central Bank Communications
by: Jones, Dominic Zaun Eu
Published: (2025)
by: Jones, Dominic Zaun Eu
Published: (2025)
Multi-Method Validation of Large Language Model Medical Translation Across High- and Low-Resource Languages
by: Anyaegbuna, Chukwuebuka, et al.
Published: (2026)
by: Anyaegbuna, Chukwuebuka, et al.
Published: (2026)
The Table of Media Bias Elements: A sentence-level taxonomy of media bias types and propaganda techniques
by: Menzner, Tim, et al.
Published: (2026)
by: Menzner, Tim, et al.
Published: (2026)
From Helpfulness to Toxic Proactivity: Diagnosing Behavioral Misalignment in LLM Agents
by: Wang, Xinyue, et al.
Published: (2026)
by: Wang, Xinyue, et al.
Published: (2026)
ParliaBench: An Evaluation and Benchmarking Framework for LLM-Generated Parliamentary Speech
by: Koniaris, Marios, et al.
Published: (2025)
by: Koniaris, Marios, et al.
Published: (2025)
Analysis of LLM as a grammatical feature tagger for African American English
by: Porwal, Rahul, et al.
Published: (2025)
by: Porwal, Rahul, et al.
Published: (2025)
Similar Items
-
BenCSSmark: Making the Social Sciences Count in LLM Research
by: Chatelain, Arnault, et al.
Published: (2026) -
How Much Does Persuasion Strategy Matter? LLM-Annotated Evidence from Charitable Donation Dialogues
by: Petrova, Tatiana, et al.
Published: (2026) -
Survey and Experiments on Mental Disorder Detection via Social Media: From Large Language Models and RAG to Agents
by: Ge, Zhuohan, et al.
Published: (2025) -
What distinguishes conspiracy from critical narratives? A computational analysis of oppositional discourse
by: Korenčić, Damir, et al.
Published: (2024) -
Using Letter Positional Probabilities to Assess Word Complexity
by: Dalvean, Michael
Published: (2024)