SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Moskovskiy, Daniil, Sushko, Nikita, Pletenev, Sergey, Tutubalina, Elena, Panchenko, Alexander |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
Multilingual and Explainable Text Detoxification with Parallel Corpora
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
DetoxLLM: A Framework for Detoxification with Explanations
von: Khondaker, Md Tawkat Islam, et al.
Veröffentlicht: (2024)
von: Khondaker, Md Tawkat Islam, et al.
Veröffentlicht: (2024)
Anatomy of Unlearning: The Dual Impact of Fact Salience and Model Fine-Tuning
von: Borisiuk, Anna, et al.
Veröffentlicht: (2026)
von: Borisiuk, Anna, et al.
Veröffentlicht: (2026)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2026)
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2026)
Demarked: A Strategy for Enhanced Abusive Speech Moderation through Counterspeech, Detoxification, and Message Management
von: Yimam, Seid Muhie, et al.
Veröffentlicht: (2024)
von: Yimam, Seid Muhie, et al.
Veröffentlicht: (2024)
CoRoVA: Compressed Representations for Vector-Augmented Code Completion
von: Cherniuk, Daria, et al.
Veröffentlicht: (2025)
von: Cherniuk, Daria, et al.
Veröffentlicht: (2025)
GemDetox at TextDetox CLEF 2025: Enhancing a Massively Multilingual Model for Text Detoxification on Low-resource Languages
von: Dang, Trung Duc Anh, et al.
Veröffentlicht: (2025)
von: Dang, Trung Duc Anh, et al.
Veröffentlicht: (2025)
CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification
von: Wang, Yian, et al.
Veröffentlicht: (2026)
von: Wang, Yian, et al.
Veröffentlicht: (2026)
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2025)
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2025)
UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
von: Lu, Huimin, et al.
Veröffentlicht: (2025)
von: Lu, Huimin, et al.
Veröffentlicht: (2025)
Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
Evaluating Text Style Transfer: A Nine-Language Benchmark for Text Detoxification
von: Protasov, Vitaly, et al.
Veröffentlicht: (2025)
von: Protasov, Vitaly, et al.
Veröffentlicht: (2025)
SmurfCat at PAN 2024 TextDetox: Alignment of Multilingual Transformers for Text Detoxification
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
von: Afonin, Nikita, et al.
Veröffentlicht: (2025)
von: Afonin, Nikita, et al.
Veröffentlicht: (2025)
LLM-Independent Adaptive RAG: Let the Question Speak for Itself
von: Marina, Maria, et al.
Veröffentlicht: (2025)
von: Marina, Maria, et al.
Veröffentlicht: (2025)
Team Anotheroption at SemEval-2025 Task 8: Bridging the Gap Between Open-Source and Proprietary LLMs in Table QA
von: Evkarpidi, Nikolas, et al.
Veröffentlicht: (2025)
von: Evkarpidi, Nikolas, et al.
Veröffentlicht: (2025)
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
von: Alekseev, Artem, et al.
Veröffentlicht: (2025)
von: Alekseev, Artem, et al.
Veröffentlicht: (2025)
Adaptive Retrieval Without Self-Knowledge? Bringing Uncertainty Back Home
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
von: Moskvoretskii, Viktor, et al.
Veröffentlicht: (2025)
Geopolitical biases in LLMs: what are the "good" and the "bad" countries according to contemporary language models
von: Salnikov, Mikhail, et al.
Veröffentlicht: (2025)
von: Salnikov, Mikhail, et al.
Veröffentlicht: (2025)
LLMs on Drugs: Language Models Are Few-Shot Consumers
von: Doudkin, Alexander
Veröffentlicht: (2025)
von: Doudkin, Alexander
Veröffentlicht: (2025)
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2024)
von: Kulkarni, Atharva, et al.
Veröffentlicht: (2024)
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
Confidence Estimation for Error Detection in Text-to-SQL Systems
von: Somov, Oleg, et al.
Veröffentlicht: (2025)
von: Somov, Oleg, et al.
Veröffentlicht: (2025)
Cost-Efficient Subjective Task Annotation and Modeling through Few-Shot Annotator Adaptation
von: Golazizian, Preni, et al.
Veröffentlicht: (2024)
von: Golazizian, Preni, et al.
Veröffentlicht: (2024)
One Task Vector is not Enough: A Large-Scale Study for In-Context Learning
von: Tikhonov, Pavel, et al.
Veröffentlicht: (2025)
von: Tikhonov, Pavel, et al.
Veröffentlicht: (2025)
Bilingual Rhetorical Structure Parsing with Large Parallel Annotations
von: Chistova, Elena
Veröffentlicht: (2024)
von: Chistova, Elena
Veröffentlicht: (2024)
Breaking Bad Tokens: Detoxification of LLMs Using Sparse Autoencoders
von: Goyal, Agam, et al.
Veröffentlicht: (2025)
von: Goyal, Agam, et al.
Veröffentlicht: (2025)
HatePRISM: Policies, Platforms, and Research Integration. Advancing NLP for Hate Speech Proactive Mitigation
von: Rizwan, Naquee, et al.
Veröffentlicht: (2025)
von: Rizwan, Naquee, et al.
Veröffentlicht: (2025)
MERA: A Comprehensive LLM Evaluation in Russian
von: Fenogenova, Alena, et al.
Veröffentlicht: (2024)
von: Fenogenova, Alena, et al.
Veröffentlicht: (2024)
nach0-pc: Multi-task Language Model with Molecular Point Cloud Encoder
von: Kuznetsov, Maksim, et al.
Veröffentlicht: (2024)
von: Kuznetsov, Maksim, et al.
Veröffentlicht: (2024)
Cross-Lingual Transfer of Debiasing and Detoxification in Multilingual LLMs: An Extensive Investigation
von: Neplenbroek, Vera, et al.
Veröffentlicht: (2024)
von: Neplenbroek, Vera, et al.
Veröffentlicht: (2024)
How to Solve Few-Shot Abusive Content Detection Using the Data We Actually Have
von: Hangya, Viktor, et al.
Veröffentlicht: (2023)
von: Hangya, Viktor, et al.
Veröffentlicht: (2023)
SmurfCat at SemEval-2024 Task 6: Leveraging Synthetic Data for Hallucination Detection
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
InfoSynth: Information-Guided Benchmark Synthesis for LLMs
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models
von: Gu, Ken, et al.
Veröffentlicht: (2025)
von: Gu, Ken, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025) -
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024) -
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025) -
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026) -
Multilingual and Explainable Text Detoxification with Parallel Corpora
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)