<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Pletenev, Sergey, Moskovskiy, Daniil, Panchenko, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators
by: Moskovskiy, Daniil, et al.
Published: (2025)
by: Moskovskiy, Daniil, et al.
Published: (2025)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
by: Vazhentsev, Artem, et al.
Published: (2026)
by: Vazhentsev, Artem, et al.
Published: (2026)
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
by: Pletenev, Sergey, et al.
Published: (2025)
by: Pletenev, Sergey, et al.
Published: (2025)
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
by: Zheng, Danna, et al.
Published: (2024)
by: Zheng, Danna, et al.
Published: (2024)
Scaling up the think-aloud method
by: Wurgaft, Daniel, et al.
Published: (2025)
by: Wurgaft, Daniel, et al.
Published: (2025)
Do not think about pink elephant!
by: Hwang, Kyomin, et al.
Published: (2024)
by: Hwang, Kyomin, et al.
Published: (2024)
LLM should think and action as a human
by: Leung, Haun, et al.
Published: (2025)
by: Leung, Haun, et al.
Published: (2025)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
by: Seleznyov, Mikhail, et al.
Published: (2026)
by: Seleznyov, Mikhail, et al.
Published: (2026)
Multilingual and Explainable Text Detoxification with Parallel Corpora
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
by: Wan, Ziyu, et al.
Published: (2025)
by: Wan, Ziyu, et al.
Published: (2025)
Humans can learn to detect AI-generated texts, or at least learn when they can't
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
ReadCtrl: Personalizing text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2024)
by: Tran, Hieu, et al.
Published: (2024)
Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
by: Min, Yingqian, et al.
Published: (2024)
by: Min, Yingqian, et al.
Published: (2024)
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
S3: A Simple Strong Sample-effective Multimodal Dialog System
by: Rykov, Elisei, et al.
Published: (2024)
by: Rykov, Elisei, et al.
Published: (2024)
StyleBench: Evaluating thinking styles in Large Language Models
by: Guo, Junyu, et al.
Published: (2025)
by: Guo, Junyu, et al.
Published: (2025)
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
by: Seleznyov, Mikhail, et al.
Published: (2025)
by: Seleznyov, Mikhail, et al.
Published: (2025)
AIstorian lets AI be a historian: A KG-powered multi-agent system for accurate biography generation
by: Li, Fengyu, et al.
Published: (2025)
by: Li, Fengyu, et al.
Published: (2025)
SmurfCat at SemEval-2024 Task 6: Leveraging Synthetic Data for Hallucination Detection
by: Rykov, Elisei, et al.
Published: (2024)
by: Rykov, Elisei, et al.
Published: (2024)
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
People will agree what I think: Investigating LLM's False Consensus Effect
by: Choi, Junhyuk, et al.
Published: (2024)
by: Choi, Junhyuk, et al.
Published: (2024)
MERA: A Comprehensive LLM Evaluation in Russian
by: Fenogenova, Alena, et al.
Published: (2024)
by: Fenogenova, Alena, et al.
Published: (2024)
Abusive text transformation using LLMs
by: Chandra, Rohitash, et al.
Published: (2025)
by: Chandra, Rohitash, et al.
Published: (2025)
Benchmark of stylistic variation in LLM-generated texts
by: Milička, Jiří, et al.
Published: (2025)
by: Milička, Jiří, et al.
Published: (2025)
Are generative AI text annotations systematically biased?
by: Stolwijk, Sjoerd B., et al.
Published: (2025)
by: Stolwijk, Sjoerd B., et al.
Published: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
by: Badanin, Ilia, et al.
Published: (2026)
by: Badanin, Ilia, et al.
Published: (2026)
Detection of tortured phrases in scientific literature
by: Martel, Eléna, et al.
Published: (2024)
by: Martel, Eléna, et al.
Published: (2024)
Do great minds think alike? Investigating Human-AI Complementarity in Question Answering with CAIMIRA
by: Gor, Maharshi, et al.
Published: (2024)
by: Gor, Maharshi, et al.
Published: (2024)
Can professional translators identify machine-generated text?
by: Farrell, Michael
Published: (2026)
by: Farrell, Michael
Published: (2026)
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
by: Leidinger, Alina, et al.
Published: (2024)
by: Leidinger, Alina, et al.
Published: (2024)
System 2 thinking in OpenAI's o1-preview model: Near-perfect performance on a mathematics exam
by: de Winter, Joost, et al.
Published: (2024)
by: de Winter, Joost, et al.
Published: (2024)
Can LLMs replace Neil deGrasse Tyson? Evaluating the Reliability of LLMs as Science Communicators
by: Bajpai, Prasoon, et al.
Published: (2024)
by: Bajpai, Prasoon, et al.
Published: (2024)
Adapting Multilingual LLMs to Low-Resource Languages with Knowledge Graphs via Adapters
by: Gurgurov, Daniil, et al.
Published: (2024)
by: Gurgurov, Daniil, et al.
Published: (2024)
Can postgraduate translation students identify machine-generated text?
by: Farrell, Michael
Published: (2025)
by: Farrell, Michael
Published: (2025)
Differentially-private text generation degrades output language quality
by: Çano, Erion, et al.
Published: (2025)
by: Çano, Erion, et al.
Published: (2025)
SparseGrad: A Selective Method for Efficient Fine-tuning of MLP Layers
by: Chekalina, Viktoriia, et al.
Published: (2024)
by: Chekalina, Viktoriia, et al.
Published: (2024)
Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
by: Qian, Shun, et al.
Published: (2026)
by: Qian, Shun, et al.
Published: (2026)
Beyond speculation: Measuring the growing presence of LLM-generated texts in multilingual disinformation
by: Macko, Dominik, et al.
Published: (2025)
by: Macko, Dominik, et al.
Published: (2025)
From text to multimodal: a survey of adversarial example generation in question answering systems
by: Yigit, Gulsum, et al.
Published: (2023)
by: Yigit, Gulsum, et al.
Published: (2023)
ARC-Encoder: learning compressed text representations for large language models
by: Pilchen, Hippolyte, et al.
Published: (2025)
by: Pilchen, Hippolyte, et al.
Published: (2025)
Similar Items
-
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators
by: Moskovskiy, Daniil, et al.
Published: (2025) -
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
by: Vazhentsev, Artem, et al.
Published: (2026) -
How Much Knowledge Can You Pack into a LoRA Adapter without Harming LLM?
by: Pletenev, Sergey, et al.
Published: (2025) -
How Reliable are LLMs as Knowledge Bases? Re-thinking Facutality and Consistency
by: Zheng, Danna, et al.
Published: (2024) -
Scaling up the think-aloud method
by: Wurgaft, Daniel, et al.
Published: (2025)