A Multilingual, Large-Scale Study of the Interplay between LLM Safeguards, Personalisation, and Disinformation
Fuente:
arXiv
Saved in:
| Main Authors: | Leite, João A., Arora, Arnav, Gargova, Silvia, Luz, João, Sampaio, Gustavo, Roberts, Ian, Scarton, Carolina, Bontcheva, Kalina |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EUvsDisinfo: A Dataset for Multilingual Detection of Pro-Kremlin Disinformation in News Articles
by: Leite, João A., et al.
Published: (2024)
by: Leite, João A., et al.
Published: (2024)
A Cross-Domain Study of the Use of Persuasion Techniques in Online Disinformation
by: Leite, João A., et al.
Published: (2024)
by: Leite, João A., et al.
Published: (2024)
Weakly supervised veracity classification with LLM-predicted credibility signals
by: Leite, João Augusto, et al.
Published: (2025)
by: Leite, João Augusto, et al.
Published: (2025)
LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems
by: Leite, João A., et al.
Published: (2026)
by: Leite, João A., et al.
Published: (2026)
Weakly Supervised Veracity Classification with LLM-Predicted Credibility Signals
by: Leite, João A., et al.
Published: (2023)
by: Leite, João A., et al.
Published: (2023)
GateNLP at SemEval-2025 Task 10: Hierarchical Three-Step Prompting for Multilingual Narrative Classification
by: Singh, Iknoor, et al.
Published: (2025)
by: Singh, Iknoor, et al.
Published: (2025)
Lying Blindly: Bypassing ChatGPT's Safeguards to Generate Hard-to-Detect Disinformation Claims
by: Heppell, Freddy, et al.
Published: (2024)
by: Heppell, Freddy, et al.
Published: (2024)
The False COVID-19 Narratives That Keep Being Debunked: A Spatiotemporal Analysis
by: Singh, Iknoor, et al.
Published: (2021)
by: Singh, Iknoor, et al.
Published: (2021)
A Lightweight Approach for User and Keyword Classification in Controversial Topics
by: Zareie, Ahmad, et al.
Published: (2025)
by: Zareie, Ahmad, et al.
Published: (2025)
Breaking Language Barriers with MMTweets: Advancing Cross-Lingual Debunked Narrative Retrieval for Fact-Checking
by: Singh, Iknoor, et al.
Published: (2023)
by: Singh, Iknoor, et al.
Published: (2023)
Comparison between parameter-efficient techniques and full fine-tuning: A case study on multilingual news article classification
by: Razuvayevskaya, Olesya, et al.
Published: (2023)
by: Razuvayevskaya, Olesya, et al.
Published: (2023)
ExU: AI Models for Examining Multilingual Disinformation Narratives and Understanding their Spread
by: Vasilakes, Jake, et al.
Published: (2024)
by: Vasilakes, Jake, et al.
Published: (2024)
Navigating Prompt Complexity for Zero-Shot Classification: A Study of Large Language Models in Computational Social Science
by: Mu, Yida, et al.
Published: (2023)
by: Mu, Yida, et al.
Published: (2023)
Truth with a Twist: The Rhetoric of Persuasion in Professional vs. Community-Authored Fact-Checks
by: Razuvayevskaya, Olesya, et al.
Published: (2026)
by: Razuvayevskaya, Olesya, et al.
Published: (2026)
UKElectionNarratives: A Dataset of Misleading Narratives Surrounding Recent UK General Elections
by: Haouari, Fatima, et al.
Published: (2025)
by: Haouari, Fatima, et al.
Published: (2025)
A Survey on Automatic Credibility Assessment Using Textual Credibility Signals in the Era of Large Language Models
by: Srba, Ivan, et al.
Published: (2024)
by: Srba, Ivan, et al.
Published: (2024)
Large Language Models Offer an Alternative to the Traditional Approach of Topic Modelling
by: Mu, Yida, et al.
Published: (2024)
by: Mu, Yida, et al.
Published: (2024)
Addressing Topic Granularity and Hallucination in Large Language Models for Topic Modelling
by: Mu, Yida, et al.
Published: (2024)
by: Mu, Yida, et al.
Published: (2024)
Exploring Vision Language Models for Multimodal and Multilingual Stance Detection
by: Vasilakes, Jake, et al.
Published: (2025)
by: Vasilakes, Jake, et al.
Published: (2025)
Examining the Limitations of Computational Rumor Detection Models Trained on Static Datasets
by: Mu, Yida, et al.
Published: (2023)
by: Mu, Yida, et al.
Published: (2023)
Examining Temporalities on Stance Detection towards COVID-19 Vaccination
by: Mu, Yida, et al.
Published: (2023)
by: Mu, Yida, et al.
Published: (2023)
Hostility Detection in UK Politics: A Dataset on Online Abuse Targeting MPs
by: Pandya, Mugdha, et al.
Published: (2024)
by: Pandya, Mugdha, et al.
Published: (2024)
Can We Identify Stance Without Target Arguments? A Study for Rumour Stance Classification
by: Li, Yue, et al.
Published: (2023)
by: Li, Yue, et al.
Published: (2023)
Reference-less Analysis of Context Specificity in Translation with Personalised Language Models
by: Vincent, Sebastian, et al.
Published: (2023)
by: Vincent, Sebastian, et al.
Published: (2023)
A Browser-based Open Source Assistant for Multimodal Content Verification
by: Milner, Rosanna, et al.
Published: (2026)
by: Milner, Rosanna, et al.
Published: (2026)
SCRum-9: Multilingual Stance Classification over Rumours on Social Media
by: Li, Yue, et al.
Published: (2025)
by: Li, Yue, et al.
Published: (2025)
Why Should This Article Be Deleted? Transparent Stance Detection in Multilingual Wikipedia Editor Discussions
by: Kaffee, Lucie-Aimée, et al.
Published: (2023)
by: Kaffee, Lucie-Aimée, et al.
Published: (2023)
Safeguarding Marketing Research: The Generation, Identification, and Mitigation of AI-Fabricated Disinformation
by: Mukherjee, Anirban
Published: (2024)
by: Mukherjee, Anirban
Published: (2024)
Leveraging Large Language Models for Zero-shot Lay Summarisation in Biomedicine and Beyond
by: Goldsack, Tomas, et al.
Published: (2025)
by: Goldsack, Tomas, et al.
Published: (2025)
A Dataset for Analysing News Framing in Chinese Media
by: Cook, Owen, et al.
Published: (2025)
by: Cook, Owen, et al.
Published: (2025)
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
by: Fonseca, Joao, et al.
Published: (2025)
by: Fonseca, Joao, et al.
Published: (2025)
Beyond the Battlefield: Framing Analysis of Media Coverage in Conflict Reporting
by: Kaur, Avneet, et al.
Published: (2025)
by: Kaur, Avneet, et al.
Published: (2025)
Timeliness, Consensus, and Composition of the Crowd: Community Notes on X
by: Razuvayevskaya, Olesya, et al.
Published: (2025)
by: Razuvayevskaya, Olesya, et al.
Published: (2025)
FedMAP: Personalised Federated Learning for Real Large-Scale Healthcare Systems
by: Zhang, Fan, et al.
Published: (2024)
by: Zhang, Fan, et al.
Published: (2024)
Analysis of Patent Examination Effort Distribution based on the Queuing Theory
by: João Gilberto Sampaio
Published: (2008)
by: João Gilberto Sampaio
Published: (2008)
How Value Induction Reshapes LLM Behaviour
by: Arora, Arnav, et al.
Published: (2026)
by: Arora, Arnav, et al.
Published: (2026)
SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for LLM Software Systems
by: Shan, Wenliang, et al.
Published: (2025)
by: Shan, Wenliang, et al.
Published: (2025)
Presumed Cultural Identity: How Names Shape LLM Responses
by: Pawar, Siddhesh, et al.
Published: (2025)
by: Pawar, Siddhesh, et al.
Published: (2025)
Narratives at Conflict: Computational Analysis of News Framing in Multilingual Disinformation Campaigns
by: Sinelnik, Antonina, et al.
Published: (2024)
by: Sinelnik, Antonina, et al.
Published: (2024)
Interplay between Electroweak Symmetry Breaking and Higgs Portal Dark Matter
by: Chakraborti, Sreemanti, et al.
Published: (2025)
by: Chakraborti, Sreemanti, et al.
Published: (2025)
Similar Items
-
EUvsDisinfo: A Dataset for Multilingual Detection of Pro-Kremlin Disinformation in News Articles
by: Leite, João A., et al.
Published: (2024) -
A Cross-Domain Study of the Use of Persuasion Techniques in Online Disinformation
by: Leite, João A., et al.
Published: (2024) -
Weakly supervised veracity classification with LLM-predicted credibility signals
by: Leite, João Augusto, et al.
Published: (2025) -
LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems
by: Leite, João A., et al.
Published: (2026) -
Weakly Supervised Veracity Classification with LLM-Predicted Credibility Signals
by: Leite, João A., et al.
Published: (2023)