DoDo Learning: DOmain-DemOgraphic Transfer in Language Models for Detecting Abuse Targeted at Public Figures
Fuente:
arXiv
Saved in:
| Main Authors: | Williams, Angus R., Kirk, Hannah Rose, Burke, Liam, Chung, Yi-Ling, Debono, Ivan, Johansson, Pica, Stevens, Francesca, Bright, Jonathan, Hale, Scott A. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Journalists are most likely to receive abuse: Analysing online abuse of UK public figures across sport, politics, and journalism on Twitter
by: Burke-Moore, Liam, et al.
Published: (2024)
by: Burke-Moore, Liam, et al.
Published: (2024)
Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
by: Castro-Gonzalez, Leonardo, et al.
Published: (2024)
by: Castro-Gonzalez, Leonardo, et al.
Published: (2024)
Understanding engagement with platform safety technology for reducing exposure to online harms
by: Bright, Jonathan, et al.
Published: (2024)
by: Bright, Jonathan, et al.
Published: (2024)
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025)
by: Rystrøm, Jonathan, et al.
Published: (2025)
Understanding gender differences in experiences and concerns surrounding online harms: A short report on a nationally representative survey of UK adults
by: Enock, Florence E., et al.
Published: (2024)
by: Enock, Florence E., et al.
Published: (2024)
Gendered Inequalities in Online Harms: Fear, Safety Work, and Online Participation
by: Enock, Florence E., et al.
Published: (2024)
by: Enock, Florence E., et al.
Published: (2024)
Exploring responsible applications of Synthetic Data to advance Online Safety Research and Development
by: Johansson, Pica, et al.
Published: (2024)
by: Johansson, Pica, et al.
Published: (2024)
Indian-BhED: A Dataset for Measuring India-Centric Biases in Large Language Models
by: Khandelwal, Khyati, et al.
Published: (2023)
by: Khandelwal, Khyati, et al.
Published: (2023)
Large language models can consistently generate high-quality content for election disinformation operations
by: Williams, Angus R., et al.
Published: (2024)
by: Williams, Angus R., et al.
Published: (2024)
How Well Do Large Language Models Disambiguate Swedish Words?
by: Johansson, Richard
Published: (2024)
by: Johansson, Richard
Published: (2024)
Why human-AI relationships need socioaffective alignment
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users
by: Kirk, Hannah Rose, et al.
Published: (2026)
by: Kirk, Hannah Rose, et al.
Published: (2026)
DoDo-Code: an Efficient Levenshtein Distance Embedding-based Code for 4-ary IDS Channel
by: Guo, Alan J. X., et al.
Published: (2023)
by: Guo, Alan J. X., et al.
Published: (2023)
Prompto: An open source library for asynchronous querying of LLM endpoints
by: Chan, Ryan Sze-Yin, et al.
Published: (2024)
by: Chan, Ryan Sze-Yin, et al.
Published: (2024)
What Is AI Safety? What Do We Want It to Be?
by: Harding, Jacqueline, et al.
Published: (2025)
by: Harding, Jacqueline, et al.
Published: (2025)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
by: Vidgen, Bertie, et al.
Published: (2023)
by: Vidgen, Bertie, et al.
Published: (2023)
Neural steering vectors reveal dose and exposure-dependent impacts of human-AI relationships
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
Measuring and Mitigating Persona Distortions from AI Writing Assistance
by: Röttger, Paul, et al.
Published: (2026)
by: Röttger, Paul, et al.
Published: (2026)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
by: Cyberey, Hannah, et al.
Published: (2024)
by: Cyberey, Hannah, et al.
Published: (2024)
LINGOLY: A Benchmark of Olympiad-Level Linguistic Reasoning Puzzles in Low-Resource and Extinct Languages
by: Bean, Andrew M., et al.
Published: (2024)
by: Bean, Andrew M., et al.
Published: (2024)
Beyond the Binary: Capturing Diverse Preferences With Reward Regularization
by: Padmakumar, Vishakh, et al.
Published: (2024)
by: Padmakumar, Vishakh, et al.
Published: (2024)
A Deep Learning Framework for Visual Attention Prediction and Analysis of News Interfaces
by: Kenely, Matthew, et al.
Published: (2025)
by: Kenely, Matthew, et al.
Published: (2025)
Disclosure By Design: Identity Transparency as a Behavioural Property of Conversational AI Models
by: Gausen, Anna, et al.
Published: (2026)
by: Gausen, Anna, et al.
Published: (2026)
DemOpts: Fairness corrections in COVID-19 case prediction models
by: Awasthi, Naman, et al.
Published: (2024)
by: Awasthi, Naman, et al.
Published: (2024)
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
by: Kirk, Hannah Rose, et al.
Published: (2024)
by: Kirk, Hannah Rose, et al.
Published: (2024)
The AI Community Building the Future? A Quantitative Analysis of Development Activity on Hugging Face Hub
by: Osborne, Cailean, et al.
Published: (2024)
by: Osborne, Cailean, et al.
Published: (2024)
Evidence of a log scaling law for political persuasion with large language models
by: Hackenburg, Kobi, et al.
Published: (2024)
by: Hackenburg, Kobi, et al.
Published: (2024)
DiaDem: Advancing Dialogue Descriptions in Audiovisual Video Captioning for Multimodal Large Language Models
by: Chen, Xinlong, et al.
Published: (2026)
by: Chen, Xinlong, et al.
Published: (2026)
Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models?
by: Modica, Luca, et al.
Published: (2026)
by: Modica, Luca, et al.
Published: (2026)
Hostility Detection in UK Politics: A Dataset on Online Abuse Targeting MPs
by: Pandya, Mugdha, et al.
Published: (2024)
by: Pandya, Mugdha, et al.
Published: (2024)
Do Large Language Models know who did what to whom?
by: Denning, Joseph M., et al.
Published: (2025)
by: Denning, Joseph M., et al.
Published: (2025)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
by: Attimonelli, Matteo, et al.
Published: (2026)
by: Attimonelli, Matteo, et al.
Published: (2026)
Do as We Do, Not as You Think: the Conformity of Large Language Models
by: Weng, Zhiyuan, et al.
Published: (2025)
by: Weng, Zhiyuan, et al.
Published: (2025)
LLMs Do Not See Age: Assessing Demographic Bias in Automated Systematic Review Synthesis
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
by: Aghaebe, Favour Yahdii, et al.
Published: (2025)
Doing Good or Doing Right? Exploring the Weakness of Commonsense Causal Reasoning Models
by: Han, Mingyue, et al.
Published: (2021)
by: Han, Mingyue, et al.
Published: (2021)
Where Do the Joules Go? Diagnosing Inference Energy Consumption
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
Large Language Models in the Abuse Detection Pipeline
by: Kath, Suraj, et al.
Published: (2026)
by: Kath, Suraj, et al.
Published: (2026)
Reference As Others Do It.
by: Coffman, Steve
Published: (1999)
by: Coffman, Steve
Published: (1999)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
by: Röttger, Paul, et al.
Published: (2023)
by: Röttger, Paul, et al.
Published: (2023)
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
by: Hsu, Ting-Yao E., et al.
Published: (2025)
by: Hsu, Ting-Yao E., et al.
Published: (2025)
Similar Items
-
Journalists are most likely to receive abuse: Analysing online abuse of UK public figures across sport, politics, and journalism on Twitter
by: Burke-Moore, Liam, et al.
Published: (2024) -
Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
by: Castro-Gonzalez, Leonardo, et al.
Published: (2024) -
Understanding engagement with platform safety technology for reducing exposure to online harms
by: Bright, Jonathan, et al.
Published: (2024) -
Multilingual != Multicultural: Evaluating Gaps Between Multilingual Capabilities and Cultural Alignment in LLMs
by: Rystrøm, Jonathan, et al.
Published: (2025) -
Understanding gender differences in experiences and concerns surrounding online harms: A short report on a nationally representative survey of UK adults
by: Enock, Florence E., et al.
Published: (2024)