Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
Fuente:
arXiv
Saved in:
| Main Authors: | Yadav, Neemesh, Liu, Jiarui, Ortu, Francesco, Ensafi, Roya, Jin, Zhijing, Mihalcea, Rada |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are LLMs Good Safety Agents or a Propaganda Engine?
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
The Curious Case of Curiosity across Human Cultures and LLMs
by: Borah, Angana, et al.
Published: (2025)
by: Borah, Angana, et al.
Published: (2025)
Language Model Alignment in Multilingual Trolley Problems
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
by: Arif, Samee, et al.
Published: (2026)
by: Arif, Samee, et al.
Published: (2026)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
by: Ortu, Francesco, et al.
Published: (2024)
by: Ortu, Francesco, et al.
Published: (2024)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Causality for Natural Language Processing
by: Jin, Zhijing
Published: (2025)
by: Jin, Zhijing
Published: (2025)
Implicit Personalization in Language Models: A Systematic Study
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Voices of Her: Analyzing Gender Differences in the AI Publication World
by: Ding, Yiwen, et al.
Published: (2023)
by: Ding, Yiwen, et al.
Published: (2023)
Are Language Models Consequentialist or Deontological Moral Reasoners?
by: Samway, Keenan, et al.
Published: (2025)
by: Samway, Keenan, et al.
Published: (2025)
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis
by: Lyu, Zhiheng, et al.
Published: (2024)
by: Lyu, Zhiheng, et al.
Published: (2024)
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias
by: Chen, Yuen, et al.
Published: (2022)
by: Chen, Yuen, et al.
Published: (2022)
When Do Language Models Endorse Limitations on Human Rights Principles?
by: Samway, Keenan, et al.
Published: (2026)
by: Samway, Keenan, et al.
Published: (2026)
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
by: Piatti, Giorgio, et al.
Published: (2024)
by: Piatti, Giorgio, et al.
Published: (2024)
Are Human Interactions Replicable by Generative Agents? A Case Study on Pronoun Usage in Hierarchical Interactions
by: Deng, Naihao, et al.
Published: (2025)
by: Deng, Naihao, et al.
Published: (2025)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
by: Borah, Angana, et al.
Published: (2024)
by: Borah, Angana, et al.
Published: (2024)
The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models
by: Arif, Samee, et al.
Published: (2026)
by: Arif, Samee, et al.
Published: (2026)
Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification
by: Abzaliev, Artem, et al.
Published: (2024)
by: Abzaliev, Artem, et al.
Published: (2024)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
by: Backmann, Steffen, et al.
Published: (2025)
by: Backmann, Steffen, et al.
Published: (2025)
Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
by: Ortu, Francesco, et al.
Published: (2026)
by: Ortu, Francesco, et al.
Published: (2026)
Rethinking Table Instruction Tuning
by: Deng, Naihao, et al.
Published: (2025)
by: Deng, Naihao, et al.
Published: (2025)
Cross-cultural Inspiration Detection and Analysis in Real and LLM-generated Social Media Data
by: Ignat, Oana, et al.
Published: (2024)
by: Ignat, Oana, et al.
Published: (2024)
Patient-Centered RAG for Oncology Visit Aid Following the Ottawa Decision Guide
by: Liu, Siyang, et al.
Published: (2025)
by: Liu, Siyang, et al.
Published: (2025)
Towards Region-aware Bias Evaluation Metrics
by: Borah, Angana, et al.
Published: (2024)
by: Borah, Angana, et al.
Published: (2024)
Whose wife is it anyway? Assessing bias against same-gender relationships in machine translation
by: Stewart, Ian, et al.
Published: (2024)
by: Stewart, Ian, et al.
Published: (2024)
CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation
by: Deng, Naihao, et al.
Published: (2025)
by: Deng, Naihao, et al.
Published: (2025)
Mind the (Belief) Gap: Group Identity in the World of LLMs
by: Borah, Angana, et al.
Published: (2025)
by: Borah, Angana, et al.
Published: (2025)
MAiDE-up: Multilingual Deception Detection of GPT-generated Hotel Reviews
by: Ignat, Oana, et al.
Published: (2024)
by: Ignat, Oana, et al.
Published: (2024)
Persuasion at Play: Understanding Misinformation Dynamics in Demographic-Aware Human-LLM Interactions
by: Borah, Angana, et al.
Published: (2025)
by: Borah, Angana, et al.
Published: (2025)
Automatic Generation of Model and Data Cards: A Step Towards Responsible AI
by: Liu, Jiarui, et al.
Published: (2024)
by: Liu, Jiarui, et al.
Published: (2024)
CLoVe: Encoding Compositional Language in Contrastive Vision-Language Models
by: Castro, Santiago, et al.
Published: (2024)
by: Castro, Santiago, et al.
Published: (2024)
QUENCH: Measuring the gap between Indic and Non-Indic Contextual General Reasoning in LLMs
by: Khan, Mohammad Aflah, et al.
Published: (2024)
by: Khan, Mohammad Aflah, et al.
Published: (2024)
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
by: Lee, Suhyun, et al.
Published: (2026)
by: Lee, Suhyun, et al.
Published: (2026)
Taming Object Hallucinations with Verified Atomic Confidence Estimation
by: Liu, Jiarui, et al.
Published: (2025)
by: Liu, Jiarui, et al.
Published: (2025)
Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
by: Ortu, Francesco, et al.
Published: (2025)
by: Ortu, Francesco, et al.
Published: (2025)
The Hidden Language of Harm: Examining the Role of Emojis in Harmful Online Communication and Content Moderation
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
by: Shen, Siqi, et al.
Published: (2024)
by: Shen, Siqi, et al.
Published: (2024)
Similar Items
-
Are LLMs Good Safety Agents or a Propaganda Engine?
by: Yadav, Neemesh, et al.
Published: (2025) -
The Curious Case of Curiosity across Human Cultures and LLMs
by: Borah, Angana, et al.
Published: (2025) -
Language Model Alignment in Multilingual Trolley Problems
by: Jin, Zhijing, et al.
Published: (2024) -
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
by: Arif, Samee, et al.
Published: (2026) -
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
by: Piedrahita, David Guzman, et al.
Published: (2025)