Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Selvaganapathy, Sanjeeevan, Nasim, Mehwish |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Competing LLM Agents in a Non-Cooperative Game of Opinion Polarisation
by: Qasmi, Amin, et al.
Published: (2025)
by: Qasmi, Amin, et al.
Published: (2025)
CogCanvas: Verbatim-Grounded Artifact Extraction for Long LLM Conversations
by: An, Tao
Published: (2025)
by: An, Tao
Published: (2025)
Simulating Influence Dynamics with LLM Agents
by: Nasim, Mehwish, et al.
Published: (2025)
by: Nasim, Mehwish, et al.
Published: (2025)
Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation
by: Xie, Pengzhen, et al.
Published: (2025)
by: Xie, Pengzhen, et al.
Published: (2025)
ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks
by: Tanguturi, Samuel Sameer
Published: (2026)
by: Tanguturi, Samuel Sameer
Published: (2026)
ATANT: An Evaluation Framework for AI Continuity
by: Tanguturi, Samuel Sameer
Published: (2026)
by: Tanguturi, Samuel Sameer
Published: (2026)
Do LLMs Use Cultural Knowledge Without Being Told? A Multilingual Evaluation of Implicit Pragmatic Adaptation
by: Nasim, Mehwish, et al.
Published: (2026)
by: Nasim, Mehwish, et al.
Published: (2026)
LLMs in the Loop: Leveraging Large Language Model Annotations for Active Learning in Low-Resource Languages
by: Kholodna, Nataliia, et al.
Published: (2024)
by: Kholodna, Nataliia, et al.
Published: (2024)
LLM Reasoning for Cold-Start Item Recommendation
by: Li, Shijun, et al.
Published: (2025)
by: Li, Shijun, et al.
Published: (2025)
A Graph-based RAG for Energy Efficiency Question Answering
by: Campi, Riccardo, et al.
Published: (2025)
by: Campi, Riccardo, et al.
Published: (2025)
Improving RAG Retrieval via Propositional Content Extraction: a Speech Act Theory Approach
by: Lima, João Alberto de Oliveira
Published: (2025)
by: Lima, João Alberto de Oliveira
Published: (2025)
ConQRet: Benchmarking Fine-Grained Evaluation of Retrieval Augmented Argumentation with LLM Judges
by: Dhole, Kaustubh D., et al.
Published: (2024)
by: Dhole, Kaustubh D., et al.
Published: (2024)
A Semantic Approach to Negation Detection and Word Disambiguation with Natural Language Processing
by: Okpala, Izunna, et al.
Published: (2023)
by: Okpala, Izunna, et al.
Published: (2023)
Comparative Performance of Advanced NLP Models and LLMs in Multilingual Geo-Entity Detection
by: Kopanov, Kalin
Published: (2024)
by: Kopanov, Kalin
Published: (2024)
Improving the Performance of Sequential Recommendation Systems with an Extended Large Language Model
by: Choi, Sinnyum, et al.
Published: (2025)
by: Choi, Sinnyum, et al.
Published: (2025)
PatenTEB: A Comprehensive Benchmark and Model Family for Patent Text Embedding
by: Ayaou, Iliass, et al.
Published: (2025)
by: Ayaou, Iliass, et al.
Published: (2025)
Real-Time RAG for the Identification of Supply Chain Vulnerabilities
by: Ponnock, Jesse, et al.
Published: (2025)
by: Ponnock, Jesse, et al.
Published: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
GraphCompliance: Aligning Policy and Context Graphs for LLM-Based Regulatory Compliance
by: Chung, Jiseong, et al.
Published: (2025)
by: Chung, Jiseong, et al.
Published: (2025)
Leveraging Large Language Models to Extract and Translate Medical Information in Doctors' Notes for Health Records and Diagnostic Billing Codes
by: Hartnett, Peter, et al.
Published: (2026)
by: Hartnett, Peter, et al.
Published: (2026)
BiasScanner: Automatic Detection and Classification of News Bias to Strengthen Democracy
by: Menzner, Tim, et al.
Published: (2024)
by: Menzner, Tim, et al.
Published: (2024)
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
by: Agarwal, Mahak, et al.
Published: (2025)
by: Agarwal, Mahak, et al.
Published: (2025)
KGiRAG: An Iterative GraphRAG Approach for Responding Sensemaking Queries
by: Iacob, Isabela, et al.
Published: (2026)
by: Iacob, Isabela, et al.
Published: (2026)
Teaching a Language Model to Speak the Language of Tools
by: Emanuilov, Simeon
Published: (2025)
by: Emanuilov, Simeon
Published: (2025)
From Knowledge Generation to Knowledge Verification: Examining the BioMedical Generative Capabilities of ChatGPT
by: Hamed, Ahmed Abdeen, et al.
Published: (2025)
by: Hamed, Ahmed Abdeen, et al.
Published: (2025)
FinGround: Detecting and Grounding Financial Hallucinations via Atomic Claim Verification
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Overview of the ClinIQLink 2025 Shared Task on Medical Question-Answering
by: Colelough, Brandon, et al.
Published: (2025)
by: Colelough, Brandon, et al.
Published: (2025)
EnterpriseEM: Fine-tuned Embeddings for Enterprise Semantic Search
by: Rathinasamy, Kamalkumar, et al.
Published: (2024)
by: Rathinasamy, Kamalkumar, et al.
Published: (2024)
DeepSlide: From Artifacts to Presentation Delivery
by: Yang, Ming, et al.
Published: (2026)
by: Yang, Ming, et al.
Published: (2026)
SomaliWeb v1: A Quality-Filtered Somali Web Corpus with a Matched Tokenizer and a Public Language-Identification Benchmark
by: Dahir, Khalid Yusuf
Published: (2026)
by: Dahir, Khalid Yusuf
Published: (2026)
SegNSP: Revisiting Next Sentence Prediction for Linear Text Segmentation
by: Isidro, José, et al.
Published: (2026)
by: Isidro, José, et al.
Published: (2026)
A Comparative Analysis of Retrieval-Augmented Generation Techniques for Bengali Standard-to-Dialect Machine Translation Using LLMs
by: Sami, K. M. Jubair, et al.
Published: (2025)
by: Sami, K. M. Jubair, et al.
Published: (2025)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
by: Cai, Weibin, et al.
Published: (2025)
by: Cai, Weibin, et al.
Published: (2025)
Can LLM Agents Maintain a Persona in Discourse?
by: Bhandari, Pranav, et al.
Published: (2025)
by: Bhandari, Pranav, et al.
Published: (2025)
FinQAPT: Empowering Financial Decisions with End-to-End LLM-driven Question Answering Pipeline
by: Singh, Kuldeep, et al.
Published: (2024)
by: Singh, Kuldeep, et al.
Published: (2024)
Calibrated Confidence Estimation for Tabular Question Answering
by: Voss, Lukas
Published: (2026)
by: Voss, Lukas
Published: (2026)
SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation
by: Qiu, Jingxi, et al.
Published: (2026)
by: Qiu, Jingxi, et al.
Published: (2026)
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
by: Wu, Shuai, et al.
Published: (2026)
by: Wu, Shuai, et al.
Published: (2026)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
by: Bian, Zhipeng, et al.
Published: (2025)
by: Bian, Zhipeng, et al.
Published: (2025)
Mixture of Experts Approaches in Dense Retrieval Tasks
by: Sokli, Effrosyni, et al.
Published: (2025)
by: Sokli, Effrosyni, et al.
Published: (2025)
Similar Items
-
Competing LLM Agents in a Non-Cooperative Game of Opinion Polarisation
by: Qasmi, Amin, et al.
Published: (2025) -
CogCanvas: Verbatim-Grounded Artifact Extraction for Long LLM Conversations
by: An, Tao
Published: (2025) -
Simulating Influence Dynamics with LLM Agents
by: Nasim, Mehwish, et al.
Published: (2025) -
Graph Your Way to Inspiration: Integrating Co-Author Graphs with Retrieval-Augmented Generation for Large Language Model Based Scientific Idea Generation
by: Xie, Pengzhen, et al.
Published: (2025) -
ATANT v1.1: Positioning Continuity Evaluation Against Memory, Long-Context, and Agentic-Memory Benchmarks
by: Tanguturi, Samuel Sameer
Published: (2026)