Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach
Fuente:
arXiv
Guardado en:
| Autores principales: | Wu, Zekun, Bulathwela, Sahan, Perez-Ortiz, Maria, Koshiyama, Adriano Soares |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Tool Calling is Linearly Readable and Steerable in Language Models
por: Wu, Zekun, et al.
Publicado: (2026)
por: Wu, Zekun, et al.
Publicado: (2026)
Explainable Collaborative Problem Solving Diagnosis with BERT using SHAP and its Implications for Teacher Adoption
por: Wong, Kester, et al.
Publicado: (2025)
por: Wong, Kester, et al.
Publicado: (2025)
Eliciting Personality Traits in Large Language Models
por: Hilliard, Airlie, et al.
Publicado: (2024)
por: Hilliard, Airlie, et al.
Publicado: (2024)
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
por: King, Theo, et al.
Publicado: (2024)
por: King, Theo, et al.
Publicado: (2024)
Exploring Human-AI Complementarity in CPS Diagnosis Using Unimodal and Multimodal BERT Models
por: Wong, Kester, et al.
Publicado: (2025)
por: Wong, Kester, et al.
Publicado: (2025)
Next Token Knowledge Tracing: Exploiting Pretrained LLM Representations to Decode Student Behaviour
por: Norris, Max, et al.
Publicado: (2025)
por: Norris, Max, et al.
Publicado: (2025)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2026)
por: Cho, Seonglae, et al.
Publicado: (2026)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
por: Demchak, Nathaniel, et al.
Publicado: (2024)
por: Demchak, Nathaniel, et al.
Publicado: (2024)
Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning
por: Imran, Moiz, et al.
Publicado: (2026)
por: Imran, Moiz, et al.
Publicado: (2026)
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
por: Wang, Ze, et al.
Publicado: (2024)
por: Wang, Ze, et al.
Publicado: (2024)
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
por: Cho, Seonglae, et al.
Publicado: (2026)
por: Cho, Seonglae, et al.
Publicado: (2026)
A Novel Approach to Scalable and Automatic Topic-Controlled Question Generation in Education
por: Li, Ziqing, et al.
Publicado: (2025)
por: Li, Ziqing, et al.
Publicado: (2025)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
por: Cho, Seonglae, et al.
Publicado: (2025)
por: Cho, Seonglae, et al.
Publicado: (2025)
Mix and Match: Context Pairing for Scalable Topic-Controlled Educational Summarisation
por: Yodthapa, Nathikan, et al.
Publicado: (2026)
por: Yodthapa, Nathikan, et al.
Publicado: (2026)
MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion
por: Guan, Xin, et al.
Publicado: (2025)
por: Guan, Xin, et al.
Publicado: (2025)
Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?
por: Dubreuil, Anthony, et al.
Publicado: (2025)
por: Dubreuil, Anthony, et al.
Publicado: (2025)
AI-generated Text Detection: A Multifaceted Approach to Binary and Multiclass Classification
por: Abburi, Harika, et al.
Publicado: (2025)
por: Abburi, Harika, et al.
Publicado: (2025)
Can We Locate and Prevent Stereotypes in LLMs?
por: D'Souza, Alex
Publicado: (2026)
por: D'Souza, Alex
Publicado: (2026)
Rethinking the Potential of Multimodality in Collaborative Problem Solving Diagnosis with Large Language Models
por: Wong, K., et al.
Publicado: (2025)
por: Wong, K., et al.
Publicado: (2025)
WebWalker: Benchmarking LLMs in Web Traversal
por: Wu, Jialong, et al.
Publicado: (2025)
por: Wu, Jialong, et al.
Publicado: (2025)
Towards Synergistic Teacher-AI Interactions with Generative Artificial Intelligence
por: Cukurova, Mutlu, et al.
Publicado: (2025)
por: Cukurova, Mutlu, et al.
Publicado: (2025)
Responsible AI in NLP: GUS-Net Span-Level Bias Detection Dataset and Benchmark for Generalizations, Unfairness, and Stereotypes
por: Powers, Maximus, et al.
Publicado: (2024)
por: Powers, Maximus, et al.
Publicado: (2024)
CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark
por: Zhang, Junzhao, et al.
Publicado: (2026)
por: Zhang, Junzhao, et al.
Publicado: (2026)
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets
por: Zakizadeh, Mahdi, et al.
Publicado: (2025)
por: Zakizadeh, Mahdi, et al.
Publicado: (2025)
Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training
por: Keisha, Figarri, et al.
Publicado: (2025)
por: Keisha, Figarri, et al.
Publicado: (2025)
Benchmarking the Detection of LLMs-Generated Modern Chinese Poetry
por: Wang, Shanshan, et al.
Publicado: (2025)
por: Wang, Shanshan, et al.
Publicado: (2025)
Detecting Linguistic Indicators for Stereotype Assessment with Large Language Models
por: Görge, Rebekka, et al.
Publicado: (2025)
por: Görge, Rebekka, et al.
Publicado: (2025)
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
por: Jeune, Pierre Le, et al.
Publicado: (2026)
por: Jeune, Pierre Le, et al.
Publicado: (2026)
Beyond Accuracy: An Explainability-Driven Analysis of Harmful Content Detection
por: Dhara, Trishita, et al.
Publicado: (2026)
por: Dhara, Trishita, et al.
Publicado: (2026)
LLMs for Explainable Business Decision-Making: A Reinforcement Learning Fine-Tuning Approach
por: Cheng, Xiang, et al.
Publicado: (2025)
por: Cheng, Xiang, et al.
Publicado: (2025)
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
por: Handa, Gunmay, et al.
Publicado: (2025)
por: Handa, Gunmay, et al.
Publicado: (2025)
Hierarchical Sentiment Analysis Framework for Hate Speech Detection: Implementing Binary and Multiclass Classification Strategy
por: Naznin, Faria, et al.
Publicado: (2024)
por: Naznin, Faria, et al.
Publicado: (2024)
Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection
por: Wang, Bing, et al.
Publicado: (2026)
por: Wang, Bing, et al.
Publicado: (2026)
LLMs for Explainable AI: A Comprehensive Survey
por: Bilal, Ahsan, et al.
Publicado: (2025)
por: Bilal, Ahsan, et al.
Publicado: (2025)
RealMem: Benchmarking LLMs in Real-World Memory-Driven Interaction
por: Bian, Haonan, et al.
Publicado: (2026)
por: Bian, Haonan, et al.
Publicado: (2026)
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
por: Wicaksono, Ilham, et al.
Publicado: (2025)
por: Wicaksono, Ilham, et al.
Publicado: (2025)
A Benchmark for the Detection of Metalinguistic Disagreements between LLMs and Knowledge Graphs
por: Allen, Bradley P., et al.
Publicado: (2025)
por: Allen, Bradley P., et al.
Publicado: (2025)
From Text to Emoji: How PEFT-Driven Personality Manipulation Unleashes the Emoji Potential in LLMs
por: Jain, Navya, et al.
Publicado: (2024)
por: Jain, Navya, et al.
Publicado: (2024)
TrueReason: An Exemplar Personalised Learning System Integrating Reasoning with Foundational Models
por: Bulathwela, Sahan, et al.
Publicado: (2025)
por: Bulathwela, Sahan, et al.
Publicado: (2025)
BELL: Benchmarking the Explainability of Large Language Models
por: Ahmed, Syed Quiser, et al.
Publicado: (2025)
por: Ahmed, Syed Quiser, et al.
Publicado: (2025)
Ejemplares similares
-
Tool Calling is Linearly Readable and Steerable in Language Models
por: Wu, Zekun, et al.
Publicado: (2026) -
Explainable Collaborative Problem Solving Diagnosis with BERT using SHAP and its Implications for Teacher Adoption
por: Wong, Kester, et al.
Publicado: (2025) -
Eliciting Personality Traits in Large Language Models
por: Hilliard, Airlie, et al.
Publicado: (2024) -
HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
por: King, Theo, et al.
Publicado: (2024) -
Exploring Human-AI Complementarity in CPS Diagnosis Using Unimodal and Multimodal BERT Models
por: Wong, Kester, et al.
Publicado: (2025)