HEARTS: A Holistic Framework for Explainable, Sustainable and Robust Text Stereotype Detection
Fuente:
arXiv
Saved in:
| Main Authors: | King, Theo, Wu, Zekun, Koshiyama, Adriano, Kazim, Emre, Treleaven, Philip |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Text to Emoji: How PEFT-Driven Personality Manipulation Unleashes the Emoji Potential in LLMs
by: Jain, Navya, et al.
Published: (2024)
by: Jain, Navya, et al.
Published: (2024)
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
by: Wicaksono, Ilham, et al.
Published: (2025)
by: Wicaksono, Ilham, et al.
Published: (2025)
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
by: Handa, Gunmay, et al.
Published: (2025)
by: Handa, Gunmay, et al.
Published: (2025)
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach
by: Wu, Zekun, et al.
Published: (2024)
by: Wu, Zekun, et al.
Published: (2024)
THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models
by: Liang, Mengfei, et al.
Published: (2024)
by: Liang, Mengfei, et al.
Published: (2024)
Knowledge Collapse in LLMs: When Fluency Survives but Facts Fail under Recursive Synthetic Training
by: Keisha, Figarri, et al.
Published: (2025)
by: Keisha, Figarri, et al.
Published: (2025)
Assessing Bias in Metric Models for LLM Open-Ended Generation Bias Benchmarks
by: Demchak, Nathaniel, et al.
Published: (2024)
by: Demchak, Nathaniel, et al.
Published: (2024)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
by: Guan, Xin, et al.
Published: (2024)
by: Guan, Xin, et al.
Published: (2024)
LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries
by: Wu, Zekun, et al.
Published: (2025)
by: Wu, Zekun, et al.
Published: (2025)
MPF: Aligning and Debiasing Language Models post Deployment via Multi Perspective Fusion
by: Guan, Xin, et al.
Published: (2025)
by: Guan, Xin, et al.
Published: (2025)
Mind the Gap: Comparing Model- vs Agentic-Level Red Teaming with Action-Graph Observability on GPT-OSS-20B
by: Wicaksono, Ilham, et al.
Published: (2025)
by: Wicaksono, Ilham, et al.
Published: (2025)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2025)
by: Cho, Seonglae, et al.
Published: (2025)
Eliciting Personality Traits in Large Language Models
by: Hilliard, Airlie, et al.
Published: (2024)
by: Hilliard, Airlie, et al.
Published: (2024)
The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models
by: Wang, Ze, et al.
Published: (2024)
by: Wang, Ze, et al.
Published: (2024)
BERT vs GPT for financial engineering
by: Sharkey, Edward, et al.
Published: (2024)
by: Sharkey, Edward, et al.
Published: (2024)
HyPA-RAG: A Hybrid Parameter Adaptive Retrieval-Augmented Generation System for AI Legal and Policy Applications
by: Kalra, Rishi, et al.
Published: (2024)
by: Kalra, Rishi, et al.
Published: (2024)
The Effect of Model Size on LLM Post-hoc Explainability via LIME
by: Heyen, Henning, et al.
Published: (2024)
by: Heyen, Henning, et al.
Published: (2024)
Tool Calling is Linearly Readable and Steerable in Language Models
by: Wu, Zekun, et al.
Published: (2026)
by: Wu, Zekun, et al.
Published: (2026)
StyleDecipher: Robust and Explainable Detection of LLM-Generated Texts with Stylistic Analysis
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
LM$^2$otifs : An Explainable Framework for Machine-Generated Texts Detection
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
Cultural Alignment in Large Language Models Using Soft Prompt Tuning
by: Masoud, Reem I., et al.
Published: (2025)
by: Masoud, Reem I., et al.
Published: (2025)
A Survey on Stereotype Detection in Natural Language Processing
by: Cignarella, Alessandra Teresa, et al.
Published: (2025)
by: Cignarella, Alessandra Teresa, et al.
Published: (2025)
Robust Detection of LLM-Generated Text: A Comparative Analysis
by: Su, Yongye, et al.
Published: (2024)
by: Su, Yongye, et al.
Published: (2024)
FairMonitor: A Dual-framework for Detecting Stereotypes and Biases in Large Language Models
by: Bai, Yanhong, et al.
Published: (2024)
by: Bai, Yanhong, et al.
Published: (2024)
Prompting Away Stereotypes? Evaluating Bias in Text-to-Image Models for Occupations
by: Raza, Shaina, et al.
Published: (2025)
by: Raza, Shaina, et al.
Published: (2025)
Simulating Identity, Propagating Bias: Abstraction and Stereotypes in LLM-Generated Text
by: Sommerauer, Pia, et al.
Published: (2025)
by: Sommerauer, Pia, et al.
Published: (2025)
A Modular LLM Framework for Explainable Price Outlier Detection
by: Sartipi, Shadi, et al.
Published: (2026)
by: Sartipi, Shadi, et al.
Published: (2026)
Cyberbullying Detection in Hinglish Text Using MURIL and Explainable AI
by: Kumar, Devesh
Published: (2025)
by: Kumar, Devesh
Published: (2025)
Applying Ensemble Methods to Model-Agnostic Machine-Generated Text Detection
by: Ong, Ivan, et al.
Published: (2024)
by: Ong, Ivan, et al.
Published: (2024)
Challenging Negative Gender Stereotypes: A Study on the Effectiveness of Automated Counter-Stereotypes
by: Nejadgholi, Isar, et al.
Published: (2024)
by: Nejadgholi, Isar, et al.
Published: (2024)
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach
by: Tomar, Aditya, et al.
Published: (2025)
by: Tomar, Aditya, et al.
Published: (2025)
Cultural Alignment in Large Language Models: An Explanatory Analysis Based on Hofstede's Cultural Dimensions
by: Masoud, Reem I., et al.
Published: (2023)
by: Masoud, Reem I., et al.
Published: (2023)
SKETCH: Structured Knowledge Enhanced Text Comprehension for Holistic Retrieval
by: Mahalingam, Aakash, et al.
Published: (2024)
by: Mahalingam, Aakash, et al.
Published: (2024)
Quantifying Stereotypes in Language
by: Liu, Yang
Published: (2024)
by: Liu, Yang
Published: (2024)
An Entropy-based Text Watermarking Detection Method
by: Lu, Yijian, et al.
Published: (2024)
by: Lu, Yijian, et al.
Published: (2024)
Explainable Semantic Text Relations: A Question-Answering Framework for Comparing Document Content
by: Aperstein, Yehudit, et al.
Published: (2025)
by: Aperstein, Yehudit, et al.
Published: (2025)
A Grey-box Text Attack Framework using Explainable AI
by: Chiramal, Esther, et al.
Published: (2025)
by: Chiramal, Esther, et al.
Published: (2025)
Are Stereotypes Leading LLMs' Zero-Shot Stance Detection ?
by: Dubreuil, Anthony, et al.
Published: (2025)
by: Dubreuil, Anthony, et al.
Published: (2025)
Similar Items
-
From Text to Emoji: How PEFT-Driven Personality Manipulation Unleashes the Emoji Potential in LLMs
by: Jain, Navya, et al.
Published: (2024) -
Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs
by: Wicaksono, Ilham, et al.
Published: (2025) -
Personality as a Probe for LLM Evaluation: Method Trade-offs and Downstream Effects
by: Handa, Gunmay, et al.
Published: (2025) -
Stereotype Detection in LLMs: A Multiclass, Explainable, and Benchmark-Driven Approach
by: Wu, Zekun, et al.
Published: (2024) -
THaMES: An End-to-End Tool for Hallucination Mitigation and Evaluation in Large Language Models
by: Liang, Mengfei, et al.
Published: (2024)