Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
Fuente:
arXiv
Saved in:
| Main Authors: | Ahsan, Hiba, Wallace, Byron C. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025)
by: Muhamed, Aashiq, et al.
Published: (2025)
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024)
by: Lawson, Tim, et al.
Published: (2024)
Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
by: Ahsan, Hiba, et al.
Published: (2025)
by: Ahsan, Hiba, et al.
Published: (2025)
Resa: Transparent Reasoning Models via SAEs
by: Wang, Shangshang, et al.
Published: (2025)
by: Wang, Shangshang, et al.
Published: (2025)
SAEs Are Good for Steering -- If You Select the Right Features
by: Arad, Dana, et al.
Published: (2025)
by: Arad, Dana, et al.
Published: (2025)
Teach Old SAEs New Domain Tricks with Boosting
by: Koriagin, Nikita, et al.
Published: (2025)
by: Koriagin, Nikita, et al.
Published: (2025)
Token Erasure as a Footprint of Implicit Vocabulary Items in LLMs
by: Feucht, Sheridan, et al.
Published: (2024)
by: Feucht, Sheridan, et al.
Published: (2024)
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs
by: Song, Xiangchen, et al.
Published: (2025)
by: Song, Xiangchen, et al.
Published: (2025)
Compared to What? Baselines and Metrics for Counterfactual Prompting
by: Yang, Zihao, et al.
Published: (2026)
by: Yang, Zihao, et al.
Published: (2026)
Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank
by: Roy, Debjyoti Saha, et al.
Published: (2024)
by: Roy, Debjyoti Saha, et al.
Published: (2024)
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State
by: Pal, Koyena, et al.
Published: (2023)
by: Pal, Koyena, et al.
Published: (2023)
Retrieving Evidence from EHRs with LLMs: Possibilities and Challenges
by: Ahsan, Hiba, et al.
Published: (2023)
by: Ahsan, Hiba, et al.
Published: (2023)
Evaluating the Factuality of Zero-shot Summarizers Across Varied Domains
by: Ramprasad, Sanjana, et al.
Published: (2024)
by: Ramprasad, Sanjana, et al.
Published: (2024)
Do Activation Verbalization Methods Convey Privileged Information?
by: Li, Millicent, et al.
Published: (2025)
by: Li, Millicent, et al.
Published: (2025)
GenAudit: Fixing Factual Errors in Language Model Outputs with Evidence
by: Krishna, Kundan, et al.
Published: (2024)
by: Krishna, Kundan, et al.
Published: (2024)
Can LLMs subtract numbers?
by: Jobanputra, Mayank, et al.
Published: (2025)
by: Jobanputra, Mayank, et al.
Published: (2025)
Function Vectors in Large Language Models
by: Todd, Eric, et al.
Published: (2023)
by: Todd, Eric, et al.
Published: (2023)
Can Multimodal LLMs Perform Time Series Anomaly Detection?
by: Xu, Xiongxiao, et al.
Published: (2025)
by: Xu, Xiongxiao, et al.
Published: (2025)
Exploration Hacking: Can LLMs Learn to Resist RL Training?
by: Jang, Eyon, et al.
Published: (2026)
by: Jang, Eyon, et al.
Published: (2026)
Can Public LLMs be used for Self-Diagnosis of Medical Conditions ?
by: Balasubramanian, Nikil Sharan Prabahar, et al.
Published: (2024)
by: Balasubramanian, Nikil Sharan Prabahar, et al.
Published: (2024)
Can Knowledge Graphs Reduce Hallucinations in LLMs? : A Survey
by: Agrawal, Garima, et al.
Published: (2023)
by: Agrawal, Garima, et al.
Published: (2023)
Can we Soft Prompt LLMs for Graph Learning Tasks?
by: Liu, Zheyuan, et al.
Published: (2024)
by: Liu, Zheyuan, et al.
Published: (2024)
Can LLMs Learn New Concepts Incrementally without Forgetting?
by: Zheng, Junhao, et al.
Published: (2024)
by: Zheng, Junhao, et al.
Published: (2024)
Coercing LLMs to do and reveal (almost) anything
by: Geiping, Jonas, et al.
Published: (2024)
by: Geiping, Jonas, et al.
Published: (2024)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
by: Roy, Tiasa Singha, et al.
Published: (2025)
by: Roy, Tiasa Singha, et al.
Published: (2025)
Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs
by: Gao, Xin, et al.
Published: (2025)
by: Gao, Xin, et al.
Published: (2025)
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
by: Fan, Dongyang, et al.
Published: (2025)
by: Fan, Dongyang, et al.
Published: (2025)
Can LLMs Follow Simple Rules?
by: Mu, Norman, et al.
Published: (2023)
by: Mu, Norman, et al.
Published: (2023)
RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
by: de Wynter, Adrian, et al.
Published: (2024)
by: de Wynter, Adrian, et al.
Published: (2024)
Relational inductive biases on attention mechanisms
by: Mijangos, Víctor, et al.
Published: (2025)
by: Mijangos, Víctor, et al.
Published: (2025)
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200x Less Data?
by: Wang, Shaobo, et al.
Published: (2025)
by: Wang, Shaobo, et al.
Published: (2025)
Can LLMs Separate Instructions From Data? And What Do We Even Mean By That?
by: Zverev, Egor, et al.
Published: (2024)
by: Zverev, Egor, et al.
Published: (2024)
Decoding AI Authorship: Can LLMs Truly Mimic Human Style Across Literature and Politics?
by: Alsadhan, Nasser A
Published: (2026)
by: Alsadhan, Nasser A
Published: (2026)
Language steering in latent space to mitigate unintended code-switching
by: Goncharov, Andrey, et al.
Published: (2025)
by: Goncharov, Andrey, et al.
Published: (2025)
Climbing the Ladder of Reasoning: What LLMs Can-and Still Can't-Solve after SFT?
by: Sun, Yiyou, et al.
Published: (2025)
by: Sun, Yiyou, et al.
Published: (2025)
Can GRPO Help LLMs Transcend Their Pretraining Origin?
by: Ni, Kangqi, et al.
Published: (2025)
by: Ni, Kangqi, et al.
Published: (2025)
Can LLMs Convert Graphs to Text-Attributed Graphs?
by: Wang, Zehong, et al.
Published: (2024)
by: Wang, Zehong, et al.
Published: (2024)
Can Post-Training Transform LLMs into Causal Reasoners?
by: Chen, Junqi, et al.
Published: (2026)
by: Chen, Junqi, et al.
Published: (2026)
Can LLMs Help Uncover Insights about LLMs? A Large-Scale, Evolving Literature Analysis of Frontier LLMs
by: Park, Jungsoo, et al.
Published: (2025)
by: Park, Jungsoo, et al.
Published: (2025)
Towards Reducing Diagnostic Errors with Interpretable Risk Prediction
by: McInerney, Denis Jered, et al.
Published: (2024)
by: McInerney, Denis Jered, et al.
Published: (2024)
Similar Items
-
SAEs $\textit{Can}$ Improve Unlearning: Dynamic Sparse Autoencoder Guardrails for Precision Unlearning in LLMs
by: Muhamed, Aashiq, et al.
Published: (2025) -
Residual Stream Analysis with Multi-Layer SAEs
by: Lawson, Tim, et al.
Published: (2024) -
Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
by: Ahsan, Hiba, et al.
Published: (2025) -
Resa: Transparent Reasoning Models via SAEs
by: Wang, Shangshang, et al.
Published: (2025) -
SAEs Are Good for Steering -- If You Select the Right Features
by: Arad, Dana, et al.
Published: (2025)