Mechanistic Interpretability of Socio-Political Frames in Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Asghari, Hadi, Nenno, Sami |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
by: Wang, Xing, et al.
Published: (2025)
by: Wang, Xing, et al.
Published: (2025)
The Polite Liar: Epistemic Pathology in Language Models
by: DeVilling, Bentley
Published: (2025)
by: DeVilling, Bentley
Published: (2025)
Large Language Models in Politics and Democracy: A Comprehensive Survey
by: Aoki, Goshi
Published: (2024)
by: Aoki, Goshi
Published: (2024)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
by: Bai, Xiaoyan, et al.
Published: (2026)
by: Bai, Xiaoyan, et al.
Published: (2026)
STAR: SocioTechnical Approach to Red Teaming Language Models
by: Weidinger, Laura, et al.
Published: (2024)
by: Weidinger, Laura, et al.
Published: (2024)
DarkBench: Benchmarking Dark Patterns in Large Language Models
by: Kran, Esben, et al.
Published: (2025)
by: Kran, Esben, et al.
Published: (2025)
Detection and Mitigation of Hallucination in Large Reasoning Models: A Mechanistic Perspective
by: Sun, Zhongxiang, et al.
Published: (2025)
by: Sun, Zhongxiang, et al.
Published: (2025)
Political Alignment in Large Language Models: A Multidimensional Audit of Psychometric Identity and Behavioral Bias
by: Sakhawat, Adib, et al.
Published: (2026)
by: Sakhawat, Adib, et al.
Published: (2026)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
by: Greco, Candida M., et al.
Published: (2026)
by: Greco, Candida M., et al.
Published: (2026)
Open Source Language Models Can Provide Feedback: Evaluating LLMs' Ability to Help Students Using GPT-4-As-A-Judge
by: Koutcheme, Charles, et al.
Published: (2024)
by: Koutcheme, Charles, et al.
Published: (2024)
Interpreting Public Sentiment in Diplomacy Events: A Counterfactual Analysis Framework Using Large Language Models
by: Ouyang, Leyi
Published: (2025)
by: Ouyang, Leyi
Published: (2025)
Evaluating Large Language Models Against Human Annotators in Latent Content Analysis: Sentiment, Political Leaning, Emotional Intensity, and Sarcasm
by: Bojic, Ljubisa, et al.
Published: (2025)
by: Bojic, Ljubisa, et al.
Published: (2025)
Large Language Models Can Be a Viable Substitute for Expert Political Surveys When a Shock Disrupts Traditional Measurement Approaches
by: Wu, Patrick Y.
Published: (2025)
by: Wu, Patrick Y.
Published: (2025)
Many LLMs Are More Utilitarian Than One
by: Keshmirian, Anita, et al.
Published: (2025)
by: Keshmirian, Anita, et al.
Published: (2025)
The Political Preferences of LLMs
by: Rozado, David
Published: (2024)
by: Rozado, David
Published: (2024)
How English Print Media Frames Human-Elephant Conflicts in India
by: Punith, Bonala Sai, et al.
Published: (2026)
by: Punith, Bonala Sai, et al.
Published: (2026)
Are LLMs (Really) Ideological? An IRT-based Analysis and Alignment Tool for Perceived Socio-Economic Bias in LLMs
by: Wachter, Jasmin, et al.
Published: (2025)
by: Wachter, Jasmin, et al.
Published: (2025)
LLMs as Strategic Actors: Behavioral Alignment, Risk Calibration, and Argumentation Framing in Geopolitical Simulations
by: Solopova, Veronika, et al.
Published: (2026)
by: Solopova, Veronika, et al.
Published: (2026)
Mechanistic Interpretability of Emotion Inference in Large Language Models
by: Tak, Ala N., et al.
Published: (2025)
by: Tak, Ala N., et al.
Published: (2025)
Interpretability Framework for LLMs in Undergraduate Calculus
by: Dakshit, Sagnik, et al.
Published: (2025)
by: Dakshit, Sagnik, et al.
Published: (2025)
Statutory Construction and Interpretation for Artificial Intelligence
by: He, Luxi, et al.
Published: (2025)
by: He, Luxi, et al.
Published: (2025)
Measuring Political Preferences in AI Systems: An Integrative Approach
by: Rozado, David
Published: (2025)
by: Rozado, David
Published: (2025)
Perceived Political Bias in LLMs Reduces Persuasive Abilities
by: DiGiuseppe, Matthew, et al.
Published: (2026)
by: DiGiuseppe, Matthew, et al.
Published: (2026)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
by: Nghiem, Huy, et al.
Published: (2026)
by: Nghiem, Huy, et al.
Published: (2026)
What Is The Political Content in LLMs' Pre- and Post-Training Data?
by: Ceron, Tanise, et al.
Published: (2025)
by: Ceron, Tanise, et al.
Published: (2025)
Cross-Language Bias Examination in Large Language Models
by: Liang, Yuxuan, et al.
Published: (2025)
by: Liang, Yuxuan, et al.
Published: (2025)
Interpretable Recognition of Cognitive Distortions in Natural Language Texts
by: Kolonin, Anton, et al.
Published: (2025)
by: Kolonin, Anton, et al.
Published: (2025)
On the Creativity of Large Language Models
by: Franceschelli, Giorgio, et al.
Published: (2023)
by: Franceschelli, Giorgio, et al.
Published: (2023)
The Algorithmic Caricature: Auditing LLM-Generated Political Discourse Across Crisis Events
by: Gunjan, et al.
Published: (2026)
by: Gunjan, et al.
Published: (2026)
Linear Representations of Political Perspective Emerge in Large Language Models
by: Kim, Junsol, et al.
Published: (2025)
by: Kim, Junsol, et al.
Published: (2025)
Evaluating Large Language Models for Detecting Antisemitism
by: Patel, Jay, et al.
Published: (2025)
by: Patel, Jay, et al.
Published: (2025)
How Large Language Models are Designed to Hallucinate
by: Ackermann, Richard, et al.
Published: (2025)
by: Ackermann, Richard, et al.
Published: (2025)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
by: Jenny, David F., et al.
Published: (2023)
by: Jenny, David F., et al.
Published: (2023)
Anticipating Innovation Using Large Language Models
by: Fenoaltea, Enrico Maria, et al.
Published: (2026)
by: Fenoaltea, Enrico Maria, et al.
Published: (2026)
Open-Ended Wargames with Large Language Models
by: Hogan, Daniel P., et al.
Published: (2024)
by: Hogan, Daniel P., et al.
Published: (2024)
Large Language Models for Education: A Survey
by: Xu, Hanyi, et al.
Published: (2024)
by: Xu, Hanyi, et al.
Published: (2024)
Large Language Models as Misleading Assistants in Conversation
by: Hou, Betty Li, et al.
Published: (2024)
by: Hou, Betty Li, et al.
Published: (2024)
Evaluating Psychological Safety of Large Language Models
by: Li, Xingxuan, et al.
Published: (2022)
by: Li, Xingxuan, et al.
Published: (2022)
Large Language Models for Medicine: A Survey
by: Zheng, Yanxin, et al.
Published: (2024)
by: Zheng, Yanxin, et al.
Published: (2024)
Are Large Language Models Good Essay Graders?
by: Kundu, Anindita, et al.
Published: (2024)
by: Kundu, Anindita, et al.
Published: (2024)
Similar Items
-
Large Language Models' Complicit Responses to Illicit Instructions across Socio-Legal Contexts
by: Wang, Xing, et al.
Published: (2025) -
The Polite Liar: Epistemic Pathology in Language Models
by: DeVilling, Bentley
Published: (2025) -
Large Language Models in Politics and Democracy: A Comprehensive Survey
by: Aoki, Goshi
Published: (2024) -
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
by: Bai, Xiaoyan, et al.
Published: (2026) -
STAR: SocioTechnical Approach to Red Teaming Language Models
by: Weidinger, Laura, et al.
Published: (2024)