Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Ortu, Francesco, Yook, Joeun, Pandey, Punya Syon, Samway, Keenan, Schölkopf, Bernhard, Cazzaniga, Alberto, Mihalcea, Rada, Jin, Zhijing |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are LLMs Good Safety Agents or a Propaganda Engine?
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Are Language Models Consequentialist or Deontological Moral Reasoners?
by: Samway, Keenan, et al.
Published: (2025)
by: Samway, Keenan, et al.
Published: (2025)
When Do Language Models Endorse Limitations on Human Rights Principles?
by: Samway, Keenan, et al.
Published: (2026)
by: Samway, Keenan, et al.
Published: (2026)
BinaryPPO: Efficient Policy Optimization for Binary Classification
by: Pandey, Punya Syon, et al.
Published: (2026)
by: Pandey, Punya Syon, et al.
Published: (2026)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
by: Harrasse, Abir, et al.
Published: (2025)
by: Harrasse, Abir, et al.
Published: (2025)
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
by: Ortu, Francesco, et al.
Published: (2024)
by: Ortu, Francesco, et al.
Published: (2024)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
by: Ceraolo, Roberto, et al.
Published: (2024)
by: Ceraolo, Roberto, et al.
Published: (2024)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
by: Backmann, Steffen, et al.
Published: (2025)
by: Backmann, Steffen, et al.
Published: (2025)
Test of Time: Rethinking Temporal Signal of Benchmark Contamination
by: Zhang, Terry Jingchen, et al.
Published: (2025)
by: Zhang, Terry Jingchen, et al.
Published: (2025)
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
by: Ortu, Francesco, et al.
Published: (2025)
by: Ortu, Francesco, et al.
Published: (2025)
Implicit Personalization in Language Models: A Systematic Study
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
by: Piedrahita, David Guzman, et al.
Published: (2025)
by: Piedrahita, David Guzman, et al.
Published: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
by: Pandey, Punya Syon, et al.
Published: (2025)
by: Pandey, Punya Syon, et al.
Published: (2025)
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
by: Yadav, Neemesh, et al.
Published: (2025)
by: Yadav, Neemesh, et al.
Published: (2025)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
by: Borah, Angana, et al.
Published: (2024)
by: Borah, Angana, et al.
Published: (2024)
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis
by: Lyu, Zhiheng, et al.
Published: (2024)
by: Lyu, Zhiheng, et al.
Published: (2024)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
by: Piatti, Giorgio, et al.
Published: (2024)
by: Piatti, Giorgio, et al.
Published: (2024)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
by: Jenny, David F., et al.
Published: (2023)
by: Jenny, David F., et al.
Published: (2023)
Language Model Alignment in Multilingual Trolley Problems
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
The Curious Case of Curiosity across Human Cultures and LLMs
by: Borah, Angana, et al.
Published: (2025)
by: Borah, Angana, et al.
Published: (2025)
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
by: Draye, Florent, et al.
Published: (2026)
by: Draye, Florent, et al.
Published: (2026)
Why AI Is WEIRD and Should Not Be This Way: Towards AI For Everyone, With Everyone, By Everyone
by: Mihalcea, Rada, et al.
Published: (2024)
by: Mihalcea, Rada, et al.
Published: (2024)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
by: Nwatu, Joan, et al.
Published: (2024)
by: Nwatu, Joan, et al.
Published: (2024)
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025)
by: Simko, Samuel, et al.
Published: (2025)
Causality for Natural Language Processing
by: Jin, Zhijing
Published: (2025)
by: Jin, Zhijing
Published: (2025)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
by: Arif, Samee, et al.
Published: (2026)
by: Arif, Samee, et al.
Published: (2026)
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data
by: Mori, Shinka, et al.
Published: (2024)
by: Mori, Shinka, et al.
Published: (2024)
CausalCite: A Causal Formulation of Paper Citations
by: Kumar, Ishan, et al.
Published: (2023)
by: Kumar, Ishan, et al.
Published: (2023)
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
by: Ignat, Oana, et al.
Published: (2023)
by: Ignat, Oana, et al.
Published: (2023)
Causality can systematically address the monsters under the bench(marks)
by: Leeb, Felix, et al.
Published: (2025)
by: Leeb, Felix, et al.
Published: (2025)
Causal Responsibility Attribution for Human-AI Collaboration
by: Qi, Yahang, et al.
Published: (2024)
by: Qi, Yahang, et al.
Published: (2024)
Which Humans? Inclusivity and Representation in Human-Centered AI
by: Mihalcea, Rada, et al.
Published: (2025)
by: Mihalcea, Rada, et al.
Published: (2025)
Future of Pandemic Prevention and Response CCC Workshop Report
by: Danks, David, et al.
Published: (2024)
by: Danks, David, et al.
Published: (2024)
Assessing Historical Structural Oppression Worldwide via Rule-Guided Prompting of Large Language Models
by: Chatterjee, Sreejato, et al.
Published: (2025)
by: Chatterjee, Sreejato, et al.
Published: (2025)
FIGURA: A Modular Prompt Engineering Method for Artistic Figure Photography in Safety-Filtered Text-to-Image Models
by: Cazzaniga, Luca
Published: (2026)
by: Cazzaniga, Luca
Published: (2026)
Blocking Mechanism of Porn Website in India: Claim and Truth
by: Pandey, Saurabh, et al.
Published: (2019)
by: Pandey, Saurabh, et al.
Published: (2019)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
by: Nwatu, Joan, et al.
Published: (2025)
by: Nwatu, Joan, et al.
Published: (2025)
Similar Items
-
Are LLMs Good Safety Agents or a Propaganda Engine?
by: Yadav, Neemesh, et al.
Published: (2025) -
Are Language Models Consequentialist or Deontological Moral Reasoners?
by: Samway, Keenan, et al.
Published: (2025) -
When Do Language Models Endorse Limitations on Human Rights Principles?
by: Samway, Keenan, et al.
Published: (2026) -
BinaryPPO: Efficient Policy Optimization for Binary Classification
by: Pandey, Punya Syon, et al.
Published: (2026) -
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
by: Pandey, Punya Syon, et al.
Published: (2025)