Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Ortu, Francesco, Yook, Joeun, Pandey, Punya Syon, Samway, Keenan, Schölkopf, Bernhard, Cazzaniga, Alberto, Mihalcea, Rada, Jin, Zhijing |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Are LLMs Good Safety Agents or a Propaganda Engine?
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
Are Language Models Consequentialist or Deontological Moral Reasoners?
di: Samway, Keenan, et al.
Pubblicazione: (2025)
di: Samway, Keenan, et al.
Pubblicazione: (2025)
When Do Language Models Endorse Limitations on Human Rights Principles?
di: Samway, Keenan, et al.
Pubblicazione: (2026)
di: Samway, Keenan, et al.
Pubblicazione: (2026)
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026)
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
di: Harrasse, Abir, et al.
Pubblicazione: (2025)
Competition of Mechanisms: Tracing How Language Models Handle Facts and Counterfactuals
di: Ortu, Francesco, et al.
Pubblicazione: (2024)
di: Ortu, Francesco, et al.
Pubblicazione: (2024)
Quriosity: Analyzing Human Questioning Behavior and Causal Inquiry through Curiosity-Driven Queries
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
di: Ceraolo, Roberto, et al.
Pubblicazione: (2024)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
Test of Time: Rethinking Temporal Signal of Benchmark Contamination
di: Zhang, Terry Jingchen, et al.
Pubblicazione: (2025)
di: Zhang, Terry Jingchen, et al.
Pubblicazione: (2025)
When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
di: Ortu, Francesco, et al.
Pubblicazione: (2025)
di: Ortu, Francesco, et al.
Pubblicazione: (2025)
Implicit Personalization in Language Models: A Systematic Study
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
di: Piedrahita, David Guzman, et al.
Pubblicazione: (2025)
Accidental Vulnerability: Factors in Fine-Tuning that Shift Model Safeguards
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
CORE: Measuring Multi-Agent LLM Interaction Quality under Game-Theoretic Pressures
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
di: Yadav, Neemesh, et al.
Pubblicazione: (2025)
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions
di: Borah, Angana, et al.
Pubblicazione: (2024)
di: Borah, Angana, et al.
Pubblicazione: (2024)
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis
di: Lyu, Zhiheng, et al.
Pubblicazione: (2024)
di: Lyu, Zhiheng, et al.
Pubblicazione: (2024)
Can Large Language Models Infer Causation from Correlation?
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
di: Jin, Zhijing, et al.
Pubblicazione: (2023)
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
di: Piatti, Giorgio, et al.
Pubblicazione: (2024)
di: Piatti, Giorgio, et al.
Pubblicazione: (2024)
Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
di: Jenny, David F., et al.
Pubblicazione: (2023)
di: Jenny, David F., et al.
Pubblicazione: (2023)
Language Model Alignment in Multilingual Trolley Problems
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
di: Jin, Zhijing, et al.
Pubblicazione: (2024)
The Curious Case of Curiosity across Human Cultures and LLMs
di: Borah, Angana, et al.
Pubblicazione: (2025)
di: Borah, Angana, et al.
Pubblicazione: (2025)
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
di: Draye, Florent, et al.
Pubblicazione: (2026)
di: Draye, Florent, et al.
Pubblicazione: (2026)
Why AI Is WEIRD and Should Not Be This Way: Towards AI For Everyone, With Everyone, By Everyone
di: Mihalcea, Rada, et al.
Pubblicazione: (2024)
di: Mihalcea, Rada, et al.
Pubblicazione: (2024)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
Improving Large Language Model Safety with Contrastive Representation Learning
di: Simko, Samuel, et al.
Pubblicazione: (2025)
di: Simko, Samuel, et al.
Pubblicazione: (2025)
Causality for Natural Language Processing
di: Jin, Zhijing
Pubblicazione: (2025)
di: Jin, Zhijing
Pubblicazione: (2025)
One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety
di: Arif, Samee, et al.
Pubblicazione: (2026)
di: Arif, Samee, et al.
Pubblicazione: (2026)
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data
di: Mori, Shinka, et al.
Pubblicazione: (2024)
di: Mori, Shinka, et al.
Pubblicazione: (2024)
CausalCite: A Causal Formulation of Paper Citations
di: Kumar, Ishan, et al.
Pubblicazione: (2023)
di: Kumar, Ishan, et al.
Pubblicazione: (2023)
Human Action Co-occurrence in Lifestyle Vlogs using Graph Link Prediction
di: Ignat, Oana, et al.
Pubblicazione: (2023)
di: Ignat, Oana, et al.
Pubblicazione: (2023)
Causality can systematically address the monsters under the bench(marks)
di: Leeb, Felix, et al.
Pubblicazione: (2025)
di: Leeb, Felix, et al.
Pubblicazione: (2025)
Causal Responsibility Attribution for Human-AI Collaboration
di: Qi, Yahang, et al.
Pubblicazione: (2024)
di: Qi, Yahang, et al.
Pubblicazione: (2024)
Which Humans? Inclusivity and Representation in Human-Centered AI
di: Mihalcea, Rada, et al.
Pubblicazione: (2025)
di: Mihalcea, Rada, et al.
Pubblicazione: (2025)
Future of Pandemic Prevention and Response CCC Workshop Report
di: Danks, David, et al.
Pubblicazione: (2024)
di: Danks, David, et al.
Pubblicazione: (2024)
Assessing Historical Structural Oppression Worldwide via Rule-Guided Prompting of Large Language Models
di: Chatterjee, Sreejato, et al.
Pubblicazione: (2025)
di: Chatterjee, Sreejato, et al.
Pubblicazione: (2025)
FIGURA: A Modular Prompt Engineering Method for Artistic Figure Photography in Safety-Filtered Text-to-Image Models
di: Cazzaniga, Luca
Pubblicazione: (2026)
di: Cazzaniga, Luca
Pubblicazione: (2026)
Blocking Mechanism of Porn Website in India: Claim and Truth
di: Pandey, Saurabh, et al.
Pubblicazione: (2019)
di: Pandey, Saurabh, et al.
Pubblicazione: (2019)
Culture Affordance Atlas: Reconciling Object Diversity Through Functional Mapping
di: Nwatu, Joan, et al.
Pubblicazione: (2025)
di: Nwatu, Joan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Are LLMs Good Safety Agents or a Propaganda Engine?
di: Yadav, Neemesh, et al.
Pubblicazione: (2025) -
Are Language Models Consequentialist or Deontological Moral Reasoners?
di: Samway, Keenan, et al.
Pubblicazione: (2025) -
When Do Language Models Endorse Limitations on Human Rights Principles?
di: Samway, Keenan, et al.
Pubblicazione: (2026) -
BinaryPPO: Efficient Policy Optimization for Binary Classification
di: Pandey, Punya Syon, et al.
Pubblicazione: (2026) -
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
di: Pandey, Punya Syon, et al.
Pubblicazione: (2025)