It's the humans, not the data: Geopolitical bias in LLMs originates in post-training, amplified by the language of the prompt
Fuente:
arXiv
Saved in:
| Main Authors: | Bladon, Stuart, Bent, Brinnae |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feature Visualization Recovers Known Cortical Selectivity from TRIBE v2
by: Bladon, Stuart, et al.
Published: (2026)
by: Bladon, Stuart, et al.
Published: (2026)
Semantic Approach to Quantifying the Consistency of Diffusion Model Image Generation
by: Bent, Brinnae
Published: (2024)
by: Bent, Brinnae
Published: (2024)
The Term 'Agent' Has Been Diluted Beyond Utility and Requires Redefinition
by: Bent, Brinnae
Published: (2025)
by: Bent, Brinnae
Published: (2025)
Efficient multi-prompt evaluation of LLMs
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
by: Gan, Xingwei, et al.
Published: (2026)
by: Gan, Xingwei, et al.
Published: (2026)
RL in Name Only? Analyzing the Structural Assumptions in RL post-training for LLMs
by: Samineni, Soumya Rani, et al.
Published: (2025)
by: Samineni, Soumya Rani, et al.
Published: (2025)
Post-training makes large language models less human-like
by: Binz, Marcel, et al.
Published: (2026)
by: Binz, Marcel, et al.
Published: (2026)
Bayes-PD: Exploring a Sequence to Binding Bayesian Neural Network model trained on Phage Display data
by: Amiaud-Plachy, Ilann, et al.
Published: (2026)
by: Amiaud-Plachy, Ilann, et al.
Published: (2026)
torchtune: PyTorch native post-training library
by: Obozov, Mark, et al.
Published: (2026)
by: Obozov, Mark, et al.
Published: (2026)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
PII-Compass: Guiding LLM training data extraction prompts towards the target PII via grounding
by: Nakka, Krishna Kanth, et al.
Published: (2024)
by: Nakka, Krishna Kanth, et al.
Published: (2024)
Causal prompting model-based offline reinforcement learning
by: Yu, Xuehui, et al.
Published: (2024)
by: Yu, Xuehui, et al.
Published: (2024)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
FoGE: Fock Space inspired encoding for graph prompting
by: Chytas, Sotirios Panagiotis, et al.
Published: (2025)
by: Chytas, Sotirios Panagiotis, et al.
Published: (2025)
From Simulation to Enaction: Post-trained language models recognize and react to their own generations
by: G., Asvin, et al.
Published: (2026)
by: G., Asvin, et al.
Published: (2026)
Benchmarking zero-shot stance detection with FlanT5-XXL: Insights from training data, prompting, and decoding strategies into its near-SoTA performance
by: Aiyappa, Rachith, et al.
Published: (2024)
by: Aiyappa, Rachith, et al.
Published: (2024)
When2Speak: A Dataset for Temporal Participation and Turn-Taking in Multi-Party Conversations for Large Language Models
by: Nama, Vihaan, et al.
Published: (2026)
by: Nama, Vihaan, et al.
Published: (2026)
Aviary: training language agents on challenging scientific tasks
by: Narayanan, Siddharth, et al.
Published: (2024)
by: Narayanan, Siddharth, et al.
Published: (2024)
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
by: Motwani, Sumeet Ramesh, et al.
Published: (2026)
Prompt replay: speeding up grpo with on-policy reuse of high-signal prompts
by: Baroian, Andrei, et al.
Published: (2026)
by: Baroian, Andrei, et al.
Published: (2026)
Where does output diversity collapse in post-training?
by: Karouzos, Constantinos, et al.
Published: (2026)
by: Karouzos, Constantinos, et al.
Published: (2026)
Discriminative protein sequence modelling with Latent Space Diffusion
by: Quinn, Eoin, et al.
Published: (2025)
by: Quinn, Eoin, et al.
Published: (2025)
Introducing HALC: A general pipeline for finding optimal prompting strategies for automated coding with LLMs in the computational social sciences
by: Reich, Andreas, et al.
Published: (2025)
by: Reich, Andreas, et al.
Published: (2025)
Hierarchical Context Merging: Better Long Context Understanding for Pre-trained LLMs
by: Song, Woomin, et al.
Published: (2024)
by: Song, Woomin, et al.
Published: (2024)
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
by: Wijk, Hjalmar, et al.
Published: (2024)
by: Wijk, Hjalmar, et al.
Published: (2024)
Aleph-Alpha-GermanWeb: Improving German-language LLM pre-training with model-based data curation and synthetic data generation
by: Burns, Thomas F, et al.
Published: (2025)
by: Burns, Thomas F, et al.
Published: (2025)
Prior-informed optimization of treatment recommendation via bandit algorithms trained on large language model-processed historical records
by: Nessari, Saman, et al.
Published: (2025)
by: Nessari, Saman, et al.
Published: (2025)
Towards a more efficient bias detection in financial language models
by: Kacem, Firas Hadj, et al.
Published: (2026)
by: Kacem, Firas Hadj, et al.
Published: (2026)
P2DT: Mitigating Forgetting in task-incremental Learning with progressive prompt Decision Transformer
by: Wang, Zhiyuan, et al.
Published: (2024)
by: Wang, Zhiyuan, et al.
Published: (2024)
Zero-shot data citation function classification using transformer-based large language models (LLMs)
by: Byers, Neil, et al.
Published: (2025)
by: Byers, Neil, et al.
Published: (2025)
Brittlebench: Quantifying LLM robustness via prompt sensitivity
by: Romanou, Angelika, et al.
Published: (2026)
by: Romanou, Angelika, et al.
Published: (2026)
Can large language models replace humans in the systematic review process? Evaluating GPT-4's efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages
by: Khraisha, Qusai, et al.
Published: (2023)
by: Khraisha, Qusai, et al.
Published: (2023)
Data filtering methods for training language models
by: Shevchenko, Egor, et al.
Published: (2026)
by: Shevchenko, Egor, et al.
Published: (2026)
Improving training time and GPU utilization in geo-distributed language model training
by: Palak, et al.
Published: (2024)
by: Palak, et al.
Published: (2024)
Hybrid Forecasting of Geopolitical Events
by: Benjamin, Daniel M., et al.
Published: (2024)
by: Benjamin, Daniel M., et al.
Published: (2024)
Detecting labeling bias using influence functions
by: Jørgensen, Frida, et al.
Published: (2026)
by: Jørgensen, Frida, et al.
Published: (2026)
E-Globe: Scalable $ε$-Global Verification of Neural Networks via Tight Upper Bounds and Pattern-Aware Branching
by: Li, Wenting, et al.
Published: (2026)
by: Li, Wenting, et al.
Published: (2026)
Understanding the dynamics of the frequency bias in neural networks
by: Molina, Juan, et al.
Published: (2024)
by: Molina, Juan, et al.
Published: (2024)
Self-Improving Pretraining: using post-trained models to pretrain better models
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
by: Tan, Ellen Xiaoqing, et al.
Published: (2026)
On the origin of neural scaling laws: from random graphs to natural language
by: Barkeshli, Maissam, et al.
Published: (2026)
by: Barkeshli, Maissam, et al.
Published: (2026)
Similar Items
-
Feature Visualization Recovers Known Cortical Selectivity from TRIBE v2
by: Bladon, Stuart, et al.
Published: (2026) -
Semantic Approach to Quantifying the Consistency of Diffusion Model Image Generation
by: Bent, Brinnae
Published: (2024) -
The Term 'Agent' Has Been Diluted Beyond Utility and Requires Redefinition
by: Bent, Brinnae
Published: (2025) -
Efficient multi-prompt evaluation of LLMs
by: Polo, Felipe Maia, et al.
Published: (2024) -
Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
by: Gan, Xingwei, et al.
Published: (2026)