How far can bias go? Tracing bias from pretraining data to alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Thaler, Marion, Köksal, Abdullatif, Leidinger, Alina, Korhonen, Anna, Schütze, Hinrich |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2024)
by: Köksal, Abdullatif, et al.
Published: (2024)
LongForm: Effective Instruction Tuning with Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2023)
by: Köksal, Abdullatif, et al.
Published: (2023)
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish
by: Yüksel, Arda, et al.
Published: (2024)
by: Yüksel, Arda, et al.
Published: (2024)
Consistent Document-Level Relation Extraction via Counterfactuals
by: Modarressi, Ali, et al.
Published: (2024)
by: Modarressi, Ali, et al.
Published: (2024)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
by: Weissweiler, Leonie, et al.
Published: (2024)
by: Weissweiler, Leonie, et al.
Published: (2024)
SYNTHEVAL: Hybrid Behavioral Testing of NLP Models with Synthetic CheckLists
by: Zhao, Raoyuan, et al.
Published: (2024)
by: Zhao, Raoyuan, et al.
Published: (2024)
CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation
by: Ziegler, Ingo, et al.
Published: (2024)
by: Ziegler, Ingo, et al.
Published: (2024)
MemLLM: Finetuning LLMs to Use An Explicit Read-Write Memory
by: Modarressi, Ali, et al.
Published: (2024)
by: Modarressi, Ali, et al.
Published: (2024)
Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing
by: Zhao, Raoyuan, et al.
Published: (2025)
by: Zhao, Raoyuan, et al.
Published: (2025)
How Are LLMs Mitigating Stereotyping Harms? Learning from Search Engine Studies
by: Leidinger, Alina, et al.
Published: (2024)
by: Leidinger, Alina, et al.
Published: (2024)
GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
by: Veloso, Leonor, et al.
Published: (2026)
by: Veloso, Leonor, et al.
Published: (2026)
Understanding Gated Neurons in Transformers from Their Input-Output Functionality
by: Gerstner, Sebastian, et al.
Published: (2025)
by: Gerstner, Sebastian, et al.
Published: (2025)
Are LLMs classical or nonmonotonic reasoners? Lessons from generics
by: Leidinger, Alina, et al.
Published: (2024)
by: Leidinger, Alina, et al.
Published: (2024)
GLUScope: A Tool for Analyzing GLU Neurons in Transformer Language Models
by: Gerstner, Sebastian, et al.
Published: (2026)
by: Gerstner, Sebastian, et al.
Published: (2026)
How Programming Concepts and Neurons Are Shared in Code Language Models
by: Kargaran, Amir Hossein, et al.
Published: (2025)
by: Kargaran, Amir Hossein, et al.
Published: (2025)
Breaking the Script Barrier in Multilingual Pre-Trained Language Models with Transliteration-Based Post-Training Alignment
by: Xhelili, Orgest, et al.
Published: (2024)
by: Xhelili, Orgest, et al.
Published: (2024)
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models
by: Nie, Ercong, et al.
Published: (2025)
by: Nie, Ercong, et al.
Published: (2025)
Evaluating Contextually Mediated Factual Recall in Multilingual Large Language Models
by: Liu, Yihong, et al.
Published: (2026)
by: Liu, Yihong, et al.
Published: (2026)
Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders
by: Veitsman, Yana, et al.
Published: (2026)
by: Veitsman, Yana, et al.
Published: (2026)
The Anatomy of an Edit: Mechanism-Guided Activation Steering for Knowledge Editing
by: Cao, Yuan, et al.
Published: (2026)
by: Cao, Yuan, et al.
Published: (2026)
MaskLID: Code-Switching Language Identification through Iterative Masking
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
XAMPLER: Learning to Retrieve Cross-Lingual In-Context Examples
by: Lin, Peiqin, et al.
Published: (2024)
by: Lin, Peiqin, et al.
Published: (2024)
A Recipe of Parallel Corpora Exploitation for Multilingual Large Language Models
by: Lin, Peiqin, et al.
Published: (2024)
by: Lin, Peiqin, et al.
Published: (2024)
GlotScript: A Resource and Tool for Low Resource Writing System Identification
by: Kargaran, Amir Hossein, et al.
Published: (2023)
by: Kargaran, Amir Hossein, et al.
Published: (2023)
Inducing anxiety in large language models can induce bias
by: Coda-Forno, Julian, et al.
Published: (2023)
by: Coda-Forno, Julian, et al.
Published: (2023)
Relational Linearity is a Predictor of Hallucinations
by: Lu, Yuetian, et al.
Published: (2026)
by: Lu, Yuetian, et al.
Published: (2026)
TransMI: A Framework to Create Strong Baselines from Multilingual Pretrained Language Models for Transliterated Data
by: Liu, Yihong, et al.
Published: (2024)
by: Liu, Yihong, et al.
Published: (2024)
Tracing Multilingual Factual Knowledge Acquisition in Pretraining
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
To Bias or Not to Bias: Detecting bias in News with bias-detector
by: Ghosh, Himel, et al.
Published: (2025)
by: Ghosh, Himel, et al.
Published: (2025)
Joint Lemmatization and Morphological Tagging with LEMMING
by: Muller, Thomas, et al.
Published: (2024)
by: Muller, Thomas, et al.
Published: (2024)
Labeled Morphological Segmentation with Semi-Markov Models
by: Cotterell, Ryan, et al.
Published: (2024)
by: Cotterell, Ryan, et al.
Published: (2024)
TransliCo: A Contrastive Learning Framework to Address the Script Barrier in Multilingual Pretrained Language Models
by: Liu, Yihong, et al.
Published: (2024)
by: Liu, Yihong, et al.
Published: (2024)
What Do Dialect Speakers Want? A Survey of Attitudes Towards Language Technology for German Dialects
by: Blaschke, Verena, et al.
Published: (2024)
by: Blaschke, Verena, et al.
Published: (2024)
RET-LLM: Towards a General Read-Write Memory for Large Language Models
by: Modarressi, Ali, et al.
Published: (2023)
by: Modarressi, Ali, et al.
Published: (2023)
Left, Right, or Center? Evaluating LLM Framing in News Classification and Generation
by: Kennedy, Molly, et al.
Published: (2026)
by: Kennedy, Molly, et al.
Published: (2026)
OFA: A Framework of Initializing Unseen Subword Embeddings for Efficient Large-scale Multilingual Continued Pretraining
by: Liu, Yihong, et al.
Published: (2023)
by: Liu, Yihong, et al.
Published: (2023)
SLAyiNG: Towards Queer Language Processing
by: Veloso, Leonor, et al.
Published: (2025)
by: Veloso, Leonor, et al.
Published: (2025)
On the Entity-Level Alignment in Crosslingual Consistency
by: Liu, Yihong, et al.
Published: (2025)
by: Liu, Yihong, et al.
Published: (2025)
GlotCC: An Open Broad-Coverage CommonCrawl Corpus and Pipeline for Minority Languages
by: Kargaran, Amir Hossein, et al.
Published: (2024)
by: Kargaran, Amir Hossein, et al.
Published: (2024)
Actuation without production bias
by: Kirby, James, et al.
Published: (2024)
by: Kirby, James, et al.
Published: (2024)
Similar Items
-
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2024) -
LongForm: Effective Instruction Tuning with Reverse Instructions
by: Köksal, Abdullatif, et al.
Published: (2023) -
TurkishMMLU: Measuring Massive Multitask Language Understanding in Turkish
by: Yüksel, Arda, et al.
Published: (2024) -
Consistent Document-Level Relation Extraction via Counterfactuals
by: Modarressi, Ali, et al.
Published: (2024) -
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
by: Weissweiler, Leonie, et al.
Published: (2024)