Evaluating Gender Bias of Pre-trained Language Models in Natural Language Inference by Considering All Labels
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anantaprayoon, Panatchakorn, Kaneko, Masahiro, Okazaki, Naoaki |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
von: Anantaprayoon, Panatchakorn, et al.
Veröffentlicht: (2025)
von: Anantaprayoon, Panatchakorn, et al.
Veröffentlicht: (2025)
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
Social Bias Evaluation for Large Language Models Requires Prompt Variations
von: Hida, Rem, et al.
Veröffentlicht: (2024)
von: Hida, Rem, et al.
Veröffentlicht: (2024)
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
von: Shiotani, Taihei, et al.
Veröffentlicht: (2026)
von: Shiotani, Taihei, et al.
Veröffentlicht: (2026)
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
von: Oi, Masanari, et al.
Veröffentlicht: (2024)
von: Oi, Masanari, et al.
Veröffentlicht: (2024)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
von: Ohi, Masanari, et al.
Veröffentlicht: (2024)
von: Ohi, Masanari, et al.
Veröffentlicht: (2024)
Sampling-based Pseudo-Likelihood for Membership Inference Attacks
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples
von: Koike, Ryuto, et al.
Veröffentlicht: (2023)
von: Koike, Ryuto, et al.
Veröffentlicht: (2023)
How You Prompt Matters! Even Task-Oriented Constraints in Instructions Affect LLM-Generated Text Detection
von: Koike, Ryuto, et al.
Veröffentlicht: (2023)
von: Koike, Ryuto, et al.
Veröffentlicht: (2023)
Solving NLP Problems through Human-System Collaboration: A Discussion-based Approach
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2023)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2023)
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks
von: Loem, Mengsay, et al.
Veröffentlicht: (2023)
von: Loem, Mengsay, et al.
Veröffentlicht: (2023)
In-Contextual Gender Bias Suppression for Large Language Models
von: Oba, Daisuke, et al.
Veröffentlicht: (2023)
von: Oba, Daisuke, et al.
Veröffentlicht: (2023)
The Gaps between Pre-train and Downstream Settings in Bias Evaluation and Debiasing
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
Dynamic Alignment for Collective Agency: Toward a Scalable Self-Improving Framework for Open-Ended LLM Alignment
von: Anantaprayoon, Panatchakorn, et al.
Veröffentlicht: (2025)
von: Anantaprayoon, Panatchakorn, et al.
Veröffentlicht: (2025)
Learning to Negotiate: Multi-Agent Deliberation for Collective Value Alignment in LLMs
von: Anantaprayoon, Panatchakorn, et al.
Veröffentlicht: (2026)
von: Anantaprayoon, Panatchakorn, et al.
Veröffentlicht: (2026)
Machine Text Detectors are Membership Inference Attacks
von: Koike, Ryuto, et al.
Veröffentlicht: (2025)
von: Koike, Ryuto, et al.
Veröffentlicht: (2025)
LLM Output Detectability and Task Performance Can be Jointly Optimized
von: Saito, Koshiro, et al.
Veröffentlicht: (2026)
von: Saito, Koshiro, et al.
Veröffentlicht: (2026)
Knowledge of Pretrained Language Models on Surface Information of Tokens
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
von: Hiraoka, Tatsuya, et al.
Veröffentlicht: (2024)
Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding
von: Oba, Daisuke, et al.
Veröffentlicht: (2026)
von: Oba, Daisuke, et al.
Veröffentlicht: (2026)
From Interpretability to Performance: Optimizing Retrieval Heads for Long-Context Language Models
von: Ma, Youmi, et al.
Veröffentlicht: (2026)
von: Ma, Youmi, et al.
Veröffentlicht: (2026)
ExaGPT: Example-Based Machine-Generated Text Detection for Human Interpretability
von: Koike, Ryuto, et al.
Veröffentlicht: (2025)
von: Koike, Ryuto, et al.
Veröffentlicht: (2025)
Projective Methods for Mitigating Gender Bias in Pre-trained Language Models
von: Dawkins, Hillary, et al.
Veröffentlicht: (2024)
von: Dawkins, Hillary, et al.
Veröffentlicht: (2024)
Drifting Objectives for Refining Discrete Diffusion Language Models
von: Oba, Daisuke, et al.
Veröffentlicht: (2026)
von: Oba, Daisuke, et al.
Veröffentlicht: (2026)
Evaluating Gender Bias Transfer between Pre-trained and Prompt-Adapted Language Models
von: Mackraz, Natalie, et al.
Veröffentlicht: (2024)
von: Mackraz, Natalie, et al.
Veröffentlicht: (2024)
Diffusion-State Policy Optimization for Masked Diffusion Language Models
von: Oba, Daisuke, et al.
Veröffentlicht: (2026)
von: Oba, Daisuke, et al.
Veröffentlicht: (2026)
JUBAKU: An Adversarial Benchmark for Exposing Culturally Grounded Stereotypes in Japanese LLMs
von: Shiotani, Taihei, et al.
Veröffentlicht: (2026)
von: Shiotani, Taihei, et al.
Veröffentlicht: (2026)
Evaluating Gender Bias in Large Language Models
von: Döll, Michael, et al.
Veröffentlicht: (2024)
von: Döll, Michael, et al.
Veröffentlicht: (2024)
Tokenization as Finite-State Transduction
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
Towards Label-Only Membership Inference Attack against Pre-trained Large Language Models
von: He, Yu, et al.
Veröffentlicht: (2025)
von: He, Yu, et al.
Veröffentlicht: (2025)
A Little Leak Will Sink a Great Ship: Survey of Transparency for Large Language Models from Start to Finish
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024)
On the Alignment of Large Language Models with Global Human Opinion
von: Liu, Yang, et al.
Veröffentlicht: (2025)
von: Liu, Yang, et al.
Veröffentlicht: (2025)
Building a Japanese Document-Level Relation Extraction Dataset Assisted by Cross-Lingual Transfer
von: Ma, Youmi, et al.
Veröffentlicht: (2024)
von: Ma, Youmi, et al.
Veröffentlicht: (2024)
Evaluating Discourse Cohesion in Pre-trained Language Models
von: He, Jie, et al.
Veröffentlicht: (2025)
von: He, Jie, et al.
Veröffentlicht: (2025)
Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model
von: Sasagawa, Keito, et al.
Veröffentlicht: (2024)
von: Sasagawa, Keito, et al.
Veröffentlicht: (2024)
Cross-Care: Assessing the Healthcare Implications of Pre-training Data on Language Model Bias
von: Chen, Shan, et al.
Veröffentlicht: (2024)
von: Chen, Shan, et al.
Veröffentlicht: (2024)
Balanced Multi-Factor In-Context Learning for Multilingual Large Language Models
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2025)
What is Your Favorite Gender, MLM? Gender Bias Evaluation in Multilingual Masked Language Models
von: Yu, Jeongrok, et al.
Veröffentlicht: (2024)
von: Yu, Jeongrok, et al.
Veröffentlicht: (2024)
Chinese Sequence Labeling with Semi-Supervised Boundary-Aware Language Model Pre-training
von: Zhang, Longhui, et al.
Veröffentlicht: (2024)
von: Zhang, Longhui, et al.
Veröffentlicht: (2024)
Bit-level BPE: Below the byte boundary
von: Moon, Sangwhan, et al.
Veröffentlicht: (2025)
von: Moon, Sangwhan, et al.
Veröffentlicht: (2025)
Distributional Properties of Subword Regularization
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
von: Cognetta, Marco, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Intent-Aware Self-Correction for Mitigating Social Biases in Large Language Models
von: Anantaprayoon, Panatchakorn, et al.
Veröffentlicht: (2025) -
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting
von: Kaneko, Masahiro, et al.
Veröffentlicht: (2024) -
Social Bias Evaluation for Large Language Models Requires Prompt Variations
von: Hida, Rem, et al.
Veröffentlicht: (2024) -
A Japanese Benchmark for Evaluating Social Bias in Reasoning Based on Attribution Theory
von: Shiotani, Taihei, et al.
Veröffentlicht: (2026) -
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
von: Oi, Masanari, et al.
Veröffentlicht: (2024)