PBBQ: A Persian Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Farsi, Farhan, Bali, Shayan, Valeh, Fatemeh, Ghofrani, Parsa, Pakniat, Alireza, Kashfipour, Kian, Payberah, Amir H. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language
by: Farsi, Farhan, et al.
Published: (2025)
by: Farsi, Farhan, et al.
Published: (2025)
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
by: Rad, Mohammad Heydari, et al.
Published: (2024)
by: Rad, Mohammad Heydari, et al.
Published: (2024)
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025)
by: Hosseini, Mohammad, et al.
Published: (2025)
Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM
by: Puhach, Dariia, et al.
Published: (2025)
by: Puhach, Dariia, et al.
Published: (2025)
An Empirical Study of Accuracy-Robustness Tradeoff and Training Efficiency in Self-Supervised Learning
by: Ghofrani, Fatemeh, et al.
Published: (2025)
by: Ghofrani, Fatemeh, et al.
Published: (2025)
LRW-Persian: Lip-reading in the Wild Dataset for Persian Language
by: Taghizadeh, Zahra, et al.
Published: (2025)
by: Taghizadeh, Zahra, et al.
Published: (2025)
Human-AI Collaborative Uncertainty Quantification
by: Noorani, Sima, et al.
Published: (2025)
by: Noorani, Sima, et al.
Published: (2025)
ParsiPy: NLP Toolkit for Historical Persian Texts in Python
by: Farsi, Farhan, et al.
Published: (2025)
by: Farsi, Farhan, et al.
Published: (2025)
Single molecule localization microscopy challenge: a biologically inspired benchmark for long-sequence modeling
by: Valeh, Fatemeh, et al.
Published: (2026)
by: Valeh, Fatemeh, et al.
Published: (2026)
Features of manufacturing case of hydraulic cylinders of structural powder steel
by: Aynur Valeh Sharifova
Published: (2024)
by: Aynur Valeh Sharifova
Published: (2024)
LVLM-COUNT: Enhancing the Counting Ability of Large Vision-Language Models
by: Qharabagh, Muhammad Fetrat, et al.
Published: (2024)
by: Qharabagh, Muhammad Fetrat, et al.
Published: (2024)
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
by: Mirbagheri, Mohammad Reza, et al.
Published: (2025)
by: Mirbagheri, Mohammad Reza, et al.
Published: (2025)
PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian
by: Mozafari, Jamshid, et al.
Published: (2026)
by: Mozafari, Jamshid, et al.
Published: (2026)
Multi-Round Human-AI Collaboration with User-Specified Requirements
by: Noorani, Sima, et al.
Published: (2026)
by: Noorani, Sima, et al.
Published: (2026)
CombiGraph-Vis: A Curated Multimodal Olympiad Benchmark for Discrete Mathematical Reasoning
by: Mahdavi, Hamed, et al.
Published: (2025)
by: Mahdavi, Hamed, et al.
Published: (2025)
Accelerate Model Parallel Training by Using Efficient Graph Traversal Order in Device Placement
by: Wang, Tianze, et al.
Published: (2022)
by: Wang, Tianze, et al.
Published: (2022)
ELAB: Extensive LLM Alignment Benchmark in Persian Language
by: Pourbahman, Zahra, et al.
Published: (2025)
by: Pourbahman, Zahra, et al.
Published: (2025)
A Comparative Evaluation of Large Language Models for Persian Sentiment Analysis and Emotion Detection in Social Media Texts
by: Tohidi, Kian, et al.
Published: (2025)
by: Tohidi, Kian, et al.
Published: (2025)
Combining Trained Models in Reinforcement Learning
by: Patil, Ujjwal, et al.
Published: (2026)
by: Patil, Ujjwal, et al.
Published: (2026)
Desde la Literatura Vanguardista hasta el Diseño cultural: Entrevista a Amir Parsa. Un dialogo sobre el Alzheimer ’s Project y el Programa Meet me at MoMA
by: Amir Parsa
Published: (2014)
by: Amir Parsa
Published: (2014)
Designing Human-AI Systems: Anthropomorphism and Framing Bias on Human-AI Collaboration
by: Olszewski, Samuel Aleksander Sánchez
Published: (2024)
by: Olszewski, Samuel Aleksander Sánchez
Published: (2024)
When Data is the Algorithm: A Systematic Study and Curation of Preference Optimization Datasets
by: Djuhera, Aladin, et al.
Published: (2025)
by: Djuhera, Aladin, et al.
Published: (2025)
DCA-Bench: A Benchmark for Dataset Curation Agents
by: Huang, Benhao, et al.
Published: (2024)
by: Huang, Benhao, et al.
Published: (2024)
Guided By AI: Navigating Trust, Bias, and Data Exploration in AI-Guided Visual Analytics
by: Ha, Sunwoo, et al.
Published: (2024)
by: Ha, Sunwoo, et al.
Published: (2024)
Guided By AI: Navigating Trust, Bias, and Data Exploration in AI‐Guided Visual Analytics
by: Sunwoo Ha, et al.
Published: (2024)
by: Sunwoo Ha, et al.
Published: (2024)
BERTCaps: BERT Capsule for Persian Multi-Domain Sentiment Analysis
by: Memari, Mohammadali, et al.
Published: (2024)
by: Memari, Mohammadali, et al.
Published: (2024)
BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
by: Yang, Chih-Hsuan, et al.
Published: (2024)
by: Yang, Chih-Hsuan, et al.
Published: (2024)
An Optimization Framework for the Time-Dependent Electric Vehicle Routing Problem with Shared Mobility: A Step Toward Smart Cities
by: Yazdiani, Alireza, et al.
Published: (2025)
by: Yazdiani, Alireza, et al.
Published: (2025)
Rhodium(II)‐Catalyzed Denitrogenative Transformations of N‐Sulfonyl‐1,2,3‐Triazoles
by: Fatemeh Doraghi, et al.
Published: (2024)
by: Fatemeh Doraghi, et al.
Published: (2024)
OPSD: an Offensive Persian Social media Dataset and its baseline evaluations
by: Safayani, Mehran, et al.
Published: (2024)
by: Safayani, Mehran, et al.
Published: (2024)
Combined Cycle Power Plants Water Purification Unit Life Cycle Assessment for Replacing Water Wastage
by: Alireza A. Majidy, et al.
Published: (2025)
by: Alireza A. Majidy, et al.
Published: (2025)
AquaCluster: Using Satellite Images And Self-supervised Machine Learning Networks To Detect Water Hidden Under Vegetation
by: Iakovidis, Ioannis, et al.
Published: (2025)
by: Iakovidis, Ioannis, et al.
Published: (2025)
Using a Human-AI Teaming Approach to Create and Curate Scientific Datasets with the SCILIRE System
by: Bölücü, Necva, et al.
Published: (2026)
by: Bölücü, Necva, et al.
Published: (2026)
Charging While Driving Lanes: A Boon to Electric Vehicle Owners or a Disruption to Traffic Flow
by: Bafandkar, Shayan, et al.
Published: (2025)
by: Bafandkar, Shayan, et al.
Published: (2025)
Publish‐Review‐Curate Modelling for Data Paper and Dataset: A Collaborative Approach
by: Youngim Jung, et al.
Published: (2025)
by: Youngim Jung, et al.
Published: (2025)
PAPPL: Personalized AI-Powered Progressive Learning Platform
by: Bafandkar, Shayan, et al.
Published: (2025)
by: Bafandkar, Shayan, et al.
Published: (2025)
Postmodern rendition of myth in Ashbery's ‘Syringa’
by: Roghayeh Farsi
Published: (2022)
by: Roghayeh Farsi
Published: (2022)
Reliable Curation of EHR Dataset via Large Language Models under Environmental Constraints
by: Xiong, Raymond M., et al.
Published: (2025)
by: Xiong, Raymond M., et al.
Published: (2025)
SynRXN: An Open Benchmark and Curated Dataset for Computational Reaction Modeling
by: Phan, Tieu-Long, et al.
Published: (2026)
by: Phan, Tieu-Long, et al.
Published: (2026)
MIMIC-Sepsis: A Curated Benchmark for Modeling and Learning from Sepsis Trajectories in the ICU
by: Huang, Yong, et al.
Published: (2025)
by: Huang, Yong, et al.
Published: (2025)
Similar Items
-
MELAC: Massive Evaluation of Large Language Models with Alignment of Culture in Persian Language
by: Farsi, Farhan, et al.
Published: (2025) -
RFBES at SemEval-2024 Task 8: Investigating Syntactic and Semantic Features for Distinguishing AI-Generated and Human-Written Texts
by: Rad, Mohammad Heydari, et al.
Published: (2024) -
PerHalluEval: Persian Hallucination Evaluation Benchmark for Large Language Models
by: Hosseini, Mohammad, et al.
Published: (2025) -
Who Gets the Mic? Investigating Gender Bias in the Speaker Assignment of a Speech-LLM
by: Puhach, Dariia, et al.
Published: (2025) -
An Empirical Study of Accuracy-Robustness Tradeoff and Training Efficiency in Self-Supervised Learning
by: Ghofrani, Fatemeh, et al.
Published: (2025)