LocalBench: Benchmarking LLMs on County-Level Local Knowledge and Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Gao, Zihan, Xu, Yifei, Thebault-Spieker, Jacob |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Collective Narrative Grounding: Community-Coordinated Data Contributions to Improve Local AI Systems
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
A Turing Test for ''Localness'': Conceptualizing, Defining, and Recognizing Localness in People and Machines
by: Gao, Zihan, et al.
Published: (2025)
by: Gao, Zihan, et al.
Published: (2025)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026)
by: Kabir, Mohsinul, et al.
Published: (2026)
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
by: Wang, Tianyu, et al.
Published: (2026)
by: Wang, Tianyu, et al.
Published: (2026)
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
by: Meadows, Gwenyth Isobel, et al.
Published: (2024)
by: Meadows, Gwenyth Isobel, et al.
Published: (2024)
PLawBench: A Rubric-Based Benchmark for Evaluating LLMs in Real-World Legal Practice
by: Shi, Yuzhen, et al.
Published: (2026)
by: Shi, Yuzhen, et al.
Published: (2026)
Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition
by: Kim, Kyuhee, et al.
Published: (2025)
by: Kim, Kyuhee, et al.
Published: (2025)
ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
by: Li, Yuchong, et al.
Published: (2025)
by: Li, Yuchong, et al.
Published: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
by: Shin, Jisu, et al.
Published: (2025)
by: Shin, Jisu, et al.
Published: (2025)
HugAgent: Benchmarking LLMs for Simulation of Individualized Human Reasoning
by: Li, Chance Jiajie, et al.
Published: (2025)
by: Li, Chance Jiajie, et al.
Published: (2025)
DarkBench: Benchmarking Dark Patterns in Large Language Models
by: Kran, Esben, et al.
Published: (2025)
by: Kran, Esben, et al.
Published: (2025)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Knowledge-Augmented Reasoning for EUAIA Compliance and Adversarial Robustness of LLMs
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
by: Momcilovic, Tomas Bueno, et al.
Published: (2024)
PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations
by: Li, Jiatong, et al.
Published: (2024)
by: Li, Jiatong, et al.
Published: (2024)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
by: Ghosh, Rajarshi, et al.
Published: (2025)
by: Ghosh, Rajarshi, et al.
Published: (2025)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
by: Liu, Chaoqun, et al.
Published: (2025)
by: Liu, Chaoqun, et al.
Published: (2025)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
InterveneBench: Benchmarking LLMs for Intervention Reasoning and Causal Study Design in Real Social Systems
by: Shi, Shaojie, et al.
Published: (2026)
by: Shi, Shaojie, et al.
Published: (2026)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
Unveiling LLMs: The Evolution of Latent Representations in a Dynamic Knowledge Graph
by: Bronzini, Marco, et al.
Published: (2024)
by: Bronzini, Marco, et al.
Published: (2024)
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
by: Juvekar, Kush, et al.
Published: (2025)
by: Juvekar, Kush, et al.
Published: (2025)
MIRA: A Bilingual Benchmark for Medical Information Response Audit
by: Xu, Mengyu, et al.
Published: (2026)
by: Xu, Mengyu, et al.
Published: (2026)
Knowledge Acquisition on Mass-shooting Events via LLMs for AI-Driven Justice
by: Ihugba, Benign John, et al.
Published: (2025)
by: Ihugba, Benign John, et al.
Published: (2025)
HypoBench: Towards Systematic and Principled Benchmarking for Hypothesis Generation
by: Liu, Haokun, et al.
Published: (2025)
by: Liu, Haokun, et al.
Published: (2025)
ChatBench: From Static Benchmarks to Human-AI Evaluation
by: Chang, Serina, et al.
Published: (2025)
by: Chang, Serina, et al.
Published: (2025)
EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI
by: Kasu, Sai Kartheek Reddy
Published: (2025)
by: Kasu, Sai Kartheek Reddy
Published: (2025)
White Men Lead, Black Women Help? Benchmarking and Mitigating Language Agency Social Biases in LLMs
by: Wan, Yixin, et al.
Published: (2024)
by: Wan, Yixin, et al.
Published: (2024)
oMeBench: Towards Robust Benchmarking of LLMs in Organic Mechanism Elucidation and Reasoning
by: Xu, Ruiling, et al.
Published: (2025)
by: Xu, Ruiling, et al.
Published: (2025)
SimBench: Benchmarking the Ability of Large Language Models to Simulate Human Behaviors
by: Hu, Tiancheng, et al.
Published: (2025)
by: Hu, Tiancheng, et al.
Published: (2025)
ArabLegalEval: A Multitask Benchmark for Assessing Arabic Legal Knowledge in Large Language Models
by: Hijazi, Faris, et al.
Published: (2024)
by: Hijazi, Faris, et al.
Published: (2024)
OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution
by: La Cava, Lucio, et al.
Published: (2025)
by: La Cava, Lucio, et al.
Published: (2025)
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
by: Gao, Cheng, et al.
Published: (2025)
by: Gao, Cheng, et al.
Published: (2025)
QueerBench: Quantifying Discrimination in Language Models Toward Queer Identities
by: Sosto, Mae, et al.
Published: (2024)
by: Sosto, Mae, et al.
Published: (2024)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
by: Hamman, Faisal, et al.
Published: (2025)
by: Hamman, Faisal, et al.
Published: (2025)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy
by: Zhang, Mian, et al.
Published: (2024)
by: Zhang, Mian, et al.
Published: (2024)
Can Large Language Models Replace Human Coders? Introducing ContentBench
by: Haman, Michael
Published: (2026)
by: Haman, Michael
Published: (2026)
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models
by: Xu, Xin, et al.
Published: (2025)
by: Xu, Xin, et al.
Published: (2025)
Similar Items
-
Collective Narrative Grounding: Community-Coordinated Data Contributions to Improve Local AI Systems
by: Gao, Zihan, et al.
Published: (2025) -
A Turing Test for ''Localness'': Conceptualizing, Defining, and Recognizing Localness in People and Machines
by: Gao, Zihan, et al.
Published: (2025) -
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
by: Kabir, Mohsinul, et al.
Published: (2026) -
Can LLMs Think Like Consumers? Benchmarking Crowd-Level Reaction Reconstruction with ConsumerSimBench
by: Wang, Tianyu, et al.
Published: (2026) -
LocalValueBench: A Collaboratively Built and Extensible Benchmark for Evaluating Localized Value Alignment and Ethical Safety in Large Language Models
by: Meadows, Gwenyth Isobel, et al.
Published: (2024)