I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations
Fuente:
arXiv
Saved in:
| Main Authors: | Kharchenko, Julia, Roosta, Tanya, Chadha, Aman, Shah, Chirag |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions
by: Kharchenko, Julia, et al.
Published: (2024)
by: Kharchenko, Julia, et al.
Published: (2024)
ClaimVer: Explainable Claim-Level Verification and Evidence Attribution of Text Through Knowledge Graphs
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024)
I Think, Therefore I am: Benchmarking Awareness of Large Language Models Using AwareBench
by: Li, Yuan, et al.
Published: (2024)
by: Li, Yuan, et al.
Published: (2024)
AuditLLM: A Tool for Auditing Large Language Models Using Multiprobe Approach
by: Amirizaniani, Maryam, et al.
Published: (2024)
by: Amirizaniani, Maryam, et al.
Published: (2024)
iAgentBench: Benchmarking Sensemaking Capabilities of Information-Seeking Agents on High-Traffic Topics
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2026)
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2026)
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
by: Liu, Yuhang, et al.
Published: (2025)
by: Liu, Yuhang, et al.
Published: (2025)
AssoCiAm: A Benchmark for Evaluating Association Thinking while Circumventing Ambiguity
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Simulating Meaning, Nevermore! Introducing ICR: A Semiotic-Hermeneutic Metric for Evaluating Meaning in LLM Text Summaries
by: Perez, Natalie, et al.
Published: (2026)
by: Perez, Natalie, et al.
Published: (2026)
Synthetic Users, Real Differences: an Evaluation Framework for User Simulation in Multi-Turn Conversations
by: Liu, Yu Lu, et al.
Published: (2026)
by: Liu, Yu Lu, et al.
Published: (2026)
I Am Aligned, But With Whom? MENA Values Benchmark for Evaluating Cultural Alignment and Multilingual Bias in LLMs
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
by: Zahraei, Pardis Sadat, et al.
Published: (2025)
I Think Therefore I Am: Building Ethical AI Through Relational Training
by: Willoughby, Samuel James
Published: (2026)
by: Willoughby, Samuel James
Published: (2026)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
by: Rani, Anku, et al.
Published: (2023)
by: Rani, Anku, et al.
Published: (2023)
Hire a Linguist!: Learning Endangered Languages with In-Context Linguistic Descriptions
by: Zhang, Kexun, et al.
Published: (2024)
by: Zhang, Kexun, et al.
Published: (2024)
Rudder: Steering Prefetching in Distributed GNN Training using LLM Agents
by: Sarkar, Aishwarya, et al.
Published: (2026)
by: Sarkar, Aishwarya, et al.
Published: (2026)
We Think, Therefore We Align LLMs to Helpful, Harmless and Honest Before They Go Wrong
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
by: Kashyap, Gautam Siddharth, et al.
Published: (2025)
The Order Effect: Investigating Prompt Sensitivity to Input Order in LLMs
by: Guan, Bryan, et al.
Published: (2025)
by: Guan, Bryan, et al.
Published: (2025)
Therefore I am. I Think
by: Esakkiraja, Esakkivel, et al.
Published: (2026)
by: Esakkiraja, Esakkivel, et al.
Published: (2026)
LLMAuditor: A Framework for Auditing Large Language Models Using Human-in-the-Loop
by: Amirizaniani, Maryam, et al.
Published: (2024)
by: Amirizaniani, Maryam, et al.
Published: (2024)
GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning
by: Yu, Qingchen, et al.
Published: (2025)
by: Yu, Qingchen, et al.
Published: (2025)
Breaking Language Barriers: A Question Answering Dataset for Hindi and Marathi
by: Sabane, Maithili, et al.
Published: (2023)
by: Sabane, Maithili, et al.
Published: (2023)
Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
by: Chopra, Harshita, et al.
Published: (2026)
by: Chopra, Harshita, et al.
Published: (2026)
Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
by: Anzenberg, Eitan, et al.
Published: (2025)
by: Anzenberg, Eitan, et al.
Published: (2025)
RealWebAssist: A Benchmark for Long-Horizon Web Assistance with Real-World Users
by: Ye, Suyu, et al.
Published: (2025)
by: Ye, Suyu, et al.
Published: (2025)
Human-Readable Adversarial Prompts: An Investigation into LLM Vulnerabilities Using Situational Context
by: Das, Nilanjana, et al.
Published: (2024)
by: Das, Nilanjana, et al.
Published: (2024)
Am I More Pointwise or Pairwise? Revealing Position Bias in Rubric-Based LLM-as-a-Judge
by: Xu, Yuzheng, et al.
Published: (2026)
by: Xu, Yuzheng, et al.
Published: (2026)
Unboxing Occupational Bias: Grounded Debiasing of LLMs with U.S. Labor Data
by: Gorti, Atmika, et al.
Published: (2024)
by: Gorti, Atmika, et al.
Published: (2024)
Generative Data Augmentation using LLMs improves Distributional Robustness in Question Answering
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
by: Chowdhury, Arijit Ghosh, et al.
Published: (2023)
Personality Shapes Gender Bias in Persona-Conditioned LLM Narratives Across English and Hindi: An Empirical Investigation
by: Kumar, Tanay, et al.
Published: (2026)
by: Kumar, Tanay, et al.
Published: (2026)
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
by: Haider, Batool, et al.
Published: (2025)
by: Haider, Batool, et al.
Published: (2025)
Hybrid Human-LLM Corpus Construction and LLM Evaluation for Rare Linguistic Phenomena
by: Weissweiler, Leonie, et al.
Published: (2024)
by: Weissweiler, Leonie, et al.
Published: (2024)
What Am I Missing? Question-Answering as Hidden State Probing
by: Luo, Chu Fei, et al.
Published: (2026)
by: Luo, Chu Fei, et al.
Published: (2026)
ChaI-TeA: A Benchmark for Evaluating Autocompletion of Interactions with LLM-based Chatbots
by: Goren, Shani, et al.
Published: (2024)
by: Goren, Shani, et al.
Published: (2024)
Gender and Positional Biases in LLM-Based Hiring Decisions: Evidence from Comparative CV/Résumé Evaluations
by: Rozado, David
Published: (2025)
by: Rozado, David
Published: (2025)
Bias in the Tails: How Name-conditioned Evaluative Framing in Resume Summaries Destabilizes LLM-based Hiring
by: Nghiem, Huy, et al.
Published: (2026)
by: Nghiem, Huy, et al.
Published: (2026)
Benchmark^2: Systematic Evaluation of LLM Benchmarks
by: Qian, Qi, et al.
Published: (2026)
by: Qian, Qi, et al.
Published: (2026)
Assessing LLM Reliability on Temporally Recent Open-Domain Questions
by: Krishnappa, Pushwitha, et al.
Published: (2026)
by: Krishnappa, Pushwitha, et al.
Published: (2026)
ClaimDB: A Fact Verification Benchmark over Large Structured Data
by: Theologitis, Michael, et al.
Published: (2026)
by: Theologitis, Michael, et al.
Published: (2026)
"You Are Rejected!": An Empirical Study of Large Language Models Taking Hiring Evaluations
by: Fu, Dingjie, et al.
Published: (2025)
by: Fu, Dingjie, et al.
Published: (2025)
Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues
by: Duan, Zhangqi, et al.
Published: (2026)
by: Duan, Zhangqi, et al.
Published: (2026)
Similar Items
-
How Well Do LLMs Represent Values Across Cultures? Empirical Analysis of LLM Responses Based on Hofstede Cultural Dimensions
by: Kharchenko, Julia, et al.
Published: (2024) -
ClaimVer: Explainable Claim-Level Verification and Evidence Attribution of Text Through Knowledge Graphs
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2024) -
I Think, Therefore I am: Benchmarking Awareness of Large Language Models Using AwareBench
by: Li, Yuan, et al.
Published: (2024) -
AuditLLM: A Tool for Auditing Large Language Models Using Multiprobe Approach
by: Amirizaniani, Maryam, et al.
Published: (2024) -
iAgentBench: Benchmarking Sensemaking Capabilities of Information-Seeking Agents on High-Traffic Topics
by: Dammu, Preetam Prabhu Srikar, et al.
Published: (2026)