Evaluating the Promise and Pitfalls of LLMs in Hiring Decisions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Anzenberg, Eitan, Samajpati, Arunava, Chandrasekar, Sivasankaran, Kacholia, Varun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
von: Gekker, Gil, et al.
Veröffentlicht: (2025)
von: Gekker, Gil, et al.
Veröffentlicht: (2025)
RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
von: de Wynter, Adrian, et al.
Veröffentlicht: (2024)
von: de Wynter, Adrian, et al.
Veröffentlicht: (2024)
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs
von: Chen, Ruxiao, et al.
Veröffentlicht: (2025)
von: Chen, Ruxiao, et al.
Veröffentlicht: (2025)
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025)
ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
von: Phutane, Mahika, et al.
Veröffentlicht: (2025)
von: Phutane, Mahika, et al.
Veröffentlicht: (2025)
Promises and Pitfalls of Generative Masked Language Modeling: Theoretical Framework and Practical Guidelines
von: Li, Yuchen, et al.
Veröffentlicht: (2024)
von: Li, Yuchen, et al.
Veröffentlicht: (2024)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
von: Chen, Yen-Shan, et al.
Veröffentlicht: (2026)
Towards Robust Evaluation of Unlearning in LLMs via Data Transformations
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
The Compliance Gap: Why AI Systems Promise to Follow Process Instructions but Don't
von: Shin, Kwan Soo
Veröffentlicht: (2026)
von: Shin, Kwan Soo
Veröffentlicht: (2026)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
Template-Based Probes Are Imperfect Lenses for Counterfactual Bias Evaluation in LLMs
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2024)
von: Kohankhaki, Farnaz, et al.
Veröffentlicht: (2024)
GRILE: A Benchmark for Grammar Reasoning and Explanation in Romanian LLMs
von: Dumitran, Adrian-Marius, et al.
Veröffentlicht: (2025)
von: Dumitran, Adrian-Marius, et al.
Veröffentlicht: (2025)
LatentQA: Teaching LLMs to Decode Activations Into Natural Language
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
von: Pan, Alexander, et al.
Veröffentlicht: (2024)
Exploring Knowledge Tracing in Tutor-Student Dialogues using LLMs
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2024)
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
von: Patil, Avinash, et al.
Veröffentlicht: (2025)
von: Patil, Avinash, et al.
Veröffentlicht: (2025)
The ProLiFIC dataset: Leveraging LLMs to Unveil the Italian Lawmaking Process
von: Contestabile, Matilde, et al.
Veröffentlicht: (2025)
von: Contestabile, Matilde, et al.
Veröffentlicht: (2025)
Exploring Consciousness in LLMs: A Systematic Survey of Theories, Implementations, and Frontier Risks
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
von: Chen, Sirui, et al.
Veröffentlicht: (2025)
Exploring the Potential of the Large Language Models (LLMs) in Identifying Misleading News Headlines
von: Rony, Md Main Uddin, et al.
Veröffentlicht: (2024)
von: Rony, Md Main Uddin, et al.
Veröffentlicht: (2024)
FoundationalASSIST: An Educational Dataset for Foundational Knowledge Tracing and Pedagogical Grounding of LLMs
von: Worden, Eamon, et al.
Veröffentlicht: (2026)
von: Worden, Eamon, et al.
Veröffentlicht: (2026)
Audit Me If You Can: Query-Efficient Active Fairness Auditing of Black-Box LLMs
von: Hartmann, David, et al.
Veröffentlicht: (2026)
von: Hartmann, David, et al.
Veröffentlicht: (2026)
Group Fairness Meets the Black Box: Enabling Fair Algorithms on Closed LLMs via Post-Processing
von: Xian, Ruicheng, et al.
Veröffentlicht: (2025)
von: Xian, Ruicheng, et al.
Veröffentlicht: (2025)
GLOCON Database: Design Decisions and User Manual (v1.0)
von: Hürriyetoğlu, Ali, et al.
Veröffentlicht: (2024)
von: Hürriyetoğlu, Ali, et al.
Veröffentlicht: (2024)
Toward Cultural Interpretability: A Linguistic Anthropological Framework for Describing and Evaluating Large Language Models (LLMs)
von: Jones, Graham M., et al.
Veröffentlicht: (2024)
von: Jones, Graham M., et al.
Veröffentlicht: (2024)
Towards Explainable Evaluation Metrics for Machine Translation
von: Leiter, Christoph, et al.
Veröffentlicht: (2023)
von: Leiter, Christoph, et al.
Veröffentlicht: (2023)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
PRSM: A Measure to Evaluate CLIP's Robustness Against Paraphrases
von: Schlegel, Udo, et al.
Veröffentlicht: (2025)
von: Schlegel, Udo, et al.
Veröffentlicht: (2025)
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams
von: Caraeni, Adriana, et al.
Veröffentlicht: (2024)
von: Caraeni, Adriana, et al.
Veröffentlicht: (2024)
Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
von: Friedrich, Felix, et al.
Veröffentlicht: (2025)
von: Friedrich, Felix, et al.
Veröffentlicht: (2025)
ELMES: An Automated Framework for Evaluating Large Language Models in Educational Scenarios
von: Wei, Shou'ang, et al.
Veröffentlicht: (2025)
von: Wei, Shou'ang, et al.
Veröffentlicht: (2025)
RDBE: Reasoning Distillation-Based Evaluation Enhances Automatic Essay Scoring
von: Mohammadkhani, Ali Ghiasvand
Veröffentlicht: (2024)
von: Mohammadkhani, Ali Ghiasvand
Veröffentlicht: (2024)
Generalization in Healthcare AI: Evaluation of a Clinical Large Language Model
von: Rahman, Salman, et al.
Veröffentlicht: (2024)
von: Rahman, Salman, et al.
Veröffentlicht: (2024)
Evaluation of LLMs for Process Model Analysis and Optimization
von: Kumar, Akhil, et al.
Veröffentlicht: (2025)
von: Kumar, Akhil, et al.
Veröffentlicht: (2025)
FairPair: A Robust Evaluation of Biases in Language Models through Paired Perturbations
von: Dwivedi-Yu, Jane, et al.
Veröffentlicht: (2024)
von: Dwivedi-Yu, Jane, et al.
Veröffentlicht: (2024)
Language Agents as Digital Representatives in Collective Decision-Making
von: Jarrett, Daniel, et al.
Veröffentlicht: (2025)
von: Jarrett, Daniel, et al.
Veröffentlicht: (2025)
Accept or Deny? Evaluating LLM Fairness and Performance in Loan Approval across Table-to-Text Serialization Approaches
von: Azime, Israel Abebe, et al.
Veröffentlicht: (2025)
von: Azime, Israel Abebe, et al.
Veröffentlicht: (2025)
Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
von: Zhao, Yibo, et al.
Veröffentlicht: (2025)
von: Zhao, Yibo, et al.
Veröffentlicht: (2025)
On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
von: Nguyen, Dang, et al.
Veröffentlicht: (2025)
Gender and Positional Biases in LLM-Based Hiring Decisions: Evidence from Comparative CV/Résumé Evaluations
von: Rozado, David
Veröffentlicht: (2025)
von: Rozado, David
Veröffentlicht: (2025)
Wikipedia in the Era of LLMs: Evolution and Risks
von: Huang, Siming, et al.
Veröffentlicht: (2025)
von: Huang, Siming, et al.
Veröffentlicht: (2025)
Can LLMs $\textit{understand}$ Math? -- Exploring the Pitfalls in Mathematical Reasoning
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
von: Roy, Tiasa Singha, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
What Makes an Evaluation Useful? Common Pitfalls and Best Practices
von: Gekker, Gil, et al.
Veröffentlicht: (2025) -
RTP-LX: Can LLMs Evaluate Toxicity in Multilingual Scenarios?
von: de Wynter, Adrian, et al.
Veröffentlicht: (2024) -
From Perceptions to Decisions: Wildfire Evacuation Decision Prediction with Behavioral Theory-informed LLMs
von: Chen, Ruxiao, et al.
Veröffentlicht: (2025) -
Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study
von: Majumdar, Ayan, et al.
Veröffentlicht: (2025) -
ABLEIST: Intersectional Disability Bias in LLM-Generated Hiring Scenarios
von: Phutane, Mahika, et al.
Veröffentlicht: (2025)