Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Runtao, Wan, Guangya, Gabriel, Saadia, Li, Sheng, Gates, Alexander J, Sap, Maarten, Hartvigsen, Thomas |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025)
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
von: He, Zhonghao, et al.
Veröffentlicht: (2025)
von: He, Zhonghao, et al.
Veröffentlicht: (2025)
ModelCitizens: Representing Community Voices in Online Safety
von: Suvarna, Ashima, et al.
Veröffentlicht: (2025)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2025)
COMPASS: Enhancing Agent Long-Horizon Reasoning with Evolving Context
von: Wan, Guangya, et al.
Veröffentlicht: (2025)
von: Wan, Guangya, et al.
Veröffentlicht: (2025)
Rejected Dialects: Biases Against African American Language in Reward Models
von: Mire, Joel, et al.
Veröffentlicht: (2025)
von: Mire, Joel, et al.
Veröffentlicht: (2025)
PolygloToxicityPrompts: Multilingual Evaluation of Neural Toxic Degeneration in Large Language Models
von: Jain, Devansh, et al.
Veröffentlicht: (2024)
von: Jain, Devansh, et al.
Veröffentlicht: (2024)
Large Language Models for Causal Discovery: Current Landscape and Future Directions
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
von: Su, Zhe, et al.
Veröffentlicht: (2024)
von: Su, Zhe, et al.
Veröffentlicht: (2024)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
von: Zhou, Xuhui, et al.
Veröffentlicht: (2024)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
von: Mun, Jimin, et al.
Veröffentlicht: (2026)
von: Mun, Jimin, et al.
Veröffentlicht: (2026)
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
von: Park, Eunkyu, et al.
Veröffentlicht: (2025)
von: Park, Eunkyu, et al.
Veröffentlicht: (2025)
Is the Pope Catholic? Yes, the Pope is Catholic. Generative Evaluation of Non-Literal Intent Resolution in LLMs
von: Yerukola, Akhila, et al.
Veröffentlicht: (2024)
von: Yerukola, Akhila, et al.
Veröffentlicht: (2024)
UTMath: Math Evaluation with Unit Test via Reasoning-to-Coding Thoughts
von: Yang, Bo, et al.
Veröffentlicht: (2024)
von: Yang, Bo, et al.
Veröffentlicht: (2024)
Math Neurosurgery: Isolating Language Models' Math Reasoning Abilities Using Only Forward Passes
von: Christ, Bryan R., et al.
Veröffentlicht: (2024)
von: Christ, Bryan R., et al.
Veröffentlicht: (2024)
Training Proactive and Personalized LLM Agents
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
von: Sun, Weiwei, et al.
Veröffentlicht: (2025)
Relying on the Unreliable: The Impact of Language Models' Reluctance to Express Uncertainty
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2024)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2024)
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
von: Suvarna, Ashima, et al.
Veröffentlicht: (2026)
PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2025)
von: Kumar, Priyanshu, et al.
Veröffentlicht: (2025)
Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations
von: Zheng, Mingqian, et al.
Veröffentlicht: (2026)
von: Zheng, Mingqian, et al.
Veröffentlicht: (2026)
Reinforcing Stereotypes of Anger: Emotion AI on African American Vernacular English
von: Dorn, Rebecca, et al.
Veröffentlicht: (2025)
von: Dorn, Rebecca, et al.
Veröffentlicht: (2025)
1-2-3 Check: Enhancing Contextual Privacy in LLM via Multi-Agent Reasoning
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
von: Li, Wenkai, et al.
Veröffentlicht: (2025)
Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
von: Park, Eunkyu, et al.
Veröffentlicht: (2025)
von: Park, Eunkyu, et al.
Veröffentlicht: (2025)
Large Language Model for Patent Concept Generation
von: Ren, Runtao, et al.
Veröffentlicht: (2024)
von: Ren, Runtao, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Generation Systems for Intellectual Property via Synthetic Multi-Angle Fine-tuning
von: Ren, Runtao, et al.
Veröffentlicht: (2025)
von: Ren, Runtao, et al.
Veröffentlicht: (2025)
Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use
von: Pang, Renning, et al.
Veröffentlicht: (2026)
von: Pang, Renning, et al.
Veröffentlicht: (2026)
DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English
von: Oh, Jio, et al.
Veröffentlicht: (2026)
von: Oh, Jio, et al.
Veröffentlicht: (2026)
Efficient Knowledge Editing via Minimal Precomputation
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
von: Gupta, Akshat, et al.
Veröffentlicht: (2025)
Reinforcement Learning vs. Distillation: Understanding Accuracy and Capability in LLM Reasoning
von: Kim, Minwu, et al.
Veröffentlicht: (2025)
von: Kim, Minwu, et al.
Veröffentlicht: (2025)
Exploring Big Five Personality and AI Capability Effects in LLM-Simulated Negotiation Dialogues
von: Cohen, Myke C., et al.
Veröffentlicht: (2025)
von: Cohen, Myke C., et al.
Veröffentlicht: (2025)
The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages
von: Onyame, Eric, et al.
Veröffentlicht: (2026)
von: Onyame, Eric, et al.
Veröffentlicht: (2026)
Rel-A.I.: An Interaction-Centered Approach To Measuring Human-LM Reliance
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2024)
von: Zhou, Kaitlyn, et al.
Veröffentlicht: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
von: Mireshghallah, Niloofar, et al.
Veröffentlicht: (2023)
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
von: Jiao, Rui, et al.
Veröffentlicht: (2025)
von: Jiao, Rui, et al.
Veröffentlicht: (2025)
Analysis of LLM as a grammatical feature tagger for African American English
von: Porwal, Rahul, et al.
Veröffentlicht: (2025)
von: Porwal, Rahul, et al.
Veröffentlicht: (2025)
Reasoning with Natural Language Explanations
von: Valentino, Marco, et al.
Veröffentlicht: (2024)
von: Valentino, Marco, et al.
Veröffentlicht: (2024)
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
von: Zheng, Mingqian, et al.
Veröffentlicht: (2025)
von: Zheng, Mingqian, et al.
Veröffentlicht: (2025)
MindMerger: Efficient Boosting LLM Reasoning in non-English Languages
von: Huang, Zixian, et al.
Veröffentlicht: (2024)
von: Huang, Zixian, et al.
Veröffentlicht: (2024)
Derailer-Rerailer: Adaptive Verification for Efficient and Reliable Language Model Reasoning
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
von: Wan, Guangya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Mind the Gesture: Evaluating AI Sensitivity to Culturally Offensive Non-Verbal Gestures
von: Yerukola, Akhila, et al.
Veröffentlicht: (2025) -
Reasoning Aware Self-Consistency: Leveraging Reasoning Paths for Efficient LLM Sampling
von: Wan, Guangya, et al.
Veröffentlicht: (2024) -
Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025) -
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
von: He, Zhonghao, et al.
Veröffentlicht: (2025) -
ModelCitizens: Representing Community Voices in Online Safety
von: Suvarna, Ashima, et al.
Veröffentlicht: (2025)