AAVENUE: Detecting LLM Biases on NLU Tasks in AAVE via a Novel Benchmark
Fuente:
arXiv
Guardado en:
| Autores principales: | Gupta, Abhay, Meng, Philip, Yurtseven, Ece, O'Brien, Sean, Zhu, Kevin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
por: Gupta, Abhay, et al.
Publicado: (2025)
por: Gupta, Abhay, et al.
Publicado: (2025)
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
por: Yugeswardeenoo, Dharunish, et al.
Publicado: (2024)
por: Yugeswardeenoo, Dharunish, et al.
Publicado: (2024)
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models
por: Gupta, Abhay, et al.
Publicado: (2025)
por: Gupta, Abhay, et al.
Publicado: (2025)
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
por: Mirza, Imran, et al.
Publicado: (2025)
por: Mirza, Imran, et al.
Publicado: (2025)
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
por: Singh, Ishneet Sukhvinder, et al.
Publicado: (2024)
por: Singh, Ishneet Sukhvinder, et al.
Publicado: (2024)
Error Reflection Prompting: Can Large Language Models Successfully Understand Errors?
por: Li, Jason, et al.
Publicado: (2025)
por: Li, Jason, et al.
Publicado: (2025)
CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning
por: Rufail, Andrew, et al.
Publicado: (2025)
por: Rufail, Andrew, et al.
Publicado: (2025)
DiversityMedQA: Assessing Demographic Biases in Medical Diagnosis using Large Language Models
por: Rawat, Rajat, et al.
Publicado: (2024)
por: Rawat, Rajat, et al.
Publicado: (2024)
Semantic Self-Consistency: Enhancing Language Model Reasoning via Semantic Weighting
por: Knappe, Tim, et al.
Publicado: (2024)
por: Knappe, Tim, et al.
Publicado: (2024)
Improving LLM Abilities in Idiomatic Translation
por: Donthi, Sundesh, et al.
Publicado: (2024)
por: Donthi, Sundesh, et al.
Publicado: (2024)
Pause-Tuning for Long-Context Comprehension: A Lightweight Approach to LLM Attention Recalibration
por: Begin, James, et al.
Publicado: (2025)
por: Begin, James, et al.
Publicado: (2025)
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning
por: Baek, Shaun, et al.
Publicado: (2025)
por: Baek, Shaun, et al.
Publicado: (2025)
Parallel Multi-Circuit Quantum Feature Fusion in Hybrid Quantum-Classical Convolutional Neural Networks for Breast Tumor Classification
por: Yurtseven, Ece
Publicado: (2025)
por: Yurtseven, Ece
Publicado: (2025)
Sarc7: Evaluating Sarcasm Detection and Generation with Seven Types and Emotion-Informed Techniques
por: Xiong, Lang, et al.
Publicado: (2025)
por: Xiong, Lang, et al.
Publicado: (2025)
RexUniNLU: Recursive Method with Explicit Schema Instructor for Universal NLU
por: Liu, Chengyuan, et al.
Publicado: (2024)
por: Liu, Chengyuan, et al.
Publicado: (2024)
From Bias to Balance: Detecting Facial Expression Recognition Biases in Large Multimodal Foundation Models
por: Chhua, Kaylee, et al.
Publicado: (2024)
por: Chhua, Kaylee, et al.
Publicado: (2024)
Causal Language Control in Multilingual Transformers via Sparse Feature Steering
por: Chou, Cheng-Ting, et al.
Publicado: (2025)
por: Chou, Cheng-Ting, et al.
Publicado: (2025)
FAIRE: Assessing Racial and Gender Bias in AI-Driven Resume Evaluations
por: Wen, Athena, et al.
Publicado: (2025)
por: Wen, Athena, et al.
Publicado: (2025)
ArabicNLU 2024: The First Arabic Natural Language Understanding Shared Task
por: Khalilia, Mohammed, et al.
Publicado: (2024)
por: Khalilia, Mohammed, et al.
Publicado: (2024)
OpenAutoNLU: Open Source AutoML Library for NLU
por: Arshinov, Grigory, et al.
Publicado: (2026)
por: Arshinov, Grigory, et al.
Publicado: (2026)
IITR-CIOL@NLU of Devanagari Script Languages 2025: Multilingual Hate Speech Detection and Target Identification in Devanagari-Scripted Languages
por: Gupta, Siddhant, et al.
Publicado: (2024)
por: Gupta, Siddhant, et al.
Publicado: (2024)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
por: Ghosh, Rajarshi, et al.
Publicado: (2025)
Introducing MAPO: Momentum-Aided Gradient Descent Prompt Optimization
por: Cui, Anthony, et al.
Publicado: (2024)
por: Cui, Anthony, et al.
Publicado: (2024)
Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT
por: Do, Timothy, et al.
Publicado: (2025)
por: Do, Timothy, et al.
Publicado: (2025)
Direct Confidence Alignment: Aligning Verbalized Confidence with Internal Confidence In Large Language Models
por: Zhang, Glenn, et al.
Publicado: (2025)
por: Zhang, Glenn, et al.
Publicado: (2025)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
por: Csizmadia, Daniel, et al.
Publicado: (2025)
por: Csizmadia, Daniel, et al.
Publicado: (2025)
ERGO: Entropy-guided Resetting for Generation Optimization in Multi-turn Language Models
por: Khalid, Haziq Mohammad, et al.
Publicado: (2025)
por: Khalid, Haziq Mohammad, et al.
Publicado: (2025)
Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
por: Le, Duy, et al.
Publicado: (2025)
por: Le, Duy, et al.
Publicado: (2025)
Survey of NLU Benchmarks Diagnosing Linguistic Phenomena: Why not Standardize Diagnostics Benchmarks?
por: Jallad, Khloud AL, et al.
Publicado: (2025)
por: Jallad, Khloud AL, et al.
Publicado: (2025)
Chain-of-Thought Augmentation with Logit Contrast for Enhanced Reasoning in Language Models
por: Shim, Jay, et al.
Publicado: (2024)
por: Shim, Jay, et al.
Publicado: (2024)
TRUTH DECAY: Quantifying Multi-Turn Sycophancy in Language Models
por: Liu, Joshua, et al.
Publicado: (2025)
por: Liu, Joshua, et al.
Publicado: (2025)
Probing Audio-Generation Capabilities of Text-Based Language Models
por: Anbazhagan, Arjun Prasaath, et al.
Publicado: (2025)
por: Anbazhagan, Arjun Prasaath, et al.
Publicado: (2025)
MatheMagic: Generating Dynamic Mathematics Benchmarks Robust to Memorization
por: O'Brien, Dayyán, et al.
Publicado: (2025)
por: O'Brien, Dayyán, et al.
Publicado: (2025)
TookaBERT: A Step Forward for Persian NLU
por: SadraeiJavaheri, MohammadAli, et al.
Publicado: (2024)
por: SadraeiJavaheri, MohammadAli, et al.
Publicado: (2024)
The Impact of Language Adapters in Cross-Lingual Transfer for NLU
por: Kunz, Jenny, et al.
Publicado: (2024)
por: Kunz, Jenny, et al.
Publicado: (2024)
LLM-Augmented Symbolic NLU System for More Reliable Continuous Causal Statement Interpretation
por: Lian, Xin, et al.
Publicado: (2025)
por: Lian, Xin, et al.
Publicado: (2025)
BioMistral-NLU: Towards More Generalizable Medical Language Understanding through Instruction Tuning
por: Fu, Yujuan Velvin, et al.
Publicado: (2024)
por: Fu, Yujuan Velvin, et al.
Publicado: (2024)
From Directions to Cones: Exploring Multidimensional Representations of Propositional Facts in LLMs
por: Yu, Stanley, et al.
Publicado: (2025)
por: Yu, Stanley, et al.
Publicado: (2025)
Sparse-IFT: Sparse Iso-FLOP Transformations for Maximizing Training Efficiency
por: Thangarasa, Vithursan, et al.
Publicado: (2023)
por: Thangarasa, Vithursan, et al.
Publicado: (2023)
DeKeyNLU: Enhancing Natural Language to SQL Generation through Task Decomposition and Keyword Extraction
por: Chen, Jian, et al.
Publicado: (2025)
por: Chen, Jian, et al.
Publicado: (2025)
Ejemplares similares
-
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
por: Gupta, Abhay, et al.
Publicado: (2025) -
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks
por: Yugeswardeenoo, Dharunish, et al.
Publicado: (2024) -
EnDive: A Cross-Dialect Benchmark for Fairness and Performance in Large Language Models
por: Gupta, Abhay, et al.
Publicado: (2025) -
MALIBU Benchmark: Multi-Agent LLM Implicit Bias Uncovered
por: Mirza, Imran, et al.
Publicado: (2025) -
ChunkRAG: Novel LLM-Chunk Filtering Method for RAG Systems
por: Singh, Ishneet Sukhvinder, et al.
Publicado: (2024)