When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
Fuente:
arXiv
Guardado en:
| Autores principales: | Basu, Abhinaba, Chakraborty, Pavan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
por: Basu, Abhinaba, et al.
Publicado: (2026)
por: Basu, Abhinaba, et al.
Publicado: (2026)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
por: Basu, Abhinaba
Publicado: (2026)
por: Basu, Abhinaba
Publicado: (2026)
ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs
por: Basu, Abhinaba, et al.
Publicado: (2026)
por: Basu, Abhinaba, et al.
Publicado: (2026)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
por: Lei, Xiang, et al.
Publicado: (2025)
por: Lei, Xiang, et al.
Publicado: (2025)
The AI Fiction Paradox
por: Elkins, Katherine
Publicado: (2026)
por: Elkins, Katherine
Publicado: (2026)
ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code
por: Madan, Kapil
Publicado: (2025)
por: Madan, Kapil
Publicado: (2025)
Named entity recognition for Serbian legal documents: Design, methodology and dataset development
por: Kalušev, Vladimir, et al.
Publicado: (2025)
por: Kalušev, Vladimir, et al.
Publicado: (2025)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
por: Cai, Weibin, et al.
Publicado: (2025)
por: Cai, Weibin, et al.
Publicado: (2025)
Understanding the Uncertainty of LLM Explanations: A Perspective Based on Reasoning Topology
por: Da, Longchao, et al.
Publicado: (2025)
por: Da, Longchao, et al.
Publicado: (2025)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
por: Hari, Vishnu, et al.
Publicado: (2025)
por: Hari, Vishnu, et al.
Publicado: (2025)
$ϕ^{\infty}$: Clause Purification, Embedding Realignment, and the Total Suppression of the Em Dash in Autoregressive Language Models
por: Kilictas, Bugra, et al.
Publicado: (2025)
por: Kilictas, Bugra, et al.
Publicado: (2025)
Thinking Longer, Not Always Smarter: Evaluating LLM Capabilities in Hierarchical Legal Reasoning
por: Zhang, Li, et al.
Publicado: (2025)
por: Zhang, Li, et al.
Publicado: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
por: Haque, Md. Asraful, et al.
Publicado: (2026)
por: Haque, Md. Asraful, et al.
Publicado: (2026)
Enhancing Ultra-Low-Bit Quantization of Large Language Models Through Saliency-Aware Partial Retraining
por: Cao, Deyu, et al.
Publicado: (2025)
por: Cao, Deyu, et al.
Publicado: (2025)
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
por: Zhang, Li, et al.
Publicado: (2026)
por: Zhang, Li, et al.
Publicado: (2026)
Rule Extraction in Machine Learning: Chat Incremental Pattern Constructor
por: Nwokocha, Caleb Princewill
Publicado: (2022)
por: Nwokocha, Caleb Princewill
Publicado: (2022)
Position: Uncertainty Quantification in LLMs is Just Unsupervised Clustering
por: Chen, Tiejin, et al.
Publicado: (2026)
por: Chen, Tiejin, et al.
Publicado: (2026)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
por: Kaiser, Daniel, et al.
Publicado: (2025)
por: Kaiser, Daniel, et al.
Publicado: (2025)
On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection
por: Schappacher-Tilp, Gudrun, et al.
Publicado: (2026)
por: Schappacher-Tilp, Gudrun, et al.
Publicado: (2026)
Formal Proofs as Structured Explanations: Proposing Several Tasks on Explainable Natural Language Inference
por: Abzianidze, Lasha
Publicado: (2023)
por: Abzianidze, Lasha
Publicado: (2023)
Conformal Path Reasoning: Trustworthy Knowledge Graph Question Answering via Path-Level Calibration
por: Lin, Shuhang, et al.
Publicado: (2026)
por: Lin, Shuhang, et al.
Publicado: (2026)
Sliced-Wasserstein Distribution Alignment Loss Improves the Ultra-Low-Bit Quantization of Large Language Models
por: Cao, Deyu, et al.
Publicado: (2026)
por: Cao, Deyu, et al.
Publicado: (2026)
Do LLMs Truly Understand When a Precedent Is Overruled?
por: Zhang, Li, et al.
Publicado: (2025)
por: Zhang, Li, et al.
Publicado: (2025)
Alpay Algebra V: Multi-Layered Semantic Games and Transfinite Fixed-Point Simulation
por: Kilictas, Bugra, et al.
Publicado: (2025)
por: Kilictas, Bugra, et al.
Publicado: (2025)
Automated Feedback Generation for Undergraduate Mathematics: Development and Evaluation of an AI Teaching Assistant
por: Gohr, Aron, et al.
Publicado: (2026)
por: Gohr, Aron, et al.
Publicado: (2026)
Alpay Algebra IV: Symbiotic Semantics and the Fixed-Point Convergence of Observer Embeddings
por: Kilictas, Bugra, et al.
Publicado: (2025)
por: Kilictas, Bugra, et al.
Publicado: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
por: Pather, Kaviraj, et al.
Publicado: (2025)
por: Pather, Kaviraj, et al.
Publicado: (2025)
How much do LLMs learn from negative examples?
por: Hamdan, Shadi, et al.
Publicado: (2025)
por: Hamdan, Shadi, et al.
Publicado: (2025)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
por: Breneur, Oleksandr Marchenko, et al.
Publicado: (2026)
por: Breneur, Oleksandr Marchenko, et al.
Publicado: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
por: Lasbordes, Maxence, et al.
Publicado: (2026)
por: Lasbordes, Maxence, et al.
Publicado: (2026)
HEFT: A Coarse-to-Fine Hierarchy for Enhancing the Efficiency and Accuracy of Language Model Reasoning
por: Hill, Brennen
Publicado: (2025)
por: Hill, Brennen
Publicado: (2025)
Hopscotch: Discovering and Skipping Redundancies in Language Models
por: Eyceoz, Mustafa, et al.
Publicado: (2025)
por: Eyceoz, Mustafa, et al.
Publicado: (2025)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
por: Guan, Xin, et al.
Publicado: (2024)
por: Guan, Xin, et al.
Publicado: (2024)
Smotrom tvoja pa ander drogoj verden! Resurrecting Dead Pidgin with Generative Models: Russenorsk Case Study
por: Tikhonov, Alexey, et al.
Publicado: (2025)
por: Tikhonov, Alexey, et al.
Publicado: (2025)
It's 2025 -- Narrative Learning is the new baseline to beat for explainable machine learning
por: Baker, Gregory D.
Publicado: (2025)
por: Baker, Gregory D.
Publicado: (2025)
When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
por: Gautam, Sushant, et al.
Publicado: (2026)
por: Gautam, Sushant, et al.
Publicado: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
por: Yang, Yibo
Publicado: (2025)
por: Yang, Yibo
Publicado: (2025)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
por: Borisov, Vadim
Publicado: (2026)
por: Borisov, Vadim
Publicado: (2026)
Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
por: Alpay, Faruk, et al.
Publicado: (2025)
por: Alpay, Faruk, et al.
Publicado: (2025)
Ejemplares similares
-
Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
por: Basu, Abhinaba, et al.
Publicado: (2026) -
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
por: Basu, Abhinaba
Publicado: (2026) -
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
por: Basu, Abhinaba
Publicado: (2026) -
ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs
por: Basu, Abhinaba, et al.
Publicado: (2026) -
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
por: Lei, Xiang, et al.
Publicado: (2025)